Document (#32319)

Kaufmann, E.
¬Das Indexieren von natürlichsprachlichen Dokumenten und die inverse Seitenhäufigkeit
Die Lizentiatsarbeit gibt im ersten theoretischen Teil einen Überblick über das Indexieren von Dokumenten. Sie zeigt die verschiedenen Typen von Indexen sowie die wichtigsten Aspekte bezüglich einer Indexsprache auf. Diverse manuelle und automatische Indexierungsverfahren werden präsentiert. Spezielle Aufmerksamkeit innerhalb des ersten Teils gilt den Schlagwortregistern, deren charakteristische Merkmale und Eigenheiten erörtert werden. Zusätzlich werden die gängigen Kriterien zur Bewertung von Indexen sowie die Masse zur Evaluation von Indexierungsverfahren und Indexierungsergebnissen vorgestellt. Im zweiten Teil der Arbeit werden fünf reale Bücher einer statistischen Untersuchung unterzogen. Zum einen werden die lexikalischen und syntaktischen Bestandteile der fünf Buchregister ermittelt, um den Inhalt von Schlagwortregistern zu erschliessen. Andererseits werden aus den Textausschnitten der Bücher Indexterme maschinell extrahiert und mit den Schlagworteinträgen in den Buchregistern verglichen. Das Hauptziel der Untersuchungen besteht darin, eine Indexierungsmethode, die auf linguistikorientierter Extraktion der Indexterme und Termhäufigkeitsgewichtung basiert, im Hinblick auf ihren Gebrauchswert für eine automatische Indexierung zu testen. Die Gewichtungsmethode ist die inverse Seitenhäufigkeit, eine Methode, welche von der inversen Dokumentfrequenz abgeleitet wurde, zur automatischen Erstellung von Schlagwortregistern für deutschsprachige Texte. Die Prüfung der Methode im statistischen Teil führte nicht zu zufriedenstellenden Resultaten.
Lizentiatsarbeit der Philosphischen Fakultät der Universität Zürich, - Vgl. auch:
Automatisches Indexieren

Similar documents (author)

  1. Kaufmann, N.C.: Kommt das Domainsterben? : Rechtsprechung gibt beschreibende Internet-Adressen zum Abschuss frei (2000) 6.01
    6.010904 = sum of:
      6.010904 = weight(author_txt:kaufmann in 4278) [ClassicSimilarity], result of:
        6.010904 = fieldWeight in 4278, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          9.617446 = idf(docFreq=7, maxDocs=44218)
          0.625 = fieldNorm(doc=4278)
  2. Kaufmann, T.: Googeln wie die Profis : Perfekte Suche (2004) 6.01
    6.010904 = sum of:
      6.010904 = weight(author_txt:kaufmann in 1926) [ClassicSimilarity], result of:
        6.010904 = fieldWeight in 1926, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          9.617446 = idf(docFreq=7, maxDocs=44218)
          0.625 = fieldNorm(doc=1926)
  3. Havemann, F.; Kaufmann, A.: ¬Der Wandel des Benutzerverhaltens in Zeiten des Internet : Ergebnisse von Befragungen an 13 Bibliotheken (2006) 4.81
    4.808723 = sum of:
      4.808723 = weight(author_txt:kaufmann in 24) [ClassicSimilarity], result of:
        4.808723 = fieldWeight in 24, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          9.617446 = idf(docFreq=7, maxDocs=44218)
          0.5 = fieldNorm(doc=24)
  4. Kaufmann, J.-C.: Wenn ICH ein anderer ist (2010) 4.81
    4.808723 = sum of:
      4.808723 = weight(author_txt:kaufmann in 3638) [ClassicSimilarity], result of:
        4.808723 = fieldWeight in 3638, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          9.617446 = idf(docFreq=7, maxDocs=44218)
          0.5 = fieldNorm(doc=3638)
  5. Kaufmann, J.-C.: ¬Die Erfindung des Ich : eine Theorie der Identität (2005) 4.81
    4.808723 = sum of:
      4.808723 = weight(author_txt:kaufmann in 3639) [ClassicSimilarity], result of:
        4.808723 = fieldWeight in 3639, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          9.617446 = idf(docFreq=7, maxDocs=44218)
          0.5 = fieldNorm(doc=3639)

Similar documents (content)

  1. Lepsky, K.: Automatisches Indexieren (2023) 0.28
    0.2835994 = sum of:
      0.2835994 = product of:
        1.1816642 = sum of:
          0.024176331 = weight(abstract_txt:eine in 781) [ClassicSimilarity], result of:
            0.024176331 = score(doc=781,freq=1.0), product of:
              0.07390804 = queryWeight, product of:
                1.1886243 = boost
                3.4892128 = idf(docFreq=3668, maxDocs=44218)
                0.01782049 = queryNorm
              0.3271137 = fieldWeight in 781, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.4892128 = idf(docFreq=3668, maxDocs=44218)
                0.09375 = fieldNorm(doc=781)
          0.10155645 = weight(abstract_txt:dokumenten in 781) [ClassicSimilarity], result of:
            0.10155645 = score(doc=781,freq=1.0), product of:
              0.16808902 = queryWeight, product of:
                1.4636014 = boost
                6.444614 = idf(docFreq=190, maxDocs=44218)
                0.01782049 = queryNorm
              0.60418254 = fieldWeight in 781, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.444614 = idf(docFreq=190, maxDocs=44218)
                0.09375 = fieldNorm(doc=781)
          0.12560704 = weight(abstract_txt:automatische in 781) [ClassicSimilarity], result of:
            0.12560704 = score(doc=781,freq=1.0), product of:
              0.19367653 = queryWeight, product of:
                1.5710558 = boost
                6.9177637 = idf(docFreq=118, maxDocs=44218)
                0.01782049 = queryNorm
              0.6485404 = fieldWeight in 781, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.9177637 = idf(docFreq=118, maxDocs=44218)
                0.09375 = fieldNorm(doc=781)
          0.28196114 = weight(abstract_txt:indexierungsverfahren in 781) [ClassicSimilarity], result of:
            0.28196114 = score(doc=781,freq=1.0), product of:
              0.33204263 = queryWeight, product of:
                2.0570745 = boost
                9.05783 = idf(docFreq=13, maxDocs=44218)
                0.01782049 = queryNorm
              0.8491715 = fieldWeight in 781, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                9.05783 = idf(docFreq=13, maxDocs=44218)
                0.09375 = fieldNorm(doc=781)
          0.5633807 = weight(abstract_txt:indexterme in 781) [ClassicSimilarity], result of:
            0.5633807 = score(doc=781,freq=3.0), product of:
              0.36522618 = queryWeight, product of:
                2.1574168 = boost
                9.499662 = idf(docFreq=8, maxDocs=44218)
                0.01782049 = queryNorm
              1.542553 = fieldWeight in 781, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                9.499662 = idf(docFreq=8, maxDocs=44218)
                0.09375 = fieldNorm(doc=781)
          0.08498248 = weight(abstract_txt:werden in 781) [ClassicSimilarity], result of:
            0.08498248 = score(doc=781,freq=3.0), product of:
              0.1492636 = queryWeight, product of:
                2.38886 = boost
                3.5062556 = idf(docFreq=3606, maxDocs=44218)
                0.01782049 = queryNorm
              0.56934494 = fieldWeight in 781, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                3.5062556 = idf(docFreq=3606, maxDocs=44218)
                0.09375 = fieldNorm(doc=781)
        0.24 = coord(6/25)
  2. Halip, I.: Automatische Extrahierung von Schlagworten aus unstrukturierten Texten (2005) 0.19
    0.18690273 = sum of:
      0.18690273 = product of:
        0.58407104 = sum of:
          0.07558158 = weight(abstract_txt:manuelle in 861) [ClassicSimilarity], result of:
            0.07558158 = score(doc=861,freq=1.0), product of:
              0.15693644 = queryWeight, product of:
                8.806516 = idf(docFreq=17, maxDocs=44218)
                0.01782049 = queryNorm
              0.48160633 = fieldWeight in 861, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                8.806516 = idf(docFreq=17, maxDocs=44218)
                0.0546875 = fieldNorm(doc=861)
          0.021828575 = weight(abstract_txt:sowie in 861) [ClassicSimilarity], result of:
            0.021828575 = score(doc=861,freq=1.0), product of:
              0.08639197 = queryWeight, product of:
                1.0492761 = boost
                4.6202335 = idf(docFreq=1183, maxDocs=44218)
                0.01782049 = queryNorm
              0.25266904 = fieldWeight in 861, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.6202335 = idf(docFreq=1183, maxDocs=44218)
                0.0546875 = fieldNorm(doc=861)
          0.019944455 = weight(abstract_txt:eine in 861) [ClassicSimilarity], result of:
            0.019944455 = score(doc=861,freq=2.0), product of:
              0.07390804 = queryWeight, product of:
                1.1886243 = boost
                3.4892128 = idf(docFreq=3668, maxDocs=44218)
                0.01782049 = queryNorm
              0.26985502 = fieldWeight in 861, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.4892128 = idf(docFreq=3668, maxDocs=44218)
                0.0546875 = fieldNorm(doc=861)
          0.102608874 = weight(abstract_txt:dokumenten in 861) [ClassicSimilarity], result of:
            0.102608874 = score(doc=861,freq=3.0), product of:
              0.16808902 = queryWeight, product of:
                1.4636014 = boost
                6.444614 = idf(docFreq=190, maxDocs=44218)
                0.01782049 = queryNorm
              0.61044365 = fieldWeight in 861, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                6.444614 = idf(docFreq=190, maxDocs=44218)
                0.0546875 = fieldNorm(doc=861)
          0.073270775 = weight(abstract_txt:automatische in 861) [ClassicSimilarity], result of:
            0.073270775 = score(doc=861,freq=1.0), product of:
              0.19367653 = queryWeight, product of:
                1.5710558 = boost
                6.9177637 = idf(docFreq=118, maxDocs=44218)
                0.01782049 = queryNorm
              0.3783152 = fieldWeight in 861, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.9177637 = idf(docFreq=118, maxDocs=44218)
                0.0546875 = fieldNorm(doc=861)
          0.056252465 = weight(abstract_txt:teil in 861) [ClassicSimilarity], result of:
            0.056252465 = score(doc=861,freq=1.0), product of:
              0.18588653 = queryWeight, product of:
                1.8850492 = boost
                5.533572 = idf(docFreq=474, maxDocs=44218)
                0.01782049 = queryNorm
              0.30261722 = fieldWeight in 861, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.533572 = idf(docFreq=474, maxDocs=44218)
                0.0546875 = fieldNorm(doc=861)
          0.16447733 = weight(abstract_txt:indexierungsverfahren in 861) [ClassicSimilarity], result of:
            0.16447733 = score(doc=861,freq=1.0), product of:
              0.33204263 = queryWeight, product of:
                2.0570745 = boost
                9.05783 = idf(docFreq=13, maxDocs=44218)
                0.01782049 = queryNorm
              0.49535006 = fieldWeight in 861, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                9.05783 = idf(docFreq=13, maxDocs=44218)
                0.0546875 = fieldNorm(doc=861)
          0.07010697 = weight(abstract_txt:werden in 861) [ClassicSimilarity], result of:
            0.07010697 = score(doc=861,freq=6.0), product of:
              0.1492636 = queryWeight, product of:
                2.38886 = boost
                3.5062556 = idf(docFreq=3606, maxDocs=44218)
                0.01782049 = queryNorm
              0.4696856 = fieldWeight in 861, product of:
                2.4494898 = tf(freq=6.0), with freq of:
                  6.0 = termFreq=6.0
                3.5062556 = idf(docFreq=3606, maxDocs=44218)
                0.0546875 = fieldNorm(doc=861)
        0.32 = coord(8/25)
  3. Bredack, J.: Automatische Extraktion fachterminologischer Mehrwortbegriffe : ein Verfahrensvergleich (2016) 0.15
    0.15057199 = sum of:
      0.15057199 = product of:
        0.6273833 = sum of:
          0.093987055 = weight(abstract_txt:extrahiert in 3194) [ClassicSimilarity], result of:
            0.093987055 = score(doc=3194,freq=1.0), product of:
              0.16602132 = queryWeight, product of:
                1.0285373 = boost
                9.05783 = idf(docFreq=13, maxDocs=44218)
                0.01782049 = queryNorm
              0.56611437 = fieldWeight in 3194, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                9.05783 = idf(docFreq=13, maxDocs=44218)
                0.0625 = fieldNorm(doc=3194)
          0.022793664 = weight(abstract_txt:eine in 3194) [ClassicSimilarity], result of:
            0.022793664 = score(doc=3194,freq=2.0), product of:
              0.07390804 = queryWeight, product of:
                1.1886243 = boost
                3.4892128 = idf(docFreq=3668, maxDocs=44218)
                0.01782049 = queryNorm
              0.30840576 = fieldWeight in 3194, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.4892128 = idf(docFreq=3668, maxDocs=44218)
                0.0625 = fieldNorm(doc=3194)
          0.08373803 = weight(abstract_txt:automatische in 3194) [ClassicSimilarity], result of:
            0.08373803 = score(doc=3194,freq=1.0), product of:
              0.19367653 = queryWeight, product of:
                1.5710558 = boost
                6.9177637 = idf(docFreq=118, maxDocs=44218)
                0.01782049 = queryNorm
              0.43236023 = fieldWeight in 3194, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.9177637 = idf(docFreq=118, maxDocs=44218)
                0.0625 = fieldNorm(doc=3194)
          0.12989697 = weight(abstract_txt:statistischen in 3194) [ClassicSimilarity], result of:
            0.12989697 = score(doc=3194,freq=1.0), product of:
              0.25953415 = queryWeight, product of:
                1.8186553 = boost
                8.008008 = idf(docFreq=39, maxDocs=44218)
                0.01782049 = queryNorm
              0.5005005 = fieldWeight in 3194, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                8.008008 = idf(docFreq=39, maxDocs=44218)
                0.0625 = fieldNorm(doc=3194)
          0.21684533 = weight(abstract_txt:indexterme in 3194) [ClassicSimilarity], result of:
            0.21684533 = score(doc=3194,freq=1.0), product of:
              0.36522618 = queryWeight, product of:
                2.1574168 = boost
                9.499662 = idf(docFreq=8, maxDocs=44218)
                0.01782049 = queryNorm
              0.5937289 = fieldWeight in 3194, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                9.499662 = idf(docFreq=8, maxDocs=44218)
                0.0625 = fieldNorm(doc=3194)
          0.080122255 = weight(abstract_txt:werden in 3194) [ClassicSimilarity], result of:
            0.080122255 = score(doc=3194,freq=6.0), product of:
              0.1492636 = queryWeight, product of:
                2.38886 = boost
                3.5062556 = idf(docFreq=3606, maxDocs=44218)
                0.01782049 = queryNorm
              0.5367836 = fieldWeight in 3194, product of:
                2.4494898 = tf(freq=6.0), with freq of:
                  6.0 = termFreq=6.0
                3.5062556 = idf(docFreq=3606, maxDocs=44218)
                0.0625 = fieldNorm(doc=3194)
        0.24 = coord(6/25)
  4. Leonhardt, H.A.: Systematik "Ästhetische Kulturwissenschaft" an der Universitätsbibliothek Hildesheim : ein Innovationsbericht (2018) 0.13
    0.13210687 = sum of:
      0.13210687 = product of:
        0.6605343 = sum of:
          0.03223511 = weight(abstract_txt:eine in 4490) [ClassicSimilarity], result of:
            0.03223511 = score(doc=4490,freq=1.0), product of:
              0.07390804 = queryWeight, product of:
                1.1886243 = boost
                3.4892128 = idf(docFreq=3668, maxDocs=44218)
                0.01782049 = queryNorm
              0.4361516 = fieldWeight in 4490, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.4892128 = idf(docFreq=3668, maxDocs=44218)
                0.125 = fieldNorm(doc=4490)
          0.15213066 = weight(abstract_txt:bücher in 4490) [ClassicSimilarity], result of:
            0.15213066 = score(doc=4490,freq=1.0), product of:
              0.18165736 = queryWeight, product of:
                1.5215268 = boost
                6.699675 = idf(docFreq=147, maxDocs=44218)
                0.01782049 = queryNorm
              0.8374594 = fieldWeight in 4490, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.699675 = idf(docFreq=147, maxDocs=44218)
                0.125 = fieldNorm(doc=4490)
          0.2550743 = weight(abstract_txt:indexieren in 4490) [ClassicSimilarity], result of:
            0.2550743 = score(doc=4490,freq=1.0), product of:
              0.25638127 = queryWeight, product of:
                1.8075747 = boost
                7.9592175 = idf(docFreq=41, maxDocs=44218)
                0.01782049 = queryNorm
              0.9949022 = fieldWeight in 4490, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.9592175 = idf(docFreq=41, maxDocs=44218)
                0.125 = fieldNorm(doc=4490)
          0.12857707 = weight(abstract_txt:teil in 4490) [ClassicSimilarity], result of:
            0.12857707 = score(doc=4490,freq=1.0), product of:
              0.18588653 = queryWeight, product of:
                1.8850492 = boost
                5.533572 = idf(docFreq=474, maxDocs=44218)
                0.01782049 = queryNorm
              0.6916965 = fieldWeight in 4490, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.533572 = idf(docFreq=474, maxDocs=44218)
                0.125 = fieldNorm(doc=4490)
          0.09251721 = weight(abstract_txt:werden in 4490) [ClassicSimilarity], result of:
            0.09251721 = score(doc=4490,freq=2.0), product of:
              0.1492636 = queryWeight, product of:
                2.38886 = boost
                3.5062556 = idf(docFreq=3606, maxDocs=44218)
                0.01782049 = queryNorm
              0.6198243 = fieldWeight in 4490, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.5062556 = idf(docFreq=3606, maxDocs=44218)
                0.125 = fieldNorm(doc=4490)
        0.2 = coord(5/25)
  5. Larroche-Boutet, V.; Pöhl, K.: ¬Das Nominalsyntagna : über die Nutzbarmachung eines logico-semantischen Konzeptes für dokumentarische Fragestellungen (1993) 0.11
    0.11483621 = sum of:
      0.11483621 = product of:
        0.574181 = sum of:
          0.093987055 = weight(abstract_txt:extrahiert in 5282) [ClassicSimilarity], result of:
            0.093987055 = score(doc=5282,freq=1.0), product of:
              0.16602132 = queryWeight, product of:
                1.0285373 = boost
                9.05783 = idf(docFreq=13, maxDocs=44218)
                0.01782049 = queryNorm
              0.56611437 = fieldWeight in 5282, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                9.05783 = idf(docFreq=13, maxDocs=44218)
                0.0625 = fieldNorm(doc=5282)
          0.022793664 = weight(abstract_txt:eine in 5282) [ClassicSimilarity], result of:
            0.022793664 = score(doc=5282,freq=2.0), product of:
              0.07390804 = queryWeight, product of:
                1.1886243 = boost
                3.4892128 = idf(docFreq=3668, maxDocs=44218)
                0.01782049 = queryNorm
              0.30840576 = fieldWeight in 5282, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.4892128 = idf(docFreq=3668, maxDocs=44218)
                0.0625 = fieldNorm(doc=5282)
          0.11842346 = weight(abstract_txt:automatische in 5282) [ClassicSimilarity], result of:
            0.11842346 = score(doc=5282,freq=2.0), product of:
              0.19367653 = queryWeight, product of:
                1.5710558 = boost
                6.9177637 = idf(docFreq=118, maxDocs=44218)
                0.01782049 = queryNorm
              0.6114497 = fieldWeight in 5282, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                6.9177637 = idf(docFreq=118, maxDocs=44218)
                0.0625 = fieldNorm(doc=5282)
          0.26583552 = weight(abstract_txt:indexierungsverfahren in 5282) [ClassicSimilarity], result of:
            0.26583552 = score(doc=5282,freq=2.0), product of:
              0.33204263 = queryWeight, product of:
                2.0570745 = boost
                9.05783 = idf(docFreq=13, maxDocs=44218)
                0.01782049 = queryNorm
              0.8006066 = fieldWeight in 5282, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                9.05783 = idf(docFreq=13, maxDocs=44218)
                0.0625 = fieldNorm(doc=5282)
          0.07314128 = weight(abstract_txt:werden in 5282) [ClassicSimilarity], result of:
            0.07314128 = score(doc=5282,freq=5.0), product of:
              0.1492636 = queryWeight, product of:
                2.38886 = boost
                3.5062556 = idf(docFreq=3606, maxDocs=44218)
                0.01782049 = queryNorm
              0.49001414 = fieldWeight in 5282, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                3.5062556 = idf(docFreq=3606, maxDocs=44218)
                0.0625 = fieldNorm(doc=5282)
        0.2 = coord(5/25)