Document (#32320)

Author
Kaufmann, E.
Title
¬Das Indexieren von natürlichsprachlichen Dokumenten und die inverse Seitenhäufigkeit
Source
http://www.ifi.unizh.ch/cl/study/lizarbeiten/lizkaufmann.pdf
Year
2001
Abstract
Die Lizentiatsarbeit gibt im ersten theoretischen Teil einen Überblick über das Indexieren von Dokumenten. Sie zeigt die verschiedenen Typen von Indexen sowie die wichtigsten Aspekte bezüglich einer Indexsprache auf. Diverse manuelle und automatische Indexierungsverfahren werden präsentiert. Spezielle Aufmerksamkeit innerhalb des ersten Teils gilt den Schlagwortregistern, deren charakteristische Merkmale und Eigenheiten erörtert werden. Zusätzlich werden die gängigen Kriterien zur Bewertung von Indexen sowie die Masse zur Evaluation von Indexierungsverfahren und Indexierungsergebnissen vorgestellt. Im zweiten Teil der Arbeit werden fünf reale Bücher einer statistischen Untersuchung unterzogen. Zum einen werden die lexikalischen und syntaktischen Bestandteile der fünf Buchregister ermittelt, um den Inhalt von Schlagwortregistern zu erschliessen. Andererseits werden aus den Textausschnitten der Bücher Indexterme maschinell extrahiert und mit den Schlagworteinträgen in den Buchregistern verglichen. Das Hauptziel der Untersuchungen besteht darin, eine Indexierungsmethode, die auf linguistikorientierter Extraktion der Indexterme und Termhäufigkeitsgewichtung basiert, im Hinblick auf ihren Gebrauchswert für eine automatische Indexierung zu testen. Die Gewichtungsmethode ist die inverse Seitenhäufigkeit, eine Methode, welche von der inversen Dokumentfrequenz abgeleitet wurde, zur automatischen Erstellung von Schlagwortregistern für deutschsprachige Texte. Die Prüfung der Methode im statistischen Teil führte nicht zu zufriedenstellenden Resultaten.
Content
Lizentiatsarbeit der Philosphischen Fakultät der Universität Zürich, - Vgl. auch: http://www.ifi.unizh.ch/cl/study/lizarbeiten/lizkaufmann.pdf.
Theme
Automatisches Indexieren
Register

Similar documents (author)

  1. Kaufmann, N.C.: Kommt das Domainsterben? : Rechtsprechung gibt beschreibende Internet-Adressen zum Abschuss frei (2000) 5.99
    5.9875464 = sum of:
      5.9875464 = weight(author_txt:kaufmann in 5279) [ClassicSimilarity], result of:
        5.9875464 = fieldWeight in 5279, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          9.580074 = idf(docFreq=7, maxDocs=42596)
          0.625 = fieldNorm(doc=5279)
    
  2. Kaufmann, T.: Googeln wie die Profis : Perfekte Suche (2004) 5.99
    5.9875464 = sum of:
      5.9875464 = weight(author_txt:kaufmann in 2927) [ClassicSimilarity], result of:
        5.9875464 = fieldWeight in 2927, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          9.580074 = idf(docFreq=7, maxDocs=42596)
          0.625 = fieldNorm(doc=2927)
    
  3. Havemann, F.; Kaufmann, A.: ¬Der Wandel des Benutzerverhaltens in Zeiten des Internet : Ergebnisse von Befragungen an 13 Bibliotheken (2006) 4.79
    4.790037 = sum of:
      4.790037 = weight(author_txt:kaufmann in 329) [ClassicSimilarity], result of:
        4.790037 = fieldWeight in 329, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          9.580074 = idf(docFreq=7, maxDocs=42596)
          0.5 = fieldNorm(doc=329)
    
  4. Kaufmann, J.-C.: Wenn ICH ein anderer ist (2010) 4.79
    4.790037 = sum of:
      4.790037 = weight(author_txt:kaufmann in 4818) [ClassicSimilarity], result of:
        4.790037 = fieldWeight in 4818, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          9.580074 = idf(docFreq=7, maxDocs=42596)
          0.5 = fieldNorm(doc=4818)
    
  5. Kaufmann, J.-C.: ¬Die Erfindung des Ich : eine Theorie der Identität (2005) 4.79
    4.790037 = sum of:
      4.790037 = weight(author_txt:kaufmann in 4819) [ClassicSimilarity], result of:
        4.790037 = fieldWeight in 4819, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          9.580074 = idf(docFreq=7, maxDocs=42596)
          0.5 = fieldNorm(doc=4819)
    

Similar documents (content)

  1. Halip, I.: Automatische Extrahierung von Schlagworten aus unstrukturierten Texten (2005) 0.19
    0.19072658 = sum of:
      0.19072658 = product of:
        0.5960206 = sum of:
          0.076190874 = weight(abstract_txt:manuelle in 1166) [ClassicSimilarity], result of:
            0.076190874 = score(doc=1166,freq=1.0), product of:
              0.1578469 = queryWeight, product of:
                1.0065181 = boost
                8.826303 = idf(docFreq=16, maxDocs=42596)
                0.01776788 = queryNorm
              0.48268843 = fieldWeight in 1166, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                8.826303 = idf(docFreq=16, maxDocs=42596)
                0.0546875 = fieldNorm(doc=1166)
          0.022296593 = weight(abstract_txt:sowie in 1166) [ClassicSimilarity], result of:
            0.022296593 = score(doc=1166,freq=1.0), product of:
              0.08766033 = queryWeight, product of:
                1.0607673 = boost
                4.6510105 = idf(docFreq=1105, maxDocs=42596)
                0.01776788 = queryNorm
              0.25435215 = fieldWeight in 1166, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.6510105 = idf(docFreq=1105, maxDocs=42596)
                0.0546875 = fieldNorm(doc=1166)
          0.020726396 = weight(abstract_txt:eine in 1166) [ClassicSimilarity], result of:
            0.020726396 = score(doc=1166,freq=2.0), product of:
              0.07586015 = queryWeight, product of:
                1.208568 = boost
                3.532702 = idf(docFreq=3383, maxDocs=42596)
                0.01776788 = queryNorm
              0.27321848 = fieldWeight in 1166, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.532702 = idf(docFreq=3383, maxDocs=42596)
                0.0546875 = fieldNorm(doc=1166)
          0.10352854 = weight(abstract_txt:dokumenten in 1166) [ClassicSimilarity], result of:
            0.10352854 = score(doc=1166,freq=3.0), product of:
              0.16916496 = queryWeight, product of:
                1.4735802 = boost
                6.4610186 = idf(docFreq=180, maxDocs=42596)
                0.01776788 = queryNorm
              0.61199754 = fieldWeight in 1166, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                6.4610186 = idf(docFreq=180, maxDocs=42596)
                0.0546875 = fieldNorm(doc=1166)
          0.07410811 = weight(abstract_txt:automatische in 1166) [ClassicSimilarity], result of:
            0.07410811 = score(doc=1166,freq=1.0), product of:
              0.1952336 = queryWeight, product of:
                1.5830545 = boost
                6.9410167 = idf(docFreq=111, maxDocs=42596)
                0.01776788 = queryNorm
              0.37958685 = fieldWeight in 1166, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.9410167 = idf(docFreq=111, maxDocs=42596)
                0.0546875 = fieldNorm(doc=1166)
          0.056498013 = weight(abstract_txt:teil in 1166) [ClassicSimilarity], result of:
            0.056498013 = score(doc=1166,freq=1.0), product of:
              0.18650764 = queryWeight, product of:
                1.8950144 = boost
                5.5392184 = idf(docFreq=454, maxDocs=42596)
                0.01776788 = queryNorm
              0.302926 = fieldWeight in 1166, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.5392184 = idf(docFreq=454, maxDocs=42596)
                0.0546875 = fieldNorm(doc=1166)
          0.17114303 = weight(abstract_txt:indexierungsverfahren in 1166) [ClassicSimilarity], result of:
            0.17114303 = score(doc=1166,freq=1.0), product of:
              0.34110144 = queryWeight, product of:
                2.0924754 = boost
                9.174609 = idf(docFreq=11, maxDocs=42596)
                0.01776788 = queryNorm
              0.50173646 = fieldWeight in 1166, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                9.174609 = idf(docFreq=11, maxDocs=42596)
                0.0546875 = fieldNorm(doc=1166)
          0.071529 = weight(abstract_txt:werden in 1166) [ClassicSimilarity], result of:
            0.071529 = score(doc=1166,freq=6.0), product of:
              0.15134062 = queryWeight, product of:
                2.4141095 = boost
                3.528279 = idf(docFreq=3398, maxDocs=42596)
                0.01776788 = queryNorm
              0.47263584 = fieldWeight in 1166, product of:
                2.4494898 = tf(freq=6.0), with freq of:
                  6.0 = termFreq=6.0
                3.528279 = idf(docFreq=3398, maxDocs=42596)
                0.0546875 = fieldNorm(doc=1166)
        0.32 = coord(8/25)
    
  2. Bredack, J.: Automatische Extraktion fachterminologischer Mehrwortbegriffe : ein Verfahrensvergleich (2016) 0.15
    0.15301307 = sum of:
      0.15301307 = product of:
        0.63755447 = sum of:
          0.095258646 = weight(abstract_txt:extrahiert in 4195) [ClassicSimilarity], result of:
            0.095258646 = score(doc=4195,freq=1.0), product of:
              0.1675878 = queryWeight, product of:
                1.03711 = boost
                9.094566 = idf(docFreq=12, maxDocs=42596)
                0.01776788 = queryNorm
              0.5684104 = fieldWeight in 4195, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                9.094566 = idf(docFreq=12, maxDocs=42596)
                0.0625 = fieldNorm(doc=4195)
          0.023687309 = weight(abstract_txt:eine in 4195) [ClassicSimilarity], result of:
            0.023687309 = score(doc=4195,freq=2.0), product of:
              0.07586015 = queryWeight, product of:
                1.208568 = boost
                3.532702 = idf(docFreq=3383, maxDocs=42596)
                0.01776788 = queryNorm
              0.3122497 = fieldWeight in 4195, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.532702 = idf(docFreq=3383, maxDocs=42596)
                0.0625 = fieldNorm(doc=4195)
          0.08469498 = weight(abstract_txt:automatische in 4195) [ClassicSimilarity], result of:
            0.08469498 = score(doc=4195,freq=1.0), product of:
              0.1952336 = queryWeight, product of:
                1.5830545 = boost
                6.9410167 = idf(docFreq=111, maxDocs=42596)
                0.01776788 = queryNorm
              0.43381354 = fieldWeight in 4195, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.9410167 = idf(docFreq=111, maxDocs=42596)
                0.0625 = fieldNorm(doc=4195)
          0.12947898 = weight(abstract_txt:statistischen in 4195) [ClassicSimilarity], result of:
            0.12947898 = score(doc=4195,freq=1.0), product of:
              0.259089 = queryWeight, product of:
                1.8236567 = boost
                7.995954 = idf(docFreq=38, maxDocs=42596)
                0.01776788 = queryNorm
              0.49974713 = fieldWeight in 4195, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.995954 = idf(docFreq=38, maxDocs=42596)
                0.0625 = fieldNorm(doc=4195)
          0.22268711 = weight(abstract_txt:indexterme in 4195) [ClassicSimilarity], result of:
            0.22268711 = score(doc=4195,freq=1.0), product of:
              0.37191713 = queryWeight, product of:
                2.1849508 = boost
                9.580074 = idf(docFreq=7, maxDocs=42596)
                0.01776788 = queryNorm
              0.59875464 = fieldWeight in 4195, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                9.580074 = idf(docFreq=7, maxDocs=42596)
                0.0625 = fieldNorm(doc=4195)
          0.08174743 = weight(abstract_txt:werden in 4195) [ClassicSimilarity], result of:
            0.08174743 = score(doc=4195,freq=6.0), product of:
              0.15134062 = queryWeight, product of:
                2.4141095 = boost
                3.528279 = idf(docFreq=3398, maxDocs=42596)
                0.01776788 = queryNorm
              0.54015523 = fieldWeight in 4195, product of:
                2.4494898 = tf(freq=6.0), with freq of:
                  6.0 = termFreq=6.0
                3.528279 = idf(docFreq=3398, maxDocs=42596)
                0.0625 = fieldNorm(doc=4195)
        0.24 = coord(6/25)
    
  3. Leonhardt, H.A.: Systematik "Ästhetische Kulturwissenschaft" an der Universitätsbibliothek Hildesheim : ein Innovationsbericht (2018) 0.13
    0.13200678 = sum of:
      0.13200678 = product of:
        0.6600339 = sum of:
          0.033498913 = weight(abstract_txt:eine in 89) [ClassicSimilarity], result of:
            0.033498913 = score(doc=89,freq=1.0), product of:
              0.07586015 = queryWeight, product of:
                1.208568 = boost
                3.532702 = idf(docFreq=3383, maxDocs=42596)
                0.01776788 = queryNorm
              0.44158775 = fieldWeight in 89, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.532702 = idf(docFreq=3383, maxDocs=42596)
                0.125 = fieldNorm(doc=89)
          0.15117857 = weight(abstract_txt:bücher in 89) [ClassicSimilarity], result of:
            0.15117857 = score(doc=89,freq=1.0), product of:
              0.18097681 = queryWeight, product of:
                1.5241582 = boost
                6.6827817 = idf(docFreq=144, maxDocs=42596)
                0.01776788 = queryNorm
              0.8353477 = fieldWeight in 89, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.6827817 = idf(docFreq=144, maxDocs=42596)
                0.125 = fieldNorm(doc=89)
          0.2518243 = weight(abstract_txt:indexieren in 89) [ClassicSimilarity], result of:
            0.2518243 = score(doc=89,freq=1.0), product of:
              0.2543087 = queryWeight, product of:
                1.8067547 = boost
                7.921846 = idf(docFreq=41, maxDocs=42596)
                0.01776788 = queryNorm
              0.99023074 = fieldWeight in 89, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.921846 = idf(docFreq=41, maxDocs=42596)
                0.125 = fieldNorm(doc=89)
          0.12913832 = weight(abstract_txt:teil in 89) [ClassicSimilarity], result of:
            0.12913832 = score(doc=89,freq=1.0), product of:
              0.18650764 = queryWeight, product of:
                1.8950144 = boost
                5.5392184 = idf(docFreq=454, maxDocs=42596)
                0.01776788 = queryNorm
              0.6924023 = fieldWeight in 89, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.5392184 = idf(docFreq=454, maxDocs=42596)
                0.125 = fieldNorm(doc=89)
          0.09439379 = weight(abstract_txt:werden in 89) [ClassicSimilarity], result of:
            0.09439379 = score(doc=89,freq=2.0), product of:
              0.15134062 = queryWeight, product of:
                2.4141095 = boost
                3.528279 = idf(docFreq=3398, maxDocs=42596)
                0.01776788 = queryNorm
              0.6237175 = fieldWeight in 89, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.528279 = idf(docFreq=3398, maxDocs=42596)
                0.125 = fieldNorm(doc=89)
        0.2 = coord(5/25)
    
  4. Larroche-Boutet, V.; Pöhl, K.: ¬Das Nominalsyntagna : über die Nutzbarmachung eines logico-semantischen Konzeptes für dokumentarische Fragestellungen (1993) 0.12
    0.11799129 = sum of:
      0.11799129 = product of:
        0.58995646 = sum of:
          0.095258646 = weight(abstract_txt:extrahiert in 6283) [ClassicSimilarity], result of:
            0.095258646 = score(doc=6283,freq=1.0), product of:
              0.1675878 = queryWeight, product of:
                1.03711 = boost
                9.094566 = idf(docFreq=12, maxDocs=42596)
                0.01776788 = queryNorm
              0.5684104 = fieldWeight in 6283, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                9.094566 = idf(docFreq=12, maxDocs=42596)
                0.0625 = fieldNorm(doc=6283)
          0.023687309 = weight(abstract_txt:eine in 6283) [ClassicSimilarity], result of:
            0.023687309 = score(doc=6283,freq=2.0), product of:
              0.07586015 = queryWeight, product of:
                1.208568 = boost
                3.532702 = idf(docFreq=3383, maxDocs=42596)
                0.01776788 = queryNorm
              0.3122497 = fieldWeight in 6283, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.532702 = idf(docFreq=3383, maxDocs=42596)
                0.0625 = fieldNorm(doc=6283)
          0.11977679 = weight(abstract_txt:automatische in 6283) [ClassicSimilarity], result of:
            0.11977679 = score(doc=6283,freq=2.0), product of:
              0.1952336 = queryWeight, product of:
                1.5830545 = boost
                6.9410167 = idf(docFreq=111, maxDocs=42596)
                0.01776788 = queryNorm
              0.613505 = fieldWeight in 6283, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                6.9410167 = idf(docFreq=111, maxDocs=42596)
                0.0625 = fieldNorm(doc=6283)
          0.27660888 = weight(abstract_txt:indexierungsverfahren in 6283) [ClassicSimilarity], result of:
            0.27660888 = score(doc=6283,freq=2.0), product of:
              0.34110144 = queryWeight, product of:
                2.0924754 = boost
                9.174609 = idf(docFreq=11, maxDocs=42596)
                0.01776788 = queryNorm
              0.8109285 = fieldWeight in 6283, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                9.174609 = idf(docFreq=11, maxDocs=42596)
                0.0625 = fieldNorm(doc=6283)
          0.07462485 = weight(abstract_txt:werden in 6283) [ClassicSimilarity], result of:
            0.07462485 = score(doc=6283,freq=5.0), product of:
              0.15134062 = queryWeight, product of:
                2.4141095 = boost
                3.528279 = idf(docFreq=3398, maxDocs=42596)
                0.01776788 = queryNorm
              0.493092 = fieldWeight in 6283, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                3.528279 = idf(docFreq=3398, maxDocs=42596)
                0.0625 = fieldNorm(doc=6283)
        0.2 = coord(5/25)
    
  5. Peters, G.; Gaese, V.: ¬Das DocCat-System in der Textdokumentation von G+J (2003) 0.11
    0.112144716 = sum of:
      0.112144716 = product of:
        0.40051684 = sum of:
          0.035530962 = weight(abstract_txt:eine in 2508) [ClassicSimilarity], result of:
            0.035530962 = score(doc=2508,freq=8.0), product of:
              0.07586015 = queryWeight, product of:
                1.208568 = boost
                3.532702 = idf(docFreq=3383, maxDocs=42596)
                0.01776788 = queryNorm
              0.46837455 = fieldWeight in 2508, product of:
                2.828427 = tf(freq=8.0), with freq of:
                  8.0 = termFreq=8.0
                3.532702 = idf(docFreq=3383, maxDocs=42596)
                0.046875 = fieldNorm(doc=2508)
          0.051233344 = weight(abstract_txt:dokumenten in 2508) [ClassicSimilarity], result of:
            0.051233344 = score(doc=2508,freq=1.0), product of:
              0.16916496 = queryWeight, product of:
                1.4735802 = boost
                6.4610186 = idf(docFreq=180, maxDocs=42596)
                0.01776788 = queryNorm
              0.30286026 = fieldWeight in 2508, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.4610186 = idf(docFreq=180, maxDocs=42596)
                0.046875 = fieldNorm(doc=2508)
          0.063521236 = weight(abstract_txt:automatische in 2508) [ClassicSimilarity], result of:
            0.063521236 = score(doc=2508,freq=1.0), product of:
              0.1952336 = queryWeight, product of:
                1.5830545 = boost
                6.9410167 = idf(docFreq=111, maxDocs=42596)
                0.01776788 = queryNorm
              0.32536015 = fieldWeight in 2508, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.9410167 = idf(docFreq=111, maxDocs=42596)
                0.046875 = fieldNorm(doc=2508)
          0.06401722 = weight(abstract_txt:methode in 2508) [ClassicSimilarity], result of:
            0.06401722 = score(doc=2508,freq=1.0), product of:
              0.19624858 = queryWeight, product of:
                1.5871642 = boost
                6.9590354 = idf(docFreq=109, maxDocs=42596)
                0.01776788 = queryNorm
              0.32620478 = fieldWeight in 2508, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.9590354 = idf(docFreq=109, maxDocs=42596)
                0.046875 = fieldNorm(doc=2508)
          0.094434105 = weight(abstract_txt:indexieren in 2508) [ClassicSimilarity], result of:
            0.094434105 = score(doc=2508,freq=1.0), product of:
              0.2543087 = queryWeight, product of:
                1.8067547 = boost
                7.921846 = idf(docFreq=41, maxDocs=42596)
                0.01776788 = queryNorm
              0.37133652 = fieldWeight in 2508, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.921846 = idf(docFreq=41, maxDocs=42596)
                0.046875 = fieldNorm(doc=2508)
          0.04842687 = weight(abstract_txt:teil in 2508) [ClassicSimilarity], result of:
            0.04842687 = score(doc=2508,freq=1.0), product of:
              0.18650764 = queryWeight, product of:
                1.8950144 = boost
                5.5392184 = idf(docFreq=454, maxDocs=42596)
                0.01776788 = queryNorm
              0.25965086 = fieldWeight in 2508, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.5392184 = idf(docFreq=454, maxDocs=42596)
                0.046875 = fieldNorm(doc=2508)
          0.043353118 = weight(abstract_txt:werden in 2508) [ClassicSimilarity], result of:
            0.043353118 = score(doc=2508,freq=3.0), product of:
              0.15134062 = queryWeight, product of:
                2.4141095 = boost
                3.528279 = idf(docFreq=3398, maxDocs=42596)
                0.01776788 = queryNorm
              0.28646055 = fieldWeight in 2508, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                3.528279 = idf(docFreq=3398, maxDocs=42596)
                0.046875 = fieldNorm(doc=2508)
        0.28 = coord(7/25)