Document (#32949)

Author
Dunlavy, D.M.
O'Leary, D.P.
Conroy, J.M.
Schlesinger, J.D.
Title
QCS: A system for querying, clustering and summarizing documents
Source
Information processing and management. 43(2007) no.6, S.1588-1605
Year
2007
Abstract
Information retrieval systems consist of many complicated components. Research and development of such systems is often hampered by the difficulty in evaluating how each particular component would behave across multiple systems. We present a novel integrated information retrieval system-the Query, Cluster, Summarize (QCS) system-which is portable, modular, and permits experimentation with different instantiations of each of the constituent text analysis components. Most importantly, the combination of the three types of methods in the QCS design improves retrievals by providing users more focused information organized by topic. We demonstrate the improved performance by a series of experiments using standard test sets from the Document Understanding Conferences (DUC) as measured by the best known automatic metric for summarization system evaluation, ROUGE. Although the DUC data and evaluations were originally designed to test multidocument summarization, we developed a framework to extend it to the task of evaluation for each of the three components: query, clustering, and summarization. Under this framework, we then demonstrate that the QCS system (end-to-end) achieves performance as good as or better than the best summarization engines. Given a query, QCS retrieves relevant documents, separates the retrieved documents into topic clusters, and creates a single summary for each cluster. In the current implementation, Latent Semantic Indexing is used for retrieval, generalized spherical k-means is used for the document clustering, and a method coupling sentence "trimming" and a hidden Markov model, followed by a pivoted QR decomposition, is used to create a single extract summary for each cluster. The user interface is designed to provide access to detailed information in a compact and useful format. Our system demonstrates the feasibility of assembling an effective IR system from existing software libraries, the usefulness of the modularity of the design, and the value of this particular combination of modules.
Theme
Automatisches Abstracting

Similar documents (author)

  1. O'Leary, M.: FirstSearch takes the lead (1992) 5.44
    5.442653 = sum of:
      5.442653 = weight(author_txt:o'leary in 3121) [ClassicSimilarity], result of:
        5.442653 = fieldWeight in 3121, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          8.708245 = idf(docFreq=18, maxDocs=42306)
          0.625 = fieldNorm(doc=3121)
    
  2. O'Leary, M.: DIALOG Select : DIALOG for knowledge worker-on the Web (1997) 5.44
    5.442653 = sum of:
      5.442653 = weight(author_txt:o'leary in 6489) [ClassicSimilarity], result of:
        5.442653 = fieldWeight in 6489, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          8.708245 = idf(docFreq=18, maxDocs=42306)
          0.625 = fieldNorm(doc=6489)
    
  3. O'Leary, M.: DIALOG TARGET's new age searching (1993) 5.44
    5.442653 = sum of:
      5.442653 = weight(author_txt:o'leary in 7951) [ClassicSimilarity], result of:
        5.442653 = fieldWeight in 7951, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          8.708245 = idf(docFreq=18, maxDocs=42306)
          0.625 = fieldNorm(doc=7951)
    
  4. O'Leary, M.: Images on CompuServe : the best of 'Bettmann' online (1995) 5.44
    5.442653 = sum of:
      5.442653 = weight(author_txt:o'leary in 2016) [ClassicSimilarity], result of:
        5.442653 = fieldWeight in 2016, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          8.708245 = idf(docFreq=18, maxDocs=42306)
          0.625 = fieldNorm(doc=2016)
    
  5. O'Leary, D.E.: Using AI in knowledge management : knowledge bases and ontologies (1998) 5.44
    5.442653 = sum of:
      5.442653 = weight(author_txt:o'leary in 2644) [ClassicSimilarity], result of:
        5.442653 = fieldWeight in 2644, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          8.708245 = idf(docFreq=18, maxDocs=42306)
          0.625 = fieldNorm(doc=2644)
    

Similar documents (content)

  1. Ouyang, Y.; Li, W.; Li, S.; Lu, Q.: Intertopic information mining for query-based summarization (2010) 0.46
    0.45888567 = sum of:
      0.45888567 = product of:
        0.95601183 = sum of:
          0.02894772 = weight(abstract_txt:evaluation in 460) [ClassicSimilarity], result of:
            0.02894772 = score(doc=460,freq=1.0), product of:
              0.10307658 = queryWeight, product of:
                4.4933925 = idf(docFreq=1285, maxDocs=42306)
                0.022939589 = queryNorm
              0.28083703 = fieldWeight in 460, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.4933925 = idf(docFreq=1285, maxDocs=42306)
                0.0625 = fieldNorm(doc=460)
          0.031775363 = weight(abstract_txt:framework in 460) [ClassicSimilarity], result of:
            0.031775363 = score(doc=460,freq=1.0), product of:
              0.10968421 = queryWeight, product of:
                1.0315542 = boost
                4.635178 = idf(docFreq=1115, maxDocs=42306)
                0.022939589 = queryNorm
              0.28969863 = fieldWeight in 460, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.635178 = idf(docFreq=1115, maxDocs=42306)
                0.0625 = fieldNorm(doc=460)
          0.024402253 = weight(abstract_txt:information in 460) [ClassicSimilarity], result of:
            0.024402253 = score(doc=460,freq=7.0), product of:
              0.06058257 = queryWeight, product of:
                1.0841986 = boost
                2.435865 = idf(docFreq=10064, maxDocs=42306)
                0.022939589 = queryNorm
              0.4027933 = fieldWeight in 460, product of:
                2.6457512 = tf(freq=7.0), with freq of:
                  7.0 = termFreq=7.0
                2.435865 = idf(docFreq=10064, maxDocs=42306)
                0.0625 = fieldNorm(doc=460)
          0.04119478 = weight(abstract_txt:best in 460) [ClassicSimilarity], result of:
            0.04119478 = score(doc=460,freq=1.0), product of:
              0.1304103 = queryWeight, product of:
                1.1248016 = boost
                5.0541754 = idf(docFreq=733, maxDocs=42306)
                0.022939589 = queryNorm
              0.31588596 = fieldWeight in 460, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.0541754 = idf(docFreq=733, maxDocs=42306)
                0.0625 = fieldNorm(doc=460)
          0.018499946 = weight(abstract_txt:used in 460) [ClassicSimilarity], result of:
            0.018499946 = score(doc=460,freq=1.0), product of:
              0.087544285 = queryWeight, product of:
                1.1287026 = boost
                3.381136 = idf(docFreq=3910, maxDocs=42306)
                0.022939589 = queryNorm
              0.211321 = fieldWeight in 460, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.381136 = idf(docFreq=3910, maxDocs=42306)
                0.0625 = fieldNorm(doc=460)
          0.08487348 = weight(abstract_txt:topic in 460) [ClassicSimilarity], result of:
            0.08487348 = score(doc=460,freq=4.0), product of:
              0.13301842 = queryWeight, product of:
                1.1359936 = boost
                5.104465 = idf(docFreq=697, maxDocs=42306)
                0.022939589 = queryNorm
              0.6380581 = fieldWeight in 460, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                5.104465 = idf(docFreq=697, maxDocs=42306)
                0.0625 = fieldNorm(doc=460)
          0.019100329 = weight(abstract_txt:systems in 460) [ClassicSimilarity], result of:
            0.019100329 = score(doc=460,freq=1.0), product of:
              0.089428246 = queryWeight, product of:
                1.1407828 = boost
                3.4173236 = idf(docFreq=3771, maxDocs=42306)
                0.022939589 = queryNorm
              0.21358272 = fieldWeight in 460, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.4173236 = idf(docFreq=3771, maxDocs=42306)
                0.0625 = fieldNorm(doc=460)
          0.033368852 = weight(abstract_txt:documents in 460) [ClassicSimilarity], result of:
            0.033368852 = score(doc=460,freq=1.0), product of:
              0.12972042 = queryWeight, product of:
                1.3739464 = boost
                4.115787 = idf(docFreq=1875, maxDocs=42306)
                0.022939589 = queryNorm
              0.2572367 = fieldWeight in 460, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.115787 = idf(docFreq=1875, maxDocs=42306)
                0.0625 = fieldNorm(doc=460)
          0.050776668 = weight(abstract_txt:query in 460) [ClassicSimilarity], result of:
            0.050776668 = score(doc=460,freq=1.0), product of:
              0.17161568 = queryWeight, product of:
                1.5803167 = boost
                4.733989 = idf(docFreq=1010, maxDocs=42306)
                0.022939589 = queryNorm
              0.2958743 = fieldWeight in 460, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.733989 = idf(docFreq=1010, maxDocs=42306)
                0.0625 = fieldNorm(doc=460)
          0.11560397 = weight(abstract_txt:clustering in 460) [ClassicSimilarity], result of:
            0.11560397 = score(doc=460,freq=1.0), product of:
              0.29700425 = queryWeight, product of:
                2.078964 = boost
                6.227734 = idf(docFreq=226, maxDocs=42306)
                0.022939589 = queryNorm
              0.38923338 = fieldWeight in 460, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.227734 = idf(docFreq=226, maxDocs=42306)
                0.0625 = fieldNorm(doc=460)
          0.04254783 = weight(abstract_txt:system in 460) [ClassicSimilarity], result of:
            0.04254783 = score(doc=460,freq=1.0), product of:
              0.20231342 = queryWeight, product of:
                2.6209962 = boost
                3.3649042 = idf(docFreq=3974, maxDocs=42306)
                0.022939589 = queryNorm
              0.21030651 = fieldWeight in 460, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.3649042 = idf(docFreq=3974, maxDocs=42306)
                0.0625 = fieldNorm(doc=460)
          0.46492064 = weight(abstract_txt:summarization in 460) [ClassicSimilarity], result of:
            0.46492064 = score(doc=460,freq=4.0), product of:
              0.5207864 = queryWeight, product of:
                3.1788118 = boost
                7.1418247 = idf(docFreq=90, maxDocs=42306)
                0.022939589 = queryNorm
              0.8927281 = fieldWeight in 460, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                7.1418247 = idf(docFreq=90, maxDocs=42306)
                0.0625 = fieldNorm(doc=460)
        0.48 = coord(12/25)
    
  2. Vanderwende, L.; Suzuki, H.; Brockett, J.M.; Nenkova, A.: Beyond SumBasic : task-focused summarization with sentence simplification and lexical expansion (2007) 0.45
    0.44611004 = sum of:
      0.44611004 = product of:
        1.0138865 = sum of:
          0.05013892 = weight(abstract_txt:evaluation in 2949) [ClassicSimilarity], result of:
            0.05013892 = score(doc=2949,freq=3.0), product of:
              0.10307658 = queryWeight, product of:
                4.4933925 = idf(docFreq=1285, maxDocs=42306)
                0.022939589 = queryNorm
              0.486424 = fieldWeight in 2949, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                4.4933925 = idf(docFreq=1285, maxDocs=42306)
                0.0625 = fieldNorm(doc=2949)
          0.009223185 = weight(abstract_txt:information in 2949) [ClassicSimilarity], result of:
            0.009223185 = score(doc=2949,freq=1.0), product of:
              0.06058257 = queryWeight, product of:
                1.0841986 = boost
                2.435865 = idf(docFreq=10064, maxDocs=42306)
                0.022939589 = queryNorm
              0.15224156 = fieldWeight in 2949, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                2.435865 = idf(docFreq=10064, maxDocs=42306)
                0.0625 = fieldNorm(doc=2949)
          0.054876164 = weight(abstract_txt:designed in 2949) [ClassicSimilarity], result of:
            0.054876164 = score(doc=2949,freq=2.0), product of:
              0.12531303 = queryWeight, product of:
                1.1026003 = boost
                4.9544163 = idf(docFreq=810, maxDocs=42306)
                0.022939589 = queryNorm
              0.43791267 = fieldWeight in 2949, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                4.9544163 = idf(docFreq=810, maxDocs=42306)
                0.0625 = fieldNorm(doc=2949)
          0.018499946 = weight(abstract_txt:used in 2949) [ClassicSimilarity], result of:
            0.018499946 = score(doc=2949,freq=1.0), product of:
              0.087544285 = queryWeight, product of:
                1.1287026 = boost
                3.381136 = idf(docFreq=3910, maxDocs=42306)
                0.022939589 = queryNorm
              0.211321 = fieldWeight in 2949, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.381136 = idf(docFreq=3910, maxDocs=42306)
                0.0625 = fieldNorm(doc=2949)
          0.094891444 = weight(abstract_txt:topic in 2949) [ClassicSimilarity], result of:
            0.094891444 = score(doc=2949,freq=5.0), product of:
              0.13301842 = queryWeight, product of:
                1.1359936 = boost
                5.104465 = idf(docFreq=697, maxDocs=42306)
                0.022939589 = queryNorm
              0.7133707 = fieldWeight in 2949, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                5.104465 = idf(docFreq=697, maxDocs=42306)
                0.0625 = fieldNorm(doc=2949)
          0.027011942 = weight(abstract_txt:systems in 2949) [ClassicSimilarity], result of:
            0.027011942 = score(doc=2949,freq=2.0), product of:
              0.089428246 = queryWeight, product of:
                1.1407828 = boost
                3.4173236 = idf(docFreq=3771, maxDocs=42306)
                0.022939589 = queryNorm
              0.30205157 = fieldWeight in 2949, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.4173236 = idf(docFreq=3771, maxDocs=42306)
                0.0625 = fieldNorm(doc=2949)
          0.08622973 = weight(abstract_txt:summary in 2949) [ClassicSimilarity], result of:
            0.08622973 = score(doc=2949,freq=1.0), product of:
              0.21339706 = queryWeight, product of:
                1.4388456 = boost
                6.465298 = idf(docFreq=178, maxDocs=42306)
                0.022939589 = queryNorm
              0.40408114 = fieldWeight in 2949, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.465298 = idf(docFreq=178, maxDocs=42306)
                0.0625 = fieldNorm(doc=2949)
          0.11845754 = weight(abstract_txt:components in 2949) [ClassicSimilarity], result of:
            0.11845754 = score(doc=2949,freq=2.0), product of:
              0.23959586 = queryWeight, product of:
                1.8672621 = boost
                5.593561 = idf(docFreq=427, maxDocs=42306)
                0.022939589 = queryNorm
              0.49440563 = fieldWeight in 2949, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                5.593561 = idf(docFreq=427, maxDocs=42306)
                0.0625 = fieldNorm(doc=2949)
          0.0567846 = weight(abstract_txt:each in 2949) [ClassicSimilarity], result of:
            0.0567846 = score(doc=2949,freq=1.0), product of:
              0.219222 = queryWeight, product of:
                2.3058555 = boost
                4.1444454 = idf(docFreq=1822, maxDocs=42306)
                0.022939589 = queryNorm
              0.25902784 = fieldWeight in 2949, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.1444454 = idf(docFreq=1822, maxDocs=42306)
                0.0625 = fieldNorm(doc=2949)
          0.095139846 = weight(abstract_txt:system in 2949) [ClassicSimilarity], result of:
            0.095139846 = score(doc=2949,freq=5.0), product of:
              0.20231342 = queryWeight, product of:
                2.6209962 = boost
                3.3649042 = idf(docFreq=3974, maxDocs=42306)
                0.022939589 = queryNorm
              0.47025967 = fieldWeight in 2949, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                3.3649042 = idf(docFreq=3974, maxDocs=42306)
                0.0625 = fieldNorm(doc=2949)
          0.40263307 = weight(abstract_txt:summarization in 2949) [ClassicSimilarity], result of:
            0.40263307 = score(doc=2949,freq=3.0), product of:
              0.5207864 = queryWeight, product of:
                3.1788118 = boost
                7.1418247 = idf(docFreq=90, maxDocs=42306)
                0.022939589 = queryNorm
              0.7731252 = fieldWeight in 2949, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                7.1418247 = idf(docFreq=90, maxDocs=42306)
                0.0625 = fieldNorm(doc=2949)
        0.44 = coord(11/25)
    
  3. Kar, M.; Nunes, S.; Ribeiro, C.: Summarization of changes in dynamic text collections using Latent Dirichlet Allocation model (2015) 0.38
    0.3829429 = sum of:
      0.3829429 = product of:
        0.95735717 = sum of:
          0.13038567 = weight(abstract_txt:rouge in 4677) [ClassicSimilarity], result of:
            0.13038567 = score(doc=4677,freq=2.0), product of:
              0.21454059 = queryWeight, product of:
                1.0201399 = boost
                9.167778 = idf(docFreq=11, maxDocs=42306)
                0.022939589 = queryNorm
              0.60774356 = fieldWeight in 4677, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                9.167778 = idf(docFreq=11, maxDocs=42306)
                0.046875 = fieldNorm(doc=4677)
          0.011981268 = weight(abstract_txt:information in 4677) [ClassicSimilarity], result of:
            0.011981268 = score(doc=4677,freq=3.0), product of:
              0.06058257 = queryWeight, product of:
                1.0841986 = boost
                2.435865 = idf(docFreq=10064, maxDocs=42306)
                0.022939589 = queryNorm
              0.19776759 = fieldWeight in 4677, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                2.435865 = idf(docFreq=10064, maxDocs=42306)
                0.046875 = fieldNorm(doc=4677)
          0.024032133 = weight(abstract_txt:used in 4677) [ClassicSimilarity], result of:
            0.024032133 = score(doc=4677,freq=3.0), product of:
              0.087544285 = queryWeight, product of:
                1.1287026 = boost
                3.381136 = idf(docFreq=3910, maxDocs=42306)
                0.022939589 = queryNorm
              0.27451402 = fieldWeight in 4677, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                3.381136 = idf(docFreq=3910, maxDocs=42306)
                0.046875 = fieldNorm(doc=4677)
          0.04501096 = weight(abstract_txt:topic in 4677) [ClassicSimilarity], result of:
            0.04501096 = score(doc=4677,freq=2.0), product of:
              0.13301842 = queryWeight, product of:
                1.1359936 = boost
                5.104465 = idf(docFreq=697, maxDocs=42306)
                0.022939589 = queryNorm
              0.3383814 = fieldWeight in 4677, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                5.104465 = idf(docFreq=697, maxDocs=42306)
                0.046875 = fieldNorm(doc=4677)
          0.021075042 = weight(abstract_txt:retrieval in 4677) [ClassicSimilarity], result of:
            0.021075042 = score(doc=4677,freq=2.0), product of:
              0.09181401 = queryWeight, product of:
                1.1558996 = boost
                3.4626071 = idf(docFreq=3604, maxDocs=42306)
                0.022939589 = queryNorm
              0.22954059 = fieldWeight in 4677, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.4626071 = idf(docFreq=3604, maxDocs=42306)
                0.046875 = fieldNorm(doc=4677)
          0.050053276 = weight(abstract_txt:documents in 4677) [ClassicSimilarity], result of:
            0.050053276 = score(doc=4677,freq=4.0), product of:
              0.12972042 = queryWeight, product of:
                1.3739464 = boost
                4.115787 = idf(docFreq=1875, maxDocs=42306)
                0.022939589 = queryNorm
              0.38585502 = fieldWeight in 4677, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                4.115787 = idf(docFreq=1875, maxDocs=42306)
                0.046875 = fieldNorm(doc=4677)
          0.14461164 = weight(abstract_txt:summary in 4677) [ClassicSimilarity], result of:
            0.14461164 = score(doc=4677,freq=5.0), product of:
              0.21339706 = queryWeight, product of:
                1.4388456 = boost
                6.465298 = idf(docFreq=178, maxDocs=42306)
                0.022939589 = queryNorm
              0.67766464 = fieldWeight in 4677, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                6.465298 = idf(docFreq=178, maxDocs=42306)
                0.046875 = fieldNorm(doc=4677)
          0.09523067 = weight(abstract_txt:each in 4677) [ClassicSimilarity], result of:
            0.09523067 = score(doc=4677,freq=5.0), product of:
              0.219222 = queryWeight, product of:
                2.3058555 = boost
                4.1444454 = idf(docFreq=1822, maxDocs=42306)
                0.022939589 = queryNorm
              0.43440288 = fieldWeight in 4677, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                4.1444454 = idf(docFreq=1822, maxDocs=42306)
                0.046875 = fieldNorm(doc=4677)
          0.04512879 = weight(abstract_txt:system in 4677) [ClassicSimilarity], result of:
            0.04512879 = score(doc=4677,freq=2.0), product of:
              0.20231342 = queryWeight, product of:
                2.6209962 = boost
                3.3649042 = idf(docFreq=3974, maxDocs=42306)
                0.022939589 = queryNorm
              0.22306374 = fieldWeight in 4677, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.3649042 = idf(docFreq=3974, maxDocs=42306)
                0.046875 = fieldNorm(doc=4677)
          0.38984782 = weight(abstract_txt:summarization in 4677) [ClassicSimilarity], result of:
            0.38984782 = score(doc=4677,freq=5.0), product of:
              0.5207864 = queryWeight, product of:
                3.1788118 = boost
                7.1418247 = idf(docFreq=90, maxDocs=42306)
                0.022939589 = queryNorm
              0.74857527 = fieldWeight in 4677, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                7.1418247 = idf(docFreq=90, maxDocs=42306)
                0.046875 = fieldNorm(doc=4677)
        0.4 = coord(10/25)
    
  4. Pons-Porrata, A.; Berlanga-Llavori, R.; Ruiz-Shulcloper, J.: Topic discovery based on text mining techniques (2007) 0.34
    0.3365566 = sum of:
      0.3365566 = product of:
        1.0517395 = sum of:
          0.1186143 = weight(abstract_txt:topic in 2917) [ClassicSimilarity], result of:
            0.1186143 = score(doc=2917,freq=5.0), product of:
              0.13301842 = queryWeight, product of:
                1.1359936 = boost
                5.104465 = idf(docFreq=697, maxDocs=42306)
                0.022939589 = queryNorm
              0.8917134 = fieldWeight in 2917, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                5.104465 = idf(docFreq=697, maxDocs=42306)
                0.078125 = fieldNorm(doc=2917)
          0.064713635 = weight(abstract_txt:demonstrate in 2917) [ClassicSimilarity], result of:
            0.064713635 = score(doc=2917,freq=1.0), product of:
              0.1518708 = queryWeight, product of:
                1.213828 = boost
                5.4542055 = idf(docFreq=491, maxDocs=42306)
                0.022939589 = queryNorm
              0.4261098 = fieldWeight in 2917, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.4542055 = idf(docFreq=491, maxDocs=42306)
                0.078125 = fieldNorm(doc=2917)
          0.058988355 = weight(abstract_txt:documents in 2917) [ClassicSimilarity], result of:
            0.058988355 = score(doc=2917,freq=2.0), product of:
              0.12972042 = queryWeight, product of:
                1.3739464 = boost
                4.115787 = idf(docFreq=1875, maxDocs=42306)
                0.022939589 = queryNorm
              0.45473453 = fieldWeight in 2917, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                4.115787 = idf(docFreq=1875, maxDocs=42306)
                0.078125 = fieldNorm(doc=2917)
          0.10778716 = weight(abstract_txt:summary in 2917) [ClassicSimilarity], result of:
            0.10778716 = score(doc=2917,freq=1.0), product of:
              0.21339706 = queryWeight, product of:
                1.4388456 = boost
                6.465298 = idf(docFreq=178, maxDocs=42306)
                0.022939589 = queryNorm
              0.50510144 = fieldWeight in 2917, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.465298 = idf(docFreq=178, maxDocs=42306)
                0.078125 = fieldNorm(doc=2917)
          0.14450496 = weight(abstract_txt:clustering in 2917) [ClassicSimilarity], result of:
            0.14450496 = score(doc=2917,freq=1.0), product of:
              0.29700425 = queryWeight, product of:
                2.078964 = boost
                6.227734 = idf(docFreq=226, maxDocs=42306)
                0.022939589 = queryNorm
              0.48654172 = fieldWeight in 2917, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.227734 = idf(docFreq=226, maxDocs=42306)
                0.078125 = fieldNorm(doc=2917)
          0.07098075 = weight(abstract_txt:each in 2917) [ClassicSimilarity], result of:
            0.07098075 = score(doc=2917,freq=1.0), product of:
              0.219222 = queryWeight, product of:
                2.3058555 = boost
                4.1444454 = idf(docFreq=1822, maxDocs=42306)
                0.022939589 = queryNorm
              0.3237848 = fieldWeight in 2917, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.1444454 = idf(docFreq=1822, maxDocs=42306)
                0.078125 = fieldNorm(doc=2917)
          0.075214654 = weight(abstract_txt:system in 2917) [ClassicSimilarity], result of:
            0.075214654 = score(doc=2917,freq=2.0), product of:
              0.20231342 = queryWeight, product of:
                2.6209962 = boost
                3.3649042 = idf(docFreq=3974, maxDocs=42306)
                0.022939589 = queryNorm
              0.37177292 = fieldWeight in 2917, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.3649042 = idf(docFreq=3974, maxDocs=42306)
                0.078125 = fieldNorm(doc=2917)
          0.4109357 = weight(abstract_txt:summarization in 2917) [ClassicSimilarity], result of:
            0.4109357 = score(doc=2917,freq=2.0), product of:
              0.5207864 = queryWeight, product of:
                3.1788118 = boost
                7.1418247 = idf(docFreq=90, maxDocs=42306)
                0.022939589 = queryNorm
              0.7890676 = fieldWeight in 2917, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                7.1418247 = idf(docFreq=90, maxDocs=42306)
                0.078125 = fieldNorm(doc=2917)
        0.32 = coord(8/25)
    
  5. Cai, X.; Li, W.: Enhancing sentence-level clustering with integrated and interactive frameworks for theme-based summarization (2011) 0.34
    0.3360075 = sum of:
      0.3360075 = product of:
        1.0500234 = sum of:
          0.04093826 = weight(abstract_txt:evaluation in 1771) [ClassicSimilarity], result of:
            0.04093826 = score(doc=1771,freq=2.0), product of:
              0.10307658 = queryWeight, product of:
                4.4933925 = idf(docFreq=1285, maxDocs=42306)
                0.022939589 = queryNorm
              0.39716354 = fieldWeight in 1771, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                4.4933925 = idf(docFreq=1285, maxDocs=42306)
                0.0625 = fieldNorm(doc=1771)
          0.031775363 = weight(abstract_txt:framework in 1771) [ClassicSimilarity], result of:
            0.031775363 = score(doc=1771,freq=1.0), product of:
              0.10968421 = queryWeight, product of:
                1.0315542 = boost
                4.635178 = idf(docFreq=1115, maxDocs=42306)
                0.022939589 = queryNorm
              0.28969863 = fieldWeight in 1771, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.635178 = idf(docFreq=1115, maxDocs=42306)
                0.0625 = fieldNorm(doc=1771)
          0.009223185 = weight(abstract_txt:information in 1771) [ClassicSimilarity], result of:
            0.009223185 = score(doc=1771,freq=1.0), product of:
              0.06058257 = queryWeight, product of:
                1.0841986 = boost
                2.435865 = idf(docFreq=10064, maxDocs=42306)
                0.022939589 = queryNorm
              0.15224156 = fieldWeight in 1771, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                2.435865 = idf(docFreq=10064, maxDocs=42306)
                0.0625 = fieldNorm(doc=1771)
          0.018499946 = weight(abstract_txt:used in 1771) [ClassicSimilarity], result of:
            0.018499946 = score(doc=1771,freq=1.0), product of:
              0.087544285 = queryWeight, product of:
                1.1287026 = boost
                3.381136 = idf(docFreq=3910, maxDocs=42306)
                0.022939589 = queryNorm
              0.211321 = fieldWeight in 1771, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.381136 = idf(docFreq=3910, maxDocs=42306)
                0.0625 = fieldNorm(doc=1771)
          0.04243674 = weight(abstract_txt:topic in 1771) [ClassicSimilarity], result of:
            0.04243674 = score(doc=1771,freq=1.0), product of:
              0.13301842 = queryWeight, product of:
                1.1359936 = boost
                5.104465 = idf(docFreq=697, maxDocs=42306)
                0.022939589 = queryNorm
              0.31902906 = fieldWeight in 1771, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.104465 = idf(docFreq=697, maxDocs=42306)
                0.0625 = fieldNorm(doc=1771)
          0.44773227 = weight(abstract_txt:clustering in 1771) [ClassicSimilarity], result of:
            0.44773227 = score(doc=1771,freq=15.0), product of:
              0.29700425 = queryWeight, product of:
                2.078964 = boost
                6.227734 = idf(docFreq=226, maxDocs=42306)
                0.022939589 = queryNorm
              1.5074944 = fieldWeight in 1771, product of:
                3.8729835 = tf(freq=15.0), with freq of:
                  15.0 = termFreq=15.0
                6.227734 = idf(docFreq=226, maxDocs=42306)
                0.0625 = fieldNorm(doc=1771)
          0.0567846 = weight(abstract_txt:each in 1771) [ClassicSimilarity], result of:
            0.0567846 = score(doc=1771,freq=1.0), product of:
              0.219222 = queryWeight, product of:
                2.3058555 = boost
                4.1444454 = idf(docFreq=1822, maxDocs=42306)
                0.022939589 = queryNorm
              0.25902784 = fieldWeight in 1771, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.1444454 = idf(docFreq=1822, maxDocs=42306)
                0.0625 = fieldNorm(doc=1771)
          0.40263307 = weight(abstract_txt:summarization in 1771) [ClassicSimilarity], result of:
            0.40263307 = score(doc=1771,freq=3.0), product of:
              0.5207864 = queryWeight, product of:
                3.1788118 = boost
                7.1418247 = idf(docFreq=90, maxDocs=42306)
                0.022939589 = queryNorm
              0.7731252 = fieldWeight in 1771, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                7.1418247 = idf(docFreq=90, maxDocs=42306)
                0.0625 = fieldNorm(doc=1771)
        0.32 = coord(8/25)