Document (#31345)

Author
Wu, S.
McClean, S.I.
Title
Improving high accuracy retrieval by eliminating the uneven correlation effect in data fusion
Source
Journal of the American Society for Information Science and Technology. 57(2006) no.14, S.1962-1973
Year
2006
Abstract
The aim of this research is twofold. On the one hand, high accuracy retrieval has been a concern of the information retrieval community for some time. We aim to investigate this issue via data fusion. On the other hand, the correlation among component results has been proven harmful to data fusion, but it has not been taken into account in data fusion algorithms. In the hope of achieving better performance, we propose a group of algorithms to eliminate the effect of uneven correlation among component results by assigning different weights to all component results or their combinations. Then the linear combination method or a variation is used for fusion. Extensive experimentation is carried out to evaluate the performances of these algorithms with six groups of component results, which are the top 10 systems submitted to Text REtrieval Conference (TREC) 6, 7, 8, 9, 2001, and 2002. The experimental results show that all eight data fusion methods involved outperform the best component system on average. Therefore, we demonstrate that the data fusion technique in general is effective with accurate retrieval results. The experimental results also demonstrate that all six methods presented in this article are effective for eliminating the effect of uneven correlation among component results. All of them outperform CombSum and five of them outperform CombMNZ on average.

Similar documents (content)

  1. Beitzel, S.M.; Jensen, E.C.; Chowdhury, A.; Grossman, D.; Frieder, O; Goharian, N.: Fusion of effective retrieval strategies in the same information retrieval system (2004) 0.28
    0.28409567 = sum of:
      0.28409567 = product of:
        1.0146275 = sum of:
          0.037500843 = weight(abstract_txt:effective in 2502) [ClassicSimilarity], result of:
            0.037500843 = score(doc=2502,freq=2.0), product of:
              0.07018263 = queryWeight, product of:
                1.2223957 = boost
                4.8362236 = idf(docFreq=953, maxDocs=44218)
                0.01187166 = queryNorm
              0.5343323 = fieldWeight in 2502, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                4.8362236 = idf(docFreq=953, maxDocs=44218)
                0.078125 = fieldNorm(doc=2502)
          0.036610097 = weight(abstract_txt:demonstrate in 2502) [ClassicSimilarity], result of:
            0.036610097 = score(doc=2502,freq=1.0), product of:
              0.08701875 = queryWeight, product of:
                1.3611419 = boost
                5.3851523 = idf(docFreq=550, maxDocs=44218)
                0.01187166 = queryNorm
              0.42071503 = fieldWeight in 2502, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.3851523 = idf(docFreq=550, maxDocs=44218)
                0.078125 = fieldNorm(doc=2502)
          0.023543296 = weight(abstract_txt:been in 2502) [ClassicSimilarity], result of:
            0.023543296 = score(doc=2502,freq=2.0), product of:
              0.05890392 = queryWeight, product of:
                1.3715597 = boost
                3.617579 = idf(docFreq=3226, maxDocs=44218)
                0.01187166 = queryNorm
              0.3996898 = fieldWeight in 2502, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.617579 = idf(docFreq=3226, maxDocs=44218)
                0.078125 = fieldNorm(doc=2502)
          0.069567844 = weight(abstract_txt:retrieval in 2502) [ClassicSimilarity], result of:
            0.069567844 = score(doc=2502,freq=8.0), product of:
              0.090594396 = queryWeight, product of:
                2.1959257 = boost
                3.4751394 = idf(docFreq=3720, maxDocs=44218)
                0.01187166 = queryNorm
              0.7679045 = fieldWeight in 2502, product of:
                2.828427 = tf(freq=8.0), with freq of:
                  8.0 = termFreq=8.0
                3.4751394 = idf(docFreq=3720, maxDocs=44218)
                0.078125 = fieldNorm(doc=2502)
          0.045237932 = weight(abstract_txt:data in 2502) [ClassicSimilarity], result of:
            0.045237932 = score(doc=2502,freq=3.0), product of:
              0.10020301 = queryWeight, product of:
                2.5298688 = boost
                3.3363478 = idf(docFreq=4274, maxDocs=44218)
                0.01187166 = queryNorm
              0.4514628 = fieldWeight in 2502, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                3.3363478 = idf(docFreq=4274, maxDocs=44218)
                0.078125 = fieldNorm(doc=2502)
          0.039601456 = weight(abstract_txt:results in 2502) [ClassicSimilarity], result of:
            0.039601456 = score(doc=2502,freq=1.0), product of:
              0.1455592 = queryWeight, product of:
                3.5208442 = boost
                3.482422 = idf(docFreq=3693, maxDocs=44218)
                0.01187166 = queryNorm
              0.27206424 = fieldWeight in 2502, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.482422 = idf(docFreq=3693, maxDocs=44218)
                0.078125 = fieldNorm(doc=2502)
          0.76256603 = weight(abstract_txt:fusion in 2502) [ClassicSimilarity], result of:
            0.76256603 = score(doc=2502,freq=4.0), product of:
              0.6300861 = queryWeight, product of:
                6.852215 = boost
                7.7456436 = idf(docFreq=51, maxDocs=44218)
                0.01187166 = queryNorm
              1.2102568 = fieldWeight in 2502, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                7.7456436 = idf(docFreq=51, maxDocs=44218)
                0.078125 = fieldNorm(doc=2502)
        0.28 = coord(7/25)
    
  2. Larsen, B.; Ingwersen, P.; Lund, B.: Data fusion according to the principle of polyrepresentation (2009) 0.23
    0.22839433 = sum of:
      0.22839433 = product of:
        0.95164305 = sum of:
          0.016547946 = weight(abstract_txt:methods in 2752) [ClassicSimilarity], result of:
            0.016547946 = score(doc=2752,freq=2.0), product of:
              0.051598012 = queryWeight, product of:
                1.048126 = boost
                4.146752 = idf(docFreq=1900, maxDocs=44218)
                0.01187166 = queryNorm
              0.320709 = fieldWeight in 2752, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                4.146752 = idf(docFreq=1900, maxDocs=44218)
                0.0546875 = fieldNorm(doc=2752)
          0.03849875 = weight(abstract_txt:retrieval in 2752) [ClassicSimilarity], result of:
            0.03849875 = score(doc=2752,freq=5.0), product of:
              0.090594396 = queryWeight, product of:
                2.1959257 = boost
                3.4751394 = idf(docFreq=3720, maxDocs=44218)
                0.01187166 = queryNorm
              0.4249573 = fieldWeight in 2752, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                3.4751394 = idf(docFreq=3720, maxDocs=44218)
                0.0546875 = fieldNorm(doc=2752)
          0.040881343 = weight(abstract_txt:data in 2752) [ClassicSimilarity], result of:
            0.040881343 = score(doc=2752,freq=5.0), product of:
              0.10020301 = queryWeight, product of:
                2.5298688 = boost
                3.3363478 = idf(docFreq=4274, maxDocs=44218)
                0.01187166 = queryNorm
              0.40798518 = fieldWeight in 2752, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                3.3363478 = idf(docFreq=4274, maxDocs=44218)
                0.0546875 = fieldNorm(doc=2752)
          0.11036557 = weight(abstract_txt:outperform in 2752) [ClassicSimilarity], result of:
            0.11036557 = score(doc=2752,freq=1.0), product of:
              0.26367345 = queryWeight, product of:
                2.9018557 = boost
                7.653836 = idf(docFreq=56, maxDocs=44218)
                0.01187166 = queryNorm
              0.41856915 = fieldWeight in 2752, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.653836 = idf(docFreq=56, maxDocs=44218)
                0.0546875 = fieldNorm(doc=2752)
          0.03920344 = weight(abstract_txt:results in 2752) [ClassicSimilarity], result of:
            0.03920344 = score(doc=2752,freq=2.0), product of:
              0.1455592 = queryWeight, product of:
                3.5208442 = boost
                3.482422 = idf(docFreq=3693, maxDocs=44218)
                0.01187166 = queryNorm
              0.26932985 = fieldWeight in 2752, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.482422 = idf(docFreq=3693, maxDocs=44218)
                0.0546875 = fieldNorm(doc=2752)
          0.706146 = weight(abstract_txt:fusion in 2752) [ClassicSimilarity], result of:
            0.706146 = score(doc=2752,freq=7.0), product of:
              0.6300861 = queryWeight, product of:
                6.852215 = boost
                7.7456436 = idf(docFreq=51, maxDocs=44218)
                0.01187166 = queryNorm
              1.1207135 = fieldWeight in 2752, product of:
                2.6457512 = tf(freq=7.0), with freq of:
                  7.0 = termFreq=7.0
                7.7456436 = idf(docFreq=51, maxDocs=44218)
                0.0546875 = fieldNorm(doc=2752)
        0.24 = coord(6/25)
    
  3. Liu, X.; Yu, S.; Janssens, F.; Glänzel, W.; Moreau, Y.; Moor, B.de: Weighted hybrid clustering by combining text mining and bibliometrics on a large-scale journal database (2010) 0.23
    0.2268878 = sum of:
      0.2268878 = product of:
        0.7090244 = sum of:
          0.028952876 = weight(abstract_txt:methods in 3464) [ClassicSimilarity], result of:
            0.028952876 = score(doc=3464,freq=3.0), product of:
              0.051598012 = queryWeight, product of:
                1.048126 = boost
                4.146752 = idf(docFreq=1900, maxDocs=44218)
                0.01187166 = queryNorm
              0.56112385 = fieldWeight in 3464, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                4.146752 = idf(docFreq=1900, maxDocs=44218)
                0.078125 = fieldNorm(doc=3464)
          0.036610097 = weight(abstract_txt:demonstrate in 3464) [ClassicSimilarity], result of:
            0.036610097 = score(doc=3464,freq=1.0), product of:
              0.08701875 = queryWeight, product of:
                1.3611419 = boost
                5.3851523 = idf(docFreq=550, maxDocs=44218)
                0.01187166 = queryNorm
              0.42071503 = fieldWeight in 3464, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.3851523 = idf(docFreq=550, maxDocs=44218)
                0.078125 = fieldNorm(doc=3464)
          0.036871474 = weight(abstract_txt:experimental in 3464) [ClassicSimilarity], result of:
            0.036871474 = score(doc=3464,freq=1.0), product of:
              0.087432444 = queryWeight, product of:
                1.3643736 = boost
                5.397938 = idf(docFreq=543, maxDocs=44218)
                0.01187166 = queryNorm
              0.4217139 = fieldWeight in 3464, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.397938 = idf(docFreq=543, maxDocs=44218)
                0.078125 = fieldNorm(doc=3464)
          0.056288853 = weight(abstract_txt:effect in 3464) [ClassicSimilarity], result of:
            0.056288853 = score(doc=3464,freq=1.0), product of:
              0.13269593 = queryWeight, product of:
                2.0585973 = boost
                5.4296865 = idf(docFreq=526, maxDocs=44218)
                0.01187166 = queryNorm
              0.42419428 = fieldWeight in 3464, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.4296865 = idf(docFreq=526, maxDocs=44218)
                0.078125 = fieldNorm(doc=3464)
          0.09248006 = weight(abstract_txt:algorithms in 3464) [ClassicSimilarity], result of:
            0.09248006 = score(doc=3464,freq=2.0), product of:
              0.14664416 = queryWeight, product of:
                2.1640885 = boost
                5.707926 = idf(docFreq=398, maxDocs=44218)
                0.01187166 = queryNorm
              0.63064265 = fieldWeight in 3464, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                5.707926 = idf(docFreq=398, maxDocs=44218)
                0.078125 = fieldNorm(doc=3464)
          0.036936615 = weight(abstract_txt:data in 3464) [ClassicSimilarity], result of:
            0.036936615 = score(doc=3464,freq=2.0), product of:
              0.10020301 = queryWeight, product of:
                2.5298688 = boost
                3.3363478 = idf(docFreq=4274, maxDocs=44218)
                0.01187166 = queryNorm
              0.36861783 = fieldWeight in 3464, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.3363478 = idf(docFreq=4274, maxDocs=44218)
                0.078125 = fieldNorm(doc=3464)
          0.039601456 = weight(abstract_txt:results in 3464) [ClassicSimilarity], result of:
            0.039601456 = score(doc=3464,freq=1.0), product of:
              0.1455592 = queryWeight, product of:
                3.5208442 = boost
                3.482422 = idf(docFreq=3693, maxDocs=44218)
                0.01187166 = queryNorm
              0.27206424 = fieldWeight in 3464, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.482422 = idf(docFreq=3693, maxDocs=44218)
                0.078125 = fieldNorm(doc=3464)
          0.38128302 = weight(abstract_txt:fusion in 3464) [ClassicSimilarity], result of:
            0.38128302 = score(doc=3464,freq=1.0), product of:
              0.6300861 = queryWeight, product of:
                6.852215 = boost
                7.7456436 = idf(docFreq=51, maxDocs=44218)
                0.01187166 = queryNorm
              0.6051284 = fieldWeight in 3464, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.7456436 = idf(docFreq=51, maxDocs=44218)
                0.078125 = fieldNorm(doc=3464)
        0.32 = coord(8/25)
    
  4. Wu, S.; Li, J.; Zeng, X.; Bi, Y.: Adaptive data fusion methods in information retrieval (2014) 0.18
    0.18410891 = sum of:
      0.18410891 = product of:
        0.92054456 = sum of:
          0.0334319 = weight(abstract_txt:methods in 1500) [ClassicSimilarity], result of:
            0.0334319 = score(doc=1500,freq=4.0), product of:
              0.051598012 = queryWeight, product of:
                1.048126 = boost
                4.146752 = idf(docFreq=1900, maxDocs=44218)
                0.01187166 = queryNorm
              0.64792997 = fieldWeight in 1500, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                4.146752 = idf(docFreq=1900, maxDocs=44218)
                0.078125 = fieldNorm(doc=1500)
          0.023543296 = weight(abstract_txt:been in 1500) [ClassicSimilarity], result of:
            0.023543296 = score(doc=1500,freq=2.0), product of:
              0.05890392 = queryWeight, product of:
                1.3715597 = boost
                3.617579 = idf(docFreq=3226, maxDocs=44218)
                0.01187166 = queryNorm
              0.3996898 = fieldWeight in 1500, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.617579 = idf(docFreq=3226, maxDocs=44218)
                0.078125 = fieldNorm(doc=1500)
          0.042601433 = weight(abstract_txt:retrieval in 1500) [ClassicSimilarity], result of:
            0.042601433 = score(doc=1500,freq=3.0), product of:
              0.090594396 = queryWeight, product of:
                2.1959257 = boost
                3.4751394 = idf(docFreq=3720, maxDocs=44218)
                0.01187166 = queryNorm
              0.47024357 = fieldWeight in 1500, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                3.4751394 = idf(docFreq=3720, maxDocs=44218)
                0.078125 = fieldNorm(doc=1500)
          0.058401916 = weight(abstract_txt:data in 1500) [ClassicSimilarity], result of:
            0.058401916 = score(doc=1500,freq=5.0), product of:
              0.10020301 = queryWeight, product of:
                2.5298688 = boost
                3.3363478 = idf(docFreq=4274, maxDocs=44218)
                0.01187166 = queryNorm
              0.582836 = fieldWeight in 1500, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                3.3363478 = idf(docFreq=4274, maxDocs=44218)
                0.078125 = fieldNorm(doc=1500)
          0.76256603 = weight(abstract_txt:fusion in 1500) [ClassicSimilarity], result of:
            0.76256603 = score(doc=1500,freq=4.0), product of:
              0.6300861 = queryWeight, product of:
                6.852215 = boost
                7.7456436 = idf(docFreq=51, maxDocs=44218)
                0.01187166 = queryNorm
              1.2102568 = fieldWeight in 1500, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                7.7456436 = idf(docFreq=51, maxDocs=44218)
                0.078125 = fieldNorm(doc=1500)
        0.2 = coord(5/25)
    
  5. Wu, M.; Hawking, D.; Turpin, A.; Scholer, F.: Using anchor text for homepage and topic distillation search tasks (2012) 0.17
    0.16607189 = sum of:
      0.16607189 = product of:
        0.5931139 = sum of:
          0.0231623 = weight(abstract_txt:methods in 257) [ClassicSimilarity], result of:
            0.0231623 = score(doc=257,freq=3.0), product of:
              0.051598012 = queryWeight, product of:
                1.048126 = boost
                4.146752 = idf(docFreq=1900, maxDocs=44218)
                0.01187166 = queryNorm
              0.44889906 = fieldWeight in 257, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                4.146752 = idf(docFreq=1900, maxDocs=44218)
                0.0625 = fieldNorm(doc=257)
          0.02121368 = weight(abstract_txt:effective in 257) [ClassicSimilarity], result of:
            0.02121368 = score(doc=257,freq=1.0), product of:
              0.07018263 = queryWeight, product of:
                1.2223957 = boost
                4.8362236 = idf(docFreq=953, maxDocs=44218)
                0.01187166 = queryNorm
              0.30226398 = fieldWeight in 257, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.8362236 = idf(docFreq=953, maxDocs=44218)
                0.0625 = fieldNorm(doc=257)
          0.02949718 = weight(abstract_txt:experimental in 257) [ClassicSimilarity], result of:
            0.02949718 = score(doc=257,freq=1.0), product of:
              0.087432444 = queryWeight, product of:
                1.3643736 = boost
                5.397938 = idf(docFreq=543, maxDocs=44218)
                0.01187166 = queryNorm
              0.3373711 = fieldWeight in 257, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.397938 = idf(docFreq=543, maxDocs=44218)
                0.0625 = fieldNorm(doc=257)
          0.013318099 = weight(abstract_txt:been in 257) [ClassicSimilarity], result of:
            0.013318099 = score(doc=257,freq=1.0), product of:
              0.05890392 = queryWeight, product of:
                1.3715597 = boost
                3.617579 = idf(docFreq=3226, maxDocs=44218)
                0.01187166 = queryNorm
              0.22609869 = fieldWeight in 257, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.617579 = idf(docFreq=3226, maxDocs=44218)
                0.0625 = fieldNorm(doc=257)
          0.01967676 = weight(abstract_txt:retrieval in 257) [ClassicSimilarity], result of:
            0.01967676 = score(doc=257,freq=1.0), product of:
              0.090594396 = queryWeight, product of:
                2.1959257 = boost
                3.4751394 = idf(docFreq=3720, maxDocs=44218)
                0.01187166 = queryNorm
              0.21719621 = fieldWeight in 257, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.4751394 = idf(docFreq=3720, maxDocs=44218)
                0.0625 = fieldNorm(doc=257)
          0.054873385 = weight(abstract_txt:results in 257) [ClassicSimilarity], result of:
            0.054873385 = score(doc=257,freq=3.0), product of:
              0.1455592 = queryWeight, product of:
                3.5208442 = boost
                3.482422 = idf(docFreq=3693, maxDocs=44218)
                0.01187166 = queryNorm
              0.37698326 = fieldWeight in 257, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                3.482422 = idf(docFreq=3693, maxDocs=44218)
                0.0625 = fieldNorm(doc=257)
          0.43137246 = weight(abstract_txt:fusion in 257) [ClassicSimilarity], result of:
            0.43137246 = score(doc=257,freq=2.0), product of:
              0.6300861 = queryWeight, product of:
                6.852215 = boost
                7.7456436 = idf(docFreq=51, maxDocs=44218)
                0.01187166 = queryNorm
              0.6846246 = fieldWeight in 257, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                7.7456436 = idf(docFreq=51, maxDocs=44218)
                0.0625 = fieldNorm(doc=257)
        0.28 = coord(7/25)