Document (#28460)

Author
Ding, C.H.Q.
Title
¬A probabilistic model for Latent Semantic Indexing
Source
Journal of the American Society for Information Science and Technology. 56(2005) no.6, S.597-608
Year
2005
Abstract
Latent Semantic Indexing (LSI), when applied to semantic space built an text collections, improves information retrieval, information filtering, and word sense disambiguation. A new dual probability model based an the similarity concepts is introduced to provide deeper understanding of LSI. Semantic associations can be quantitatively characterized by their statistical significance, the likelihood. Semantic dimensions containing redundant and noisy information can be separated out and should be ignored because their negative contribution to the overall statistical significance. LSI is the optimal solution of the model. The peak in the likelihood curve indicates the existence of an intrinsic semantic dimension. The importance of LSI dimensions follows the Zipf-distribution, indicating that LSI dimensions represent latent concepts. Document frequency of words follows the Zipf distribution, and the number of distinct words follows log-normal distribution. Experiments an five standard document collections confirm and illustrate the analysis.
Theme
Retrievalstudien
Object
Latent Semantic Indexing

Similar documents (author)

  1. Ding, Y.: Visualization of intellectual structure in information retrieval : author cocitation analysis (1998) 4.77
    4.7727776 = sum of:
      4.7727776 = weight(author_txt:ding in 2792) [ClassicSimilarity], result of:
        4.7727776 = fieldWeight in 2792, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          7.636444 = idf(docFreq=57, maxDocs=44218)
          0.625 = fieldNorm(doc=2792)
    
  2. Ding, Y.: Scholarly communication and bibliometrics : Part 1: The scholarly communication model: literature review (1998) 4.77
    4.7727776 = sum of:
      4.7727776 = weight(author_txt:ding in 3995) [ClassicSimilarity], result of:
        4.7727776 = fieldWeight in 3995, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          7.636444 = idf(docFreq=57, maxDocs=44218)
          0.625 = fieldNorm(doc=3995)
    
  3. Ding, Y.: ¬A review of ontologies with the Semantic Web in view (2001) 4.77
    4.7727776 = sum of:
      4.7727776 = weight(author_txt:ding in 4152) [ClassicSimilarity], result of:
        4.7727776 = fieldWeight in 4152, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          7.636444 = idf(docFreq=57, maxDocs=44218)
          0.625 = fieldNorm(doc=4152)
    
  4. Ding, Y.: Applying weighted PageRank to author citation networks (2011) 4.77
    4.7727776 = sum of:
      4.7727776 = weight(author_txt:ding in 4188) [ClassicSimilarity], result of:
        4.7727776 = fieldWeight in 4188, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          7.636444 = idf(docFreq=57, maxDocs=44218)
          0.625 = fieldNorm(doc=4188)
    
  5. Ding, Y.: Topic-based PageRank on author cocitation networks (2011) 4.77
    4.7727776 = sum of:
      4.7727776 = weight(author_txt:ding in 4348) [ClassicSimilarity], result of:
        4.7727776 = fieldWeight in 4348, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          7.636444 = idf(docFreq=57, maxDocs=44218)
          0.625 = fieldNorm(doc=4348)
    

Similar documents (content)

  1. Zhu, W.Z.; Allen, R.B.: Document clustering using the LSI subspace signature model (2013) 0.31
    0.3065915 = sum of:
      0.3065915 = product of:
        0.95809853 = sum of:
          0.051591482 = weight(abstract_txt:document in 690) [ClassicSimilarity], result of:
            0.051591482 = score(doc=690,freq=3.0), product of:
              0.08881904 = queryWeight, product of:
                1.1504045 = boost
                4.2926083 = idf(docFreq=1642, maxDocs=44218)
                0.017985987 = queryNorm
              0.5808606 = fieldWeight in 690, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                4.2926083 = idf(docFreq=1642, maxDocs=44218)
                0.078125 = fieldNorm(doc=690)
          0.043824077 = weight(abstract_txt:indexing in 690) [ClassicSimilarity], result of:
            0.043824077 = score(doc=690,freq=2.0), product of:
              0.09119262 = queryWeight, product of:
                1.1656747 = boost
                4.3495874 = idf(docFreq=1551, maxDocs=44218)
                0.017985987 = queryNorm
              0.48056605 = fieldWeight in 690, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                4.3495874 = idf(docFreq=1551, maxDocs=44218)
                0.078125 = fieldNorm(doc=690)
          0.06424806 = weight(abstract_txt:statistical in 690) [ClassicSimilarity], result of:
            0.06424806 = score(doc=690,freq=1.0), product of:
              0.14827497 = queryWeight, product of:
                1.4863855 = boost
                5.5462847 = idf(docFreq=468, maxDocs=44218)
                0.017985987 = queryNorm
              0.43330348 = fieldWeight in 690, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.5462847 = idf(docFreq=468, maxDocs=44218)
                0.078125 = fieldNorm(doc=690)
          0.071558826 = weight(abstract_txt:model in 690) [ClassicSimilarity], result of:
            0.071558826 = score(doc=690,freq=4.0), product of:
              0.11488952 = queryWeight, product of:
                1.6024458 = boost
                3.986234 = idf(docFreq=2231, maxDocs=44218)
                0.017985987 = queryNorm
              0.62284905 = fieldWeight in 690, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                3.986234 = idf(docFreq=2231, maxDocs=44218)
                0.078125 = fieldNorm(doc=690)
          0.099017106 = weight(abstract_txt:distribution in 690) [ClassicSimilarity], result of:
            0.099017106 = score(doc=690,freq=1.0), product of:
              0.22646359 = queryWeight, product of:
                2.2497919 = boost
                5.596568 = idf(docFreq=445, maxDocs=44218)
                0.017985987 = queryNorm
              0.4372319 = fieldWeight in 690, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.596568 = idf(docFreq=445, maxDocs=44218)
                0.078125 = fieldNorm(doc=690)
          0.11751898 = weight(abstract_txt:dimensions in 690) [ClassicSimilarity], result of:
            0.11751898 = score(doc=690,freq=1.0), product of:
              0.25386155 = queryWeight, product of:
                2.3819993 = boost
                5.925446 = idf(docFreq=320, maxDocs=44218)
                0.017985987 = queryNorm
              0.46292546 = fieldWeight in 690, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.925446 = idf(docFreq=320, maxDocs=44218)
                0.078125 = fieldNorm(doc=690)
          0.33506623 = weight(abstract_txt:latent in 690) [ClassicSimilarity], result of:
            0.33506623 = score(doc=690,freq=3.0), product of:
              0.35391986 = queryWeight, product of:
                2.81252 = boost
                6.996407 = idf(docFreq=109, maxDocs=44218)
                0.017985987 = queryNorm
              0.9467291 = fieldWeight in 690, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                6.996407 = idf(docFreq=109, maxDocs=44218)
                0.078125 = fieldNorm(doc=690)
          0.17527379 = weight(abstract_txt:semantic in 690) [ClassicSimilarity], result of:
            0.17527379 = score(doc=690,freq=3.0), product of:
              0.28949374 = queryWeight, product of:
                3.5973089 = boost
                4.4743214 = idf(docFreq=1369, maxDocs=44218)
                0.017985987 = queryNorm
              0.6054493 = fieldWeight in 690, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                4.4743214 = idf(docFreq=1369, maxDocs=44218)
                0.078125 = fieldNorm(doc=690)
        0.32 = coord(8/25)
    
  2. Li, D.; Kwong, C.-P.; Lee, D.L.: Unified linear subspace approach to semantic analysis (2009) 0.20
    0.1953743 = sum of:
      0.1953743 = product of:
        0.69776535 = sum of:
          0.101288125 = weight(abstract_txt:dual in 3321) [ClassicSimilarity], result of:
            0.101288125 = score(doc=3321,freq=2.0), product of:
              0.14682056 = queryWeight, product of:
                1.0458658 = boost
                7.805067 = idf(docFreq=48, maxDocs=44218)
                0.017985987 = queryNorm
              0.689877 = fieldWeight in 3321, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                7.805067 = idf(docFreq=48, maxDocs=44218)
                0.0625 = fieldNorm(doc=3321)
          0.033699416 = weight(abstract_txt:document in 3321) [ClassicSimilarity], result of:
            0.033699416 = score(doc=3321,freq=2.0), product of:
              0.08881904 = queryWeight, product of:
                1.1504045 = boost
                4.2926083 = idf(docFreq=1642, maxDocs=44218)
                0.017985987 = queryNorm
              0.37941656 = fieldWeight in 3321, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                4.2926083 = idf(docFreq=1642, maxDocs=44218)
                0.0625 = fieldNorm(doc=3321)
          0.02479064 = weight(abstract_txt:indexing in 3321) [ClassicSimilarity], result of:
            0.02479064 = score(doc=3321,freq=1.0), product of:
              0.09119262 = queryWeight, product of:
                1.1656747 = boost
                4.3495874 = idf(docFreq=1551, maxDocs=44218)
                0.017985987 = queryNorm
              0.27184922 = fieldWeight in 3321, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.3495874 = idf(docFreq=1551, maxDocs=44218)
                0.0625 = fieldNorm(doc=3321)
          0.031154688 = weight(abstract_txt:collections in 3321) [ClassicSimilarity], result of:
            0.031154688 = score(doc=3321,freq=1.0), product of:
              0.10619811 = queryWeight, product of:
                1.2579284 = boost
                4.693822 = idf(docFreq=1099, maxDocs=44218)
                0.017985987 = queryNorm
              0.29336387 = fieldWeight in 3321, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.693822 = idf(docFreq=1099, maxDocs=44218)
                0.0625 = fieldNorm(doc=3321)
          0.040479783 = weight(abstract_txt:model in 3321) [ClassicSimilarity], result of:
            0.040479783 = score(doc=3321,freq=2.0), product of:
              0.11488952 = queryWeight, product of:
                1.6024458 = boost
                3.986234 = idf(docFreq=2231, maxDocs=44218)
                0.017985987 = queryNorm
              0.35233662 = fieldWeight in 3321, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.986234 = idf(docFreq=2231, maxDocs=44218)
                0.0625 = fieldNorm(doc=3321)
          0.268053 = weight(abstract_txt:latent in 3321) [ClassicSimilarity], result of:
            0.268053 = score(doc=3321,freq=3.0), product of:
              0.35391986 = queryWeight, product of:
                2.81252 = boost
                6.996407 = idf(docFreq=109, maxDocs=44218)
                0.017985987 = queryNorm
              0.7573833 = fieldWeight in 3321, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                6.996407 = idf(docFreq=109, maxDocs=44218)
                0.0625 = fieldNorm(doc=3321)
          0.19829968 = weight(abstract_txt:semantic in 3321) [ClassicSimilarity], result of:
            0.19829968 = score(doc=3321,freq=6.0), product of:
              0.28949374 = queryWeight, product of:
                3.5973089 = boost
                4.4743214 = idf(docFreq=1369, maxDocs=44218)
                0.017985987 = queryNorm
              0.6849878 = fieldWeight in 3321, product of:
                2.4494898 = tf(freq=6.0), with freq of:
                  6.0 = termFreq=6.0
                4.4743214 = idf(docFreq=1369, maxDocs=44218)
                0.0625 = fieldNorm(doc=3321)
        0.28 = coord(7/25)
    
  3. Choi, Y.: ¬A complete assessment of tagging quality : a consolidated methodology (2015) 0.19
    0.19316386 = sum of:
      0.19316386 = product of:
        0.6898709 = sum of:
          0.029786356 = weight(abstract_txt:document in 1730) [ClassicSimilarity], result of:
            0.029786356 = score(doc=1730,freq=1.0), product of:
              0.08881904 = queryWeight, product of:
                1.1504045 = boost
                4.2926083 = idf(docFreq=1642, maxDocs=44218)
                0.017985987 = queryNorm
              0.33536002 = fieldWeight in 1730, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.2926083 = idf(docFreq=1642, maxDocs=44218)
                0.078125 = fieldNorm(doc=1730)
          0.06929195 = weight(abstract_txt:indexing in 1730) [ClassicSimilarity], result of:
            0.06929195 = score(doc=1730,freq=5.0), product of:
              0.09119262 = queryWeight, product of:
                1.1656747 = boost
                4.3495874 = idf(docFreq=1551, maxDocs=44218)
                0.017985987 = queryNorm
              0.7598417 = fieldWeight in 1730, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                4.3495874 = idf(docFreq=1551, maxDocs=44218)
                0.078125 = fieldNorm(doc=1730)
          0.06424806 = weight(abstract_txt:statistical in 1730) [ClassicSimilarity], result of:
            0.06424806 = score(doc=1730,freq=1.0), product of:
              0.14827497 = queryWeight, product of:
                1.4863855 = boost
                5.5462847 = idf(docFreq=468, maxDocs=44218)
                0.017985987 = queryNorm
              0.43330348 = fieldWeight in 1730, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.5462847 = idf(docFreq=468, maxDocs=44218)
                0.078125 = fieldNorm(doc=1730)
          0.035779413 = weight(abstract_txt:model in 1730) [ClassicSimilarity], result of:
            0.035779413 = score(doc=1730,freq=1.0), product of:
              0.11488952 = queryWeight, product of:
                1.6024458 = boost
                3.986234 = idf(docFreq=2231, maxDocs=44218)
                0.017985987 = queryNorm
              0.31142452 = fieldWeight in 1730, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.986234 = idf(docFreq=2231, maxDocs=44218)
                0.078125 = fieldNorm(doc=1730)
          0.094925776 = weight(abstract_txt:significance in 1730) [ClassicSimilarity], result of:
            0.094925776 = score(doc=1730,freq=1.0), product of:
              0.19234635 = queryWeight, product of:
                1.6929319 = boost
                6.31699 = idf(docFreq=216, maxDocs=44218)
                0.017985987 = queryNorm
              0.49351484 = fieldWeight in 1730, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.31699 = idf(docFreq=216, maxDocs=44218)
                0.078125 = fieldNorm(doc=1730)
          0.19345059 = weight(abstract_txt:latent in 1730) [ClassicSimilarity], result of:
            0.19345059 = score(doc=1730,freq=1.0), product of:
              0.35391986 = queryWeight, product of:
                2.81252 = boost
                6.996407 = idf(docFreq=109, maxDocs=44218)
                0.017985987 = queryNorm
              0.5465943 = fieldWeight in 1730, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.996407 = idf(docFreq=109, maxDocs=44218)
                0.078125 = fieldNorm(doc=1730)
          0.20238875 = weight(abstract_txt:semantic in 1730) [ClassicSimilarity], result of:
            0.20238875 = score(doc=1730,freq=4.0), product of:
              0.28949374 = queryWeight, product of:
                3.5973089 = boost
                4.4743214 = idf(docFreq=1369, maxDocs=44218)
                0.017985987 = queryNorm
              0.6991127 = fieldWeight in 1730, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                4.4743214 = idf(docFreq=1369, maxDocs=44218)
                0.078125 = fieldNorm(doc=1730)
        0.28 = coord(7/25)
    
  4. He, X.; Cai, D.; Liu, H.; Ma, W.Y.: Locality preserving indexing for document representation (2004) 0.16
    0.16246513 = sum of:
      0.16246513 = product of:
        1.3538761 = sum of:
          0.1752963 = weight(abstract_txt:indexing in 4079) [ClassicSimilarity], result of:
            0.1752963 = score(doc=4079,freq=2.0), product of:
              0.09119262 = queryWeight, product of:
                1.1656747 = boost
                4.3495874 = idf(docFreq=1551, maxDocs=44218)
                0.017985987 = queryNorm
              1.9222642 = fieldWeight in 4079, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                4.3495874 = idf(docFreq=1551, maxDocs=44218)
                0.3125 = fieldNorm(doc=4079)
          0.77380234 = weight(abstract_txt:latent in 4079) [ClassicSimilarity], result of:
            0.77380234 = score(doc=4079,freq=1.0), product of:
              0.35391986 = queryWeight, product of:
                2.81252 = boost
                6.996407 = idf(docFreq=109, maxDocs=44218)
                0.017985987 = queryNorm
              2.1863773 = fieldWeight in 4079, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.996407 = idf(docFreq=109, maxDocs=44218)
                0.3125 = fieldNorm(doc=4079)
          0.4047775 = weight(abstract_txt:semantic in 4079) [ClassicSimilarity], result of:
            0.4047775 = score(doc=4079,freq=1.0), product of:
              0.28949374 = queryWeight, product of:
                3.5973089 = boost
                4.4743214 = idf(docFreq=1369, maxDocs=44218)
                0.017985987 = queryNorm
              1.3982254 = fieldWeight in 4079, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.4743214 = idf(docFreq=1369, maxDocs=44218)
                0.3125 = fieldNorm(doc=4079)
        0.12 = coord(3/25)
    
  5. Story, R.E.: ¬An explanation of the effectiveness of latent semantic indexing by means of a Baysian regression model (1996) 0.16
    0.16041657 = sum of:
      0.16041657 = product of:
        0.66840243 = sum of:
          0.0417009 = weight(abstract_txt:document in 1943) [ClassicSimilarity], result of:
            0.0417009 = score(doc=1943,freq=1.0), product of:
              0.08881904 = queryWeight, product of:
                1.1504045 = boost
                4.2926083 = idf(docFreq=1642, maxDocs=44218)
                0.017985987 = queryNorm
              0.46950403 = fieldWeight in 1943, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.2926083 = idf(docFreq=1642, maxDocs=44218)
                0.109375 = fieldNorm(doc=1943)
          0.043383624 = weight(abstract_txt:indexing in 1943) [ClassicSimilarity], result of:
            0.043383624 = score(doc=1943,freq=1.0), product of:
              0.09119262 = queryWeight, product of:
                1.1656747 = boost
                4.3495874 = idf(docFreq=1551, maxDocs=44218)
                0.017985987 = queryNorm
              0.47573614 = fieldWeight in 1943, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.3495874 = idf(docFreq=1551, maxDocs=44218)
                0.109375 = fieldNorm(doc=1943)
          0.080867685 = weight(abstract_txt:words in 1943) [ClassicSimilarity], result of:
            0.080867685 = score(doc=1943,freq=1.0), product of:
              0.13812083 = queryWeight, product of:
                1.4345877 = boost
                5.353007 = idf(docFreq=568, maxDocs=44218)
                0.017985987 = queryNorm
              0.5854851 = fieldWeight in 1943, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.353007 = idf(docFreq=568, maxDocs=44218)
                0.109375 = fieldNorm(doc=1943)
          0.08994729 = weight(abstract_txt:statistical in 1943) [ClassicSimilarity], result of:
            0.08994729 = score(doc=1943,freq=1.0), product of:
              0.14827497 = queryWeight, product of:
                1.4863855 = boost
                5.5462847 = idf(docFreq=468, maxDocs=44218)
                0.017985987 = queryNorm
              0.6066249 = fieldWeight in 1943, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.5462847 = idf(docFreq=468, maxDocs=44218)
                0.109375 = fieldNorm(doc=1943)
          0.2708308 = weight(abstract_txt:latent in 1943) [ClassicSimilarity], result of:
            0.2708308 = score(doc=1943,freq=1.0), product of:
              0.35391986 = queryWeight, product of:
                2.81252 = boost
                6.996407 = idf(docFreq=109, maxDocs=44218)
                0.017985987 = queryNorm
              0.765232 = fieldWeight in 1943, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.996407 = idf(docFreq=109, maxDocs=44218)
                0.109375 = fieldNorm(doc=1943)
          0.14167213 = weight(abstract_txt:semantic in 1943) [ClassicSimilarity], result of:
            0.14167213 = score(doc=1943,freq=1.0), product of:
              0.28949374 = queryWeight, product of:
                3.5973089 = boost
                4.4743214 = idf(docFreq=1369, maxDocs=44218)
                0.017985987 = queryNorm
              0.4893789 = fieldWeight in 1943, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.4743214 = idf(docFreq=1369, maxDocs=44218)
                0.109375 = fieldNorm(doc=1943)
        0.24 = coord(6/25)