Document (#37958)

Author
Li, C.
Sun, A.
Datta, A.
Title
TSDW: Two-stage word sense disambiguation using Wikipedia
Source
Journal of the American Society for Information Science and Technology. 64(2013) no.6, S.1203-1223
Year
2013
Abstract
The semantic knowledge of Wikipedia has proved to be useful for many tasks, for example, named entity disambiguation. Among these applications, the task of identifying the word sense based on Wikipedia is a crucial component because the output of this component is often used in subsequent tasks. In this article, we present a two-stage framework (called TSDW) for word sense disambiguation using knowledge latent in Wikipedia. The disambiguation of a given phrase is applied through a two-stage disambiguation process: (a) The first-stage disambiguation explores the contextual semantic information, where the noisy information is pruned for better effectiveness and efficiency; and (b) the second-stage disambiguation explores the disambiguated phrases of high confidence from the first stage to achieve better redisambiguation decisions for the phrases that are difficult to disambiguate in the first stage. Moreover, existing studies have addressed the disambiguation problem for English text only. Considering the popular usage of Wikipedia in different languages, we study the performance of TSDW and the existing state-of-the-art approaches over both English and Traditional Chinese articles. The experimental results show that TSDW generalizes well to different semantic relatedness measures and text in different languages. More important, TSDW significantly outperforms the state-of-the-art approaches with both better effectiveness and efficiency.
Object
Wikipedia

Similar documents (author)

  1. Datta, S.; Farradane, J.E.L.: ¬A psychological basis for general classification (1974) 4.73
    4.72773 = sum of:
      4.72773 = weight(author_txt:datta in 1196) [ClassicSimilarity], result of:
        4.72773 = fieldWeight in 1196, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          9.45546 = idf(docFreq=8, maxDocs=42306)
          0.5 = fieldNorm(doc=1196)
    
  2. Meso, P.; Datta, P.; Mbarika, V.: Moderating information and communication technologies' influences on socioeconomic development with good governance : a study of the developing countries (2006) 3.55
    3.5457973 = sum of:
      3.5457973 = weight(author_txt:datta in 920) [ClassicSimilarity], result of:
        3.5457973 = fieldWeight in 920, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          9.45546 = idf(docFreq=8, maxDocs=42306)
          0.375 = fieldNorm(doc=920)
    
  3. Kifle, M.; Mbarika, V.W.A.; Datta, P.: Telemedicine in sub-Saharan Africa : the case of teleophthalmology and eye care in Ethiopia (2006) 3.55
    3.5457973 = sum of:
      3.5457973 = weight(author_txt:datta in 911) [ClassicSimilarity], result of:
        3.5457973 = fieldWeight in 911, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          9.45546 = idf(docFreq=8, maxDocs=42306)
          0.375 = fieldNorm(doc=911)
    
  4. Pal, D.; Mitra, M.; Datta, K.: Improving query expansion using WordNet (2014) 3.55
    3.5457973 = sum of:
      3.5457973 = weight(author_txt:datta in 3546) [ClassicSimilarity], result of:
        3.5457973 = fieldWeight in 3546, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          9.45546 = idf(docFreq=8, maxDocs=42306)
          0.375 = fieldNorm(doc=3546)
    
  5. Datta, A.; Yong, J.T.T.; Braghin, S.: ¬The zen of multidisciplinary team recommendation (2014) 3.55
    3.5457973 = sum of:
      3.5457973 = weight(author_txt:datta in 3551) [ClassicSimilarity], result of:
        3.5457973 = fieldWeight in 3551, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          9.45546 = idf(docFreq=8, maxDocs=42306)
          0.375 = fieldNorm(doc=3551)
    

Similar documents (content)

  1. Zhao, G.; Wu, J.; Wang, D.; Li, T.: Entity disambiguation to Wikipedia using collective ranking (2016) 0.23
    0.230208 = sum of:
      0.230208 = product of:
        0.9592 = sum of:
          0.02552475 = weight(abstract_txt:existing in 185) [ClassicSimilarity], result of:
            0.02552475 = score(doc=185,freq=1.0), product of:
              0.06937308 = queryWeight, product of:
                1.132041 = boost
                4.709562 = idf(docFreq=1035, maxDocs=42306)
                0.013012127 = queryNorm
              0.36793453 = fieldWeight in 185, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.709562 = idf(docFreq=1035, maxDocs=42306)
                0.078125 = fieldNorm(doc=185)
          0.032803398 = weight(abstract_txt:effectiveness in 185) [ClassicSimilarity], result of:
            0.032803398 = score(doc=185,freq=1.0), product of:
              0.08200289 = queryWeight, product of:
                1.2307824 = boost
                5.12035 = idf(docFreq=686, maxDocs=42306)
                0.013012127 = queryNorm
              0.40002733 = fieldWeight in 185, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.12035 = idf(docFreq=686, maxDocs=42306)
                0.078125 = fieldNorm(doc=185)
          0.027503021 = weight(abstract_txt:first in 185) [ClassicSimilarity], result of:
            0.027503021 = score(doc=185,freq=1.0), product of:
              0.08346428 = queryWeight, product of:
                1.5207669 = boost
                4.2178364 = idf(docFreq=1693, maxDocs=42306)
                0.013012127 = queryNorm
              0.32951847 = fieldWeight in 185, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.2178364 = idf(docFreq=1693, maxDocs=42306)
                0.078125 = fieldNorm(doc=185)
          0.033374585 = weight(abstract_txt:semantic in 185) [ClassicSimilarity], result of:
            0.033374585 = score(doc=185,freq=1.0), product of:
              0.09495641 = queryWeight, product of:
                1.6220882 = boost
                4.4988503 = idf(docFreq=1278, maxDocs=42306)
                0.013012127 = queryNorm
              0.35147268 = fieldWeight in 185, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.4988503 = idf(docFreq=1278, maxDocs=42306)
                0.078125 = fieldNorm(doc=185)
          0.15281026 = weight(abstract_txt:wikipedia in 185) [ClassicSimilarity], result of:
            0.15281026 = score(doc=185,freq=1.0), product of:
              0.3104309 = queryWeight, product of:
                3.7863362 = boost
                6.300826 = idf(docFreq=210, maxDocs=42306)
                0.013012127 = queryNorm
              0.49225205 = fieldWeight in 185, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.300826 = idf(docFreq=210, maxDocs=42306)
                0.078125 = fieldNorm(doc=185)
          0.68718404 = weight(abstract_txt:disambiguation in 185) [ClassicSimilarity], result of:
            0.68718404 = score(doc=185,freq=3.0), product of:
              0.68587494 = queryWeight, product of:
                7.1190023 = boost
                7.404189 = idf(docFreq=69, maxDocs=42306)
                0.013012127 = queryNorm
              1.0019087 = fieldWeight in 185, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                7.404189 = idf(docFreq=69, maxDocs=42306)
                0.078125 = fieldNorm(doc=185)
        0.24 = coord(6/25)
    
  2. Green, R.: WordNet (2009) 0.20
    0.20121898 = sum of:
      0.20121898 = product of:
        0.8384124 = sum of:
          0.041817214 = weight(abstract_txt:tasks in 1697) [ClassicSimilarity], result of:
            0.041817214 = score(doc=1697,freq=1.0), product of:
              0.08537535 = queryWeight, product of:
                1.255836 = boost
                5.224579 = idf(docFreq=618, maxDocs=42306)
                0.013012127 = queryNorm
              0.48980427 = fieldWeight in 1697, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.224579 = idf(docFreq=618, maxDocs=42306)
                0.09375 = fieldNorm(doc=1697)
          0.073404424 = weight(abstract_txt:english in 1697) [ClassicSimilarity], result of:
            0.073404424 = score(doc=1697,freq=2.0), product of:
              0.09860538 = queryWeight, product of:
                1.349637 = boost
                5.6148133 = idf(docFreq=418, maxDocs=42306)
                0.013012127 = queryNorm
              0.74442613 = fieldWeight in 1697, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                5.6148133 = idf(docFreq=418, maxDocs=42306)
                0.09375 = fieldNorm(doc=1697)
          0.05663855 = weight(abstract_txt:semantic in 1697) [ClassicSimilarity], result of:
            0.05663855 = score(doc=1697,freq=2.0), product of:
              0.09495641 = queryWeight, product of:
                1.6220882 = boost
                4.4988503 = idf(docFreq=1278, maxDocs=42306)
                0.013012127 = queryNorm
              0.5964689 = fieldWeight in 1697, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                4.4988503 = idf(docFreq=1278, maxDocs=42306)
                0.09375 = fieldNorm(doc=1697)
          0.10126567 = weight(abstract_txt:word in 1697) [ClassicSimilarity], result of:
            0.10126567 = score(doc=1697,freq=2.0), product of:
              0.13988067 = queryWeight, product of:
                1.9687527 = boost
                5.460322 = idf(docFreq=488, maxDocs=42306)
                0.013012127 = queryNorm
              0.72394323 = fieldWeight in 1697, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                5.460322 = idf(docFreq=488, maxDocs=42306)
                0.09375 = fieldNorm(doc=1697)
          0.08919143 = weight(abstract_txt:sense in 1697) [ClassicSimilarity], result of:
            0.08919143 = score(doc=1697,freq=1.0), product of:
              0.16193533 = queryWeight, product of:
                2.1182787 = boost
                5.875032 = idf(docFreq=322, maxDocs=42306)
                0.013012127 = queryNorm
              0.55078423 = fieldWeight in 1697, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.875032 = idf(docFreq=322, maxDocs=42306)
                0.09375 = fieldNorm(doc=1697)
          0.47609508 = weight(abstract_txt:disambiguation in 1697) [ClassicSimilarity], result of:
            0.47609508 = score(doc=1697,freq=1.0), product of:
              0.68587494 = queryWeight, product of:
                7.1190023 = boost
                7.404189 = idf(docFreq=69, maxDocs=42306)
                0.013012127 = queryNorm
              0.6941427 = fieldWeight in 1697, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.404189 = idf(docFreq=69, maxDocs=42306)
                0.09375 = fieldNorm(doc=1697)
        0.24 = coord(6/25)
    
  3. Ng, H.T.; Zelle, J.: Corpus-based approaches to semantic interpretation in natural language processing (1997) 0.19
    0.18833598 = sum of:
      0.18833598 = product of:
        1.1771 = sum of:
          0.115612954 = weight(abstract_txt:semantic in 4253) [ClassicSimilarity], result of:
            0.115612954 = score(doc=4253,freq=3.0), product of:
              0.09495641 = queryWeight, product of:
                1.6220882 = boost
                4.4988503 = idf(docFreq=1278, maxDocs=42306)
                0.013012127 = queryNorm
              1.217537 = fieldWeight in 4253, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                4.4988503 = idf(docFreq=1278, maxDocs=42306)
                0.15625 = fieldNorm(doc=4253)
          0.11934273 = weight(abstract_txt:word in 4253) [ClassicSimilarity], result of:
            0.11934273 = score(doc=4253,freq=1.0), product of:
              0.13988067 = queryWeight, product of:
                1.9687527 = boost
                5.460322 = idf(docFreq=488, maxDocs=42306)
                0.013012127 = queryNorm
              0.8531753 = fieldWeight in 4253, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.460322 = idf(docFreq=488, maxDocs=42306)
                0.15625 = fieldNorm(doc=4253)
          0.14865239 = weight(abstract_txt:sense in 4253) [ClassicSimilarity], result of:
            0.14865239 = score(doc=4253,freq=1.0), product of:
              0.16193533 = queryWeight, product of:
                2.1182787 = boost
                5.875032 = idf(docFreq=322, maxDocs=42306)
                0.013012127 = queryNorm
              0.91797376 = fieldWeight in 4253, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.875032 = idf(docFreq=322, maxDocs=42306)
                0.15625 = fieldNorm(doc=4253)
          0.79349184 = weight(abstract_txt:disambiguation in 4253) [ClassicSimilarity], result of:
            0.79349184 = score(doc=4253,freq=1.0), product of:
              0.68587494 = queryWeight, product of:
                7.1190023 = boost
                7.404189 = idf(docFreq=69, maxDocs=42306)
                0.013012127 = queryNorm
              1.1569046 = fieldWeight in 4253, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.404189 = idf(docFreq=69, maxDocs=42306)
                0.15625 = fieldNorm(doc=4253)
        0.16 = coord(4/25)
    
  4. Vlachidis, A.; Tudhope, D.: ¬A knowledge-based approach to information extraction for semantic interoperability in the archaeology domain (2016) 0.18
    0.18151619 = sum of:
      0.18151619 = product of:
        0.75631744 = sum of:
          0.027878141 = weight(abstract_txt:tasks in 4896) [ClassicSimilarity], result of:
            0.027878141 = score(doc=4896,freq=1.0), product of:
              0.08537535 = queryWeight, product of:
                1.255836 = boost
                5.224579 = idf(docFreq=618, maxDocs=42306)
                0.013012127 = queryNorm
              0.32653618 = fieldWeight in 4896, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.224579 = idf(docFreq=618, maxDocs=42306)
                0.0625 = fieldNorm(doc=4896)
          0.06540056 = weight(abstract_txt:semantic in 4896) [ClassicSimilarity], result of:
            0.06540056 = score(doc=4896,freq=6.0), product of:
              0.09495641 = queryWeight, product of:
                1.6220882 = boost
                4.4988503 = idf(docFreq=1278, maxDocs=42306)
                0.013012127 = queryNorm
              0.688743 = fieldWeight in 4896, product of:
                2.4494898 = tf(freq=6.0), with freq of:
                  6.0 = termFreq=6.0
                4.4988503 = idf(docFreq=1278, maxDocs=42306)
                0.0625 = fieldNorm(doc=4896)
          0.062571056 = weight(abstract_txt:phrases in 4896) [ClassicSimilarity], result of:
            0.062571056 = score(doc=4896,freq=1.0), product of:
              0.14635435 = queryWeight, product of:
                1.6442562 = boost
                6.8405 = idf(docFreq=122, maxDocs=42306)
                0.013012127 = queryNorm
              0.42753124 = fieldWeight in 4896, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.8405 = idf(docFreq=122, maxDocs=42306)
                0.0625 = fieldNorm(doc=4896)
          0.06751044 = weight(abstract_txt:word in 4896) [ClassicSimilarity], result of:
            0.06751044 = score(doc=4896,freq=2.0), product of:
              0.13988067 = queryWeight, product of:
                1.9687527 = boost
                5.460322 = idf(docFreq=488, maxDocs=42306)
                0.013012127 = queryNorm
              0.48262882 = fieldWeight in 4896, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                5.460322 = idf(docFreq=488, maxDocs=42306)
                0.0625 = fieldNorm(doc=4896)
          0.084090486 = weight(abstract_txt:sense in 4896) [ClassicSimilarity], result of:
            0.084090486 = score(doc=4896,freq=2.0), product of:
              0.16193533 = queryWeight, product of:
                2.1182787 = boost
                5.875032 = idf(docFreq=322, maxDocs=42306)
                0.013012127 = queryNorm
              0.51928437 = fieldWeight in 4896, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                5.875032 = idf(docFreq=322, maxDocs=42306)
                0.0625 = fieldNorm(doc=4896)
          0.44886675 = weight(abstract_txt:disambiguation in 4896) [ClassicSimilarity], result of:
            0.44886675 = score(doc=4896,freq=2.0), product of:
              0.68587494 = queryWeight, product of:
                7.1190023 = boost
                7.404189 = idf(docFreq=69, maxDocs=42306)
                0.013012127 = queryNorm
              0.65444404 = fieldWeight in 4896, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                7.404189 = idf(docFreq=69, maxDocs=42306)
                0.0625 = fieldNorm(doc=4896)
        0.24 = coord(6/25)
    
  5. Brychcín, T.; Konopík, M.: HPS: High precision stemmer (2015) 0.16
    0.1614494 = sum of:
      0.1614494 = product of:
        0.5045294 = sum of:
          0.031985313 = weight(abstract_txt:state in 4687) [ClassicSimilarity], result of:
            0.031985313 = score(doc=4687,freq=2.0), product of:
              0.07426435 = queryWeight, product of:
                1.1712695 = boost
                4.872762 = idf(docFreq=879, maxDocs=42306)
                0.013012127 = queryNorm
              0.43069538 = fieldWeight in 4687, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                4.872762 = idf(docFreq=879, maxDocs=42306)
                0.0625 = fieldNorm(doc=4687)
          0.027322812 = weight(abstract_txt:languages in 4687) [ClassicSimilarity], result of:
            0.027322812 = score(doc=4687,freq=1.0), product of:
              0.08423778 = queryWeight, product of:
                1.2474413 = boost
                5.189655 = idf(docFreq=640, maxDocs=42306)
                0.013012127 = queryNorm
              0.32435343 = fieldWeight in 4687, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.189655 = idf(docFreq=640, maxDocs=42306)
                0.0625 = fieldNorm(doc=4687)
          0.039425645 = weight(abstract_txt:tasks in 4687) [ClassicSimilarity], result of:
            0.039425645 = score(doc=4687,freq=2.0), product of:
              0.08537535 = queryWeight, product of:
                1.255836 = boost
                5.224579 = idf(docFreq=618, maxDocs=42306)
                0.013012127 = queryNorm
              0.46179187 = fieldWeight in 4687, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                5.224579 = idf(docFreq=618, maxDocs=42306)
                0.0625 = fieldNorm(doc=4687)
          0.034603175 = weight(abstract_txt:english in 4687) [ClassicSimilarity], result of:
            0.034603175 = score(doc=4687,freq=1.0), product of:
              0.09860538 = queryWeight, product of:
                1.349637 = boost
                5.6148133 = idf(docFreq=418, maxDocs=42306)
                0.013012127 = queryNorm
              0.35092583 = fieldWeight in 4687, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.6148133 = idf(docFreq=418, maxDocs=42306)
                0.0625 = fieldNorm(doc=4687)
          0.022002418 = weight(abstract_txt:first in 4687) [ClassicSimilarity], result of:
            0.022002418 = score(doc=4687,freq=1.0), product of:
              0.08346428 = queryWeight, product of:
                1.5207669 = boost
                4.2178364 = idf(docFreq=1693, maxDocs=42306)
                0.013012127 = queryNorm
              0.26361477 = fieldWeight in 4687, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.2178364 = idf(docFreq=1693, maxDocs=42306)
                0.0625 = fieldNorm(doc=4687)
          0.026699668 = weight(abstract_txt:semantic in 4687) [ClassicSimilarity], result of:
            0.026699668 = score(doc=4687,freq=1.0), product of:
              0.09495641 = queryWeight, product of:
                1.6220882 = boost
                4.4988503 = idf(docFreq=1278, maxDocs=42306)
                0.013012127 = queryNorm
              0.28117815 = fieldWeight in 4687, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.4988503 = idf(docFreq=1278, maxDocs=42306)
                0.0625 = fieldNorm(doc=4687)
          0.04773709 = weight(abstract_txt:word in 4687) [ClassicSimilarity], result of:
            0.04773709 = score(doc=4687,freq=1.0), product of:
              0.13988067 = queryWeight, product of:
                1.9687527 = boost
                5.460322 = idf(docFreq=488, maxDocs=42306)
                0.013012127 = queryNorm
              0.34127012 = fieldWeight in 4687, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.460322 = idf(docFreq=488, maxDocs=42306)
                0.0625 = fieldNorm(doc=4687)
          0.2747533 = weight(abstract_txt:stage in 4687) [ClassicSimilarity], result of:
            0.2747533 = score(doc=4687,freq=3.0), product of:
              0.41314343 = queryWeight, product of:
                5.1683407 = boost
                6.143296 = idf(docFreq=246, maxDocs=42306)
                0.013012127 = queryNorm
              0.66503125 = fieldWeight in 4687, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                6.143296 = idf(docFreq=246, maxDocs=42306)
                0.0625 = fieldNorm(doc=4687)
        0.32 = coord(8/25)