Document (#37957)

Author
Li, C.
Sun, A.
Datta, A.
Title
TSDW: Two-stage word sense disambiguation using Wikipedia
Source
Journal of the American Society for Information Science and Technology. 64(2013) no.6, S.1203-1223
Year
2013
Abstract
The semantic knowledge of Wikipedia has proved to be useful for many tasks, for example, named entity disambiguation. Among these applications, the task of identifying the word sense based on Wikipedia is a crucial component because the output of this component is often used in subsequent tasks. In this article, we present a two-stage framework (called TSDW) for word sense disambiguation using knowledge latent in Wikipedia. The disambiguation of a given phrase is applied through a two-stage disambiguation process: (a) The first-stage disambiguation explores the contextual semantic information, where the noisy information is pruned for better effectiveness and efficiency; and (b) the second-stage disambiguation explores the disambiguated phrases of high confidence from the first stage to achieve better redisambiguation decisions for the phrases that are difficult to disambiguate in the first stage. Moreover, existing studies have addressed the disambiguation problem for English text only. Considering the popular usage of Wikipedia in different languages, we study the performance of TSDW and the existing state-of-the-art approaches over both English and Traditional Chinese articles. The experimental results show that TSDW generalizes well to different semantic relatedness measures and text in different languages. More important, TSDW significantly outperforms the state-of-the-art approaches with both better effectiveness and efficiency.
Object
Wikipedia

Similar documents (author)

  1. Datta, S.; Farradane, J.E.L.: ¬A psychological basis for general classification (1974) 4.75
    4.749831 = sum of:
      4.749831 = weight(author_txt:datta in 1196) [ClassicSimilarity], result of:
        4.749831 = fieldWeight in 1196, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          9.499662 = idf(docFreq=8, maxDocs=44218)
          0.5 = fieldNorm(doc=1196)
    
  2. Meso, P.; Datta, P.; Mbarika, V.: Moderating information and communication technologies' influences on socioeconomic development with good governance : a study of the developing countries (2006) 3.56
    3.5623734 = sum of:
      3.5623734 = weight(author_txt:datta in 4919) [ClassicSimilarity], result of:
        3.5623734 = fieldWeight in 4919, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          9.499662 = idf(docFreq=8, maxDocs=44218)
          0.375 = fieldNorm(doc=4919)
    
  3. Kifle, M.; Mbarika, V.W.A.; Datta, P.: Telemedicine in sub-Saharan Africa : the case of teleophthalmology and eye care in Ethiopia (2006) 3.56
    3.5623734 = sum of:
      3.5623734 = weight(author_txt:datta in 5910) [ClassicSimilarity], result of:
        3.5623734 = fieldWeight in 5910, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          9.499662 = idf(docFreq=8, maxDocs=44218)
          0.375 = fieldNorm(doc=5910)
    
  4. Pal, D.; Mitra, M.; Datta, K.: Improving query expansion using WordNet (2014) 3.56
    3.5623734 = sum of:
      3.5623734 = weight(author_txt:datta in 1545) [ClassicSimilarity], result of:
        3.5623734 = fieldWeight in 1545, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          9.499662 = idf(docFreq=8, maxDocs=44218)
          0.375 = fieldNorm(doc=1545)
    
  5. Datta, A.; Yong, J.T.T.; Braghin, S.: ¬The zen of multidisciplinary team recommendation (2014) 3.56
    3.5623734 = sum of:
      3.5623734 = weight(author_txt:datta in 1550) [ClassicSimilarity], result of:
        3.5623734 = fieldWeight in 1550, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          9.499662 = idf(docFreq=8, maxDocs=44218)
          0.375 = fieldNorm(doc=1550)
    

Similar documents (content)

  1. Zhao, G.; Wu, J.; Wang, D.; Li, T.: Entity disambiguation to Wikipedia using collective ranking (2016) 0.23
    0.22809178 = sum of:
      0.22809178 = product of:
        0.9503824 = sum of:
          0.024957046 = weight(abstract_txt:existing in 3266) [ClassicSimilarity], result of:
            0.024957046 = score(doc=3266,freq=1.0), product of:
              0.06868256 = queryWeight, product of:
                1.1343647 = boost
                4.6511106 = idf(docFreq=1147, maxDocs=44218)
                0.013017785 = queryNorm
              0.36336803 = fieldWeight in 3266, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.6511106 = idf(docFreq=1147, maxDocs=44218)
                0.078125 = fieldNorm(doc=3266)
          0.03287148 = weight(abstract_txt:effectiveness in 3266) [ClassicSimilarity], result of:
            0.03287148 = score(doc=3266,freq=1.0), product of:
              0.08252722 = queryWeight, product of:
                1.2434493 = boost
                5.098378 = idf(docFreq=733, maxDocs=44218)
                0.013017785 = queryNorm
              0.39831078 = fieldWeight in 3266, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.098378 = idf(docFreq=733, maxDocs=44218)
                0.078125 = fieldNorm(doc=3266)
          0.026940344 = weight(abstract_txt:first in 3266) [ClassicSimilarity], result of:
            0.026940344 = score(doc=3266,freq=1.0), product of:
              0.08273391 = queryWeight, product of:
                1.524814 = boost
                4.168018 = idf(docFreq=1860, maxDocs=44218)
                0.013017785 = queryNorm
              0.3256264 = fieldWeight in 3266, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.168018 = idf(docFreq=1860, maxDocs=44218)
                0.078125 = fieldNorm(doc=3266)
          0.03332698 = weight(abstract_txt:semantic in 3266) [ClassicSimilarity], result of:
            0.03332698 = score(doc=3266,freq=1.0), product of:
              0.09534079 = queryWeight, product of:
                1.6368711 = boost
                4.4743214 = idf(docFreq=1369, maxDocs=44218)
                0.013017785 = queryNorm
              0.34955636 = fieldWeight in 3266, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.4743214 = idf(docFreq=1369, maxDocs=44218)
                0.078125 = fieldNorm(doc=3266)
          0.1526704 = weight(abstract_txt:wikipedia in 3266) [ClassicSimilarity], result of:
            0.1526704 = score(doc=3266,freq=1.0), product of:
              0.3117939 = queryWeight, product of:
                3.8214948 = boost
                6.2675414 = idf(docFreq=227, maxDocs=44218)
                0.013017785 = queryNorm
              0.48965168 = fieldWeight in 3266, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.2675414 = idf(docFreq=227, maxDocs=44218)
                0.078125 = fieldNorm(doc=3266)
          0.67961615 = weight(abstract_txt:disambiguation in 3266) [ClassicSimilarity], result of:
            0.67961615 = score(doc=3266,freq=3.0), product of:
              0.68423676 = queryWeight, product of:
                7.1608186 = boost
                7.3401785 = idf(docFreq=77, maxDocs=44218)
                0.013017785 = queryNorm
              0.99324703 = fieldWeight in 3266, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                7.3401785 = idf(docFreq=77, maxDocs=44218)
                0.078125 = fieldNorm(doc=3266)
        0.24 = coord(6/25)
    
  2. Green, R.: WordNet (2009) 0.20
    0.19945067 = sum of:
      0.19945067 = product of:
        0.8310445 = sum of:
          0.04096714 = weight(abstract_txt:tasks in 4696) [ClassicSimilarity], result of:
            0.04096714 = score(doc=4696,freq=1.0), product of:
              0.08463577 = queryWeight, product of:
                1.2592341 = boost
                5.1630983 = idf(docFreq=687, maxDocs=44218)
                0.013017785 = queryNorm
              0.48404047 = fieldWeight in 4696, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.1630983 = idf(docFreq=687, maxDocs=44218)
                0.09375 = fieldNorm(doc=4696)
          0.07291426 = weight(abstract_txt:english in 4696) [ClassicSimilarity], result of:
            0.07291426 = score(doc=4696,freq=2.0), product of:
              0.09865714 = queryWeight, product of:
                1.3595455 = boost
                5.574394 = idf(docFreq=455, maxDocs=44218)
                0.013017785 = queryNorm
              0.7390672 = fieldWeight in 4696, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                5.574394 = idf(docFreq=455, maxDocs=44218)
                0.09375 = fieldNorm(doc=4696)
          0.056557756 = weight(abstract_txt:semantic in 4696) [ClassicSimilarity], result of:
            0.056557756 = score(doc=4696,freq=2.0), product of:
              0.09534079 = queryWeight, product of:
                1.6368711 = boost
                4.4743214 = idf(docFreq=1369, maxDocs=44218)
                0.013017785 = queryNorm
              0.5932168 = fieldWeight in 4696, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                4.4743214 = idf(docFreq=1369, maxDocs=44218)
                0.09375 = fieldNorm(doc=4696)
          0.1013921 = weight(abstract_txt:word in 4696) [ClassicSimilarity], result of:
            0.1013921 = score(doc=4696,freq=2.0), product of:
              0.1406976 = queryWeight, product of:
                1.9884673 = boost
                5.4353957 = idf(docFreq=523, maxDocs=44218)
                0.013017785 = queryNorm
              0.72063845 = fieldWeight in 4696, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                5.4353957 = idf(docFreq=523, maxDocs=44218)
                0.09375 = fieldNorm(doc=4696)
          0.08836142 = weight(abstract_txt:sense in 4696) [ClassicSimilarity], result of:
            0.08836142 = score(doc=4696,freq=1.0), product of:
              0.1617344 = queryWeight, product of:
                2.1319466 = boost
                5.8275905 = idf(docFreq=353, maxDocs=44218)
                0.013017785 = queryNorm
              0.5463366 = fieldWeight in 4696, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.8275905 = idf(docFreq=353, maxDocs=44218)
                0.09375 = fieldNorm(doc=4696)
          0.47085184 = weight(abstract_txt:disambiguation in 4696) [ClassicSimilarity], result of:
            0.47085184 = score(doc=4696,freq=1.0), product of:
              0.68423676 = queryWeight, product of:
                7.1608186 = boost
                7.3401785 = idf(docFreq=77, maxDocs=44218)
                0.013017785 = queryNorm
              0.6881417 = fieldWeight in 4696, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.3401785 = idf(docFreq=77, maxDocs=44218)
                0.09375 = fieldNorm(doc=4696)
        0.24 = coord(6/25)
    
  3. Ng, H.T.; Zelle, J.: Corpus-based approaches to semantic interpretation in natural language processing (1997) 0.19
    0.1867139 = sum of:
      0.1867139 = product of:
        1.1669619 = sum of:
          0.11544803 = weight(abstract_txt:semantic in 3252) [ClassicSimilarity], result of:
            0.11544803 = score(doc=3252,freq=3.0), product of:
              0.09534079 = queryWeight, product of:
                1.6368711 = boost
                4.4743214 = idf(docFreq=1369, maxDocs=44218)
                0.013017785 = queryNorm
              1.2108986 = fieldWeight in 3252, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                4.4743214 = idf(docFreq=1369, maxDocs=44218)
                0.15625 = fieldNorm(doc=3252)
          0.11949174 = weight(abstract_txt:word in 3252) [ClassicSimilarity], result of:
            0.11949174 = score(doc=3252,freq=1.0), product of:
              0.1406976 = queryWeight, product of:
                1.9884673 = boost
                5.4353957 = idf(docFreq=523, maxDocs=44218)
                0.013017785 = queryNorm
              0.8492806 = fieldWeight in 3252, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.4353957 = idf(docFreq=523, maxDocs=44218)
                0.15625 = fieldNorm(doc=3252)
          0.14726904 = weight(abstract_txt:sense in 3252) [ClassicSimilarity], result of:
            0.14726904 = score(doc=3252,freq=1.0), product of:
              0.1617344 = queryWeight, product of:
                2.1319466 = boost
                5.8275905 = idf(docFreq=353, maxDocs=44218)
                0.013017785 = queryNorm
              0.910561 = fieldWeight in 3252, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.8275905 = idf(docFreq=353, maxDocs=44218)
                0.15625 = fieldNorm(doc=3252)
          0.78475314 = weight(abstract_txt:disambiguation in 3252) [ClassicSimilarity], result of:
            0.78475314 = score(doc=3252,freq=1.0), product of:
              0.68423676 = queryWeight, product of:
                7.1608186 = boost
                7.3401785 = idf(docFreq=77, maxDocs=44218)
                0.013017785 = queryNorm
              1.1469029 = fieldWeight in 3252, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.3401785 = idf(docFreq=77, maxDocs=44218)
                0.15625 = fieldNorm(doc=3252)
        0.16 = coord(4/25)
    
  4. Vlachidis, A.; Tudhope, D.: ¬A knowledge-based approach to information extraction for semantic interoperability in the archaeology domain (2016) 0.18
    0.18041882 = sum of:
      0.18041882 = product of:
        0.7517451 = sum of:
          0.027311426 = weight(abstract_txt:tasks in 2895) [ClassicSimilarity], result of:
            0.027311426 = score(doc=2895,freq=1.0), product of:
              0.08463577 = queryWeight, product of:
                1.2592341 = boost
                5.1630983 = idf(docFreq=687, maxDocs=44218)
                0.013017785 = queryNorm
              0.32269365 = fieldWeight in 2895, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.1630983 = idf(docFreq=687, maxDocs=44218)
                0.0625 = fieldNorm(doc=2895)
          0.065307274 = weight(abstract_txt:semantic in 2895) [ClassicSimilarity], result of:
            0.065307274 = score(doc=2895,freq=6.0), product of:
              0.09534079 = queryWeight, product of:
                1.6368711 = boost
                4.4743214 = idf(docFreq=1369, maxDocs=44218)
                0.013017785 = queryNorm
              0.6849878 = fieldWeight in 2895, product of:
                2.4494898 = tf(freq=6.0), with freq of:
                  6.0 = termFreq=6.0
                4.4743214 = idf(docFreq=1369, maxDocs=44218)
                0.0625 = fieldNorm(doc=2895)
          0.06430028 = weight(abstract_txt:phrases in 2895) [ClassicSimilarity], result of:
            0.06430028 = score(doc=2895,freq=1.0), product of:
              0.1497843 = queryWeight, product of:
                1.6751844 = boost
                6.8685737 = idf(docFreq=124, maxDocs=44218)
                0.013017785 = queryNorm
              0.42928585 = fieldWeight in 2895, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.8685737 = idf(docFreq=124, maxDocs=44218)
                0.0625 = fieldNorm(doc=2895)
          0.06759473 = weight(abstract_txt:word in 2895) [ClassicSimilarity], result of:
            0.06759473 = score(doc=2895,freq=2.0), product of:
              0.1406976 = queryWeight, product of:
                1.9884673 = boost
                5.4353957 = idf(docFreq=523, maxDocs=44218)
                0.013017785 = queryNorm
              0.48042563 = fieldWeight in 2895, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                5.4353957 = idf(docFreq=523, maxDocs=44218)
                0.0625 = fieldNorm(doc=2895)
          0.083307944 = weight(abstract_txt:sense in 2895) [ClassicSimilarity], result of:
            0.083307944 = score(doc=2895,freq=2.0), product of:
              0.1617344 = queryWeight, product of:
                2.1319466 = boost
                5.8275905 = idf(docFreq=353, maxDocs=44218)
                0.013017785 = queryNorm
              0.51509106 = fieldWeight in 2895, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                5.8275905 = idf(docFreq=353, maxDocs=44218)
                0.0625 = fieldNorm(doc=2895)
          0.4439234 = weight(abstract_txt:disambiguation in 2895) [ClassicSimilarity], result of:
            0.4439234 = score(doc=2895,freq=2.0), product of:
              0.68423676 = queryWeight, product of:
                7.1608186 = boost
                7.3401785 = idf(docFreq=77, maxDocs=44218)
                0.013017785 = queryNorm
              0.64878625 = fieldWeight in 2895, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                7.3401785 = idf(docFreq=77, maxDocs=44218)
                0.0625 = fieldNorm(doc=2895)
        0.24 = coord(6/25)
    
  5. Kim, J.; Kim, J.; Owen-Smith, J.: Ethnicity-based name partitioning for author name disambiguation using supervised machine learning (2021) 0.17
    0.1664808 = sum of:
      0.1664808 = product of:
        0.83240396 = sum of:
          0.064266935 = weight(abstract_txt:disambiguate in 311) [ClassicSimilarity], result of:
            0.064266935 = score(doc=311,freq=1.0), product of:
              0.118842766 = queryWeight, product of:
                1.0551176 = boost
                8.652365 = idf(docFreq=20, maxDocs=44218)
                0.013017785 = queryNorm
              0.5407728 = fieldWeight in 311, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                8.652365 = idf(docFreq=20, maxDocs=44218)
                0.0625 = fieldNorm(doc=311)
          0.019421639 = weight(abstract_txt:approaches in 311) [ClassicSimilarity], result of:
            0.019421639 = score(doc=311,freq=1.0), product of:
              0.067429245 = queryWeight, product of:
                1.1239672 = boost
                4.6084785 = idf(docFreq=1197, maxDocs=44218)
                0.013017785 = queryNorm
              0.2880299 = fieldWeight in 311, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.6084785 = idf(docFreq=1197, maxDocs=44218)
                0.0625 = fieldNorm(doc=311)
          0.0146590145 = weight(abstract_txt:different in 311) [ClassicSimilarity], result of:
            0.0146590145 = score(doc=311,freq=1.0), product of:
              0.063986935 = queryWeight, product of:
                1.3409753 = boost
                3.6655018 = idf(docFreq=3075, maxDocs=44218)
                0.013017785 = queryNorm
              0.22909386 = fieldWeight in 311, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.6655018 = idf(docFreq=3075, maxDocs=44218)
                0.0625 = fieldNorm(doc=311)
          0.032151897 = weight(abstract_txt:better in 311) [ClassicSimilarity], result of:
            0.032151897 = score(doc=311,freq=1.0), product of:
              0.1080171 = queryWeight, product of:
                1.7422937 = boost
                4.76249 = idf(docFreq=1026, maxDocs=44218)
                0.013017785 = queryNorm
              0.2976556 = fieldWeight in 311, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.76249 = idf(docFreq=1026, maxDocs=44218)
                0.0625 = fieldNorm(doc=311)
          0.7019045 = weight(abstract_txt:disambiguation in 311) [ClassicSimilarity], result of:
            0.7019045 = score(doc=311,freq=5.0), product of:
              0.68423676 = queryWeight, product of:
                7.1608186 = boost
                7.3401785 = idf(docFreq=77, maxDocs=44218)
                0.013017785 = queryNorm
              1.0258211 = fieldWeight in 311, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                7.3401785 = idf(docFreq=77, maxDocs=44218)
                0.0625 = fieldNorm(doc=311)
        0.2 = coord(5/25)