Document (#26684)

Author
Yang, C.C.
Li, K.W.
Title
Automatic construction of English/Chinese parallel corpora
Source
Journal of the American Society for Information Science and technology. 54(2003) no.8, S.730-742
Year
2003
Abstract
As the demand for global information increases significantly, multilingual corpora has become a valuable linguistic resource for applications to cross-lingual information retrieval and natural language processing. In order to cross the boundaries that exist between different languages, dictionaries are the most typical tools. However, the general-purpose dictionary is less sensitive in both genre and domain. It is also impractical to manually construct tailored bilingual dictionaries or sophisticated multilingual thesauri for large applications. Corpusbased approaches, which do not have the limitation of dictionaries, provide a statistical translation model with which to cross the language boundary. There are many domain-specific parallel or comparable corpora that are employed in machine translation and cross-lingual information retrieval. Most of these are corpora between Indo-European languages, such as English/French and English/Spanish. The Asian/Indo-European corpus, especially English/Chinese corpus, is relatively sparse. The objective of the present research is to construct English/ Chinese parallel corpus automatically from the World Wide Web. In this paper, an alignment method is presented which is based an dynamic programming to identify the one-to-one Chinese and English title pairs. The method includes alignment at title level, word level and character level. The longest common subsequence (LCS) is applied to find the most reliabie Chinese translation of an English word. As one word for a language may translate into two or more words repetitively in another language, the edit operation, deletion, is used to resolve redundancy. A score function is then proposed to determine the optimal title pairs. Experiments have been conducted to investigate the performance of the proposed method using the daily press release articles by the Hong Kong SAR government as the test bed. The precision of the result is 0.998 while the recall is 0.806. The release articles and speech articles, published by Hongkong & Shanghai Banking Corporation Limited, are also used to test our method, the precision is 1.00, and the recall is 0.948.
Theme
Computerlinguistik

Similar documents (author)

  1. Yang, S.C.: ¬An interpretive and situated approach to an evaluation of Perseus digital libraries (2001) 4.50
    4.4981737 = sum of:
      4.4981737 = weight(author_txt:yang in 6933) [ClassicSimilarity], result of:
        4.4981737 = fieldWeight in 6933, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          7.1970778 = idf(docFreq=89, maxDocs=44218)
          0.625 = fieldNorm(doc=6933)
    
  2. Yang, K.: Information retrieval on the Web (2004) 4.50
    4.4981737 = sum of:
      4.4981737 = weight(author_txt:yang in 4278) [ClassicSimilarity], result of:
        4.4981737 = fieldWeight in 4278, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          7.1970778 = idf(docFreq=89, maxDocs=44218)
          0.625 = fieldNorm(doc=4278)
    
  3. Yang, C.C.: Content-based image retrievaI : a comparison between query by example and image browsing map approaches (2005) 4.50
    4.4981737 = sum of:
      4.4981737 = weight(author_txt:yang in 4649) [ClassicSimilarity], result of:
        4.4981737 = fieldWeight in 4649, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          7.1970778 = idf(docFreq=89, maxDocs=44218)
          0.625 = fieldNorm(doc=4649)
    
  4. Salton, G.; Yang, C.S.: On the specification of term values in automatic indexing (1973) 3.60
    3.5985389 = sum of:
      3.5985389 = weight(author_txt:yang in 5476) [ClassicSimilarity], result of:
        3.5985389 = fieldWeight in 5476, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          7.1970778 = idf(docFreq=89, maxDocs=44218)
          0.5 = fieldNorm(doc=5476)
    
  5. Yang, Y.; Chute, C.G.A.: ¬A schematic analysis of the Unified Medical Language System (1992) 3.60
    3.5985389 = sum of:
      3.5985389 = weight(author_txt:yang in 6445) [ClassicSimilarity], result of:
        3.5985389 = fieldWeight in 6445, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          7.1970778 = idf(docFreq=89, maxDocs=44218)
          0.5 = fieldNorm(doc=6445)
    

Similar documents (content)

  1. Li, K.W.; Yang, C.C.: Conceptual analysis of parallel corpus collected from the Web (2006) 1.60
    1.5986034 = sum of:
      1.5986034 = product of:
        2.2202823 = sum of:
          0.06124737 = weight(abstract_txt:languages in 5051) [ClassicSimilarity], result of:
            0.06124737 = score(doc=5051,freq=4.0), product of:
              0.094442524 = queryWeight, product of:
                5.188118 = idf(docFreq=670, maxDocs=44218)
                0.01820362 = queryNorm
              0.64851475 = fieldWeight in 5051, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                5.188118 = idf(docFreq=670, maxDocs=44218)
                0.0625 = fieldNorm(doc=5051)
          0.036988672 = weight(abstract_txt:precision in 5051) [ClassicSimilarity], result of:
            0.036988672 = score(doc=5051,freq=1.0), product of:
              0.1071129 = queryWeight, product of:
                1.0649693 = boost
                5.5251865 = idf(docFreq=478, maxDocs=44218)
                0.01820362 = queryNorm
              0.34532416 = fieldWeight in 5051, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.5251865 = idf(docFreq=478, maxDocs=44218)
                0.0625 = fieldNorm(doc=5051)
          0.05722246 = weight(abstract_txt:european in 5051) [ClassicSimilarity], result of:
            0.05722246 = score(doc=5051,freq=2.0), product of:
              0.113718286 = queryWeight, product of:
                1.0973151 = boost
                5.6930003 = idf(docFreq=404, maxDocs=44218)
                0.01820362 = queryNorm
              0.50319487 = fieldWeight in 5051, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                5.6930003 = idf(docFreq=404, maxDocs=44218)
                0.0625 = fieldNorm(doc=5051)
          0.041665 = weight(abstract_txt:recall in 5051) [ClassicSimilarity], result of:
            0.041665 = score(doc=5051,freq=1.0), product of:
              0.11596054 = queryWeight, product of:
                1.1080805 = boost
                5.7488523 = idf(docFreq=382, maxDocs=44218)
                0.01820362 = queryNorm
              0.35930327 = fieldWeight in 5051, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.7488523 = idf(docFreq=382, maxDocs=44218)
                0.0625 = fieldNorm(doc=5051)
          0.051346365 = weight(abstract_txt:construct in 5051) [ClassicSimilarity], result of:
            0.051346365 = score(doc=5051,freq=1.0), product of:
              0.1332915 = queryWeight, product of:
                1.1880027 = boost
                6.163498 = idf(docFreq=252, maxDocs=44218)
                0.01820362 = queryNorm
              0.38521862 = fieldWeight in 5051, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.163498 = idf(docFreq=252, maxDocs=44218)
                0.0625 = fieldNorm(doc=5051)
          0.07749984 = weight(abstract_txt:multilingual in 5051) [ClassicSimilarity], result of:
            0.07749984 = score(doc=5051,freq=2.0), product of:
              0.13920446 = queryWeight, product of:
                1.2140673 = boost
                6.2987247 = idf(docFreq=220, maxDocs=44218)
                0.01820362 = queryNorm
              0.55673385 = fieldWeight in 5051, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                6.2987247 = idf(docFreq=220, maxDocs=44218)
                0.0625 = fieldNorm(doc=5051)
          0.09747362 = weight(abstract_txt:pairs in 5051) [ClassicSimilarity], result of:
            0.09747362 = score(doc=5051,freq=2.0), product of:
              0.16219746 = queryWeight, product of:
                1.3105036 = boost
                6.7990475 = idf(docFreq=133, maxDocs=44218)
                0.01820362 = queryNorm
              0.60095656 = fieldWeight in 5051, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                6.7990475 = idf(docFreq=133, maxDocs=44218)
                0.0625 = fieldNorm(doc=5051)
          0.19704978 = weight(abstract_txt:alignment in 5051) [ClassicSimilarity], result of:
            0.19704978 = score(doc=5051,freq=5.0), product of:
              0.19106886 = queryWeight, product of:
                1.4223653 = boost
                7.3793993 = idf(docFreq=74, maxDocs=44218)
                0.01820362 = queryNorm
              1.0313025 = fieldWeight in 5051, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                7.3793993 = idf(docFreq=74, maxDocs=44218)
                0.0625 = fieldNorm(doc=5051)
          0.11712011 = weight(abstract_txt:lingual in 5051) [ClassicSimilarity], result of:
            0.11712011 = score(doc=5051,freq=1.0), product of:
              0.23096718 = queryWeight, product of:
                1.5638365 = boost
                8.113368 = idf(docFreq=35, maxDocs=44218)
                0.01820362 = queryNorm
              0.5070855 = fieldWeight in 5051, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                8.113368 = idf(docFreq=35, maxDocs=44218)
                0.0625 = fieldNorm(doc=5051)
          0.08720132 = weight(abstract_txt:title in 5051) [ClassicSimilarity], result of:
            0.08720132 = score(doc=5051,freq=2.0), product of:
              0.17238459 = queryWeight, product of:
                1.6546687 = boost
                5.723078 = idf(docFreq=392, maxDocs=44218)
                0.01820362 = queryNorm
              0.50585335 = fieldWeight in 5051, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                5.723078 = idf(docFreq=392, maxDocs=44218)
                0.0625 = fieldNorm(doc=5051)
          0.08942429 = weight(abstract_txt:method in 5051) [ClassicSimilarity], result of:
            0.08942429 = score(doc=5051,freq=5.0), product of:
              0.1421629 = queryWeight, product of:
                1.7350993 = boost
                4.50095 = idf(docFreq=1333, maxDocs=44218)
                0.01820362 = queryNorm
              0.6290269 = fieldWeight in 5051, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                4.50095 = idf(docFreq=1333, maxDocs=44218)
                0.0625 = fieldNorm(doc=5051)
          0.07460723 = weight(abstract_txt:corpus in 5051) [ClassicSimilarity], result of:
            0.07460723 = score(doc=5051,freq=1.0), product of:
              0.19574034 = queryWeight, product of:
                1.7632017 = boost
                6.0984654 = idf(docFreq=269, maxDocs=44218)
                0.01820362 = queryNorm
              0.3811541 = fieldWeight in 5051, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.0984654 = idf(docFreq=269, maxDocs=44218)
                0.0625 = fieldNorm(doc=5051)
          0.07746759 = weight(abstract_txt:translation in 5051) [ClassicSimilarity], result of:
            0.07746759 = score(doc=5051,freq=1.0), product of:
              0.20071188 = queryWeight, product of:
                1.7854528 = boost
                6.1754265 = idf(docFreq=249, maxDocs=44218)
                0.01820362 = queryNorm
              0.38596416 = fieldWeight in 5051, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.1754265 = idf(docFreq=249, maxDocs=44218)
                0.0625 = fieldNorm(doc=5051)
          0.17356883 = weight(abstract_txt:parallel in 5051) [ClassicSimilarity], result of:
            0.17356883 = score(doc=5051,freq=4.0), product of:
              0.21649817 = queryWeight, product of:
                1.8543382 = boost
                6.4136834 = idf(docFreq=196, maxDocs=44218)
                0.01820362 = queryNorm
              0.8017104 = fieldWeight in 5051, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                6.4136834 = idf(docFreq=196, maxDocs=44218)
                0.0625 = fieldNorm(doc=5051)
          0.077067144 = weight(abstract_txt:cross in 5051) [ClassicSimilarity], result of:
            0.077067144 = score(doc=5051,freq=1.0), product of:
              0.22015007 = queryWeight, product of:
                2.1591887 = boost
                5.601063 = idf(docFreq=443, maxDocs=44218)
                0.01820362 = queryNorm
              0.35006642 = fieldWeight in 5051, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.601063 = idf(docFreq=443, maxDocs=44218)
                0.0625 = fieldNorm(doc=5051)
          0.33986562 = weight(abstract_txt:corpora in 5051) [ClassicSimilarity], result of:
            0.33986562 = score(doc=5051,freq=5.0), product of:
              0.34622157 = queryWeight, product of:
                2.7077482 = boost
                7.0240583 = idf(docFreq=106, maxDocs=44218)
                0.01820362 = queryNorm
              0.981642 = fieldWeight in 5051, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                7.0240583 = idf(docFreq=106, maxDocs=44218)
                0.0625 = fieldNorm(doc=5051)
          0.23780675 = weight(abstract_txt:chinese in 5051) [ClassicSimilarity], result of:
            0.23780675 = score(doc=5051,freq=3.0), product of:
              0.34851247 = queryWeight, product of:
                3.0373538 = boost
                6.30326 = idf(docFreq=219, maxDocs=44218)
                0.01820362 = queryNorm
              0.6823479 = fieldWeight in 5051, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                6.30326 = idf(docFreq=219, maxDocs=44218)
                0.0625 = fieldNorm(doc=5051)
          0.3256602 = weight(abstract_txt:english in 5051) [ClassicSimilarity], result of:
            0.3256602 = score(doc=5051,freq=6.0), product of:
              0.38160262 = queryWeight, product of:
                3.7605891 = boost
                5.574394 = idf(docFreq=455, maxDocs=44218)
                0.01820362 = queryNorm
              0.85340136 = fieldWeight in 5051, product of:
                2.4494898 = tf(freq=6.0), with freq of:
                  6.0 = termFreq=6.0
                5.574394 = idf(docFreq=455, maxDocs=44218)
                0.0625 = fieldNorm(doc=5051)
        0.72 = coord(18/25)
    
  2. Xu, J.; Weischedel, R.: Empirical studies on the impact of lexical resources on CLIR performance (2005) 0.57
    0.5697967 = sum of:
      0.5697967 = product of:
        1.5827686 = sum of:
          0.14640014 = weight(abstract_txt:lingual in 1020) [ClassicSimilarity], result of:
            0.14640014 = score(doc=1020,freq=1.0), product of:
              0.23096718 = queryWeight, product of:
                1.5638365 = boost
                8.113368 = idf(docFreq=35, maxDocs=44218)
                0.01820362 = queryNorm
              0.6338569 = fieldWeight in 1020, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                8.113368 = idf(docFreq=35, maxDocs=44218)
                0.078125 = fieldNorm(doc=1020)
          0.05671034 = weight(abstract_txt:language in 1020) [ClassicSimilarity], result of:
            0.05671034 = score(doc=1020,freq=2.0), product of:
              0.12273379 = queryWeight, product of:
                1.612179 = boost
                4.1820874 = idf(docFreq=1834, maxDocs=44218)
                0.01820362 = queryNorm
              0.46205974 = fieldWeight in 1020, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                4.1820874 = idf(docFreq=1834, maxDocs=44218)
                0.078125 = fieldNorm(doc=1020)
          0.13188821 = weight(abstract_txt:corpus in 1020) [ClassicSimilarity], result of:
            0.13188821 = score(doc=1020,freq=2.0), product of:
              0.19574034 = queryWeight, product of:
                1.7632017 = boost
                6.0984654 = idf(docFreq=269, maxDocs=44218)
                0.01820362 = queryNorm
              0.67379165 = fieldWeight in 1020, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                6.0984654 = idf(docFreq=269, maxDocs=44218)
                0.078125 = fieldNorm(doc=1020)
          0.09683449 = weight(abstract_txt:translation in 1020) [ClassicSimilarity], result of:
            0.09683449 = score(doc=1020,freq=1.0), product of:
              0.20071188 = queryWeight, product of:
                1.7854528 = boost
                6.1754265 = idf(docFreq=249, maxDocs=44218)
                0.01820362 = queryNorm
              0.4824552 = fieldWeight in 1020, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.1754265 = idf(docFreq=249, maxDocs=44218)
                0.078125 = fieldNorm(doc=1020)
          0.26572195 = weight(abstract_txt:parallel in 1020) [ClassicSimilarity], result of:
            0.26572195 = score(doc=1020,freq=6.0), product of:
              0.21649817 = queryWeight, product of:
                1.8543382 = boost
                6.4136834 = idf(docFreq=196, maxDocs=44218)
                0.01820362 = queryNorm
              1.2273635 = fieldWeight in 1020, product of:
                2.4494898 = tf(freq=6.0), with freq of:
                  6.0 = termFreq=6.0
                6.4136834 = idf(docFreq=196, maxDocs=44218)
                0.078125 = fieldNorm(doc=1020)
          0.096333936 = weight(abstract_txt:cross in 1020) [ClassicSimilarity], result of:
            0.096333936 = score(doc=1020,freq=1.0), product of:
              0.22015007 = queryWeight, product of:
                2.1591887 = boost
                5.601063 = idf(docFreq=443, maxDocs=44218)
                0.01820362 = queryNorm
              0.43758303 = fieldWeight in 1020, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.601063 = idf(docFreq=443, maxDocs=44218)
                0.078125 = fieldNorm(doc=1020)
          0.37998134 = weight(abstract_txt:corpora in 1020) [ClassicSimilarity], result of:
            0.37998134 = score(doc=1020,freq=4.0), product of:
              0.34622157 = queryWeight, product of:
                2.7077482 = boost
                7.0240583 = idf(docFreq=106, maxDocs=44218)
                0.01820362 = queryNorm
              1.0975091 = fieldWeight in 1020, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                7.0240583 = idf(docFreq=106, maxDocs=44218)
                0.078125 = fieldNorm(doc=1020)
          0.24271047 = weight(abstract_txt:chinese in 1020) [ClassicSimilarity], result of:
            0.24271047 = score(doc=1020,freq=2.0), product of:
              0.34851247 = queryWeight, product of:
                3.0373538 = boost
                6.30326 = idf(docFreq=219, maxDocs=44218)
                0.01820362 = queryNorm
              0.69641834 = fieldWeight in 1020, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                6.30326 = idf(docFreq=219, maxDocs=44218)
                0.078125 = fieldNorm(doc=1020)
          0.16618776 = weight(abstract_txt:english in 1020) [ClassicSimilarity], result of:
            0.16618776 = score(doc=1020,freq=1.0), product of:
              0.38160262 = queryWeight, product of:
                3.7605891 = boost
                5.574394 = idf(docFreq=455, maxDocs=44218)
                0.01820362 = queryNorm
              0.43549955 = fieldWeight in 1020, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.574394 = idf(docFreq=455, maxDocs=44218)
                0.078125 = fieldNorm(doc=1020)
        0.36 = coord(9/25)
    
  3. Yang, C.C.; Luk, J.: Automatic generation of English/Chinese thesaurus based on a parallel corpus in laws (2003) 0.54
    0.5425796 = sum of:
      0.5425796 = product of:
        1.0434223 = sum of:
          0.038279604 = weight(abstract_txt:languages in 1616) [ClassicSimilarity], result of:
            0.038279604 = score(doc=1616,freq=4.0), product of:
              0.094442524 = queryWeight, product of:
                5.188118 = idf(docFreq=670, maxDocs=44218)
                0.01820362 = queryNorm
              0.40532172 = fieldWeight in 1616, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                5.188118 = idf(docFreq=670, maxDocs=44218)
                0.0390625 = fieldNorm(doc=1616)
          0.02311792 = weight(abstract_txt:precision in 1616) [ClassicSimilarity], result of:
            0.02311792 = score(doc=1616,freq=1.0), product of:
              0.1071129 = queryWeight, product of:
                1.0649693 = boost
                5.5251865 = idf(docFreq=478, maxDocs=44218)
                0.01820362 = queryNorm
              0.2158276 = fieldWeight in 1616, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.5251865 = idf(docFreq=478, maxDocs=44218)
                0.0390625 = fieldNorm(doc=1616)
          0.025288993 = weight(abstract_txt:european in 1616) [ClassicSimilarity], result of:
            0.025288993 = score(doc=1616,freq=1.0), product of:
              0.113718286 = queryWeight, product of:
                1.0973151 = boost
                5.6930003 = idf(docFreq=404, maxDocs=44218)
                0.01820362 = queryNorm
              0.22238283 = fieldWeight in 1616, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.6930003 = idf(docFreq=404, maxDocs=44218)
                0.0390625 = fieldNorm(doc=1616)
          0.026040625 = weight(abstract_txt:recall in 1616) [ClassicSimilarity], result of:
            0.026040625 = score(doc=1616,freq=1.0), product of:
              0.11596054 = queryWeight, product of:
                1.1080805 = boost
                5.7488523 = idf(docFreq=382, maxDocs=44218)
                0.01820362 = queryNorm
              0.22456454 = fieldWeight in 1616, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.7488523 = idf(docFreq=382, maxDocs=44218)
                0.0390625 = fieldNorm(doc=1616)
          0.012609807 = weight(abstract_txt:most in 1616) [ClassicSimilarity], result of:
            0.012609807 = score(doc=1616,freq=1.0), product of:
              0.08185502 = queryWeight, product of:
                1.1402091 = boost
                3.943693 = idf(docFreq=2328, maxDocs=44218)
                0.01820362 = queryNorm
              0.1540505 = fieldWeight in 1616, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.943693 = idf(docFreq=2328, maxDocs=44218)
                0.0390625 = fieldNorm(doc=1616)
          0.16368033 = weight(abstract_txt:lingual in 1616) [ClassicSimilarity], result of:
            0.16368033 = score(doc=1616,freq=5.0), product of:
              0.23096718 = queryWeight, product of:
                1.5638365 = boost
                8.113368 = idf(docFreq=35, maxDocs=44218)
                0.01820362 = queryNorm
              0.70867354 = fieldWeight in 1616, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                8.113368 = idf(docFreq=35, maxDocs=44218)
                0.0390625 = fieldNorm(doc=1616)
          0.034727853 = weight(abstract_txt:language in 1616) [ClassicSimilarity], result of:
            0.034727853 = score(doc=1616,freq=3.0), product of:
              0.12273379 = queryWeight, product of:
                1.612179 = boost
                4.1820874 = idf(docFreq=1834, maxDocs=44218)
                0.01820362 = queryNorm
              0.28295267 = fieldWeight in 1616, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                4.1820874 = idf(docFreq=1834, maxDocs=44218)
                0.0390625 = fieldNorm(doc=1616)
          0.065944105 = weight(abstract_txt:corpus in 1616) [ClassicSimilarity], result of:
            0.065944105 = score(doc=1616,freq=2.0), product of:
              0.19574034 = queryWeight, product of:
                1.7632017 = boost
                6.0984654 = idf(docFreq=269, maxDocs=44218)
                0.01820362 = queryNorm
              0.33689582 = fieldWeight in 1616, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                6.0984654 = idf(docFreq=269, maxDocs=44218)
                0.0390625 = fieldNorm(doc=1616)
          0.048417244 = weight(abstract_txt:translation in 1616) [ClassicSimilarity], result of:
            0.048417244 = score(doc=1616,freq=1.0), product of:
              0.20071188 = queryWeight, product of:
                1.7854528 = boost
                6.1754265 = idf(docFreq=249, maxDocs=44218)
                0.01820362 = queryNorm
              0.2412276 = fieldWeight in 1616, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.1754265 = idf(docFreq=249, maxDocs=44218)
                0.0390625 = fieldNorm(doc=1616)
          0.07670732 = weight(abstract_txt:parallel in 1616) [ClassicSimilarity], result of:
            0.07670732 = score(doc=1616,freq=2.0), product of:
              0.21649817 = queryWeight, product of:
                1.8543382 = boost
                6.4136834 = idf(docFreq=196, maxDocs=44218)
                0.01820362 = queryNorm
              0.35430932 = fieldWeight in 1616, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                6.4136834 = idf(docFreq=196, maxDocs=44218)
                0.0390625 = fieldNorm(doc=1616)
          0.10770461 = weight(abstract_txt:cross in 1616) [ClassicSimilarity], result of:
            0.10770461 = score(doc=1616,freq=5.0), product of:
              0.22015007 = queryWeight, product of:
                2.1591887 = boost
                5.601063 = idf(docFreq=443, maxDocs=44218)
                0.01820362 = queryNorm
              0.4892327 = fieldWeight in 1616, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                5.601063 = idf(docFreq=443, maxDocs=44218)
                0.0390625 = fieldNorm(doc=1616)
          0.17162225 = weight(abstract_txt:chinese in 1616) [ClassicSimilarity], result of:
            0.17162225 = score(doc=1616,freq=4.0), product of:
              0.34851247 = queryWeight, product of:
                3.0373538 = boost
                6.30326 = idf(docFreq=219, maxDocs=44218)
                0.01820362 = queryNorm
              0.4924422 = fieldWeight in 1616, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                6.30326 = idf(docFreq=219, maxDocs=44218)
                0.0390625 = fieldNorm(doc=1616)
          0.24928164 = weight(abstract_txt:english in 1616) [ClassicSimilarity], result of:
            0.24928164 = score(doc=1616,freq=9.0), product of:
              0.38160262 = queryWeight, product of:
                3.7605891 = boost
                5.574394 = idf(docFreq=455, maxDocs=44218)
                0.01820362 = queryNorm
              0.6532493 = fieldWeight in 1616, product of:
                3.0 = tf(freq=9.0), with freq of:
                  9.0 = termFreq=9.0
                5.574394 = idf(docFreq=455, maxDocs=44218)
                0.0390625 = fieldNorm(doc=1616)
        0.52 = coord(13/25)
    
  4. Lam, W.; Chan, K.; Radev, D.; Saggion, H.; Teufel, S.: Context-based generic cross-lingual retrieval of documents and automated summaries (2005) 0.50
    0.49969196 = sum of:
      0.49969196 = product of:
        1.3880332 = sum of:
          0.08615533 = weight(abstract_txt:pairs in 1965) [ClassicSimilarity], result of:
            0.08615533 = score(doc=1965,freq=1.0), product of:
              0.16219746 = queryWeight, product of:
                1.3105036 = boost
                6.7990475 = idf(docFreq=133, maxDocs=44218)
                0.01820362 = queryNorm
              0.5311756 = fieldWeight in 1965, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.7990475 = idf(docFreq=133, maxDocs=44218)
                0.078125 = fieldNorm(doc=1965)
          0.29280028 = weight(abstract_txt:lingual in 1965) [ClassicSimilarity], result of:
            0.29280028 = score(doc=1965,freq=4.0), product of:
              0.23096718 = queryWeight, product of:
                1.5638365 = boost
                8.113368 = idf(docFreq=35, maxDocs=44218)
                0.01820362 = queryNorm
              1.2677138 = fieldWeight in 1965, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                8.113368 = idf(docFreq=35, maxDocs=44218)
                0.078125 = fieldNorm(doc=1965)
          0.040100265 = weight(abstract_txt:language in 1965) [ClassicSimilarity], result of:
            0.040100265 = score(doc=1965,freq=1.0), product of:
              0.12273379 = queryWeight, product of:
                1.612179 = boost
                4.1820874 = idf(docFreq=1834, maxDocs=44218)
                0.01820362 = queryNorm
              0.32672557 = fieldWeight in 1965, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.1820874 = idf(docFreq=1834, maxDocs=44218)
                0.078125 = fieldNorm(doc=1965)
          0.09325904 = weight(abstract_txt:corpus in 1965) [ClassicSimilarity], result of:
            0.09325904 = score(doc=1965,freq=1.0), product of:
              0.19574034 = queryWeight, product of:
                1.7632017 = boost
                6.0984654 = idf(docFreq=269, maxDocs=44218)
                0.01820362 = queryNorm
              0.4764426 = fieldWeight in 1965, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.0984654 = idf(docFreq=269, maxDocs=44218)
                0.078125 = fieldNorm(doc=1965)
          0.09683449 = weight(abstract_txt:translation in 1965) [ClassicSimilarity], result of:
            0.09683449 = score(doc=1965,freq=1.0), product of:
              0.20071188 = queryWeight, product of:
                1.7854528 = boost
                6.1754265 = idf(docFreq=249, maxDocs=44218)
                0.01820362 = queryNorm
              0.4824552 = fieldWeight in 1965, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.1754265 = idf(docFreq=249, maxDocs=44218)
                0.078125 = fieldNorm(doc=1965)
          0.10848052 = weight(abstract_txt:parallel in 1965) [ClassicSimilarity], result of:
            0.10848052 = score(doc=1965,freq=1.0), product of:
              0.21649817 = queryWeight, product of:
                1.8543382 = boost
                6.4136834 = idf(docFreq=196, maxDocs=44218)
                0.01820362 = queryNorm
              0.501069 = fieldWeight in 1965, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.4136834 = idf(docFreq=196, maxDocs=44218)
                0.078125 = fieldNorm(doc=1965)
          0.19266787 = weight(abstract_txt:cross in 1965) [ClassicSimilarity], result of:
            0.19266787 = score(doc=1965,freq=4.0), product of:
              0.22015007 = queryWeight, product of:
                2.1591887 = boost
                5.601063 = idf(docFreq=443, maxDocs=44218)
                0.01820362 = queryNorm
              0.87516606 = fieldWeight in 1965, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                5.601063 = idf(docFreq=443, maxDocs=44218)
                0.078125 = fieldNorm(doc=1965)
          0.24271047 = weight(abstract_txt:chinese in 1965) [ClassicSimilarity], result of:
            0.24271047 = score(doc=1965,freq=2.0), product of:
              0.34851247 = queryWeight, product of:
                3.0373538 = boost
                6.30326 = idf(docFreq=219, maxDocs=44218)
                0.01820362 = queryNorm
              0.69641834 = fieldWeight in 1965, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                6.30326 = idf(docFreq=219, maxDocs=44218)
                0.078125 = fieldNorm(doc=1965)
          0.23502499 = weight(abstract_txt:english in 1965) [ClassicSimilarity], result of:
            0.23502499 = score(doc=1965,freq=2.0), product of:
              0.38160262 = queryWeight, product of:
                3.7605891 = boost
                5.574394 = idf(docFreq=455, maxDocs=44218)
                0.01820362 = queryNorm
              0.6158894 = fieldWeight in 1965, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                5.574394 = idf(docFreq=455, maxDocs=44218)
                0.078125 = fieldNorm(doc=1965)
        0.36 = coord(9/25)
    
  5. Talvensaari, T.; Laurikkala, J.; Järvelin, K.; Juhola, M.: ¬A study on automatic creation of a comparable document collection in cross-language information retrieval (2006) 0.49
    0.48809305 = sum of:
      0.48809305 = product of:
        1.1093024 = sum of:
          0.05304178 = weight(abstract_txt:languages in 5601) [ClassicSimilarity], result of:
            0.05304178 = score(doc=5601,freq=3.0), product of:
              0.094442524 = queryWeight, product of:
                5.188118 = idf(docFreq=670, maxDocs=44218)
                0.01820362 = queryNorm
              0.56163025 = fieldWeight in 5601, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                5.188118 = idf(docFreq=670, maxDocs=44218)
                0.0625 = fieldNorm(doc=5601)
          0.119380325 = weight(abstract_txt:pairs in 5601) [ClassicSimilarity], result of:
            0.119380325 = score(doc=5601,freq=3.0), product of:
              0.16219746 = queryWeight, product of:
                1.3105036 = boost
                6.7990475 = idf(docFreq=133, maxDocs=44218)
                0.01820362 = queryNorm
              0.7360185 = fieldWeight in 5601, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                6.7990475 = idf(docFreq=133, maxDocs=44218)
                0.0625 = fieldNorm(doc=5601)
          0.05087513 = weight(abstract_txt:articles in 5601) [ClassicSimilarity], result of:
            0.05087513 = score(doc=5601,freq=2.0), product of:
              0.12036127 = queryWeight, product of:
                1.3826276 = boost
                4.7821565 = idf(docFreq=1006, maxDocs=44218)
                0.01820362 = queryNorm
              0.4226869 = fieldWeight in 5601, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                4.7821565 = idf(docFreq=1006, maxDocs=44218)
                0.0625 = fieldNorm(doc=5601)
          0.1526341 = weight(abstract_txt:alignment in 5601) [ClassicSimilarity], result of:
            0.1526341 = score(doc=5601,freq=3.0), product of:
              0.19106886 = queryWeight, product of:
                1.4223653 = boost
                7.3793993 = idf(docFreq=74, maxDocs=44218)
                0.01820362 = queryNorm
              0.7988434 = fieldWeight in 5601, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                7.3793993 = idf(docFreq=74, maxDocs=44218)
                0.0625 = fieldNorm(doc=5601)
          0.032080214 = weight(abstract_txt:language in 5601) [ClassicSimilarity], result of:
            0.032080214 = score(doc=5601,freq=1.0), product of:
              0.12273379 = queryWeight, product of:
                1.612179 = boost
                4.1820874 = idf(docFreq=1834, maxDocs=44218)
                0.01820362 = queryNorm
              0.26138046 = fieldWeight in 5601, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.1820874 = idf(docFreq=1834, maxDocs=44218)
                0.0625 = fieldNorm(doc=5601)
          0.07998351 = weight(abstract_txt:method in 5601) [ClassicSimilarity], result of:
            0.07998351 = score(doc=5601,freq=4.0), product of:
              0.1421629 = queryWeight, product of:
                1.7350993 = boost
                4.50095 = idf(docFreq=1333, maxDocs=44218)
                0.01820362 = queryNorm
              0.56261873 = fieldWeight in 5601, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                4.50095 = idf(docFreq=1333, maxDocs=44218)
                0.0625 = fieldNorm(doc=5601)
          0.109555714 = weight(abstract_txt:translation in 5601) [ClassicSimilarity], result of:
            0.109555714 = score(doc=5601,freq=2.0), product of:
              0.20071188 = queryWeight, product of:
                1.7854528 = boost
                6.1754265 = idf(docFreq=249, maxDocs=44218)
                0.01820362 = queryNorm
              0.54583573 = fieldWeight in 5601, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                6.1754265 = idf(docFreq=249, maxDocs=44218)
                0.0625 = fieldNorm(doc=5601)
          0.086784415 = weight(abstract_txt:parallel in 5601) [ClassicSimilarity], result of:
            0.086784415 = score(doc=5601,freq=1.0), product of:
              0.21649817 = queryWeight, product of:
                1.8543382 = boost
                6.4136834 = idf(docFreq=196, maxDocs=44218)
                0.01820362 = queryNorm
              0.4008552 = fieldWeight in 5601, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.4136834 = idf(docFreq=196, maxDocs=44218)
                0.0625 = fieldNorm(doc=5601)
          0.077067144 = weight(abstract_txt:cross in 5601) [ClassicSimilarity], result of:
            0.077067144 = score(doc=5601,freq=1.0), product of:
              0.22015007 = queryWeight, product of:
                2.1591887 = boost
                5.601063 = idf(docFreq=443, maxDocs=44218)
                0.01820362 = queryNorm
              0.35006642 = fieldWeight in 5601, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.601063 = idf(docFreq=443, maxDocs=44218)
                0.0625 = fieldNorm(doc=5601)
          0.21494989 = weight(abstract_txt:corpora in 5601) [ClassicSimilarity], result of:
            0.21494989 = score(doc=5601,freq=2.0), product of:
              0.34622157 = queryWeight, product of:
                2.7077482 = boost
                7.0240583 = idf(docFreq=106, maxDocs=44218)
                0.01820362 = queryNorm
              0.6208449 = fieldWeight in 5601, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                7.0240583 = idf(docFreq=106, maxDocs=44218)
                0.0625 = fieldNorm(doc=5601)
          0.13295022 = weight(abstract_txt:english in 5601) [ClassicSimilarity], result of:
            0.13295022 = score(doc=5601,freq=1.0), product of:
              0.38160262 = queryWeight, product of:
                3.7605891 = boost
                5.574394 = idf(docFreq=455, maxDocs=44218)
                0.01820362 = queryNorm
              0.34839964 = fieldWeight in 5601, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.574394 = idf(docFreq=455, maxDocs=44218)
                0.0625 = fieldNorm(doc=5601)
        0.44 = coord(11/25)