Search (1 results, page 1 of 1)

Did you mean:
rvk_ss%3a%2200 75400 allgemeines %2f buch- und bibliothekswesen%2c informationswissenschaft %2f bibliothekswesen %2f sacherschlie%c3%9fung in bibliotheken %2f schlagwortregeln%2c schlagwortverzeichnis%22 1
rvk_ss%3a%2200 75000 allgemeines %2f buch- und bibliothekswesen%2c informationswissenschaft %2f bibliothekswesen %2f sacherschlie%c3%9fung in bibliotheken %2f schlagwortregeln%2c schlagwortverzeichnis%22 1
rvk_ss%3a%2200 75400 allgemeinen %2f buch- und bibliothekswesen%2c informationswissenschaft %2f bibliothekswesen %2f sacherschlie%c3%9fung in bibliotheken %2f schlagwortregeln%2c schlagwortverzeichnis%22 1
rvk_ss%3a%2200 75400 allgemeines %2f buch- und bibliothekswesen%2c informationswissenschaft %2f bibliothekswesen %2f sacherschlie%c3%9fung in bibliotheken %2f schlagwortregeln%2c sschlagwortverzeichnis%22 1
rvk_ss%3a%2223 75400 allgemeines %2f buch- und bibliothekswesen%2c informationswissenschaft %2f bibliothekswesen %2f sacherschlie%c3%9fung in bibliotheken %2f schlagwortregeln%2c schlagwortverzeichnis%22 1

Muneer, I.; Sharjeel, M.; Iqbal, M.; Adeel Nawab, R.M.; Rayson, P.: CLEU - A Cross-language english-urdu corpus and benchmark for text reuse experiments (2019) 0.00
```
3.3567476E-4 = product of:
  0.005706471 = sum of:
    0.005706471 = weight(_text_:in in 5299) [ClassicSimilarity], result of:
      0.005706471 = score(doc=5299,freq=10.0), product of:
        0.033961542 = queryWeight, product of:
          1.3602545 = idf(docFreq=30841, maxDocs=44218)
          0.024967048 = queryNorm
        0.16802745 = fieldWeight in 5299, product of:
          3.1622777 = tf(freq=10.0), with freq of:
            10.0 = termFreq=10.0
          1.3602545 = idf(docFreq=30841, maxDocs=44218)
          0.0390625 = fieldNorm(doc=5299)
  0.05882353 = coord(1/17)
```
Abstract

Text reuse is becoming a serious issue in many fields and research shows that it is much harder to detect when it occurs across languages. The recent rise in multi-lingual content on the Web has increased cross-language text reuse to an unprecedented scale. Although researchers have proposed methods to detect it, one major drawback is the unavailability of large-scale gold standard evaluation resources built on real cases. To overcome this problem, we propose a cross-language sentence/passage level text reuse corpus for the English-Urdu language pair. The Cross-Language English-Urdu Corpus (CLEU) has source text in English whereas the derived text is in Urdu. It contains in total 3,235 sentence/passage pairs manually tagged into three categories that is near copy, paraphrased copy, and independently written. Further, as a second contribution, we evaluate the Translation plus Mono-lingual Analysis method using three sets of experiments on the proposed dataset to highlight its usefulness. Evaluation results (f1=0.732 binary, f1=0.552 ternary classification) indicate that it is harder to detect cross-language real cases of text reuse, especially when the language pairs have unrelated scripts. The corpus is a useful benchmark resource for the future development and assessment of cross-language text reuse detection systems for the English-Urdu language pair.