Search (6 results, page 1 of 1)

Järvelin, A.; Keskustalo, H.; Sormunen, E.; Saastamoinen, M.; Kettunen, K.: Information retrieval from historical newspaper collections in highly inflectional languages : a query expansion approach (2016) 0.02
```
0.015536643 = product of:
  0.07768321 = sum of:
    0.07768321 = weight(_text_:index in 3223) [ClassicSimilarity], result of:
      0.07768321 = score(doc=3223,freq=6.0), product of:
        0.18579477 = queryWeight, product of:
          4.369764 = idf(docFreq=1520, maxDocs=44218)
          0.04251826 = queryNorm
        0.418113 = fieldWeight in 3223, product of:
          2.4494898 = tf(freq=6.0), with freq of:
            6.0 = termFreq=6.0
          4.369764 = idf(docFreq=1520, maxDocs=44218)
          0.0390625 = fieldNorm(doc=3223)
  0.2 = coord(1/5)
```
Abstract

The aim of the study was to test whether query expansion by approximate string matching methods is beneficial in retrieval from historical newspaper collections in a language rich with compounds and inflectional forms (Finnish). First, approximate string matching methods were used to generate lists of index words most similar to contemporary query terms in a digitized newspaper collection from the 1800s. Top index word variants were categorized to estimate the appropriate query expansion ranges in the retrieval test. Second, the effectiveness of approximate string matching methods, automatically generated inflectional forms, and their combinations were measured in a Cranfield-style test. Finally, a detailed topic-level analysis of test results was conducted. In the index of historical newspaper collection the occurrences of a word typically spread to many linguistic and historical variants along with optical character recognition (OCR) errors. All query expansion methods improved the baseline results. Extensive expansion of around 30 variants for each query word was required to achieve the highest performance improvement. Query expansion based on approximate string matching was superior to using the inflectional forms of the query words, showing that coverage of the different types of variation is more important than precision in handling one type of variation.

Sormunen, E.: Free-text searching in full-text databases : probing system limits (1993) 0.01

0.013047772 = product of:
  0.06523886 = sum of:
    0.06523886 = weight(_text_:system in 7120) [ClassicSimilarity], result of:
      0.06523886 = score(doc=7120,freq=2.0), product of:
        0.13391352 = queryWeight, product of:
          3.1495528 = idf(docFreq=5152, maxDocs=44218)
          0.04251826 = queryNorm
        0.4871716 = fieldWeight in 7120, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.1495528 = idf(docFreq=5152, maxDocs=44218)
          0.109375 = fieldNorm(doc=7120)
  0.2 = coord(1/5)

Sormunen, E.; Kekäläinen, J.; Koivisto, J.; Järvelin, K.: Document text characteristics affect the ranking of the most relevant documents by expanded structured queries (2001) 0.01
```
0.0065901205 = product of:
  0.032950602 = sum of:
    0.032950602 = weight(_text_:system in 4487) [ClassicSimilarity], result of:
      0.032950602 = score(doc=4487,freq=4.0), product of:
        0.13391352 = queryWeight, product of:
          3.1495528 = idf(docFreq=5152, maxDocs=44218)
          0.04251826 = queryNorm
        0.24605882 = fieldWeight in 4487, product of:
          2.0 = tf(freq=4.0), with freq of:
            4.0 = termFreq=4.0
          3.1495528 = idf(docFreq=5152, maxDocs=44218)
          0.0390625 = fieldNorm(doc=4487)
  0.2 = coord(1/5)
```
Abstract

The increasing flood of documentary information through the Internet and other information sources challenges the developers of information retrieval systems. It is not enough that an IR system is able to make a distinction between relevant and non-relevant documents. The reduction of information overload requires that IR systems provide the capability of screening the most valuable documents out of the mass of potentially or marginally relevant documents. This paper introduces a new concept-based method to analyse the text characteristics of documents at varying relevance levels. The results of the document analysis were applied in an experiment on query expansion (QE) in a probabilistic IR system. Statistical differences in textual characteristics of highly relevant and less relevant documents were investigated by applying a facet analysis technique. In highly relevant documents a larger number of aspects of the request were discussed, searchable expressions for the aspects were distributed over a larger set of text paragraphs, and a larger set of unique expressions were used per aspect than in marginally relevant documents. A query expansion experiment verified that the findings of the text analysis can be exploited in formulating more effective queries for best match retrieval in the search for highly relevant documents. The results revealed that expanded queries with concept-based structures performed better than unexpanded queries or Ñnatural languageÒ queries. Further, it was shown that highly relevant documents benefit essentially more from the concept-based QE in ranking than marginally relevant documents.
Vakkari, P.; Jones, S.; MacFarlane, A.; Sormunen, E.: Query exhaustivity, relevance feedback and search success in automatic and interactive query expansion (2004) 0.01
```
0.0055919024 = product of:
  0.027959513 = sum of:
    0.027959513 = weight(_text_:system in 4435) [ClassicSimilarity], result of:
      0.027959513 = score(doc=4435,freq=2.0), product of:
        0.13391352 = queryWeight, product of:
          3.1495528 = idf(docFreq=5152, maxDocs=44218)
          0.04251826 = queryNorm
        0.20878783 = fieldWeight in 4435, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.1495528 = idf(docFreq=5152, maxDocs=44218)
          0.046875 = fieldNorm(doc=4435)
  0.2 = coord(1/5)
```
Abstract

This study explored how the expression of search facets and relevance feedback (RF) by users was related to search success in interactive and automatic query expansion in the course of the search process. Search success was measured both in the number of relevant documents retrieved, whether identified by users or not. Research design consisted of 26 users searching for four TREC topics in Okapi IR system, half of the searchers using interactive and half automatic query expansion based on RF. The search logs were recorded, and the users filled in questionnaires for each topic concerning various features of searching. The results showed that the exhaustivity of the query was the most significant predictor of search success. Interactive expansion led to better search success than automatic expansion if all retrieved relevant items were counted, but there was no difference between the methods if only those items recognised relevant by users were observed. The analysis showed that the difference was facilitated by the liberal relevance criterion used in TREC not favouring highly relevant documents in evaluation.

Alkula, R.; Sormunen, E.: Problems and guidelines for database descriptions (1989) 0.00

0.0027127003 = product of:
  0.013563501 = sum of:
    0.013563501 = product of:
      0.0406905 = sum of:
        0.0406905 = weight(_text_:29 in 2397) [ClassicSimilarity], result of:
          0.0406905 = score(doc=2397,freq=2.0), product of:
            0.14956595 = queryWeight, product of:
              3.5176873 = idf(docFreq=3565, maxDocs=44218)
              0.04251826 = queryNorm
            0.27205724 = fieldWeight in 2397, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              3.5176873 = idf(docFreq=3565, maxDocs=44218)
              0.0546875 = fieldNorm(doc=2397)
      0.33333334 = coord(1/3)
  0.2 = coord(1/5)

Pages: S.29-37

Järvelin, K.; Kristensen, J.; Niemi, T.; Sormunen, E.; Keskustalo, H.: ¬A deductive data model for query expansion (1996) 0.00

0.0023042548 = product of:
  0.011521274 = sum of:
    0.011521274 = product of:
      0.03456382 = sum of:
        0.03456382 = weight(_text_:22 in 2230) [ClassicSimilarity], result of:
          0.03456382 = score(doc=2230,freq=2.0), product of:
            0.1488917 = queryWeight, product of:
              3.5018296 = idf(docFreq=3622, maxDocs=44218)
              0.04251826 = queryNorm
            0.23214069 = fieldWeight in 2230, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              3.5018296 = idf(docFreq=3622, maxDocs=44218)
              0.046875 = fieldNorm(doc=2230)
      0.33333334 = coord(1/3)
  0.2 = coord(1/5)

Source: Proceedings of the 19th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (ACM SIGIR '96), Zürich, Switzerland, August 18-22, 1996. Eds.: H.P. Frei et al

Search (6 results, page 1 of 1)

Authors

Years

Themes