Search (116 results, page 1 of 6)

Robin, J.; McKeown, K.: Empirically designing and evaluating a new revision-based model for summary generation (1996) 0.06

0.060332857 = product of:
  0.120665714 = sum of:
    0.120665714 = sum of:
      0.058716543 = weight(_text_:k in 6751) [ClassicSimilarity], result of:
        0.058716543 = score(doc=6751,freq=2.0), product of:
          0.18609051 = queryWeight, product of:
            3.569778 = idf(docFreq=3384, maxDocs=44218)
            0.052129436 = queryNorm
          0.31552678 = fieldWeight in 6751, product of:
            1.4142135 = tf(freq=2.0), with freq of:
              2.0 = termFreq=2.0
            3.569778 = idf(docFreq=3384, maxDocs=44218)
            0.0625 = fieldNorm(doc=6751)
      0.0054466184 = weight(_text_:s in 6751) [ClassicSimilarity], result of:
        0.0054466184 = score(doc=6751,freq=2.0), product of:
          0.056677084 = queryWeight, product of:
            1.0872376 = idf(docFreq=40523, maxDocs=44218)
            0.052129436 = queryNorm
          0.09609913 = fieldWeight in 6751, product of:
            1.4142135 = tf(freq=2.0), with freq of:
              2.0 = termFreq=2.0
            1.0872376 = idf(docFreq=40523, maxDocs=44218)
            0.0625 = fieldNorm(doc=6751)
      0.05650255 = weight(_text_:22 in 6751) [ClassicSimilarity], result of:
        0.05650255 = score(doc=6751,freq=2.0), product of:
          0.1825484 = queryWeight, product of:
            3.5018296 = idf(docFreq=3622, maxDocs=44218)
            0.052129436 = queryNorm
          0.30952093 = fieldWeight in 6751, product of:
            1.4142135 = tf(freq=2.0), with freq of:
              2.0 = termFreq=2.0
            3.5018296 = idf(docFreq=3622, maxDocs=44218)
            0.0625 = fieldNorm(doc=6751)
  0.5 = coord(1/2)

Date: 6. 3.1997 16:22:15
Source: Artificial intelligence. 85(1996) nos.1/2, S.135-179

Automatic summarizing : introduction (1995) 0.04

0.041711923 = product of:
  0.083423845 = sum of:
    0.083423845 = product of:
      0.12513576 = sum of:
        0.117433086 = weight(_text_:k in 626) [ClassicSimilarity], result of:
          0.117433086 = score(doc=626,freq=8.0), product of:
            0.18609051 = queryWeight, product of:
              3.569778 = idf(docFreq=3384, maxDocs=44218)
              0.052129436 = queryNorm
            0.63105357 = fieldWeight in 626, product of:
              2.828427 = tf(freq=8.0), with freq of:
                8.0 = termFreq=8.0
              3.569778 = idf(docFreq=3384, maxDocs=44218)
              0.0625 = fieldNorm(doc=626)
        0.007702682 = weight(_text_:s in 626) [ClassicSimilarity], result of:
          0.007702682 = score(doc=626,freq=4.0), product of:
            0.056677084 = queryWeight, product of:
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.052129436 = queryNorm
            0.1359047 = fieldWeight in 626, product of:
              2.0 = tf(freq=4.0), with freq of:
                4.0 = termFreq=4.0
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.0625 = fieldNorm(doc=626)
      0.6666667 = coord(2/3)
  0.5 = coord(1/2)

Content: Enthält u.a. Beiträge von: J. BATEMAN u. E. TEICH; R. BRANDOW, K. MITZE u. L.F. RAU; B. ENDRES-NIGGEMEYER, E. MAIER u. A. SIGEL; M.T. MAYBURY; K. McKEOWN, J. ROBIN u. K. KUKICH; A. ROTHKEGEL
Editor: Sparck Jones, K. u. B. Endres-Niggemeyer
Source: Information processing and management. 31(1995) no.5, S.625-630
Type: s

Wu, Y.-f.B.; Li, Q.; Bot, R.S.; Chen, X.: Finding nuggets in documents : a machine learning approach (2006) 0.04

0.040774576 = sum of:
  0.01496242 = product of:
    0.05984968 = sum of:
      0.05984968 = weight(_text_:authors in 5290) [ClassicSimilarity], result of:
        0.05984968 = score(doc=5290,freq=2.0), product of:
          0.23764841 = queryWeight, product of:
            4.558814 = idf(docFreq=1258, maxDocs=44218)
            0.052129436 = queryNorm
          0.25184128 = fieldWeight in 5290, product of:
            1.4142135 = tf(freq=2.0), with freq of:
              2.0 = termFreq=2.0
            4.558814 = idf(docFreq=1258, maxDocs=44218)
            0.0390625 = fieldNorm(doc=5290)
    0.25 = coord(1/4)
  0.025812155 = product of:
    0.03871823 = sum of:
      0.0034041367 = weight(_text_:s in 5290) [ClassicSimilarity], result of:
        0.0034041367 = score(doc=5290,freq=2.0), product of:
          0.056677084 = queryWeight, product of:
            1.0872376 = idf(docFreq=40523, maxDocs=44218)
            0.052129436 = queryNorm
          0.060061958 = fieldWeight in 5290, product of:
            1.4142135 = tf(freq=2.0), with freq of:
              2.0 = termFreq=2.0
            1.0872376 = idf(docFreq=40523, maxDocs=44218)
            0.0390625 = fieldNorm(doc=5290)
      0.035314094 = weight(_text_:22 in 5290) [ClassicSimilarity], result of:
        0.035314094 = score(doc=5290,freq=2.0), product of:
          0.1825484 = queryWeight, product of:
            3.5018296 = idf(docFreq=3622, maxDocs=44218)
            0.052129436 = queryNorm
          0.19345059 = fieldWeight in 5290, product of:
            1.4142135 = tf(freq=2.0), with freq of:
              2.0 = termFreq=2.0
            3.5018296 = idf(docFreq=3622, maxDocs=44218)
            0.0390625 = fieldNorm(doc=5290)
    0.6666667 = coord(2/3)

Abstract: Document keyphrases provide a concise summary of a document's content, offering semantic metadata summarizing a document. They can be used in many applications related to knowledge management and text mining, such as automatic text summarization, development of search engines, document clustering, document classification, thesaurus construction, and browsing interfaces. Because only a small portion of documents have keyphrases assigned by authors, and it is time-consuming and costly to manually assign keyphrases to documents, it is necessary to develop an algorithm to automatically generate keyphrases for documents. This paper describes a Keyphrase Identification Program (KIP), which extracts document keyphrases by using prior positive samples of human identified phrases to assign weights to the candidate keyphrases. The logic of our algorithm is: The more keywords a candidate keyphrase contains and the more significant these keywords are, the more likely this candidate phrase is a keyphrase. KIP's learning function can enrich the glossary database by automatically adding new identified keyphrases to the database. KIP's personalization feature will let the user build a glossary database specifically suitable for the area of his/her interest. The evaluation results show that KIP's performance is better than the systems we compared to and that the learning function is effective.
Date: 22. 7.2006 17:25:48
Source: Journal of the American Society for Information Science and Technology. 57(2006) no.6, S.740-752

McKeown, K.; Robin, J.; Kukich, K.: Generating concise natural language summaries (1995) 0.04

0.03686848 = product of:
  0.07373696 = sum of:
    0.07373696 = product of:
      0.11060543 = sum of:
        0.10379716 = weight(_text_:k in 2932) [ClassicSimilarity], result of:
          0.10379716 = score(doc=2932,freq=4.0), product of:
            0.18609051 = queryWeight, product of:
              3.569778 = idf(docFreq=3384, maxDocs=44218)
              0.052129436 = queryNorm
            0.5577778 = fieldWeight in 2932, product of:
              2.0 = tf(freq=4.0), with freq of:
                4.0 = termFreq=4.0
              3.569778 = idf(docFreq=3384, maxDocs=44218)
              0.078125 = fieldNorm(doc=2932)
        0.0068082735 = weight(_text_:s in 2932) [ClassicSimilarity], result of:
          0.0068082735 = score(doc=2932,freq=2.0), product of:
            0.056677084 = queryWeight, product of:
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.052129436 = queryNorm
            0.120123915 = fieldWeight in 2932, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.078125 = fieldNorm(doc=2932)
      0.6666667 = coord(2/3)
  0.5 = coord(1/2)

Source: Information processing and management. 31(1995) no.5, S.703-733

Atanassova, I.; Bertin, M.; Larivière, V.: On the composition of scientific abstracts (2016) 0.03
```
0.031529564 = sum of:
  0.02992484 = product of:
    0.11969936 = sum of:
      0.11969936 = weight(_text_:authors in 3028) [ClassicSimilarity], result of:
        0.11969936 = score(doc=3028,freq=8.0), product of:
          0.23764841 = queryWeight, product of:
            4.558814 = idf(docFreq=1258, maxDocs=44218)
            0.052129436 = queryNorm
          0.50368255 = fieldWeight in 3028, product of:
            2.828427 = tf(freq=8.0), with freq of:
              8.0 = termFreq=8.0
            4.558814 = idf(docFreq=1258, maxDocs=44218)
            0.0390625 = fieldNorm(doc=3028)
    0.25 = coord(1/4)
  0.0016047254 = product of:
    0.004814176 = sum of:
      0.004814176 = weight(_text_:s in 3028) [ClassicSimilarity], result of:
        0.004814176 = score(doc=3028,freq=4.0), product of:
          0.056677084 = queryWeight, product of:
            1.0872376 = idf(docFreq=40523, maxDocs=44218)
            0.052129436 = queryNorm
          0.08494043 = fieldWeight in 3028, product of:
            2.0 = tf(freq=4.0), with freq of:
              4.0 = termFreq=4.0
            1.0872376 = idf(docFreq=40523, maxDocs=44218)
            0.0390625 = fieldNorm(doc=3028)
    0.33333334 = coord(1/3)
```
Abstract

Purpose - Scientific abstracts reproduce only part of the information and the complexity of argumentation in a scientific article. The purpose of this paper provides a first analysis of the similarity between the text of scientific abstracts and the body of articles, using sentences as the basic textual unit. It contributes to the understanding of the structure of abstracts. Design/methodology/approach - Using sentence-based similarity metrics, the authors quantify the phenomenon of text re-use in abstracts and examine the positions of the sentences that are similar to sentences in abstracts in the introduction, methods, results and discussion structure, using a corpus of over 85,000 research articles published in the seven Public Library of Science journals. Findings - The authors provide evidence that 84 percent of abstract have at least one sentence in common with the body of the paper. Studying the distributions of sentences in the body of the articles that are re-used in abstracts, the authors show that there exists a strong relation between the rhetorical structure of articles and the zones that authors re-use when writing abstracts, with sentences mainly coming from the beginning of the introduction and the end of the conclusion. Originality/value - Scientific abstracts contain what is considered by the author(s) as information that best describe documents' content. This is a first study that examines the relation between the contents of abstracts and the rhetorical structure of scientific articles. The work might provide new insight for improving automatic abstracting tools as well as information retrieval approaches, in which text organization and structure are important features.

Source

Journal of documentation. 72(2016) no.4, S.636-647

Johnson, F.C.; Paice, C.D.; Black, W.J.; Neal, A.P.: ¬The application of linguistic processing to automatic abstract generation (1993) 0.03

0.027674675 = product of:
  0.05534935 = sum of:
    0.05534935 = product of:
      0.083024025 = sum of:
        0.07339568 = weight(_text_:k in 2290) [ClassicSimilarity], result of:
          0.07339568 = score(doc=2290,freq=2.0), product of:
            0.18609051 = queryWeight, product of:
              3.569778 = idf(docFreq=3384, maxDocs=44218)
              0.052129436 = queryNorm
            0.39440846 = fieldWeight in 2290, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              3.569778 = idf(docFreq=3384, maxDocs=44218)
              0.078125 = fieldNorm(doc=2290)
        0.009628352 = weight(_text_:s in 2290) [ClassicSimilarity], result of:
          0.009628352 = score(doc=2290,freq=4.0), product of:
            0.056677084 = queryWeight, product of:
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.052129436 = queryNorm
            0.16988087 = fieldWeight in 2290, product of:
              2.0 = tf(freq=4.0), with freq of:
                4.0 = termFreq=4.0
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.078125 = fieldNorm(doc=2290)
      0.6666667 = coord(2/3)
  0.5 = coord(1/2)

Footnote: Wiederabgedruckt in: Readings in information retrieval. Ed.: K. Sparck Jones u. P. Willett. San Francisco: Morgan Kaufmann 1997. S.538-552.
Source: Journal of document and text management. 1(1993), S.215-241

Salton, G.; Allan, J.; Buckley, C.; Singhal, A.: Automatic analysis, theme generation, and summarization of machine readable texts (1994) 0.03

0.027674675 = product of:
  0.05534935 = sum of:
    0.05534935 = product of:
      0.083024025 = sum of:
        0.07339568 = weight(_text_:k in 1949) [ClassicSimilarity], result of:
          0.07339568 = score(doc=1949,freq=2.0), product of:
            0.18609051 = queryWeight, product of:
              3.569778 = idf(docFreq=3384, maxDocs=44218)
              0.052129436 = queryNorm
            0.39440846 = fieldWeight in 1949, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              3.569778 = idf(docFreq=3384, maxDocs=44218)
              0.078125 = fieldNorm(doc=1949)
        0.009628352 = weight(_text_:s in 1949) [ClassicSimilarity], result of:
          0.009628352 = score(doc=1949,freq=4.0), product of:
            0.056677084 = queryWeight, product of:
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.052129436 = queryNorm
            0.16988087 = fieldWeight in 1949, product of:
              2.0 = tf(freq=4.0), with freq of:
                4.0 = termFreq=4.0
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.078125 = fieldNorm(doc=1949)
      0.6666667 = coord(2/3)
  0.5 = coord(1/2)

Footnote: Wiederabgedruckt in: Readings in information retrieval. Ed.: K. Sparck Jones u. P. Willett. San Francisco: Morgan Kaufmann 1997. S.478-483.
Source: Science. 264(1994), S.1421-1426

Marsh, E.: ¬A production rule system for message summarisation (1984) 0.03

0.027674675 = product of:
  0.05534935 = sum of:
    0.05534935 = product of:
      0.083024025 = sum of:
        0.07339568 = weight(_text_:k in 1956) [ClassicSimilarity], result of:
          0.07339568 = score(doc=1956,freq=2.0), product of:
            0.18609051 = queryWeight, product of:
              3.569778 = idf(docFreq=3384, maxDocs=44218)
              0.052129436 = queryNorm
            0.39440846 = fieldWeight in 1956, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              3.569778 = idf(docFreq=3384, maxDocs=44218)
              0.078125 = fieldNorm(doc=1956)
        0.009628352 = weight(_text_:s in 1956) [ClassicSimilarity], result of:
          0.009628352 = score(doc=1956,freq=4.0), product of:
            0.056677084 = queryWeight, product of:
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.052129436 = queryNorm
            0.16988087 = fieldWeight in 1956, product of:
              2.0 = tf(freq=4.0), with freq of:
                4.0 = termFreq=4.0
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.078125 = fieldNorm(doc=1956)
      0.6666667 = coord(2/3)
  0.5 = coord(1/2)

Footnote: Wiederabgedruckt in: Readings in information retrieval. Ed.: K. Sparck Jones u. P. Willett. San Francisco: Morgan Kaufmann 1997. S.534-537.
Pages: S.243-246

Ouyang, Y.; Li, W.; Li, S.; Lu, Q.: Intertopic information mining for query-based summarization (2010) 0.02
```
0.022764783 = sum of:
  0.021160059 = product of:
    0.084640235 = sum of:
      0.084640235 = weight(_text_:authors in 3459) [ClassicSimilarity], result of:
        0.084640235 = score(doc=3459,freq=4.0), product of:
          0.23764841 = queryWeight, product of:
            4.558814 = idf(docFreq=1258, maxDocs=44218)
            0.052129436 = queryNorm
          0.35615736 = fieldWeight in 3459, product of:
            2.0 = tf(freq=4.0), with freq of:
              4.0 = termFreq=4.0
            4.558814 = idf(docFreq=1258, maxDocs=44218)
            0.0390625 = fieldNorm(doc=3459)
    0.25 = coord(1/4)
  0.0016047254 = product of:
    0.004814176 = sum of:
      0.004814176 = weight(_text_:s in 3459) [ClassicSimilarity], result of:
        0.004814176 = score(doc=3459,freq=4.0), product of:
          0.056677084 = queryWeight, product of:
            1.0872376 = idf(docFreq=40523, maxDocs=44218)
            0.052129436 = queryNorm
          0.08494043 = fieldWeight in 3459, product of:
            2.0 = tf(freq=4.0), with freq of:
              4.0 = termFreq=4.0
            1.0872376 = idf(docFreq=40523, maxDocs=44218)
            0.0390625 = fieldNorm(doc=3459)
    0.33333334 = coord(1/3)
```
Abstract

In this article, the authors address the problem of sentence ranking in summarization. Although most existing summarization approaches are concerned with the information embodied in a particular topic (including a set of documents and an associated query) for sentence ranking, they propose a novel ranking approach that incorporates intertopic information mining. Intertopic information, in contrast to intratopic information, is able to reveal pairwise topic relationships and thus can be considered as the bridge across different topics. In this article, the intertopic information is used for transferring word importance learned from known topics to unknown topics under a learning-based summarization framework. To mine this information, the authors model the topic relationship by clustering all the words in both known and unknown topics according to various kinds of word conceptual labels, which indicate the roles of the words in the topic. Based on the mined relationships, we develop a probabilistic model using manually generated summaries provided for known topics to predict ranking scores for sentences in unknown topics. A series of experiments have been conducted on the Document Understanding Conference (DUC) 2006 data set. The evaluation results show that intertopic information is indeed effective for sentence ranking and the resultant summarization system performs comparably well to the best-performing DUC participating systems on the same data set.

Source

Journal of the American Society for Information Science and Technology. 61(2010) no.5, S.1062-1072

Brandow, R.; Mitze, K.; Rau, L.F.: Automatic condensation of electronic publications by sentence selection (1995) 0.02

0.021387722 = product of:
  0.042775445 = sum of:
    0.042775445 = product of:
      0.06416316 = sum of:
        0.058716543 = weight(_text_:k in 2929) [ClassicSimilarity], result of:
          0.058716543 = score(doc=2929,freq=2.0), product of:
            0.18609051 = queryWeight, product of:
              3.569778 = idf(docFreq=3384, maxDocs=44218)
              0.052129436 = queryNorm
            0.31552678 = fieldWeight in 2929, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              3.569778 = idf(docFreq=3384, maxDocs=44218)
              0.0625 = fieldNorm(doc=2929)
        0.0054466184 = weight(_text_:s in 2929) [ClassicSimilarity], result of:
          0.0054466184 = score(doc=2929,freq=2.0), product of:
            0.056677084 = queryWeight, product of:
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.052129436 = queryNorm
            0.09609913 = fieldWeight in 2929, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.0625 = fieldNorm(doc=2929)
      0.6666667 = coord(2/3)
  0.5 = coord(1/2)

Source: Information processing and management. 31(1995) no.5, S.675-685

Sparck Jones, K.; Endres-Niggemeyer, B.: Introduction: automatic summarizing (1995) 0.02

0.021387722 = product of:
  0.042775445 = sum of:
    0.042775445 = product of:
      0.06416316 = sum of:
        0.058716543 = weight(_text_:k in 2931) [ClassicSimilarity], result of:
          0.058716543 = score(doc=2931,freq=2.0), product of:
            0.18609051 = queryWeight, product of:
              3.569778 = idf(docFreq=3384, maxDocs=44218)
              0.052129436 = queryNorm
            0.31552678 = fieldWeight in 2931, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              3.569778 = idf(docFreq=3384, maxDocs=44218)
              0.0625 = fieldNorm(doc=2931)
        0.0054466184 = weight(_text_:s in 2931) [ClassicSimilarity], result of:
          0.0054466184 = score(doc=2931,freq=2.0), product of:
            0.056677084 = queryWeight, product of:
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.052129436 = queryNorm
            0.09609913 = fieldWeight in 2931, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.0625 = fieldNorm(doc=2931)
      0.6666667 = coord(2/3)
  0.5 = coord(1/2)

Source: Information processing and management. 31(1995) no.5, S.625-630

Ahmad, K.: Text summarisation : the role of lexical cohesion analysis (1995) 0.02

0.021387722 = product of:
  0.042775445 = sum of:
    0.042775445 = product of:
      0.06416316 = sum of:
        0.058716543 = weight(_text_:k in 5795) [ClassicSimilarity], result of:
          0.058716543 = score(doc=5795,freq=2.0), product of:
            0.18609051 = queryWeight, product of:
              3.569778 = idf(docFreq=3384, maxDocs=44218)
              0.052129436 = queryNorm
            0.31552678 = fieldWeight in 5795, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              3.569778 = idf(docFreq=3384, maxDocs=44218)
              0.0625 = fieldNorm(doc=5795)
        0.0054466184 = weight(_text_:s in 5795) [ClassicSimilarity], result of:
          0.0054466184 = score(doc=5795,freq=2.0), product of:
            0.056677084 = queryWeight, product of:
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.052129436 = queryNorm
            0.09609913 = fieldWeight in 5795, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.0625 = fieldNorm(doc=5795)
      0.6666667 = coord(2/3)
  0.5 = coord(1/2)

Source: New review of document and text management. 1995, no.1, S.321-335

Gomez, J.; Allen, K.; Matney, M.; Awopetu, T.; Shafer, S.: Experimenting with a machine generated annotations pipeline (2020) 0.02

0.021387722 = product of:
  0.042775445 = sum of:
    0.042775445 = product of:
      0.06416316 = sum of:
        0.058716543 = weight(_text_:k in 657) [ClassicSimilarity], result of:
          0.058716543 = score(doc=657,freq=2.0), product of:
            0.18609051 = queryWeight, product of:
              3.569778 = idf(docFreq=3384, maxDocs=44218)
              0.052129436 = queryNorm
            0.31552678 = fieldWeight in 657, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              3.569778 = idf(docFreq=3384, maxDocs=44218)
              0.0625 = fieldNorm(doc=657)
        0.0054466184 = weight(_text_:s in 657) [ClassicSimilarity], result of:
          0.0054466184 = score(doc=657,freq=2.0), product of:
            0.056677084 = queryWeight, product of:
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.052129436 = queryNorm
            0.09609913 = fieldWeight in 657, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.0625 = fieldNorm(doc=657)
      0.6666667 = coord(2/3)
  0.5 = coord(1/2)

Goh, A.; Hui, S.C.: TES: a text extraction system (1996) 0.02

0.020649724 = product of:
  0.041299447 = sum of:
    0.041299447 = product of:
      0.06194917 = sum of:
        0.0054466184 = weight(_text_:s in 6599) [ClassicSimilarity], result of:
          0.0054466184 = score(doc=6599,freq=2.0), product of:
            0.056677084 = queryWeight, product of:
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.052129436 = queryNorm
            0.09609913 = fieldWeight in 6599, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.0625 = fieldNorm(doc=6599)
        0.05650255 = weight(_text_:22 in 6599) [ClassicSimilarity], result of:
          0.05650255 = score(doc=6599,freq=2.0), product of:
            0.1825484 = queryWeight, product of:
              3.5018296 = idf(docFreq=3622, maxDocs=44218)
              0.052129436 = queryNorm
            0.30952093 = fieldWeight in 6599, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              3.5018296 = idf(docFreq=3622, maxDocs=44218)
              0.0625 = fieldNorm(doc=6599)
      0.6666667 = coord(2/3)
  0.5 = coord(1/2)

Date: 26. 2.1997 10:22:43
Source: Microcomputers for information management. 13(1996) no.1, S.41-55

Jones, P.A.; Bradbeer, P.V.G.: Discovery of optimal weights in a concept selection system (1996) 0.02

0.020649724 = product of:
  0.041299447 = sum of:
    0.041299447 = product of:
      0.06194917 = sum of:
        0.0054466184 = weight(_text_:s in 6974) [ClassicSimilarity], result of:
          0.0054466184 = score(doc=6974,freq=2.0), product of:
            0.056677084 = queryWeight, product of:
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.052129436 = queryNorm
            0.09609913 = fieldWeight in 6974, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.0625 = fieldNorm(doc=6974)
        0.05650255 = weight(_text_:22 in 6974) [ClassicSimilarity], result of:
          0.05650255 = score(doc=6974,freq=2.0), product of:
            0.1825484 = queryWeight, product of:
              3.5018296 = idf(docFreq=3622, maxDocs=44218)
              0.052129436 = queryNorm
            0.30952093 = fieldWeight in 6974, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              3.5018296 = idf(docFreq=3622, maxDocs=44218)
              0.0625 = fieldNorm(doc=6974)
      0.6666667 = coord(2/3)
  0.5 = coord(1/2)

Pages: S.145-153
Source: Information retrieval: new systems and current research. Proceedings of the 16th Research Colloquium of the British Computer Society Information Retrieval Specialist Group, Drymen, Scotland, 22-23 Mar 94. Ed.: R. Leon

Wang, W.; Hwang, D.: Abstraction Assistant : an automatic text abstraction system (2010) 0.02

0.019316558 = sum of:
  0.017954903 = product of:
    0.07181961 = sum of:
      0.07181961 = weight(_text_:authors in 3981) [ClassicSimilarity], result of:
        0.07181961 = score(doc=3981,freq=2.0), product of:
          0.23764841 = queryWeight, product of:
            4.558814 = idf(docFreq=1258, maxDocs=44218)
            0.052129436 = queryNorm
          0.30220953 = fieldWeight in 3981, product of:
            1.4142135 = tf(freq=2.0), with freq of:
              2.0 = termFreq=2.0
            4.558814 = idf(docFreq=1258, maxDocs=44218)
            0.046875 = fieldNorm(doc=3981)
    0.25 = coord(1/4)
  0.0013616546 = product of:
    0.004084964 = sum of:
      0.004084964 = weight(_text_:s in 3981) [ClassicSimilarity], result of:
        0.004084964 = score(doc=3981,freq=2.0), product of:
          0.056677084 = queryWeight, product of:
            1.0872376 = idf(docFreq=40523, maxDocs=44218)
            0.052129436 = queryNorm
          0.072074346 = fieldWeight in 3981, product of:
            1.4142135 = tf(freq=2.0), with freq of:
              2.0 = termFreq=2.0
            1.0872376 = idf(docFreq=40523, maxDocs=44218)
            0.046875 = fieldNorm(doc=3981)
    0.33333334 = coord(1/3)

Abstract: In the interest of standardization and quality assurance, it is desirable for authors and staff of access services to follow the American National Standards Institute (ANSI) guidelines in preparing abstracts. Using the statistical approach an extraction system (the Abstraction Assistant) was developed to generate informative abstracts to meet the ANSI guidelines for structural content elements. The system performance is evaluated by comparing the system-generated abstracts with the author's original abstracts and the manually enhanced system abstracts on three criteria: balance (satisfaction of the ANSI standards), fluency (text coherence), and understandability (clarity). The results suggest that it is possible to use the system output directly without manual modification, but there are issues that need to be addressed in further studies to make the system a better tool.
Source: Journal of the American Society for Information Science and Technology. 61(2010) no.9, S.1790-1799

Lam, W.; Chan, K.; Radev, D.; Saggion, H.; Teufel, S.: Context-based generic cross-lingual retrieval of documents and automated summaries (2005) 0.02

0.016604807 = product of:
  0.033209614 = sum of:
    0.033209614 = product of:
      0.049814418 = sum of:
        0.044037405 = weight(_text_:k in 1965) [ClassicSimilarity], result of:
          0.044037405 = score(doc=1965,freq=2.0), product of:
            0.18609051 = queryWeight, product of:
              3.569778 = idf(docFreq=3384, maxDocs=44218)
              0.052129436 = queryNorm
            0.23664509 = fieldWeight in 1965, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              3.569778 = idf(docFreq=3384, maxDocs=44218)
              0.046875 = fieldNorm(doc=1965)
        0.0057770116 = weight(_text_:s in 1965) [ClassicSimilarity], result of:
          0.0057770116 = score(doc=1965,freq=4.0), product of:
            0.056677084 = queryWeight, product of:
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.052129436 = queryNorm
            0.101928525 = fieldWeight in 1965, product of:
              2.0 = tf(freq=4.0), with freq of:
                4.0 = termFreq=4.0
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.046875 = fieldNorm(doc=1965)
      0.6666667 = coord(2/3)
  0.5 = coord(1/2)

Source: Journal of the American Society for Information Science and Technology. 56(2005) no.2, S.129-139

Wang, S.; Koopman, R.: Embed first, then predict (2019) 0.02

0.016567145 = sum of:
  0.01496242 = product of:
    0.05984968 = sum of:
      0.05984968 = weight(_text_:authors in 5400) [ClassicSimilarity], result of:
        0.05984968 = score(doc=5400,freq=2.0), product of:
          0.23764841 = queryWeight, product of:
            4.558814 = idf(docFreq=1258, maxDocs=44218)
            0.052129436 = queryNorm
          0.25184128 = fieldWeight in 5400, product of:
            1.4142135 = tf(freq=2.0), with freq of:
              2.0 = termFreq=2.0
            4.558814 = idf(docFreq=1258, maxDocs=44218)
            0.0390625 = fieldNorm(doc=5400)
    0.25 = coord(1/4)
  0.0016047254 = product of:
    0.004814176 = sum of:
      0.004814176 = weight(_text_:s in 5400) [ClassicSimilarity], result of:
        0.004814176 = score(doc=5400,freq=4.0), product of:
          0.056677084 = queryWeight, product of:
            1.0872376 = idf(docFreq=40523, maxDocs=44218)
            0.052129436 = queryNorm
          0.08494043 = fieldWeight in 5400, product of:
            2.0 = tf(freq=4.0), with freq of:
              4.0 = termFreq=4.0
            1.0872376 = idf(docFreq=40523, maxDocs=44218)
            0.0390625 = fieldNorm(doc=5400)
    0.33333334 = coord(1/3)

Abstract: Automatic subject prediction is a desirable feature for modern digital library systems, as manual indexing can no longer cope with the rapid growth of digital collections. It is also desirable to be able to identify a small set of entities (e.g., authors, citations, bibliographic records) which are most relevant to a query. This gets more difficult when the amount of data increases dramatically. Data sparsity and model scalability are the major challenges to solving this type of extreme multilabel classification problem automatically. In this paper, we propose to address this problem in two steps: we first embed different types of entities into the same semantic space, where similarity could be computed easily; second, we propose a novel non-parametric method to identify the most relevant entities in addition to direct semantic similarities. We show how effectively this approach predicts even very specialised subjects, which are associated with few documents in the training set and are more problematic for a classifier.
Source: Knowledge organization. 46(2019) no.5, S.364-370

Sparck Jones, K.: Automatic summarising : the state of the art (2007) 0.02

0.01604079 = product of:
  0.03208158 = sum of:
    0.03208158 = product of:
      0.04812237 = sum of:
        0.044037405 = weight(_text_:k in 932) [ClassicSimilarity], result of:
          0.044037405 = score(doc=932,freq=2.0), product of:
            0.18609051 = queryWeight, product of:
              3.569778 = idf(docFreq=3384, maxDocs=44218)
              0.052129436 = queryNorm
            0.23664509 = fieldWeight in 932, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              3.569778 = idf(docFreq=3384, maxDocs=44218)
              0.046875 = fieldNorm(doc=932)
        0.004084964 = weight(_text_:s in 932) [ClassicSimilarity], result of:
          0.004084964 = score(doc=932,freq=2.0), product of:
            0.056677084 = queryWeight, product of:
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.052129436 = queryNorm
            0.072074346 = fieldWeight in 932, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.046875 = fieldNorm(doc=932)
      0.6666667 = coord(2/3)
  0.5 = coord(1/2)

Source: Information processing and management. 43(2007) no.6, S.1449-1481

Nomoto, T.: Discriminative sentence compression with conditional random fields (2007) 0.02

0.01604079 = product of:
  0.03208158 = sum of:
    0.03208158 = product of:
      0.04812237 = sum of:
        0.044037405 = weight(_text_:k in 945) [ClassicSimilarity], result of:
          0.044037405 = score(doc=945,freq=2.0), product of:
            0.18609051 = queryWeight, product of:
              3.569778 = idf(docFreq=3384, maxDocs=44218)
              0.052129436 = queryNorm
            0.23664509 = fieldWeight in 945, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              3.569778 = idf(docFreq=3384, maxDocs=44218)
              0.046875 = fieldNorm(doc=945)
        0.004084964 = weight(_text_:s in 945) [ClassicSimilarity], result of:
          0.004084964 = score(doc=945,freq=2.0), product of:
            0.056677084 = queryWeight, product of:
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.052129436 = queryNorm
            0.072074346 = fieldWeight in 945, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              1.0872376 = idf(docFreq=40523, maxDocs=44218)
              0.046875 = fieldNorm(doc=945)
      0.6666667 = coord(2/3)
  0.5 = coord(1/2)

Abstract: The paper focuses on a particular approach to automatic sentence compression which makes use of a discriminative sequence classifier known as Conditional Random Fields (CRF). We devise several features for CRF that allow it to incorporate information on nonlinear relations among words. Along with that, we address the issue of data paucity by collecting data from RSS feeds available on the Internet, and turning them into training data for use with CRF, drawing on techniques from biology and information retrieval. We also discuss a recursive application of CRF on the syntactic structure of a sentence as a way of improving the readability of the compression it generates. Experiments found that our approach works reasonably well compared to the state-of-the-art system [Knight, K., & Marcu, D. (2002). Summarization beyond sentence extraction: A probabilistic approach to sentence compression. Artificial Intelligence 139, 91-107.].
Source: Information processing and management. 43(2007) no.6, S.1571-1587

Search (116 results, page 1 of 6)

Authors

Years

Languages

Types

Themes

Subjects