Search (4 results, page 1 of 1)

  • × theme_ss:"Automatisches Abstracting"
  • × theme_ss:"Automatisches Indexieren"
  1. Wang, S.; Koopman, R.: Embed first, then predict (2019) 0.01
    0.0078013875 = sum of:
      0.005581817 = product of:
        0.044654537 = sum of:
          0.044654537 = weight(_text_:authors in 5400) [ClassicSimilarity], result of:
            0.044654537 = score(doc=5400,freq=2.0), product of:
              0.17731223 = queryWeight, product of:
                4.558814 = idf(docFreq=1258, maxDocs=44218)
                0.038894374 = queryNorm
              0.25184128 = fieldWeight in 5400, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                4.558814 = idf(docFreq=1258, maxDocs=44218)
                0.0390625 = fieldNorm(doc=5400)
        0.125 = coord(1/8)
      0.0022195703 = product of:
        0.0044391407 = sum of:
          0.0044391407 = weight(_text_:e in 5400) [ClassicSimilarity], result of:
            0.0044391407 = score(doc=5400,freq=2.0), product of:
              0.055905603 = queryWeight, product of:
                1.43737 = idf(docFreq=28552, maxDocs=44218)
                0.038894374 = queryNorm
              0.07940422 = fieldWeight in 5400, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                1.43737 = idf(docFreq=28552, maxDocs=44218)
                0.0390625 = fieldNorm(doc=5400)
        0.5 = coord(1/2)
    
    Abstract
    Automatic subject prediction is a desirable feature for modern digital library systems, as manual indexing can no longer cope with the rapid growth of digital collections. It is also desirable to be able to identify a small set of entities (e.g., authors, citations, bibliographic records) which are most relevant to a query. This gets more difficult when the amount of data increases dramatically. Data sparsity and model scalability are the major challenges to solving this type of extreme multilabel classification problem automatically. In this paper, we propose to address this problem in two steps: we first embed different types of entities into the same semantic space, where similarity could be computed easily; second, we propose a novel non-parametric method to identify the most relevant entities in addition to direct semantic similarities. We show how effectively this approach predicts even very specialised subjects, which are associated with few documents in the training set and are more problematic for a classifier.
    Language
    e
  2. Salton, G.; Allan, J.; Buckley, C.; Singhal, A.: Automatic analysis, theme generation, and summarization of machine readable texts (1994) 0.00
    0.0022195703 = product of:
      0.0044391407 = sum of:
        0.0044391407 = product of:
          0.008878281 = sum of:
            0.008878281 = weight(_text_:e in 1949) [ClassicSimilarity], result of:
              0.008878281 = score(doc=1949,freq=2.0), product of:
                0.055905603 = queryWeight, product of:
                  1.43737 = idf(docFreq=28552, maxDocs=44218)
                  0.038894374 = queryNorm
                0.15880844 = fieldWeight in 1949, product of:
                  1.4142135 = tf(freq=2.0), with freq of:
                    2.0 = termFreq=2.0
                  1.43737 = idf(docFreq=28552, maxDocs=44218)
                  0.078125 = fieldNorm(doc=1949)
          0.5 = coord(1/2)
      0.5 = coord(1/2)
    
    Language
    e
  3. Moens, M.F.: Automatic indexing and abstracting of document texts (2000) 0.00
    0.0022195703 = product of:
      0.0044391407 = sum of:
        0.0044391407 = product of:
          0.008878281 = sum of:
            0.008878281 = weight(_text_:e in 6892) [ClassicSimilarity], result of:
              0.008878281 = score(doc=6892,freq=2.0), product of:
                0.055905603 = queryWeight, product of:
                  1.43737 = idf(docFreq=28552, maxDocs=44218)
                  0.038894374 = queryNorm
                0.15880844 = fieldWeight in 6892, product of:
                  1.4142135 = tf(freq=2.0), with freq of:
                    2.0 = termFreq=2.0
                  1.43737 = idf(docFreq=28552, maxDocs=44218)
                  0.078125 = fieldNorm(doc=6892)
          0.5 = coord(1/2)
      0.5 = coord(1/2)
    
    Language
    e
  4. Jones, S.; Paynter, G.W.: Automatic extractionof document keyphrases for use in digital libraries : evaluations and applications (2002) 0.00
    0.0011097852 = product of:
      0.0022195703 = sum of:
        0.0022195703 = product of:
          0.0044391407 = sum of:
            0.0044391407 = weight(_text_:e in 601) [ClassicSimilarity], result of:
              0.0044391407 = score(doc=601,freq=2.0), product of:
                0.055905603 = queryWeight, product of:
                  1.43737 = idf(docFreq=28552, maxDocs=44218)
                  0.038894374 = queryNorm
                0.07940422 = fieldWeight in 601, product of:
                  1.4142135 = tf(freq=2.0), with freq of:
                    2.0 = termFreq=2.0
                  1.43737 = idf(docFreq=28552, maxDocs=44218)
                  0.0390625 = fieldNorm(doc=601)
          0.5 = coord(1/2)
      0.5 = coord(1/2)
    
    Language
    e