Search (69 results, page 1 of 4)

Hotho, A.; Bloehdorn, S.: Data Mining 2004 : Text classification by boosting weak learners based on terms and concepts (2004) 0.12

0.11659854 = product of:
  0.29149634 = sum of:
    0.24901254 = weight(_text_:3a in 562) [ClassicSimilarity], result of:
      0.24901254 = score(doc=562,freq=2.0), product of:
        0.4430686 = queryWeight, product of:
          8.478011 = idf(docFreq=24, maxDocs=44218)
          0.052260913 = queryNorm
        0.56201804 = fieldWeight in 562, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          8.478011 = idf(docFreq=24, maxDocs=44218)
          0.046875 = fieldNorm(doc=562)
    0.042483795 = weight(_text_:22 in 562) [ClassicSimilarity], result of:
      0.042483795 = score(doc=562,freq=2.0), product of:
        0.18300882 = queryWeight, product of:
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.052260913 = queryNorm
        0.23214069 = fieldWeight in 562, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.046875 = fieldNorm(doc=562)
  0.4 = coord(2/5)

Content: Vgl.: http://www.google.de/url?sa=t&rct=j&q=&esrc=s&source=web&cd=1&cad=rja&ved=0CEAQFjAA&url=http%3A%2F%2Fciteseerx.ist.psu.edu%2Fviewdoc%2Fdownload%3Fdoi%3D10.1.1.91.4940%26rep%3Drep1%26type%3Dpdf&ei=dOXrUMeIDYHDtQahsIGACg&usg=AFQjCNHFWVh6gNPvnOrOS9R3rkrXCNVD-A&sig2=5I2F5evRfMnsttSgFF9g7Q&bvm=bv.1357316858,d.Yms.
Date: 8. 1.2013 10:22:32

Liu, R.-L.: Context recognition for hierarchical text classification (2009) 0.03
```
0.033387464 = product of:
  0.08346866 = sum of:
    0.04098487 = weight(_text_:it in 2760) [ClassicSimilarity], result of:
      0.04098487 = score(doc=2760,freq=4.0), product of:
        0.15115225 = queryWeight, product of:
          2.892262 = idf(docFreq=6664, maxDocs=44218)
          0.052260913 = queryNorm
        0.27114958 = fieldWeight in 2760, product of:
          2.0 = tf(freq=4.0), with freq of:
            4.0 = termFreq=4.0
          2.892262 = idf(docFreq=6664, maxDocs=44218)
          0.046875 = fieldNorm(doc=2760)
    0.042483795 = weight(_text_:22 in 2760) [ClassicSimilarity], result of:
      0.042483795 = score(doc=2760,freq=2.0), product of:
        0.18300882 = queryWeight, product of:
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.052260913 = queryNorm
        0.23214069 = fieldWeight in 2760, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.046875 = fieldNorm(doc=2760)
  0.4 = coord(2/5)
```
Abstract

Information is often organized as a text hierarchy. A hierarchical text-classification system is thus essential for the management, sharing, and dissemination of information. It aims to automatically classify each incoming document into zero, one, or several categories in the text hierarchy. In this paper, we present a technique called CRHTC (context recognition for hierarchical text classification) that performs hierarchical text classification by recognizing the context of discussion (COD) of each category. A category's COD is governed by its ancestor categories, whose contents indicate contextual backgrounds of the category. A document may be classified into a category only if its content matches the category's COD. CRHTC does not require any trials to manually set parameters, and hence is more portable and easier to implement than other methods. It is empirically evaluated under various conditions. The results show that CRHTC achieves both better and more stable performance than several hierarchical and nonhierarchical text-classification methodologies.

Date

22. 3.2009 19:11:54

Jenkins, C.: Automatic classification of Web resources using Java and Dewey Decimal Classification (1998) 0.03

0.033350088 = product of:
  0.083375216 = sum of:
    0.03381079 = weight(_text_:it in 1673) [ClassicSimilarity], result of:
      0.03381079 = score(doc=1673,freq=2.0), product of:
        0.15115225 = queryWeight, product of:
          2.892262 = idf(docFreq=6664, maxDocs=44218)
          0.052260913 = queryNorm
        0.22368698 = fieldWeight in 1673, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          2.892262 = idf(docFreq=6664, maxDocs=44218)
          0.0546875 = fieldNorm(doc=1673)
    0.04956443 = weight(_text_:22 in 1673) [ClassicSimilarity], result of:
      0.04956443 = score(doc=1673,freq=2.0), product of:
        0.18300882 = queryWeight, product of:
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.052260913 = queryNorm
        0.2708308 = fieldWeight in 1673, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.0546875 = fieldNorm(doc=1673)
  0.4 = coord(2/5)

Abstract: The Wolverhampton Web Library (WWLib) is a WWW search engine that provides access to UK based information. The experimental version developed in 1995, was a success but highlighted the need for a much higher degree of automation. An interesting feature of the experimental WWLib was that it organised information according to DDC. Discusses the advantages of classification and describes the automatic classifier that is being developed in Java as part of the new, fully automated WWLib
Date: 1. 8.1996 22:08:06

Khoo, C.S.G.; Ng, K.; Ou, S.: ¬An exploratory study of human clustering of Web pages (2003) 0.02
```
0.019057194 = product of:
  0.047642983 = sum of:
    0.019320453 = weight(_text_:it in 2741) [ClassicSimilarity], result of:
      0.019320453 = score(doc=2741,freq=2.0), product of:
        0.15115225 = queryWeight, product of:
          2.892262 = idf(docFreq=6664, maxDocs=44218)
          0.052260913 = queryNorm
        0.12782113 = fieldWeight in 2741, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          2.892262 = idf(docFreq=6664, maxDocs=44218)
          0.03125 = fieldNorm(doc=2741)
    0.02832253 = weight(_text_:22 in 2741) [ClassicSimilarity], result of:
      0.02832253 = score(doc=2741,freq=2.0), product of:
        0.18300882 = queryWeight, product of:
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.052260913 = queryNorm
        0.15476047 = fieldWeight in 2741, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.03125 = fieldNorm(doc=2741)
  0.4 = coord(2/5)
```
Abstract

This study seeks to find out how human beings cluster Web pages naturally. Twenty Web pages retrieved by the Northem Light search engine for each of 10 queries were sorted by 3 subjects into categories that were natural or meaningful to them. lt was found that different subjects clustered the same set of Web pages quite differently and created different categories. The average inter-subject similarity of the clusters created was a low 0.27. Subjects created an average of 5.4 clusters for each sorting. The categories constructed can be divided into 10 types. About 1/3 of the categories created were topical. Another 20% of the categories relate to the degree of relevance or usefulness. The rest of the categories were subject-independent categories such as format, purpose, authoritativeness and direction to other sources. The authors plan to develop automatic methods for categorizing Web pages using the common categories created by the subjects. lt is hoped that the techniques developed can be used by Web search engines to automatically organize Web pages retrieved into categories that are natural to users. 1. Introduction The World Wide Web is an increasingly important source of information for people globally because of its ease of access, the ease of publishing, its ability to transcend geographic and national boundaries, its flexibility and heterogeneity and its dynamic nature. However, Web users also find it increasingly difficult to locate relevant and useful information in this vast information storehouse. Web search engines, despite their scope and power, appear to be quite ineffective. They retrieve too many pages, and though they attempt to rank retrieved pages in order of probable relevance, often the relevant documents do not appear in the top-ranked 10 or 20 documents displayed. Several studies have found that users do not know how to use the advanced features of Web search engines, and do not know how to formulate and re-formulate queries. Users also typically exert minimal effort in performing, evaluating and refining their searches, and are unwilling to scan more than 10 or 20 items retrieved (Jansen, Spink, Bateman & Saracevic, 1998). This suggests that the conventional ranked-list display of search results does not satisfy user requirements, and that better ways of presenting and summarizing search results have to be developed. One promising approach is to group retrieved pages into clusters or categories to allow users to navigate immediately to the "promising" clusters where the most useful Web pages are likely to be located. This approach has been adopted by a number of search engines (notably Northem Light) and search agents.

Date

12. 9.2004 9:56:22

Subramanian, S.; Shafer, K.E.: Clustering (2001) 0.02

0.016993519 = product of:
  0.08496759 = sum of:
    0.08496759 = weight(_text_:22 in 1046) [ClassicSimilarity], result of:
      0.08496759 = score(doc=1046,freq=2.0), product of:
        0.18300882 = queryWeight, product of:
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.052260913 = queryNorm
        0.46428138 = fieldWeight in 1046, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.09375 = fieldNorm(doc=1046)
  0.2 = coord(1/5)

Date: 5. 5.2003 14:17:22

Reiner, U.: Automatische DDC-Klassifizierung von bibliografischen Titeldatensätzen (2009) 0.01

0.014161265 = product of:
  0.070806324 = sum of:
    0.070806324 = weight(_text_:22 in 611) [ClassicSimilarity], result of:
      0.070806324 = score(doc=611,freq=2.0), product of:
        0.18300882 = queryWeight, product of:
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.052260913 = queryNorm
        0.38690117 = fieldWeight in 611, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.078125 = fieldNorm(doc=611)
  0.2 = coord(1/5)

Date: 22. 8.2009 12:54:24

HaCohen-Kerner, Y. et al.: Classification using various machine learning methods and combinations of key-phrases and visual features (2016) 0.01

0.014161265 = product of:
  0.070806324 = sum of:
    0.070806324 = weight(_text_:22 in 2748) [ClassicSimilarity], result of:
      0.070806324 = score(doc=2748,freq=2.0), product of:
        0.18300882 = queryWeight, product of:
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.052260913 = queryNorm
        0.38690117 = fieldWeight in 2748, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.078125 = fieldNorm(doc=2748)
  0.2 = coord(1/5)

Date: 1. 2.2016 18:25:22

Bianchini, C.; Bargioni, S.: Automated classification using linked open data : a case study on faceted classification and Wikidata (2021) 0.01
```
0.011712402 = product of:
  0.05856201 = sum of:
    0.05856201 = weight(_text_:it in 724) [ClassicSimilarity], result of:
      0.05856201 = score(doc=724,freq=6.0), product of:
        0.15115225 = queryWeight, product of:
          2.892262 = idf(docFreq=6664, maxDocs=44218)
          0.052260913 = queryNorm
        0.38743722 = fieldWeight in 724, product of:
          2.4494898 = tf(freq=6.0), with freq of:
            6.0 = termFreq=6.0
          2.892262 = idf(docFreq=6664, maxDocs=44218)
          0.0546875 = fieldNorm(doc=724)
  0.2 = coord(1/5)
```
Abstract

The Wikidata gadget, CCLitBox, for the automated classification of literary authors and works by a faceted classification and using Linked Open Data (LOD) is presented. The tool reproduces the classification algorithm of class O Literature of the Colon Classification and uses data freely available in Wikidata to create Colon Classification class numbers. CCLitBox is totally free and enables any user to classify literary authors and their works; it is easily accessible to everybody; it uses LOD from Wikidata but missing data for classification can be freely added if necessary; it is readymade for any cooperative and networked project.
Drori, O.; Alon, N.: Using document classification for displaying search results (2003) 0.01
```
0.011592272 = product of:
  0.057961356 = sum of:
    0.057961356 = weight(_text_:it in 1565) [ClassicSimilarity], result of:
      0.057961356 = score(doc=1565,freq=8.0), product of:
        0.15115225 = queryWeight, product of:
          2.892262 = idf(docFreq=6664, maxDocs=44218)
          0.052260913 = queryNorm
        0.38346338 = fieldWeight in 1565, product of:
          2.828427 = tf(freq=8.0), with freq of:
            8.0 = termFreq=8.0
          2.892262 = idf(docFreq=6664, maxDocs=44218)
          0.046875 = fieldNorm(doc=1565)
  0.2 = coord(1/5)
```
Abstract

In this paper, four self-developed user interfaces that display document search results using different methods were compared. In order to create the four interfaces, two information elements: document categories and lines from the document were used. A user study compared the four interfaces. It was found that the category addition to the interface was beneficial in both measurable and subjective measures. It was also found that displaying the relevant lines from the document increased the effectiveness and shortened the search time in all cases and tasks. It was found that the participants preferred the interface containing categories and relevant lines to all other interfaces checked. It was also the fastest in the objective time measurement. Another sub-research that was conducted showed that the most important parameter for the users was the confidence level that the answer was accurate, and the least important parameter was the feeling of comfort while conducting a search

Bock, H.-H.: Datenanalyse zur Strukturierung und Ordnung von Information (1989) 0.01

0.009912886 = product of:
  0.04956443 = sum of:
    0.04956443 = weight(_text_:22 in 141) [ClassicSimilarity], result of:
      0.04956443 = score(doc=141,freq=2.0), product of:
        0.18300882 = queryWeight, product of:
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.052260913 = queryNorm
        0.2708308 = fieldWeight in 141, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.0546875 = fieldNorm(doc=141)
  0.2 = coord(1/5)

Pages: S.1-22

Dubin, D.: Dimensions and discriminability (1998) 0.01

0.009912886 = product of:
  0.04956443 = sum of:
    0.04956443 = weight(_text_:22 in 2338) [ClassicSimilarity], result of:
      0.04956443 = score(doc=2338,freq=2.0), product of:
        0.18300882 = queryWeight, product of:
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.052260913 = queryNorm
        0.2708308 = fieldWeight in 2338, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.0546875 = fieldNorm(doc=2338)
  0.2 = coord(1/5)

Date: 22. 9.1997 19:16:05

Automatic classification research at OCLC (2002) 0.01

0.009912886 = product of:
  0.04956443 = sum of:
    0.04956443 = weight(_text_:22 in 1563) [ClassicSimilarity], result of:
      0.04956443 = score(doc=1563,freq=2.0), product of:
        0.18300882 = queryWeight, product of:
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.052260913 = queryNorm
        0.2708308 = fieldWeight in 1563, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.0546875 = fieldNorm(doc=1563)
  0.2 = coord(1/5)

Date: 5. 5.2003 9:22:09

Yoon, Y.; Lee, C.; Lee, G.G.: ¬An effective procedure for constructing a hierarchical text classification system (2006) 0.01

0.009912886 = product of:
  0.04956443 = sum of:
    0.04956443 = weight(_text_:22 in 5273) [ClassicSimilarity], result of:
      0.04956443 = score(doc=5273,freq=2.0), product of:
        0.18300882 = queryWeight, product of:
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.052260913 = queryNorm
        0.2708308 = fieldWeight in 5273, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.0546875 = fieldNorm(doc=5273)
  0.2 = coord(1/5)

Date: 22. 7.2006 16:24:52

Yi, K.: Automatic text classification using library classification schemes : trends, issues and challenges (2007) 0.01

0.009912886 = product of:
  0.04956443 = sum of:
    0.04956443 = weight(_text_:22 in 2560) [ClassicSimilarity], result of:
      0.04956443 = score(doc=2560,freq=2.0), product of:
        0.18300882 = queryWeight, product of:
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.052260913 = queryNorm
        0.2708308 = fieldWeight in 2560, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.0546875 = fieldNorm(doc=2560)
  0.2 = coord(1/5)

Date: 22. 9.2008 18:31:54

Dolin, R.; Agrawal, D.; El Abbadi, A.; Pearlman, J.: Using automated classification for summarizing and selecting heterogeneous information sources (1998) 0.01
```
0.009164495 = product of:
  0.045822475 = sum of:
    0.045822475 = weight(_text_:it in 1253) [ClassicSimilarity], result of:
      0.045822475 = score(doc=1253,freq=20.0), product of:
        0.15115225 = queryWeight, product of:
          2.892262 = idf(docFreq=6664, maxDocs=44218)
          0.052260913 = queryNorm
        0.30315444 = fieldWeight in 1253, product of:
          4.472136 = tf(freq=20.0), with freq of:
            20.0 = termFreq=20.0
          2.892262 = idf(docFreq=6664, maxDocs=44218)
          0.0234375 = fieldNorm(doc=1253)
  0.2 = coord(1/5)
```
Abstract

Information retrieval over the Internet increasingly requires the filtering of thousands of heterogeneous information sources. Important sources of information include not only traditional databases with structured data and queries, but also increasing numbers of non-traditional, semi- or unstructured collections such as Web sites, FTP archives, etc. As the number and variability of sources increases, new ways of automatically summarizing, discovering, and selecting collections relevant to a user's query are needed. One such method involves the use of classification schemes, such as the Library of Congress Classification (LCC), within which a collection may be represented based on its content, irrespective of the structure of the actual data or documents. For such a system to be useful in a large-scale distributed environment, it must be easy to use for both collection managers and users. As a result, it must be possible to classify documents automatically within a classification scheme. Furthermore, there must be a straightforward and intuitive interface with which the user may use the scheme to assist in information retrieval (IR). Our work with the Alexandria Digital Library (ADL) Project focuses on geo-referenced information, whether text, maps, aerial photographs, or satellite images. As a result, we have emphasized techniques which work with both text and non-text, such as combined textual and graphical queries, multi-dimensional indexing, and IR methods which are not solely dependent on words or phrases. Part of this work involves locating relevant online sources of information. In particular, we have designed and are currently testing aspects of an architecture, Pharos, which we believe will scale up to 1.000.000 heterogeneous sources. Pharos accommodates heterogeneity in content and format, both among multiple sources as well as within a single source. That is, we consider sources to include Web sites, FTP archives, newsgroups, and full digital libraries; all of these systems can include a wide variety of content and multimedia data formats. Pharos is based on the use of hierarchical classification schemes. These include not only well-known 'subject' (or 'concept') based schemes such as the Dewey Decimal System and the LCC, but also, for example, geographic classifications, which might be constructed as layers of smaller and smaller hierarchical longitude/latitude boxes. Pharos is designed to work with sophisticated queries which utilize subjects, geographical locations, temporal specifications, and other types of information domains. The Pharos architecture requires that hierarchically structured collection metadata be extracted so that it can be partitioned in such a way as to greatly enhance scalability. Automated classification is important to Pharos because it allows information sources to extract the requisite collection metadata automatically that must be distributed.
We are currently experimenting with newsgroups as collections. We have built an initial prototype which automatically classifies and summarizes newsgroups within the LCC. (The prototype can be tested below, and more details may be found at http://pharos.alexandria.ucsb.edu/). The prototype uses electronic library catalog records as a `training set' and Latent Semantic Indexing (LSI) for IR. We use the training set to build a rich set of classification terminology, and associate these terms with the relevant categories in the LCC. This association between terms and classification categories allows us to relate users' queries to nodes in the LCC so that users can select appropriate query categories. Newsgroups are similarly associated with classification categories. Pharos then matches the categories selected by users to relevant newsgroups. In principle, this approach allows users to exclude newsgroups that might have been selected based on an unintended meaning of a query term, and to include newsgroups with relevant content even though the exact query terms may not have been used. This work is extensible to other types of classification, including geographical, temporal, and image feature. Before discussing the methodology of the collection summarization and selection, we first present an online demonstration below. The demonstration is not intended to be a complete end-user interface. Rather, it is intended merely to offer a view of the process to suggest the "look and feel" of the prototype. The demo works as follows. First supply it with a few keywords of interest. The system will then use those terms to try to return to you the most relevant subject categories within the LCC. Assuming that the system recognizes any of your terms (it has over 400,000 terms indexed), it will give you a list of 15 LCC categories sorted by relevancy ranking. From there, you have two choices. The first choice, by clicking on the "News" links, is to get a list of newsgroups which the system has identified as relevant to the LCC category you select. The other choice, by clicking on the LCC ID links, is to enter the LCC hierarchy starting at the category of your choice and navigate the tree until you locate the best category for your query. From there, again, you can get a list of newsgroups by clicking on the "News" links. After having shown this demonstration to many people, we would like to suggest that you first give it easier examples before trying to break it. For example, "prostate cancer" (discussed below), "remote sensing", "investment banking", and "gershwin" all work reasonably well.

Pfeffer, M.: Automatische Vergabe von RVK-Notationen mittels fallbasiertem Schließen (2009) 0.01

0.008496759 = product of:
  0.042483795 = sum of:
    0.042483795 = weight(_text_:22 in 3051) [ClassicSimilarity], result of:
      0.042483795 = score(doc=3051,freq=2.0), product of:
        0.18300882 = queryWeight, product of:
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.052260913 = queryNorm
        0.23214069 = fieldWeight in 3051, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.046875 = fieldNorm(doc=3051)
  0.2 = coord(1/5)

Date: 22. 8.2009 19:51:28

Zhu, W.Z.; Allen, R.B.: Document clustering using the LSI subspace signature model (2013) 0.01

0.008496759 = product of:
  0.042483795 = sum of:
    0.042483795 = weight(_text_:22 in 690) [ClassicSimilarity], result of:
      0.042483795 = score(doc=690,freq=2.0), product of:
        0.18300882 = queryWeight, product of:
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.052260913 = queryNorm
        0.23214069 = fieldWeight in 690, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.046875 = fieldNorm(doc=690)
  0.2 = coord(1/5)

Date: 23. 3.2013 13:22:36

Egbert, J.; Biber, D.; Davies, M.: Developing a bottom-up, user-based method of web register classification (2015) 0.01

0.008496759 = product of:
  0.042483795 = sum of:
    0.042483795 = weight(_text_:22 in 2158) [ClassicSimilarity], result of:
      0.042483795 = score(doc=2158,freq=2.0), product of:
        0.18300882 = queryWeight, product of:
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.052260913 = queryNorm
        0.23214069 = fieldWeight in 2158, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5018296 = idf(docFreq=3622, maxDocs=44218)
          0.046875 = fieldNorm(doc=2158)
  0.2 = coord(1/5)

Date: 4. 8.2015 19:22:04

Mu, T.; Goulermas, J.Y.; Korkontzelos, I.; Ananiadou, S.: Descriptive document clustering via discriminant learning in a co-embedded space of multilevel similarities (2016) 0.01
```
0.008366001 = product of:
  0.041830003 = sum of:
    0.041830003 = weight(_text_:it in 2496) [ClassicSimilarity], result of:
      0.041830003 = score(doc=2496,freq=6.0), product of:
        0.15115225 = queryWeight, product of:
          2.892262 = idf(docFreq=6664, maxDocs=44218)
          0.052260913 = queryNorm
        0.27674085 = fieldWeight in 2496, product of:
          2.4494898 = tf(freq=6.0), with freq of:
            6.0 = termFreq=6.0
          2.892262 = idf(docFreq=6664, maxDocs=44218)
          0.0390625 = fieldNorm(doc=2496)
  0.2 = coord(1/5)
```
Abstract

Descriptive document clustering aims at discovering clusters of semantically interrelated documents together with meaningful labels to summarize the content of each document cluster. In this work, we propose a novel descriptive clustering framework, referred to as CEDL. It relies on the formulation and generation of 2 types of heterogeneous objects, which correspond to documents and candidate phrases, using multilevel similarity information. CEDL is composed of 5 main processing stages. First, it simultaneously maps the documents and candidate phrases into a common co-embedded space that preserves higher-order, neighbor-based proximities between the combined sets of documents and phrases. Then, it discovers an approximate cluster structure of documents in the common space. The third stage extracts promising topic phrases by constructing a discriminant model where documents along with their cluster memberships are used as training instances. Subsequently, the final cluster labels are selected from the topic phrases using a ranking scheme using multiple scores based on the extracted co-embedding information and the discriminant output. The final stage polishes the initial clusters to reduce noise and accommodate the multitopic nature of documents. The effectiveness and competitiveness of CEDL is demonstrated qualitatively and quantitatively with experiments using document databases from different application fields.
Wang, H.; Hong, M.: Supervised Hebb rule based feature selection for text classification (2019) 0.01
```
0.008366001 = product of:
  0.041830003 = sum of:
    0.041830003 = weight(_text_:it in 5036) [ClassicSimilarity], result of:
      0.041830003 = score(doc=5036,freq=6.0), product of:
        0.15115225 = queryWeight, product of:
          2.892262 = idf(docFreq=6664, maxDocs=44218)
          0.052260913 = queryNorm
        0.27674085 = fieldWeight in 5036, product of:
          2.4494898 = tf(freq=6.0), with freq of:
            6.0 = termFreq=6.0
          2.892262 = idf(docFreq=6664, maxDocs=44218)
          0.0390625 = fieldNorm(doc=5036)
  0.2 = coord(1/5)
```
Abstract

Text documents usually contain high dimensional non-discriminative (irrelevant and noisy) terms which lead to steep computational costs and poor learning performance of text classification. One of the effective solutions for this problem is feature selection which aims to identify discriminative terms from text data. This paper proposes a method termed "Hebb rule based feature selection (HRFS)". HRFS is based on supervised Hebb rule and assumes that terms and classes are neurons and select terms under the assumption that a term is discriminative if it keeps "exciting" the corresponding classes. This assumption can be explained as "a term is highly correlated with a class if it is able to keep "exciting" the class according to the original Hebb postulate. Six benchmarking datasets are used to compare HRFS with other seven feature selection methods. Experimental results indicate that HRFS is effective to achieve better performance than the compared methods. HRFS can identify discriminative terms in the view of synapse between neurons. Moreover, HRFS is also efficient because it can be described in the view of matrix operation to decrease complexity of feature selection.

Search (69 results, page 1 of 4)

Authors

Years

Languages

Types

Themes