Document (#35464)

Author
Kishida, K.
Title
High-speed rough clustering for very large document collections
Source
Journal of the American Society for Information Science and Technology. 61(2010) no.6, S.1092-1104
Year
2010
Abstract
Document clustering is an important tool, but it is not yet widely used in practice probably because of its high computational complexity. This article explores techniques of high-speed rough clustering of documents, assuming that it is sometimes necessary to obtain a clustering result in a shorter time, although the result is just an approximate outline of document clusters. A promising approach for such clustering is to reduce the number of documents to be checked for generating cluster vectors in the leader-follower clustering algorithm. Based on this idea, the present article proposes a modified Crouch algorithm and incomplete single-pass leader-follower algorithm. Also, a two-stage grouping technique, in which the first stage attempts to decrease the number of documents to be processed in the second stage by applying a quick merging technique, is developed. An experiment using a part of the Reuters corpus RCV1 showed empirically that both the modified Crouch and the incomplete single-pass leader-follower algorithms achieve clustering results more efficiently than the original methods, and also improved the effectiveness of clustering results. On the other hand, the two-stage grouping technique did not reduce the processing time in this experiment.
Theme
Automatisches Klassifizieren

Similar documents (content)

  1. Zamir, O.; Etzioni, O.: Grouper : a dynamic clustering interface to Web search results (1999) 0.23
    0.23384278 = sum of:
      0.23384278 = product of:
        0.83515275 = sum of:
          0.019854072 = weight(abstract_txt:time in 6207) [ClassicSimilarity], result of:
            0.019854072 = score(doc=6207,freq=1.0), product of:
              0.061261296 = queryWeight, product of:
                1.0629841 = boost
                4.148331 = idf(docFreq=1897, maxDocs=44218)
                0.013892679 = queryNorm
              0.32408836 = fieldWeight in 6207, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.148331 = idf(docFreq=1897, maxDocs=44218)
                0.078125 = fieldNorm(doc=6207)
          0.029202778 = weight(abstract_txt:documents in 6207) [ClassicSimilarity], result of:
            0.029202778 = score(doc=6207,freq=1.0), product of:
              0.0906984 = queryWeight, product of:
                1.5840874 = boost
                4.1213026 = idf(docFreq=1949, maxDocs=44218)
                0.013892679 = queryNorm
              0.32197678 = fieldWeight in 6207, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.1213026 = idf(docFreq=1949, maxDocs=44218)
                0.078125 = fieldNorm(doc=6207)
          0.04666587 = weight(abstract_txt:document in 6207) [ClassicSimilarity], result of:
            0.04666587 = score(doc=6207,freq=2.0), product of:
              0.09839501 = queryWeight, product of:
                1.6499313 = boost
                4.2926083 = idf(docFreq=1642, maxDocs=44218)
                0.013892679 = queryNorm
              0.4742707 = fieldWeight in 6207, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                4.2926083 = idf(docFreq=1642, maxDocs=44218)
                0.078125 = fieldNorm(doc=6207)
          0.07962909 = weight(abstract_txt:speed in 6207) [ClassicSimilarity], result of:
            0.07962909 = score(doc=6207,freq=1.0), product of:
              0.15464443 = queryWeight, product of:
                1.688888 = boost
                6.590942 = idf(docFreq=164, maxDocs=44218)
                0.013892679 = queryNorm
              0.5149173 = fieldWeight in 6207, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.590942 = idf(docFreq=164, maxDocs=44218)
                0.078125 = fieldNorm(doc=6207)
          0.047875125 = weight(abstract_txt:high in 6207) [ClassicSimilarity], result of:
            0.047875125 = score(doc=6207,freq=1.0), product of:
              0.12610243 = queryWeight, product of:
                1.8678459 = boost
                4.8595543 = idf(docFreq=931, maxDocs=44218)
                0.013892679 = queryNorm
              0.37965268 = fieldWeight in 6207, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.8595543 = idf(docFreq=931, maxDocs=44218)
                0.078125 = fieldNorm(doc=6207)
          0.077479035 = weight(abstract_txt:algorithm in 6207) [ClassicSimilarity], result of:
            0.077479035 = score(doc=6207,freq=1.0), product of:
              0.17382263 = queryWeight, product of:
                2.1929688 = boost
                5.705423 = idf(docFreq=399, maxDocs=44218)
                0.013892679 = queryNorm
              0.44573617 = fieldWeight in 6207, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.705423 = idf(docFreq=399, maxDocs=44218)
                0.078125 = fieldNorm(doc=6207)
          0.5344468 = weight(abstract_txt:clustering in 6207) [ClassicSimilarity], result of:
            0.5344468 = score(doc=6207,freq=4.0), product of:
              0.550245 = queryWeight, product of:
                6.3715005 = boost
                6.2162485 = idf(docFreq=239, maxDocs=44218)
                0.013892679 = queryNorm
              0.9712888 = fieldWeight in 6207, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                6.2162485 = idf(docFreq=239, maxDocs=44218)
                0.078125 = fieldNorm(doc=6207)
        0.28 = coord(7/25)
    
  2. Lee, Y.-H.; Wei, C.-P.; Hu, P.J.-H.: ¬An ontology-based technique for preserving user preferences in document-category evolutions (2011) 0.20
    0.19852391 = sum of:
      0.19852391 = product of:
        0.70901394 = sum of:
          0.04628362 = weight(abstract_txt:vectors in 4353) [ClassicSimilarity], result of:
            0.04628362 = score(doc=4353,freq=1.0), product of:
              0.10843329 = queryWeight, product of:
                7.805067 = idf(docFreq=48, maxDocs=44218)
                0.013892679 = queryNorm
              0.4268396 = fieldWeight in 4353, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.805067 = idf(docFreq=48, maxDocs=44218)
                0.0546875 = fieldNorm(doc=4353)
          0.013897852 = weight(abstract_txt:time in 4353) [ClassicSimilarity], result of:
            0.013897852 = score(doc=4353,freq=1.0), product of:
              0.061261296 = queryWeight, product of:
                1.0629841 = boost
                4.148331 = idf(docFreq=1897, maxDocs=44218)
                0.013892679 = queryNorm
              0.22686186 = fieldWeight in 4353, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.148331 = idf(docFreq=1897, maxDocs=44218)
                0.0546875 = fieldNorm(doc=4353)
          0.035406485 = weight(abstract_txt:documents in 4353) [ClassicSimilarity], result of:
            0.035406485 = score(doc=4353,freq=3.0), product of:
              0.0906984 = queryWeight, product of:
                1.5840874 = boost
                4.1213026 = idf(docFreq=1949, maxDocs=44218)
                0.013892679 = queryNorm
              0.39037606 = fieldWeight in 4353, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                4.1213026 = idf(docFreq=1949, maxDocs=44218)
                0.0546875 = fieldNorm(doc=4353)
          0.06533222 = weight(abstract_txt:document in 4353) [ClassicSimilarity], result of:
            0.06533222 = score(doc=4353,freq=8.0), product of:
              0.09839501 = queryWeight, product of:
                1.6499313 = boost
                4.2926083 = idf(docFreq=1642, maxDocs=44218)
                0.013892679 = queryNorm
              0.663979 = fieldWeight in 4353, product of:
                2.828427 = tf(freq=8.0), with freq of:
                  8.0 = termFreq=8.0
                4.2926083 = idf(docFreq=1642, maxDocs=44218)
                0.0546875 = fieldNorm(doc=4353)
          0.08915305 = weight(abstract_txt:grouping in 4353) [ClassicSimilarity], result of:
            0.08915305 = score(doc=4353,freq=1.0), product of:
              0.21150073 = queryWeight, product of:
                1.9751024 = boost
                7.7079034 = idf(docFreq=53, maxDocs=44218)
                0.013892679 = queryNorm
              0.42152596 = fieldWeight in 4353, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.7079034 = idf(docFreq=53, maxDocs=44218)
                0.0546875 = fieldNorm(doc=4353)
          0.13494956 = weight(abstract_txt:technique in 4353) [ClassicSimilarity], result of:
            0.13494956 = score(doc=4353,freq=7.0), product of:
              0.16685265 = queryWeight, product of:
                2.148552 = boost
                5.5898643 = idf(docFreq=448, maxDocs=44218)
                0.013892679 = queryNorm
              0.8087948 = fieldWeight in 4353, product of:
                2.6457512 = tf(freq=7.0), with freq of:
                  7.0 = termFreq=7.0
                5.5898643 = idf(docFreq=448, maxDocs=44218)
                0.0546875 = fieldNorm(doc=4353)
          0.32399115 = weight(abstract_txt:clustering in 4353) [ClassicSimilarity], result of:
            0.32399115 = score(doc=4353,freq=3.0), product of:
              0.550245 = queryWeight, product of:
                6.3715005 = boost
                6.2162485 = idf(docFreq=239, maxDocs=44218)
                0.013892679 = queryNorm
              0.58881253 = fieldWeight in 4353, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                6.2162485 = idf(docFreq=239, maxDocs=44218)
                0.0546875 = fieldNorm(doc=4353)
        0.28 = coord(7/25)
    
  3. Zhan, J.; Loh, H.T.: Using latent semantic indexing to improve the accuracy of document clustering (2007) 0.18
    0.1819868 = sum of:
      0.1819868 = product of:
        0.909934 = sum of:
          0.019854072 = weight(abstract_txt:time in 264) [ClassicSimilarity], result of:
            0.019854072 = score(doc=264,freq=1.0), product of:
              0.061261296 = queryWeight, product of:
                1.0629841 = boost
                4.148331 = idf(docFreq=1897, maxDocs=44218)
                0.013892679 = queryNorm
              0.32408836 = fieldWeight in 264, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.148331 = idf(docFreq=1897, maxDocs=44218)
                0.078125 = fieldNorm(doc=264)
          0.06920259 = weight(abstract_txt:reduce in 264) [ClassicSimilarity], result of:
            0.06920259 = score(doc=264,freq=1.0), product of:
              0.14083199 = queryWeight, product of:
                1.6117005 = boost
                6.2897153 = idf(docFreq=222, maxDocs=44218)
                0.013892679 = queryNorm
              0.491384 = fieldWeight in 264, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.2897153 = idf(docFreq=222, maxDocs=44218)
                0.078125 = fieldNorm(doc=264)
          0.06599551 = weight(abstract_txt:document in 264) [ClassicSimilarity], result of:
            0.06599551 = score(doc=264,freq=4.0), product of:
              0.09839501 = queryWeight, product of:
                1.6499313 = boost
                4.2926083 = idf(docFreq=1642, maxDocs=44218)
                0.013892679 = queryNorm
              0.67072004 = fieldWeight in 264, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                4.2926083 = idf(docFreq=1642, maxDocs=44218)
                0.078125 = fieldNorm(doc=264)
          0.047875125 = weight(abstract_txt:high in 264) [ClassicSimilarity], result of:
            0.047875125 = score(doc=264,freq=1.0), product of:
              0.12610243 = queryWeight, product of:
                1.8678459 = boost
                4.8595543 = idf(docFreq=931, maxDocs=44218)
                0.013892679 = queryNorm
              0.37965268 = fieldWeight in 264, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.8595543 = idf(docFreq=931, maxDocs=44218)
                0.078125 = fieldNorm(doc=264)
          0.7070067 = weight(abstract_txt:clustering in 264) [ClassicSimilarity], result of:
            0.7070067 = score(doc=264,freq=7.0), product of:
              0.550245 = queryWeight, product of:
                6.3715005 = boost
                6.2162485 = idf(docFreq=239, maxDocs=44218)
                0.013892679 = queryNorm
              1.2848943 = fieldWeight in 264, product of:
                2.6457512 = tf(freq=7.0), with freq of:
                  7.0 = termFreq=7.0
                6.2162485 = idf(docFreq=239, maxDocs=44218)
                0.078125 = fieldNorm(doc=264)
        0.2 = coord(5/25)
    
  4. Mu, T.; Goulermas, J.Y.; Korkontzelos, I.; Ananiadou, S.: Descriptive document clustering via discriminant learning in a co-embedded space of multilevel similarities (2016) 0.16
    0.15921357 = sum of:
      0.15921357 = product of:
        0.6633899 = sum of:
          0.054645896 = weight(abstract_txt:approximate in 2496) [ClassicSimilarity], result of:
            0.054645896 = score(doc=2496,freq=1.0), product of:
              0.11081234 = queryWeight, product of:
                1.0109106 = boost
                7.890225 = idf(docFreq=44, maxDocs=44218)
                0.013892679 = queryNorm
              0.49313906 = fieldWeight in 2496, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.890225 = idf(docFreq=44, maxDocs=44218)
                0.0625 = fieldNorm(doc=2496)
          0.061810624 = weight(abstract_txt:documents in 2496) [ClassicSimilarity], result of:
            0.061810624 = score(doc=2496,freq=7.0), product of:
              0.0906984 = queryWeight, product of:
                1.5840874 = boost
                4.1213026 = idf(docFreq=1949, maxDocs=44218)
                0.013892679 = queryNorm
              0.6814963 = fieldWeight in 2496, product of:
                2.6457512 = tf(freq=7.0), with freq of:
                  7.0 = termFreq=7.0
                4.1213026 = idf(docFreq=1949, maxDocs=44218)
                0.0625 = fieldNorm(doc=2496)
          0.05536207 = weight(abstract_txt:reduce in 2496) [ClassicSimilarity], result of:
            0.05536207 = score(doc=2496,freq=1.0), product of:
              0.14083199 = queryWeight, product of:
                1.6117005 = boost
                6.2897153 = idf(docFreq=222, maxDocs=44218)
                0.013892679 = queryNorm
              0.3931072 = fieldWeight in 2496, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.2897153 = idf(docFreq=222, maxDocs=44218)
                0.0625 = fieldNorm(doc=2496)
          0.04572303 = weight(abstract_txt:document in 2496) [ClassicSimilarity], result of:
            0.04572303 = score(doc=2496,freq=3.0), product of:
              0.09839501 = queryWeight, product of:
                1.6499313 = boost
                4.2926083 = idf(docFreq=1642, maxDocs=44218)
                0.013892679 = queryNorm
              0.46468848 = fieldWeight in 2496, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                4.2926083 = idf(docFreq=1642, maxDocs=44218)
                0.0625 = fieldNorm(doc=2496)
          0.1435195 = weight(abstract_txt:stage in 2496) [ClassicSimilarity], result of:
            0.1435195 = score(doc=2496,freq=2.0), product of:
              0.2657666 = queryWeight, product of:
                3.131114 = boost
                6.1096387 = idf(docFreq=266, maxDocs=44218)
                0.013892679 = queryNorm
              0.5400209 = fieldWeight in 2496, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                6.1096387 = idf(docFreq=266, maxDocs=44218)
                0.0625 = fieldNorm(doc=2496)
          0.30232877 = weight(abstract_txt:clustering in 2496) [ClassicSimilarity], result of:
            0.30232877 = score(doc=2496,freq=2.0), product of:
              0.550245 = queryWeight, product of:
                6.3715005 = boost
                6.2162485 = idf(docFreq=239, maxDocs=44218)
                0.013892679 = queryNorm
              0.5494439 = fieldWeight in 2496, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                6.2162485 = idf(docFreq=239, maxDocs=44218)
                0.0625 = fieldNorm(doc=2496)
        0.24 = coord(6/25)
    
  5. Guerrero, V.P.; Moya Anegón, F. de: Reduction of the dimension of a document space using the fuzzified output of a Kohonen network (2001) 0.16
    0.15753983 = sum of:
      0.15753983 = product of:
        0.656416 = sum of:
          0.11220844 = weight(abstract_txt:vectors in 6935) [ClassicSimilarity], result of:
            0.11220844 = score(doc=6935,freq=2.0), product of:
              0.10843329 = queryWeight, product of:
                7.805067 = idf(docFreq=48, maxDocs=44218)
                0.013892679 = queryNorm
              1.0348154 = fieldWeight in 6935, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                7.805067 = idf(docFreq=48, maxDocs=44218)
                0.09375 = fieldNorm(doc=6935)
          0.023555709 = weight(abstract_txt:number in 6935) [ClassicSimilarity], result of:
            0.023555709 = score(doc=6935,freq=1.0), product of:
              0.06079899 = queryWeight, product of:
                1.0589657 = boost
                4.132649 = idf(docFreq=1927, maxDocs=44218)
                0.013892679 = queryNorm
              0.38743585 = fieldWeight in 6935, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.132649 = idf(docFreq=1927, maxDocs=44218)
                0.09375 = fieldNorm(doc=6935)
          0.04955876 = weight(abstract_txt:documents in 6935) [ClassicSimilarity], result of:
            0.04955876 = score(doc=6935,freq=2.0), product of:
              0.0906984 = queryWeight, product of:
                1.5840874 = boost
                4.1213026 = idf(docFreq=1949, maxDocs=44218)
                0.013892679 = queryNorm
              0.5464127 = fieldWeight in 6935, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                4.1213026 = idf(docFreq=1949, maxDocs=44218)
                0.09375 = fieldNorm(doc=6935)
          0.057450153 = weight(abstract_txt:high in 6935) [ClassicSimilarity], result of:
            0.057450153 = score(doc=6935,freq=1.0), product of:
              0.12610243 = queryWeight, product of:
                1.8678459 = boost
                4.8595543 = idf(docFreq=931, maxDocs=44218)
                0.013892679 = queryNorm
              0.4555832 = fieldWeight in 6935, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.8595543 = idf(docFreq=931, maxDocs=44218)
                0.09375 = fieldNorm(doc=6935)
          0.092974834 = weight(abstract_txt:algorithm in 6935) [ClassicSimilarity], result of:
            0.092974834 = score(doc=6935,freq=1.0), product of:
              0.17382263 = queryWeight, product of:
                2.1929688 = boost
                5.705423 = idf(docFreq=399, maxDocs=44218)
                0.013892679 = queryNorm
              0.5348834 = fieldWeight in 6935, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.705423 = idf(docFreq=399, maxDocs=44218)
                0.09375 = fieldNorm(doc=6935)
          0.3206681 = weight(abstract_txt:clustering in 6935) [ClassicSimilarity], result of:
            0.3206681 = score(doc=6935,freq=1.0), product of:
              0.550245 = queryWeight, product of:
                6.3715005 = boost
                6.2162485 = idf(docFreq=239, maxDocs=44218)
                0.013892679 = queryNorm
              0.5827733 = fieldWeight in 6935, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.2162485 = idf(docFreq=239, maxDocs=44218)
                0.09375 = fieldNorm(doc=6935)
        0.24 = coord(6/25)