Document (#30264)

Author
Hsu, C.-N.
Chang, C.-H.
Hsieh, C.-H.
Lu, J.-J.
Chang, C.-C.
Title
Reconfigurable Web wrapper agents for biological information integration
Source
Journal of the American Society for Information Science and Technology. 56(2005) no.5, S.505-517
Year
2005
Abstract
A variety of biological data is transferred and exchanged in overwhelming volumes on the World Wide Web. How to rapidly capture, utilize, and integrate the information on the Internet to discover valuable biological knowledge is one of the most critical issues in bioinformatics. Many information integration systems have been proposed for integrating biological data. These systems usually rely on an intermediate software layer called wrappers to access connected information sources. Wrapper construction for Web data sources is often specially hand coded to accommodate the differences between each Web site. However, programming a Web wrapper requires substantial programming skill, and is time-consuming and hard to maintain. In this article we provide a solution for rapidly building software agents that can serve as Web wrappers for biological information integration. We define an XML-based language called Web Navigation Description Language (WNDL), to model a Web-browsing session. A WNDL script describes how to locate the data, extract the data, and combine the data. By executing different WNDL scripts, we can automate virtually all types of Web-browsing sessions. We also describe IEPAD (Information Extraction Based on Pattern Discovery), a data extractor based on pattern discovery techniques. IEPAD allows our software agents to automatically discover the extraction rules to extract the contents of a structurally formatted Web page. With a programming-by-example authoring tool, a user can generate a complete Web wrapper agent by browsing the target Web sites. We built a variety of biological applications to demonstrate the feasibility of our approach.
Footnote
Beitrag in einem special issue on bioinformatics

Similar documents (author)

  1. Yang, T.-H.; Hsieh, Y.-L.; Liu, S.-H.; Chang, Y.-C.; Hsu, W.-L.: ¬A flexible template generation and matching method with applications for publication reference metadata extraction (2021) 2.75
    2.7481346 = sum of:
      2.7481346 = sum of:
        1.1214445 = weight(author_txt:hsieh in 63) [ClassicSimilarity], result of:
          1.1214445 = score(doc=63,freq=1.0), product of:
            0.52657187 = queryWeight, product of:
              8.518833 = idf(docFreq=23, maxDocs=44218)
              0.061812673 = queryNorm
            2.1297083 = fieldWeight in 63, product of:
              1.0 = tf(freq=1.0), with freq of:
                1.0 = termFreq=1.0
              8.518833 = idf(docFreq=23, maxDocs=44218)
              0.25 = fieldNorm(doc=63)
        1.62669 = weight(author_txt:chang in 63) [ClassicSimilarity], result of:
          1.62669 = score(doc=63,freq=1.0), product of:
            0.8501306 = queryWeight, product of:
              1.7969211 = boost
              7.653836 = idf(docFreq=56, maxDocs=44218)
              0.061812673 = queryNorm
            1.913459 = fieldWeight in 63, product of:
              1.0 = tf(freq=1.0), with freq of:
                1.0 = termFreq=1.0
              7.653836 = idf(docFreq=56, maxDocs=44218)
              0.25 = fieldNorm(doc=63)
    
  2. Chang, R.: DBase, relational data models, and MARC records (1992) 2.03
    2.0333626 = sum of:
      2.0333626 = product of:
        4.0667253 = sum of:
          4.0667253 = weight(author_txt:chang in 5057) [ClassicSimilarity], result of:
            4.0667253 = score(doc=5057,freq=1.0), product of:
              0.8501306 = queryWeight, product of:
                1.7969211 = boost
                7.653836 = idf(docFreq=56, maxDocs=44218)
                0.061812673 = queryNorm
              4.7836475 = fieldWeight in 5057, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.653836 = idf(docFreq=56, maxDocs=44218)
                0.625 = fieldNorm(doc=5057)
        0.5 = coord(1/2)
    
  3. Chang, R.: ¬The development of indexing technology (1993) 2.03
    2.0333626 = sum of:
      2.0333626 = product of:
        4.0667253 = sum of:
          4.0667253 = weight(author_txt:chang in 7024) [ClassicSimilarity], result of:
            4.0667253 = score(doc=7024,freq=1.0), product of:
              0.8501306 = queryWeight, product of:
                1.7969211 = boost
                7.653836 = idf(docFreq=56, maxDocs=44218)
                0.061812673 = queryNorm
              4.7836475 = fieldWeight in 7024, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.653836 = idf(docFreq=56, maxDocs=44218)
                0.625 = fieldNorm(doc=7024)
        0.5 = coord(1/2)
    
  4. Chang, R.: Keyword searching and indexing (1993) 2.03
    2.0333626 = sum of:
      2.0333626 = product of:
        4.0667253 = sum of:
          4.0667253 = weight(author_txt:chang in 7223) [ClassicSimilarity], result of:
            4.0667253 = score(doc=7223,freq=1.0), product of:
              0.8501306 = queryWeight, product of:
                1.7969211 = boost
                7.653836 = idf(docFreq=56, maxDocs=44218)
                0.061812673 = queryNorm
              4.7836475 = fieldWeight in 7223, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.653836 = idf(docFreq=56, maxDocs=44218)
                0.625 = fieldNorm(doc=7223)
        0.5 = coord(1/2)
    
  5. Chang, R.H.: To classify or not to classify? : a new look at an old problem (1989) 2.03
    2.0333626 = sum of:
      2.0333626 = product of:
        4.0667253 = sum of:
          4.0667253 = weight(author_txt:chang in 2510) [ClassicSimilarity], result of:
            4.0667253 = score(doc=2510,freq=1.0), product of:
              0.8501306 = queryWeight, product of:
                1.7969211 = boost
                7.653836 = idf(docFreq=56, maxDocs=44218)
                0.061812673 = queryNorm
              4.7836475 = fieldWeight in 2510, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.653836 = idf(docFreq=56, maxDocs=44218)
                0.625 = fieldNorm(doc=2510)
        0.5 = coord(1/2)
    

Similar documents (content)

  1. Haslhofer, B.: ¬A Web-based mapping technique for establishing metadata interoperability (2008) 0.17
    0.17460002 = sum of:
      0.17460002 = product of:
        0.6235715 = sum of:
          0.026102198 = weight(abstract_txt:sources in 3173) [ClassicSimilarity], result of:
            0.026102198 = score(doc=3173,freq=4.0), product of:
              0.070396766 = queryWeight, product of:
                1.1142541 = boost
                4.7460723 = idf(docFreq=1043, maxDocs=44218)
                0.013311718 = queryNorm
              0.3707869 = fieldWeight in 3173, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                4.7460723 = idf(docFreq=1043, maxDocs=44218)
                0.0390625 = fieldNorm(doc=3173)
          0.010275934 = weight(abstract_txt:based in 3173) [ClassicSimilarity], result of:
            0.010275934 = score(doc=3173,freq=3.0), product of:
              0.047642242 = queryWeight, product of:
                1.1226636 = boost
                3.1879277 = idf(docFreq=4958, maxDocs=44218)
                0.013311718 = queryNorm
              0.21568955 = fieldWeight in 3173, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                3.1879277 = idf(docFreq=4958, maxDocs=44218)
                0.0390625 = fieldNorm(doc=3173)
          0.03063324 = weight(abstract_txt:discovery in 3173) [ClassicSimilarity], result of:
            0.03063324 = score(doc=3173,freq=2.0), product of:
              0.09868245 = queryWeight, product of:
                1.3192523 = boost
                5.619245 = idf(docFreq=435, maxDocs=44218)
                0.013311718 = queryNorm
              0.31042236 = fieldWeight in 3173, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                5.619245 = idf(docFreq=435, maxDocs=44218)
                0.0390625 = fieldNorm(doc=3173)
          0.007349127 = weight(abstract_txt:information in 3173) [ClassicSimilarity], result of:
            0.007349127 = score(doc=3173,freq=2.0), product of:
              0.054950997 = queryWeight, product of:
                1.7051253 = boost
                2.4209464 = idf(docFreq=10677, maxDocs=44218)
                0.013311718 = queryNorm
              0.13373965 = fieldWeight in 3173, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                2.4209464 = idf(docFreq=10677, maxDocs=44218)
                0.0390625 = fieldNorm(doc=3173)
          0.0471741 = weight(abstract_txt:integration in 3173) [ClassicSimilarity], result of:
            0.0471741 = score(doc=3173,freq=3.0), product of:
              0.13159733 = queryWeight, product of:
                1.8658514 = boost
                5.298292 = idf(docFreq=600, maxDocs=44218)
                0.013311718 = queryNorm
              0.3584731 = fieldWeight in 3173, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                5.298292 = idf(docFreq=600, maxDocs=44218)
                0.0390625 = fieldNorm(doc=3173)
          0.027484426 = weight(abstract_txt:data in 3173) [ClassicSimilarity], result of:
            0.027484426 = score(doc=3173,freq=3.0), product of:
              0.12175721 = queryWeight, product of:
                2.7415063 = boost
                3.3363478 = idf(docFreq=4274, maxDocs=44218)
                0.013311718 = queryNorm
              0.2257314 = fieldWeight in 3173, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                3.3363478 = idf(docFreq=4274, maxDocs=44218)
                0.0390625 = fieldNorm(doc=3173)
          0.47455245 = weight(abstract_txt:wrapper in 3173) [ClassicSimilarity], result of:
            0.47455245 = score(doc=3173,freq=4.0), product of:
              0.6132451 = queryWeight, product of:
                4.6509314 = boost
                9.905128 = idf(docFreq=5, maxDocs=44218)
                0.013311718 = queryNorm
              0.7738381 = fieldWeight in 3173, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                9.905128 = idf(docFreq=5, maxDocs=44218)
                0.0390625 = fieldNorm(doc=3173)
        0.28 = coord(7/25)
    
  2. Rodríguez, A.; Carazo, J.M.; Trelles-Salazar, O.: Mining association rules from biological databases (2005) 0.14
    0.13768524 = sum of:
      0.13768524 = product of:
        0.5736885 = sum of:
          0.086672686 = weight(abstract_txt:bioinformatics in 5261) [ClassicSimilarity], result of:
            0.086672686 = score(doc=5261,freq=2.0), product of:
              0.11453622 = queryWeight, product of:
                1.004996 = boost
                8.561393 = idf(docFreq=22, maxDocs=44218)
                0.013311718 = queryNorm
              0.75672734 = fieldWeight in 5261, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                8.561393 = idf(docFreq=22, maxDocs=44218)
                0.0625 = fieldNorm(doc=5261)
          0.034657553 = weight(abstract_txt:discovery in 5261) [ClassicSimilarity], result of:
            0.034657553 = score(doc=5261,freq=1.0), product of:
              0.09868245 = queryWeight, product of:
                1.3192523 = boost
                5.619245 = idf(docFreq=435, maxDocs=44218)
                0.013311718 = queryNorm
              0.35120282 = fieldWeight in 5261, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.619245 = idf(docFreq=435, maxDocs=44218)
                0.0625 = fieldNorm(doc=5261)
          0.046362128 = weight(abstract_txt:extraction in 5261) [ClassicSimilarity], result of:
            0.046362128 = score(doc=5261,freq=1.0), product of:
              0.11980738 = queryWeight, product of:
                1.4536159 = boost
                6.1915555 = idf(docFreq=245, maxDocs=44218)
                0.013311718 = queryNorm
              0.38697222 = fieldWeight in 5261, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.1915555 = idf(docFreq=245, maxDocs=44218)
                0.0625 = fieldNorm(doc=5261)
          0.047592483 = weight(abstract_txt:pattern in 5261) [ClassicSimilarity], result of:
            0.047592483 = score(doc=5261,freq=1.0), product of:
              0.12191774 = queryWeight, product of:
                1.4663625 = boost
                6.2458487 = idf(docFreq=232, maxDocs=44218)
                0.013311718 = queryNorm
              0.39036554 = fieldWeight in 5261, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.2458487 = idf(docFreq=232, maxDocs=44218)
                0.0625 = fieldNorm(doc=5261)
          0.025389025 = weight(abstract_txt:data in 5261) [ClassicSimilarity], result of:
            0.025389025 = score(doc=5261,freq=1.0), product of:
              0.12175721 = queryWeight, product of:
                2.7415063 = boost
                3.3363478 = idf(docFreq=4274, maxDocs=44218)
                0.013311718 = queryNorm
              0.20852174 = fieldWeight in 5261, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.3363478 = idf(docFreq=4274, maxDocs=44218)
                0.0625 = fieldNorm(doc=5261)
          0.3330146 = weight(abstract_txt:biological in 5261) [ClassicSimilarity], result of:
            0.3330146 = score(doc=5261,freq=2.0), product of:
              0.5105606 = queryWeight, product of:
                5.197472 = boost
                7.3793993 = idf(docFreq=74, maxDocs=44218)
                0.013311718 = queryNorm
              0.6522529 = fieldWeight in 5261, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                7.3793993 = idf(docFreq=74, maxDocs=44218)
                0.0625 = fieldNorm(doc=5261)
        0.24 = coord(6/25)
    
  3. Mahoui, M.; Miled, Z.B.; Godse, A.; Kulkarni, H.; Li, N.: BioFacets : faceted classification for biological information (2006) 0.12
    0.117209874 = sum of:
      0.117209874 = product of:
        0.7325617 = sum of:
          0.011865627 = weight(abstract_txt:based in 779) [ClassicSimilarity], result of:
            0.011865627 = score(doc=779,freq=1.0), product of:
              0.047642242 = queryWeight, product of:
                1.1226636 = boost
                3.1879277 = idf(docFreq=4958, maxDocs=44218)
                0.013311718 = queryNorm
              0.24905685 = fieldWeight in 779, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.1879277 = idf(docFreq=4958, maxDocs=44218)
                0.078125 = fieldNorm(doc=779)
          0.07703499 = weight(abstract_txt:integration in 779) [ClassicSimilarity], result of:
            0.07703499 = score(doc=779,freq=2.0), product of:
              0.13159733 = queryWeight, product of:
                1.8658514 = boost
                5.298292 = idf(docFreq=600, maxDocs=44218)
                0.013311718 = queryNorm
              0.58538413 = fieldWeight in 779, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                5.298292 = idf(docFreq=600, maxDocs=44218)
                0.078125 = fieldNorm(doc=779)
          0.054968853 = weight(abstract_txt:data in 779) [ClassicSimilarity], result of:
            0.054968853 = score(doc=779,freq=3.0), product of:
              0.12175721 = queryWeight, product of:
                2.7415063 = boost
                3.3363478 = idf(docFreq=4274, maxDocs=44218)
                0.013311718 = queryNorm
              0.4514628 = fieldWeight in 779, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                3.3363478 = idf(docFreq=4274, maxDocs=44218)
                0.078125 = fieldNorm(doc=779)
          0.58869225 = weight(abstract_txt:biological in 779) [ClassicSimilarity], result of:
            0.58869225 = score(doc=779,freq=4.0), product of:
              0.5105606 = queryWeight, product of:
                5.197472 = boost
                7.3793993 = idf(docFreq=74, maxDocs=44218)
                0.013311718 = queryNorm
              1.1530311 = fieldWeight in 779, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                7.3793993 = idf(docFreq=74, maxDocs=44218)
                0.078125 = fieldNorm(doc=779)
        0.16 = coord(4/25)
    
  4. Handbook of metadata, semantics and ontologies (2014) 0.09
    0.08877933 = sum of:
      0.08877933 = product of:
        0.36991388 = sum of:
          0.0134244235 = weight(abstract_txt:based in 5134) [ClassicSimilarity], result of:
            0.0134244235 = score(doc=5134,freq=2.0), product of:
              0.047642242 = queryWeight, product of:
                1.1226636 = boost
                3.1879277 = idf(docFreq=4958, maxDocs=44218)
                0.013311718 = queryNorm
              0.28177565 = fieldWeight in 5134, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.1879277 = idf(docFreq=4958, maxDocs=44218)
                0.0625 = fieldNorm(doc=5134)
          0.026681572 = weight(abstract_txt:variety in 5134) [ClassicSimilarity], result of:
            0.026681572 = score(doc=5134,freq=1.0), product of:
              0.08289257 = queryWeight, product of:
                1.2091097 = boost
                5.1501017 = idf(docFreq=696, maxDocs=44218)
                0.013311718 = queryNorm
              0.32188135 = fieldWeight in 5134, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.1501017 = idf(docFreq=696, maxDocs=44218)
                0.0625 = fieldNorm(doc=5134)
          0.02910658 = weight(abstract_txt:called in 5134) [ClassicSimilarity], result of:
            0.02910658 = score(doc=5134,freq=1.0), product of:
              0.08784197 = queryWeight, product of:
                1.2446835 = boost
                5.3016257 = idf(docFreq=598, maxDocs=44218)
                0.013311718 = queryNorm
              0.3313516 = fieldWeight in 5134, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.3016257 = idf(docFreq=598, maxDocs=44218)
                0.0625 = fieldNorm(doc=5134)
          0.053465817 = weight(abstract_txt:rapidly in 5134) [ClassicSimilarity], result of:
            0.053465817 = score(doc=5134,freq=1.0), product of:
              0.13175248 = queryWeight, product of:
                1.5243591 = boost
                6.4928803 = idf(docFreq=181, maxDocs=44218)
                0.013311718 = queryNorm
              0.40580502 = fieldWeight in 5134, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.4928803 = idf(docFreq=181, maxDocs=44218)
                0.0625 = fieldNorm(doc=5134)
          0.011758604 = weight(abstract_txt:information in 5134) [ClassicSimilarity], result of:
            0.011758604 = score(doc=5134,freq=2.0), product of:
              0.054950997 = queryWeight, product of:
                1.7051253 = boost
                2.4209464 = idf(docFreq=10677, maxDocs=44218)
                0.013311718 = queryNorm
              0.21398345 = fieldWeight in 5134, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                2.4209464 = idf(docFreq=10677, maxDocs=44218)
                0.0625 = fieldNorm(doc=5134)
          0.2354769 = weight(abstract_txt:biological in 5134) [ClassicSimilarity], result of:
            0.2354769 = score(doc=5134,freq=1.0), product of:
              0.5105606 = queryWeight, product of:
                5.197472 = boost
                7.3793993 = idf(docFreq=74, maxDocs=44218)
                0.013311718 = queryNorm
              0.46121246 = fieldWeight in 5134, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.3793993 = idf(docFreq=74, maxDocs=44218)
                0.0625 = fieldNorm(doc=5134)
        0.24 = coord(6/25)
    
  5. Park, H.; You, S.; Wolfram, D.: Informal data citation for data sharing and reuse is more common than formal data citation in biomedical fields (2018) 0.09
    0.08505397 = sum of:
      0.08505397 = product of:
        0.42526984 = sum of:
          0.020881759 = weight(abstract_txt:sources in 4544) [ClassicSimilarity], result of:
            0.020881759 = score(doc=4544,freq=1.0), product of:
              0.070396766 = queryWeight, product of:
                1.1142541 = boost
                4.7460723 = idf(docFreq=1043, maxDocs=44218)
                0.013311718 = queryNorm
              0.29662952 = fieldWeight in 4544, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.7460723 = idf(docFreq=1043, maxDocs=44218)
                0.0625 = fieldNorm(doc=4544)
          0.046362128 = weight(abstract_txt:extraction in 4544) [ClassicSimilarity], result of:
            0.046362128 = score(doc=4544,freq=1.0), product of:
              0.11980738 = queryWeight, product of:
                1.4536159 = boost
                6.1915555 = idf(docFreq=245, maxDocs=44218)
                0.013311718 = queryNorm
              0.38697222 = fieldWeight in 4544, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.1915555 = idf(docFreq=245, maxDocs=44218)
                0.0625 = fieldNorm(doc=4544)
          0.024217775 = weight(abstract_txt:software in 4544) [ClassicSimilarity], result of:
            0.024217775 = score(doc=4544,freq=1.0), product of:
              0.08895313 = queryWeight, product of:
                1.534031 = boost
                4.3560514 = idf(docFreq=1541, maxDocs=44218)
                0.013311718 = queryNorm
              0.27225322 = fieldWeight in 4544, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.3560514 = idf(docFreq=1541, maxDocs=44218)
                0.0625 = fieldNorm(doc=4544)
          0.09833128 = weight(abstract_txt:data in 4544) [ClassicSimilarity], result of:
            0.09833128 = score(doc=4544,freq=15.0), product of:
              0.12175721 = queryWeight, product of:
                2.7415063 = boost
                3.3363478 = idf(docFreq=4274, maxDocs=44218)
                0.013311718 = queryNorm
              0.8076013 = fieldWeight in 4544, product of:
                3.8729835 = tf(freq=15.0), with freq of:
                  15.0 = termFreq=15.0
                3.3363478 = idf(docFreq=4274, maxDocs=44218)
                0.0625 = fieldNorm(doc=4544)
          0.2354769 = weight(abstract_txt:biological in 4544) [ClassicSimilarity], result of:
            0.2354769 = score(doc=4544,freq=1.0), product of:
              0.5105606 = queryWeight, product of:
                5.197472 = boost
                7.3793993 = idf(docFreq=74, maxDocs=44218)
                0.013311718 = queryNorm
              0.46121246 = fieldWeight in 4544, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.3793993 = idf(docFreq=74, maxDocs=44218)
                0.0625 = fieldNorm(doc=4544)
        0.2 = coord(5/25)