Document (#38537)

Author
Nagy T., I.
Title
Detecting multiword expressions and named entities in natural language texts
Imprint
Szeged : University of Szeged, Faculty of Science and Informatics, Doctoral School of Computer Science
Year
2014
Pages
XVIII,
Abstract
Multiword expressions (MWEs) are lexical items that can be decomposed into single words and display lexical, syntactic, semantic, pragmatic and/or statistical idiosyncrasy (Sag et al., 2002; Kim, 2008; Calzolari et al., 2002). The proper treatment of multiword expressions such as rock 'n' roll and make a decision is essential for many natural language processing (NLP) applications like information extraction and retrieval, terminology extraction and machine translation, and it is important to identify multiword expressions in context. For example, in machine translation we must know that MWEs form one semantic unit, hence their parts should not be translated separately. For this, multiword expressions should be identified first in the text to be translated. The chief aim of this thesis is to develop machine learning-based approaches for the automatic detection of different types of multiword expressions in English and Hungarian natural language texts. In our investigations, we pay attention to the characteristics of different types of multiword expressions such as nominal compounds, multiword named entities and light verb constructions, and we apply novel methods to identify MWEs in raw texts. In the thesis it will be demonstrated that nominal compounds and multiword amed entities may require a similar approach for their automatic detection as they behave in the same way from a linguistic point of view. Furthermore, it will be shown that the automatic detection of light verb constructions can be carried out using two effective machine learning-based approaches.
In this thesis, we focused on the automatic detection of multiword expressions in natural language texts. On the basis of the main contributions, we can argue that: - Supervised machine learning methods can be successfully applied for the automatic detection of different types of multiword expressions in natural language texts. - Machine learning-based multiword expression detection can be successfully carried out for English as well as for Hungarian. - Our supervised machine learning-based model was successfully applied to the automatic detection of nominal compounds from English raw texts. - We developed a Wikipedia-based dictionary labeling method to automatically detect English nominal compounds. - A prior knowledge of nominal compounds can enhance Named Entity Recognition, while previously identified named entities can assist the nominal compound identification process. - The machine learning-based method can also provide acceptable results when it was trained on an automatically generated silver standard corpus. - As named entities form one semantic unit and may consist of more than one word and function as a noun, we can treat them in a similar way to nominal compounds. - Our sequence labelling-based tool can be successfully applied for identifying verbal light verb constructions in two typologically different languages, namely English and Hungarian. - Domain adaptation techniques may help diminish the distance between domains in the automatic detection of light verb constructions. - Our syntax-based method can be successfully applied for the full-coverage identification of light verb constructions. As a first step, a data-driven candidate extraction method can be utilized. After, a machine learning approach that makes use of an extended and rich feature set selects LVCs among extracted candidates. - When a precise syntactic parser is available for the actual domain, the full-coverage identification can be performed better. In other cases, the usage of the sequence labeling method is recommended.
Content
Vgl.: http://doktori.bibl.u-szeged.hu/2434/1/main.pdf.
Theme
Computerlinguistik

Similar documents (content)

  1. Ramisch, C.: Multiword expressions acquisition : a generic and open framework (2015) 0.41
    0.41014746 = sum of:
      0.41014746 = product of:
        1.2817109 = sum of:
          0.029335098 = weight(abstract_txt:language in 1649) [ClassicSimilarity], result of:
            0.029335098 = score(doc=1649,freq=4.0), product of:
              0.056115706 = queryWeight, product of:
                1.1257503 = boost
                4.1820874 = idf(docFreq=1834, maxDocs=44218)
                0.011919258 = queryNorm
              0.5227609 = fieldWeight in 1649, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                4.1820874 = idf(docFreq=1834, maxDocs=44218)
                0.0625 = fieldNorm(doc=1649)
          0.026280712 = weight(abstract_txt:natural in 1649) [ClassicSimilarity], result of:
            0.026280712 = score(doc=1649,freq=1.0), product of:
              0.0827823 = queryWeight, product of:
                1.367315 = boost
                5.0794845 = idf(docFreq=747, maxDocs=44218)
                0.011919258 = queryNorm
              0.31746778 = fieldWeight in 1649, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.0794845 = idf(docFreq=747, maxDocs=44218)
                0.0625 = fieldNorm(doc=1649)
          0.061517756 = weight(abstract_txt:texts in 1649) [ClassicSimilarity], result of:
            0.061517756 = score(doc=1649,freq=2.0), product of:
              0.123092085 = queryWeight, product of:
                1.8264406 = boost
                5.6542544 = idf(docFreq=420, maxDocs=44218)
                0.011919258 = queryNorm
              0.4997702 = fieldWeight in 1649, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                5.6542544 = idf(docFreq=420, maxDocs=44218)
                0.0625 = fieldNorm(doc=1649)
          0.039374296 = weight(abstract_txt:automatic in 1649) [ClassicSimilarity], result of:
            0.039374296 = score(doc=1649,freq=1.0), product of:
              0.12125433 = queryWeight, product of:
                1.9579992 = boost
                5.1955976 = idf(docFreq=665, maxDocs=44218)
                0.011919258 = queryNorm
              0.32472485 = fieldWeight in 1649, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.1955976 = idf(docFreq=665, maxDocs=44218)
                0.0625 = fieldNorm(doc=1649)
          0.1221969 = weight(abstract_txt:constructions in 1649) [ClassicSimilarity], result of:
            0.1221969 = score(doc=1649,freq=1.0), product of:
              0.23061427 = queryWeight, product of:
                2.2821436 = boost
                8.478011 = idf(docFreq=24, maxDocs=44218)
                0.011919258 = queryNorm
              0.5298757 = fieldWeight in 1649, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                8.478011 = idf(docFreq=24, maxDocs=44218)
                0.0625 = fieldNorm(doc=1649)
          0.05308696 = weight(abstract_txt:machine in 1649) [ClassicSimilarity], result of:
            0.05308696 = score(doc=1649,freq=1.0), product of:
              0.1609146 = queryWeight, product of:
                2.5576072 = boost
                5.2785225 = idf(docFreq=612, maxDocs=44218)
                0.011919258 = queryNorm
              0.32990766 = fieldWeight in 1649, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.2785225 = idf(docFreq=612, maxDocs=44218)
                0.0625 = fieldNorm(doc=1649)
          0.25284505 = weight(abstract_txt:expressions in 1649) [ClassicSimilarity], result of:
            0.25284505 = score(doc=1649,freq=5.0), product of:
              0.26638913 = queryWeight, product of:
                3.290746 = boost
                6.7916126 = idf(docFreq=134, maxDocs=44218)
                0.011919258 = queryNorm
              0.94915676 = fieldWeight in 1649, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                6.7916126 = idf(docFreq=134, maxDocs=44218)
                0.0625 = fieldNorm(doc=1649)
          0.6970741 = weight(abstract_txt:multiword in 1649) [ClassicSimilarity], result of:
            0.6970741 = score(doc=1649,freq=5.0), product of:
              0.5764732 = queryWeight, product of:
                5.589784 = boost
                8.652365 = idf(docFreq=20, maxDocs=44218)
                0.011919258 = queryNorm
              1.2092048 = fieldWeight in 1649, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                8.652365 = idf(docFreq=20, maxDocs=44218)
                0.0625 = fieldNorm(doc=1649)
        0.32 = coord(8/25)
    
  2. Gödert, W.: Detecting multiword phrases in mathematical text corpora (2012) 0.25
    0.24647203 = sum of:
      0.24647203 = product of:
        1.0269668 = sum of:
          0.03878796 = weight(abstract_txt:method in 466) [ClassicSimilarity], result of:
            0.03878796 = score(doc=466,freq=2.0), product of:
              0.06499898 = queryWeight, product of:
                1.2115829 = boost
                4.50095 = idf(docFreq=1333, maxDocs=44218)
                0.011919258 = queryNorm
              0.5967473 = fieldWeight in 466, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                4.50095 = idf(docFreq=1333, maxDocs=44218)
                0.09375 = fieldNorm(doc=466)
          0.022051077 = weight(abstract_txt:based in 466) [ClassicSimilarity], result of:
            0.022051077 = score(doc=466,freq=2.0), product of:
              0.05217171 = queryWeight, product of:
                1.3730217 = boost
                3.1879277 = idf(docFreq=4958, maxDocs=44218)
                0.011919258 = queryNorm
              0.42266348 = fieldWeight in 466, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.1879277 = idf(docFreq=4958, maxDocs=44218)
                0.09375 = fieldNorm(doc=466)
          0.09548726 = weight(abstract_txt:named in 466) [ClassicSimilarity], result of:
            0.09548726 = score(doc=466,freq=1.0), product of:
              0.14930768 = queryWeight, product of:
                1.8362887 = boost
                6.82169 = idf(docFreq=130, maxDocs=44218)
                0.011919258 = queryNorm
              0.63953346 = fieldWeight in 466, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.82169 = idf(docFreq=130, maxDocs=44218)
                0.09375 = fieldNorm(doc=466)
          0.05906144 = weight(abstract_txt:automatic in 466) [ClassicSimilarity], result of:
            0.05906144 = score(doc=466,freq=1.0), product of:
              0.12125433 = queryWeight, product of:
                1.9579992 = boost
                5.1955976 = idf(docFreq=665, maxDocs=44218)
                0.011919258 = queryNorm
              0.48708728 = fieldWeight in 466, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.1955976 = idf(docFreq=665, maxDocs=44218)
                0.09375 = fieldNorm(doc=466)
          0.15027665 = weight(abstract_txt:detection in 466) [ClassicSimilarity], result of:
            0.15027665 = score(doc=466,freq=1.0), product of:
              0.23627597 = queryWeight, product of:
                2.921929 = boost
                6.784232 = idf(docFreq=135, maxDocs=44218)
                0.011919258 = queryNorm
              0.63602173 = fieldWeight in 466, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.784232 = idf(docFreq=135, maxDocs=44218)
                0.09375 = fieldNorm(doc=466)
          0.6613025 = weight(abstract_txt:multiword in 466) [ClassicSimilarity], result of:
            0.6613025 = score(doc=466,freq=2.0), product of:
              0.5764732 = queryWeight, product of:
                5.589784 = boost
                8.652365 = idf(docFreq=20, maxDocs=44218)
                0.011919258 = queryNorm
              1.1471523 = fieldWeight in 466, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                8.652365 = idf(docFreq=20, maxDocs=44218)
                0.09375 = fieldNorm(doc=466)
        0.24 = coord(6/25)
    
  3. Snajder, J.; Almic, P.: Modeling semantic compositionality of Croatian multiword expressions (2015) 0.24
    0.24089487 = sum of:
      0.24089487 = product of:
        1.0037286 = sum of:
          0.022001324 = weight(abstract_txt:language in 2920) [ClassicSimilarity], result of:
            0.022001324 = score(doc=2920,freq=1.0), product of:
              0.056115706 = queryWeight, product of:
                1.1257503 = boost
                4.1820874 = idf(docFreq=1834, maxDocs=44218)
                0.011919258 = queryNorm
              0.3920707 = fieldWeight in 2920, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.1820874 = idf(docFreq=1834, maxDocs=44218)
                0.09375 = fieldNorm(doc=2920)
          0.039421067 = weight(abstract_txt:natural in 2920) [ClassicSimilarity], result of:
            0.039421067 = score(doc=2920,freq=1.0), product of:
              0.0827823 = queryWeight, product of:
                1.367315 = boost
                5.0794845 = idf(docFreq=747, maxDocs=44218)
                0.011919258 = queryNorm
              0.47620165 = fieldWeight in 2920, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.0794845 = idf(docFreq=747, maxDocs=44218)
                0.09375 = fieldNorm(doc=2920)
          0.027006945 = weight(abstract_txt:based in 2920) [ClassicSimilarity], result of:
            0.027006945 = score(doc=2920,freq=3.0), product of:
              0.05217171 = queryWeight, product of:
                1.3730217 = boost
                3.1879277 = idf(docFreq=4958, maxDocs=44218)
                0.011919258 = queryNorm
              0.51765496 = fieldWeight in 2920, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                3.1879277 = idf(docFreq=4958, maxDocs=44218)
                0.09375 = fieldNorm(doc=2920)
          0.27807415 = weight(abstract_txt:mwes in 2920) [ClassicSimilarity], result of:
            0.27807415 = score(doc=2920,freq=3.0), product of:
              0.17806105 = queryWeight, product of:
                1.5533165 = boost
                9.617446 = idf(docFreq=7, maxDocs=44218)
                0.011919258 = queryNorm
              1.5616786 = fieldWeight in 2920, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                9.617446 = idf(docFreq=7, maxDocs=44218)
                0.09375 = fieldNorm(doc=2920)
          0.1696136 = weight(abstract_txt:expressions in 2920) [ClassicSimilarity], result of:
            0.1696136 = score(doc=2920,freq=1.0), product of:
              0.26638913 = queryWeight, product of:
                3.290746 = boost
                6.7916126 = idf(docFreq=134, maxDocs=44218)
                0.011919258 = queryNorm
              0.6367137 = fieldWeight in 2920, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.7916126 = idf(docFreq=134, maxDocs=44218)
                0.09375 = fieldNorm(doc=2920)
          0.46761152 = weight(abstract_txt:multiword in 2920) [ClassicSimilarity], result of:
            0.46761152 = score(doc=2920,freq=1.0), product of:
              0.5764732 = queryWeight, product of:
                5.589784 = boost
                8.652365 = idf(docFreq=20, maxDocs=44218)
                0.011919258 = queryNorm
              0.8111592 = fieldWeight in 2920, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                8.652365 = idf(docFreq=20, maxDocs=44218)
                0.09375 = fieldNorm(doc=2920)
        0.24 = coord(6/25)
    
  4. Cruys, T. van de; Moirón, B.V.: Semantics-based multiword expression extraction (2007) 0.23
    0.2307574 = sum of:
      0.2307574 = product of:
        0.9614892 = sum of:
          0.0428371 = weight(abstract_txt:extraction in 2919) [ClassicSimilarity], result of:
            0.0428371 = score(doc=2919,freq=1.0), product of:
              0.073798746 = queryWeight, product of:
                6.1915555 = idf(docFreq=245, maxDocs=44218)
                0.011919258 = queryNorm
              0.58045834 = fieldWeight in 2919, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.1915555 = idf(docFreq=245, maxDocs=44218)
                0.09375 = fieldNorm(doc=2919)
          0.03878796 = weight(abstract_txt:method in 2919) [ClassicSimilarity], result of:
            0.03878796 = score(doc=2919,freq=2.0), product of:
              0.06499898 = queryWeight, product of:
                1.2115829 = boost
                4.50095 = idf(docFreq=1333, maxDocs=44218)
                0.011919258 = queryNorm
              0.5967473 = fieldWeight in 2919, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                4.50095 = idf(docFreq=1333, maxDocs=44218)
                0.09375 = fieldNorm(doc=2919)
          0.015592467 = weight(abstract_txt:based in 2919) [ClassicSimilarity], result of:
            0.015592467 = score(doc=2919,freq=1.0), product of:
              0.05217171 = queryWeight, product of:
                1.3730217 = boost
                3.1879277 = idf(docFreq=4958, maxDocs=44218)
                0.011919258 = queryNorm
              0.29886824 = fieldWeight in 2919, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.1879277 = idf(docFreq=4958, maxDocs=44218)
                0.09375 = fieldNorm(doc=2919)
          0.22704658 = weight(abstract_txt:mwes in 2919) [ClassicSimilarity], result of:
            0.22704658 = score(doc=2919,freq=2.0), product of:
              0.17806105 = queryWeight, product of:
                1.5533165 = boost
                9.617446 = idf(docFreq=7, maxDocs=44218)
                0.011919258 = queryNorm
              1.2751052 = fieldWeight in 2919, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                9.617446 = idf(docFreq=7, maxDocs=44218)
                0.09375 = fieldNorm(doc=2919)
          0.1696136 = weight(abstract_txt:expressions in 2919) [ClassicSimilarity], result of:
            0.1696136 = score(doc=2919,freq=1.0), product of:
              0.26638913 = queryWeight, product of:
                3.290746 = boost
                6.7916126 = idf(docFreq=134, maxDocs=44218)
                0.011919258 = queryNorm
              0.6367137 = fieldWeight in 2919, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.7916126 = idf(docFreq=134, maxDocs=44218)
                0.09375 = fieldNorm(doc=2919)
          0.46761152 = weight(abstract_txt:multiword in 2919) [ClassicSimilarity], result of:
            0.46761152 = score(doc=2919,freq=1.0), product of:
              0.5764732 = queryWeight, product of:
                5.589784 = boost
                8.652365 = idf(docFreq=20, maxDocs=44218)
                0.011919258 = queryNorm
              0.8111592 = fieldWeight in 2919, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                8.652365 = idf(docFreq=20, maxDocs=44218)
                0.09375 = fieldNorm(doc=2919)
        0.24 = coord(6/25)
    
  5. Nissim, M.; Zaninello, A,: Modeling the internal variability of multiword expressions through a pattern-based method (2013) 0.21
    0.21257576 = sum of:
      0.21257576 = product of:
        0.75919914 = sum of:
          0.040387202 = weight(abstract_txt:extraction in 990) [ClassicSimilarity], result of:
            0.040387202 = score(doc=990,freq=2.0), product of:
              0.073798746 = queryWeight, product of:
                6.1915555 = idf(docFreq=245, maxDocs=44218)
                0.011919258 = queryNorm
              0.54726136 = fieldWeight in 990, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                6.1915555 = idf(docFreq=245, maxDocs=44218)
                0.0625 = fieldNorm(doc=990)
          0.02585864 = weight(abstract_txt:method in 990) [ClassicSimilarity], result of:
            0.02585864 = score(doc=990,freq=2.0), product of:
              0.06499898 = queryWeight, product of:
                1.2115829 = boost
                4.50095 = idf(docFreq=1333, maxDocs=44218)
                0.011919258 = queryNorm
              0.3978315 = fieldWeight in 990, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                4.50095 = idf(docFreq=1333, maxDocs=44218)
                0.0625 = fieldNorm(doc=990)
          0.014700718 = weight(abstract_txt:based in 990) [ClassicSimilarity], result of:
            0.014700718 = score(doc=990,freq=2.0), product of:
              0.05217171 = queryWeight, product of:
                1.3730217 = boost
                3.1879277 = idf(docFreq=4958, maxDocs=44218)
                0.011919258 = queryNorm
              0.28177565 = fieldWeight in 990, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.1879277 = idf(docFreq=4958, maxDocs=44218)
                0.0625 = fieldNorm(doc=990)
          0.21406157 = weight(abstract_txt:mwes in 990) [ClassicSimilarity], result of:
            0.21406157 = score(doc=990,freq=4.0), product of:
              0.17806105 = queryWeight, product of:
                1.5533165 = boost
                9.617446 = idf(docFreq=7, maxDocs=44218)
                0.011919258 = queryNorm
              1.2021807 = fieldWeight in 990, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                9.617446 = idf(docFreq=7, maxDocs=44218)
                0.0625 = fieldNorm(doc=990)
          0.039374296 = weight(abstract_txt:automatic in 990) [ClassicSimilarity], result of:
            0.039374296 = score(doc=990,freq=1.0), product of:
              0.12125433 = queryWeight, product of:
                1.9579992 = boost
                5.1955976 = idf(docFreq=665, maxDocs=44218)
                0.011919258 = queryNorm
              0.32472485 = fieldWeight in 990, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.1955976 = idf(docFreq=665, maxDocs=44218)
                0.0625 = fieldNorm(doc=990)
          0.11307573 = weight(abstract_txt:expressions in 990) [ClassicSimilarity], result of:
            0.11307573 = score(doc=990,freq=1.0), product of:
              0.26638913 = queryWeight, product of:
                3.290746 = boost
                6.7916126 = idf(docFreq=134, maxDocs=44218)
                0.011919258 = queryNorm
              0.4244758 = fieldWeight in 990, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.7916126 = idf(docFreq=134, maxDocs=44218)
                0.0625 = fieldNorm(doc=990)
          0.31174102 = weight(abstract_txt:multiword in 990) [ClassicSimilarity], result of:
            0.31174102 = score(doc=990,freq=1.0), product of:
              0.5764732 = queryWeight, product of:
                5.589784 = boost
                8.652365 = idf(docFreq=20, maxDocs=44218)
                0.011919258 = queryNorm
              0.5407728 = fieldWeight in 990, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                8.652365 = idf(docFreq=20, maxDocs=44218)
                0.0625 = fieldNorm(doc=990)
        0.28 = coord(7/25)