A statistical model for grammar mapping.

Saved in:
Bibliographic Details
Title: A statistical model for grammar mapping.
Authors: BASIRAT, A.1,2 ali.basirat@lingfil.uu.se, FAILI, H.1,3 h.faili@ut.ac.ir, NIVRE, J.2 joakim.nivre@lingfil.uu.se
Source: Natural Language Engineering. Mar2016, Vol. 22 Issue 2, p215-255. 41p.
Subjects: Grammar checkers (Computer software), Statistical models, Grammatical categories, Language ability, Grammar handbooks
Abstract: The two main classes of grammars are (a) hand-crafted grammars, which are developed by language experts, and (b) data-driven grammars, which are extracted from annotated corpora. This paper introduces a statistical method for mapping the elementary structures of a data-driven grammar onto the elementary structures of a hand-crafted grammar in order to combine their advantages. The idea is employed in the context of Lexicalized Tree-Adjoining Grammars (LTAG) and tested on two LTAGs of English: the hand-crafted LTAG developed in the XTAG project, and the data-driven LTAG, which is automatically extracted from the Penn Treebank and used by the MICA parser. We propose a statistical model for mapping any elementary tree sequence of the MICA grammar onto a proper elementary tree sequence of the XTAG grammar. The model has been tested on three subsets of the WSJ corpus that have average lengths of 10, 16, and 18 words, respectively. The experimental results show that full-parse trees with average F1-scores of 72.49, 64.80, and 62.30 points could be built from 94.97%, 96.01%, and 90.25% of the XTAG elementary tree sequences assigned to the subsets, respectively. Moreover, by reducing the amount of syntactic lexical ambiguity of sentences, the proposed model significantly improves the efficiency of parsing in the XTAG system. [ABSTRACT FROM PUBLISHER]
Copyright of Natural Language Engineering is the property of Cambridge University Press and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 112852094
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: A statistical model for grammar mapping.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22BASIRAT%2C+A%2E%22">BASIRAT, A.</searchLink><relatesTo>1,2</relatesTo><i> ali.basirat@lingfil.uu.se</i><br /><searchLink fieldCode="AR" term="%22FAILI%2C+H%2E%22">FAILI, H.</searchLink><relatesTo>1,3</relatesTo><i> h.faili@ut.ac.ir</i><br /><searchLink fieldCode="AR" term="%22NIVRE%2C+J%2E%22">NIVRE, J.</searchLink><relatesTo>2</relatesTo><i> joakim.nivre@lingfil.uu.se</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Natural+Language+Engineering%22">Natural Language Engineering</searchLink>. Mar2016, Vol. 22 Issue 2, p215-255. 41p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Grammar+checkers+%28Computer+software%29%22">Grammar checkers (Computer software)</searchLink><br /><searchLink fieldCode="DE" term="%22Statistical+models%22">Statistical models</searchLink><br /><searchLink fieldCode="DE" term="%22Grammatical+categories%22">Grammatical categories</searchLink><br /><searchLink fieldCode="DE" term="%22Language+ability%22">Language ability</searchLink><br /><searchLink fieldCode="DE" term="%22Grammar+handbooks%22">Grammar handbooks</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: The two main classes of grammars are (a) hand-crafted grammars, which are developed by language experts, and (b) data-driven grammars, which are extracted from annotated corpora. This paper introduces a statistical method for mapping the elementary structures of a data-driven grammar onto the elementary structures of a hand-crafted grammar in order to combine their advantages. The idea is employed in the context of Lexicalized Tree-Adjoining Grammars (LTAG) and tested on two LTAGs of English: the hand-crafted LTAG developed in the XTAG project, and the data-driven LTAG, which is automatically extracted from the Penn Treebank and used by the MICA parser. We propose a statistical model for mapping any elementary tree sequence of the MICA grammar onto a proper elementary tree sequence of the XTAG grammar. The model has been tested on three subsets of the WSJ corpus that have average lengths of 10, 16, and 18 words, respectively. The experimental results show that full-parse trees with average F1-scores of 72.49, 64.80, and 62.30 points could be built from 94.97%, 96.01%, and 90.25% of the XTAG elementary tree sequences assigned to the subsets, respectively. Moreover, by reducing the amount of syntactic lexical ambiguity of sentences, the proposed model significantly improves the efficiency of parsing in the XTAG system. [ABSTRACT FROM PUBLISHER]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Natural Language Engineering is the property of Cambridge University Press and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=112852094
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1017/S1351324915000017
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 41
        StartPage: 215
    Subjects:
      – SubjectFull: Grammar checkers (Computer software)
        Type: general
      – SubjectFull: Statistical models
        Type: general
      – SubjectFull: Grammatical categories
        Type: general
      – SubjectFull: Language ability
        Type: general
      – SubjectFull: Grammar handbooks
        Type: general
    Titles:
      – TitleFull: A statistical model for grammar mapping.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: BASIRAT, A.
      – PersonEntity:
          Name:
            NameFull: FAILI, H.
      – PersonEntity:
          Name:
            NameFull: NIVRE, J.
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 03
              Text: Mar2016
              Type: published
              Y: 2016
          Identifiers:
            – Type: issn-print
              Value: 13513249
          Numbering:
            – Type: volume
              Value: 22
            – Type: issue
              Value: 2
          Titles:
            – TitleFull: Natural Language Engineering
              Type: main
ResultId 1