Robust multilingual Named Entity Recognition with shallow semi-supervised features.

Saved in:
Bibliographic Details
Title: Robust multilingual Named Entity Recognition with shallow semi-supervised features.
Authors: Agerri, Rodrigo1 rodrigo.agerri@ehu.eus, Rigau, German1 german.rigau@ehu.eus
Source: Artificial Intelligence. Sep2016, Vol. 238, p63-82. 20p.
Subjects: Natural language processing, Supervised learning, Data mining, Document clustering, Task performance
Abstract: We present a multilingual Named Entity Recognition approach based on a robust and general set of features across languages and datasets. Our system combines shallow local information with clustering semi-supervised features induced on large amounts of unlabeled text. Understanding via empirical experimentation how to effectively combine various types of clustering features allows us to seamlessly export our system to other datasets and languages. The result is a simple but highly competitive system which obtains state of the art results across five languages and twelve datasets. The results are reported on standard shared task evaluation data such as CoNLL for English, Spanish and Dutch. Furthermore, and despite the lack of linguistically motivated features, we also report best results for languages such as Basque and German. In addition, we demonstrate that our method also obtains very competitive results even when the amount of supervised data is cut by half, alleviating the dependency on manually annotated data. Finally, the results show that our emphasis on clustering features is crucial to develop robust out-of-domain models. The system and models are freely available to facilitate its use and guarantee the reproducibility of results. [ABSTRACT FROM AUTHOR]
Copyright of Artificial Intelligence is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 118151381
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Robust multilingual Named Entity Recognition with shallow semi-supervised features.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Agerri%2C+Rodrigo%22">Agerri, Rodrigo</searchLink><relatesTo>1</relatesTo><i> rodrigo.agerri@ehu.eus</i><br /><searchLink fieldCode="AR" term="%22Rigau%2C+German%22">Rigau, German</searchLink><relatesTo>1</relatesTo><i> german.rigau@ehu.eus</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink>. Sep2016, Vol. 238, p63-82. 20p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Natural+language+processing%22">Natural language processing</searchLink><br /><searchLink fieldCode="DE" term="%22Supervised+learning%22">Supervised learning</searchLink><br /><searchLink fieldCode="DE" term="%22Data+mining%22">Data mining</searchLink><br /><searchLink fieldCode="DE" term="%22Document+clustering%22">Document clustering</searchLink><br /><searchLink fieldCode="DE" term="%22Task+performance%22">Task performance</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: We present a multilingual Named Entity Recognition approach based on a robust and general set of features across languages and datasets. Our system combines shallow local information with clustering semi-supervised features induced on large amounts of unlabeled text. Understanding via empirical experimentation how to effectively combine various types of clustering features allows us to seamlessly export our system to other datasets and languages. The result is a simple but highly competitive system which obtains state of the art results across five languages and twelve datasets. The results are reported on standard shared task evaluation data such as CoNLL for English, Spanish and Dutch. Furthermore, and despite the lack of linguistically motivated features, we also report best results for languages such as Basque and German. In addition, we demonstrate that our method also obtains very competitive results even when the amount of supervised data is cut by half, alleviating the dependency on manually annotated data. Finally, the results show that our emphasis on clustering features is crucial to develop robust out-of-domain models. The system and models are freely available to facilitate its use and guarantee the reproducibility of results. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Artificial Intelligence is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=118151381
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1016/j.artint.2016.05.003
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 20
        StartPage: 63
    Subjects:
      – SubjectFull: Natural language processing
        Type: general
      – SubjectFull: Supervised learning
        Type: general
      – SubjectFull: Data mining
        Type: general
      – SubjectFull: Document clustering
        Type: general
      – SubjectFull: Task performance
        Type: general
    Titles:
      – TitleFull: Robust multilingual Named Entity Recognition with shallow semi-supervised features.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Agerri, Rodrigo
      – PersonEntity:
          Name:
            NameFull: Rigau, German
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 09
              Text: Sep2016
              Type: published
              Y: 2016
          Identifiers:
            – Type: issn-print
              Value: 00043702
          Numbering:
            – Type: volume
              Value: 238
          Titles:
            – TitleFull: Artificial Intelligence
              Type: main
ResultId 1