Robust multilingual Named Entity Recognition with shallow semi-supervised features.
Saved in:
| Title: | Robust multilingual Named Entity Recognition with shallow semi-supervised features. |
|---|---|
| Authors: | Agerri, Rodrigo1 rodrigo.agerri@ehu.eus, Rigau, German1 german.rigau@ehu.eus |
| Source: | Artificial Intelligence. Sep2016, Vol. 238, p63-82. 20p. |
| Subjects: | Natural language processing, Supervised learning, Data mining, Document clustering, Task performance |
| Abstract: | We present a multilingual Named Entity Recognition approach based on a robust and general set of features across languages and datasets. Our system combines shallow local information with clustering semi-supervised features induced on large amounts of unlabeled text. Understanding via empirical experimentation how to effectively combine various types of clustering features allows us to seamlessly export our system to other datasets and languages. The result is a simple but highly competitive system which obtains state of the art results across five languages and twelve datasets. The results are reported on standard shared task evaluation data such as CoNLL for English, Spanish and Dutch. Furthermore, and despite the lack of linguistically motivated features, we also report best results for languages such as Basque and German. In addition, we demonstrate that our method also obtains very competitive results even when the amount of supervised data is cut by half, alleviating the dependency on manually annotated data. Finally, the results show that our emphasis on clustering features is crucial to develop robust out-of-domain models. The system and models are freely available to facilitate its use and guarantee the reproducibility of results. [ABSTRACT FROM AUTHOR] |
| Copyright of Artificial Intelligence is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 118151381 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Robust multilingual Named Entity Recognition with shallow semi-supervised features. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Agerri%2C+Rodrigo%22">Agerri, Rodrigo</searchLink><relatesTo>1</relatesTo><i> rodrigo.agerri@ehu.eus</i><br /><searchLink fieldCode="AR" term="%22Rigau%2C+German%22">Rigau, German</searchLink><relatesTo>1</relatesTo><i> german.rigau@ehu.eus</i> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink>. Sep2016, Vol. 238, p63-82. 20p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Natural+language+processing%22">Natural language processing</searchLink><br /><searchLink fieldCode="DE" term="%22Supervised+learning%22">Supervised learning</searchLink><br /><searchLink fieldCode="DE" term="%22Data+mining%22">Data mining</searchLink><br /><searchLink fieldCode="DE" term="%22Document+clustering%22">Document clustering</searchLink><br /><searchLink fieldCode="DE" term="%22Task+performance%22">Task performance</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: We present a multilingual Named Entity Recognition approach based on a robust and general set of features across languages and datasets. Our system combines shallow local information with clustering semi-supervised features induced on large amounts of unlabeled text. Understanding via empirical experimentation how to effectively combine various types of clustering features allows us to seamlessly export our system to other datasets and languages. The result is a simple but highly competitive system which obtains state of the art results across five languages and twelve datasets. The results are reported on standard shared task evaluation data such as CoNLL for English, Spanish and Dutch. Furthermore, and despite the lack of linguistically motivated features, we also report best results for languages such as Basque and German. In addition, we demonstrate that our method also obtains very competitive results even when the amount of supervised data is cut by half, alleviating the dependency on manually annotated data. Finally, the results show that our emphasis on clustering features is crucial to develop robust out-of-domain models. The system and models are freely available to facilitate its use and guarantee the reproducibility of results. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Artificial Intelligence is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=118151381 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1016/j.artint.2016.05.003 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 20 StartPage: 63 Subjects: – SubjectFull: Natural language processing Type: general – SubjectFull: Supervised learning Type: general – SubjectFull: Data mining Type: general – SubjectFull: Document clustering Type: general – SubjectFull: Task performance Type: general Titles: – TitleFull: Robust multilingual Named Entity Recognition with shallow semi-supervised features. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Agerri, Rodrigo – PersonEntity: Name: NameFull: Rigau, German IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 09 Text: Sep2016 Type: published Y: 2016 Identifiers: – Type: issn-print Value: 00043702 Numbering: – Type: volume Value: 238 Titles: – TitleFull: Artificial Intelligence Type: main |
| ResultId | 1 |