NewsReader: Using knowledge resources in a cross-lingual reading machine to generate more knowledge from massive streams of news.
Saved in:
| Title: | NewsReader: Using knowledge resources in a cross-lingual reading machine to generate more knowledge from massive streams of news. |
|---|---|
| Authors: | Vossen, Piek1 piek.vossen@vu.nl, Agerri, Rodrigo2, Aldabe, Itziar2, Cybulska, Agata1, van Erp, Marieke1, Fokkens, Antske1, Laparra, Egoitz2, Minard, Anne-Lyse3, Palmero Aprosio, Alessio3, Rigau, German2, Rospocher, Marco3, Segers, Roxane1 |
| Source: | Knowledge-Based Systems. Oct2016, Vol. 110, p60-85. 26p. |
| Subjects: | Reading machines (Data processing equipment), Streaming technology, RDF (Document markup language), Multilingualism, Semantic Web, Performance evaluation |
| Abstract: | In this article, we describe a system that reads news articles in four different languages and detects what happened, who is involved, where and when. This event-centric information is represented as episodic situational knowledge on individuals in an interoperable RDF format that allows for reasoning on the implications of the events. Our system covers the complete path from unstructured text to structured knowledge, for which we defined a formal model that links interpreted textual mentions of things to their representation as instances. The model forms the skeleton for interoperable interpretation across different sources and languages. The real content, however, is defined using multilingual and cross-lingual knowledge resources, both semantic and episodic. We explain how these knowledge resources are used for the processing of text and ultimately define the actual content of the episodic situational knowledge that is reported in the news. The knowledge and model in our system can be seen as an example how the Semantic Web helps NLP. However, our systems also generate massive episodic knowledge of the same type as the Semantic Web is built on. We thus envision a cycle of knowledge acquisition and NLP improvement on a massive scale. This article reports on the details of the system but also on the performance of various high-level components. We demonstrate that our system performs at state-of-the-art level for various subtasks in the four languages of the project, but that we also consider the full integration of these tasks in an overall system with the purpose of reading text. We applied our system to millions of news articles, generating billions of triples expressing formal semantic properties. This shows the capacity of the system to perform at an unprecedented scale. [ABSTRACT FROM AUTHOR] |
| Copyright of Knowledge-Based Systems is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 118026164 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: NewsReader: Using knowledge resources in a cross-lingual reading machine to generate more knowledge from massive streams of news. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Vossen%2C+Piek%22">Vossen, Piek</searchLink><relatesTo>1</relatesTo><i> piek.vossen@vu.nl</i><br /><searchLink fieldCode="AR" term="%22Agerri%2C+Rodrigo%22">Agerri, Rodrigo</searchLink><relatesTo>2</relatesTo><br /><searchLink fieldCode="AR" term="%22Aldabe%2C+Itziar%22">Aldabe, Itziar</searchLink><relatesTo>2</relatesTo><br /><searchLink fieldCode="AR" term="%22Cybulska%2C+Agata%22">Cybulska, Agata</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22van+Erp%2C+Marieke%22">van Erp, Marieke</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Fokkens%2C+Antske%22">Fokkens, Antske</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Laparra%2C+Egoitz%22">Laparra, Egoitz</searchLink><relatesTo>2</relatesTo><br /><searchLink fieldCode="AR" term="%22Minard%2C+Anne-Lyse%22">Minard, Anne-Lyse</searchLink><relatesTo>3</relatesTo><br /><searchLink fieldCode="AR" term="%22Palmero+Aprosio%2C+Alessio%22">Palmero Aprosio, Alessio</searchLink><relatesTo>3</relatesTo><br /><searchLink fieldCode="AR" term="%22Rigau%2C+German%22">Rigau, German</searchLink><relatesTo>2</relatesTo><br /><searchLink fieldCode="AR" term="%22Rospocher%2C+Marco%22">Rospocher, Marco</searchLink><relatesTo>3</relatesTo><br /><searchLink fieldCode="AR" term="%22Segers%2C+Roxane%22">Segers, Roxane</searchLink><relatesTo>1</relatesTo> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Knowledge-Based+Systems%22">Knowledge-Based Systems</searchLink>. Oct2016, Vol. 110, p60-85. 26p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Reading+machines+%28Data+processing+equipment%29%22">Reading machines (Data processing equipment)</searchLink><br /><searchLink fieldCode="DE" term="%22Streaming+technology%22">Streaming technology</searchLink><br /><searchLink fieldCode="DE" term="%22RDF+%28Document+markup+language%29%22">RDF (Document markup language)</searchLink><br /><searchLink fieldCode="DE" term="%22Multilingualism%22">Multilingualism</searchLink><br /><searchLink fieldCode="DE" term="%22Semantic+Web%22">Semantic Web</searchLink><br /><searchLink fieldCode="DE" term="%22Performance+evaluation%22">Performance evaluation</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: In this article, we describe a system that reads news articles in four different languages and detects what happened, who is involved, where and when. This event-centric information is represented as episodic situational knowledge on individuals in an interoperable RDF format that allows for reasoning on the implications of the events. Our system covers the complete path from unstructured text to structured knowledge, for which we defined a formal model that links interpreted textual mentions of things to their representation as instances. The model forms the skeleton for interoperable interpretation across different sources and languages. The real content, however, is defined using multilingual and cross-lingual knowledge resources, both semantic and episodic. We explain how these knowledge resources are used for the processing of text and ultimately define the actual content of the episodic situational knowledge that is reported in the news. The knowledge and model in our system can be seen as an example how the Semantic Web helps NLP. However, our systems also generate massive episodic knowledge of the same type as the Semantic Web is built on. We thus envision a cycle of knowledge acquisition and NLP improvement on a massive scale. This article reports on the details of the system but also on the performance of various high-level components. We demonstrate that our system performs at state-of-the-art level for various subtasks in the four languages of the project, but that we also consider the full integration of these tasks in an overall system with the purpose of reading text. We applied our system to millions of news articles, generating billions of triples expressing formal semantic properties. This shows the capacity of the system to perform at an unprecedented scale. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Knowledge-Based Systems is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=118026164 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1016/j.knosys.2016.07.013 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 26 StartPage: 60 Subjects: – SubjectFull: Reading machines (Data processing equipment) Type: general – SubjectFull: Streaming technology Type: general – SubjectFull: RDF (Document markup language) Type: general – SubjectFull: Multilingualism Type: general – SubjectFull: Semantic Web Type: general – SubjectFull: Performance evaluation Type: general Titles: – TitleFull: NewsReader: Using knowledge resources in a cross-lingual reading machine to generate more knowledge from massive streams of news. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Vossen, Piek – PersonEntity: Name: NameFull: Agerri, Rodrigo – PersonEntity: Name: NameFull: Aldabe, Itziar – PersonEntity: Name: NameFull: Cybulska, Agata – PersonEntity: Name: NameFull: van Erp, Marieke – PersonEntity: Name: NameFull: Fokkens, Antske – PersonEntity: Name: NameFull: Laparra, Egoitz – PersonEntity: Name: NameFull: Minard, Anne-Lyse – PersonEntity: Name: NameFull: Palmero Aprosio, Alessio – PersonEntity: Name: NameFull: Rigau, German – PersonEntity: Name: NameFull: Rospocher, Marco – PersonEntity: Name: NameFull: Segers, Roxane IsPartOfRelationships: – BibEntity: Dates: – D: 15 M: 10 Text: Oct2016 Type: published Y: 2016 Identifiers: – Type: issn-print Value: 09507051 Numbering: – Type: volume Value: 110 Titles: – TitleFull: Knowledge-Based Systems Type: main |
| ResultId | 1 |