Big data for Natural Language Processing: A streaming approach.
Saved in:
| Title: | Big data for Natural Language Processing: A streaming approach. |
|---|---|
| Authors: | Agerri, Rodrigo1 rodrigo.agerri@ehu.es, Artola, Xabier1 xabier.artola@ehu.es, Beloki, Zuhaitz1 zuhaitz.beloki@ehu.es, Rigau, German1 german.rigau@ehu.es, Soroa, Aitor1 a.soroa@ehu.es |
| Source: | Knowledge-Based Systems. May2015, Vol. 79, p36-42. 7p. |
| Subjects: | Big data, Natural language processing, Electronic data processing, Computer architecture, Distributed computing |
| Abstract: | Requirements in computational power have grown dramatically in recent years. This is also the case in many language processing tasks, due to the overwhelming and ever increasing amount of textual information that must be processed in a reasonable time frame. This scenario has led to a paradigm shift in the computing architectures and large-scale data processing strategies used in the Natural Language Processing field. In this paper we present a new distributed architecture and technology for scaling up text analysis running a complete chain of linguistic processors on several virtual machines. Furthermore, we also describe a series of experiments carried out with the goal of analyzing the scaling capabilities of the language processing pipeline used in this setting. We explore the use of Storm in a new approach for scalable distributed language processing across multiple machines and evaluate its effectiveness and efficiency when processing documents on a medium and large scale. The experiments have shown that there is a big room for improvement regarding language processing performance when adopting parallel architectures, and that we might expect even better results with the use of large clusters with many processing nodes. [ABSTRACT FROM AUTHOR] |
| Copyright of Knowledge-Based Systems is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 101941841 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Big data for Natural Language Processing: A streaming approach. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Agerri%2C+Rodrigo%22">Agerri, Rodrigo</searchLink><relatesTo>1</relatesTo><i> rodrigo.agerri@ehu.es</i><br /><searchLink fieldCode="AR" term="%22Artola%2C+Xabier%22">Artola, Xabier</searchLink><relatesTo>1</relatesTo><i> xabier.artola@ehu.es</i><br /><searchLink fieldCode="AR" term="%22Beloki%2C+Zuhaitz%22">Beloki, Zuhaitz</searchLink><relatesTo>1</relatesTo><i> zuhaitz.beloki@ehu.es</i><br /><searchLink fieldCode="AR" term="%22Rigau%2C+German%22">Rigau, German</searchLink><relatesTo>1</relatesTo><i> german.rigau@ehu.es</i><br /><searchLink fieldCode="AR" term="%22Soroa%2C+Aitor%22">Soroa, Aitor</searchLink><relatesTo>1</relatesTo><i> a.soroa@ehu.es</i> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Knowledge-Based+Systems%22">Knowledge-Based Systems</searchLink>. May2015, Vol. 79, p36-42. 7p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Big+data%22">Big data</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+language+processing%22">Natural language processing</searchLink><br /><searchLink fieldCode="DE" term="%22Electronic+data+processing%22">Electronic data processing</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+architecture%22">Computer architecture</searchLink><br /><searchLink fieldCode="DE" term="%22Distributed+computing%22">Distributed computing</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Requirements in computational power have grown dramatically in recent years. This is also the case in many language processing tasks, due to the overwhelming and ever increasing amount of textual information that must be processed in a reasonable time frame. This scenario has led to a paradigm shift in the computing architectures and large-scale data processing strategies used in the Natural Language Processing field. In this paper we present a new distributed architecture and technology for scaling up text analysis running a complete chain of linguistic processors on several virtual machines. Furthermore, we also describe a series of experiments carried out with the goal of analyzing the scaling capabilities of the language processing pipeline used in this setting. We explore the use of Storm in a new approach for scalable distributed language processing across multiple machines and evaluate its effectiveness and efficiency when processing documents on a medium and large scale. The experiments have shown that there is a big room for improvement regarding language processing performance when adopting parallel architectures, and that we might expect even better results with the use of large clusters with many processing nodes. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Knowledge-Based Systems is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=101941841 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1016/j.knosys.2014.11.007 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 7 StartPage: 36 Subjects: – SubjectFull: Big data Type: general – SubjectFull: Natural language processing Type: general – SubjectFull: Electronic data processing Type: general – SubjectFull: Computer architecture Type: general – SubjectFull: Distributed computing Type: general Titles: – TitleFull: Big data for Natural Language Processing: A streaming approach. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Agerri, Rodrigo – PersonEntity: Name: NameFull: Artola, Xabier – PersonEntity: Name: NameFull: Beloki, Zuhaitz – PersonEntity: Name: NameFull: Rigau, German – PersonEntity: Name: NameFull: Soroa, Aitor IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 05 Text: May2015 Type: published Y: 2015 Identifiers: – Type: issn-print Value: 09507051 Numbering: – Type: volume Value: 79 Titles: – TitleFull: Knowledge-Based Systems Type: main |
| ResultId | 1 |