Big data for Natural Language Processing: A streaming approach.

Saved in:
Bibliographic Details
Title: Big data for Natural Language Processing: A streaming approach.
Authors: Agerri, Rodrigo1 rodrigo.agerri@ehu.es, Artola, Xabier1 xabier.artola@ehu.es, Beloki, Zuhaitz1 zuhaitz.beloki@ehu.es, Rigau, German1 german.rigau@ehu.es, Soroa, Aitor1 a.soroa@ehu.es
Source: Knowledge-Based Systems. May2015, Vol. 79, p36-42. 7p.
Subjects: Big data, Natural language processing, Electronic data processing, Computer architecture, Distributed computing
Abstract: Requirements in computational power have grown dramatically in recent years. This is also the case in many language processing tasks, due to the overwhelming and ever increasing amount of textual information that must be processed in a reasonable time frame. This scenario has led to a paradigm shift in the computing architectures and large-scale data processing strategies used in the Natural Language Processing field. In this paper we present a new distributed architecture and technology for scaling up text analysis running a complete chain of linguistic processors on several virtual machines. Furthermore, we also describe a series of experiments carried out with the goal of analyzing the scaling capabilities of the language processing pipeline used in this setting. We explore the use of Storm in a new approach for scalable distributed language processing across multiple machines and evaluate its effectiveness and efficiency when processing documents on a medium and large scale. The experiments have shown that there is a big room for improvement regarding language processing performance when adopting parallel architectures, and that we might expect even better results with the use of large clusters with many processing nodes. [ABSTRACT FROM AUTHOR]
Copyright of Knowledge-Based Systems is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 101941841
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Big data for Natural Language Processing: A streaming approach.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Agerri%2C+Rodrigo%22">Agerri, Rodrigo</searchLink><relatesTo>1</relatesTo><i> rodrigo.agerri@ehu.es</i><br /><searchLink fieldCode="AR" term="%22Artola%2C+Xabier%22">Artola, Xabier</searchLink><relatesTo>1</relatesTo><i> xabier.artola@ehu.es</i><br /><searchLink fieldCode="AR" term="%22Beloki%2C+Zuhaitz%22">Beloki, Zuhaitz</searchLink><relatesTo>1</relatesTo><i> zuhaitz.beloki@ehu.es</i><br /><searchLink fieldCode="AR" term="%22Rigau%2C+German%22">Rigau, German</searchLink><relatesTo>1</relatesTo><i> german.rigau@ehu.es</i><br /><searchLink fieldCode="AR" term="%22Soroa%2C+Aitor%22">Soroa, Aitor</searchLink><relatesTo>1</relatesTo><i> a.soroa@ehu.es</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Knowledge-Based+Systems%22">Knowledge-Based Systems</searchLink>. May2015, Vol. 79, p36-42. 7p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Big+data%22">Big data</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+language+processing%22">Natural language processing</searchLink><br /><searchLink fieldCode="DE" term="%22Electronic+data+processing%22">Electronic data processing</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+architecture%22">Computer architecture</searchLink><br /><searchLink fieldCode="DE" term="%22Distributed+computing%22">Distributed computing</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Requirements in computational power have grown dramatically in recent years. This is also the case in many language processing tasks, due to the overwhelming and ever increasing amount of textual information that must be processed in a reasonable time frame. This scenario has led to a paradigm shift in the computing architectures and large-scale data processing strategies used in the Natural Language Processing field. In this paper we present a new distributed architecture and technology for scaling up text analysis running a complete chain of linguistic processors on several virtual machines. Furthermore, we also describe a series of experiments carried out with the goal of analyzing the scaling capabilities of the language processing pipeline used in this setting. We explore the use of Storm in a new approach for scalable distributed language processing across multiple machines and evaluate its effectiveness and efficiency when processing documents on a medium and large scale. The experiments have shown that there is a big room for improvement regarding language processing performance when adopting parallel architectures, and that we might expect even better results with the use of large clusters with many processing nodes. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Knowledge-Based Systems is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=101941841
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1016/j.knosys.2014.11.007
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 7
        StartPage: 36
    Subjects:
      – SubjectFull: Big data
        Type: general
      – SubjectFull: Natural language processing
        Type: general
      – SubjectFull: Electronic data processing
        Type: general
      – SubjectFull: Computer architecture
        Type: general
      – SubjectFull: Distributed computing
        Type: general
    Titles:
      – TitleFull: Big data for Natural Language Processing: A streaming approach.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Agerri, Rodrigo
      – PersonEntity:
          Name:
            NameFull: Artola, Xabier
      – PersonEntity:
          Name:
            NameFull: Beloki, Zuhaitz
      – PersonEntity:
          Name:
            NameFull: Rigau, German
      – PersonEntity:
          Name:
            NameFull: Soroa, Aitor
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 05
              Text: May2015
              Type: published
              Y: 2015
          Identifiers:
            – Type: issn-print
              Value: 09507051
          Numbering:
            – Type: volume
              Value: 79
          Titles:
            – TitleFull: Knowledge-Based Systems
              Type: main
ResultId 1