Semi-automatic generation of multilingual datasets for stance detection in Twitter.

Saved in:
Bibliographic Details
Title: Semi-automatic generation of multilingual datasets for stance detection in Twitter.
Authors: Zotova, Elena1 (AUTHOR) ezotova@vicomtech.org, Agerri, Rodrigo1,2 (AUTHOR) rodrigo.agerri@ehu.eus, Rigau, German2 (AUTHOR) german.rigau@ehu.eus
Source: Expert Systems with Applications. May2021, Vol. 170, pN.PAG-N.PAG. 1p.
Subjects: Natural language processing, Skewness (Probability theory), Natural languages, Social interaction, Social media
Abstract: • New method to semi-automatically build labeled stance detection datasets from Twitter. • Translation strategies outperform zero-shot approaches when data is translated to a high-resourced language. • User-based information helps to label individual tweets. • Our method is applicable to quickly and cheaply generate labeled Twitter-based data. Popular social media networks provide the perfect environment to study the opinions and attitudes expressed by users. While interactions in social media such as Twitter occur in many natural languages, research on stance detection (the position or attitude expressed with respect to a specific topic) within the Natural Language Processing field has largely been done for English. Although some efforts have recently been made to develop annotated data in other languages, there is a telling lack of resources to facilitate multilingual and crosslingual research on stance detection. This is partially due to the fact that manually annotating a corpus of social media texts is a difficult, slow and costly process. Furthermore, as stance is a highly domain- and topic-specific phenomenon, the need for annotated data is specially demanding. As a result, most of the manually labeled resources are hindered by their relatively small size and skewed class distribution. This paper presents a method to obtain multilingual datasets for stance detection in Twitter. Instead of manually annotating on a per tweet basis, we leverage user-based information to semi-automatically label large amounts of tweets. Empirical monolingual and cross-lingual experimentation and qualitative analysis show that our method helps to overcome the aforementioned difficulties to build large, balanced and multilingual labeled corpora. We believe that our method can be easily adapted to easily generate labeled social media data for other Natural Language Processing tasks and domains. [ABSTRACT FROM AUTHOR]
Copyright of Expert Systems with Applications is the property of Pergamon Press - An Imprint of Elsevier Science and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 148986708
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Semi-automatic generation of multilingual datasets for stance detection in Twitter.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Zotova%2C+Elena%22">Zotova, Elena</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> ezotova@vicomtech.org</i><br /><searchLink fieldCode="AR" term="%22Agerri%2C+Rodrigo%22">Agerri, Rodrigo</searchLink><relatesTo>1,2</relatesTo> (AUTHOR)<i> rodrigo.agerri@ehu.eus</i><br /><searchLink fieldCode="AR" term="%22Rigau%2C+German%22">Rigau, German</searchLink><relatesTo>2</relatesTo> (AUTHOR)<i> german.rigau@ehu.eus</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Expert+Systems+with+Applications%22">Expert Systems with Applications</searchLink>. May2021, Vol. 170, pN.PAG-N.PAG. 1p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Natural+language+processing%22">Natural language processing</searchLink><br /><searchLink fieldCode="DE" term="%22Skewness+%28Probability+theory%29%22">Skewness (Probability theory)</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+languages%22">Natural languages</searchLink><br /><searchLink fieldCode="DE" term="%22Social+interaction%22">Social interaction</searchLink><br /><searchLink fieldCode="DE" term="%22Social+media%22">Social media</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: • New method to semi-automatically build labeled stance detection datasets from Twitter. • Translation strategies outperform zero-shot approaches when data is translated to a high-resourced language. • User-based information helps to label individual tweets. • Our method is applicable to quickly and cheaply generate labeled Twitter-based data. Popular social media networks provide the perfect environment to study the opinions and attitudes expressed by users. While interactions in social media such as Twitter occur in many natural languages, research on stance detection (the position or attitude expressed with respect to a specific topic) within the Natural Language Processing field has largely been done for English. Although some efforts have recently been made to develop annotated data in other languages, there is a telling lack of resources to facilitate multilingual and crosslingual research on stance detection. This is partially due to the fact that manually annotating a corpus of social media texts is a difficult, slow and costly process. Furthermore, as stance is a highly domain- and topic-specific phenomenon, the need for annotated data is specially demanding. As a result, most of the manually labeled resources are hindered by their relatively small size and skewed class distribution. This paper presents a method to obtain multilingual datasets for stance detection in Twitter. Instead of manually annotating on a per tweet basis, we leverage user-based information to semi-automatically label large amounts of tweets. Empirical monolingual and cross-lingual experimentation and qualitative analysis show that our method helps to overcome the aforementioned difficulties to build large, balanced and multilingual labeled corpora. We believe that our method can be easily adapted to easily generate labeled social media data for other Natural Language Processing tasks and domains. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Expert Systems with Applications is the property of Pergamon Press - An Imprint of Elsevier Science and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=148986708
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1016/j.eswa.2020.114547
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 1
        StartPage: N.PAG
    Subjects:
      – SubjectFull: Natural language processing
        Type: general
      – SubjectFull: Skewness (Probability theory)
        Type: general
      – SubjectFull: Natural languages
        Type: general
      – SubjectFull: Social interaction
        Type: general
      – SubjectFull: Social media
        Type: general
    Titles:
      – TitleFull: Semi-automatic generation of multilingual datasets for stance detection in Twitter.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Zotova, Elena
      – PersonEntity:
          Name:
            NameFull: Agerri, Rodrigo
      – PersonEntity:
          Name:
            NameFull: Rigau, German
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 15
              M: 05
              Text: May2021
              Type: published
              Y: 2021
          Identifiers:
            – Type: issn-print
              Value: 09574174
          Numbering:
            – Type: volume
              Value: 170
          Titles:
            – TitleFull: Expert Systems with Applications
              Type: main
ResultId 1