Optimal strategies to perform multilingual analysis of social content for a novel dataset in the tourism domain.

Saved in:
Bibliographic Details
Title: Optimal strategies to perform multilingual analysis of social content for a novel dataset in the tourism domain.
Authors: Masson, Maxime1,2 (AUTHOR) maxime.masson@univ-pau.fr, Agerri, Rodrigo2 (AUTHOR) rodrigo.agerri@ehu.eus, Sallaberry, Christian1 (AUTHOR) christian.sallaberry@univ-pau.fr, Bessagnet, Marie-Noelle1 (AUTHOR) marie-noelle.bessagnet@univ-pau.fr, Le Parc Lacayrelle, Annig1 (AUTHOR) annig.lacayrelle@univ-pau.fr, Roose, Philippe1 (AUTHOR) philippe.roose@univ-pau.fr
Source: Knowledge-Based Systems. Sep2025, Vol. 326, pN.PAG-N.PAG. 1p.
Subjects: Tourism, Natural language processing, Acquisition of data, Social media, Data mining, Sentiment analysis
Abstract: The rising influence of social media platforms in various domains, including tourism, has highlighted the growing need for efficient and automated Natural Language Processing (NLP) strategies to take advantage of this valuable resource. However, the transformation of multilingual, unstructured, and informal texts into structured knowledge still poses significant challenges, most notably the never-ending requirement for manually annotated data to train deep learning classifiers. In this work, we study different NLP techniques to establish the best ones to obtain competitive performances while keeping the need for training annotated data to a minimum. To do so, we built the first publicly available multilingual dataset (French, English, and Spanish) for the tourism domain, composed of tourism-related tweets. The dataset includes multilayered, manually revised annotations for Named Entity Recognition (NER) for Locations and Fine-grained Thematic Concepts Extraction mapped to the Thesaurus of Tourism and Leisure Activities of the World Tourism Organization, as well as for Sentiment Analysis at the tweet level. Extensive experimentation comparing various few-shot and fine-tuning techniques with modern language models demonstrate that modern few-shot techniques allow us to obtain competitive results for all three tasks with very little annotation data: 5 tweets per label (15 in total) for Sentiment Analysis, 30 tweets for Named Entity Recognition of Locations and 1K tweets annotated with fine-grained thematic concepts, a highly fine-grained sequence labeling task based on an inventory of 315 classes. We believe that our results, grounded in a novel dataset, pave the way for applying NLP to new domain-specific applications, reducing the need for manual annotations and circumventing the complexities of rule-based, ad-hoc solutions. [ABSTRACT FROM AUTHOR]
Copyright of Knowledge-Based Systems is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 188445548
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Optimal strategies to perform multilingual analysis of social content for a novel dataset in the tourism domain.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Masson%2C+Maxime%22">Masson, Maxime</searchLink><relatesTo>1,2</relatesTo> (AUTHOR)<i> maxime.masson@univ-pau.fr</i><br /><searchLink fieldCode="AR" term="%22Agerri%2C+Rodrigo%22">Agerri, Rodrigo</searchLink><relatesTo>2</relatesTo> (AUTHOR)<i> rodrigo.agerri@ehu.eus</i><br /><searchLink fieldCode="AR" term="%22Sallaberry%2C+Christian%22">Sallaberry, Christian</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> christian.sallaberry@univ-pau.fr</i><br /><searchLink fieldCode="AR" term="%22Bessagnet%2C+Marie-Noelle%22">Bessagnet, Marie-Noelle</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> marie-noelle.bessagnet@univ-pau.fr</i><br /><searchLink fieldCode="AR" term="%22Le+Parc+Lacayrelle%2C+Annig%22">Le Parc Lacayrelle, Annig</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> annig.lacayrelle@univ-pau.fr</i><br /><searchLink fieldCode="AR" term="%22Roose%2C+Philippe%22">Roose, Philippe</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> philippe.roose@univ-pau.fr</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Knowledge-Based+Systems%22">Knowledge-Based Systems</searchLink>. Sep2025, Vol. 326, pN.PAG-N.PAG. 1p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Tourism%22">Tourism</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+language+processing%22">Natural language processing</searchLink><br /><searchLink fieldCode="DE" term="%22Acquisition+of+data%22">Acquisition of data</searchLink><br /><searchLink fieldCode="DE" term="%22Social+media%22">Social media</searchLink><br /><searchLink fieldCode="DE" term="%22Data+mining%22">Data mining</searchLink><br /><searchLink fieldCode="DE" term="%22Sentiment+analysis%22">Sentiment analysis</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: The rising influence of social media platforms in various domains, including tourism, has highlighted the growing need for efficient and automated Natural Language Processing (NLP) strategies to take advantage of this valuable resource. However, the transformation of multilingual, unstructured, and informal texts into structured knowledge still poses significant challenges, most notably the never-ending requirement for manually annotated data to train deep learning classifiers. In this work, we study different NLP techniques to establish the best ones to obtain competitive performances while keeping the need for training annotated data to a minimum. To do so, we built the first publicly available multilingual dataset (French, English, and Spanish) for the tourism domain, composed of tourism-related tweets. The dataset includes multilayered, manually revised annotations for Named Entity Recognition (NER) for Locations and Fine-grained Thematic Concepts Extraction mapped to the Thesaurus of Tourism and Leisure Activities of the World Tourism Organization, as well as for Sentiment Analysis at the tweet level. Extensive experimentation comparing various few-shot and fine-tuning techniques with modern language models demonstrate that modern few-shot techniques allow us to obtain competitive results for all three tasks with very little annotation data: 5 tweets per label (15 in total) for Sentiment Analysis, 30 tweets for Named Entity Recognition of Locations and 1K tweets annotated with fine-grained thematic concepts, a highly fine-grained sequence labeling task based on an inventory of 315 classes. We believe that our results, grounded in a novel dataset, pave the way for applying NLP to new domain-specific applications, reducing the need for manual annotations and circumventing the complexities of rule-based, ad-hoc solutions. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Knowledge-Based Systems is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=188445548
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1016/j.knosys.2025.114001
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 1
        StartPage: N.PAG
    Subjects:
      – SubjectFull: Tourism
        Type: general
      – SubjectFull: Natural language processing
        Type: general
      – SubjectFull: Acquisition of data
        Type: general
      – SubjectFull: Social media
        Type: general
      – SubjectFull: Data mining
        Type: general
      – SubjectFull: Sentiment analysis
        Type: general
    Titles:
      – TitleFull: Optimal strategies to perform multilingual analysis of social content for a novel dataset in the tourism domain.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Masson, Maxime
      – PersonEntity:
          Name:
            NameFull: Agerri, Rodrigo
      – PersonEntity:
          Name:
            NameFull: Sallaberry, Christian
      – PersonEntity:
          Name:
            NameFull: Bessagnet, Marie-Noelle
      – PersonEntity:
          Name:
            NameFull: Le Parc Lacayrelle, Annig
      – PersonEntity:
          Name:
            NameFull: Roose, Philippe
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 27
              M: 09
              Text: Sep2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 09507051
          Numbering:
            – Type: volume
              Value: 326
          Titles:
            – TitleFull: Knowledge-Based Systems
              Type: main
ResultId 1