C-BiLDA extracting cross-lingual topics from non-parallel texts by distinguishing shared from unshared content.

Saved in:
Bibliographic Details
Title: C-BiLDA extracting cross-lingual topics from non-parallel texts by distinguishing shared from unshared content.
Authors: Heyman, Geert1 geert.heyman@cs.kuleuven.be, Vulić, Ivan1, Moens, Marie-Francine1
Source: Data Mining & Knowledge Discovery. Sep2016, Vol. 30 Issue 5, p1299-1323. 25p.
Subjects: Multilingual computing, Text mining, Benchmark problems (Computer science), Knowledge transfer, Approximation algorithms
Abstract: We study the problem of extracting cross-lingual topics from non-parallel multilingual text datasets with partially overlapping thematic content (e.g., aligned Wikipedia articles in two different languages). To this end, we develop a new bilingual probabilistic topic model called comparable bilingual latent Dirichlet allocation (C-BiLDA), which is able to deal with such comparable data, and, unlike the standard bilingual LDA model (BiLDA), does not assume the availability of document pairs with identical topic distributions. We present a full overview of C-BiLDA, and show its utility in the task of cross-lingual knowledge transfer for multi-class document classification on two benchmarking datasets for three language pairs. The proposed model outperforms the baseline LDA model, as well as the standard BiLDA model and two standard low-rank approximation methods (CL-LSI and CL-KCCA) used in previous work on this task. [ABSTRACT FROM AUTHOR]
Copyright of Data Mining & Knowledge Discovery is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Links:
  – Type: pdflink
Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 117633272
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: C-BiLDA extracting cross-lingual topics from non-parallel texts by distinguishing shared from unshared content.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Heyman%2C+Geert%22">Heyman, Geert</searchLink><relatesTo>1</relatesTo><i> geert.heyman@cs.kuleuven.be</i><br /><searchLink fieldCode="AR" term="%22Vulić%2C+Ivan%22">Vulić, Ivan</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Moens%2C+Marie-Francine%22">Moens, Marie-Francine</searchLink><relatesTo>1</relatesTo>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Data+Mining+%26+Knowledge+Discovery%22">Data Mining & Knowledge Discovery</searchLink>. Sep2016, Vol. 30 Issue 5, p1299-1323. 25p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Multilingual+computing%22">Multilingual computing</searchLink><br /><searchLink fieldCode="DE" term="%22Text+mining%22">Text mining</searchLink><br /><searchLink fieldCode="DE" term="%22Benchmark+problems+%28Computer+science%29%22">Benchmark problems (Computer science)</searchLink><br /><searchLink fieldCode="DE" term="%22Knowledge+transfer%22">Knowledge transfer</searchLink><br /><searchLink fieldCode="DE" term="%22Approximation+algorithms%22">Approximation algorithms</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: We study the problem of extracting cross-lingual topics from non-parallel multilingual text datasets with partially overlapping thematic content (e.g., aligned Wikipedia articles in two different languages). To this end, we develop a new bilingual probabilistic topic model called comparable bilingual latent Dirichlet allocation (C-BiLDA), which is able to deal with such comparable data, and, unlike the standard bilingual LDA model (BiLDA), does not assume the availability of document pairs with identical topic distributions. We present a full overview of C-BiLDA, and show its utility in the task of cross-lingual knowledge transfer for multi-class document classification on two benchmarking datasets for three language pairs. The proposed model outperforms the baseline LDA model, as well as the standard BiLDA model and two standard low-rank approximation methods (CL-LSI and CL-KCCA) used in previous work on this task. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Data Mining & Knowledge Discovery is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=117633272
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1007/s10618-015-0442-x
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 25
        StartPage: 1299
    Subjects:
      – SubjectFull: Multilingual computing
        Type: general
      – SubjectFull: Text mining
        Type: general
      – SubjectFull: Benchmark problems (Computer science)
        Type: general
      – SubjectFull: Knowledge transfer
        Type: general
      – SubjectFull: Approximation algorithms
        Type: general
    Titles:
      – TitleFull: C-BiLDA extracting cross-lingual topics from non-parallel texts by distinguishing shared from unshared content.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Heyman, Geert
      – PersonEntity:
          Name:
            NameFull: Vulić, Ivan
      – PersonEntity:
          Name:
            NameFull: Moens, Marie-Francine
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 09
              Text: Sep2016
              Type: published
              Y: 2016
          Identifiers:
            – Type: issn-print
              Value: 13845810
          Numbering:
            – Type: volume
              Value: 30
            – Type: issue
              Value: 5
          Titles:
            – TitleFull: Data Mining & Knowledge Discovery
              Type: main
ResultId 1