OPTIMIZING BERTSNN TO ENHANCE SOURCE-TARGET DOMAIN SIMILARITY SCORING FOR CROSS-DOMAIN SENTIMENT CLASSIFICATION OF PRODUCT REVIEWS.

Saved in:
Bibliographic Details
Title: OPTIMIZING BERTSNN TO ENHANCE SOURCE-TARGET DOMAIN SIMILARITY SCORING FOR CROSS-DOMAIN SENTIMENT CLASSIFICATION OF PRODUCT REVIEWS.
Authors: Haitao Zhao1 andrewzhao@student.usm.my, Yan, Jasy Liew Suet1 jasyliew@usm.my
Source: Malaysian Journal of Computer Science. 2025 Special Issue, Vol. 38, p1-19. 19p.
Subjects: Sentiment analysis, Domain specificity, Product reviews, Language models, Artificial neural networks, Machine learning
Abstract: Cross-domain sentiment analysis (CDSA) predicts sentiment polarity in a target domain using knowledge from source domains but existing CDSA methods lack effective source domain selection strategies. This study investigates BertSNN, which combines pre-trained BERT embeddings, a Siamese neural network, and various distance metrics to measure domain similarity and optimize source domain selection for CDSA. First, we experiment with document-level (DocBERT) and sentence-level (SentenceBERT) embeddings with BiLSTM and BiLSTM + CNN neural network configurations to identify the best combination for BertSNN. Second, we explore two distance metrics--Euclidean and Manhattan--alongside shifted cosine similarity to determine the most effective choice for domain similarity scoring. Using product reviews, we test on 25 target domains, examining whether using multiple top most similar source domains improves cross-domain sentiment classification compared to a single most similar source domain. Results indicate that document-level embeddings, BiLSTM and shifted cosine similarity produce the most optimal BertSNN that can select high-quality similar source domains to train a cross-domain sentiment classifier for a target domain, beating two other traditional baseline methods (i.e., bagof-words and TF-IDF representations). Our findings also show that using top five most similar source domains (k = 5) for training generally improves cross-domain sentiment classification performance as opposed to using a single most similar source domain (k = 1). This study contributes to CDSA by advancing the understanding of embedding choices and distance metrics within a Siamese neural network for source-target domain similarity scoring and providing actionable insights on domain selection strategies to improve sentiment analysis models. [ABSTRACT FROM AUTHOR]
Copyright of Malaysian Journal of Computer Science is the property of University of Malaysia, Faculty of Computer Science & Information Technology and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Links:
  – Type: pdflink
Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 189320523
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: OPTIMIZING BERTSNN TO ENHANCE SOURCE-TARGET DOMAIN SIMILARITY SCORING FOR CROSS-DOMAIN SENTIMENT CLASSIFICATION OF PRODUCT REVIEWS.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Haitao+Zhao%22">Haitao Zhao</searchLink><relatesTo>1</relatesTo><i> andrewzhao@student.usm.my</i><br /><searchLink fieldCode="AR" term="%22Yan%2C+Jasy+Liew+Suet%22">Yan, Jasy Liew Suet</searchLink><relatesTo>1</relatesTo><i> jasyliew@usm.my</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Malaysian+Journal+of+Computer+Science%22">Malaysian Journal of Computer Science</searchLink>. 2025 Special Issue, Vol. 38, p1-19. 19p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Sentiment+analysis%22">Sentiment analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Domain+specificity%22">Domain specificity</searchLink><br /><searchLink fieldCode="DE" term="%22Product+reviews%22">Product reviews</searchLink><br /><searchLink fieldCode="DE" term="%22Language+models%22">Language models</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+neural+networks%22">Artificial neural networks</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Cross-domain sentiment analysis (CDSA) predicts sentiment polarity in a target domain using knowledge from source domains but existing CDSA methods lack effective source domain selection strategies. This study investigates BertSNN, which combines pre-trained BERT embeddings, a Siamese neural network, and various distance metrics to measure domain similarity and optimize source domain selection for CDSA. First, we experiment with document-level (DocBERT) and sentence-level (SentenceBERT) embeddings with BiLSTM and BiLSTM + CNN neural network configurations to identify the best combination for BertSNN. Second, we explore two distance metrics--Euclidean and Manhattan--alongside shifted cosine similarity to determine the most effective choice for domain similarity scoring. Using product reviews, we test on 25 target domains, examining whether using multiple top most similar source domains improves cross-domain sentiment classification compared to a single most similar source domain. Results indicate that document-level embeddings, BiLSTM and shifted cosine similarity produce the most optimal BertSNN that can select high-quality similar source domains to train a cross-domain sentiment classifier for a target domain, beating two other traditional baseline methods (i.e., bagof-words and TF-IDF representations). Our findings also show that using top five most similar source domains (k = 5) for training generally improves cross-domain sentiment classification performance as opposed to using a single most similar source domain (k = 1). This study contributes to CDSA by advancing the understanding of embedding choices and distance metrics within a Siamese neural network for source-target domain similarity scoring and providing actionable insights on domain selection strategies to improve sentiment analysis models. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Malaysian Journal of Computer Science is the property of University of Malaysia, Faculty of Computer Science & Information Technology and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=189320523
RecordInfo BibRecord:
  BibEntity:
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 19
        StartPage: 1
    Subjects:
      – SubjectFull: Sentiment analysis
        Type: general
      – SubjectFull: Domain specificity
        Type: general
      – SubjectFull: Product reviews
        Type: general
      – SubjectFull: Language models
        Type: general
      – SubjectFull: Artificial neural networks
        Type: general
      – SubjectFull: Machine learning
        Type: general
    Titles:
      – TitleFull: OPTIMIZING BERTSNN TO ENHANCE SOURCE-TARGET DOMAIN SIMILARITY SCORING FOR CROSS-DOMAIN SENTIMENT CLASSIFICATION OF PRODUCT REVIEWS.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Haitao Zhao
      – PersonEntity:
          Name:
            NameFull: Yan, Jasy Liew Suet
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 02
              M: 01
              Text: 2025 Special Issue
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 01279084
          Numbering:
            – Type: volume
              Value: 38
          Titles:
            – TitleFull: Malaysian Journal of Computer Science
              Type: main
ResultId 1