CoFFEe-Qwen: A Large Language Model for Chinese Financial Sentiment Analysis Using the Contrastive Learning and Fine-Tuning Paradigm.

Saved in:
Bibliographic Details
Title: CoFFEe-Qwen: A Large Language Model for Chinese Financial Sentiment Analysis Using the Contrastive Learning and Fine-Tuning Paradigm.
Authors: Feng, Wenfang1 1036784024@qq.com, Yang, Chen2 867320505@qq.com, Zhao, Minrui2 2904104078@qq.com, Xia, Zhiyuan2 xia15094905773@163.com, Wang, Fufu2 wangfufu2001@163.com
Source: IAENG International Journal of Computer Science. Jul2026, Vol. 53 Issue 7, p2526-2539. 14p.
Subjects: Contrastive learning, Spelling errors, Machine learning, Language models, Market sentiment
Abstract: Financial sentiment analysis is increasingly recognized as a pivotal component of stock market research, providing a more accurate and efficient means of quantifying and interpreting market sentiment. However, stock commentaries--one of the most prevalent forms of financial text on social media--are often unstructured and rife with typographical errors and colloquial expressions. To mitigate the lack of publicly available Chinese financial sentiment analysis datasets, we constructed a dedicated expert-annotated corpus and further validated the model's robustness on the large-scale Eastmoney Guba benchmark. To address challenges such as misclassification arising from the non-standard nature of stock commentary, we introduce CoFFEe-Qwen (Contrastive Fine-tuned Financial Embedding for Qwen), a novel framework that integrates supervised contrastive learning with parameter-efficient fine-tuning of large language models (LLMs). This approach fully exploits the transfer learning capabilities of LLMs. Furthermore, the CoFFEe mechanism enables the incorporation of heterogeneous domain data, thereby improving the model's representational capacity for financial texts with limited samples, alleviating the scarcity of social media financial data, and enhancing the overall quality of learned feature representations. In addition, to handle typographical errors and colloquial expressions commonly found in stock commentaries, we design a homophonic character perturbation mechanism that improves the model's robustness and encoding effectiveness when processing noisy input text. Experimental results show that CoFFEe-Qwen consistently outperforms baseline models across multiple metrics--including accuracy, F1-score, precision, and recall--demonstrating its clear superiority in sentiment analysis tasks involving stock commentaries with typographical noise and informal language. [ABSTRACT FROM AUTHOR]
Copyright of IAENG International Journal of Computer Science is the property of International Association of Engineers (IAENG) and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Description
Abstract:Financial sentiment analysis is increasingly recognized as a pivotal component of stock market research, providing a more accurate and efficient means of quantifying and interpreting market sentiment. However, stock commentaries--one of the most prevalent forms of financial text on social media--are often unstructured and rife with typographical errors and colloquial expressions. To mitigate the lack of publicly available Chinese financial sentiment analysis datasets, we constructed a dedicated expert-annotated corpus and further validated the model's robustness on the large-scale Eastmoney Guba benchmark. To address challenges such as misclassification arising from the non-standard nature of stock commentary, we introduce CoFFEe-Qwen (Contrastive Fine-tuned Financial Embedding for Qwen), a novel framework that integrates supervised contrastive learning with parameter-efficient fine-tuning of large language models (LLMs). This approach fully exploits the transfer learning capabilities of LLMs. Furthermore, the CoFFEe mechanism enables the incorporation of heterogeneous domain data, thereby improving the model's representational capacity for financial texts with limited samples, alleviating the scarcity of social media financial data, and enhancing the overall quality of learned feature representations. In addition, to handle typographical errors and colloquial expressions commonly found in stock commentaries, we design a homophonic character perturbation mechanism that improves the model's robustness and encoding effectiveness when processing noisy input text. Experimental results show that CoFFEe-Qwen consistently outperforms baseline models across multiple metrics--including accuracy, F1-score, precision, and recall--demonstrating its clear superiority in sentiment analysis tasks involving stock commentaries with typographical noise and informal language. [ABSTRACT FROM AUTHOR]
ISSN:1819656X