Bibliographic Details
| Title: |
A Transformer Model for Manifesto Classification Using Cross-Context Training: An Ecuadorian Case Study. |
| Authors: |
Barzallo, Fernanda1 (AUTHOR), Baldeon-Calisto, Maria1,2 (AUTHOR) mbaldeonc@usfq.edu.ec, Pérez, Margorie1 (AUTHOR), Moscoso, Maria Emilia1 (AUTHOR), Navarrete, Danny1 (AUTHOR), Riofrío, Daniel2 (AUTHOR), Medina-Peréz, Pablo3 (AUTHOR), Lai-Yuen, Susana K4 (AUTHOR), Benítez, Diego2 (AUTHOR), Peréz, Noel2 (AUTHOR), Moyano, Ricardo Flores2 (AUTHOR), Fierro, Mateo3 (AUTHOR) |
| Source: |
Social Science Computer Review. Jun2025, Vol. 43 Issue 3, p578-603. 26p. |
| Subject Terms: |
Natural language processing, Transformer models, Databases, Political manifestoes, Factorial experiment designs |
| Abstract: |
Content analysis of political manifestos is necessary to understand the policies and proposed actions of a party. However, manually labeling political texts is time-consuming and labor-intensive. Transformer networks have become essential tools for automating this task. Nevertheless, these models require extensive datasets to achieve good performance. This can be a limitation in manifesto classification, where the availability of publicly labeled datasets can be scarce. To address this challenge, in this work, we developed a Transformer network for the classification of manifestos using a cross-domain training strategy. Using the database of the Comparative Manifesto Project, we implemented a fractional factorial experimental design to determine which Spanish-written manifestos form the best training set for Ecuadorian manifesto labeling. Furthermore, we statistically analyzed which Transformer architecture and preprocessing operations improve the model accuracy. The results indicate that creating a training set with manifestos from Spain and Uruguay, along with implementing stemming and lemmatization preprocessing operations, produces the highest classification accuracy. In addition, we found that the DistilBERT and RoBERTa transformer networks perform statistically similarly and consistently well in manifesto classification. Using the cross-context training strategy, DistilBERT and RoBERTa achieve 60.05% and 57.64% accuracy, respectively, in the classification of the Ecuadorian manifesto. Finally, we investigated the effect of the composition of the training set on performance. The experiments demonstrate that training DistilBERT solely with Ecuadorian manifestos achieves the highest accuracy and F1-score. Furthermore, in the absence of the Ecuadorian dataset, competitive performance is achieved by training the model with datasets from Spain and Uruguay. [ABSTRACT FROM AUTHOR] |
|
Copyright of Social Science Computer Review is the property of Sage Publications Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) |
| Database: |
Education Research Complete |