Cross-Project Generalization Challenges in Transformer-Based Code Smell Detection: An Empirical Study.

Saved in:
Bibliographic Details
Title: Cross-Project Generalization Challenges in Transformer-Based Code Smell Detection: An Empirical Study.
Authors: Burra, Bhavana Chowdary1 burrabhavana@gmail.com, Shukla, Seema1, Goyal, Mayank Kumar1
Source: International Journal of Performability Engineering. Jun2026, Vol. 22 Issue 6, p318-330. 13p.
Subjects: Transformer models, Machine learning, Computer software quality control, Maintainability (Engineering), Optimization algorithms, Skewness (Probability theory)
Abstract: Detecting code smells is very important for increasing software maintainability and lowering the technical debt of large-scale software systems. Traditional machine learning methods rely heavily on manually engineered features and, as a result, can struggle to generalize across projects due to domain differences and class imbalance in the datasets. However, although transformer-based pre-trained models have shown great promise in understanding the semantics of source code, there has been limited investigation into how well they perform across different datasets, particularly balanced versus imbalanced ones. In this study, we compare the performance of baseline machine learning models and transformer-based models for detecting multiple types of code smells on two heterogeneous datasets with different distribution properties. From the analysis, we see that the degree of imbalance in the datasets and the differences between the two domains significantly affect the performance and generalization of the various models. Our experimental results show that whilst transformer-based models outperform baseline machine learning models, the extent of their advantage varies with dataset characteristics; therefore, transformer-based models do not generalize well across projects. We have also found that providing domain-specific fine-tuning strategies can improve adaptability and detection performance in real-world use. This study provides insights into dataset characteristics, model behavior across domains, and the need for adaptive learning approaches to develop robust, generalized code smell detection systems. [ABSTRACT FROM AUTHOR]
Copyright of International Journal of Performability Engineering is the property of Totem Publisher, Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Description
Abstract:Detecting code smells is very important for increasing software maintainability and lowering the technical debt of large-scale software systems. Traditional machine learning methods rely heavily on manually engineered features and, as a result, can struggle to generalize across projects due to domain differences and class imbalance in the datasets. However, although transformer-based pre-trained models have shown great promise in understanding the semantics of source code, there has been limited investigation into how well they perform across different datasets, particularly balanced versus imbalanced ones. In this study, we compare the performance of baseline machine learning models and transformer-based models for detecting multiple types of code smells on two heterogeneous datasets with different distribution properties. From the analysis, we see that the degree of imbalance in the datasets and the differences between the two domains significantly affect the performance and generalization of the various models. Our experimental results show that whilst transformer-based models outperform baseline machine learning models, the extent of their advantage varies with dataset characteristics; therefore, transformer-based models do not generalize well across projects. We have also found that providing domain-specific fine-tuning strategies can improve adaptability and detection performance in real-world use. This study provides insights into dataset characteristics, model behavior across domains, and the need for adaptive learning approaches to develop robust, generalized code smell detection systems. [ABSTRACT FROM AUTHOR]
ISSN:09731318
DOI:10.23940/ijpe.26.06.p3.318330