AI and Measurement Concerns: Dealing with Imbalanced Data in Autoscoring
Saved in:
| Title: | AI and Measurement Concerns: Dealing with Imbalanced Data in Autoscoring |
|---|---|
| Language: | English |
| Authors: | Yunting Liu (ORCID |
| Source: | Journal of Educational Measurement. 2026 63(1). |
| Availability: | Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us |
| Peer Reviewed: | Y |
| Page Count: | 21 |
| Publication Date: | 2026 |
| Document Type: | Journal Articles Reports - Research |
| Descriptors: | Automation, Scores, Bias, Accuracy, Prediction, Classification, Algorithms, Data Analysis, Measurement, Methods, Evaluation Methods, Sampling, Technology Uses in Education, Artificial Intelligence |
| DOI: | 10.1111/jedm.70031 |
| ISSN: | 0022-0655 1745-3984 |
| Abstract: | Unbiasedness for proficiency estimates is important for autoscoring engines since the outcome might be used for future learning or placement. Imbalanced training data may lead to certain biases and lower the prediction accuracy for classification algorithms. In this article, we investigated several data augmentation methods to lower the negative effect of imbalanced data in measurement settings. Four approaches were examined: (1) Resampling methods, either oversampling or undersampling; (2) Active resampling methods, where the resampling weight is based on representativeness in the training set; (3) Data expansion methods using synonym Replacement, slightly changing the meaning or semantics of the original answers; and (4) Content recreation method using Generative AI (e.g., ChatGPT) to create responses for less populated scores. We compared the performance (e.g., Accuracy, QWK, F1) as well as the distance metric for different combinations of the methods. Two datasets with different imbalanced distributions were used. Results show that all four methods can help to mitigate the bias issue and the efficacy was influenced by the imbalance level, representativeness of the original data and the level of increment in the variety of the response (i.e., lexical diversity). In general, resampling and GenAI with active resampling showed the best overall performance. |
| Abstractor: | As Provided |
| Notes: | https://osf.io/cr2s6/?view_only=a356d94ef26342aaa5b19674c558169c |
| Entry Date: | 2026 |
| Accession Number: | EJ1501394 |
| Database: | ERIC |
| Abstract: | Unbiasedness for proficiency estimates is important for autoscoring engines since the outcome might be used for future learning or placement. Imbalanced training data may lead to certain biases and lower the prediction accuracy for classification algorithms. In this article, we investigated several data augmentation methods to lower the negative effect of imbalanced data in measurement settings. Four approaches were examined: (1) Resampling methods, either oversampling or undersampling; (2) Active resampling methods, where the resampling weight is based on representativeness in the training set; (3) Data expansion methods using synonym Replacement, slightly changing the meaning or semantics of the original answers; and (4) Content recreation method using Generative AI (e.g., ChatGPT) to create responses for less populated scores. We compared the performance (e.g., Accuracy, QWK, F1) as well as the distance metric for different combinations of the methods. Two datasets with different imbalanced distributions were used. Results show that all four methods can help to mitigate the bias issue and the efficacy was influenced by the imbalance level, representativeness of the original data and the level of increment in the variety of the response (i.e., lexical diversity). In general, resampling and GenAI with active resampling showed the best overall performance. |
|---|---|
| ISSN: | 0022-0655 1745-3984 |
| DOI: | 10.1111/jedm.70031 |