AI and Measurement Concerns: Dealing with Imbalanced Data in Autoscoring

Saved in:
Bibliographic Details
Title: AI and Measurement Concerns: Dealing with Imbalanced Data in Autoscoring
Language: English
Authors: Yunting Liu (ORCID 0009-0004-9594-9661), Yijun Xiang, Xutao Feng, Mark Wilson (ORCID 0000-0002-0425-5305)
Source: Journal of Educational Measurement. 2026 63(1).
Availability: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 21
Publication Date: 2026
Document Type: Journal Articles
Reports - Research
Descriptors: Automation, Scores, Bias, Accuracy, Prediction, Classification, Algorithms, Data Analysis, Measurement, Methods, Evaluation Methods, Sampling, Technology Uses in Education, Artificial Intelligence
DOI: 10.1111/jedm.70031
ISSN: 0022-0655
1745-3984
Abstract: Unbiasedness for proficiency estimates is important for autoscoring engines since the outcome might be used for future learning or placement. Imbalanced training data may lead to certain biases and lower the prediction accuracy for classification algorithms. In this article, we investigated several data augmentation methods to lower the negative effect of imbalanced data in measurement settings. Four approaches were examined: (1) Resampling methods, either oversampling or undersampling; (2) Active resampling methods, where the resampling weight is based on representativeness in the training set; (3) Data expansion methods using synonym Replacement, slightly changing the meaning or semantics of the original answers; and (4) Content recreation method using Generative AI (e.g., ChatGPT) to create responses for less populated scores. We compared the performance (e.g., Accuracy, QWK, F1) as well as the distance metric for different combinations of the methods. Two datasets with different imbalanced distributions were used. Results show that all four methods can help to mitigate the bias issue and the efficacy was influenced by the imbalance level, representativeness of the original data and the level of increment in the variety of the response (i.e., lexical diversity). In general, resampling and GenAI with active resampling showed the best overall performance.
Abstractor: As Provided
Notes: https://osf.io/cr2s6/?view_only=a356d94ef26342aaa5b19674c558169c
Entry Date: 2026
Accession Number: EJ1501394
Database: ERIC
Description
Abstract:Unbiasedness for proficiency estimates is important for autoscoring engines since the outcome might be used for future learning or placement. Imbalanced training data may lead to certain biases and lower the prediction accuracy for classification algorithms. In this article, we investigated several data augmentation methods to lower the negative effect of imbalanced data in measurement settings. Four approaches were examined: (1) Resampling methods, either oversampling or undersampling; (2) Active resampling methods, where the resampling weight is based on representativeness in the training set; (3) Data expansion methods using synonym Replacement, slightly changing the meaning or semantics of the original answers; and (4) Content recreation method using Generative AI (e.g., ChatGPT) to create responses for less populated scores. We compared the performance (e.g., Accuracy, QWK, F1) as well as the distance metric for different combinations of the methods. Two datasets with different imbalanced distributions were used. Results show that all four methods can help to mitigate the bias issue and the efficacy was influenced by the imbalance level, representativeness of the original data and the level of increment in the variety of the response (i.e., lexical diversity). In general, resampling and GenAI with active resampling showed the best overall performance.
ISSN:0022-0655
1745-3984
DOI:10.1111/jedm.70031