AI and Measurement Concerns: Dealing with Imbalanced Data in Autoscoring
Saved in:
| Title: | AI and Measurement Concerns: Dealing with Imbalanced Data in Autoscoring |
|---|---|
| Language: | English |
| Authors: | Yunting Liu (ORCID |
| Source: | Journal of Educational Measurement. 2026 63(1). |
| Availability: | Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us |
| Peer Reviewed: | Y |
| Page Count: | 21 |
| Publication Date: | 2026 |
| Document Type: | Journal Articles Reports - Research |
| Descriptors: | Automation, Scores, Bias, Accuracy, Prediction, Classification, Algorithms, Data Analysis, Measurement, Methods, Evaluation Methods, Sampling, Technology Uses in Education, Artificial Intelligence |
| DOI: | 10.1111/jedm.70031 |
| ISSN: | 0022-0655 1745-3984 |
| Abstract: | Unbiasedness for proficiency estimates is important for autoscoring engines since the outcome might be used for future learning or placement. Imbalanced training data may lead to certain biases and lower the prediction accuracy for classification algorithms. In this article, we investigated several data augmentation methods to lower the negative effect of imbalanced data in measurement settings. Four approaches were examined: (1) Resampling methods, either oversampling or undersampling; (2) Active resampling methods, where the resampling weight is based on representativeness in the training set; (3) Data expansion methods using synonym Replacement, slightly changing the meaning or semantics of the original answers; and (4) Content recreation method using Generative AI (e.g., ChatGPT) to create responses for less populated scores. We compared the performance (e.g., Accuracy, QWK, F1) as well as the distance metric for different combinations of the methods. Two datasets with different imbalanced distributions were used. Results show that all four methods can help to mitigate the bias issue and the efficacy was influenced by the imbalance level, representativeness of the original data and the level of increment in the variety of the response (i.e., lexical diversity). In general, resampling and GenAI with active resampling showed the best overall performance. |
| Abstractor: | As Provided |
| Notes: | https://osf.io/cr2s6/?view_only=a356d94ef26342aaa5b19674c558169c |
| Entry Date: | 2026 |
| Accession Number: | EJ1501394 |
| Database: | ERIC |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1501394 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: AI and Measurement Concerns: Dealing with Imbalanced Data in Autoscoring – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Yunting+Liu%22">Yunting Liu</searchLink> (ORCID <externalLink term="https://orcid.org/0009-0004-9594-9661">0009-0004-9594-9661</externalLink>)<br /><searchLink fieldCode="AR" term="%22Yijun+Xiang%22">Yijun Xiang</searchLink><br /><searchLink fieldCode="AR" term="%22Xutao+Feng%22">Xutao Feng</searchLink><br /><searchLink fieldCode="AR" term="%22Mark+Wilson%22">Mark Wilson</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-0425-5305">0000-0002-0425-5305</externalLink>) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Journal+of+Educational+Measurement%22"><i>Journal of Educational Measurement</i></searchLink>. 2026 63(1). – Name: Avail Label: Availability Group: Avail Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 21 – Name: DatePubCY Label: Publication Date Group: Date Data: 2026 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Automation%22">Automation</searchLink><br /><searchLink fieldCode="DE" term="%22Scores%22">Scores</searchLink><br /><searchLink fieldCode="DE" term="%22Bias%22">Bias</searchLink><br /><searchLink fieldCode="DE" term="%22Accuracy%22">Accuracy</searchLink><br /><searchLink fieldCode="DE" term="%22Prediction%22">Prediction</searchLink><br /><searchLink fieldCode="DE" term="%22Classification%22">Classification</searchLink><br /><searchLink fieldCode="DE" term="%22Algorithms%22">Algorithms</searchLink><br /><searchLink fieldCode="DE" term="%22Data+Analysis%22">Data Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Measurement%22">Measurement</searchLink><br /><searchLink fieldCode="DE" term="%22Methods%22">Methods</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluation+Methods%22">Evaluation Methods</searchLink><br /><searchLink fieldCode="DE" term="%22Sampling%22">Sampling</searchLink><br /><searchLink fieldCode="DE" term="%22Technology+Uses+in+Education%22">Technology Uses in Education</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1111/jedm.70031 – Name: ISSN Label: ISSN Group: ISSN Data: 0022-0655<br />1745-3984 – Name: Abstract Label: Abstract Group: Ab Data: Unbiasedness for proficiency estimates is important for autoscoring engines since the outcome might be used for future learning or placement. Imbalanced training data may lead to certain biases and lower the prediction accuracy for classification algorithms. In this article, we investigated several data augmentation methods to lower the negative effect of imbalanced data in measurement settings. Four approaches were examined: (1) Resampling methods, either oversampling or undersampling; (2) Active resampling methods, where the resampling weight is based on representativeness in the training set; (3) Data expansion methods using synonym Replacement, slightly changing the meaning or semantics of the original answers; and (4) Content recreation method using Generative AI (e.g., ChatGPT) to create responses for less populated scores. We compared the performance (e.g., Accuracy, QWK, F1) as well as the distance metric for different combinations of the methods. Two datasets with different imbalanced distributions were used. Results show that all four methods can help to mitigate the bias issue and the efficacy was influenced by the imbalance level, representativeness of the original data and the level of increment in the variety of the response (i.e., lexical diversity). In general, resampling and GenAI with active resampling showed the best overall performance. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: Note Label: Notes Group: Note Data: https://osf.io/cr2s6/?view_only=a356d94ef26342aaa5b19674c558169c – Name: DateEntry Label: Entry Date Group: Date Data: 2026 – Name: AN Label: Accession Number Group: ID Data: EJ1501394 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1501394 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1111/jedm.70031 Languages: – Text: English PhysicalDescription: Pagination: PageCount: 21 Subjects: – SubjectFull: Automation Type: general – SubjectFull: Scores Type: general – SubjectFull: Bias Type: general – SubjectFull: Accuracy Type: general – SubjectFull: Prediction Type: general – SubjectFull: Classification Type: general – SubjectFull: Algorithms Type: general – SubjectFull: Data Analysis Type: general – SubjectFull: Measurement Type: general – SubjectFull: Methods Type: general – SubjectFull: Evaluation Methods Type: general – SubjectFull: Sampling Type: general – SubjectFull: Technology Uses in Education Type: general – SubjectFull: Artificial Intelligence Type: general Titles: – TitleFull: AI and Measurement Concerns: Dealing with Imbalanced Data in Autoscoring Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Yunting Liu – PersonEntity: Name: NameFull: Yijun Xiang – PersonEntity: Name: NameFull: Xutao Feng – PersonEntity: Name: NameFull: Mark Wilson IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 03 Type: published Y: 2026 Identifiers: – Type: issn-print Value: 0022-0655 – Type: issn-electronic Value: 1745-3984 Numbering: – Type: volume Value: 63 – Type: issue Value: 1 Titles: – TitleFull: Journal of Educational Measurement Type: main |
| ResultId | 1 |