AI and Measurement Concerns: Dealing with Imbalanced Data in Autoscoring

Saved in:
Bibliographic Details
Title: AI and Measurement Concerns: Dealing with Imbalanced Data in Autoscoring
Language: English
Authors: Yunting Liu (ORCID 0009-0004-9594-9661), Yijun Xiang, Xutao Feng, Mark Wilson (ORCID 0000-0002-0425-5305)
Source: Journal of Educational Measurement. 2026 63(1).
Availability: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 21
Publication Date: 2026
Document Type: Journal Articles
Reports - Research
Descriptors: Automation, Scores, Bias, Accuracy, Prediction, Classification, Algorithms, Data Analysis, Measurement, Methods, Evaluation Methods, Sampling, Technology Uses in Education, Artificial Intelligence
DOI: 10.1111/jedm.70031
ISSN: 0022-0655
1745-3984
Abstract: Unbiasedness for proficiency estimates is important for autoscoring engines since the outcome might be used for future learning or placement. Imbalanced training data may lead to certain biases and lower the prediction accuracy for classification algorithms. In this article, we investigated several data augmentation methods to lower the negative effect of imbalanced data in measurement settings. Four approaches were examined: (1) Resampling methods, either oversampling or undersampling; (2) Active resampling methods, where the resampling weight is based on representativeness in the training set; (3) Data expansion methods using synonym Replacement, slightly changing the meaning or semantics of the original answers; and (4) Content recreation method using Generative AI (e.g., ChatGPT) to create responses for less populated scores. We compared the performance (e.g., Accuracy, QWK, F1) as well as the distance metric for different combinations of the methods. Two datasets with different imbalanced distributions were used. Results show that all four methods can help to mitigate the bias issue and the efficacy was influenced by the imbalance level, representativeness of the original data and the level of increment in the variety of the response (i.e., lexical diversity). In general, resampling and GenAI with active resampling showed the best overall performance.
Abstractor: As Provided
Notes: https://osf.io/cr2s6/?view_only=a356d94ef26342aaa5b19674c558169c
Entry Date: 2026
Accession Number: EJ1501394
Database: ERIC
FullText Text:
  Availability: 0
Header DbId: eric
DbLabel: ERIC
An: EJ1501394
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: AI and Measurement Concerns: Dealing with Imbalanced Data in Autoscoring
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Yunting+Liu%22">Yunting Liu</searchLink> (ORCID <externalLink term="https://orcid.org/0009-0004-9594-9661">0009-0004-9594-9661</externalLink>)<br /><searchLink fieldCode="AR" term="%22Yijun+Xiang%22">Yijun Xiang</searchLink><br /><searchLink fieldCode="AR" term="%22Xutao+Feng%22">Xutao Feng</searchLink><br /><searchLink fieldCode="AR" term="%22Mark+Wilson%22">Mark Wilson</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-0425-5305">0000-0002-0425-5305</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Journal+of+Educational+Measurement%22"><i>Journal of Educational Measurement</i></searchLink>. 2026 63(1).
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 21
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2026
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Automation%22">Automation</searchLink><br /><searchLink fieldCode="DE" term="%22Scores%22">Scores</searchLink><br /><searchLink fieldCode="DE" term="%22Bias%22">Bias</searchLink><br /><searchLink fieldCode="DE" term="%22Accuracy%22">Accuracy</searchLink><br /><searchLink fieldCode="DE" term="%22Prediction%22">Prediction</searchLink><br /><searchLink fieldCode="DE" term="%22Classification%22">Classification</searchLink><br /><searchLink fieldCode="DE" term="%22Algorithms%22">Algorithms</searchLink><br /><searchLink fieldCode="DE" term="%22Data+Analysis%22">Data Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Measurement%22">Measurement</searchLink><br /><searchLink fieldCode="DE" term="%22Methods%22">Methods</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluation+Methods%22">Evaluation Methods</searchLink><br /><searchLink fieldCode="DE" term="%22Sampling%22">Sampling</searchLink><br /><searchLink fieldCode="DE" term="%22Technology+Uses+in+Education%22">Technology Uses in Education</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1111/jedm.70031
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0022-0655<br />1745-3984
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Unbiasedness for proficiency estimates is important for autoscoring engines since the outcome might be used for future learning or placement. Imbalanced training data may lead to certain biases and lower the prediction accuracy for classification algorithms. In this article, we investigated several data augmentation methods to lower the negative effect of imbalanced data in measurement settings. Four approaches were examined: (1) Resampling methods, either oversampling or undersampling; (2) Active resampling methods, where the resampling weight is based on representativeness in the training set; (3) Data expansion methods using synonym Replacement, slightly changing the meaning or semantics of the original answers; and (4) Content recreation method using Generative AI (e.g., ChatGPT) to create responses for less populated scores. We compared the performance (e.g., Accuracy, QWK, F1) as well as the distance metric for different combinations of the methods. Two datasets with different imbalanced distributions were used. Results show that all four methods can help to mitigate the bias issue and the efficacy was influenced by the imbalance level, representativeness of the original data and the level of increment in the variety of the response (i.e., lexical diversity). In general, resampling and GenAI with active resampling showed the best overall performance.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: Note
  Label: Notes
  Group: Note
  Data: https://osf.io/cr2s6/?view_only=a356d94ef26342aaa5b19674c558169c
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2026
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1501394
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1501394
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1111/jedm.70031
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 21
    Subjects:
      – SubjectFull: Automation
        Type: general
      – SubjectFull: Scores
        Type: general
      – SubjectFull: Bias
        Type: general
      – SubjectFull: Accuracy
        Type: general
      – SubjectFull: Prediction
        Type: general
      – SubjectFull: Classification
        Type: general
      – SubjectFull: Algorithms
        Type: general
      – SubjectFull: Data Analysis
        Type: general
      – SubjectFull: Measurement
        Type: general
      – SubjectFull: Methods
        Type: general
      – SubjectFull: Evaluation Methods
        Type: general
      – SubjectFull: Sampling
        Type: general
      – SubjectFull: Technology Uses in Education
        Type: general
      – SubjectFull: Artificial Intelligence
        Type: general
    Titles:
      – TitleFull: AI and Measurement Concerns: Dealing with Imbalanced Data in Autoscoring
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Yunting Liu
      – PersonEntity:
          Name:
            NameFull: Yijun Xiang
      – PersonEntity:
          Name:
            NameFull: Xutao Feng
      – PersonEntity:
          Name:
            NameFull: Mark Wilson
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 03
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 0022-0655
            – Type: issn-electronic
              Value: 1745-3984
          Numbering:
            – Type: volume
              Value: 63
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Journal of Educational Measurement
              Type: main
ResultId 1