Finding Words Associated with DIF: Predicting Differential Item Functioning Using LLMs and Explainable AI

Saved in:
Bibliographic Details
Title: Finding Words Associated with DIF: Predicting Differential Item Functioning Using LLMs and Explainable AI
Language: English
Authors: Hotaka Maeda (ORCID 0009-0000-9498-786X), Yikai Lu (ORCID 0000-0003-4410-2589)
Source: Journal of Educational Measurement. 2025 62(4):883-906.
Availability: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 24
Publication Date: 2025
Document Type: Journal Articles
Reports - Research
Education Level: Elementary Secondary Education
Descriptors: Artificial Intelligence, Prediction, Test Bias, Test Items, Vocabulary, Language Arts, Mathematics, Summative Evaluation, Elementary Secondary Education, Test Content
DOI: 10.1111/jedm.70017
ISSN: 0022-0655
1745-3984
Abstract: We fine-tuned and compared several encoder-based Transformer large language models (LLM) to predict differential item functioning (DIF) from the item text. We then applied explainable artificial intelligence (XAI) methods to identify specific words associated with the DIF prediction. The data included 42,180 items designed for English language arts and mathematics summative state assessments among students in grades 3 to 11. Prediction R[superscript 2] ranged from 0.04 to 0.32 among eight focal and reference group pairs. Our findings suggest that many words associated with DIF reflect minor subdomains included in the test blueprint by design, rather than construct-irrelevant content that may need to be removed from assessments. This may explain why qualitative reviews of DIF items often yield inconclusive results. Our approach can be used to (1) screen words associated with DIF during the item-writing process for immediate revision to reduce preventable adverse DIF, (2) assist traditional DIF item reviews by highlighting key words, or (3) use DIF prediction as an alternative when obtaining sufficient sample size for traditional DIF analyses is impossible. Extensions of this research can enhance the assessment fairness, especially programs that lack resources to build high-quality items, and among smaller subpopulations with insufficient sample sizes for traditional DIF analyses.
Abstractor: As Provided
Entry Date: 2026
Accession Number: EJ1491342
Database: ERIC
Description
Abstract:We fine-tuned and compared several encoder-based Transformer large language models (LLM) to predict differential item functioning (DIF) from the item text. We then applied explainable artificial intelligence (XAI) methods to identify specific words associated with the DIF prediction. The data included 42,180 items designed for English language arts and mathematics summative state assessments among students in grades 3 to 11. Prediction R[superscript 2] ranged from 0.04 to 0.32 among eight focal and reference group pairs. Our findings suggest that many words associated with DIF reflect minor subdomains included in the test blueprint by design, rather than construct-irrelevant content that may need to be removed from assessments. This may explain why qualitative reviews of DIF items often yield inconclusive results. Our approach can be used to (1) screen words associated with DIF during the item-writing process for immediate revision to reduce preventable adverse DIF, (2) assist traditional DIF item reviews by highlighting key words, or (3) use DIF prediction as an alternative when obtaining sufficient sample size for traditional DIF analyses is impossible. Extensions of this research can enhance the assessment fairness, especially programs that lack resources to build high-quality items, and among smaller subpopulations with insufficient sample sizes for traditional DIF analyses.
ISSN:0022-0655
1745-3984
DOI:10.1111/jedm.70017