RACES: reward-aligned consistent essay scoring with large language models.
Saved in:
| Title: | RACES: reward-aligned consistent essay scoring with large language models. |
|---|---|
| Authors: | Zhang, Zhenxin1 (AUTHOR) zzx@stu.xidian.edu.cn, Ding, Ziyu1 (AUTHOR) dingziyu@stu.xidian.edu.cn, Liu, Mengyun2 (AUTHOR) lmy4pub@gmail.com, Sang, Haiwei3 (AUTHOR) haiweisang@gznc.edu.cn |
| Source: | International Journal of Educational Technology in Higher Education. 6/19/2026, Vol. 23 Issue 1, p1-21. 21p. |
| Subject Terms: | *Educational evaluation, Language models, Reinforcement learning, Text mining, Statistical reliability, Reward (Psychology), Mathematical optimization |
| Abstract: | With the rapid advancement of large language models, the demand for intelligent and fine-grained automated essay scoring in educational assessment has increased significantly. However, existing methods still face challenges in maintaining scoring alignment and output consistency, making it difficult to consistently approximate human scoring standards. To address these issues, this paper proposes a unified framework named RACES (Reward-Aligned Consistent Essay Scoring), which integrates LoRA-based parameter-efficient fine-tuning, reward modeling, and proximal policy optimization reinforcement learning. The framework establishes an offline inference–feedback–optimization pipeline, enabling optimization toward proxy preference signals simulated via LLM-generated feedback while constraining policy drift through KL regularization. Experimental results on the ASAP 2.0 dataset show that RACES improves QWK and auxiliary SimCSE metrics compared with the evaluated pretrained and fine-tuned model configurations, achieving rapid convergence with limited training iterations. The framework improves scoring accuracy under the evaluated settings, while consistency is examined through KL-regularized optimization behavior and auxiliary proxy-feedback analysis rather than direct deployment-level robustness tests. These findings suggest the practical potential of RACES for supporting more controlled preliminary essay scoring in educational assessment, particularly as an auxiliary tool for reducing grading workload and improving the reliability of large-scale writing evaluation. [ABSTRACT FROM AUTHOR] |
| Copyright of International Journal of Educational Technology in Higher Education is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Education Research Complete |
|
Full text is not displayed to guests.
Login for full access.
|
|
| Abstract: | With the rapid advancement of large language models, the demand for intelligent and fine-grained automated essay scoring in educational assessment has increased significantly. However, existing methods still face challenges in maintaining scoring alignment and output consistency, making it difficult to consistently approximate human scoring standards. To address these issues, this paper proposes a unified framework named RACES (Reward-Aligned Consistent Essay Scoring), which integrates LoRA-based parameter-efficient fine-tuning, reward modeling, and proximal policy optimization reinforcement learning. The framework establishes an offline inference–feedback–optimization pipeline, enabling optimization toward proxy preference signals simulated via LLM-generated feedback while constraining policy drift through KL regularization. Experimental results on the ASAP 2.0 dataset show that RACES improves QWK and auxiliary SimCSE metrics compared with the evaluated pretrained and fine-tuned model configurations, achieving rapid convergence with limited training iterations. The framework improves scoring accuracy under the evaluated settings, while consistency is examined through KL-regularized optimization behavior and auxiliary proxy-feedback analysis rather than direct deployment-level robustness tests. These findings suggest the practical potential of RACES for supporting more controlled preliminary essay scoring in educational assessment, particularly as an auxiliary tool for reducing grading workload and improving the reliability of large-scale writing evaluation. [ABSTRACT FROM AUTHOR] |
|---|---|
| ISSN: | 23659440 |
| DOI: | 10.1186/s41239-026-00607-8 |