RACES: reward-aligned consistent essay scoring with large language models.

Saved in:
Bibliographic Details
Title: RACES: reward-aligned consistent essay scoring with large language models.
Authors: Zhang, Zhenxin1 (AUTHOR) zzx@stu.xidian.edu.cn, Ding, Ziyu1 (AUTHOR) dingziyu@stu.xidian.edu.cn, Liu, Mengyun2 (AUTHOR) lmy4pub@gmail.com, Sang, Haiwei3 (AUTHOR) haiweisang@gznc.edu.cn
Source: International Journal of Educational Technology in Higher Education. 6/19/2026, Vol. 23 Issue 1, p1-21. 21p.
Subject Terms: *Educational evaluation, Language models, Reinforcement learning, Text mining, Statistical reliability, Reward (Psychology), Mathematical optimization
Abstract: With the rapid advancement of large language models, the demand for intelligent and fine-grained automated essay scoring in educational assessment has increased significantly. However, existing methods still face challenges in maintaining scoring alignment and output consistency, making it difficult to consistently approximate human scoring standards. To address these issues, this paper proposes a unified framework named RACES (Reward-Aligned Consistent Essay Scoring), which integrates LoRA-based parameter-efficient fine-tuning, reward modeling, and proximal policy optimization reinforcement learning. The framework establishes an offline inference–feedback–optimization pipeline, enabling optimization toward proxy preference signals simulated via LLM-generated feedback while constraining policy drift through KL regularization. Experimental results on the ASAP 2.0 dataset show that RACES improves QWK and auxiliary SimCSE metrics compared with the evaluated pretrained and fine-tuned model configurations, achieving rapid convergence with limited training iterations. The framework improves scoring accuracy under the evaluated settings, while consistency is examined through KL-regularized optimization behavior and auxiliary proxy-feedback analysis rather than direct deployment-level robustness tests. These findings suggest the practical potential of RACES for supporting more controlled preliminary essay scoring in educational assessment, particularly as an auxiliary tool for reducing grading workload and improving the reliability of large-scale writing evaluation. [ABSTRACT FROM AUTHOR]
Copyright of International Journal of Educational Technology in Higher Education is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Education Research Complete
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: ehh
DbLabel: Education Research Complete
An: 194698353
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: RACES: reward-aligned consistent essay scoring with large language models.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Zhang%2C+Zhenxin%22">Zhang, Zhenxin</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> zzx@stu.xidian.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Ding%2C+Ziyu%22">Ding, Ziyu</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> dingziyu@stu.xidian.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Liu%2C+Mengyun%22">Liu, Mengyun</searchLink><relatesTo>2</relatesTo> (AUTHOR)<i> lmy4pub@gmail.com</i><br /><searchLink fieldCode="AR" term="%22Sang%2C+Haiwei%22">Sang, Haiwei</searchLink><relatesTo>3</relatesTo> (AUTHOR)<i> haiweisang@gznc.edu.cn</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22International+Journal+of+Educational+Technology+in+Higher+Education%22">International Journal of Educational Technology in Higher Education</searchLink>. 6/19/2026, Vol. 23 Issue 1, p1-21. 21p.
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: *<searchLink fieldCode="DE" term="%22Educational+evaluation%22">Educational evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Language+models%22">Language models</searchLink><br /><searchLink fieldCode="DE" term="%22Reinforcement+learning%22">Reinforcement learning</searchLink><br /><searchLink fieldCode="DE" term="%22Text+mining%22">Text mining</searchLink><br /><searchLink fieldCode="DE" term="%22Statistical+reliability%22">Statistical reliability</searchLink><br /><searchLink fieldCode="DE" term="%22Reward+%28Psychology%29%22">Reward (Psychology)</searchLink><br /><searchLink fieldCode="DE" term="%22Mathematical+optimization%22">Mathematical optimization</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: With the rapid advancement of large language models, the demand for intelligent and fine-grained automated essay scoring in educational assessment has increased significantly. However, existing methods still face challenges in maintaining scoring alignment and output consistency, making it difficult to consistently approximate human scoring standards. To address these issues, this paper proposes a unified framework named RACES (Reward-Aligned Consistent Essay Scoring), which integrates LoRA-based parameter-efficient fine-tuning, reward modeling, and proximal policy optimization reinforcement learning. The framework establishes an offline inference–feedback–optimization pipeline, enabling optimization toward proxy preference signals simulated via LLM-generated feedback while constraining policy drift through KL regularization. Experimental results on the ASAP 2.0 dataset show that RACES improves QWK and auxiliary SimCSE metrics compared with the evaluated pretrained and fine-tuned model configurations, achieving rapid convergence with limited training iterations. The framework improves scoring accuracy under the evaluated settings, while consistency is examined through KL-regularized optimization behavior and auxiliary proxy-feedback analysis rather than direct deployment-level robustness tests. These findings suggest the practical potential of RACES for supporting more controlled preliminary essay scoring in educational assessment, particularly as an auxiliary tool for reducing grading workload and improving the reliability of large-scale writing evaluation. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of International Journal of Educational Technology in Higher Education is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=ehh&AN=194698353
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1186/s41239-026-00607-8
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 21
        StartPage: 1
    Subjects:
      – SubjectFull: Educational evaluation
        Type: general
      – SubjectFull: Language models
        Type: general
      – SubjectFull: Reinforcement learning
        Type: general
      – SubjectFull: Text mining
        Type: general
      – SubjectFull: Statistical reliability
        Type: general
      – SubjectFull: Reward (Psychology)
        Type: general
      – SubjectFull: Mathematical optimization
        Type: general
    Titles:
      – TitleFull: RACES: reward-aligned consistent essay scoring with large language models.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Zhang, Zhenxin
      – PersonEntity:
          Name:
            NameFull: Ding, Ziyu
      – PersonEntity:
          Name:
            NameFull: Liu, Mengyun
      – PersonEntity:
          Name:
            NameFull: Sang, Haiwei
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 19
              M: 06
              Text: 6/19/2026
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 23659440
          Numbering:
            – Type: volume
              Value: 23
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: International Journal of Educational Technology in Higher Education
              Type: main
ResultId 1