Effectiveness of large language models in automated evaluation of argumentative essays: finetuning vs. zero-shot prompting.

Saved in:
Bibliographic Details
Title: Effectiveness of large language models in automated evaluation of argumentative essays: finetuning vs. zero-shot prompting.
Authors: Wang, Qiao1 (AUTHOR) judy.wang@aoni.waseda.jp, Gayed, John Maurice2 (AUTHOR)
Source: Computer Assisted Language Learning. Feb/Mar2026, Vol. 39 Issue 1/2, p209-237. 29p.
Subject Terms: *Machine learning, Generative pre-trained transformers, Essays, Language models
Abstract: To address the long-standing challenge facing traditional automated writing evaluation (AWE) systems in assessing higher-order thinking, this study built an AWE system for scoring argumentative essays by finetuning the GPT-3.5 Large Language Model. The system's effectiveness was compared with that of the non-finetuned GPT-3.5 and GPT-4 base models via zero-shot prompting, which involves applying the model to perform tasks without any prior specific training or examples on those tasks. The dataset used was the TOEFL Public Writing Dataset provided by Education Testing Service (ETS), containing 480 argumentative essays with ground truth scores under two essay prompts. Three finetuned models were generated: two finetuned exclusively on either prompt and one on both. All finetuned and base models were used to score the remaining essays after finetuning and their scoring effectiveness was compared with ground truth scores, i.e., benchmark scores assigned by ETS-trained human raters. The impact of the variety of finetuning prompts and the robustness of finetuned models were also explored. Results showed a 100% consistency of all models in two scoring sessions. More importantly, the finetuned models significantly outperformed the base models in accuracy and reliability. The best-performing model, finetuned on prompt 1, showed an RMSE of 0.57, a percentage agreement (score discrepancy ≤ 0.5) of 84.72% and a QWK of 0.78. Further, the model finetuned on both prompts did not exhibit enhanced performance, and the two models finetuned on one prompt remained robust when scoring essays from the alternative prompt. These results suggest (1) task-specific finetuning for AWE is beneficial; (2) finetuning does not require a large variety of essay prompts; and (3) fine-tuned models are robust to unseen essays. [ABSTRACT FROM AUTHOR]
Copyright of Computer Assisted Language Learning is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Education Research Complete
FullText Text:
  Availability: 0
Header DbId: ehh
DbLabel: Education Research Complete
An: 191948521
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Effectiveness of large language models in automated evaluation of argumentative essays: finetuning vs. zero-shot prompting.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Wang%2C+Qiao%22">Wang, Qiao</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> judy.wang@aoni.waseda.jp</i><br /><searchLink fieldCode="AR" term="%22Gayed%2C+John+Maurice%22">Gayed, John Maurice</searchLink><relatesTo>2</relatesTo> (AUTHOR)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Computer+Assisted+Language+Learning%22">Computer Assisted Language Learning</searchLink>. Feb/Mar2026, Vol. 39 Issue 1/2, p209-237. 29p.
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: *<searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Generative+pre-trained+transformers%22">Generative pre-trained transformers</searchLink><br /><searchLink fieldCode="DE" term="%22Essays%22">Essays</searchLink><br /><searchLink fieldCode="DE" term="%22Language+models%22">Language models</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: To address the long-standing challenge facing traditional automated writing evaluation (AWE) systems in assessing higher-order thinking, this study built an AWE system for scoring argumentative essays by finetuning the GPT-3.5 Large Language Model. The system's effectiveness was compared with that of the non-finetuned GPT-3.5 and GPT-4 base models via zero-shot prompting, which involves applying the model to perform tasks without any prior specific training or examples on those tasks. The dataset used was the TOEFL Public Writing Dataset provided by Education Testing Service (ETS), containing 480 argumentative essays with ground truth scores under two essay prompts. Three finetuned models were generated: two finetuned exclusively on either prompt and one on both. All finetuned and base models were used to score the remaining essays after finetuning and their scoring effectiveness was compared with ground truth scores, i.e., benchmark scores assigned by ETS-trained human raters. The impact of the variety of finetuning prompts and the robustness of finetuned models were also explored. Results showed a 100% consistency of all models in two scoring sessions. More importantly, the finetuned models significantly outperformed the base models in accuracy and reliability. The best-performing model, finetuned on prompt 1, showed an RMSE of 0.57, a percentage agreement (score discrepancy ≤ 0.5) of 84.72% and a QWK of 0.78. Further, the model finetuned on both prompts did not exhibit enhanced performance, and the two models finetuned on one prompt remained robust when scoring essays from the alternative prompt. These results suggest (1) task-specific finetuning for AWE is beneficial; (2) finetuning does not require a large variety of essay prompts; and (3) fine-tuned models are robust to unseen essays. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Computer Assisted Language Learning is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=ehh&AN=191948521
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1080/09588221.2024.2371395
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 29
        StartPage: 209
    Subjects:
      – SubjectFull: Machine learning
        Type: general
      – SubjectFull: Generative pre-trained transformers
        Type: general
      – SubjectFull: Essays
        Type: general
      – SubjectFull: Language models
        Type: general
    Titles:
      – TitleFull: Effectiveness of large language models in automated evaluation of argumentative essays: finetuning vs. zero-shot prompting.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Wang, Qiao
      – PersonEntity:
          Name:
            NameFull: Gayed, John Maurice
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 02
              Text: Feb/Mar2026
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 09588221
          Numbering:
            – Type: volume
              Value: 39
            – Type: issue
              Value: 1/2
          Titles:
            – TitleFull: Computer Assisted Language Learning
              Type: main
ResultId 1