A Comparative Study of the Human, Automated Scoring Model, and GPT-4 Ratings of Young EFL Students' Writing

Saved in:
Bibliographic Details
Title: A Comparative Study of the Human, Automated Scoring Model, and GPT-4 Ratings of Young EFL Students' Writing
Language: English
Authors: Michael Suhan (ORCID 0000-0002-4734-3794), Mikyung Kim Wolf (ORCID 0000-0002-0622-7754)
Source: Language Testing. 2026 43(1):66-78.
Availability: SAGE Publications. 2455 Teller Road, Thousand Oaks, CA 91320. Tel: 800-818-7243; Tel: 805-499-9774; Fax: 800-583-2665; e-mail: journals@sagepub.com; Web site: https://sagepub.com
Peer Reviewed: Y
Page Count: 13
Publication Date: 2026
Document Type: Journal Articles
Reports - Research
Education Level: Elementary Education
Secondary Education
Descriptors: Language Tests, Automation, Computer Assisted Testing, Scoring, Artificial Intelligence, Writing Tests, Writing Evaluation, Second Language Learning, English (Second Language), Evaluation Methods, Comparative Analysis, Natural Language Processing, Elementary School Students, Secondary School Students, Foreign Countries
Geographic Terms: China, Hong Kong, Turkey, Indonesia
Assessment and Survey Identifiers: Test of English as a Foreign Language
DOI: 10.1177/02655322251346860
ISSN: 0265-5322
1477-0946
Abstract: Large language models, such as OpenAI's GPT-4, have the potential to revolutionize automated writing evaluation (AWE). The present study examines the performance of the GPT-4 model in evaluating the writing of young English as a foreign language learners. Responses to three constructed response tasks (n = 1908) on Educational Testing Service's (ETS's) TOEFL Junior® Writing test were scored using two different GPT-4 prompting methods. Scores predicted by the two GPT-4 prompting methods and the TOEFL Junior Writing test's operational AWE model were compared against human ratings to evaluate the performance of each method in terms of several measures. The results indicated that there are inconsistencies in the performance of the GPT-4 models depending on the measures, task types, and test forms, compared to the test's AWE model and human ratings. Implications of the findings for practice and further research are discussed.
Abstractor: As Provided
Entry Date: 2026
Accession Number: EJ1493276
Database: ERIC
FullText Text:
  Availability: 0
Header DbId: eric
DbLabel: ERIC
An: EJ1493276
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: A Comparative Study of the Human, Automated Scoring Model, and GPT-4 Ratings of Young EFL Students' Writing
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Michael+Suhan%22">Michael Suhan</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-4734-3794">0000-0002-4734-3794</externalLink>)<br /><searchLink fieldCode="AR" term="%22Mikyung+Kim+Wolf%22">Mikyung Kim Wolf</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-0622-7754">0000-0002-0622-7754</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Language+Testing%22"><i>Language Testing</i></searchLink>. 2026 43(1):66-78.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: SAGE Publications. 2455 Teller Road, Thousand Oaks, CA 91320. Tel: 800-818-7243; Tel: 805-499-9774; Fax: 800-583-2665; e-mail: journals@sagepub.com; Web site: https://sagepub.com
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 13
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2026
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Audience
  Label: Education Level
  Group: Audnce
  Data: <searchLink fieldCode="EL" term="%22Elementary+Education%22">Elementary Education</searchLink><br /><searchLink fieldCode="EL" term="%22Secondary+Education%22">Secondary Education</searchLink>
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Language+Tests%22">Language Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Automation%22">Automation</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+Assisted+Testing%22">Computer Assisted Testing</searchLink><br /><searchLink fieldCode="DE" term="%22Scoring%22">Scoring</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Writing+Tests%22">Writing Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Writing+Evaluation%22">Writing Evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Second+Language+Learning%22">Second Language Learning</searchLink><br /><searchLink fieldCode="DE" term="%22English+%28Second+Language%29%22">English (Second Language)</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluation+Methods%22">Evaluation Methods</searchLink><br /><searchLink fieldCode="DE" term="%22Comparative+Analysis%22">Comparative Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+Language+Processing%22">Natural Language Processing</searchLink><br /><searchLink fieldCode="DE" term="%22Elementary+School+Students%22">Elementary School Students</searchLink><br /><searchLink fieldCode="DE" term="%22Secondary+School+Students%22">Secondary School Students</searchLink><br /><searchLink fieldCode="DE" term="%22Foreign+Countries%22">Foreign Countries</searchLink>
– Name: Subject
  Label: Geographic Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22China%22">China</searchLink><br /><searchLink fieldCode="DE" term="%22Hong+Kong%22">Hong Kong</searchLink><br /><searchLink fieldCode="DE" term="%22Turkey%22">Turkey</searchLink><br /><searchLink fieldCode="DE" term="%22Indonesia%22">Indonesia</searchLink>
– Name: SubjectThesaurus
  Label: Assessment and Survey Identifiers
  Group: Su
  Data: <searchLink fieldCode="SU" term="%22Test+of+English+as+a+Foreign+Language%22">Test of English as a Foreign Language</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1177/02655322251346860
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0265-5322<br />1477-0946
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Large language models, such as OpenAI's GPT-4, have the potential to revolutionize automated writing evaluation (AWE). The present study examines the performance of the GPT-4 model in evaluating the writing of young English as a foreign language learners. Responses to three constructed response tasks (n = 1908) on Educational Testing Service's (ETS's) TOEFL Junior® Writing test were scored using two different GPT-4 prompting methods. Scores predicted by the two GPT-4 prompting methods and the TOEFL Junior Writing test's operational AWE model were compared against human ratings to evaluate the performance of each method in terms of several measures. The results indicated that there are inconsistencies in the performance of the GPT-4 models depending on the measures, task types, and test forms, compared to the test's AWE model and human ratings. Implications of the findings for practice and further research are discussed.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2026
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1493276
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1493276
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1177/02655322251346860
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 13
        StartPage: 66
    Subjects:
      – SubjectFull: Language Tests
        Type: general
      – SubjectFull: Automation
        Type: general
      – SubjectFull: Computer Assisted Testing
        Type: general
      – SubjectFull: Scoring
        Type: general
      – SubjectFull: Artificial Intelligence
        Type: general
      – SubjectFull: Writing Tests
        Type: general
      – SubjectFull: Writing Evaluation
        Type: general
      – SubjectFull: Second Language Learning
        Type: general
      – SubjectFull: English (Second Language)
        Type: general
      – SubjectFull: Evaluation Methods
        Type: general
      – SubjectFull: Comparative Analysis
        Type: general
      – SubjectFull: Natural Language Processing
        Type: general
      – SubjectFull: Elementary School Students
        Type: general
      – SubjectFull: Secondary School Students
        Type: general
      – SubjectFull: Foreign Countries
        Type: general
      – SubjectFull: China
        Type: general
      – SubjectFull: Hong Kong
        Type: general
      – SubjectFull: Turkey
        Type: general
      – SubjectFull: Indonesia
        Type: general
      – SubjectFull: Test of English as a Foreign Language
        Type: general
    Titles:
      – TitleFull: A Comparative Study of the Human, Automated Scoring Model, and GPT-4 Ratings of Young EFL Students' Writing
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Michael Suhan
      – PersonEntity:
          Name:
            NameFull: Mikyung Kim Wolf
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 0265-5322
            – Type: issn-electronic
              Value: 1477-0946
          Numbering:
            – Type: volume
              Value: 43
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Language Testing
              Type: main
ResultId 1