A Comparative Study of the Human, Automated Scoring Model, and GPT-4 Ratings of Young EFL Students' Writing
Saved in:
| Title: | A Comparative Study of the Human, Automated Scoring Model, and GPT-4 Ratings of Young EFL Students' Writing |
|---|---|
| Language: | English |
| Authors: | Michael Suhan (ORCID |
| Source: | Language Testing. 2026 43(1):66-78. |
| Availability: | SAGE Publications. 2455 Teller Road, Thousand Oaks, CA 91320. Tel: 800-818-7243; Tel: 805-499-9774; Fax: 800-583-2665; e-mail: journals@sagepub.com; Web site: https://sagepub.com |
| Peer Reviewed: | Y |
| Page Count: | 13 |
| Publication Date: | 2026 |
| Document Type: | Journal Articles Reports - Research |
| Education Level: | Elementary Education Secondary Education |
| Descriptors: | Language Tests, Automation, Computer Assisted Testing, Scoring, Artificial Intelligence, Writing Tests, Writing Evaluation, Second Language Learning, English (Second Language), Evaluation Methods, Comparative Analysis, Natural Language Processing, Elementary School Students, Secondary School Students, Foreign Countries |
| Geographic Terms: | China, Hong Kong, Turkey, Indonesia |
| Assessment and Survey Identifiers: | Test of English as a Foreign Language |
| DOI: | 10.1177/02655322251346860 |
| ISSN: | 0265-5322 1477-0946 |
| Abstract: | Large language models, such as OpenAI's GPT-4, have the potential to revolutionize automated writing evaluation (AWE). The present study examines the performance of the GPT-4 model in evaluating the writing of young English as a foreign language learners. Responses to three constructed response tasks (n = 1908) on Educational Testing Service's (ETS's) TOEFL Junior® Writing test were scored using two different GPT-4 prompting methods. Scores predicted by the two GPT-4 prompting methods and the TOEFL Junior Writing test's operational AWE model were compared against human ratings to evaluate the performance of each method in terms of several measures. The results indicated that there are inconsistencies in the performance of the GPT-4 models depending on the measures, task types, and test forms, compared to the test's AWE model and human ratings. Implications of the findings for practice and further research are discussed. |
| Abstractor: | As Provided |
| Entry Date: | 2026 |
| Accession Number: | EJ1493276 |
| Database: | ERIC |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1493276 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: A Comparative Study of the Human, Automated Scoring Model, and GPT-4 Ratings of Young EFL Students' Writing – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Michael+Suhan%22">Michael Suhan</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-4734-3794">0000-0002-4734-3794</externalLink>)<br /><searchLink fieldCode="AR" term="%22Mikyung+Kim+Wolf%22">Mikyung Kim Wolf</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-0622-7754">0000-0002-0622-7754</externalLink>) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Language+Testing%22"><i>Language Testing</i></searchLink>. 2026 43(1):66-78. – Name: Avail Label: Availability Group: Avail Data: SAGE Publications. 2455 Teller Road, Thousand Oaks, CA 91320. Tel: 800-818-7243; Tel: 805-499-9774; Fax: 800-583-2665; e-mail: journals@sagepub.com; Web site: https://sagepub.com – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 13 – Name: DatePubCY Label: Publication Date Group: Date Data: 2026 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Audience Label: Education Level Group: Audnce Data: <searchLink fieldCode="EL" term="%22Elementary+Education%22">Elementary Education</searchLink><br /><searchLink fieldCode="EL" term="%22Secondary+Education%22">Secondary Education</searchLink> – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Language+Tests%22">Language Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Automation%22">Automation</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+Assisted+Testing%22">Computer Assisted Testing</searchLink><br /><searchLink fieldCode="DE" term="%22Scoring%22">Scoring</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Writing+Tests%22">Writing Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Writing+Evaluation%22">Writing Evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Second+Language+Learning%22">Second Language Learning</searchLink><br /><searchLink fieldCode="DE" term="%22English+%28Second+Language%29%22">English (Second Language)</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluation+Methods%22">Evaluation Methods</searchLink><br /><searchLink fieldCode="DE" term="%22Comparative+Analysis%22">Comparative Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+Language+Processing%22">Natural Language Processing</searchLink><br /><searchLink fieldCode="DE" term="%22Elementary+School+Students%22">Elementary School Students</searchLink><br /><searchLink fieldCode="DE" term="%22Secondary+School+Students%22">Secondary School Students</searchLink><br /><searchLink fieldCode="DE" term="%22Foreign+Countries%22">Foreign Countries</searchLink> – Name: Subject Label: Geographic Terms Group: Su Data: <searchLink fieldCode="DE" term="%22China%22">China</searchLink><br /><searchLink fieldCode="DE" term="%22Hong+Kong%22">Hong Kong</searchLink><br /><searchLink fieldCode="DE" term="%22Turkey%22">Turkey</searchLink><br /><searchLink fieldCode="DE" term="%22Indonesia%22">Indonesia</searchLink> – Name: SubjectThesaurus Label: Assessment and Survey Identifiers Group: Su Data: <searchLink fieldCode="SU" term="%22Test+of+English+as+a+Foreign+Language%22">Test of English as a Foreign Language</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1177/02655322251346860 – Name: ISSN Label: ISSN Group: ISSN Data: 0265-5322<br />1477-0946 – Name: Abstract Label: Abstract Group: Ab Data: Large language models, such as OpenAI's GPT-4, have the potential to revolutionize automated writing evaluation (AWE). The present study examines the performance of the GPT-4 model in evaluating the writing of young English as a foreign language learners. Responses to three constructed response tasks (n = 1908) on Educational Testing Service's (ETS's) TOEFL Junior® Writing test were scored using two different GPT-4 prompting methods. Scores predicted by the two GPT-4 prompting methods and the TOEFL Junior Writing test's operational AWE model were compared against human ratings to evaluate the performance of each method in terms of several measures. The results indicated that there are inconsistencies in the performance of the GPT-4 models depending on the measures, task types, and test forms, compared to the test's AWE model and human ratings. Implications of the findings for practice and further research are discussed. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2026 – Name: AN Label: Accession Number Group: ID Data: EJ1493276 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1493276 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1177/02655322251346860 Languages: – Text: English PhysicalDescription: Pagination: PageCount: 13 StartPage: 66 Subjects: – SubjectFull: Language Tests Type: general – SubjectFull: Automation Type: general – SubjectFull: Computer Assisted Testing Type: general – SubjectFull: Scoring Type: general – SubjectFull: Artificial Intelligence Type: general – SubjectFull: Writing Tests Type: general – SubjectFull: Writing Evaluation Type: general – SubjectFull: Second Language Learning Type: general – SubjectFull: English (Second Language) Type: general – SubjectFull: Evaluation Methods Type: general – SubjectFull: Comparative Analysis Type: general – SubjectFull: Natural Language Processing Type: general – SubjectFull: Elementary School Students Type: general – SubjectFull: Secondary School Students Type: general – SubjectFull: Foreign Countries Type: general – SubjectFull: China Type: general – SubjectFull: Hong Kong Type: general – SubjectFull: Turkey Type: general – SubjectFull: Indonesia Type: general – SubjectFull: Test of English as a Foreign Language Type: general Titles: – TitleFull: A Comparative Study of the Human, Automated Scoring Model, and GPT-4 Ratings of Young EFL Students' Writing Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Michael Suhan – PersonEntity: Name: NameFull: Mikyung Kim Wolf IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 01 Type: published Y: 2026 Identifiers: – Type: issn-print Value: 0265-5322 – Type: issn-electronic Value: 1477-0946 Numbering: – Type: volume Value: 43 – Type: issue Value: 1 Titles: – TitleFull: Language Testing Type: main |
| ResultId | 1 |