Exploring the Potential of ChatGPT for Evaluating English Essays in a Criterion-Based Assessment

Saved in:
Bibliographic Details
Title: Exploring the Potential of ChatGPT for Evaluating English Essays in a Criterion-Based Assessment
Language: English
Authors: Andrea Gjorevski, Mimi Li (ORCID 0000-0002-6705-5167), Troy L. Cox
Source: TESOL Quarterly: A Journal for Teachers of English to Speakers of Other Languages and of Standard English as a Second Dialect. 2025 59(1):S251-S279.
Availability: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 29
Publication Date: 2025
Document Type: Journal Articles
Reports - Research
Descriptors: Artificial Intelligence, Criterion Referenced Tests, Essay Tests, Automation, Writing Evaluation, Scoring, Interrater Reliability, Writing Tests, English (Second Language), International Organizations
DOI: 10.1002/tesq.70011
ISSN: 0039-8322
1545-7249
Abstract: Open access to novel AI tools offers unprecedented opportunities for human-AI collaboration in writing instruction and assessment. While research on using generative AI tools like ChatGPT in these contexts is emerging, more is needed to understand their effectiveness as Automated Writing Evaluation (AWE) tools. This study explores the potential of ChatGPT (GPT-3.5) to assist teachers and learners in the North Atlantic Treaty Organization (NATO) by evaluating English writing based on holistic scoring criteria. Using a mixed-methods approach, the study compared ChatGPT's ratings with human ratings on 100 writing tests to assess inter-rater reliability. It also analyzed the justifications provided by both human raters and ChatGPT to evaluate how well ChatGPT understood the rating criteria at different proficiency levels and whether its rationales could provide effective feedback for learners and support teacher feedback practices. Results showed strong agreement between ChatGPT's and human ratings, with ChatGPT demonstrating a similar understanding of the rating scales and offering justifications with elements of effective feedback. These findings indicate that ChatGPT holds promise as an AWE tool, providing meaningful feedback and valuable insights into holistic rating scales. This study encourages further exploration of AI in the L2 classroom and suggests leveraging AI to enhance writing pedagogy and classroom-based assessment.
Abstractor: As Provided
Entry Date: 2025
Accession Number: EJ1490674
Database: ERIC
FullText Text:
  Availability: 0
Header DbId: eric
DbLabel: ERIC
An: EJ1490674
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Exploring the Potential of ChatGPT for Evaluating English Essays in a Criterion-Based Assessment
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Andrea+Gjorevski%22">Andrea Gjorevski</searchLink><br /><searchLink fieldCode="AR" term="%22Mimi+Li%22">Mimi Li</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-6705-5167">0000-0002-6705-5167</externalLink>)<br /><searchLink fieldCode="AR" term="%22Troy+L%2E+Cox%22">Troy L. Cox</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22TESOL+Quarterly%3A+A+Journal+for+Teachers+of+English+to+Speakers+of+Other+Languages+and+of+Standard+English+as+a+Second+Dialect%22"><i>TESOL Quarterly: A Journal for Teachers of English to Speakers of Other Languages and of Standard English as a Second Dialect</i></searchLink>. 2025 59(1):S251-S279.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 29
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2025
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Criterion+Referenced+Tests%22">Criterion Referenced Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Essay+Tests%22">Essay Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Automation%22">Automation</searchLink><br /><searchLink fieldCode="DE" term="%22Writing+Evaluation%22">Writing Evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Scoring%22">Scoring</searchLink><br /><searchLink fieldCode="DE" term="%22Interrater+Reliability%22">Interrater Reliability</searchLink><br /><searchLink fieldCode="DE" term="%22Writing+Tests%22">Writing Tests</searchLink><br /><searchLink fieldCode="DE" term="%22English+%28Second+Language%29%22">English (Second Language)</searchLink><br /><searchLink fieldCode="DE" term="%22International+Organizations%22">International Organizations</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1002/tesq.70011
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0039-8322<br />1545-7249
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Open access to novel AI tools offers unprecedented opportunities for human-AI collaboration in writing instruction and assessment. While research on using generative AI tools like ChatGPT in these contexts is emerging, more is needed to understand their effectiveness as Automated Writing Evaluation (AWE) tools. This study explores the potential of ChatGPT (GPT-3.5) to assist teachers and learners in the North Atlantic Treaty Organization (NATO) by evaluating English writing based on holistic scoring criteria. Using a mixed-methods approach, the study compared ChatGPT's ratings with human ratings on 100 writing tests to assess inter-rater reliability. It also analyzed the justifications provided by both human raters and ChatGPT to evaluate how well ChatGPT understood the rating criteria at different proficiency levels and whether its rationales could provide effective feedback for learners and support teacher feedback practices. Results showed strong agreement between ChatGPT's and human ratings, with ChatGPT demonstrating a similar understanding of the rating scales and offering justifications with elements of effective feedback. These findings indicate that ChatGPT holds promise as an AWE tool, providing meaningful feedback and valuable insights into holistic rating scales. This study encourages further exploration of AI in the L2 classroom and suggests leveraging AI to enhance writing pedagogy and classroom-based assessment.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2025
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1490674
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1490674
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1002/tesq.70011
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 29
        StartPage: S251
    Subjects:
      – SubjectFull: Artificial Intelligence
        Type: general
      – SubjectFull: Criterion Referenced Tests
        Type: general
      – SubjectFull: Essay Tests
        Type: general
      – SubjectFull: Automation
        Type: general
      – SubjectFull: Writing Evaluation
        Type: general
      – SubjectFull: Scoring
        Type: general
      – SubjectFull: Interrater Reliability
        Type: general
      – SubjectFull: Writing Tests
        Type: general
      – SubjectFull: English (Second Language)
        Type: general
      – SubjectFull: International Organizations
        Type: general
    Titles:
      – TitleFull: Exploring the Potential of ChatGPT for Evaluating English Essays in a Criterion-Based Assessment
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Andrea Gjorevski
      – PersonEntity:
          Name:
            NameFull: Mimi Li
      – PersonEntity:
          Name:
            NameFull: Troy L. Cox
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 12
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 0039-8322
            – Type: issn-electronic
              Value: 1545-7249
          Numbering:
            – Type: volume
              Value: 59
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: TESOL Quarterly: A Journal for Teachers of English to Speakers of Other Languages and of Standard English as a Second Dialect
              Type: main
ResultId 1