Exploring the Potential of ChatGPT for Evaluating English Essays in a Criterion-Based Assessment
Saved in:
| Title: | Exploring the Potential of ChatGPT for Evaluating English Essays in a Criterion-Based Assessment |
|---|---|
| Language: | English |
| Authors: | Andrea Gjorevski, Mimi Li (ORCID |
| Source: | TESOL Quarterly: A Journal for Teachers of English to Speakers of Other Languages and of Standard English as a Second Dialect. 2025 59(1):S251-S279. |
| Availability: | Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us |
| Peer Reviewed: | Y |
| Page Count: | 29 |
| Publication Date: | 2025 |
| Document Type: | Journal Articles Reports - Research |
| Descriptors: | Artificial Intelligence, Criterion Referenced Tests, Essay Tests, Automation, Writing Evaluation, Scoring, Interrater Reliability, Writing Tests, English (Second Language), International Organizations |
| DOI: | 10.1002/tesq.70011 |
| ISSN: | 0039-8322 1545-7249 |
| Abstract: | Open access to novel AI tools offers unprecedented opportunities for human-AI collaboration in writing instruction and assessment. While research on using generative AI tools like ChatGPT in these contexts is emerging, more is needed to understand their effectiveness as Automated Writing Evaluation (AWE) tools. This study explores the potential of ChatGPT (GPT-3.5) to assist teachers and learners in the North Atlantic Treaty Organization (NATO) by evaluating English writing based on holistic scoring criteria. Using a mixed-methods approach, the study compared ChatGPT's ratings with human ratings on 100 writing tests to assess inter-rater reliability. It also analyzed the justifications provided by both human raters and ChatGPT to evaluate how well ChatGPT understood the rating criteria at different proficiency levels and whether its rationales could provide effective feedback for learners and support teacher feedback practices. Results showed strong agreement between ChatGPT's and human ratings, with ChatGPT demonstrating a similar understanding of the rating scales and offering justifications with elements of effective feedback. These findings indicate that ChatGPT holds promise as an AWE tool, providing meaningful feedback and valuable insights into holistic rating scales. This study encourages further exploration of AI in the L2 classroom and suggests leveraging AI to enhance writing pedagogy and classroom-based assessment. |
| Abstractor: | As Provided |
| Entry Date: | 2025 |
| Accession Number: | EJ1490674 |
| Database: | ERIC |
| Abstract: | Open access to novel AI tools offers unprecedented opportunities for human-AI collaboration in writing instruction and assessment. While research on using generative AI tools like ChatGPT in these contexts is emerging, more is needed to understand their effectiveness as Automated Writing Evaluation (AWE) tools. This study explores the potential of ChatGPT (GPT-3.5) to assist teachers and learners in the North Atlantic Treaty Organization (NATO) by evaluating English writing based on holistic scoring criteria. Using a mixed-methods approach, the study compared ChatGPT's ratings with human ratings on 100 writing tests to assess inter-rater reliability. It also analyzed the justifications provided by both human raters and ChatGPT to evaluate how well ChatGPT understood the rating criteria at different proficiency levels and whether its rationales could provide effective feedback for learners and support teacher feedback practices. Results showed strong agreement between ChatGPT's and human ratings, with ChatGPT demonstrating a similar understanding of the rating scales and offering justifications with elements of effective feedback. These findings indicate that ChatGPT holds promise as an AWE tool, providing meaningful feedback and valuable insights into holistic rating scales. This study encourages further exploration of AI in the L2 classroom and suggests leveraging AI to enhance writing pedagogy and classroom-based assessment. |
|---|---|
| ISSN: | 0039-8322 1545-7249 |
| DOI: | 10.1002/tesq.70011 |