Artificial Intelligence in International English Language Testing System Writing Assessments: A Comparative Study of Human Ratings and DeepAI
Saved in:
| Title: | Artificial Intelligence in International English Language Testing System Writing Assessments: A Comparative Study of Human Ratings and DeepAI |
|---|---|
| Language: | English |
| Authors: | Somayeh Fathali (ORCID |
| Source: | Technology in Language Teaching & Learning. 2025 7(4). |
| Availability: | Castledown Publishers. Ground Level, 470 St Kilda Road, Melbourne, 3004, Australia. Tel: +61-3-7003-8355; e-mail: contact@castledown.com; Web site: https://www.castledown.com/journals/tltl |
| Peer Reviewed: | Y |
| Page Count: | 18 |
| Publication Date: | 2025 |
| Document Type: | Journal Articles Reports - Research |
| Descriptors: | English (Second Language), Language Tests, Second Language Learning, Artificial Intelligence, Writing Tests, High Stakes Tests, Foreign Countries, Scoring, Interrater Reliability, Automation |
| Geographic Terms: | Iran (Tehran) |
| Assessment and Survey Identifiers: | International English Language Testing System |
| ISSN: | 2652-1687 |
| Abstract: | The International English Language Testing System (IELTS) is a high-stakes exam where Writing Task 2 significantly influences the overall scores, requiring reliable evaluation. While trained human raters perform this task, concerns about subjectivity and inconsistency have led to growing interest in artificial intelligence (AI)-based assessment tools. However, little empirical evidence exists on AI in high-stakes testing, and no study has examined DeepAI in this context. Accordingly, using a repeatedmeasures design, this study investigated the comparability and reliability of human and DeepAI ratings of 145 IELTS Writing Task 2 essays collected from the official IELTS Tehran Test Centre. These essays had been previously scored by certified human examiners and were subsequently rescored by DeepAI using a rubric-based prompt based on IELTS standards. Statistical analyses, including paired sample t-tests and multivariate analysis of variance, were conducted to explore rater differences and scoring alignment. The results revealed no significant differences in the overall band scores between the human and AI assessments; however, minor differences were observed in some specific criteria. Additionally, DeepAI showed strong intra-rater reliability, producing consistent scores over a two-week interval. These findings suggest that DeepAI may serve as a reliable supplementary tool in high-stakes writing assessments. However, full replacement of human judgment remains premature, and a combination of human judgment and AI support may be the most effective approach. |
| Abstractor: | As Provided |
| Entry Date: | 2025 |
| Accession Number: | EJ1489503 |
| Database: | ERIC |
| FullText | Text: Availability: 0 CustomLinks: – Url: https://eric.ed.gov/contentdelivery/servlet/ERICServlet?accno=EJ1489503 Name: ERIC Full Text Category: fullText Text: Full Text from ERIC |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1489503 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Artificial Intelligence in International English Language Testing System Writing Assessments: A Comparative Study of Human Ratings and DeepAI – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Somayeh+Fathali%22">Somayeh Fathali</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-3430-7257">0000-0003-3430-7257</externalLink>)<br /><searchLink fieldCode="AR" term="%22Fatemeh+Mohajeri%22">Fatemeh Mohajeri</searchLink> (ORCID <externalLink term="https://orcid.org/0009-0004-8395-1131">0009-0004-8395-1131</externalLink>) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Technology+in+Language+Teaching+%26+Learning%22"><i>Technology in Language Teaching & Learning</i></searchLink>. 2025 7(4). – Name: Avail Label: Availability Group: Avail Data: Castledown Publishers. Ground Level, 470 St Kilda Road, Melbourne, 3004, Australia. Tel: +61-3-7003-8355; e-mail: contact@castledown.com; Web site: https://www.castledown.com/journals/tltl – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 18 – Name: DatePubCY Label: Publication Date Group: Date Data: 2025 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22English+%28Second+Language%29%22">English (Second Language)</searchLink><br /><searchLink fieldCode="DE" term="%22Language+Tests%22">Language Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Second+Language+Learning%22">Second Language Learning</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Writing+Tests%22">Writing Tests</searchLink><br /><searchLink fieldCode="DE" term="%22High+Stakes+Tests%22">High Stakes Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Foreign+Countries%22">Foreign Countries</searchLink><br /><searchLink fieldCode="DE" term="%22Scoring%22">Scoring</searchLink><br /><searchLink fieldCode="DE" term="%22Interrater+Reliability%22">Interrater Reliability</searchLink><br /><searchLink fieldCode="DE" term="%22Automation%22">Automation</searchLink> – Name: Subject Label: Geographic Terms Group: Su Data: <searchLink fieldCode="DE" term="%22Iran+%28Tehran%29%22">Iran (Tehran)</searchLink> – Name: SubjectThesaurus Label: Assessment and Survey Identifiers Group: Su Data: <searchLink fieldCode="SU" term="%22International+English+Language+Testing+System%22">International English Language Testing System</searchLink> – Name: ISSN Label: ISSN Group: ISSN Data: 2652-1687 – Name: Abstract Label: Abstract Group: Ab Data: The International English Language Testing System (IELTS) is a high-stakes exam where Writing Task 2 significantly influences the overall scores, requiring reliable evaluation. While trained human raters perform this task, concerns about subjectivity and inconsistency have led to growing interest in artificial intelligence (AI)-based assessment tools. However, little empirical evidence exists on AI in high-stakes testing, and no study has examined DeepAI in this context. Accordingly, using a repeatedmeasures design, this study investigated the comparability and reliability of human and DeepAI ratings of 145 IELTS Writing Task 2 essays collected from the official IELTS Tehran Test Centre. These essays had been previously scored by certified human examiners and were subsequently rescored by DeepAI using a rubric-based prompt based on IELTS standards. Statistical analyses, including paired sample t-tests and multivariate analysis of variance, were conducted to explore rater differences and scoring alignment. The results revealed no significant differences in the overall band scores between the human and AI assessments; however, minor differences were observed in some specific criteria. Additionally, DeepAI showed strong intra-rater reliability, producing consistent scores over a two-week interval. These findings suggest that DeepAI may serve as a reliable supplementary tool in high-stakes writing assessments. However, full replacement of human judgment remains premature, and a combination of human judgment and AI support may be the most effective approach. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2025 – Name: AN Label: Accession Number Group: ID Data: EJ1489503 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1489503 |
| RecordInfo | BibRecord: BibEntity: Languages: – Text: English PhysicalDescription: Pagination: PageCount: 18 Subjects: – SubjectFull: English (Second Language) Type: general – SubjectFull: Language Tests Type: general – SubjectFull: Second Language Learning Type: general – SubjectFull: Artificial Intelligence Type: general – SubjectFull: Writing Tests Type: general – SubjectFull: High Stakes Tests Type: general – SubjectFull: Foreign Countries Type: general – SubjectFull: Scoring Type: general – SubjectFull: Interrater Reliability Type: general – SubjectFull: Automation Type: general – SubjectFull: Iran (Tehran) Type: general – SubjectFull: International English Language Testing System Type: general Titles: – TitleFull: Artificial Intelligence in International English Language Testing System Writing Assessments: A Comparative Study of Human Ratings and DeepAI Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Somayeh Fathali – PersonEntity: Name: NameFull: Fatemeh Mohajeri IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 01 Type: published Y: 2025 Identifiers: – Type: issn-electronic Value: 2652-1687 Numbering: – Type: volume Value: 7 – Type: issue Value: 4 Titles: – TitleFull: Technology in Language Teaching & Learning Type: main |
| ResultId | 1 |