Artificial Intelligence in International English Language Testing System Writing Assessments: A Comparative Study of Human Ratings and DeepAI

Saved in:
Bibliographic Details
Title: Artificial Intelligence in International English Language Testing System Writing Assessments: A Comparative Study of Human Ratings and DeepAI
Language: English
Authors: Somayeh Fathali (ORCID 0000-0003-3430-7257), Fatemeh Mohajeri (ORCID 0009-0004-8395-1131)
Source: Technology in Language Teaching & Learning. 2025 7(4).
Availability: Castledown Publishers. Ground Level, 470 St Kilda Road, Melbourne, 3004, Australia. Tel: +61-3-7003-8355; e-mail: contact@castledown.com; Web site: https://www.castledown.com/journals/tltl
Peer Reviewed: Y
Page Count: 18
Publication Date: 2025
Document Type: Journal Articles
Reports - Research
Descriptors: English (Second Language), Language Tests, Second Language Learning, Artificial Intelligence, Writing Tests, High Stakes Tests, Foreign Countries, Scoring, Interrater Reliability, Automation
Geographic Terms: Iran (Tehran)
Assessment and Survey Identifiers: International English Language Testing System
ISSN: 2652-1687
Abstract: The International English Language Testing System (IELTS) is a high-stakes exam where Writing Task 2 significantly influences the overall scores, requiring reliable evaluation. While trained human raters perform this task, concerns about subjectivity and inconsistency have led to growing interest in artificial intelligence (AI)-based assessment tools. However, little empirical evidence exists on AI in high-stakes testing, and no study has examined DeepAI in this context. Accordingly, using a repeatedmeasures design, this study investigated the comparability and reliability of human and DeepAI ratings of 145 IELTS Writing Task 2 essays collected from the official IELTS Tehran Test Centre. These essays had been previously scored by certified human examiners and were subsequently rescored by DeepAI using a rubric-based prompt based on IELTS standards. Statistical analyses, including paired sample t-tests and multivariate analysis of variance, were conducted to explore rater differences and scoring alignment. The results revealed no significant differences in the overall band scores between the human and AI assessments; however, minor differences were observed in some specific criteria. Additionally, DeepAI showed strong intra-rater reliability, producing consistent scores over a two-week interval. These findings suggest that DeepAI may serve as a reliable supplementary tool in high-stakes writing assessments. However, full replacement of human judgment remains premature, and a combination of human judgment and AI support may be the most effective approach.
Abstractor: As Provided
Entry Date: 2025
Accession Number: EJ1489503
Database: ERIC
FullText Text:
  Availability: 0
CustomLinks:
  – Url: https://eric.ed.gov/contentdelivery/servlet/ERICServlet?accno=EJ1489503
    Name: ERIC Full Text
    Category: fullText
    Text: Full Text from ERIC
Header DbId: eric
DbLabel: ERIC
An: EJ1489503
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Artificial Intelligence in International English Language Testing System Writing Assessments: A Comparative Study of Human Ratings and DeepAI
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Somayeh+Fathali%22">Somayeh Fathali</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-3430-7257">0000-0003-3430-7257</externalLink>)<br /><searchLink fieldCode="AR" term="%22Fatemeh+Mohajeri%22">Fatemeh Mohajeri</searchLink> (ORCID <externalLink term="https://orcid.org/0009-0004-8395-1131">0009-0004-8395-1131</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Technology+in+Language+Teaching+%26+Learning%22"><i>Technology in Language Teaching & Learning</i></searchLink>. 2025 7(4).
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Castledown Publishers. Ground Level, 470 St Kilda Road, Melbourne, 3004, Australia. Tel: +61-3-7003-8355; e-mail: contact@castledown.com; Web site: https://www.castledown.com/journals/tltl
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 18
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2025
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22English+%28Second+Language%29%22">English (Second Language)</searchLink><br /><searchLink fieldCode="DE" term="%22Language+Tests%22">Language Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Second+Language+Learning%22">Second Language Learning</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Writing+Tests%22">Writing Tests</searchLink><br /><searchLink fieldCode="DE" term="%22High+Stakes+Tests%22">High Stakes Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Foreign+Countries%22">Foreign Countries</searchLink><br /><searchLink fieldCode="DE" term="%22Scoring%22">Scoring</searchLink><br /><searchLink fieldCode="DE" term="%22Interrater+Reliability%22">Interrater Reliability</searchLink><br /><searchLink fieldCode="DE" term="%22Automation%22">Automation</searchLink>
– Name: Subject
  Label: Geographic Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Iran+%28Tehran%29%22">Iran (Tehran)</searchLink>
– Name: SubjectThesaurus
  Label: Assessment and Survey Identifiers
  Group: Su
  Data: <searchLink fieldCode="SU" term="%22International+English+Language+Testing+System%22">International English Language Testing System</searchLink>
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 2652-1687
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: The International English Language Testing System (IELTS) is a high-stakes exam where Writing Task 2 significantly influences the overall scores, requiring reliable evaluation. While trained human raters perform this task, concerns about subjectivity and inconsistency have led to growing interest in artificial intelligence (AI)-based assessment tools. However, little empirical evidence exists on AI in high-stakes testing, and no study has examined DeepAI in this context. Accordingly, using a repeatedmeasures design, this study investigated the comparability and reliability of human and DeepAI ratings of 145 IELTS Writing Task 2 essays collected from the official IELTS Tehran Test Centre. These essays had been previously scored by certified human examiners and were subsequently rescored by DeepAI using a rubric-based prompt based on IELTS standards. Statistical analyses, including paired sample t-tests and multivariate analysis of variance, were conducted to explore rater differences and scoring alignment. The results revealed no significant differences in the overall band scores between the human and AI assessments; however, minor differences were observed in some specific criteria. Additionally, DeepAI showed strong intra-rater reliability, producing consistent scores over a two-week interval. These findings suggest that DeepAI may serve as a reliable supplementary tool in high-stakes writing assessments. However, full replacement of human judgment remains premature, and a combination of human judgment and AI support may be the most effective approach.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2025
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1489503
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1489503
RecordInfo BibRecord:
  BibEntity:
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 18
    Subjects:
      – SubjectFull: English (Second Language)
        Type: general
      – SubjectFull: Language Tests
        Type: general
      – SubjectFull: Second Language Learning
        Type: general
      – SubjectFull: Artificial Intelligence
        Type: general
      – SubjectFull: Writing Tests
        Type: general
      – SubjectFull: High Stakes Tests
        Type: general
      – SubjectFull: Foreign Countries
        Type: general
      – SubjectFull: Scoring
        Type: general
      – SubjectFull: Interrater Reliability
        Type: general
      – SubjectFull: Automation
        Type: general
      – SubjectFull: Iran (Tehran)
        Type: general
      – SubjectFull: International English Language Testing System
        Type: general
    Titles:
      – TitleFull: Artificial Intelligence in International English Language Testing System Writing Assessments: A Comparative Study of Human Ratings and DeepAI
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Somayeh Fathali
      – PersonEntity:
          Name:
            NameFull: Fatemeh Mohajeri
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-electronic
              Value: 2652-1687
          Numbering:
            – Type: volume
              Value: 7
            – Type: issue
              Value: 4
          Titles:
            – TitleFull: Technology in Language Teaching & Learning
              Type: main
ResultId 1