Can AI Assess Writing Skills Like a Human? A Reliability Analysis

Saved in:
Bibliographic Details
Title: Can AI Assess Writing Skills Like a Human? A Reliability Analysis
Language: English
Authors: Hüseyin Ataseven (ORCID 0000-0001-9992-4518), Ömay Çokluk-Bökeoglu (ORCID 0000-0002-3879-9204), Fazilet Tasdemir (ORCID 0000-0002-0430-9094)
Source: Journal of Theoretical Educational Science. 2025 18(4):736-754.
Availability: Afyon Kocatepe University. ANS Kampusu, Egitim Fakultesi, Merkez, Afyonkarahisar 03200, Turkey. Tel: +90-272-2181740; Fax: +90-272-2281418; e-mail: editorkebd@gmail.com; Web site: https://dergipark.org.tr/en/pub/akukeg
Peer Reviewed: Y
Page Count: 19
Publication Date: 2025
Document Type: Journal Articles
Reports - Research
Education Level: Higher Education
Postsecondary Education
Descriptors: Artificial Intelligence, Technology Uses in Education, Writing Skills, Student Evaluation, Scoring, English (Second Language), Interrater Reliability, College Students, Foreign Countries
Geographic Terms: Turkey
ISSN: 1308-1659
Abstract: This study investigates the reliability and consistency of a custom GPT-based scoring system in comparison to trained human raters, focusing on B1-level opinion paragraphs written by English preparatory students. Addressing the limited evidence on how AI scoring systems align with human evaluations in foreign language contexts, the study provides insights into both strengths and limitations of automated writing assessment. A total of 175 student writings were evaluated twice by human raters and twice by the AI system using analytic rubric. Findings indicate excellent agreement among human raters and high consistency across AI-generated scores, but only moderate alignment between human and AI evaluations, with the AI showing a tendency to assign higher scores and overlook off-topic content. These results suggest that while AI scoring systems offer efficiency and consistency, they still lack the interpretive depth of human judgment. The study highlights the potential of AI as a complementary tool in writing assessment, with practical implications for language testing policy and classroom pedagogy.
Abstractor: As Provided
Entry Date: 2025
Accession Number: EJ1487685
Database: ERIC
FullText Text:
  Availability: 0
CustomLinks:
  – Url: https://eric.ed.gov/contentdelivery/servlet/ERICServlet?accno=EJ1487685
    Name: ERIC Full Text
    Category: fullText
    Text: Full Text from ERIC
Header DbId: eric
DbLabel: ERIC
An: EJ1487685
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Can AI Assess Writing Skills Like a Human? A Reliability Analysis
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Hüseyin+Ataseven%22">Hüseyin Ataseven</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-9992-4518">0000-0001-9992-4518</externalLink>)<br /><searchLink fieldCode="AR" term="%22Ömay+Çokluk-Bökeoglu%22">Ömay Çokluk-Bökeoglu</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-3879-9204">0000-0002-3879-9204</externalLink>)<br /><searchLink fieldCode="AR" term="%22Fazilet+Tasdemir%22">Fazilet Tasdemir</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-0430-9094">0000-0002-0430-9094</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Journal+of+Theoretical+Educational+Science%22"><i>Journal of Theoretical Educational Science</i></searchLink>. 2025 18(4):736-754.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Afyon Kocatepe University. ANS Kampusu, Egitim Fakultesi, Merkez, Afyonkarahisar 03200, Turkey. Tel: +90-272-2181740; Fax: +90-272-2281418; e-mail: editorkebd@gmail.com; Web site: https://dergipark.org.tr/en/pub/akukeg
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 19
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2025
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Audience
  Label: Education Level
  Group: Audnce
  Data: <searchLink fieldCode="EL" term="%22Higher+Education%22">Higher Education</searchLink><br /><searchLink fieldCode="EL" term="%22Postsecondary+Education%22">Postsecondary Education</searchLink>
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Technology+Uses+in+Education%22">Technology Uses in Education</searchLink><br /><searchLink fieldCode="DE" term="%22Writing+Skills%22">Writing Skills</searchLink><br /><searchLink fieldCode="DE" term="%22Student+Evaluation%22">Student Evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Scoring%22">Scoring</searchLink><br /><searchLink fieldCode="DE" term="%22English+%28Second+Language%29%22">English (Second Language)</searchLink><br /><searchLink fieldCode="DE" term="%22Interrater+Reliability%22">Interrater Reliability</searchLink><br /><searchLink fieldCode="DE" term="%22College+Students%22">College Students</searchLink><br /><searchLink fieldCode="DE" term="%22Foreign+Countries%22">Foreign Countries</searchLink>
– Name: Subject
  Label: Geographic Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Turkey%22">Turkey</searchLink>
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 1308-1659
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: This study investigates the reliability and consistency of a custom GPT-based scoring system in comparison to trained human raters, focusing on B1-level opinion paragraphs written by English preparatory students. Addressing the limited evidence on how AI scoring systems align with human evaluations in foreign language contexts, the study provides insights into both strengths and limitations of automated writing assessment. A total of 175 student writings were evaluated twice by human raters and twice by the AI system using analytic rubric. Findings indicate excellent agreement among human raters and high consistency across AI-generated scores, but only moderate alignment between human and AI evaluations, with the AI showing a tendency to assign higher scores and overlook off-topic content. These results suggest that while AI scoring systems offer efficiency and consistency, they still lack the interpretive depth of human judgment. The study highlights the potential of AI as a complementary tool in writing assessment, with practical implications for language testing policy and classroom pedagogy.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2025
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1487685
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1487685
RecordInfo BibRecord:
  BibEntity:
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 19
        StartPage: 736
    Subjects:
      – SubjectFull: Artificial Intelligence
        Type: general
      – SubjectFull: Technology Uses in Education
        Type: general
      – SubjectFull: Writing Skills
        Type: general
      – SubjectFull: Student Evaluation
        Type: general
      – SubjectFull: Scoring
        Type: general
      – SubjectFull: English (Second Language)
        Type: general
      – SubjectFull: Interrater Reliability
        Type: general
      – SubjectFull: College Students
        Type: general
      – SubjectFull: Foreign Countries
        Type: general
      – SubjectFull: Turkey
        Type: general
    Titles:
      – TitleFull: Can AI Assess Writing Skills Like a Human? A Reliability Analysis
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Hüseyin Ataseven
      – PersonEntity:
          Name:
            NameFull: Ömay Çokluk-Bökeoglu
      – PersonEntity:
          Name:
            NameFull: Fazilet Tasdemir
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-electronic
              Value: 1308-1659
          Numbering:
            – Type: volume
              Value: 18
            – Type: issue
              Value: 4
          Titles:
            – TitleFull: Journal of Theoretical Educational Science
              Type: main
ResultId 1