Can AI Assess Writing Skills Like a Human? A Reliability Analysis
Saved in:
| Title: | Can AI Assess Writing Skills Like a Human? A Reliability Analysis |
|---|---|
| Language: | English |
| Authors: | Hüseyin Ataseven (ORCID |
| Source: | Journal of Theoretical Educational Science. 2025 18(4):736-754. |
| Availability: | Afyon Kocatepe University. ANS Kampusu, Egitim Fakultesi, Merkez, Afyonkarahisar 03200, Turkey. Tel: +90-272-2181740; Fax: +90-272-2281418; e-mail: editorkebd@gmail.com; Web site: https://dergipark.org.tr/en/pub/akukeg |
| Peer Reviewed: | Y |
| Page Count: | 19 |
| Publication Date: | 2025 |
| Document Type: | Journal Articles Reports - Research |
| Education Level: | Higher Education Postsecondary Education |
| Descriptors: | Artificial Intelligence, Technology Uses in Education, Writing Skills, Student Evaluation, Scoring, English (Second Language), Interrater Reliability, College Students, Foreign Countries |
| Geographic Terms: | Turkey |
| ISSN: | 1308-1659 |
| Abstract: | This study investigates the reliability and consistency of a custom GPT-based scoring system in comparison to trained human raters, focusing on B1-level opinion paragraphs written by English preparatory students. Addressing the limited evidence on how AI scoring systems align with human evaluations in foreign language contexts, the study provides insights into both strengths and limitations of automated writing assessment. A total of 175 student writings were evaluated twice by human raters and twice by the AI system using analytic rubric. Findings indicate excellent agreement among human raters and high consistency across AI-generated scores, but only moderate alignment between human and AI evaluations, with the AI showing a tendency to assign higher scores and overlook off-topic content. These results suggest that while AI scoring systems offer efficiency and consistency, they still lack the interpretive depth of human judgment. The study highlights the potential of AI as a complementary tool in writing assessment, with practical implications for language testing policy and classroom pedagogy. |
| Abstractor: | As Provided |
| Entry Date: | 2025 |
| Accession Number: | EJ1487685 |
| Database: | ERIC |
| FullText | Text: Availability: 0 CustomLinks: – Url: https://eric.ed.gov/contentdelivery/servlet/ERICServlet?accno=EJ1487685 Name: ERIC Full Text Category: fullText Text: Full Text from ERIC |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1487685 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Can AI Assess Writing Skills Like a Human? A Reliability Analysis – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Hüseyin+Ataseven%22">Hüseyin Ataseven</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-9992-4518">0000-0001-9992-4518</externalLink>)<br /><searchLink fieldCode="AR" term="%22Ömay+Çokluk-Bökeoglu%22">Ömay Çokluk-Bökeoglu</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-3879-9204">0000-0002-3879-9204</externalLink>)<br /><searchLink fieldCode="AR" term="%22Fazilet+Tasdemir%22">Fazilet Tasdemir</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-0430-9094">0000-0002-0430-9094</externalLink>) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Journal+of+Theoretical+Educational+Science%22"><i>Journal of Theoretical Educational Science</i></searchLink>. 2025 18(4):736-754. – Name: Avail Label: Availability Group: Avail Data: Afyon Kocatepe University. ANS Kampusu, Egitim Fakultesi, Merkez, Afyonkarahisar 03200, Turkey. Tel: +90-272-2181740; Fax: +90-272-2281418; e-mail: editorkebd@gmail.com; Web site: https://dergipark.org.tr/en/pub/akukeg – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 19 – Name: DatePubCY Label: Publication Date Group: Date Data: 2025 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Audience Label: Education Level Group: Audnce Data: <searchLink fieldCode="EL" term="%22Higher+Education%22">Higher Education</searchLink><br /><searchLink fieldCode="EL" term="%22Postsecondary+Education%22">Postsecondary Education</searchLink> – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Technology+Uses+in+Education%22">Technology Uses in Education</searchLink><br /><searchLink fieldCode="DE" term="%22Writing+Skills%22">Writing Skills</searchLink><br /><searchLink fieldCode="DE" term="%22Student+Evaluation%22">Student Evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Scoring%22">Scoring</searchLink><br /><searchLink fieldCode="DE" term="%22English+%28Second+Language%29%22">English (Second Language)</searchLink><br /><searchLink fieldCode="DE" term="%22Interrater+Reliability%22">Interrater Reliability</searchLink><br /><searchLink fieldCode="DE" term="%22College+Students%22">College Students</searchLink><br /><searchLink fieldCode="DE" term="%22Foreign+Countries%22">Foreign Countries</searchLink> – Name: Subject Label: Geographic Terms Group: Su Data: <searchLink fieldCode="DE" term="%22Turkey%22">Turkey</searchLink> – Name: ISSN Label: ISSN Group: ISSN Data: 1308-1659 – Name: Abstract Label: Abstract Group: Ab Data: This study investigates the reliability and consistency of a custom GPT-based scoring system in comparison to trained human raters, focusing on B1-level opinion paragraphs written by English preparatory students. Addressing the limited evidence on how AI scoring systems align with human evaluations in foreign language contexts, the study provides insights into both strengths and limitations of automated writing assessment. A total of 175 student writings were evaluated twice by human raters and twice by the AI system using analytic rubric. Findings indicate excellent agreement among human raters and high consistency across AI-generated scores, but only moderate alignment between human and AI evaluations, with the AI showing a tendency to assign higher scores and overlook off-topic content. These results suggest that while AI scoring systems offer efficiency and consistency, they still lack the interpretive depth of human judgment. The study highlights the potential of AI as a complementary tool in writing assessment, with practical implications for language testing policy and classroom pedagogy. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2025 – Name: AN Label: Accession Number Group: ID Data: EJ1487685 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1487685 |
| RecordInfo | BibRecord: BibEntity: Languages: – Text: English PhysicalDescription: Pagination: PageCount: 19 StartPage: 736 Subjects: – SubjectFull: Artificial Intelligence Type: general – SubjectFull: Technology Uses in Education Type: general – SubjectFull: Writing Skills Type: general – SubjectFull: Student Evaluation Type: general – SubjectFull: Scoring Type: general – SubjectFull: English (Second Language) Type: general – SubjectFull: Interrater Reliability Type: general – SubjectFull: College Students Type: general – SubjectFull: Foreign Countries Type: general – SubjectFull: Turkey Type: general Titles: – TitleFull: Can AI Assess Writing Skills Like a Human? A Reliability Analysis Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Hüseyin Ataseven – PersonEntity: Name: NameFull: Ömay Çokluk-Bökeoglu – PersonEntity: Name: NameFull: Fazilet Tasdemir IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 01 Type: published Y: 2025 Identifiers: – Type: issn-electronic Value: 1308-1659 Numbering: – Type: volume Value: 18 – Type: issue Value: 4 Titles: – TitleFull: Journal of Theoretical Educational Science Type: main |
| ResultId | 1 |