AI-Based L2 Pragmatic Speaking Assessment: Evidence from Many-Facet Rasch Measurement
Saved in:
| Title: | AI-Based L2 Pragmatic Speaking Assessment: Evidence from Many-Facet Rasch Measurement |
|---|---|
| Language: | English |
| Authors: | Soo Jung Youn |
| Source: | English Teaching. 2025 80(4):3-28. |
| Availability: | Korea Association of Teachers of English. 6105 English Education Department, Chinju National University of Education, 369beon-gil 3, Jinyangho-ro, Jinju, Gyeongsangnam-do, 52673, Republic of Korea. Tel: +82-42-629-7381; Fax: +82-42-629-7320; e-mail: katejournal29@gmail.com; Web site: https://journal.kate.or.kr/ |
| Peer Reviewed: | Y |
| Page Count: | 26 |
| Publication Date: | 2025 |
| Document Type: | Journal Articles Reports - Research |
| Descriptors: | Artificial Intelligence, Second Language Learning, Computer Assisted Testing, Pragmatics, Language Tests, Oral Language, Item Response Theory, English (Second Language), Evaluators, Test Items, Difficulty Level, Evaluation Criteria, Scoring, Interrater Reliability |
| ISSN: | 1017-7108 2671-9312 |
| Abstract: | With the growing interest in generative AI (GenAI) for language assessment, its potential as a rater has been discussed. This study compares trained human raters' scores with GenAI ratings in assessing L2 pragmatic speaking performance across different task types. Fifty L2 English learners of varying proficiency levels completed pragmatic speaking test items, which were scored by five trained raters and ChatGPT-5. To examine the comparability, many-facet Rasch measurement was employed, focusing on examinees' abilities, raters' severity, item difficulty, and rating criteria functioning. Findings indicated a moderate correlation between GenAI and human ratings in terms of examinee ability. Compared to human raters, ChatGPT exhibited higher internal consistency and produced a narrower examinee ability distribution. ChatGPT ratings tended to focus on explicit features, such as specific conditions in real-life pragmatic tasks and formulaic expressions, while showing inconsistency in scoring off-task performances and implicit sociopragmatic dimensions. These findings are discussed in light of the potential of GenAI for low-stakes classroom assessment. |
| Abstractor: | As Provided |
| Entry Date: | 2026 |
| Accession Number: | EJ1493937 |
| Database: | ERIC |
| FullText | Links: – Type: pdflink Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwH1MI66VI-qNwOvX7Ktp0L9AAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDEceYEZvXTUQiu3X2wIBEICBmwl5UC-m2GLheSCJWGt4CQQXPwk_rpetCn43Bmv1QSt46LfEYLfvT2RnCPDNL7mriPubEbS0Hu8iQcqBvO1P3yY0mCIpAGFt4FfK-wp7DURUb4T3jHKgHM2MHiO1dbO2Lztc4kd5yK2jKK5tX1so2Iuk4Efy3D0pw9meBy_itDs4yRoQql5Iz9ytoYeA3dIpq2G3-rJwwHdLZkve Text: Availability: 0 CustomLinks: – Url: https://eric.ed.gov/contentdelivery/servlet/ERICServlet?accno=EJ1493937 Name: ERIC Full Text Category: fullText Text: Full Text from ERIC |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1493937 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: AI-Based L2 Pragmatic Speaking Assessment: Evidence from Many-Facet Rasch Measurement – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Soo+Jung+Youn%22">Soo Jung Youn</searchLink> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22English+Teaching%22"><i>English Teaching</i></searchLink>. 2025 80(4):3-28. – Name: Avail Label: Availability Group: Avail Data: Korea Association of Teachers of English. 6105 English Education Department, Chinju National University of Education, 369beon-gil 3, Jinyangho-ro, Jinju, Gyeongsangnam-do, 52673, Republic of Korea. Tel: +82-42-629-7381; Fax: +82-42-629-7320; e-mail: katejournal29@gmail.com; Web site: https://journal.kate.or.kr/ – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 26 – Name: DatePubCY Label: Publication Date Group: Date Data: 2025 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Second+Language+Learning%22">Second Language Learning</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+Assisted+Testing%22">Computer Assisted Testing</searchLink><br /><searchLink fieldCode="DE" term="%22Pragmatics%22">Pragmatics</searchLink><br /><searchLink fieldCode="DE" term="%22Language+Tests%22">Language Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Oral+Language%22">Oral Language</searchLink><br /><searchLink fieldCode="DE" term="%22Item+Response+Theory%22">Item Response Theory</searchLink><br /><searchLink fieldCode="DE" term="%22English+%28Second+Language%29%22">English (Second Language)</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluators%22">Evaluators</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Items%22">Test Items</searchLink><br /><searchLink fieldCode="DE" term="%22Difficulty+Level%22">Difficulty Level</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluation+Criteria%22">Evaluation Criteria</searchLink><br /><searchLink fieldCode="DE" term="%22Scoring%22">Scoring</searchLink><br /><searchLink fieldCode="DE" term="%22Interrater+Reliability%22">Interrater Reliability</searchLink> – Name: ISSN Label: ISSN Group: ISSN Data: 1017-7108<br />2671-9312 – Name: Abstract Label: Abstract Group: Ab Data: With the growing interest in generative AI (GenAI) for language assessment, its potential as a rater has been discussed. This study compares trained human raters' scores with GenAI ratings in assessing L2 pragmatic speaking performance across different task types. Fifty L2 English learners of varying proficiency levels completed pragmatic speaking test items, which were scored by five trained raters and ChatGPT-5. To examine the comparability, many-facet Rasch measurement was employed, focusing on examinees' abilities, raters' severity, item difficulty, and rating criteria functioning. Findings indicated a moderate correlation between GenAI and human ratings in terms of examinee ability. Compared to human raters, ChatGPT exhibited higher internal consistency and produced a narrower examinee ability distribution. ChatGPT ratings tended to focus on explicit features, such as specific conditions in real-life pragmatic tasks and formulaic expressions, while showing inconsistency in scoring off-task performances and implicit sociopragmatic dimensions. These findings are discussed in light of the potential of GenAI for low-stakes classroom assessment. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2026 – Name: AN Label: Accession Number Group: ID Data: EJ1493937 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1493937 |
| RecordInfo | BibRecord: BibEntity: Languages: – Text: English PhysicalDescription: Pagination: PageCount: 26 StartPage: 3 Subjects: – SubjectFull: Artificial Intelligence Type: general – SubjectFull: Second Language Learning Type: general – SubjectFull: Computer Assisted Testing Type: general – SubjectFull: Pragmatics Type: general – SubjectFull: Language Tests Type: general – SubjectFull: Oral Language Type: general – SubjectFull: Item Response Theory Type: general – SubjectFull: English (Second Language) Type: general – SubjectFull: Evaluators Type: general – SubjectFull: Test Items Type: general – SubjectFull: Difficulty Level Type: general – SubjectFull: Evaluation Criteria Type: general – SubjectFull: Scoring Type: general – SubjectFull: Interrater Reliability Type: general Titles: – TitleFull: AI-Based L2 Pragmatic Speaking Assessment: Evidence from Many-Facet Rasch Measurement Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Soo Jung Youn IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 01 Type: published Y: 2025 Identifiers: – Type: issn-print Value: 1017-7108 – Type: issn-electronic Value: 2671-9312 Numbering: – Type: volume Value: 80 – Type: issue Value: 4 Titles: – TitleFull: English Teaching Type: main |
| ResultId | 1 |