AI-Based L2 Pragmatic Speaking Assessment: Evidence from Many-Facet Rasch Measurement

Saved in:
Bibliographic Details
Title: AI-Based L2 Pragmatic Speaking Assessment: Evidence from Many-Facet Rasch Measurement
Language: English
Authors: Soo Jung Youn
Source: English Teaching. 2025 80(4):3-28.
Availability: Korea Association of Teachers of English. 6105 English Education Department, Chinju National University of Education, 369beon-gil 3, Jinyangho-ro, Jinju, Gyeongsangnam-do, 52673, Republic of Korea. Tel: +82-42-629-7381; Fax: +82-42-629-7320; e-mail: katejournal29@gmail.com; Web site: https://journal.kate.or.kr/
Peer Reviewed: Y
Page Count: 26
Publication Date: 2025
Document Type: Journal Articles
Reports - Research
Descriptors: Artificial Intelligence, Second Language Learning, Computer Assisted Testing, Pragmatics, Language Tests, Oral Language, Item Response Theory, English (Second Language), Evaluators, Test Items, Difficulty Level, Evaluation Criteria, Scoring, Interrater Reliability
ISSN: 1017-7108
2671-9312
Abstract: With the growing interest in generative AI (GenAI) for language assessment, its potential as a rater has been discussed. This study compares trained human raters' scores with GenAI ratings in assessing L2 pragmatic speaking performance across different task types. Fifty L2 English learners of varying proficiency levels completed pragmatic speaking test items, which were scored by five trained raters and ChatGPT-5. To examine the comparability, many-facet Rasch measurement was employed, focusing on examinees' abilities, raters' severity, item difficulty, and rating criteria functioning. Findings indicated a moderate correlation between GenAI and human ratings in terms of examinee ability. Compared to human raters, ChatGPT exhibited higher internal consistency and produced a narrower examinee ability distribution. ChatGPT ratings tended to focus on explicit features, such as specific conditions in real-life pragmatic tasks and formulaic expressions, while showing inconsistency in scoring off-task performances and implicit sociopragmatic dimensions. These findings are discussed in light of the potential of GenAI for low-stakes classroom assessment.
Abstractor: As Provided
Entry Date: 2026
Accession Number: EJ1493937
Database: ERIC
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwH1MI66VI-qNwOvX7Ktp0L9AAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDEceYEZvXTUQiu3X2wIBEICBmwl5UC-m2GLheSCJWGt4CQQXPwk_rpetCn43Bmv1QSt46LfEYLfvT2RnCPDNL7mriPubEbS0Hu8iQcqBvO1P3yY0mCIpAGFt4FfK-wp7DURUb4T3jHKgHM2MHiO1dbO2Lztc4kd5yK2jKK5tX1so2Iuk4Efy3D0pw9meBy_itDs4yRoQql5Iz9ytoYeA3dIpq2G3-rJwwHdLZkve
Text:
  Availability: 0
CustomLinks:
  – Url: https://eric.ed.gov/contentdelivery/servlet/ERICServlet?accno=EJ1493937
    Name: ERIC Full Text
    Category: fullText
    Text: Full Text from ERIC
Header DbId: eric
DbLabel: ERIC
An: EJ1493937
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: AI-Based L2 Pragmatic Speaking Assessment: Evidence from Many-Facet Rasch Measurement
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Soo+Jung+Youn%22">Soo Jung Youn</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22English+Teaching%22"><i>English Teaching</i></searchLink>. 2025 80(4):3-28.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Korea Association of Teachers of English. 6105 English Education Department, Chinju National University of Education, 369beon-gil 3, Jinyangho-ro, Jinju, Gyeongsangnam-do, 52673, Republic of Korea. Tel: +82-42-629-7381; Fax: +82-42-629-7320; e-mail: katejournal29@gmail.com; Web site: https://journal.kate.or.kr/
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 26
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2025
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Second+Language+Learning%22">Second Language Learning</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+Assisted+Testing%22">Computer Assisted Testing</searchLink><br /><searchLink fieldCode="DE" term="%22Pragmatics%22">Pragmatics</searchLink><br /><searchLink fieldCode="DE" term="%22Language+Tests%22">Language Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Oral+Language%22">Oral Language</searchLink><br /><searchLink fieldCode="DE" term="%22Item+Response+Theory%22">Item Response Theory</searchLink><br /><searchLink fieldCode="DE" term="%22English+%28Second+Language%29%22">English (Second Language)</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluators%22">Evaluators</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Items%22">Test Items</searchLink><br /><searchLink fieldCode="DE" term="%22Difficulty+Level%22">Difficulty Level</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluation+Criteria%22">Evaluation Criteria</searchLink><br /><searchLink fieldCode="DE" term="%22Scoring%22">Scoring</searchLink><br /><searchLink fieldCode="DE" term="%22Interrater+Reliability%22">Interrater Reliability</searchLink>
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 1017-7108<br />2671-9312
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: With the growing interest in generative AI (GenAI) for language assessment, its potential as a rater has been discussed. This study compares trained human raters' scores with GenAI ratings in assessing L2 pragmatic speaking performance across different task types. Fifty L2 English learners of varying proficiency levels completed pragmatic speaking test items, which were scored by five trained raters and ChatGPT-5. To examine the comparability, many-facet Rasch measurement was employed, focusing on examinees' abilities, raters' severity, item difficulty, and rating criteria functioning. Findings indicated a moderate correlation between GenAI and human ratings in terms of examinee ability. Compared to human raters, ChatGPT exhibited higher internal consistency and produced a narrower examinee ability distribution. ChatGPT ratings tended to focus on explicit features, such as specific conditions in real-life pragmatic tasks and formulaic expressions, while showing inconsistency in scoring off-task performances and implicit sociopragmatic dimensions. These findings are discussed in light of the potential of GenAI for low-stakes classroom assessment.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2026
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1493937
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1493937
RecordInfo BibRecord:
  BibEntity:
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 26
        StartPage: 3
    Subjects:
      – SubjectFull: Artificial Intelligence
        Type: general
      – SubjectFull: Second Language Learning
        Type: general
      – SubjectFull: Computer Assisted Testing
        Type: general
      – SubjectFull: Pragmatics
        Type: general
      – SubjectFull: Language Tests
        Type: general
      – SubjectFull: Oral Language
        Type: general
      – SubjectFull: Item Response Theory
        Type: general
      – SubjectFull: English (Second Language)
        Type: general
      – SubjectFull: Evaluators
        Type: general
      – SubjectFull: Test Items
        Type: general
      – SubjectFull: Difficulty Level
        Type: general
      – SubjectFull: Evaluation Criteria
        Type: general
      – SubjectFull: Scoring
        Type: general
      – SubjectFull: Interrater Reliability
        Type: general
    Titles:
      – TitleFull: AI-Based L2 Pragmatic Speaking Assessment: Evidence from Many-Facet Rasch Measurement
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Soo Jung Youn
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 1017-7108
            – Type: issn-electronic
              Value: 2671-9312
          Numbering:
            – Type: volume
              Value: 80
            – Type: issue
              Value: 4
          Titles:
            – TitleFull: English Teaching
              Type: main
ResultId 1