Optimizing Inter-Rater Reliability in Foreign Language Constructed-Response Assessments
Saved in:
| Title: | Optimizing Inter-Rater Reliability in Foreign Language Constructed-Response Assessments |
|---|---|
| Language: | English |
| Authors: | Jose Fabian Elizondo-Gonzalez (ORCID |
| Source: | Research in Pedagogy. 2025 15(2):471-482. |
| Availability: | Preschool Teacher Training College "Mihailo Palov" and Serbian Academy of Education in Belgrade. Omladinski Trg 1, Vrsac, 26300 Serbia. Tel: +381-832517; Fax: +381-832517; Web site: http://research.rs |
| Peer Reviewed: | Y |
| Page Count: | 12 |
| Publication Date: | 2025 |
| Document Type: | Journal Articles Reports - Research |
| Descriptors: | Interrater Reliability, Second Language Learning, English (Second Language), Language Tests, Language Proficiency, Foreign Countries, Test Reliability, Error Patterns, Scoring, Prompting, Responses, Test Construction |
| Geographic Terms: | Costa Rica |
| Assessment and Survey Identifiers: | Test of English as a Foreign Language |
| ISSN: | 2217-7337 2406-2006 |
| Abstract: | This study examines inter-rater reliability in a constructed-response English proficiency test developed by the Foreign Language Assessment Program (PELEx) in Costa Rica. Thirty university instructors completed three writing tasks aligned with A2, B2, and C1 CEFR bands, each scored by two trained raters. Inter-rater reliability was estimated using percent agreement, Cohen's weighted kappa, intraclass correlation coefficients (ICCs), and Generalizability Theory (G-Theory). While traditional estimators suggested good to excellent reliability, G-Theory revealed additional sources of error not accounted for by kappa or ICC, particularly prompt-related variability. For example, in the C1 task, person-by-prompt interaction accounted for over 20% of total variance. These findings suggest that while rater training remains important, prompt-related variability must also be addressed to ensure fairness and score comparability. Incorporating calibrated prompts and structured scoring protocols across proficiency levels may strengthen the reliability of constructed-response tasks, especially in high-stakes settings. |
| Abstractor: | As Provided |
| Entry Date: | 2026 |
| Accession Number: | EJ1498173 |
| Database: | ERIC |
| FullText | Text: Availability: 0 CustomLinks: – Url: https://eric.ed.gov/contentdelivery/servlet/ERICServlet?accno=EJ1498173 Name: ERIC Full Text Category: fullText Text: Full Text from ERIC |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1498173 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Optimizing Inter-Rater Reliability in Foreign Language Constructed-Response Assessments – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Jose+Fabian+Elizondo-Gonzalez%22">Jose Fabian Elizondo-Gonzalez</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-4819-0213">0000-0003-4819-0213</externalLink>) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Research+in+Pedagogy%22"><i>Research in Pedagogy</i></searchLink>. 2025 15(2):471-482. – Name: Avail Label: Availability Group: Avail Data: Preschool Teacher Training College "Mihailo Palov" and Serbian Academy of Education in Belgrade. Omladinski Trg 1, Vrsac, 26300 Serbia. Tel: +381-832517; Fax: +381-832517; Web site: http://research.rs – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 12 – Name: DatePubCY Label: Publication Date Group: Date Data: 2025 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Interrater+Reliability%22">Interrater Reliability</searchLink><br /><searchLink fieldCode="DE" term="%22Second+Language+Learning%22">Second Language Learning</searchLink><br /><searchLink fieldCode="DE" term="%22English+%28Second+Language%29%22">English (Second Language)</searchLink><br /><searchLink fieldCode="DE" term="%22Language+Tests%22">Language Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Language+Proficiency%22">Language Proficiency</searchLink><br /><searchLink fieldCode="DE" term="%22Foreign+Countries%22">Foreign Countries</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Reliability%22">Test Reliability</searchLink><br /><searchLink fieldCode="DE" term="%22Error+Patterns%22">Error Patterns</searchLink><br /><searchLink fieldCode="DE" term="%22Scoring%22">Scoring</searchLink><br /><searchLink fieldCode="DE" term="%22Prompting%22">Prompting</searchLink><br /><searchLink fieldCode="DE" term="%22Responses%22">Responses</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Construction%22">Test Construction</searchLink> – Name: Subject Label: Geographic Terms Group: Su Data: <searchLink fieldCode="DE" term="%22Costa+Rica%22">Costa Rica</searchLink> – Name: SubjectThesaurus Label: Assessment and Survey Identifiers Group: Su Data: <searchLink fieldCode="SU" term="%22Test+of+English+as+a+Foreign+Language%22">Test of English as a Foreign Language</searchLink> – Name: ISSN Label: ISSN Group: ISSN Data: 2217-7337<br />2406-2006 – Name: Abstract Label: Abstract Group: Ab Data: This study examines inter-rater reliability in a constructed-response English proficiency test developed by the Foreign Language Assessment Program (PELEx) in Costa Rica. Thirty university instructors completed three writing tasks aligned with A2, B2, and C1 CEFR bands, each scored by two trained raters. Inter-rater reliability was estimated using percent agreement, Cohen's weighted kappa, intraclass correlation coefficients (ICCs), and Generalizability Theory (G-Theory). While traditional estimators suggested good to excellent reliability, G-Theory revealed additional sources of error not accounted for by kappa or ICC, particularly prompt-related variability. For example, in the C1 task, person-by-prompt interaction accounted for over 20% of total variance. These findings suggest that while rater training remains important, prompt-related variability must also be addressed to ensure fairness and score comparability. Incorporating calibrated prompts and structured scoring protocols across proficiency levels may strengthen the reliability of constructed-response tasks, especially in high-stakes settings. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2026 – Name: AN Label: Accession Number Group: ID Data: EJ1498173 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1498173 |
| RecordInfo | BibRecord: BibEntity: Languages: – Text: English PhysicalDescription: Pagination: PageCount: 12 StartPage: 471 Subjects: – SubjectFull: Interrater Reliability Type: general – SubjectFull: Second Language Learning Type: general – SubjectFull: English (Second Language) Type: general – SubjectFull: Language Tests Type: general – SubjectFull: Language Proficiency Type: general – SubjectFull: Foreign Countries Type: general – SubjectFull: Test Reliability Type: general – SubjectFull: Error Patterns Type: general – SubjectFull: Scoring Type: general – SubjectFull: Prompting Type: general – SubjectFull: Responses Type: general – SubjectFull: Test Construction Type: general – SubjectFull: Costa Rica Type: general – SubjectFull: Test of English as a Foreign Language Type: general Titles: – TitleFull: Optimizing Inter-Rater Reliability in Foreign Language Constructed-Response Assessments Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Jose Fabian Elizondo-Gonzalez IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 01 Type: published Y: 2025 Identifiers: – Type: issn-print Value: 2217-7337 – Type: issn-electronic Value: 2406-2006 Numbering: – Type: volume Value: 15 – Type: issue Value: 2 Titles: – TitleFull: Research in Pedagogy Type: main |
| ResultId | 1 |