Limitations of QWK in Evaluating Automated and Human Scoring Systems
Saved in:
| Title: | Limitations of QWK in Evaluating Automated and Human Scoring Systems |
|---|---|
| Language: | English |
| Authors: | Jennifer Lewis (ORCID |
| Source: | Educational Measurement: Issues and Practice. 2026 45(1). |
| Availability: | Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us |
| Peer Reviewed: | Y |
| Page Count: | 14 |
| Publication Date: | 2026 |
| Document Type: | Journal Articles Reports - Research |
| Descriptors: | Interrater Reliability, Evaluation Methods, Measurement Techniques, Automation, Artificial Intelligence, Scores, Responses |
| DOI: | 10.1111/emip.70017 |
| ISSN: | 0731-1745 1745-3992 |
| Abstract: | To assess the interrater reliability of human ratings of constructed responses (CR), or the accuracy of scores given by automated scoring engines, concordance metrics quantify agreement between measures. This article examines the quadratic weighted kappa (QWK) in these contexts and highlights its practical limitations compared to other metrics. Both empirical and simulation study results reveal how different factors including the shape of the marginal distributions and score scale length may impact the estimates and how we can adjust for these properties of the contingency table. The results highlight the QWK's sensitivities and suggest that additional caution should be taken before decisions about whether to keep a CR item on a test form are made. If using QWK without the proper interpretive supports, such decisions may be misinformed. Consequently, we make suggestions for best practices to promote responsible evaluation of agreement in the context of CR scoring in educational testing. |
| Abstractor: | As Provided |
| Entry Date: | 2026 |
| Accession Number: | EJ1498710 |
| Database: | ERIC |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1498710 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Limitations of QWK in Evaluating Automated and Human Scoring Systems – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Jennifer+Lewis%22">Jennifer Lewis</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-6523-7883">0000-0001-6523-7883</externalLink>)<br /><searchLink fieldCode="AR" term="%22Jodi+M%2E+Casabianca%22">Jodi M. Casabianca</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-1644-6731">0000-0002-1644-6731</externalLink>) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Educational+Measurement%3A+Issues+and+Practice%22"><i>Educational Measurement: Issues and Practice</i></searchLink>. 2026 45(1). – Name: Avail Label: Availability Group: Avail Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 14 – Name: DatePubCY Label: Publication Date Group: Date Data: 2026 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Interrater+Reliability%22">Interrater Reliability</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluation+Methods%22">Evaluation Methods</searchLink><br /><searchLink fieldCode="DE" term="%22Measurement+Techniques%22">Measurement Techniques</searchLink><br /><searchLink fieldCode="DE" term="%22Automation%22">Automation</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Scores%22">Scores</searchLink><br /><searchLink fieldCode="DE" term="%22Responses%22">Responses</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1111/emip.70017 – Name: ISSN Label: ISSN Group: ISSN Data: 0731-1745<br />1745-3992 – Name: Abstract Label: Abstract Group: Ab Data: To assess the interrater reliability of human ratings of constructed responses (CR), or the accuracy of scores given by automated scoring engines, concordance metrics quantify agreement between measures. This article examines the quadratic weighted kappa (QWK) in these contexts and highlights its practical limitations compared to other metrics. Both empirical and simulation study results reveal how different factors including the shape of the marginal distributions and score scale length may impact the estimates and how we can adjust for these properties of the contingency table. The results highlight the QWK's sensitivities and suggest that additional caution should be taken before decisions about whether to keep a CR item on a test form are made. If using QWK without the proper interpretive supports, such decisions may be misinformed. Consequently, we make suggestions for best practices to promote responsible evaluation of agreement in the context of CR scoring in educational testing. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2026 – Name: AN Label: Accession Number Group: ID Data: EJ1498710 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1498710 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1111/emip.70017 Languages: – Text: English PhysicalDescription: Pagination: PageCount: 14 Subjects: – SubjectFull: Interrater Reliability Type: general – SubjectFull: Evaluation Methods Type: general – SubjectFull: Measurement Techniques Type: general – SubjectFull: Automation Type: general – SubjectFull: Artificial Intelligence Type: general – SubjectFull: Scores Type: general – SubjectFull: Responses Type: general Titles: – TitleFull: Limitations of QWK in Evaluating Automated and Human Scoring Systems Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Jennifer Lewis – PersonEntity: Name: NameFull: Jodi M. Casabianca IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 03 Type: published Y: 2026 Identifiers: – Type: issn-print Value: 0731-1745 – Type: issn-electronic Value: 1745-3992 Numbering: – Type: volume Value: 45 – Type: issue Value: 1 Titles: – TitleFull: Educational Measurement: Issues and Practice Type: main |
| ResultId | 1 |