Limitations of QWK in Evaluating Automated and Human Scoring Systems

Saved in:
Bibliographic Details
Title: Limitations of QWK in Evaluating Automated and Human Scoring Systems
Language: English
Authors: Jennifer Lewis (ORCID 0000-0001-6523-7883), Jodi M. Casabianca (ORCID 0000-0002-1644-6731)
Source: Educational Measurement: Issues and Practice. 2026 45(1).
Availability: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 14
Publication Date: 2026
Document Type: Journal Articles
Reports - Research
Descriptors: Interrater Reliability, Evaluation Methods, Measurement Techniques, Automation, Artificial Intelligence, Scores, Responses
DOI: 10.1111/emip.70017
ISSN: 0731-1745
1745-3992
Abstract: To assess the interrater reliability of human ratings of constructed responses (CR), or the accuracy of scores given by automated scoring engines, concordance metrics quantify agreement between measures. This article examines the quadratic weighted kappa (QWK) in these contexts and highlights its practical limitations compared to other metrics. Both empirical and simulation study results reveal how different factors including the shape of the marginal distributions and score scale length may impact the estimates and how we can adjust for these properties of the contingency table. The results highlight the QWK's sensitivities and suggest that additional caution should be taken before decisions about whether to keep a CR item on a test form are made. If using QWK without the proper interpretive supports, such decisions may be misinformed. Consequently, we make suggestions for best practices to promote responsible evaluation of agreement in the context of CR scoring in educational testing.
Abstractor: As Provided
Entry Date: 2026
Accession Number: EJ1498710
Database: ERIC
FullText Text:
  Availability: 0
Header DbId: eric
DbLabel: ERIC
An: EJ1498710
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Limitations of QWK in Evaluating Automated and Human Scoring Systems
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Jennifer+Lewis%22">Jennifer Lewis</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-6523-7883">0000-0001-6523-7883</externalLink>)<br /><searchLink fieldCode="AR" term="%22Jodi+M%2E+Casabianca%22">Jodi M. Casabianca</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-1644-6731">0000-0002-1644-6731</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Educational+Measurement%3A+Issues+and+Practice%22"><i>Educational Measurement: Issues and Practice</i></searchLink>. 2026 45(1).
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 14
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2026
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Interrater+Reliability%22">Interrater Reliability</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluation+Methods%22">Evaluation Methods</searchLink><br /><searchLink fieldCode="DE" term="%22Measurement+Techniques%22">Measurement Techniques</searchLink><br /><searchLink fieldCode="DE" term="%22Automation%22">Automation</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Scores%22">Scores</searchLink><br /><searchLink fieldCode="DE" term="%22Responses%22">Responses</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1111/emip.70017
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0731-1745<br />1745-3992
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: To assess the interrater reliability of human ratings of constructed responses (CR), or the accuracy of scores given by automated scoring engines, concordance metrics quantify agreement between measures. This article examines the quadratic weighted kappa (QWK) in these contexts and highlights its practical limitations compared to other metrics. Both empirical and simulation study results reveal how different factors including the shape of the marginal distributions and score scale length may impact the estimates and how we can adjust for these properties of the contingency table. The results highlight the QWK's sensitivities and suggest that additional caution should be taken before decisions about whether to keep a CR item on a test form are made. If using QWK without the proper interpretive supports, such decisions may be misinformed. Consequently, we make suggestions for best practices to promote responsible evaluation of agreement in the context of CR scoring in educational testing.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2026
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1498710
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1498710
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1111/emip.70017
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 14
    Subjects:
      – SubjectFull: Interrater Reliability
        Type: general
      – SubjectFull: Evaluation Methods
        Type: general
      – SubjectFull: Measurement Techniques
        Type: general
      – SubjectFull: Automation
        Type: general
      – SubjectFull: Artificial Intelligence
        Type: general
      – SubjectFull: Scores
        Type: general
      – SubjectFull: Responses
        Type: general
    Titles:
      – TitleFull: Limitations of QWK in Evaluating Automated and Human Scoring Systems
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Jennifer Lewis
      – PersonEntity:
          Name:
            NameFull: Jodi M. Casabianca
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 03
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 0731-1745
            – Type: issn-electronic
              Value: 1745-3992
          Numbering:
            – Type: volume
              Value: 45
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Educational Measurement: Issues and Practice
              Type: main
ResultId 1