Designing, evaluating, and deploying automated scoring systems with validity in mind: Methodological design decisions.

Saved in:
Bibliographic Details
Title: Designing, evaluating, and deploying automated scoring systems with validity in mind: Methodological design decisions.
Authors: Rupp, André A.
Source: Applied Measurement in Education. Jul-Sep2018, Vol. 31 Issue 3, p191-214. 24p.
Subjects: Conversation ability testing, Conversation method (Language teaching), Decision making, Professional education, Systems engineering
Abstract: This article discusses critical methodological design decisions for collecting, interpreting, and synthesizing empirical evidence during the design, deployment, and operational quality-control phases for automated scoring systems. The discussion is inspired by work on operational large-scale systems for automated essay scoring but many of the principles have implications for principled reasoning and workflow management for other use contexts. The overall workflow is described as a series of five phases, each one having two critical sub-phases with a large number of associated methodological design decisions. These phases involve assessment design, linguistic component design, model design, model validation, and operational deployment. Through brief examples, the various considerations for these design decisions are illustrated, which have to be carefully weighed in the overall decision-making process for the system in order to unveil the complexities that underlie this work. The article closes with reflections on resource demands as well as recommendations for best practices of interdisciplinary teams who engage in this work, underscoring how this work is a blend of scientific rigor and artful practice. [ABSTRACT FROM AUTHOR]
Copyright of Applied Measurement in Education is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Psychology and Behavioral Sciences Collection
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: pbh
DbLabel: Psychology and Behavioral Sciences Collection
An: 129702644
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Designing, evaluating, and deploying automated scoring systems with validity in mind: Methodological design decisions.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Rupp%2C+André+A%2E%22">Rupp, André A.</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Applied+Measurement+in+Education%22">Applied Measurement in Education</searchLink>. Jul-Sep2018, Vol. 31 Issue 3, p191-214. 24p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Conversation+ability+testing%22">Conversation ability testing</searchLink><br /><searchLink fieldCode="DE" term="%22Conversation+method+%28Language+teaching%29%22">Conversation method (Language teaching)</searchLink><br /><searchLink fieldCode="DE" term="%22Decision+making%22">Decision making</searchLink><br /><searchLink fieldCode="DE" term="%22Professional+education%22">Professional education</searchLink><br /><searchLink fieldCode="DE" term="%22Systems+engineering%22">Systems engineering</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: This article discusses critical methodological design decisions for collecting, interpreting, and synthesizing empirical evidence during the design, deployment, and operational quality-control phases for automated scoring systems. The discussion is inspired by work on operational large-scale systems for automated essay scoring but many of the principles have implications for principled reasoning and workflow management for other use contexts. The overall workflow is described as a series of five phases, each one having two critical sub-phases with a large number of associated methodological design decisions. These phases involve assessment design, linguistic component design, model design, model validation, and operational deployment. Through brief examples, the various considerations for these design decisions are illustrated, which have to be carefully weighed in the overall decision-making process for the system in order to unveil the complexities that underlie this work. The article closes with reflections on resource demands as well as recommendations for best practices of interdisciplinary teams who engage in this work, underscoring how this work is a blend of scientific rigor and artful practice. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Applied Measurement in Education is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=pbh&AN=129702644
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1080/08957347.2018.1464448
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 24
        StartPage: 191
    Subjects:
      – SubjectFull: Conversation ability testing
        Type: general
      – SubjectFull: Conversation method (Language teaching)
        Type: general
      – SubjectFull: Decision making
        Type: general
      – SubjectFull: Professional education
        Type: general
      – SubjectFull: Systems engineering
        Type: general
    Titles:
      – TitleFull: Designing, evaluating, and deploying automated scoring systems with validity in mind: Methodological design decisions.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Rupp, André A.
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 07
              Text: Jul-Sep2018
              Type: published
              Y: 2018
          Identifiers:
            – Type: issn-print
              Value: 08957347
          Numbering:
            – Type: volume
              Value: 31
            – Type: issue
              Value: 3
          Titles:
            – TitleFull: Applied Measurement in Education
              Type: main
ResultId 1