Designing, evaluating, and deploying automated scoring systems with validity in mind: Methodological design decisions.

Saved in:
Bibliographic Details
Title: Designing, evaluating, and deploying automated scoring systems with validity in mind: Methodological design decisions.
Authors: Rupp, André A.
Source: Applied Measurement in Education. Jul-Sep2018, Vol. 31 Issue 3, p191-214. 24p.
Subjects: Conversation ability testing, Conversation method (Language teaching), Decision making, Professional education, Systems engineering
Abstract: This article discusses critical methodological design decisions for collecting, interpreting, and synthesizing empirical evidence during the design, deployment, and operational quality-control phases for automated scoring systems. The discussion is inspired by work on operational large-scale systems for automated essay scoring but many of the principles have implications for principled reasoning and workflow management for other use contexts. The overall workflow is described as a series of five phases, each one having two critical sub-phases with a large number of associated methodological design decisions. These phases involve assessment design, linguistic component design, model design, model validation, and operational deployment. Through brief examples, the various considerations for these design decisions are illustrated, which have to be carefully weighed in the overall decision-making process for the system in order to unveil the complexities that underlie this work. The article closes with reflections on resource demands as well as recommendations for best practices of interdisciplinary teams who engage in this work, underscoring how this work is a blend of scientific rigor and artful practice. [ABSTRACT FROM AUTHOR]
Copyright of Applied Measurement in Education is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Psychology and Behavioral Sciences Collection
Full text is not displayed to guests.
Description
Abstract:This article discusses critical methodological design decisions for collecting, interpreting, and synthesizing empirical evidence during the design, deployment, and operational quality-control phases for automated scoring systems. The discussion is inspired by work on operational large-scale systems for automated essay scoring but many of the principles have implications for principled reasoning and workflow management for other use contexts. The overall workflow is described as a series of five phases, each one having two critical sub-phases with a large number of associated methodological design decisions. These phases involve assessment design, linguistic component design, model design, model validation, and operational deployment. Through brief examples, the various considerations for these design decisions are illustrated, which have to be carefully weighed in the overall decision-making process for the system in order to unveil the complexities that underlie this work. The article closes with reflections on resource demands as well as recommendations for best practices of interdisciplinary teams who engage in this work, underscoring how this work is a blend of scientific rigor and artful practice. [ABSTRACT FROM AUTHOR]
ISSN:08957347
DOI:10.1080/08957347.2018.1464448