Performance of Automated Speech Scoring on Different Low- to Medium-Entropy Item Types for Low-Proficiency English Learners. Research Report. ETS RR-17-12

Saved in:
Bibliographic Details
Title: Performance of Automated Speech Scoring on Different Low- to Medium-Entropy Item Types for Low-Proficiency English Learners. Research Report. ETS RR-17-12
Language: English
Authors: Loukina, Anastassia, Zechner, Klaus, Yoon, Su-Youn, Zhang, Mo, Tao, Jidong, Wang, Xinhao, Lee, Chong Min, Mulholland, Matthew
Source: ETS Research Report Series. Mar 2017.
Availability: Educational Testing Service. Rosedale Road, MS19-R Princeton, NJ 08541. Tel: 609-921-9000; Fax: 609-734-5410; e-mail: RDweb@ets.org; Web site: https://www.ets.org/research/policy_research_reports/ets
Peer Reviewed: Y
Page Count: 19
Publication Date: 2017
Document Type: Journal Articles
Reports - Research
Descriptors: Automation, Scoring, Speech Tests, Test Items, Language Proficiency, English (Second Language), Second Language Learning, Test Scoring Machines, Models, Comparative Analysis
ISSN: 2330-8516
Abstract: This report presents an overview of the "SpeechRater"? automated scoring engine model building and evaluation process for several item types with a focus on a low-English-proficiency test-taker population. We discuss each stage of speech scoring, including automatic speech recognition, filtering models for nonscorable responses, and scoring model building and evaluation and compare how the performance at each step differs between different item types. We conclude by discussing the effect of item type on automated scoring performance. We also give recommendations about what considerations should be taken into account when developing tests for low-proficiency English speakers to obtain reliable scores from an automatic scoring engine.
Abstractor: As Provided
Number of References: 22
Entry Date: 2018
Accession Number: EJ1168911
Database: ERIC
Description
Abstract:This report presents an overview of the "SpeechRater"? automated scoring engine model building and evaluation process for several item types with a focus on a low-English-proficiency test-taker population. We discuss each stage of speech scoring, including automatic speech recognition, filtering models for nonscorable responses, and scoring model building and evaluation and compare how the performance at each step differs between different item types. We conclude by discussing the effect of item type on automated scoring performance. We also give recommendations about what considerations should be taken into account when developing tests for low-proficiency English speakers to obtain reliable scores from an automatic scoring engine.
ISSN:2330-8516