An Automated Scoring System for the AAPPL Spanish Presentational Writing Tasks

Saved in:
Bibliographic Details
Title: An Automated Scoring System for the AAPPL Spanish Presentational Writing Tasks
Language: English
Authors: Erik Voss (ORCID 0000-0001-7011-3084), Kim Sallee, Young-A Son, Margaret E. Malone, Camelot Marshall, Celia Chomon-Zamora
Source: Foreign Language Annals. 2026 59(1):63-81.
Availability: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 19
Publication Date: 2026
Document Type: Journal Articles
Reports - Research
Education Level: Junior High Schools
Middle Schools
Secondary Education
High Schools
Descriptors: Automation, Scoring, Computer Assisted Testing, Writing Tests, Spanish, Middle School Students, High School Students, Test Construction, Natural Language Processing, Man Machine Systems, Evaluators
DOI: 10.1111/flan.70038
ISSN: 0015-718X
1944-9720
Abstract: Reliable rating of large-scale writing tests is a challenging venture. Preparing, calibrating, and providing quality assurance for effective and efficient raters is critical to the operation of these tests (Hughes, 2003). To address this challenge, testing organizations have employed different approaches to automated scoring using artificial intelligence (AI). Claims about automated scoring of such large-scale writing tests suggest that they are as or more reliable than human raters. However, there is limited research on the implementation of automated scoring for tests of writing in languages other than English. This paper describes research on the development and evaluation of an automated scoring system for a large-scale Spanish language writing test for middle and high school students. The model, consistent with the philosophy behind the test development and scoring system, focuses primarily on the features examinees can do and produce rather than a subtractive model that places more focus on examinees' mistakes. The system was trained on completed operational responses and with ratings from certified human raters. Natural Language Processing techniques were used to identify specific linguistic features. Results show equal or better agreement between machine-human scores than between two human raters.
Abstractor: As Provided
Entry Date: 2026
Accession Number: EJ1500432
Database: ERIC
Description
Abstract:Reliable rating of large-scale writing tests is a challenging venture. Preparing, calibrating, and providing quality assurance for effective and efficient raters is critical to the operation of these tests (Hughes, 2003). To address this challenge, testing organizations have employed different approaches to automated scoring using artificial intelligence (AI). Claims about automated scoring of such large-scale writing tests suggest that they are as or more reliable than human raters. However, there is limited research on the implementation of automated scoring for tests of writing in languages other than English. This paper describes research on the development and evaluation of an automated scoring system for a large-scale Spanish language writing test for middle and high school students. The model, consistent with the philosophy behind the test development and scoring system, focuses primarily on the features examinees can do and produce rather than a subtractive model that places more focus on examinees' mistakes. The system was trained on completed operational responses and with ratings from certified human raters. Natural Language Processing techniques were used to identify specific linguistic features. Results show equal or better agreement between machine-human scores than between two human raters.
ISSN:0015-718X
1944-9720
DOI:10.1111/flan.70038