An Automated Scoring System for the AAPPL Spanish Presentational Writing Tasks
Saved in:
| Title: | An Automated Scoring System for the AAPPL Spanish Presentational Writing Tasks |
|---|---|
| Language: | English |
| Authors: | Erik Voss (ORCID |
| Source: | Foreign Language Annals. 2026 59(1):63-81. |
| Availability: | Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us |
| Peer Reviewed: | Y |
| Page Count: | 19 |
| Publication Date: | 2026 |
| Document Type: | Journal Articles Reports - Research |
| Education Level: | Junior High Schools Middle Schools Secondary Education High Schools |
| Descriptors: | Automation, Scoring, Computer Assisted Testing, Writing Tests, Spanish, Middle School Students, High School Students, Test Construction, Natural Language Processing, Man Machine Systems, Evaluators |
| DOI: | 10.1111/flan.70038 |
| ISSN: | 0015-718X 1944-9720 |
| Abstract: | Reliable rating of large-scale writing tests is a challenging venture. Preparing, calibrating, and providing quality assurance for effective and efficient raters is critical to the operation of these tests (Hughes, 2003). To address this challenge, testing organizations have employed different approaches to automated scoring using artificial intelligence (AI). Claims about automated scoring of such large-scale writing tests suggest that they are as or more reliable than human raters. However, there is limited research on the implementation of automated scoring for tests of writing in languages other than English. This paper describes research on the development and evaluation of an automated scoring system for a large-scale Spanish language writing test for middle and high school students. The model, consistent with the philosophy behind the test development and scoring system, focuses primarily on the features examinees can do and produce rather than a subtractive model that places more focus on examinees' mistakes. The system was trained on completed operational responses and with ratings from certified human raters. Natural Language Processing techniques were used to identify specific linguistic features. Results show equal or better agreement between machine-human scores than between two human raters. |
| Abstractor: | As Provided |
| Entry Date: | 2026 |
| Accession Number: | EJ1500432 |
| Database: | ERIC |
| Abstract: | Reliable rating of large-scale writing tests is a challenging venture. Preparing, calibrating, and providing quality assurance for effective and efficient raters is critical to the operation of these tests (Hughes, 2003). To address this challenge, testing organizations have employed different approaches to automated scoring using artificial intelligence (AI). Claims about automated scoring of such large-scale writing tests suggest that they are as or more reliable than human raters. However, there is limited research on the implementation of automated scoring for tests of writing in languages other than English. This paper describes research on the development and evaluation of an automated scoring system for a large-scale Spanish language writing test for middle and high school students. The model, consistent with the philosophy behind the test development and scoring system, focuses primarily on the features examinees can do and produce rather than a subtractive model that places more focus on examinees' mistakes. The system was trained on completed operational responses and with ratings from certified human raters. Natural Language Processing techniques were used to identify specific linguistic features. Results show equal or better agreement between machine-human scores than between two human raters. |
|---|---|
| ISSN: | 0015-718X 1944-9720 |
| DOI: | 10.1111/flan.70038 |