An Automated Scoring System for the AAPPL Spanish Presentational Writing Tasks

Saved in:
Bibliographic Details
Title: An Automated Scoring System for the AAPPL Spanish Presentational Writing Tasks
Language: English
Authors: Erik Voss (ORCID 0000-0001-7011-3084), Kim Sallee, Young-A Son, Margaret E. Malone, Camelot Marshall, Celia Chomon-Zamora
Source: Foreign Language Annals. 2026 59(1):63-81.
Availability: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 19
Publication Date: 2026
Document Type: Journal Articles
Reports - Research
Education Level: Junior High Schools
Middle Schools
Secondary Education
High Schools
Descriptors: Automation, Scoring, Computer Assisted Testing, Writing Tests, Spanish, Middle School Students, High School Students, Test Construction, Natural Language Processing, Man Machine Systems, Evaluators
DOI: 10.1111/flan.70038
ISSN: 0015-718X
1944-9720
Abstract: Reliable rating of large-scale writing tests is a challenging venture. Preparing, calibrating, and providing quality assurance for effective and efficient raters is critical to the operation of these tests (Hughes, 2003). To address this challenge, testing organizations have employed different approaches to automated scoring using artificial intelligence (AI). Claims about automated scoring of such large-scale writing tests suggest that they are as or more reliable than human raters. However, there is limited research on the implementation of automated scoring for tests of writing in languages other than English. This paper describes research on the development and evaluation of an automated scoring system for a large-scale Spanish language writing test for middle and high school students. The model, consistent with the philosophy behind the test development and scoring system, focuses primarily on the features examinees can do and produce rather than a subtractive model that places more focus on examinees' mistakes. The system was trained on completed operational responses and with ratings from certified human raters. Natural Language Processing techniques were used to identify specific linguistic features. Results show equal or better agreement between machine-human scores than between two human raters.
Abstractor: As Provided
Entry Date: 2026
Accession Number: EJ1500432
Database: ERIC
FullText Text:
  Availability: 0
Header DbId: eric
DbLabel: ERIC
An: EJ1500432
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: An Automated Scoring System for the AAPPL Spanish Presentational Writing Tasks
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Erik+Voss%22">Erik Voss</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-7011-3084">0000-0001-7011-3084</externalLink>)<br /><searchLink fieldCode="AR" term="%22Kim+Sallee%22">Kim Sallee</searchLink><br /><searchLink fieldCode="AR" term="%22Young-A+Son%22">Young-A Son</searchLink><br /><searchLink fieldCode="AR" term="%22Margaret+E%2E+Malone%22">Margaret E. Malone</searchLink><br /><searchLink fieldCode="AR" term="%22Camelot+Marshall%22">Camelot Marshall</searchLink><br /><searchLink fieldCode="AR" term="%22Celia+Chomon-Zamora%22">Celia Chomon-Zamora</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Foreign+Language+Annals%22"><i>Foreign Language Annals</i></searchLink>. 2026 59(1):63-81.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 19
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2026
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Audience
  Label: Education Level
  Group: Audnce
  Data: <searchLink fieldCode="EL" term="%22Junior+High+Schools%22">Junior High Schools</searchLink><br /><searchLink fieldCode="EL" term="%22Middle+Schools%22">Middle Schools</searchLink><br /><searchLink fieldCode="EL" term="%22Secondary+Education%22">Secondary Education</searchLink><br /><searchLink fieldCode="EL" term="%22High+Schools%22">High Schools</searchLink>
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Automation%22">Automation</searchLink><br /><searchLink fieldCode="DE" term="%22Scoring%22">Scoring</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+Assisted+Testing%22">Computer Assisted Testing</searchLink><br /><searchLink fieldCode="DE" term="%22Writing+Tests%22">Writing Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Spanish%22">Spanish</searchLink><br /><searchLink fieldCode="DE" term="%22Middle+School+Students%22">Middle School Students</searchLink><br /><searchLink fieldCode="DE" term="%22High+School+Students%22">High School Students</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Construction%22">Test Construction</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+Language+Processing%22">Natural Language Processing</searchLink><br /><searchLink fieldCode="DE" term="%22Man+Machine+Systems%22">Man Machine Systems</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluators%22">Evaluators</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1111/flan.70038
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0015-718X<br />1944-9720
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Reliable rating of large-scale writing tests is a challenging venture. Preparing, calibrating, and providing quality assurance for effective and efficient raters is critical to the operation of these tests (Hughes, 2003). To address this challenge, testing organizations have employed different approaches to automated scoring using artificial intelligence (AI). Claims about automated scoring of such large-scale writing tests suggest that they are as or more reliable than human raters. However, there is limited research on the implementation of automated scoring for tests of writing in languages other than English. This paper describes research on the development and evaluation of an automated scoring system for a large-scale Spanish language writing test for middle and high school students. The model, consistent with the philosophy behind the test development and scoring system, focuses primarily on the features examinees can do and produce rather than a subtractive model that places more focus on examinees' mistakes. The system was trained on completed operational responses and with ratings from certified human raters. Natural Language Processing techniques were used to identify specific linguistic features. Results show equal or better agreement between machine-human scores than between two human raters.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2026
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1500432
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1500432
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1111/flan.70038
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 19
        StartPage: 63
    Subjects:
      – SubjectFull: Automation
        Type: general
      – SubjectFull: Scoring
        Type: general
      – SubjectFull: Computer Assisted Testing
        Type: general
      – SubjectFull: Writing Tests
        Type: general
      – SubjectFull: Spanish
        Type: general
      – SubjectFull: Middle School Students
        Type: general
      – SubjectFull: High School Students
        Type: general
      – SubjectFull: Test Construction
        Type: general
      – SubjectFull: Natural Language Processing
        Type: general
      – SubjectFull: Man Machine Systems
        Type: general
      – SubjectFull: Evaluators
        Type: general
    Titles:
      – TitleFull: An Automated Scoring System for the AAPPL Spanish Presentational Writing Tasks
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Erik Voss
      – PersonEntity:
          Name:
            NameFull: Kim Sallee
      – PersonEntity:
          Name:
            NameFull: Young-A Son
      – PersonEntity:
          Name:
            NameFull: Margaret E. Malone
      – PersonEntity:
          Name:
            NameFull: Camelot Marshall
      – PersonEntity:
          Name:
            NameFull: Celia Chomon-Zamora
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 03
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 0015-718X
            – Type: issn-electronic
              Value: 1944-9720
          Numbering:
            – Type: volume
              Value: 59
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Foreign Language Annals
              Type: main
ResultId 1