Automatic Prompt Engineering for Automatic Scoring

Saved in:
Bibliographic Details
Title: Automatic Prompt Engineering for Automatic Scoring
Language: English
Authors: Mingfeng Xue (ORCID 0000-0002-4801-3754), Yunting Liu (ORCID 0009-0004-9594-9661), Xingyao Xiao (ORCID 0000-0001-8430-0438), Mark Wilson (ORCID 0000-0002-0425-5305)
Source: Journal of Educational Measurement. 2025 62(4):559-587.
Availability: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 29
Publication Date: 2025
Sponsoring Agency: National Science Foundation (NSF)
Contract Number: 2010322
Document Type: Journal Articles
Reports - Research
Descriptors: Computer Assisted Testing, Prompting, Educational Assessment, Automation, Natural Language Processing, Scoring, True Scores, Accuracy, Validity, Scoring Rubrics, Innovation
DOI: 10.1111/jedm.70002
ISSN: 0022-0655
1745-3984
Abstract: Prompts play a crucial role in eliciting accurate outputs from large language models (LLMs). This study examines the effectiveness of an automatic prompt engineering (APE) framework for automatic scoring in educational measurement. We collected constructed-response data from 930 students across 11 items and used human scores as the true labels. A baseline was established by providing LLMs with the original human-scoring instructions and materials. APE was then applied to optimize prompts for each item. We found that on average, APE increased scoring accuracy by 9%; few-shot learning (i.e., giving multiple labeled examples related to the goal) increased APE performance by 2%; a high temperature (i.e., a parameter for output randomness) was needed in at least part of the APE to improve the scoring accuracy; Quadratic Weighted Kappa (QWK) showed a similar pattern. These findings support the use of APE in automatic scoring. Moreover, compared with the manual scoring instructions, APE tended to restate and reformat the scoring prompts, which could give rise to concerns about validity. Thus, the creative variability introduced by LLMs raises considerations about the balance between innovation and adherence to scoring rubrics.
Abstractor: As Provided
Entry Date: 2026
Accession Number: EJ1491371
Database: ERIC
FullText Text:
  Availability: 0
Header DbId: eric
DbLabel: ERIC
An: EJ1491371
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Automatic Prompt Engineering for Automatic Scoring
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Mingfeng+Xue%22">Mingfeng Xue</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-4801-3754">0000-0002-4801-3754</externalLink>)<br /><searchLink fieldCode="AR" term="%22Yunting+Liu%22">Yunting Liu</searchLink> (ORCID <externalLink term="https://orcid.org/0009-0004-9594-9661">0009-0004-9594-9661</externalLink>)<br /><searchLink fieldCode="AR" term="%22Xingyao+Xiao%22">Xingyao Xiao</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-8430-0438">0000-0001-8430-0438</externalLink>)<br /><searchLink fieldCode="AR" term="%22Mark+Wilson%22">Mark Wilson</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-0425-5305">0000-0002-0425-5305</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Journal+of+Educational+Measurement%22"><i>Journal of Educational Measurement</i></searchLink>. 2025 62(4):559-587.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 29
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2025
– Name: SourceSuprt
  Label: Sponsoring Agency
  Group: SrcSuprt
  Data: National Science Foundation (NSF)
– Name: NumberContract
  Label: Contract Number
  Group: NumCntrct
  Data: 2010322
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Computer+Assisted+Testing%22">Computer Assisted Testing</searchLink><br /><searchLink fieldCode="DE" term="%22Prompting%22">Prompting</searchLink><br /><searchLink fieldCode="DE" term="%22Educational+Assessment%22">Educational Assessment</searchLink><br /><searchLink fieldCode="DE" term="%22Automation%22">Automation</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+Language+Processing%22">Natural Language Processing</searchLink><br /><searchLink fieldCode="DE" term="%22Scoring%22">Scoring</searchLink><br /><searchLink fieldCode="DE" term="%22True+Scores%22">True Scores</searchLink><br /><searchLink fieldCode="DE" term="%22Accuracy%22">Accuracy</searchLink><br /><searchLink fieldCode="DE" term="%22Validity%22">Validity</searchLink><br /><searchLink fieldCode="DE" term="%22Scoring+Rubrics%22">Scoring Rubrics</searchLink><br /><searchLink fieldCode="DE" term="%22Innovation%22">Innovation</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1111/jedm.70002
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0022-0655<br />1745-3984
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Prompts play a crucial role in eliciting accurate outputs from large language models (LLMs). This study examines the effectiveness of an automatic prompt engineering (APE) framework for automatic scoring in educational measurement. We collected constructed-response data from 930 students across 11 items and used human scores as the true labels. A baseline was established by providing LLMs with the original human-scoring instructions and materials. APE was then applied to optimize prompts for each item. We found that on average, APE increased scoring accuracy by 9%; few-shot learning (i.e., giving multiple labeled examples related to the goal) increased APE performance by 2%; a high temperature (i.e., a parameter for output randomness) was needed in at least part of the APE to improve the scoring accuracy; Quadratic Weighted Kappa (QWK) showed a similar pattern. These findings support the use of APE in automatic scoring. Moreover, compared with the manual scoring instructions, APE tended to restate and reformat the scoring prompts, which could give rise to concerns about validity. Thus, the creative variability introduced by LLMs raises considerations about the balance between innovation and adherence to scoring rubrics.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2026
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1491371
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1491371
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1111/jedm.70002
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 29
        StartPage: 559
    Subjects:
      – SubjectFull: Computer Assisted Testing
        Type: general
      – SubjectFull: Prompting
        Type: general
      – SubjectFull: Educational Assessment
        Type: general
      – SubjectFull: Automation
        Type: general
      – SubjectFull: Natural Language Processing
        Type: general
      – SubjectFull: Scoring
        Type: general
      – SubjectFull: True Scores
        Type: general
      – SubjectFull: Accuracy
        Type: general
      – SubjectFull: Validity
        Type: general
      – SubjectFull: Scoring Rubrics
        Type: general
      – SubjectFull: Innovation
        Type: general
    Titles:
      – TitleFull: Automatic Prompt Engineering for Automatic Scoring
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Mingfeng Xue
      – PersonEntity:
          Name:
            NameFull: Yunting Liu
      – PersonEntity:
          Name:
            NameFull: Xingyao Xiao
      – PersonEntity:
          Name:
            NameFull: Mark Wilson
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 12
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 0022-0655
            – Type: issn-electronic
              Value: 1745-3984
          Numbering:
            – Type: volume
              Value: 62
            – Type: issue
              Value: 4
          Titles:
            – TitleFull: Journal of Educational Measurement
              Type: main
ResultId 1