Automatic Prompt Engineering for Automatic Scoring
Saved in:
| Title: | Automatic Prompt Engineering for Automatic Scoring |
|---|---|
| Language: | English |
| Authors: | Mingfeng Xue (ORCID |
| Source: | Journal of Educational Measurement. 2025 62(4):559-587. |
| Availability: | Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us |
| Peer Reviewed: | Y |
| Page Count: | 29 |
| Publication Date: | 2025 |
| Sponsoring Agency: | National Science Foundation (NSF) |
| Contract Number: | 2010322 |
| Document Type: | Journal Articles Reports - Research |
| Descriptors: | Computer Assisted Testing, Prompting, Educational Assessment, Automation, Natural Language Processing, Scoring, True Scores, Accuracy, Validity, Scoring Rubrics, Innovation |
| DOI: | 10.1111/jedm.70002 |
| ISSN: | 0022-0655 1745-3984 |
| Abstract: | Prompts play a crucial role in eliciting accurate outputs from large language models (LLMs). This study examines the effectiveness of an automatic prompt engineering (APE) framework for automatic scoring in educational measurement. We collected constructed-response data from 930 students across 11 items and used human scores as the true labels. A baseline was established by providing LLMs with the original human-scoring instructions and materials. APE was then applied to optimize prompts for each item. We found that on average, APE increased scoring accuracy by 9%; few-shot learning (i.e., giving multiple labeled examples related to the goal) increased APE performance by 2%; a high temperature (i.e., a parameter for output randomness) was needed in at least part of the APE to improve the scoring accuracy; Quadratic Weighted Kappa (QWK) showed a similar pattern. These findings support the use of APE in automatic scoring. Moreover, compared with the manual scoring instructions, APE tended to restate and reformat the scoring prompts, which could give rise to concerns about validity. Thus, the creative variability introduced by LLMs raises considerations about the balance between innovation and adherence to scoring rubrics. |
| Abstractor: | As Provided |
| Entry Date: | 2026 |
| Accession Number: | EJ1491371 |
| Database: | ERIC |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1491371 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Automatic Prompt Engineering for Automatic Scoring – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Mingfeng+Xue%22">Mingfeng Xue</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-4801-3754">0000-0002-4801-3754</externalLink>)<br /><searchLink fieldCode="AR" term="%22Yunting+Liu%22">Yunting Liu</searchLink> (ORCID <externalLink term="https://orcid.org/0009-0004-9594-9661">0009-0004-9594-9661</externalLink>)<br /><searchLink fieldCode="AR" term="%22Xingyao+Xiao%22">Xingyao Xiao</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-8430-0438">0000-0001-8430-0438</externalLink>)<br /><searchLink fieldCode="AR" term="%22Mark+Wilson%22">Mark Wilson</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-0425-5305">0000-0002-0425-5305</externalLink>) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Journal+of+Educational+Measurement%22"><i>Journal of Educational Measurement</i></searchLink>. 2025 62(4):559-587. – Name: Avail Label: Availability Group: Avail Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 29 – Name: DatePubCY Label: Publication Date Group: Date Data: 2025 – Name: SourceSuprt Label: Sponsoring Agency Group: SrcSuprt Data: National Science Foundation (NSF) – Name: NumberContract Label: Contract Number Group: NumCntrct Data: 2010322 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Computer+Assisted+Testing%22">Computer Assisted Testing</searchLink><br /><searchLink fieldCode="DE" term="%22Prompting%22">Prompting</searchLink><br /><searchLink fieldCode="DE" term="%22Educational+Assessment%22">Educational Assessment</searchLink><br /><searchLink fieldCode="DE" term="%22Automation%22">Automation</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+Language+Processing%22">Natural Language Processing</searchLink><br /><searchLink fieldCode="DE" term="%22Scoring%22">Scoring</searchLink><br /><searchLink fieldCode="DE" term="%22True+Scores%22">True Scores</searchLink><br /><searchLink fieldCode="DE" term="%22Accuracy%22">Accuracy</searchLink><br /><searchLink fieldCode="DE" term="%22Validity%22">Validity</searchLink><br /><searchLink fieldCode="DE" term="%22Scoring+Rubrics%22">Scoring Rubrics</searchLink><br /><searchLink fieldCode="DE" term="%22Innovation%22">Innovation</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1111/jedm.70002 – Name: ISSN Label: ISSN Group: ISSN Data: 0022-0655<br />1745-3984 – Name: Abstract Label: Abstract Group: Ab Data: Prompts play a crucial role in eliciting accurate outputs from large language models (LLMs). This study examines the effectiveness of an automatic prompt engineering (APE) framework for automatic scoring in educational measurement. We collected constructed-response data from 930 students across 11 items and used human scores as the true labels. A baseline was established by providing LLMs with the original human-scoring instructions and materials. APE was then applied to optimize prompts for each item. We found that on average, APE increased scoring accuracy by 9%; few-shot learning (i.e., giving multiple labeled examples related to the goal) increased APE performance by 2%; a high temperature (i.e., a parameter for output randomness) was needed in at least part of the APE to improve the scoring accuracy; Quadratic Weighted Kappa (QWK) showed a similar pattern. These findings support the use of APE in automatic scoring. Moreover, compared with the manual scoring instructions, APE tended to restate and reformat the scoring prompts, which could give rise to concerns about validity. Thus, the creative variability introduced by LLMs raises considerations about the balance between innovation and adherence to scoring rubrics. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2026 – Name: AN Label: Accession Number Group: ID Data: EJ1491371 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1491371 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1111/jedm.70002 Languages: – Text: English PhysicalDescription: Pagination: PageCount: 29 StartPage: 559 Subjects: – SubjectFull: Computer Assisted Testing Type: general – SubjectFull: Prompting Type: general – SubjectFull: Educational Assessment Type: general – SubjectFull: Automation Type: general – SubjectFull: Natural Language Processing Type: general – SubjectFull: Scoring Type: general – SubjectFull: True Scores Type: general – SubjectFull: Accuracy Type: general – SubjectFull: Validity Type: general – SubjectFull: Scoring Rubrics Type: general – SubjectFull: Innovation Type: general Titles: – TitleFull: Automatic Prompt Engineering for Automatic Scoring Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Mingfeng Xue – PersonEntity: Name: NameFull: Yunting Liu – PersonEntity: Name: NameFull: Xingyao Xiao – PersonEntity: Name: NameFull: Mark Wilson IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 12 Type: published Y: 2025 Identifiers: – Type: issn-print Value: 0022-0655 – Type: issn-electronic Value: 1745-3984 Numbering: – Type: volume Value: 62 – Type: issue Value: 4 Titles: – TitleFull: Journal of Educational Measurement Type: main |
| ResultId | 1 |