The Efficacy of Using a Large-Language Model as an Item Writing Assistant.
Saved in:
| Title: | The Efficacy of Using a Large-Language Model as an Item Writing Assistant. |
|---|---|
| Authors: | Jones, Paul E. (AUTHOR), Becker, Kirk A. (AUTHOR) |
| Source: | Applied Measurement in Education. Jul-Dec2025, Vol. 38 Issue 3/4, p217-238. 22p. |
| Subjects: | Language models, Generative pre-trained transformers |
| Abstract: | The authors investigated using a large language model (LLM) for writing test questions for a real estate licensing exam. In Study 1 items were generated by GPT-4 and rated by subject matter experts (SMEs). These items were on-topic,relevant, and generally appropriate. Item difficulty manipulation was ineffective. Cognitive level matching was harder as cognitive level increased. Study 2 compared human and LLM items using SME and content developer ratings. Human and LLM items were similar in blueprint alignment, relevance, factual errors, and key quality. LLM items had better stem quality and cognitive level matching. Human distractors had an edge in quality. In Study 3 investigated content overlap and breadth of coverage. Similar prompts frequently generated overlapping content. The range of content represented in large sets of generated items did not cover the breadth of the generating content areas. Results suggest LLMs are as good as SMEs at generating first-draft items. [ABSTRACT FROM AUTHOR] |
| Copyright of Applied Measurement in Education is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Psychology and Behavioral Sciences Collection |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Text: Availability: 1 |
|---|---|
| Header | DbId: pbh DbLabel: Psychology and Behavioral Sciences Collection An: 192312334 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: The Efficacy of Using a Large-Language Model as an Item Writing Assistant. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Jones%2C+Paul+E%2E%22">Jones, Paul E.</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Becker%2C+Kirk+A%2E%22">Becker, Kirk A.</searchLink> (AUTHOR) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Applied+Measurement+in+Education%22">Applied Measurement in Education</searchLink>. Jul-Dec2025, Vol. 38 Issue 3/4, p217-238. 22p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Language+models%22">Language models</searchLink><br /><searchLink fieldCode="DE" term="%22Generative+pre-trained+transformers%22">Generative pre-trained transformers</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: The authors investigated using a large language model (LLM) for writing test questions for a real estate licensing exam. In Study 1 items were generated by GPT-4 and rated by subject matter experts (SMEs). These items were on-topic,relevant, and generally appropriate. Item difficulty manipulation was ineffective. Cognitive level matching was harder as cognitive level increased. Study 2 compared human and LLM items using SME and content developer ratings. Human and LLM items were similar in blueprint alignment, relevance, factual errors, and key quality. LLM items had better stem quality and cognitive level matching. Human distractors had an edge in quality. In Study 3 investigated content overlap and breadth of coverage. Similar prompts frequently generated overlapping content. The range of content represented in large sets of generated items did not cover the breadth of the generating content areas. Results suggest LLMs are as good as SMEs at generating first-draft items. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Applied Measurement in Education is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=pbh&AN=192312334 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1080/08957347.2025.2580612 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 22 StartPage: 217 Subjects: – SubjectFull: Language models Type: general – SubjectFull: Generative pre-trained transformers Type: general Titles: – TitleFull: The Efficacy of Using a Large-Language Model as an Item Writing Assistant. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Jones, Paul E. – PersonEntity: Name: NameFull: Becker, Kirk A. IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 07 Text: Jul-Dec2025 Type: published Y: 2025 Identifiers: – Type: issn-print Value: 08957347 Numbering: – Type: volume Value: 38 – Type: issue Value: 3/4 Titles: – TitleFull: Applied Measurement in Education Type: main |
| ResultId | 1 |