The Efficacy of Using a Large-Language Model as an Item Writing Assistant.
Saved in:
| Title: | The Efficacy of Using a Large-Language Model as an Item Writing Assistant. |
|---|---|
| Authors: | Jones, Paul E. (AUTHOR), Becker, Kirk A. (AUTHOR) |
| Source: | Applied Measurement in Education. Jul-Dec2025, Vol. 38 Issue 3/4, p217-238. 22p. |
| Subjects: | Language models, Generative pre-trained transformers |
| Abstract: | The authors investigated using a large language model (LLM) for writing test questions for a real estate licensing exam. In Study 1 items were generated by GPT-4 and rated by subject matter experts (SMEs). These items were on-topic,relevant, and generally appropriate. Item difficulty manipulation was ineffective. Cognitive level matching was harder as cognitive level increased. Study 2 compared human and LLM items using SME and content developer ratings. Human and LLM items were similar in blueprint alignment, relevance, factual errors, and key quality. LLM items had better stem quality and cognitive level matching. Human distractors had an edge in quality. In Study 3 investigated content overlap and breadth of coverage. Similar prompts frequently generated overlapping content. The range of content represented in large sets of generated items did not cover the breadth of the generating content areas. Results suggest LLMs are as good as SMEs at generating first-draft items. [ABSTRACT FROM AUTHOR] |
| Copyright of Applied Measurement in Education is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Psychology and Behavioral Sciences Collection |
|
Full text is not displayed to guests.
Login for full access.
|
|
| Abstract: | The authors investigated using a large language model (LLM) for writing test questions for a real estate licensing exam. In Study 1 items were generated by GPT-4 and rated by subject matter experts (SMEs). These items were on-topic,relevant, and generally appropriate. Item difficulty manipulation was ineffective. Cognitive level matching was harder as cognitive level increased. Study 2 compared human and LLM items using SME and content developer ratings. Human and LLM items were similar in blueprint alignment, relevance, factual errors, and key quality. LLM items had better stem quality and cognitive level matching. Human distractors had an edge in quality. In Study 3 investigated content overlap and breadth of coverage. Similar prompts frequently generated overlapping content. The range of content represented in large sets of generated items did not cover the breadth of the generating content areas. Results suggest LLMs are as good as SMEs at generating first-draft items. [ABSTRACT FROM AUTHOR] |
|---|---|
| ISSN: | 08957347 |
| DOI: | 10.1080/08957347.2025.2580612 |