The Efficacy of Using a Large-Language Model as an Item Writing Assistant.

Saved in:
Bibliographic Details
Title: The Efficacy of Using a Large-Language Model as an Item Writing Assistant.
Authors: Jones, Paul E. (AUTHOR), Becker, Kirk A. (AUTHOR)
Source: Applied Measurement in Education. Jul-Dec2025, Vol. 38 Issue 3/4, p217-238. 22p.
Subjects: Language models, Generative pre-trained transformers
Abstract: The authors investigated using a large language model (LLM) for writing test questions for a real estate licensing exam. In Study 1 items were generated by GPT-4 and rated by subject matter experts (SMEs). These items were on-topic,relevant, and generally appropriate. Item difficulty manipulation was ineffective. Cognitive level matching was harder as cognitive level increased. Study 2 compared human and LLM items using SME and content developer ratings. Human and LLM items were similar in blueprint alignment, relevance, factual errors, and key quality. LLM items had better stem quality and cognitive level matching. Human distractors had an edge in quality. In Study 3 investigated content overlap and breadth of coverage. Similar prompts frequently generated overlapping content. The range of content represented in large sets of generated items did not cover the breadth of the generating content areas. Results suggest LLMs are as good as SMEs at generating first-draft items. [ABSTRACT FROM AUTHOR]
Copyright of Applied Measurement in Education is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Psychology and Behavioral Sciences Collection
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: pbh
DbLabel: Psychology and Behavioral Sciences Collection
An: 192312334
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: The Efficacy of Using a Large-Language Model as an Item Writing Assistant.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Jones%2C+Paul+E%2E%22">Jones, Paul E.</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Becker%2C+Kirk+A%2E%22">Becker, Kirk A.</searchLink> (AUTHOR)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Applied+Measurement+in+Education%22">Applied Measurement in Education</searchLink>. Jul-Dec2025, Vol. 38 Issue 3/4, p217-238. 22p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Language+models%22">Language models</searchLink><br /><searchLink fieldCode="DE" term="%22Generative+pre-trained+transformers%22">Generative pre-trained transformers</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: The authors investigated using a large language model (LLM) for writing test questions for a real estate licensing exam. In Study 1 items were generated by GPT-4 and rated by subject matter experts (SMEs). These items were on-topic,relevant, and generally appropriate. Item difficulty manipulation was ineffective. Cognitive level matching was harder as cognitive level increased. Study 2 compared human and LLM items using SME and content developer ratings. Human and LLM items were similar in blueprint alignment, relevance, factual errors, and key quality. LLM items had better stem quality and cognitive level matching. Human distractors had an edge in quality. In Study 3 investigated content overlap and breadth of coverage. Similar prompts frequently generated overlapping content. The range of content represented in large sets of generated items did not cover the breadth of the generating content areas. Results suggest LLMs are as good as SMEs at generating first-draft items. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Applied Measurement in Education is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=pbh&AN=192312334
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1080/08957347.2025.2580612
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 22
        StartPage: 217
    Subjects:
      – SubjectFull: Language models
        Type: general
      – SubjectFull: Generative pre-trained transformers
        Type: general
    Titles:
      – TitleFull: The Efficacy of Using a Large-Language Model as an Item Writing Assistant.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Jones, Paul E.
      – PersonEntity:
          Name:
            NameFull: Becker, Kirk A.
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 07
              Text: Jul-Dec2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 08957347
          Numbering:
            – Type: volume
              Value: 38
            – Type: issue
              Value: 3/4
          Titles:
            – TitleFull: Applied Measurement in Education
              Type: main
ResultId 1