A comparison of the psychometric properties of GPT-4 versus human novice and expert authors of clinically complex MCQs in a mock examination of Australian medical students.

Saved in:
Bibliographic Details
Title: A comparison of the psychometric properties of GPT-4 versus human novice and expert authors of clinically complex MCQs in a mock examination of Australian medical students.
Authors: Wu, Hannah1,2 (AUTHOR) hannah.wu@adelaide.edu.au, Lee, Daniel3 (AUTHOR), Zerner, Toby2 (AUTHOR), Court-Kowalski, Stefan1,2 (AUTHOR), Devitt, Peter2 (AUTHOR), Palmer, Edward3 (AUTHOR)
Source: Medical Teacher. Jan2026, Vol. 48 Issue 1, p74-84. 11p.
Subject Terms: *Generative artificial intelligence, *Data analysis, *Undergraduate programs, *Educational tests & measurements, *Medical students, *Clinical competence, *Comparative studies, Cronbach's alpha, Research evaluation, Fisher exact test, Descriptive statistics, Psychometrics, Analysis of variance, Statistics, Data analysis software
Geographic Terms: Australia
Abstract: Purpose: Creating clinically complex Multiple Choice Questions (MCQs) for medical assessment can be time-consuming. Large language models such as GPT-4, a type of generative artificial intelligence (AI), are a potential MCQ design tool. Evaluating the psychometric properties of AI-generated MCQs is essential to ensuring quality. Methods: A 120-item mock examination was constructed, containing 40 human-generated MCQs at novice item-writer level, 40 at expert level, and 40 AI-generated MCQs. int. All examination items underwent panel review to ensure they tested higher order cognitive skills and met a minimum acceptable standard. The online mock examination was administered to Australian medical students, who were blinded to each item's author. Results: 234 medical students completed the examination. Analysis showed acceptable reliability (Cronbach's 0.836). There were no differences in item difficulty or discrimination between AI, Novice, and Expert items. The mean item difficulty was 'easy' and mean item discrimination 'fair' across all groups. AI items had lower distractor efficiency (39%) compared to Novice items (55%, p = 0.035), but no difference to Expert items (48%, p = 0.382). Conclusions: The psychometric properties of AI-generated MCQs are comparable to human-generated MCQs at both novice and expert level. Item quality can be improved across all author groups. AI-generated items should undergo human review to enhance distractor efficiency. [ABSTRACT FROM AUTHOR]
Copyright of Medical Teacher is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Education Research Complete
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: ehh
DbLabel: Education Research Complete
An: 190717572
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: A comparison of the psychometric properties of GPT-4 versus human novice and expert authors of clinically complex MCQs in a mock examination of Australian medical students.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Wu%2C+Hannah%22">Wu, Hannah</searchLink><relatesTo>1,2</relatesTo> (AUTHOR)<i> hannah.wu@adelaide.edu.au</i><br /><searchLink fieldCode="AR" term="%22Lee%2C+Daniel%22">Lee, Daniel</searchLink><relatesTo>3</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Zerner%2C+Toby%22">Zerner, Toby</searchLink><relatesTo>2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Court-Kowalski%2C+Stefan%22">Court-Kowalski, Stefan</searchLink><relatesTo>1,2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Devitt%2C+Peter%22">Devitt, Peter</searchLink><relatesTo>2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Palmer%2C+Edward%22">Palmer, Edward</searchLink><relatesTo>3</relatesTo> (AUTHOR)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Medical+Teacher%22">Medical Teacher</searchLink>. Jan2026, Vol. 48 Issue 1, p74-84. 11p.
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: *<searchLink fieldCode="DE" term="%22Generative+artificial+intelligence%22">Generative artificial intelligence</searchLink><br />*<searchLink fieldCode="DE" term="%22Data+analysis%22">Data analysis</searchLink><br />*<searchLink fieldCode="DE" term="%22Undergraduate+programs%22">Undergraduate programs</searchLink><br />*<searchLink fieldCode="DE" term="%22Educational+tests+%26+measurements%22">Educational tests & measurements</searchLink><br />*<searchLink fieldCode="DE" term="%22Medical+students%22">Medical students</searchLink><br />*<searchLink fieldCode="DE" term="%22Clinical+competence%22">Clinical competence</searchLink><br />*<searchLink fieldCode="DE" term="%22Comparative+studies%22">Comparative studies</searchLink><br /><searchLink fieldCode="DE" term="%22Cronbach's+alpha%22">Cronbach's alpha</searchLink><br /><searchLink fieldCode="DE" term="%22Research+evaluation%22">Research evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Fisher+exact+test%22">Fisher exact test</searchLink><br /><searchLink fieldCode="DE" term="%22Descriptive+statistics%22">Descriptive statistics</searchLink><br /><searchLink fieldCode="DE" term="%22Psychometrics%22">Psychometrics</searchLink><br /><searchLink fieldCode="DE" term="%22Analysis+of+variance%22">Analysis of variance</searchLink><br /><searchLink fieldCode="DE" term="%22Statistics%22">Statistics</searchLink><br /><searchLink fieldCode="DE" term="%22Data+analysis+software%22">Data analysis software</searchLink>
– Name: SubjectGeographic
  Label: Geographic Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Australia%22">Australia</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Purpose: Creating clinically complex Multiple Choice Questions (MCQs) for medical assessment can be time-consuming. Large language models such as GPT-4, a type of generative artificial intelligence (AI), are a potential MCQ design tool. Evaluating the psychometric properties of AI-generated MCQs is essential to ensuring quality. Methods: A 120-item mock examination was constructed, containing 40 human-generated MCQs at novice item-writer level, 40 at expert level, and 40 AI-generated MCQs. int. All examination items underwent panel review to ensure they tested higher order cognitive skills and met a minimum acceptable standard. The online mock examination was administered to Australian medical students, who were blinded to each item's author. Results: 234 medical students completed the examination. Analysis showed acceptable reliability (Cronbach's 0.836). There were no differences in item difficulty or discrimination between AI, Novice, and Expert items. The mean item difficulty was 'easy' and mean item discrimination 'fair' across all groups. AI items had lower distractor efficiency (39%) compared to Novice items (55%, p = 0.035), but no difference to Expert items (48%, p = 0.382). Conclusions: The psychometric properties of AI-generated MCQs are comparable to human-generated MCQs at both novice and expert level. Item quality can be improved across all author groups. AI-generated items should undergo human review to enhance distractor efficiency. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Medical Teacher is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=ehh&AN=190717572
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1080/0142159X.2025.2513418
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 11
        StartPage: 74
    Subjects:
      – SubjectFull: Generative artificial intelligence
        Type: general
      – SubjectFull: Data analysis
        Type: general
      – SubjectFull: Undergraduate programs
        Type: general
      – SubjectFull: Educational tests & measurements
        Type: general
      – SubjectFull: Medical students
        Type: general
      – SubjectFull: Clinical competence
        Type: general
      – SubjectFull: Comparative studies
        Type: general
      – SubjectFull: Cronbach's alpha
        Type: general
      – SubjectFull: Research evaluation
        Type: general
      – SubjectFull: Fisher exact test
        Type: general
      – SubjectFull: Descriptive statistics
        Type: general
      – SubjectFull: Psychometrics
        Type: general
      – SubjectFull: Analysis of variance
        Type: general
      – SubjectFull: Statistics
        Type: general
      – SubjectFull: Data analysis software
        Type: general
      – SubjectFull: Australia
        Type: general
    Titles:
      – TitleFull: A comparison of the psychometric properties of GPT-4 versus human novice and expert authors of clinically complex MCQs in a mock examination of Australian medical students.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Wu, Hannah
      – PersonEntity:
          Name:
            NameFull: Lee, Daniel
      – PersonEntity:
          Name:
            NameFull: Zerner, Toby
      – PersonEntity:
          Name:
            NameFull: Court-Kowalski, Stefan
      – PersonEntity:
          Name:
            NameFull: Devitt, Peter
      – PersonEntity:
          Name:
            NameFull: Palmer, Edward
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Text: Jan2026
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 0142159X
          Numbering:
            – Type: volume
              Value: 48
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Medical Teacher
              Type: main
ResultId 1