Evaluating the value of AI-generated questions for USMLE step 1 preparation: A study using ChatGPT-3.5.

Saved in:
Bibliographic Details
Title: Evaluating the value of AI-generated questions for USMLE step 1 preparation: A study using ChatGPT-3.5.
Authors: Balu, Alan1 agb76@georgetown.edu, Prvulovic, Stefan T.1, Fernandez Perez, Claudia1, Kim, Alexander1, Donoho, Daniel A.2, Keating, Gregory3
Source: Medical Teacher. Oct2025, Vol. 47 Issue 10, p1645-1653. 9p.
Subject Terms: *Generative artificial intelligence, *Self-evaluation, *Medical education, *Educational tests & measurements, *Bloom's taxonomy, *Test design, Scientific observation, Fisher exact test, Descriptive statistics, Thematic analysis, Professional licenses, Data analysis software
Abstract: Purpose: Students are increasingly relying on artificial intelligence (AI) for medical education and exam preparation. However, the factual accuracy and content distribution of AI-generated exam questions for self-assessment have not been systematically investigated. Methods: Curated prompts were created to generate multiple-choice questions matching the USMLE Step 1 examination style. We utilized ChatGPT-3.5 to generate 50 questions and answers based upon each prompt style. We manually examined output for factual accuracy, Bloom's Taxonomy, and category within the USMLE Step 1 content outline. Results: ChatGPT-3.5 generated 150 multiple-choice case-style questions and selected an answer. Overall, 83% of generated multiple questions had no factual inaccuracies and 15% contained one to two factual inaccuracies. With simple prompting, common themes included deep venous thrombosis, myocardial infarction, and thyroid disease. Topic diversity improved by separating content topic generation from question generation, and specificity to Step 1 increased by indicating that "treatment" questions were not desired. Conclusion: We demonstrate that ChatGPT-3.5 can successfully generate Step 1 style questions with reasonable factual accuracy, and this method may be used by medical students preparing for USMLE examinations. While AI-generated questions demonstrated adequate factual accuracy, targeted prompting techniques should be used to overcome ChatGPT's bias towards particular medical conditions. [ABSTRACT FROM AUTHOR]
Copyright of Medical Teacher is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Education Research Complete
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: ehh
DbLabel: Education Research Complete
An: 188100637
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Evaluating the value of AI-generated questions for USMLE step 1 preparation: A study using ChatGPT-3.5.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Balu%2C+Alan%22">Balu, Alan</searchLink><relatesTo>1</relatesTo><i> agb76@georgetown.edu</i><br /><searchLink fieldCode="AR" term="%22Prvulovic%2C+Stefan+T%2E%22">Prvulovic, Stefan T.</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Fernandez+Perez%2C+Claudia%22">Fernandez Perez, Claudia</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Kim%2C+Alexander%22">Kim, Alexander</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Donoho%2C+Daniel+A%2E%22">Donoho, Daniel A.</searchLink><relatesTo>2</relatesTo><br /><searchLink fieldCode="AR" term="%22Keating%2C+Gregory%22">Keating, Gregory</searchLink><relatesTo>3</relatesTo>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Medical+Teacher%22">Medical Teacher</searchLink>. Oct2025, Vol. 47 Issue 10, p1645-1653. 9p.
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: *<searchLink fieldCode="DE" term="%22Generative+artificial+intelligence%22">Generative artificial intelligence</searchLink><br />*<searchLink fieldCode="DE" term="%22Self-evaluation%22">Self-evaluation</searchLink><br />*<searchLink fieldCode="DE" term="%22Medical+education%22">Medical education</searchLink><br />*<searchLink fieldCode="DE" term="%22Educational+tests+%26+measurements%22">Educational tests & measurements</searchLink><br />*<searchLink fieldCode="DE" term="%22Bloom's+taxonomy%22">Bloom's taxonomy</searchLink><br />*<searchLink fieldCode="DE" term="%22Test+design%22">Test design</searchLink><br /><searchLink fieldCode="DE" term="%22Scientific+observation%22">Scientific observation</searchLink><br /><searchLink fieldCode="DE" term="%22Fisher+exact+test%22">Fisher exact test</searchLink><br /><searchLink fieldCode="DE" term="%22Descriptive+statistics%22">Descriptive statistics</searchLink><br /><searchLink fieldCode="DE" term="%22Thematic+analysis%22">Thematic analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Professional+licenses%22">Professional licenses</searchLink><br /><searchLink fieldCode="DE" term="%22Data+analysis+software%22">Data analysis software</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Purpose: Students are increasingly relying on artificial intelligence (AI) for medical education and exam preparation. However, the factual accuracy and content distribution of AI-generated exam questions for self-assessment have not been systematically investigated. Methods: Curated prompts were created to generate multiple-choice questions matching the USMLE Step 1 examination style. We utilized ChatGPT-3.5 to generate 50 questions and answers based upon each prompt style. We manually examined output for factual accuracy, Bloom's Taxonomy, and category within the USMLE Step 1 content outline. Results: ChatGPT-3.5 generated 150 multiple-choice case-style questions and selected an answer. Overall, 83% of generated multiple questions had no factual inaccuracies and 15% contained one to two factual inaccuracies. With simple prompting, common themes included deep venous thrombosis, myocardial infarction, and thyroid disease. Topic diversity improved by separating content topic generation from question generation, and specificity to Step 1 increased by indicating that "treatment" questions were not desired. Conclusion: We demonstrate that ChatGPT-3.5 can successfully generate Step 1 style questions with reasonable factual accuracy, and this method may be used by medical students preparing for USMLE examinations. While AI-generated questions demonstrated adequate factual accuracy, targeted prompting techniques should be used to overcome ChatGPT's bias towards particular medical conditions. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Medical Teacher is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=ehh&AN=188100637
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1080/0142159X.2025.2478872
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 9
        StartPage: 1645
    Subjects:
      – SubjectFull: Generative artificial intelligence
        Type: general
      – SubjectFull: Self-evaluation
        Type: general
      – SubjectFull: Medical education
        Type: general
      – SubjectFull: Educational tests & measurements
        Type: general
      – SubjectFull: Bloom's taxonomy
        Type: general
      – SubjectFull: Test design
        Type: general
      – SubjectFull: Scientific observation
        Type: general
      – SubjectFull: Fisher exact test
        Type: general
      – SubjectFull: Descriptive statistics
        Type: general
      – SubjectFull: Thematic analysis
        Type: general
      – SubjectFull: Professional licenses
        Type: general
      – SubjectFull: Data analysis software
        Type: general
    Titles:
      – TitleFull: Evaluating the value of AI-generated questions for USMLE step 1 preparation: A study using ChatGPT-3.5.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Balu, Alan
      – PersonEntity:
          Name:
            NameFull: Prvulovic, Stefan T.
      – PersonEntity:
          Name:
            NameFull: Fernandez Perez, Claudia
      – PersonEntity:
          Name:
            NameFull: Kim, Alexander
      – PersonEntity:
          Name:
            NameFull: Donoho, Daniel A.
      – PersonEntity:
          Name:
            NameFull: Keating, Gregory
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 10
              Text: Oct2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 0142159X
          Numbering:
            – Type: volume
              Value: 47
            – Type: issue
              Value: 10
          Titles:
            – TitleFull: Medical Teacher
              Type: main
ResultId 1