Evaluating the value of AI-generated questions for USMLE step 1 preparation: A study using ChatGPT-3.5.
Saved in:
| Title: | Evaluating the value of AI-generated questions for USMLE step 1 preparation: A study using ChatGPT-3.5. |
|---|---|
| Authors: | Balu, Alan1 agb76@georgetown.edu, Prvulovic, Stefan T.1, Fernandez Perez, Claudia1, Kim, Alexander1, Donoho, Daniel A.2, Keating, Gregory3 |
| Source: | Medical Teacher. Oct2025, Vol. 47 Issue 10, p1645-1653. 9p. |
| Subject Terms: | *Generative artificial intelligence, *Self-evaluation, *Medical education, *Educational tests & measurements, *Bloom's taxonomy, *Test design, Scientific observation, Fisher exact test, Descriptive statistics, Thematic analysis, Professional licenses, Data analysis software |
| Abstract: | Purpose: Students are increasingly relying on artificial intelligence (AI) for medical education and exam preparation. However, the factual accuracy and content distribution of AI-generated exam questions for self-assessment have not been systematically investigated. Methods: Curated prompts were created to generate multiple-choice questions matching the USMLE Step 1 examination style. We utilized ChatGPT-3.5 to generate 50 questions and answers based upon each prompt style. We manually examined output for factual accuracy, Bloom's Taxonomy, and category within the USMLE Step 1 content outline. Results: ChatGPT-3.5 generated 150 multiple-choice case-style questions and selected an answer. Overall, 83% of generated multiple questions had no factual inaccuracies and 15% contained one to two factual inaccuracies. With simple prompting, common themes included deep venous thrombosis, myocardial infarction, and thyroid disease. Topic diversity improved by separating content topic generation from question generation, and specificity to Step 1 increased by indicating that "treatment" questions were not desired. Conclusion: We demonstrate that ChatGPT-3.5 can successfully generate Step 1 style questions with reasonable factual accuracy, and this method may be used by medical students preparing for USMLE examinations. While AI-generated questions demonstrated adequate factual accuracy, targeted prompting techniques should be used to overcome ChatGPT's bias towards particular medical conditions. [ABSTRACT FROM AUTHOR] |
| Copyright of Medical Teacher is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Education Research Complete |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Text: Availability: 1 |
|---|---|
| Header | DbId: ehh DbLabel: Education Research Complete An: 188100637 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Evaluating the value of AI-generated questions for USMLE step 1 preparation: A study using ChatGPT-3.5. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Balu%2C+Alan%22">Balu, Alan</searchLink><relatesTo>1</relatesTo><i> agb76@georgetown.edu</i><br /><searchLink fieldCode="AR" term="%22Prvulovic%2C+Stefan+T%2E%22">Prvulovic, Stefan T.</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Fernandez+Perez%2C+Claudia%22">Fernandez Perez, Claudia</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Kim%2C+Alexander%22">Kim, Alexander</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Donoho%2C+Daniel+A%2E%22">Donoho, Daniel A.</searchLink><relatesTo>2</relatesTo><br /><searchLink fieldCode="AR" term="%22Keating%2C+Gregory%22">Keating, Gregory</searchLink><relatesTo>3</relatesTo> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Medical+Teacher%22">Medical Teacher</searchLink>. Oct2025, Vol. 47 Issue 10, p1645-1653. 9p. – Name: Subject Label: Subject Terms Group: Su Data: *<searchLink fieldCode="DE" term="%22Generative+artificial+intelligence%22">Generative artificial intelligence</searchLink><br />*<searchLink fieldCode="DE" term="%22Self-evaluation%22">Self-evaluation</searchLink><br />*<searchLink fieldCode="DE" term="%22Medical+education%22">Medical education</searchLink><br />*<searchLink fieldCode="DE" term="%22Educational+tests+%26+measurements%22">Educational tests & measurements</searchLink><br />*<searchLink fieldCode="DE" term="%22Bloom's+taxonomy%22">Bloom's taxonomy</searchLink><br />*<searchLink fieldCode="DE" term="%22Test+design%22">Test design</searchLink><br /><searchLink fieldCode="DE" term="%22Scientific+observation%22">Scientific observation</searchLink><br /><searchLink fieldCode="DE" term="%22Fisher+exact+test%22">Fisher exact test</searchLink><br /><searchLink fieldCode="DE" term="%22Descriptive+statistics%22">Descriptive statistics</searchLink><br /><searchLink fieldCode="DE" term="%22Thematic+analysis%22">Thematic analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Professional+licenses%22">Professional licenses</searchLink><br /><searchLink fieldCode="DE" term="%22Data+analysis+software%22">Data analysis software</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Purpose: Students are increasingly relying on artificial intelligence (AI) for medical education and exam preparation. However, the factual accuracy and content distribution of AI-generated exam questions for self-assessment have not been systematically investigated. Methods: Curated prompts were created to generate multiple-choice questions matching the USMLE Step 1 examination style. We utilized ChatGPT-3.5 to generate 50 questions and answers based upon each prompt style. We manually examined output for factual accuracy, Bloom's Taxonomy, and category within the USMLE Step 1 content outline. Results: ChatGPT-3.5 generated 150 multiple-choice case-style questions and selected an answer. Overall, 83% of generated multiple questions had no factual inaccuracies and 15% contained one to two factual inaccuracies. With simple prompting, common themes included deep venous thrombosis, myocardial infarction, and thyroid disease. Topic diversity improved by separating content topic generation from question generation, and specificity to Step 1 increased by indicating that "treatment" questions were not desired. Conclusion: We demonstrate that ChatGPT-3.5 can successfully generate Step 1 style questions with reasonable factual accuracy, and this method may be used by medical students preparing for USMLE examinations. While AI-generated questions demonstrated adequate factual accuracy, targeted prompting techniques should be used to overcome ChatGPT's bias towards particular medical conditions. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Medical Teacher is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=ehh&AN=188100637 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1080/0142159X.2025.2478872 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 9 StartPage: 1645 Subjects: – SubjectFull: Generative artificial intelligence Type: general – SubjectFull: Self-evaluation Type: general – SubjectFull: Medical education Type: general – SubjectFull: Educational tests & measurements Type: general – SubjectFull: Bloom's taxonomy Type: general – SubjectFull: Test design Type: general – SubjectFull: Scientific observation Type: general – SubjectFull: Fisher exact test Type: general – SubjectFull: Descriptive statistics Type: general – SubjectFull: Thematic analysis Type: general – SubjectFull: Professional licenses Type: general – SubjectFull: Data analysis software Type: general Titles: – TitleFull: Evaluating the value of AI-generated questions for USMLE step 1 preparation: A study using ChatGPT-3.5. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Balu, Alan – PersonEntity: Name: NameFull: Prvulovic, Stefan T. – PersonEntity: Name: NameFull: Fernandez Perez, Claudia – PersonEntity: Name: NameFull: Kim, Alexander – PersonEntity: Name: NameFull: Donoho, Daniel A. – PersonEntity: Name: NameFull: Keating, Gregory IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 10 Text: Oct2025 Type: published Y: 2025 Identifiers: – Type: issn-print Value: 0142159X Numbering: – Type: volume Value: 47 – Type: issue Value: 10 Titles: – TitleFull: Medical Teacher Type: main |
| ResultId | 1 |