How Does ChatGPT Perform on the United States Medical Licensing Examination? The Implications of Large Language Models for Medical Education and Knowledge Assessment.

Saved in:
Bibliographic Details
Title: How Does ChatGPT Perform on the United States Medical Licensing Examination? The Implications of Large Language Models for Medical Education and Knowledge Assessment.
Authors: Gilson, Aidan1,2, Safranek, Conrad W.1, Huang, Thomas2, Socrates, Vimig1,3, Ling Chi1, Taylor, Richard Andrew1,2, Chartash, David1,4 david.chartash@yale.edu
Source: JMIR Medical Education. 2023, Vol. 9, p183-191. 9p.
Subject Terms: *Language & languages, *Medical education, *Data analysis, ChatGPT
Company/Entity: National Board of Medical Examiners
Abstract: Background: Chat Generative Pre-trained Transformer (ChatGPT) is a 175-billion-parameter natural language processing model that can generate conversation-style responses to user input. Objective: This study aimed to evaluate the performance of ChatGPT on questions within the scope of the United States Medical Licensing Examination Step 1 and Step 2 exams, as well as to analyze responses for user interpretability. Methods: We used 2 sets of multiple-choice questions to evaluate ChatGPT's performance, each with questions pertaining to Step 1 and Step 2. The first set was derived from AMBOSS, a commonly used question bank for medical students, which also provides statistics on question difficulty and the performance on an exam relative to the user base. The second set was the National Board of Medical Examiners (NBME) free 120 questions. ChatGPT's performance was compared to 2 other large language models, GPT-3 and InstructGPT. The text output of each ChatGPT response was evaluated across 3 qualitative metrics: logical justification of the answer selected, presence of information internal to the question, and presence of information external to the question. Results: Of the 4 data sets, AMBOSS-Step1, AMBOSS-Step2, NBME-Free-Step1, and NBME-Free-Step2, ChatGPT achieved accuracies of 44% (44/100), 42% (42/100), 64.4% (56/87), and 57.8% (59/102), respectively. ChatGPT outperformed InstructGPT by 8.15% on average across all data sets, and GPT-3 performed similarly to random chance. The model demonstrated a significant decrease in performance as question difficulty increased (P=.01) within the AMBOSS-Step1 data set. We found that logical justification for ChatGPT's answer selection was present in 100% of outputs of the NBME data sets. Internal information to the question was present in 96.8% (183/189) of all questions. The presence of information external to the question was 44.5% and 27% lower for incorrect answers relative to correct answers on the NBME-Free-Step1 (P<.001) and NBME-Free-Step2 (P=.001) data sets, respectively. Conclusions: ChatGPT marks a significant improvement in natural language processing models on the tasks of medical question answering. By performing at a greater than 60% threshold on the NBME-Free-Step-1 data set, we show that the model achieves the equivalent of a passing score for a third-year medical student. Additionally, we highlight ChatGPT's capacity to provide logic and informational context across the majority of answers. These facts taken together make a compelling case for the potential applications of ChatGPT as an interactive medical education tool to support learning. [ABSTRACT FROM AUTHOR]
Copyright of JMIR Medical Education is the property of JMIR Publications Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Education Research Complete
FullText Text:
  Availability: 0
Header DbId: ehh
DbLabel: Education Research Complete
An: 162663457
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: How Does ChatGPT Perform on the United States Medical Licensing Examination? The Implications of Large Language Models for Medical Education and Knowledge Assessment.
– Name: Author
  Label: Authors
  Group: Au
  Data: &lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Gilson%2C+Aidan%22&quot;&gt;Gilson, Aidan&lt;/searchLink&gt;&lt;relatesTo&gt;1,2&lt;/relatesTo&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Safranek%2C+Conrad+W%2E%22&quot;&gt;Safranek, Conrad W.&lt;/searchLink&gt;&lt;relatesTo&gt;1&lt;/relatesTo&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Huang%2C+Thomas%22&quot;&gt;Huang, Thomas&lt;/searchLink&gt;&lt;relatesTo&gt;2&lt;/relatesTo&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Socrates%2C+Vimig%22&quot;&gt;Socrates, Vimig&lt;/searchLink&gt;&lt;relatesTo&gt;1,3&lt;/relatesTo&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Ling+Chi%22&quot;&gt;Ling Chi&lt;/searchLink&gt;&lt;relatesTo&gt;1&lt;/relatesTo&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Taylor%2C+Richard+Andrew%22&quot;&gt;Taylor, Richard Andrew&lt;/searchLink&gt;&lt;relatesTo&gt;1,2&lt;/relatesTo&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Chartash%2C+David%22&quot;&gt;Chartash, David&lt;/searchLink&gt;&lt;relatesTo&gt;1,4&lt;/relatesTo&gt;&lt;i&gt; david.chartash@yale.edu&lt;/i&gt;
– Name: TitleSource
  Label: Source
  Group: Src
  Data: &lt;searchLink fieldCode=&quot;JN&quot; term=&quot;%22JMIR+Medical+Education%22&quot;&gt;JMIR Medical Education&lt;/searchLink&gt;. 2023, Vol. 9, p183-191. 9p.
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: *&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Language+%26+languages%22&quot;&gt;Language &amp; languages&lt;/searchLink&gt;&lt;br /&gt;*&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Medical+education%22&quot;&gt;Medical education&lt;/searchLink&gt;&lt;br /&gt;*&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Data+analysis%22&quot;&gt;Data analysis&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22ChatGPT%22&quot;&gt;ChatGPT&lt;/searchLink&gt;
– Name: SubjectCompany
  Label: Company/Entity
  Group: Su
  Data: &lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22National+Board+of+Medical+Examiners%22&quot;&gt;National Board of Medical Examiners&lt;/searchLink&gt;
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Background: Chat Generative Pre-trained Transformer (ChatGPT) is a 175-billion-parameter natural language processing model that can generate conversation-style responses to user input. Objective: This study aimed to evaluate the performance of ChatGPT on questions within the scope of the United States Medical Licensing Examination Step 1 and Step 2 exams, as well as to analyze responses for user interpretability. Methods: We used 2 sets of multiple-choice questions to evaluate ChatGPT&#39;s performance, each with questions pertaining to Step 1 and Step 2. The first set was derived from AMBOSS, a commonly used question bank for medical students, which also provides statistics on question difficulty and the performance on an exam relative to the user base. The second set was the National Board of Medical Examiners (NBME) free 120 questions. ChatGPT&#39;s performance was compared to 2 other large language models, GPT-3 and InstructGPT. The text output of each ChatGPT response was evaluated across 3 qualitative metrics: logical justification of the answer selected, presence of information internal to the question, and presence of information external to the question. Results: Of the 4 data sets, AMBOSS-Step1, AMBOSS-Step2, NBME-Free-Step1, and NBME-Free-Step2, ChatGPT achieved accuracies of 44% (44/100), 42% (42/100), 64.4% (56/87), and 57.8% (59/102), respectively. ChatGPT outperformed InstructGPT by 8.15% on average across all data sets, and GPT-3 performed similarly to random chance. The model demonstrated a significant decrease in performance as question difficulty increased (P=.01) within the AMBOSS-Step1 data set. We found that logical justification for ChatGPT&#39;s answer selection was present in 100% of outputs of the NBME data sets. Internal information to the question was present in 96.8% (183/189) of all questions. The presence of information external to the question was 44.5% and 27% lower for incorrect answers relative to correct answers on the NBME-Free-Step1 (P&lt;.001) and NBME-Free-Step2 (P=.001) data sets, respectively. Conclusions: ChatGPT marks a significant improvement in natural language processing models on the tasks of medical question answering. By performing at a greater than 60% threshold on the NBME-Free-Step-1 data set, we show that the model achieves the equivalent of a passing score for a third-year medical student. Additionally, we highlight ChatGPT&#39;s capacity to provide logic and informational context across the majority of answers. These facts taken together make a compelling case for the potential applications of ChatGPT as an interactive medical education tool to support learning. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: &lt;i&gt;Copyright of JMIR Medical Education is the property of JMIR Publications Inc. and its content may not be copied or emailed to multiple sites without the copyright holder&#39;s express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.&lt;/i&gt; (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=ehh&AN=162663457
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.2196/45312
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 9
        StartPage: 183
    Subjects:
      – SubjectFull: Language & languages
        Type: general
      – SubjectFull: Medical education
        Type: general
      – SubjectFull: Data analysis
        Type: general
      – SubjectFull: ChatGPT
        Type: general
      – SubjectFull: National Board of Medical Examiners
        Type: general
    Titles:
      – TitleFull: How Does ChatGPT Perform on the United States Medical Licensing Examination? The Implications of Large Language Models for Medical Education and Knowledge Assessment.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Gilson, Aidan
      – PersonEntity:
          Name:
            NameFull: Safranek, Conrad W.
      – PersonEntity:
          Name:
            NameFull: Huang, Thomas
      – PersonEntity:
          Name:
            NameFull: Socrates, Vimig
      – PersonEntity:
          Name:
            NameFull: Ling Chi
      – PersonEntity:
          Name:
            NameFull: Taylor, Richard Andrew
      – PersonEntity:
          Name:
            NameFull: Chartash, David
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Text: 2023
              Type: published
              Y: 2023
          Identifiers:
            – Type: issn-print
              Value: 23693762
          Numbering:
            – Type: volume
              Value: 9
          Titles:
            – TitleFull: JMIR Medical Education
              Type: main
ResultId 1