Assessing the Ability of a Large Language Model to Score Free-Text Medical Student Clinical Notes: Quantitative Study.

Saved in:
Bibliographic Details
Title: Assessing the Ability of a Large Language Model to Score Free-Text Medical Student Clinical Notes: Quantitative Study.
Authors: Burke, Harry B1 harry.burke@gmail.com, Hoang, Albert1, Lopreiato, Joseph O1, King, Heidi2, Hemmer, Paul1, Montgomery, Michael1, Gagarin, Viktoria1
Source: JMIR Medical Education. 2024, Vol. 10, p1-1656. 1656p.
Subject Terms: *Medical students, *Medical education, Language models, ChatGPT, Artificial intelligence in medicine
Abstract: Background: Teaching medical students the skills required to acquire, interpret, apply, and communicate clinical information is an integral part of medical education. A crucial aspect of this process involves providing students with feedback regarding the quality of their free-text clinical notes. Objective: The goal of this study was to assess the ability of ChatGPT 3.5, a large language model, to score medical students' free-text history and physical notes. Methods: This is a single-institution, retrospective study. Standardized patients learned a prespecified clinical case and, acting as the patient, interacted with medical students. Each student wrote a free-text history and physical note of their interaction. The students' notes were scored independently by the standardized patients and ChatGPT using a prespecified scoring rubric that consisted of 85 case elements. The measure of accuracy was percent correct. Results: The study population consisted of 168 first-year medical students. There was a total of 14,280 scores. The ChatGPT incorrect scoring rate was 1.0%, and the standardized patient incorrect scoring rate was 7.2%. The ChatGPT error rate was 86%, lower than the standardized patient error rate. The ChatGPT mean incorrect scoring rate of 12 (SD 11) was significantly lower than the standardized patient mean incorrect scoring rate of 85 (SD 74 ; P =.002). Conclusions: ChatGPT demonstrated a significantly lower error rate compared to standardized patients. This is the first study to assess the ability of a generative pretrained transformer (GPT) program to score medical students' standardized patient-based free-text clinical notes. It is expected that, in the near future, large language models will provide real-time feedback to practicing physicians regarding their free-text notes. GPT artificial intelligence programs represent an important advance in medical education and medical practice. [ABSTRACT FROM AUTHOR]
Copyright of JMIR Medical Education is the property of JMIR Publications Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Education Research Complete
FullText Text:
  Availability: 0
Header DbId: ehh
DbLabel: Education Research Complete
An: 182585541
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Assessing the Ability of a Large Language Model to Score Free-Text Medical Student Clinical Notes: Quantitative Study.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Burke%2C+Harry+B%22">Burke, Harry B</searchLink><relatesTo>1</relatesTo><i> harry.burke@gmail.com</i><br /><searchLink fieldCode="AR" term="%22Hoang%2C+Albert%22">Hoang, Albert</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Lopreiato%2C+Joseph+O%22">Lopreiato, Joseph O</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22King%2C+Heidi%22">King, Heidi</searchLink><relatesTo>2</relatesTo><br /><searchLink fieldCode="AR" term="%22Hemmer%2C+Paul%22">Hemmer, Paul</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Montgomery%2C+Michael%22">Montgomery, Michael</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Gagarin%2C+Viktoria%22">Gagarin, Viktoria</searchLink><relatesTo>1</relatesTo>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22JMIR+Medical+Education%22">JMIR Medical Education</searchLink>. 2024, Vol. 10, p1-1656. 1656p.
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: *<searchLink fieldCode="DE" term="%22Medical+students%22">Medical students</searchLink><br />*<searchLink fieldCode="DE" term="%22Medical+education%22">Medical education</searchLink><br /><searchLink fieldCode="DE" term="%22Language+models%22">Language models</searchLink><br /><searchLink fieldCode="DE" term="%22ChatGPT%22">ChatGPT</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+intelligence+in+medicine%22">Artificial intelligence in medicine</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Background: Teaching medical students the skills required to acquire, interpret, apply, and communicate clinical information is an integral part of medical education. A crucial aspect of this process involves providing students with feedback regarding the quality of their free-text clinical notes. Objective: The goal of this study was to assess the ability of ChatGPT 3.5, a large language model, to score medical students' free-text history and physical notes. Methods: This is a single-institution, retrospective study. Standardized patients learned a prespecified clinical case and, acting as the patient, interacted with medical students. Each student wrote a free-text history and physical note of their interaction. The students' notes were scored independently by the standardized patients and ChatGPT using a prespecified scoring rubric that consisted of 85 case elements. The measure of accuracy was percent correct. Results: The study population consisted of 168 first-year medical students. There was a total of 14,280 scores. The ChatGPT incorrect scoring rate was 1.0%, and the standardized patient incorrect scoring rate was 7.2%. The ChatGPT error rate was 86%, lower than the standardized patient error rate. The ChatGPT mean incorrect scoring rate of 12 (SD 11) was significantly lower than the standardized patient mean incorrect scoring rate of 85 (SD 74 ; P =.002). Conclusions: ChatGPT demonstrated a significantly lower error rate compared to standardized patients. This is the first study to assess the ability of a generative pretrained transformer (GPT) program to score medical students' standardized patient-based free-text clinical notes. It is expected that, in the near future, large language models will provide real-time feedback to practicing physicians regarding their free-text notes. GPT artificial intelligence programs represent an important advance in medical education and medical practice. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of JMIR Medical Education is the property of JMIR Publications Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=ehh&AN=182585541
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.2196/56342
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 1656
        StartPage: 1
    Subjects:
      – SubjectFull: Medical students
        Type: general
      – SubjectFull: Medical education
        Type: general
      – SubjectFull: Language models
        Type: general
      – SubjectFull: ChatGPT
        Type: general
      – SubjectFull: Artificial intelligence in medicine
        Type: general
    Titles:
      – TitleFull: Assessing the Ability of a Large Language Model to Score Free-Text Medical Student Clinical Notes: Quantitative Study.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Burke, Harry B
      – PersonEntity:
          Name:
            NameFull: Hoang, Albert
      – PersonEntity:
          Name:
            NameFull: Lopreiato, Joseph O
      – PersonEntity:
          Name:
            NameFull: King, Heidi
      – PersonEntity:
          Name:
            NameFull: Hemmer, Paul
      – PersonEntity:
          Name:
            NameFull: Montgomery, Michael
      – PersonEntity:
          Name:
            NameFull: Gagarin, Viktoria
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Text: 2024
              Type: published
              Y: 2024
          Identifiers:
            – Type: issn-print
              Value: 23693762
          Numbering:
            – Type: volume
              Value: 10
          Titles:
            – TitleFull: JMIR Medical Education
              Type: main
ResultId 1