Performance of a large language model on the reasoning tasks of a physician.

Saved in:
Bibliographic Details
Title: Performance of a large language model on the reasoning tasks of a physician.
Authors: Brodeur, Peter G. (AUTHOR), Buckley, Thomas A. (AUTHOR), Kanjee, Zahir (AUTHOR), Goh, Ethan (AUTHOR), Ling, Evelyn Bin (AUTHOR), Jain, Priyank (AUTHOR), Cabral, Stephanie (AUTHOR), Abdulnour, Raja-Elie (AUTHOR), Haimovich, Adrian D. (AUTHOR), Freed, Jason A. (AUTHOR), Olson, Andrew (AUTHOR), Morgan, Daniel J. (AUTHOR), Hom, Jason (AUTHOR), Gallo, Robert (AUTHOR), McCoy, Liam G. (AUTHOR), Mombini, Haadi (AUTHOR), Lucas, Christopher (AUTHOR), Fotoohi, Misha (AUTHOR), Gwiazdon, Matthew (AUTHOR), Restifo, Daniele (AUTHOR)
Source: Science. 4/30/2026, Vol. 392 Issue 6797, p524-527. 4p.
Subjects: Artificial intelligence, Clinical decision support systems, Clinical medicine, Medical logic, Diagnosis, Hospital emergency services, Language models
Abstract: More than 65 years ago, complex clinical diagnostic reasoning cases were introduced as the gold standard for the evaluation of expert medical computing systems, a standard that has held ever since. In this study, we report the results of a physician evaluation of a large language model (LLM) on challenging clinical cases across five experiments with a baseline of hundreds of physicians. We then report a real-world study comparing human expert and artificial intelligence (AI) second opinions in randomly selected patients in the emergency room of a major tertiary academic medical center. In all experiments, the LLM outperformed physician baselines and displayed continued improvement from prior generations of AI clinical decision support. Our study suggests that LLMs have eclipsed most benchmarks of clinical reasoning, motivating the urgent need for prospective trials. Editor's summary: Computational tools for medical decision support have been advancing over time, mainly by serving as resources for limited applications. Machine learning tools for autonomous interpretation of clinical cases have also been gradually improving over time. Brodeur et al. pitted a large language model, the OpenAI o1 series, directly against hundreds of physicians at different levels of training and experience on a variety of clinical cases ranging from published patient vignettes to evaluations of brand-new emergency room patients, as well as on clinical tasks including both diagnosis and planning of clinical management (see the Perspective by Hopkins and Cornelisse). Across a variety of scenarios and applications, the large language model outperformed both human physicians and older models, suggesting its potential utility for clinical care. —Yevgeniya Nusinovich [ABSTRACT FROM AUTHOR]
Copyright of Science is the property of American Association for the Advancement of Science and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Psychology and Behavioral Sciences Collection
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: pbh
DbLabel: Psychology and Behavioral Sciences Collection
An: 193402121
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Performance of a large language model on the reasoning tasks of a physician.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Brodeur%2C+Peter+G%2E%22">Brodeur, Peter G.</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Buckley%2C+Thomas+A%2E%22">Buckley, Thomas A.</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Kanjee%2C+Zahir%22">Kanjee, Zahir</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Goh%2C+Ethan%22">Goh, Ethan</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Ling%2C+Evelyn+Bin%22">Ling, Evelyn Bin</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Jain%2C+Priyank%22">Jain, Priyank</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Cabral%2C+Stephanie%22">Cabral, Stephanie</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Abdulnour%2C+Raja-Elie%22">Abdulnour, Raja-Elie</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Haimovich%2C+Adrian+D%2E%22">Haimovich, Adrian D.</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Freed%2C+Jason+A%2E%22">Freed, Jason A.</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Olson%2C+Andrew%22">Olson, Andrew</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Morgan%2C+Daniel+J%2E%22">Morgan, Daniel J.</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Hom%2C+Jason%22">Hom, Jason</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Gallo%2C+Robert%22">Gallo, Robert</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22McCoy%2C+Liam+G%2E%22">McCoy, Liam G.</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Mombini%2C+Haadi%22">Mombini, Haadi</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Lucas%2C+Christopher%22">Lucas, Christopher</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Fotoohi%2C+Misha%22">Fotoohi, Misha</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Gwiazdon%2C+Matthew%22">Gwiazdon, Matthew</searchLink> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Restifo%2C+Daniele%22">Restifo, Daniele</searchLink> (AUTHOR)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Science%22">Science</searchLink>. 4/30/2026, Vol. 392 Issue 6797, p524-527. 4p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Artificial+intelligence%22">Artificial intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Clinical+decision+support+systems%22">Clinical decision support systems</searchLink><br /><searchLink fieldCode="DE" term="%22Clinical+medicine%22">Clinical medicine</searchLink><br /><searchLink fieldCode="DE" term="%22Medical+logic%22">Medical logic</searchLink><br /><searchLink fieldCode="DE" term="%22Diagnosis%22">Diagnosis</searchLink><br /><searchLink fieldCode="DE" term="%22Hospital+emergency+services%22">Hospital emergency services</searchLink><br /><searchLink fieldCode="DE" term="%22Language+models%22">Language models</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: More than 65 years ago, complex clinical diagnostic reasoning cases were introduced as the gold standard for the evaluation of expert medical computing systems, a standard that has held ever since. In this study, we report the results of a physician evaluation of a large language model (LLM) on challenging clinical cases across five experiments with a baseline of hundreds of physicians. We then report a real-world study comparing human expert and artificial intelligence (AI) second opinions in randomly selected patients in the emergency room of a major tertiary academic medical center. In all experiments, the LLM outperformed physician baselines and displayed continued improvement from prior generations of AI clinical decision support. Our study suggests that LLMs have eclipsed most benchmarks of clinical reasoning, motivating the urgent need for prospective trials. Editor's summary: Computational tools for medical decision support have been advancing over time, mainly by serving as resources for limited applications. Machine learning tools for autonomous interpretation of clinical cases have also been gradually improving over time. Brodeur et al. pitted a large language model, the OpenAI o1 series, directly against hundreds of physicians at different levels of training and experience on a variety of clinical cases ranging from published patient vignettes to evaluations of brand-new emergency room patients, as well as on clinical tasks including both diagnosis and planning of clinical management (see the Perspective by Hopkins and Cornelisse). Across a variety of scenarios and applications, the large language model outperformed both human physicians and older models, suggesting its potential utility for clinical care. —Yevgeniya Nusinovich [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Science is the property of American Association for the Advancement of Science and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=pbh&AN=193402121
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1126/science.adz4433
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 4
        StartPage: 524
    Subjects:
      – SubjectFull: Artificial intelligence
        Type: general
      – SubjectFull: Clinical decision support systems
        Type: general
      – SubjectFull: Clinical medicine
        Type: general
      – SubjectFull: Medical logic
        Type: general
      – SubjectFull: Diagnosis
        Type: general
      – SubjectFull: Hospital emergency services
        Type: general
      – SubjectFull: Language models
        Type: general
    Titles:
      – TitleFull: Performance of a large language model on the reasoning tasks of a physician.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Brodeur, Peter G.
      – PersonEntity:
          Name:
            NameFull: Buckley, Thomas A.
      – PersonEntity:
          Name:
            NameFull: Kanjee, Zahir
      – PersonEntity:
          Name:
            NameFull: Goh, Ethan
      – PersonEntity:
          Name:
            NameFull: Ling, Evelyn Bin
      – PersonEntity:
          Name:
            NameFull: Jain, Priyank
      – PersonEntity:
          Name:
            NameFull: Cabral, Stephanie
      – PersonEntity:
          Name:
            NameFull: Abdulnour, Raja-Elie
      – PersonEntity:
          Name:
            NameFull: Haimovich, Adrian D.
      – PersonEntity:
          Name:
            NameFull: Freed, Jason A.
      – PersonEntity:
          Name:
            NameFull: Olson, Andrew
      – PersonEntity:
          Name:
            NameFull: Morgan, Daniel J.
      – PersonEntity:
          Name:
            NameFull: Hom, Jason
      – PersonEntity:
          Name:
            NameFull: Gallo, Robert
      – PersonEntity:
          Name:
            NameFull: McCoy, Liam G.
      – PersonEntity:
          Name:
            NameFull: Mombini, Haadi
      – PersonEntity:
          Name:
            NameFull: Lucas, Christopher
      – PersonEntity:
          Name:
            NameFull: Fotoohi, Misha
      – PersonEntity:
          Name:
            NameFull: Gwiazdon, Matthew
      – PersonEntity:
          Name:
            NameFull: Restifo, Daniele
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 30
              M: 04
              Text: 4/30/2026
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 00368075
          Numbering:
            – Type: volume
              Value: 392
            – Type: issue
              Value: 6797
          Titles:
            – TitleFull: Science
              Type: main
ResultId 1