Performance of Large Language Models in Neurology Multiple‐Choice Questions.

Saved in:
Bibliographic Details
Title: Performance of Large Language Models in Neurology Multiple‐Choice Questions.
Authors: Habibi, Gholamreza (AUTHOR), Gargari, Omid Kohandel (AUTHOR), Hosseini, Mostafa (AUTHOR), Afchangi, Kasra (AUTHOR), Saleem, Suraiya (AUTHOR)
Source: Acta Neurologica Scandinavica. 6/1/2026, Vol. 2026, p1-7. 7p.
Subjects: Generative pre-trained transformers, Neurology, Artificial intelligence in medicine, Language models, Multiple choice examinations
Abstract: Introduction: Navigating neurological disorders is complex due to overlapping symptoms and diverse diagnostic requirements. Large language models (LLMs) offer potential support for clinicians by processing vast amounts of textual data. This study evaluates the performance of four advanced LLMs—GPT‐4, GPT‐3.5, Clinical Camel, and MedALPACA—in answering neurology‐based multiple‐choice questions (MCQs). Methods: The study utilized 170 MCQs from the Comprehensive Review in Clinical Neurology book. Questions were divided into the body and answer choices, stored separately in an Excel spreadsheet. The models were prompted to select the correct answers. Generative Pretrained Transformer (GPT) models were accessed via OpenAI′s API, whereas Clinical Camel and MedALPACA were downloaded from Hugging Face. Accuracy was calculated by the number of correct answers over total questions, with subgroup analysis based on subject headings. Results: GPT‐4 achieved the highest accuracy at 84.7%, significantly outperforming GPT‐3.5 (58.8%), Clinical Camel (52.9%), and MedALPACA (40%). GPT‐4′s performance was significantly better (p < 0.001). The difference between GPT‐3.5 and Clinical Camel was not significant (p = 0.27), but both outperformed MedALPACA. Response similarity was highest between GPT‐4 and GPT‐3.5 (64.1%) and lowest between GPT‐4 and MedALPACA (41.2%). In subgroup analysis, GPT‐4 was superior across all topics, achieving full scores in six topics and its lowest in vascular neurology (40%). GPT‐3.5 performed best in eight topics, Clinical Camel in three topics, and MedALPACA was the weakest in all but two topics. Conclusion: GPT‐4 demonstrated the highest accuracy in answering neurology MCQs, outperforming GPT‐3.5, Clinical Camel, and MedALPACA. Clinical Camel′s comparable performance to GPT‐3.5 highlights the potential of specialized medical models. External evaluation datasets remain essential to avoid data leakage and ensure fair benchmarking. Further research is needed to expand topic coverage, assess reasoning processes, and include human comparison to support safe and effective clinical integration of medical LLMs. [ABSTRACT FROM AUTHOR]
Copyright of Acta Neurologica Scandinavica is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Psychology and Behavioral Sciences Collection
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: pbh
DbLabel: Psychology and Behavioral Sciences Collection
An: 194204550
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Performance of Large Language Models in Neurology Multiple‐Choice Questions.
– Name: Author
  Label: Authors
  Group: Au
  Data: &lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Habibi%2C+Gholamreza%22&quot;&gt;Habibi, Gholamreza&lt;/searchLink&gt; (AUTHOR)&lt;br /&gt;&lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Gargari%2C+Omid+Kohandel%22&quot;&gt;Gargari, Omid Kohandel&lt;/searchLink&gt; (AUTHOR)&lt;br /&gt;&lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Hosseini%2C+Mostafa%22&quot;&gt;Hosseini, Mostafa&lt;/searchLink&gt; (AUTHOR)&lt;br /&gt;&lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Afchangi%2C+Kasra%22&quot;&gt;Afchangi, Kasra&lt;/searchLink&gt; (AUTHOR)&lt;br /&gt;&lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Saleem%2C+Suraiya%22&quot;&gt;Saleem, Suraiya&lt;/searchLink&gt; (AUTHOR)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: &lt;searchLink fieldCode=&quot;JN&quot; term=&quot;%22Acta+Neurologica+Scandinavica%22&quot;&gt;Acta Neurologica Scandinavica&lt;/searchLink&gt;. 6/1/2026, Vol. 2026, p1-7. 7p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: &lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Generative+pre-trained+transformers%22&quot;&gt;Generative pre-trained transformers&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Neurology%22&quot;&gt;Neurology&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Artificial+intelligence+in+medicine%22&quot;&gt;Artificial intelligence in medicine&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Language+models%22&quot;&gt;Language models&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Multiple+choice+examinations%22&quot;&gt;Multiple choice examinations&lt;/searchLink&gt;
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Introduction: Navigating neurological disorders is complex due to overlapping symptoms and diverse diagnostic requirements. Large language models (LLMs) offer potential support for clinicians by processing vast amounts of textual data. This study evaluates the performance of four advanced LLMs—GPT‐4, GPT‐3.5, Clinical Camel, and MedALPACA—in answering neurology‐based multiple‐choice questions (MCQs). Methods: The study utilized 170 MCQs from the Comprehensive Review in Clinical Neurology book. Questions were divided into the body and answer choices, stored separately in an Excel spreadsheet. The models were prompted to select the correct answers. Generative Pretrained Transformer (GPT) models were accessed via OpenAI′s API, whereas Clinical Camel and MedALPACA were downloaded from Hugging Face. Accuracy was calculated by the number of correct answers over total questions, with subgroup analysis based on subject headings. Results: GPT‐4 achieved the highest accuracy at 84.7%, significantly outperforming GPT‐3.5 (58.8%), Clinical Camel (52.9%), and MedALPACA (40%). GPT‐4′s performance was significantly better (p &lt; 0.001). The difference between GPT‐3.5 and Clinical Camel was not significant (p = 0.27), but both outperformed MedALPACA. Response similarity was highest between GPT‐4 and GPT‐3.5 (64.1%) and lowest between GPT‐4 and MedALPACA (41.2%). In subgroup analysis, GPT‐4 was superior across all topics, achieving full scores in six topics and its lowest in vascular neurology (40%). GPT‐3.5 performed best in eight topics, Clinical Camel in three topics, and MedALPACA was the weakest in all but two topics. Conclusion: GPT‐4 demonstrated the highest accuracy in answering neurology MCQs, outperforming GPT‐3.5, Clinical Camel, and MedALPACA. Clinical Camel′s comparable performance to GPT‐3.5 highlights the potential of specialized medical models. External evaluation datasets remain essential to avoid data leakage and ensure fair benchmarking. Further research is needed to expand topic coverage, assess reasoning processes, and include human comparison to support safe and effective clinical integration of medical LLMs. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: &lt;i&gt;Copyright of Acta Neurologica Scandinavica is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder&#39;s express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.&lt;/i&gt; (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=pbh&AN=194204550
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1155/ane/5623086
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 7
        StartPage: 1
    Subjects:
      – SubjectFull: Generative pre-trained transformers
        Type: general
      – SubjectFull: Neurology
        Type: general
      – SubjectFull: Artificial intelligence in medicine
        Type: general
      – SubjectFull: Language models
        Type: general
      – SubjectFull: Multiple choice examinations
        Type: general
    Titles:
      – TitleFull: Performance of Large Language Models in Neurology Multiple‐Choice Questions.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Habibi, Gholamreza
      – PersonEntity:
          Name:
            NameFull: Gargari, Omid Kohandel
      – PersonEntity:
          Name:
            NameFull: Hosseini, Mostafa
      – PersonEntity:
          Name:
            NameFull: Afchangi, Kasra
      – PersonEntity:
          Name:
            NameFull: Saleem, Suraiya
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 06
              Text: 6/1/2026
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 00016314
          Numbering:
            – Type: volume
              Value: 2026
          Titles:
            – TitleFull: Acta Neurologica Scandinavica
              Type: main
ResultId 1