Accuracy analysis of AI chatbots GPT-3.5 and GEMINI on English NCLEX-style and Spanish EU general nursing multiple choice questions: challenges and performance insights.

Saved in:
Bibliographic Details
Title: Accuracy analysis of AI chatbots GPT-3.5 and GEMINI on English NCLEX-style and Spanish EU general nursing multiple choice questions: challenges and performance insights.
Authors: García-Rudolph, Alejandro1,2,3 (AUTHOR) alejandropablogarcia@gmail.com, Sanchez-Pinsach, David1,2,3 (AUTHOR), Fernandez-Mira, Caridad1,2,3 (AUTHOR), Cunyat, Sandra1,2,3 (AUTHOR), Opisso, Eloy1,2,3 (AUTHOR), Hernandez-Pena, Elena1,2,3 (AUTHOR)
Source: Teaching & Learning in Nursing. Jul2025, Vol. 20 Issue 3, pe730-e735. 6p.
Subject Terms: *Generative artificial intelligence, *Language & languages, *National Council Licensure Examination for Registered Nurses, *Nursing education, *Educational tests & measurements, *Computer assisted instruction, *Nursing students, Natural language processing, Descriptive statistics, Nursing licensure, Chatbots
Geographic Terms: United States
Company/Entity: European Union
Abstract: • The use of ChatGPT in nursing education comes with inaccuracies ("hallucinations"). • ChatGPT has been scarcely validated in standard exams on different cultural contexts. • Our results yielded GPT-3.5 and GEMINI accuracy < 70% in US and < 80% in spain. • We identified specific concepts (e.g., pregnancy) where both chatbots failed. Chat GPT produces factual inaccuracies ("hallucinations"), outputs must be rigorously checked before use in nursing education, given limited validation across diverse cultural contexts. To evaluate GPT-3.5 and Google GEMINI on publicly available NCLEX-style nursing exam questions in the US and official EU general nursing exam questions for Spanish nationals. We used publicly available U.S. National Council Licensure Examination for Registered Nurses (NCLEX-RN) and for Practical Nurses (NCLEX-PN) style questions, and official Spanish general nursing exam questions (CONVALIDATE-EU-SPAIN). Accuracy was the same for GPT-3.5 and GEMINI in NCLEX-PN (67.5%, 81/120), in NCLEX-RN was higher for GPT-3.5 (69.2%, 83/120) than for GEMINI (65.8%, 79/120). Regarding CONVALIDATE-EU-SPAIN accuracy was the same for both chatbots (76.7%, 92/120). By language, in English, GPT-3.5 performed slightly better (68.3%, 164/240) than GEMINI (66.7%, 160/240). In Spanish, both chatbots achieved the same accuracy (76.7%, 92/120). We identified specific NCLEX-PN concepts where both chatbots struggled (e.g., pregnancy). In the US, chatbots' accuracy was below 70%, and in Spain, below 80%, highlighting the need to assess them comprehensively across languages. [ABSTRACT FROM AUTHOR]
Copyright of Teaching & Learning in Nursing is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Education Research Complete
FullText Text:
  Availability: 0
Header DbId: ehh
DbLabel: Education Research Complete
An: 185650910
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Accuracy analysis of AI chatbots GPT-3.5 and GEMINI on English NCLEX-style and Spanish EU general nursing multiple choice questions: challenges and performance insights.
– Name: Author
  Label: Authors
  Group: Au
  Data: &lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Garc&#237;a-Rudolph%2C+Alejandro%22&quot;&gt;Garc&#237;a-Rudolph, Alejandro&lt;/searchLink&gt;&lt;relatesTo&gt;1,2,3&lt;/relatesTo&gt; (AUTHOR)&lt;i&gt; alejandropablogarcia@gmail.com&lt;/i&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Sanchez-Pinsach%2C+David%22&quot;&gt;Sanchez-Pinsach, David&lt;/searchLink&gt;&lt;relatesTo&gt;1,2,3&lt;/relatesTo&gt; (AUTHOR)&lt;br /&gt;&lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Fernandez-Mira%2C+Caridad%22&quot;&gt;Fernandez-Mira, Caridad&lt;/searchLink&gt;&lt;relatesTo&gt;1,2,3&lt;/relatesTo&gt; (AUTHOR)&lt;br /&gt;&lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Cunyat%2C+Sandra%22&quot;&gt;Cunyat, Sandra&lt;/searchLink&gt;&lt;relatesTo&gt;1,2,3&lt;/relatesTo&gt; (AUTHOR)&lt;br /&gt;&lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Opisso%2C+Eloy%22&quot;&gt;Opisso, Eloy&lt;/searchLink&gt;&lt;relatesTo&gt;1,2,3&lt;/relatesTo&gt; (AUTHOR)&lt;br /&gt;&lt;searchLink fieldCode=&quot;AR&quot; term=&quot;%22Hernandez-Pena%2C+Elena%22&quot;&gt;Hernandez-Pena, Elena&lt;/searchLink&gt;&lt;relatesTo&gt;1,2,3&lt;/relatesTo&gt; (AUTHOR)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: &lt;searchLink fieldCode=&quot;JN&quot; term=&quot;%22Teaching+%26+Learning+in+Nursing%22&quot;&gt;Teaching &amp; Learning in Nursing&lt;/searchLink&gt;. Jul2025, Vol. 20 Issue 3, pe730-e735. 6p.
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: *&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Generative+artificial+intelligence%22&quot;&gt;Generative artificial intelligence&lt;/searchLink&gt;&lt;br /&gt;*&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Language+%26+languages%22&quot;&gt;Language &amp; languages&lt;/searchLink&gt;&lt;br /&gt;*&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22National+Council+Licensure+Examination+for+Registered+Nurses%22&quot;&gt;National Council Licensure Examination for Registered Nurses&lt;/searchLink&gt;&lt;br /&gt;*&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Nursing+education%22&quot;&gt;Nursing education&lt;/searchLink&gt;&lt;br /&gt;*&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Educational+tests+%26+measurements%22&quot;&gt;Educational tests &amp; measurements&lt;/searchLink&gt;&lt;br /&gt;*&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Computer+assisted+instruction%22&quot;&gt;Computer assisted instruction&lt;/searchLink&gt;&lt;br /&gt;*&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Nursing+students%22&quot;&gt;Nursing students&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Natural+language+processing%22&quot;&gt;Natural language processing&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Descriptive+statistics%22&quot;&gt;Descriptive statistics&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Nursing+licensure%22&quot;&gt;Nursing licensure&lt;/searchLink&gt;&lt;br /&gt;&lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22Chatbots%22&quot;&gt;Chatbots&lt;/searchLink&gt;
– Name: SubjectGeographic
  Label: Geographic Terms
  Group: Su
  Data: &lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22United+States%22&quot;&gt;United States&lt;/searchLink&gt;
– Name: SubjectCompany
  Label: Company/Entity
  Group: Su
  Data: &lt;searchLink fieldCode=&quot;DE&quot; term=&quot;%22European+Union%22&quot;&gt;European Union&lt;/searchLink&gt;
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: • The use of ChatGPT in nursing education comes with inaccuracies (&quot;hallucinations&quot;). • ChatGPT has been scarcely validated in standard exams on different cultural contexts. • Our results yielded GPT-3.5 and GEMINI accuracy &lt; 70% in US and &lt; 80% in spain. • We identified specific concepts (e.g., pregnancy) where both chatbots failed. Chat GPT produces factual inaccuracies (&quot;hallucinations&quot;), outputs must be rigorously checked before use in nursing education, given limited validation across diverse cultural contexts. To evaluate GPT-3.5 and Google GEMINI on publicly available NCLEX-style nursing exam questions in the US and official EU general nursing exam questions for Spanish nationals. We used publicly available U.S. National Council Licensure Examination for Registered Nurses (NCLEX-RN) and for Practical Nurses (NCLEX-PN) style questions, and official Spanish general nursing exam questions (CONVALIDATE-EU-SPAIN). Accuracy was the same for GPT-3.5 and GEMINI in NCLEX-PN (67.5%, 81/120), in NCLEX-RN was higher for GPT-3.5 (69.2%, 83/120) than for GEMINI (65.8%, 79/120). Regarding CONVALIDATE-EU-SPAIN accuracy was the same for both chatbots (76.7%, 92/120). By language, in English, GPT-3.5 performed slightly better (68.3%, 164/240) than GEMINI (66.7%, 160/240). In Spanish, both chatbots achieved the same accuracy (76.7%, 92/120). We identified specific NCLEX-PN concepts where both chatbots struggled (e.g., pregnancy). In the US, chatbots&#39; accuracy was below 70%, and in Spain, below 80%, highlighting the need to assess them comprehensively across languages. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: &lt;i&gt;Copyright of Teaching &amp; Learning in Nursing is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder&#39;s express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.&lt;/i&gt; (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=ehh&AN=185650910
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1016/j.teln.2025.02.013
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 6
        StartPage: e730
    Subjects:
      – SubjectFull: Generative artificial intelligence
        Type: general
      – SubjectFull: Language & languages
        Type: general
      – SubjectFull: National Council Licensure Examination for Registered Nurses
        Type: general
      – SubjectFull: Nursing education
        Type: general
      – SubjectFull: Educational tests & measurements
        Type: general
      – SubjectFull: Computer assisted instruction
        Type: general
      – SubjectFull: Nursing students
        Type: general
      – SubjectFull: Natural language processing
        Type: general
      – SubjectFull: Descriptive statistics
        Type: general
      – SubjectFull: Nursing licensure
        Type: general
      – SubjectFull: Chatbots
        Type: general
      – SubjectFull: United States
        Type: general
      – SubjectFull: European Union
        Type: general
    Titles:
      – TitleFull: Accuracy analysis of AI chatbots GPT-3.5 and GEMINI on English NCLEX-style and Spanish EU general nursing multiple choice questions: challenges and performance insights.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: García-Rudolph, Alejandro
      – PersonEntity:
          Name:
            NameFull: Sanchez-Pinsach, David
      – PersonEntity:
          Name:
            NameFull: Fernandez-Mira, Caridad
      – PersonEntity:
          Name:
            NameFull: Cunyat, Sandra
      – PersonEntity:
          Name:
            NameFull: Opisso, Eloy
      – PersonEntity:
          Name:
            NameFull: Hernandez-Pena, Elena
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 07
              Text: Jul2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 15573087
          Numbering:
            – Type: volume
              Value: 20
            – Type: issue
              Value: 3
          Titles:
            – TitleFull: Teaching & Learning in Nursing
              Type: main
ResultId 1