Bibliographic Details
| Title: |
Accuracy analysis of AI chatbots GPT-3.5 and GEMINI on English NCLEX-style and Spanish EU general nursing multiple choice questions: challenges and performance insights. |
| Authors: |
García-Rudolph, Alejandro1,2,3 (AUTHOR) alejandropablogarcia@gmail.com, Sanchez-Pinsach, David1,2,3 (AUTHOR), Fernandez-Mira, Caridad1,2,3 (AUTHOR), Cunyat, Sandra1,2,3 (AUTHOR), Opisso, Eloy1,2,3 (AUTHOR), Hernandez-Pena, Elena1,2,3 (AUTHOR) |
| Source: |
Teaching & Learning in Nursing. Jul2025, Vol. 20 Issue 3, pe730-e735. 6p. |
| Subject Terms: |
*Generative artificial intelligence, *Language & languages, *National Council Licensure Examination for Registered Nurses, *Nursing education, *Educational tests & measurements, *Computer assisted instruction, *Nursing students, Natural language processing, Descriptive statistics, Nursing licensure, Chatbots |
| Geographic Terms: |
United States |
| Company/Entity: |
European Union |
| Abstract: |
• The use of ChatGPT in nursing education comes with inaccuracies ("hallucinations"). • ChatGPT has been scarcely validated in standard exams on different cultural contexts. • Our results yielded GPT-3.5 and GEMINI accuracy < 70% in US and < 80% in spain. • We identified specific concepts (e.g., pregnancy) where both chatbots failed. Chat GPT produces factual inaccuracies ("hallucinations"), outputs must be rigorously checked before use in nursing education, given limited validation across diverse cultural contexts. To evaluate GPT-3.5 and Google GEMINI on publicly available NCLEX-style nursing exam questions in the US and official EU general nursing exam questions for Spanish nationals. We used publicly available U.S. National Council Licensure Examination for Registered Nurses (NCLEX-RN) and for Practical Nurses (NCLEX-PN) style questions, and official Spanish general nursing exam questions (CONVALIDATE-EU-SPAIN). Accuracy was the same for GPT-3.5 and GEMINI in NCLEX-PN (67.5%, 81/120), in NCLEX-RN was higher for GPT-3.5 (69.2%, 83/120) than for GEMINI (65.8%, 79/120). Regarding CONVALIDATE-EU-SPAIN accuracy was the same for both chatbots (76.7%, 92/120). By language, in English, GPT-3.5 performed slightly better (68.3%, 164/240) than GEMINI (66.7%, 160/240). In Spanish, both chatbots achieved the same accuracy (76.7%, 92/120). We identified specific NCLEX-PN concepts where both chatbots struggled (e.g., pregnancy). In the US, chatbots' accuracy was below 70%, and in Spain, below 80%, highlighting the need to assess them comprehensively across languages. [ABSTRACT FROM AUTHOR] |
|
Copyright of Teaching & Learning in Nursing is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) |
| Database: |
Education Research Complete |