Conducting Medication Reviews: A Comparative Study Between ChatGPT‐4 and Healthcare Professionals.

Saved in:
Bibliographic Details
Title: Conducting Medication Reviews: A Comparative Study Between ChatGPT‐4 and Healthcare Professionals.
Authors: ten Hoope, Simone M. K., Marongiu, Sabrina, Siegert, Carl E. H., Heerdink, Eibert R., Janssen, Marjo J. A., Karapinar‐Çarkit, Fatma
Source: Journal of the American Geriatrics Society. May2026, Vol. 74 Issue 5, p1378-1385. 8p.
Subjects: Generative artificial intelligence, Academic medical centers, Medication error prevention, Patient care, Polypharmacy, Natural language processing, Retrospective studies, Descriptive statistics, Research methodology, Medical records, Acquisition of data, Statistics, Comparative studies, Hospital care of older people, Confidence intervals, Data analysis software
Geographic Terms: Netherlands
Abstract: Background: The increasing prevalence of patients with hyperpolypharmacy (> 10 medications) has made medication reviews increasingly complex. ChatGPT‐4‐Turbo (ChatGPT), a large language model, has demonstrated potential in healthcare applications and could potentially support medication reviews. Objectives: This study aimed to evaluate the agreement between medication reviews conducted by ChatGPT compared to healthcare professionals (HCPs) in older people. Secondary objectives included: the validity of additional interventions detected by ChatGPT, its ability to structure diagnoses to medication use, and laboratory target values based on patient characteristics. Methods: In this retrospective proof‐of‐concept study, 51 medication reviews previously conducted by a geriatric internist and hospital pharmacist were re‐evaluated using ChatGPT. ChatGPT was trained on polypharmacy guidelines, was then provided with the same primary data as HCPs, and was asked to perform medication reviews. Two pharmacists scored the agreement between ChatGPT and HCPs. The structuring of information and additional interventions suggested by ChatGPT were reviewed within an expert team. Descriptive statistics were used. Outcomes: The primary outcome was the percentage agreement between the interventions suggested by ChatGPT compared to HCPs. Secondary outcomes included the proportion of valid and incorrect interventions suggested by ChatGPT and its ability to structure patient information. Results: HCPs suggested 183 interventions and ChatGPT 202 interventions. ChatGPT achieved a 27.7% agreement with interventions of HCPs. It identified 19 additional valid interventions which HCPs missed (7.6%), but also proposed 84 incorrect interventions (33.7%). While ChatGPT demonstrated strong capability in structuring patient data (86.4% correct diagnoses linked to medication), it struggled with contextualizing appropriate laboratory target values based on patient characteristics (46.7%). Conclusion: ChatGPT had low agreement with HCPs, but found additional interventions that HCPs missed. ChatGPT lacks clinical decision‐making capabilities based on individual patient contexts in older people. ChatGPT may, however, serve as a support tool to structure diagnoses and medication lists. Summary: Key points ○ChatGPT‐4‐Turbo shows 27.7% agreement with interventions suggested by healthcare professionals during medication reviews for older people. It fails to take patient context into consideration.○ChatGPT‐4‐Turbo is able to provide 7.6% new valid interventions that healthcare professionals missed.○ChatGPT‐4‐Turbo can structure information and match diagnosis with medications accordingly (86.4%), but lacks the ability to contextualize appropriate laboratory target values based on patient characteristics (46.7%).Why does this paper matter? ○As the integration of artificial intelligence into healthcare continues to gain momentum, this proof‐of‐concept study highlights the potentials and limitations of large language models such as ChatGPT‐4.○While the model shows promise in organizing clinical information and supporting diagnosis‐medication alignment and finds additional valid interventions missed by healthcare professionals; it remains inadequate for direct use in clinical medication reviews due to its inability to account for patient‐specific contextual factors in older people. [ABSTRACT FROM AUTHOR]
Copyright of Journal of the American Geriatrics Society is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Psychology and Behavioral Sciences Collection
Description
Abstract:Background: The increasing prevalence of patients with hyperpolypharmacy (> 10 medications) has made medication reviews increasingly complex. ChatGPT‐4‐Turbo (ChatGPT), a large language model, has demonstrated potential in healthcare applications and could potentially support medication reviews. Objectives: This study aimed to evaluate the agreement between medication reviews conducted by ChatGPT compared to healthcare professionals (HCPs) in older people. Secondary objectives included: the validity of additional interventions detected by ChatGPT, its ability to structure diagnoses to medication use, and laboratory target values based on patient characteristics. Methods: In this retrospective proof‐of‐concept study, 51 medication reviews previously conducted by a geriatric internist and hospital pharmacist were re‐evaluated using ChatGPT. ChatGPT was trained on polypharmacy guidelines, was then provided with the same primary data as HCPs, and was asked to perform medication reviews. Two pharmacists scored the agreement between ChatGPT and HCPs. The structuring of information and additional interventions suggested by ChatGPT were reviewed within an expert team. Descriptive statistics were used. Outcomes: The primary outcome was the percentage agreement between the interventions suggested by ChatGPT compared to HCPs. Secondary outcomes included the proportion of valid and incorrect interventions suggested by ChatGPT and its ability to structure patient information. Results: HCPs suggested 183 interventions and ChatGPT 202 interventions. ChatGPT achieved a 27.7% agreement with interventions of HCPs. It identified 19 additional valid interventions which HCPs missed (7.6%), but also proposed 84 incorrect interventions (33.7%). While ChatGPT demonstrated strong capability in structuring patient data (86.4% correct diagnoses linked to medication), it struggled with contextualizing appropriate laboratory target values based on patient characteristics (46.7%). Conclusion: ChatGPT had low agreement with HCPs, but found additional interventions that HCPs missed. ChatGPT lacks clinical decision‐making capabilities based on individual patient contexts in older people. ChatGPT may, however, serve as a support tool to structure diagnoses and medication lists. Summary: Key points ○ChatGPT‐4‐Turbo shows 27.7% agreement with interventions suggested by healthcare professionals during medication reviews for older people. It fails to take patient context into consideration.○ChatGPT‐4‐Turbo is able to provide 7.6% new valid interventions that healthcare professionals missed.○ChatGPT‐4‐Turbo can structure information and match diagnosis with medications accordingly (86.4%), but lacks the ability to contextualize appropriate laboratory target values based on patient characteristics (46.7%).Why does this paper matter? ○As the integration of artificial intelligence into healthcare continues to gain momentum, this proof‐of‐concept study highlights the potentials and limitations of large language models such as ChatGPT‐4.○While the model shows promise in organizing clinical information and supporting diagnosis‐medication alignment and finds additional valid interventions missed by healthcare professionals; it remains inadequate for direct use in clinical medication reviews due to its inability to account for patient‐specific contextual factors in older people. [ABSTRACT FROM AUTHOR]
ISSN:00028614
DOI:10.1111/jgs.70415