Bibliographic Details
| Title: |
ChatGPT versus human authors: A comparative study of concept maps for clinical reasoning training with virtual patients. |
| Authors: |
Szydlak, Renata1 (AUTHOR) renata.szydlak@uj.edu.pl, Kiyak, Yavuz Selim2 (AUTHOR), Hege, Inga3 (AUTHOR), Górski, Stanisław4 (AUTHOR), Linglart, Lea5,6 (AUTHOR), Shchudrova, Tetiana7 (AUTHOR), Torre, Dario8 (AUTHOR), Kononowicz, Andrzej A.1 (AUTHOR) |
| Source: |
Medical Teacher. Apr2026, Vol. 48 Issue 4, p712-720. 9p. |
| Subject Terms: |
*Artificial intelligence, *Teaching aids, *Educational tests & measurements, *Decision making, *Learning, *Research bias, *Information retrieval, *Clinical competence, *Concepts, *Computer assisted instruction, *Comparative studies, *Educational attainment, *Inter-observer reliability, Medical logic, Psychology of physicians, T-test (Statistics), Research funding, Research evaluation, Natural language processing, Judgment sampling, Diagnosis, Descriptive statistics, Simulated patients, Statistics, Data analysis software, User interfaces |
| Abstract: |
Purpose: This study investigates whether ChatGPT can generate clinically accurate and pedagogically valuable maps for clinical reasoning (CR) training. The aim is to assess its potential as a tool for supporting the creation of high-quality educational resources for CR training. Materials and methods: We selected 10 diverse virtual patients (VPs) from the European iCoViP project. For each case, CR concept maps were generated by a custom ChatGPT model and compared to expert-created maps available in the CASUS VP system. The comparison encompassed structural metrics (number of concepts, connections, and graph density), clinical content quality (clinical expert evaluation of concept and connection validity), and pedagogical utility (medical educator assessment of clarity, abstraction, and progression). Statistical analysis included Student's t-tests and interrater reliability using weighted Cohen's kappa. Results: ChatGPT-generated maps contained significantly more concepts and connections than expert maps, indicating higher structural complexity (p < 0.001), though graph density did not differ significantly. Clinician evaluations showed comparable clinical content quality across both groups, with no statistically significant differences in concept or connection ratings. The educational review revealed that while ChatGPT maps offered comprehensive information, they lacked abstraction, prioritization, and contextual alignment, occasionally exceeding the optimal cognitive load for learners. Conclusions: ChatGPT can reliably generate concept maps that match expert-level clinical accuracy. However, limitations in educational clarity and usability underscore the need for expert refinement. With appropriate oversight, large language models (LLMs) such as ChatGPT can support efficient development of learning resources for CR education. [ABSTRACT FROM AUTHOR] |
|
Copyright of Medical Teacher is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) |
| Database: |
Education Research Complete |