Enabling micro-assessments of skills in the simulated setting using temporal artificial intelligence-models.

Saved in:
Bibliographic Details
Title: Enabling micro-assessments of skills in the simulated setting using temporal artificial intelligence-models.
Authors: Bang Andersen, Iben1,2 (AUTHOR), Søndergaard Svendsen, Morten Bo3 (AUTHOR), Risgaard, Anne Line1,2 (AUTHOR) a.risgaard@rn.dk, Sander Danstrup, Christian2,4 (AUTHOR), Todsen, Tobias3,5,6 (AUTHOR), Tolsgaard, Martin G.3,6,7 (AUTHOR), Friis, Mikkel Lønborg1 (AUTHOR)
Source: Medical Teacher. Mar2026, Vol. 48 Issue 3, p415-424. 10p.
Subject Terms: *Predictive tests, *Medical education, *Academic medical centers, *Philosophy of education, *Artificial intelligence, *Teaching methods, *Entry level employees, *Educational technology, *Medical students, *Simulation methods in education, *Research methodology, *National competency-based educational tests, *Automation, *Comparative studies, *Machine learning, Medical personnel, T-test (Statistics), Medical specialties & specialists, Receiver operating characteristic curves, Probability theory, Ultrasonic imaging, Mentoring, Descriptive statistics, Diagnostic errors, Thyroid gland, Software architecture, Data analysis software, Professional competence, Video recording, Sensitivity & specificity (Statistics)
Geographic Terms: Denmark
Abstract: Background: Assessing skills in simulated settings is resource-intensive and lacks validated metrics. Advances in AI offer the potential for automated competence assessment, addressing these limitations. This study aimed to develop and validate a machine learning AI model for automated evaluation during simulation-based thyroid ultrasound (US) training. Methods: Videos from eight experts and 21 novices performing thyroid US on a simulator were analyzed. Frames were processed into sequences of 1, 10, and 50 seconds. A convolutional neural network with a pre-trained ResNet-50 base and a long short-term memory layer analyzed these sequences. The model was trained to distinguish competence levels (competent=1, not competent=0) using fourfold cross-validation, with performance metrics including precision, recall, F1 score, and accuracy. Bayesian updating and adaptive thresholding assessed performance over time. Results: The AI model effectively differentiated expert and novice US performance. The 50-second sequences achieved the highest accuracy (70%) and F1 score (0.76). Experts showed significantly longer durations above the threshold (15.71s) compared to novices (9.31s, p=.030). Conclusions: A long short-term memory-based AI model provides near real-time, automated assessments of competence in US training. Utilizing temporal video data enables detailed micro-assessments of complex procedures, which may enhance interpretability and be applied across various procedural domains. [ABSTRACT FROM AUTHOR]
Copyright of Medical Teacher is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Education Research Complete
Full text is not displayed to guests.
Be the first to leave a comment!
You must be logged in first