T, A. D., LC, A., J, L., MM, G., CJ, M., F, B., . . . I, L. (2026). GPT-4.1 and Llama 3.3 70 fail to detect clinically relevant errors in radiology reports in zero-shot evaluation. European radiology. https://doi.org/10.1007/s00330-026-12697-z
Chicago Style (17th ed.) CitationT, Akinci D'Antonoli, et al. "GPT-4.1 and Llama 3.3 70 Fail to Detect Clinically Relevant Errors in Radiology Reports in Zero-shot Evaluation." European Radiology 2026. https://doi.org/10.1007/s00330-026-12697-z.
MLA (9th ed.) CitationT, Akinci D'Antonoli, et al. "GPT-4.1 and Llama 3.3 70 Fail to Detect Clinically Relevant Errors in Radiology Reports in Zero-shot Evaluation." European Radiology, 2026, https://doi.org/10.1007/s00330-026-12697-z.
Warning: These citations may not always be 100% accurate.