Reproducibility of methodological radiomics score (METRICS): an intra- and inter-rater reliability study endorsed by EuSoMII.

Saved in:
Bibliographic Details
Title: Reproducibility of methodological radiomics score (METRICS): an intra- and inter-rater reliability study endorsed by EuSoMII.
Authors: Akinci D'Antonoli, Tugba1 (AUTHOR) tugba.akincidantonoli@unibas.ch, Cavallo, Armando Ugo2 (AUTHOR), Kocak, Burak3 (AUTHOR), Borgheresi, Alessandra4,5 (AUTHOR), Ponsiglione, Andrea6 (AUTHOR), Stanzione, Arnaldo6 (AUTHOR), Koltsakis, Emmanouil7,8 (AUTHOR), Doniselli, Fabio Martino9 (AUTHOR), Vernuccio, Federica10 (AUTHOR), Ugga, Lorenzo6 (AUTHOR), Triantafyllou, Matthaios11 (AUTHOR), Huisman, Merel12 (AUTHOR), Klontzas, Michail E.8,13,14 (AUTHOR), Trotta, Romina15 (AUTHOR), Cannella, Roberto10 (AUTHOR), Fanni, Salvatore Claudio16 (AUTHOR), Cuocolo, Renato17 (AUTHOR)
Source: European Radiology. Aug2025, Vol. 35 Issue 8, p4533-4545. 13p.
Subjects: Radiomics, Inter-observer reliability, Statistical reliability, Standards, Reproducible research
Abstract: Objectives: To investigate the intra- and inter-rater reliability of the total methodological radiomics score (METRICS) and its items through a multi-reader analysis. Materials and methods: A total of 12 raters with different backgrounds and experience levels were recruited for the study. Based on their level of expertise, raters were randomly assigned to the following groups: two inter-rater reliability groups, and two intra-rater reliability groups, where each group included one group with and one group without a preliminary training session on the use of METRICS. Inter-rater reliability groups assessed all 34 papers, while intra-rater reliability groups completed the assessment of 17 papers twice within 21 days each time, and a "wash out" period of 60 days in between. Results: Inter-rater reliability was poor to moderate between raters of group 1 (without training; ICC = 0.393; 95% CI = 0.115–0.630; p = 0.002), and between raters of group 2 (with training; ICC = 0.433; 95% CI = 0.127–0.671; p = 0.002). The intra-rater analysis was excellent for raters 9 and 12, good to excellent for raters 8 and 10, moderate to excellent for rater 7, and poor to good for rater 11. Conclusion: The intra-rater reliability of the METRICS score was relatively good, while the inter-rater reliability was relatively low. This highlights the need for further efforts to achieve a common understanding of METRICS items, as well as resources consisting of explanations, elaborations, and examples to improve reproducibility and enhance their usability and robustness. Key Points: QuestionsGuidelines and scoring tools are necessary to improve the quality of radiomics research; however, the application of these tools is challenging for less experienced raters. FindingsIntra-rater reliability was high across all raters regardless of experience level or previous training, and inter-rater reliability was generally poor to moderate across raters. Clinical relevanceGuidelines and scoring tools are necessary for proper reporting in radiomics research and for closing the gap between research and clinical implementation. There is a need for further resources offering explanations, elaborations, and examples to enhance the usability and robustness of these guidelines. [ABSTRACT FROM AUTHOR]
Copyright of European Radiology is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
Description
Abstract:Objectives: To investigate the intra- and inter-rater reliability of the total methodological radiomics score (METRICS) and its items through a multi-reader analysis. Materials and methods: A total of 12 raters with different backgrounds and experience levels were recruited for the study. Based on their level of expertise, raters were randomly assigned to the following groups: two inter-rater reliability groups, and two intra-rater reliability groups, where each group included one group with and one group without a preliminary training session on the use of METRICS. Inter-rater reliability groups assessed all 34 papers, while intra-rater reliability groups completed the assessment of 17 papers twice within 21 days each time, and a "wash out" period of 60 days in between. Results: Inter-rater reliability was poor to moderate between raters of group 1 (without training; ICC = 0.393; 95% CI = 0.115–0.630; p = 0.002), and between raters of group 2 (with training; ICC = 0.433; 95% CI = 0.127–0.671; p = 0.002). The intra-rater analysis was excellent for raters 9 and 12, good to excellent for raters 8 and 10, moderate to excellent for rater 7, and poor to good for rater 11. Conclusion: The intra-rater reliability of the METRICS score was relatively good, while the inter-rater reliability was relatively low. This highlights the need for further efforts to achieve a common understanding of METRICS items, as well as resources consisting of explanations, elaborations, and examples to improve reproducibility and enhance their usability and robustness. Key Points: QuestionsGuidelines and scoring tools are necessary to improve the quality of radiomics research; however, the application of these tools is challenging for less experienced raters. FindingsIntra-rater reliability was high across all raters regardless of experience level or previous training, and inter-rater reliability was generally poor to moderate across raters. Clinical relevanceGuidelines and scoring tools are necessary for proper reporting in radiomics research and for closing the gap between research and clinical implementation. There is a need for further resources offering explanations, elaborations, and examples to enhance the usability and robustness of these guidelines. [ABSTRACT FROM AUTHOR]
ISSN:09387994
DOI:10.1007/s00330-025-11443-1