AI's Ability to Interpret Unlabeled Anatomy Images and Supplement Educational Research as an AI Rater
Saved in:
| Title: | AI's Ability to Interpret Unlabeled Anatomy Images and Supplement Educational Research as an AI Rater |
|---|---|
| Language: | English |
| Authors: | Lord J. Hyeamang, Tejas C. Sekhar, Emily Rush, Amy C. Beresheim (ORCID |
| Source: | Anatomical Sciences Education. 2025 18(10):1102-1113. |
| Availability: | Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us |
| Peer Reviewed: | Y |
| Page Count: | 12 |
| Publication Date: | 2025 |
| Document Type: | Journal Articles Reports - Research |
| Descriptors: | Artificial Intelligence, Anatomy, Identification, Man Machine Systems, Natural Language Processing, Likert Scales, Accuracy |
| DOI: | 10.1002/ase.70074 |
| ISSN: | 1935-9772 1935-9780 |
| Abstract: | Evidence suggests custom chatbots are superior to commercial generative artificial intelligence (GenAI) systems for text-based anatomy content inquiries. This study evaluates ChatGPT-4o's and Claude 3.5 Sonnet's capabilities to interpret unlabeled anatomical images. Secondarily, ChatGPT o1-preview was evaluated as an AI rater to grade AI-generated outputs using a rubric and was compared against human raters. Anatomical images (five musculoskeletal, five thoracic) representing diverse image-based media (e.g., illustrations, photographs, MRI) were annotated with identification markers (e.g., arrows, circles) and uploaded to each GenAI system for interpretation. Forty-five prompts (i.e., 15 first-order, 15 second-order, and 15 third-order questions) with associated images were submitted to both GenAI systems across two timepoints. Responses were graded by anatomy experts for factual accuracy and superfluity (the presence of excessive wording) on a three-point Likert scale. ChatGPT o1-preview was tested for agreement against human anatomy experts to determine its usefulness as an AI rater. Statistical analyses included inter-rater agreement, hierarchical linear modeling, and test-retest reliability. ChatGPT-4o's factual accuracy score across 45 outputs was 68.0% compared to Claude 3.5 Sonnet's score of 61.5% (p = 0.319). As an AI rater, ChatGPT o1-preview showed moderate to substantial agreement with human raters (Cohen's kappa = 0.545-0.755) for evaluating factual accuracy according to a rubric of textbook answers. Further improvements and evaluations are needed before commercial GenAI systems can be used as credible student resources in anatomy education. Similarly, ChatGPT o1-preview demonstrates promise as an AI assistant for educational research, though further investigation is warranted. |
| Abstractor: | As Provided |
| Entry Date: | 2025 |
| Accession Number: | EJ1486230 |
| Database: | ERIC |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1486230 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: AI's Ability to Interpret Unlabeled Anatomy Images and Supplement Educational Research as an AI Rater – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Lord+J%2E+Hyeamang%22">Lord J. Hyeamang</searchLink><br /><searchLink fieldCode="AR" term="%22Tejas+C%2E+Sekhar%22">Tejas C. Sekhar</searchLink><br /><searchLink fieldCode="AR" term="%22Emily+Rush%22">Emily Rush</searchLink><br /><searchLink fieldCode="AR" term="%22Amy+C%2E+Beresheim%22">Amy C. Beresheim</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-1672-1131">0000-0002-1672-1131</externalLink>)<br /><searchLink fieldCode="AR" term="%22Colleen+M%2E+Cheverko%22">Colleen M. Cheverko</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-9065-1768">0000-0002-9065-1768</externalLink>)<br /><searchLink fieldCode="AR" term="%22William+S%2E+Brooks%22">William S. Brooks</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-4745-9295">0000-0002-4745-9295</externalLink>)<br /><searchLink fieldCode="AR" term="%22Abbey+C%2E+M%2E+Breckling%22">Abbey C. M. Breckling</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-4322-4744">0000-0003-4322-4744</externalLink>)<br /><searchLink fieldCode="AR" term="%22M%2E+Nazmul+Karim%22">M. Nazmul Karim</searchLink><br /><searchLink fieldCode="AR" term="%22Christopher+Ferrigno%22">Christopher Ferrigno</searchLink><br /><searchLink fieldCode="AR" term="%22Adam+B%2E+Wilson%22">Adam B. Wilson</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-1221-5602">0000-0002-1221-5602</externalLink>) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Anatomical+Sciences+Education%22"><i>Anatomical Sciences Education</i></searchLink>. 2025 18(10):1102-1113. – Name: Avail Label: Availability Group: Avail Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 12 – Name: DatePubCY Label: Publication Date Group: Date Data: 2025 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Anatomy%22">Anatomy</searchLink><br /><searchLink fieldCode="DE" term="%22Identification%22">Identification</searchLink><br /><searchLink fieldCode="DE" term="%22Man+Machine+Systems%22">Man Machine Systems</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+Language+Processing%22">Natural Language Processing</searchLink><br /><searchLink fieldCode="DE" term="%22Likert+Scales%22">Likert Scales</searchLink><br /><searchLink fieldCode="DE" term="%22Accuracy%22">Accuracy</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1002/ase.70074 – Name: ISSN Label: ISSN Group: ISSN Data: 1935-9772<br />1935-9780 – Name: Abstract Label: Abstract Group: Ab Data: Evidence suggests custom chatbots are superior to commercial generative artificial intelligence (GenAI) systems for text-based anatomy content inquiries. This study evaluates ChatGPT-4o's and Claude 3.5 Sonnet's capabilities to interpret unlabeled anatomical images. Secondarily, ChatGPT o1-preview was evaluated as an AI rater to grade AI-generated outputs using a rubric and was compared against human raters. Anatomical images (five musculoskeletal, five thoracic) representing diverse image-based media (e.g., illustrations, photographs, MRI) were annotated with identification markers (e.g., arrows, circles) and uploaded to each GenAI system for interpretation. Forty-five prompts (i.e., 15 first-order, 15 second-order, and 15 third-order questions) with associated images were submitted to both GenAI systems across two timepoints. Responses were graded by anatomy experts for factual accuracy and superfluity (the presence of excessive wording) on a three-point Likert scale. ChatGPT o1-preview was tested for agreement against human anatomy experts to determine its usefulness as an AI rater. Statistical analyses included inter-rater agreement, hierarchical linear modeling, and test-retest reliability. ChatGPT-4o's factual accuracy score across 45 outputs was 68.0% compared to Claude 3.5 Sonnet's score of 61.5% (p = 0.319). As an AI rater, ChatGPT o1-preview showed moderate to substantial agreement with human raters (Cohen's kappa = 0.545-0.755) for evaluating factual accuracy according to a rubric of textbook answers. Further improvements and evaluations are needed before commercial GenAI systems can be used as credible student resources in anatomy education. Similarly, ChatGPT o1-preview demonstrates promise as an AI assistant for educational research, though further investigation is warranted. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2025 – Name: AN Label: Accession Number Group: ID Data: EJ1486230 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1486230 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1002/ase.70074 Languages: – Text: English PhysicalDescription: Pagination: PageCount: 12 StartPage: 1102 Subjects: – SubjectFull: Artificial Intelligence Type: general – SubjectFull: Anatomy Type: general – SubjectFull: Identification Type: general – SubjectFull: Man Machine Systems Type: general – SubjectFull: Natural Language Processing Type: general – SubjectFull: Likert Scales Type: general – SubjectFull: Accuracy Type: general Titles: – TitleFull: AI's Ability to Interpret Unlabeled Anatomy Images and Supplement Educational Research as an AI Rater Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Lord J. Hyeamang – PersonEntity: Name: NameFull: Tejas C. Sekhar – PersonEntity: Name: NameFull: Emily Rush – PersonEntity: Name: NameFull: Amy C. Beresheim – PersonEntity: Name: NameFull: Colleen M. Cheverko – PersonEntity: Name: NameFull: William S. Brooks – PersonEntity: Name: NameFull: Abbey C. M. Breckling – PersonEntity: Name: NameFull: M. Nazmul Karim – PersonEntity: Name: NameFull: Christopher Ferrigno – PersonEntity: Name: NameFull: Adam B. Wilson IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 10 Type: published Y: 2025 Identifiers: – Type: issn-print Value: 1935-9772 – Type: issn-electronic Value: 1935-9780 Numbering: – Type: volume Value: 18 – Type: issue Value: 10 Titles: – TitleFull: Anatomical Sciences Education Type: main |
| ResultId | 1 |