Large language models for generating script concordance test in obstetrics and gynecology: ChatGPT and Claude.

Saved in:
Bibliographic Details
Title: Large language models for generating script concordance test in obstetrics and gynecology: ChatGPT and Claude.
Authors: Yapıcı Coşkun, Zuhal1,2 (AUTHOR) zuhalyapici@yahoo.com, Kıyak, Yavuz Selim3 (AUTHOR), Coşkun, Özlem3 (AUTHOR), Budakoğlu, Işıl İrem3 (AUTHOR), Özdemir, Özhan4 (AUTHOR)
Source: Medical Teacher. Nov2025, Vol. 47 Issue 11, p1767-1771. 5p.
Subject Terms: *Generative artificial intelligence, *Medical education, *Undergraduates, *Educational tests & measurements, *Medical students, Medical logic, Cross-sectional method, Scale analysis (Psychology), Primary health care, Descriptive statistics, Gynecology, Chatbots, Obstetrics
Abstract: Objective: To evaluate the performance of large language models (ChatGPT-4o and Claude 3.5 Sonnet) to generate script concordance test (SCT) items for assessing clinical reasoning in obstetrics and gynecology. Methods: This cross-sectional study involved the generation of SCT items for five common diagnostic topics in obstetrics and gynecology in primary care settings. A total of 16 panelists evaluated the AI-generated SCT items against 11 predefined criteria. Descriptive statistics were used to compare the models' performance across criteria. Results: ChatGPT-4o had an overall agreement rate of 90.57% for SCT items meeting the quality criteria, while Claude 3.5 Sonnet achieved 91.48%. The criterion with the lowest scores was "The scenario is of appropriate difficulty for medical students," with ChatGPT-4o rated at 71.25% and Claude 3.5 Sonnet at 76.25%. Conclusion: Large language models can generate SCT items that effectively assess clinical reasoning; however, further refinement is required to ensure the appropriate level of difficulty for medical students. These findings highlight the potential of AI to enhance the efficiency of SCT generation in obstetrics and gynecology within primary care settings. [ABSTRACT FROM AUTHOR]
Copyright of Medical Teacher is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Education Research Complete
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: ehh
DbLabel: Education Research Complete
An: 188805000
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Large language models for generating script concordance test in obstetrics and gynecology: ChatGPT and Claude.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Yapıcı+Coşkun%2C+Zuhal%22">Yapıcı Coşkun, Zuhal</searchLink><relatesTo>1,2</relatesTo> (AUTHOR)<i> zuhalyapici@yahoo.com</i><br /><searchLink fieldCode="AR" term="%22Kıyak%2C+Yavuz+Selim%22">Kıyak, Yavuz Selim</searchLink><relatesTo>3</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Coşkun%2C+Özlem%22">Coşkun, Özlem</searchLink><relatesTo>3</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Budakoğlu%2C+Işıl+İrem%22">Budakoğlu, Işıl İrem</searchLink><relatesTo>3</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Özdemir%2C+Özhan%22">Özdemir, Özhan</searchLink><relatesTo>4</relatesTo> (AUTHOR)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Medical+Teacher%22">Medical Teacher</searchLink>. Nov2025, Vol. 47 Issue 11, p1767-1771. 5p.
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: *<searchLink fieldCode="DE" term="%22Generative+artificial+intelligence%22">Generative artificial intelligence</searchLink><br />*<searchLink fieldCode="DE" term="%22Medical+education%22">Medical education</searchLink><br />*<searchLink fieldCode="DE" term="%22Undergraduates%22">Undergraduates</searchLink><br />*<searchLink fieldCode="DE" term="%22Educational+tests+%26+measurements%22">Educational tests & measurements</searchLink><br />*<searchLink fieldCode="DE" term="%22Medical+students%22">Medical students</searchLink><br /><searchLink fieldCode="DE" term="%22Medical+logic%22">Medical logic</searchLink><br /><searchLink fieldCode="DE" term="%22Cross-sectional+method%22">Cross-sectional method</searchLink><br /><searchLink fieldCode="DE" term="%22Scale+analysis+%28Psychology%29%22">Scale analysis (Psychology)</searchLink><br /><searchLink fieldCode="DE" term="%22Primary+health+care%22">Primary health care</searchLink><br /><searchLink fieldCode="DE" term="%22Descriptive+statistics%22">Descriptive statistics</searchLink><br /><searchLink fieldCode="DE" term="%22Gynecology%22">Gynecology</searchLink><br /><searchLink fieldCode="DE" term="%22Chatbots%22">Chatbots</searchLink><br /><searchLink fieldCode="DE" term="%22Obstetrics%22">Obstetrics</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Objective: To evaluate the performance of large language models (ChatGPT-4o and Claude 3.5 Sonnet) to generate script concordance test (SCT) items for assessing clinical reasoning in obstetrics and gynecology. Methods: This cross-sectional study involved the generation of SCT items for five common diagnostic topics in obstetrics and gynecology in primary care settings. A total of 16 panelists evaluated the AI-generated SCT items against 11 predefined criteria. Descriptive statistics were used to compare the models' performance across criteria. Results: ChatGPT-4o had an overall agreement rate of 90.57% for SCT items meeting the quality criteria, while Claude 3.5 Sonnet achieved 91.48%. The criterion with the lowest scores was "The scenario is of appropriate difficulty for medical students," with ChatGPT-4o rated at 71.25% and Claude 3.5 Sonnet at 76.25%. Conclusion: Large language models can generate SCT items that effectively assess clinical reasoning; however, further refinement is required to ensure the appropriate level of difficulty for medical students. These findings highlight the potential of AI to enhance the efficiency of SCT generation in obstetrics and gynecology within primary care settings. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Medical Teacher is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=ehh&AN=188805000
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1080/0142159X.2025.2497888
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 5
        StartPage: 1767
    Subjects:
      – SubjectFull: Generative artificial intelligence
        Type: general
      – SubjectFull: Medical education
        Type: general
      – SubjectFull: Undergraduates
        Type: general
      – SubjectFull: Educational tests & measurements
        Type: general
      – SubjectFull: Medical students
        Type: general
      – SubjectFull: Medical logic
        Type: general
      – SubjectFull: Cross-sectional method
        Type: general
      – SubjectFull: Scale analysis (Psychology)
        Type: general
      – SubjectFull: Primary health care
        Type: general
      – SubjectFull: Descriptive statistics
        Type: general
      – SubjectFull: Gynecology
        Type: general
      – SubjectFull: Chatbots
        Type: general
      – SubjectFull: Obstetrics
        Type: general
    Titles:
      – TitleFull: Large language models for generating script concordance test in obstetrics and gynecology: ChatGPT and Claude.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Yapıcı Coşkun, Zuhal
      – PersonEntity:
          Name:
            NameFull: Kıyak, Yavuz Selim
      – PersonEntity:
          Name:
            NameFull: Coşkun, Özlem
      – PersonEntity:
          Name:
            NameFull: Budakoğlu, Işıl İrem
      – PersonEntity:
          Name:
            NameFull: Özdemir, Özhan
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 11
              Text: Nov2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 0142159X
          Numbering:
            – Type: volume
              Value: 47
            – Type: issue
              Value: 11
          Titles:
            – TitleFull: Medical Teacher
              Type: main
ResultId 1