ChatGPT vs. DeepSeek: A Comparative Psychometric Evaluation of AI Tools in Generating Multiple-Choice Questions

Saved in:
Bibliographic Details
Title: ChatGPT vs. DeepSeek: A Comparative Psychometric Evaluation of AI Tools in Generating Multiple-Choice Questions
Language: English
Authors: Ceylan Gündeger Kilci (ORCID 0000-0003-3572-1708)
Source: International Journal of Assessment Tools in Education. 2025 12(4):1055-1079.
Availability: International Journal of Assessment Tools in Education. Pamukkale University, Faculty of Education, Kinikli Campus, Denizli 20070, Turkey. e-mail: ijate.editor@gmail.com; Web site: https://dergipark.org.tr/en/pub/ijate
Peer Reviewed: Y
Page Count: 25
Publication Date: 2025
Document Type: Journal Articles
Reports - Research
Education Level: Higher Education
Postsecondary Education
Descriptors: Psychometrics, Multiple Choice Tests, Artificial Intelligence, Natural Language Processing, Test Items, Item Analysis, Content Validity, Difficulty Level, Statistical Significance, Scores, Undergraduate Students, Schools of Education, Technology Uses in Education, Test Reliability, Test Validity, Test Construction, Foreign Countries, Expertise, Feedback (Response)
Geographic Terms: Turkey
ISSN: 2148-7456
Abstract: This study examined the psychometric quality of multiple-choice questions generated by two AI tools, ChatGPT and DeepSeek, within the context of an undergraduate Educational Measurement and Evaluation course. Guided by ten learning outcomes (LOs) aligned with Bloom's Taxonomy, each tool was prompted to generate one five-option multiple-choice item per LO. Following expert review (Kendall's "W" = 0.58); revisions were made, and the finalized test was administered to 120 students. Item analyses revealed no statistically significant differences between the two AI models regarding item difficulty, discrimination, variance, or reliability. A few items--two from ChatGPT and one from DeepSeek--had suboptimal discrimination indices. Tetrachoric correlation analyses of item pairs generated by the two AI tools for the same LO revealed that only one pair showed a non-significant association, whereas all other pairs demonstrated statistically significant and generally moderate correlations. KR-20 and split-half reliability coefficients reflected acceptable internal consistency for a classroom-based assessment, with the DeepSeek-generated half showing a slightly stronger correlation with total scores. Expert feedback indicated that while AI tools generally produced valid stems and correct answers, most revisions focused on improving distractor quality, highlighting the need for human refinement. Generalizability and Decision studies confirmed consistency in expert ratings and recommended a minimum of seven experts for reliable evaluations. In conclusion, both AI tools demonstrated the capacity to generate psychometrically comparable items, highlighting their potential to support educators and test developers in test construction. The study concludes with practical recommendations for effectively incorporating AI into test development workflows.
Abstractor: As Provided
Entry Date: 2026
Accession Number: EJ1491386
Database: ERIC
FullText Text:
  Availability: 0
CustomLinks:
  – Url: https://eric.ed.gov/contentdelivery/servlet/ERICServlet?accno=EJ1491386
    Name: ERIC Full Text
    Category: fullText
    Text: Full Text from ERIC
Header DbId: eric
DbLabel: ERIC
An: EJ1491386
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: ChatGPT vs. DeepSeek: A Comparative Psychometric Evaluation of AI Tools in Generating Multiple-Choice Questions
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Ceylan+Gündeger+Kilci%22">Ceylan Gündeger Kilci</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-3572-1708">0000-0003-3572-1708</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22International+Journal+of+Assessment+Tools+in+Education%22"><i>International Journal of Assessment Tools in Education</i></searchLink>. 2025 12(4):1055-1079.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: International Journal of Assessment Tools in Education. Pamukkale University, Faculty of Education, Kinikli Campus, Denizli 20070, Turkey. e-mail: ijate.editor@gmail.com; Web site: https://dergipark.org.tr/en/pub/ijate
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 25
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2025
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Audience
  Label: Education Level
  Group: Audnce
  Data: <searchLink fieldCode="EL" term="%22Higher+Education%22">Higher Education</searchLink><br /><searchLink fieldCode="EL" term="%22Postsecondary+Education%22">Postsecondary Education</searchLink>
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Psychometrics%22">Psychometrics</searchLink><br /><searchLink fieldCode="DE" term="%22Multiple+Choice+Tests%22">Multiple Choice Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+Language+Processing%22">Natural Language Processing</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Items%22">Test Items</searchLink><br /><searchLink fieldCode="DE" term="%22Item+Analysis%22">Item Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Content+Validity%22">Content Validity</searchLink><br /><searchLink fieldCode="DE" term="%22Difficulty+Level%22">Difficulty Level</searchLink><br /><searchLink fieldCode="DE" term="%22Statistical+Significance%22">Statistical Significance</searchLink><br /><searchLink fieldCode="DE" term="%22Scores%22">Scores</searchLink><br /><searchLink fieldCode="DE" term="%22Undergraduate+Students%22">Undergraduate Students</searchLink><br /><searchLink fieldCode="DE" term="%22Schools+of+Education%22">Schools of Education</searchLink><br /><searchLink fieldCode="DE" term="%22Technology+Uses+in+Education%22">Technology Uses in Education</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Reliability%22">Test Reliability</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Validity%22">Test Validity</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Construction%22">Test Construction</searchLink><br /><searchLink fieldCode="DE" term="%22Foreign+Countries%22">Foreign Countries</searchLink><br /><searchLink fieldCode="DE" term="%22Expertise%22">Expertise</searchLink><br /><searchLink fieldCode="DE" term="%22Feedback+%28Response%29%22">Feedback (Response)</searchLink>
– Name: Subject
  Label: Geographic Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Turkey%22">Turkey</searchLink>
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 2148-7456
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: This study examined the psychometric quality of multiple-choice questions generated by two AI tools, ChatGPT and DeepSeek, within the context of an undergraduate Educational Measurement and Evaluation course. Guided by ten learning outcomes (LOs) aligned with Bloom's Taxonomy, each tool was prompted to generate one five-option multiple-choice item per LO. Following expert review (Kendall's "W" = 0.58); revisions were made, and the finalized test was administered to 120 students. Item analyses revealed no statistically significant differences between the two AI models regarding item difficulty, discrimination, variance, or reliability. A few items--two from ChatGPT and one from DeepSeek--had suboptimal discrimination indices. Tetrachoric correlation analyses of item pairs generated by the two AI tools for the same LO revealed that only one pair showed a non-significant association, whereas all other pairs demonstrated statistically significant and generally moderate correlations. KR-20 and split-half reliability coefficients reflected acceptable internal consistency for a classroom-based assessment, with the DeepSeek-generated half showing a slightly stronger correlation with total scores. Expert feedback indicated that while AI tools generally produced valid stems and correct answers, most revisions focused on improving distractor quality, highlighting the need for human refinement. Generalizability and Decision studies confirmed consistency in expert ratings and recommended a minimum of seven experts for reliable evaluations. In conclusion, both AI tools demonstrated the capacity to generate psychometrically comparable items, highlighting their potential to support educators and test developers in test construction. The study concludes with practical recommendations for effectively incorporating AI into test development workflows.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2026
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1491386
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1491386
RecordInfo BibRecord:
  BibEntity:
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 25
        StartPage: 1055
    Subjects:
      – SubjectFull: Psychometrics
        Type: general
      – SubjectFull: Multiple Choice Tests
        Type: general
      – SubjectFull: Artificial Intelligence
        Type: general
      – SubjectFull: Natural Language Processing
        Type: general
      – SubjectFull: Test Items
        Type: general
      – SubjectFull: Item Analysis
        Type: general
      – SubjectFull: Content Validity
        Type: general
      – SubjectFull: Difficulty Level
        Type: general
      – SubjectFull: Statistical Significance
        Type: general
      – SubjectFull: Scores
        Type: general
      – SubjectFull: Undergraduate Students
        Type: general
      – SubjectFull: Schools of Education
        Type: general
      – SubjectFull: Technology Uses in Education
        Type: general
      – SubjectFull: Test Reliability
        Type: general
      – SubjectFull: Test Validity
        Type: general
      – SubjectFull: Test Construction
        Type: general
      – SubjectFull: Foreign Countries
        Type: general
      – SubjectFull: Expertise
        Type: general
      – SubjectFull: Feedback (Response)
        Type: general
      – SubjectFull: Turkey
        Type: general
    Titles:
      – TitleFull: ChatGPT vs. DeepSeek: A Comparative Psychometric Evaluation of AI Tools in Generating Multiple-Choice Questions
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Ceylan Gündeger Kilci
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-electronic
              Value: 2148-7456
          Numbering:
            – Type: volume
              Value: 12
            – Type: issue
              Value: 4
          Titles:
            – TitleFull: International Journal of Assessment Tools in Education
              Type: main
ResultId 1