ChatGPT vs. DeepSeek: A Comparative Psychometric Evaluation of AI Tools in Generating Multiple-Choice Questions
Saved in:
| Title: | ChatGPT vs. DeepSeek: A Comparative Psychometric Evaluation of AI Tools in Generating Multiple-Choice Questions |
|---|---|
| Language: | English |
| Authors: | Ceylan Gündeger Kilci (ORCID |
| Source: | International Journal of Assessment Tools in Education. 2025 12(4):1055-1079. |
| Availability: | International Journal of Assessment Tools in Education. Pamukkale University, Faculty of Education, Kinikli Campus, Denizli 20070, Turkey. e-mail: ijate.editor@gmail.com; Web site: https://dergipark.org.tr/en/pub/ijate |
| Peer Reviewed: | Y |
| Page Count: | 25 |
| Publication Date: | 2025 |
| Document Type: | Journal Articles Reports - Research |
| Education Level: | Higher Education Postsecondary Education |
| Descriptors: | Psychometrics, Multiple Choice Tests, Artificial Intelligence, Natural Language Processing, Test Items, Item Analysis, Content Validity, Difficulty Level, Statistical Significance, Scores, Undergraduate Students, Schools of Education, Technology Uses in Education, Test Reliability, Test Validity, Test Construction, Foreign Countries, Expertise, Feedback (Response) |
| Geographic Terms: | Turkey |
| ISSN: | 2148-7456 |
| Abstract: | This study examined the psychometric quality of multiple-choice questions generated by two AI tools, ChatGPT and DeepSeek, within the context of an undergraduate Educational Measurement and Evaluation course. Guided by ten learning outcomes (LOs) aligned with Bloom's Taxonomy, each tool was prompted to generate one five-option multiple-choice item per LO. Following expert review (Kendall's "W" = 0.58); revisions were made, and the finalized test was administered to 120 students. Item analyses revealed no statistically significant differences between the two AI models regarding item difficulty, discrimination, variance, or reliability. A few items--two from ChatGPT and one from DeepSeek--had suboptimal discrimination indices. Tetrachoric correlation analyses of item pairs generated by the two AI tools for the same LO revealed that only one pair showed a non-significant association, whereas all other pairs demonstrated statistically significant and generally moderate correlations. KR-20 and split-half reliability coefficients reflected acceptable internal consistency for a classroom-based assessment, with the DeepSeek-generated half showing a slightly stronger correlation with total scores. Expert feedback indicated that while AI tools generally produced valid stems and correct answers, most revisions focused on improving distractor quality, highlighting the need for human refinement. Generalizability and Decision studies confirmed consistency in expert ratings and recommended a minimum of seven experts for reliable evaluations. In conclusion, both AI tools demonstrated the capacity to generate psychometrically comparable items, highlighting their potential to support educators and test developers in test construction. The study concludes with practical recommendations for effectively incorporating AI into test development workflows. |
| Abstractor: | As Provided |
| Entry Date: | 2026 |
| Accession Number: | EJ1491386 |
| Database: | ERIC |
| FullText | Text: Availability: 0 CustomLinks: – Url: https://eric.ed.gov/contentdelivery/servlet/ERICServlet?accno=EJ1491386 Name: ERIC Full Text Category: fullText Text: Full Text from ERIC |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1491386 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: ChatGPT vs. DeepSeek: A Comparative Psychometric Evaluation of AI Tools in Generating Multiple-Choice Questions – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Ceylan+Gündeger+Kilci%22">Ceylan Gündeger Kilci</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-3572-1708">0000-0003-3572-1708</externalLink>) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22International+Journal+of+Assessment+Tools+in+Education%22"><i>International Journal of Assessment Tools in Education</i></searchLink>. 2025 12(4):1055-1079. – Name: Avail Label: Availability Group: Avail Data: International Journal of Assessment Tools in Education. Pamukkale University, Faculty of Education, Kinikli Campus, Denizli 20070, Turkey. e-mail: ijate.editor@gmail.com; Web site: https://dergipark.org.tr/en/pub/ijate – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 25 – Name: DatePubCY Label: Publication Date Group: Date Data: 2025 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Audience Label: Education Level Group: Audnce Data: <searchLink fieldCode="EL" term="%22Higher+Education%22">Higher Education</searchLink><br /><searchLink fieldCode="EL" term="%22Postsecondary+Education%22">Postsecondary Education</searchLink> – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Psychometrics%22">Psychometrics</searchLink><br /><searchLink fieldCode="DE" term="%22Multiple+Choice+Tests%22">Multiple Choice Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+Language+Processing%22">Natural Language Processing</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Items%22">Test Items</searchLink><br /><searchLink fieldCode="DE" term="%22Item+Analysis%22">Item Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Content+Validity%22">Content Validity</searchLink><br /><searchLink fieldCode="DE" term="%22Difficulty+Level%22">Difficulty Level</searchLink><br /><searchLink fieldCode="DE" term="%22Statistical+Significance%22">Statistical Significance</searchLink><br /><searchLink fieldCode="DE" term="%22Scores%22">Scores</searchLink><br /><searchLink fieldCode="DE" term="%22Undergraduate+Students%22">Undergraduate Students</searchLink><br /><searchLink fieldCode="DE" term="%22Schools+of+Education%22">Schools of Education</searchLink><br /><searchLink fieldCode="DE" term="%22Technology+Uses+in+Education%22">Technology Uses in Education</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Reliability%22">Test Reliability</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Validity%22">Test Validity</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Construction%22">Test Construction</searchLink><br /><searchLink fieldCode="DE" term="%22Foreign+Countries%22">Foreign Countries</searchLink><br /><searchLink fieldCode="DE" term="%22Expertise%22">Expertise</searchLink><br /><searchLink fieldCode="DE" term="%22Feedback+%28Response%29%22">Feedback (Response)</searchLink> – Name: Subject Label: Geographic Terms Group: Su Data: <searchLink fieldCode="DE" term="%22Turkey%22">Turkey</searchLink> – Name: ISSN Label: ISSN Group: ISSN Data: 2148-7456 – Name: Abstract Label: Abstract Group: Ab Data: This study examined the psychometric quality of multiple-choice questions generated by two AI tools, ChatGPT and DeepSeek, within the context of an undergraduate Educational Measurement and Evaluation course. Guided by ten learning outcomes (LOs) aligned with Bloom's Taxonomy, each tool was prompted to generate one five-option multiple-choice item per LO. Following expert review (Kendall's "W" = 0.58); revisions were made, and the finalized test was administered to 120 students. Item analyses revealed no statistically significant differences between the two AI models regarding item difficulty, discrimination, variance, or reliability. A few items--two from ChatGPT and one from DeepSeek--had suboptimal discrimination indices. Tetrachoric correlation analyses of item pairs generated by the two AI tools for the same LO revealed that only one pair showed a non-significant association, whereas all other pairs demonstrated statistically significant and generally moderate correlations. KR-20 and split-half reliability coefficients reflected acceptable internal consistency for a classroom-based assessment, with the DeepSeek-generated half showing a slightly stronger correlation with total scores. Expert feedback indicated that while AI tools generally produced valid stems and correct answers, most revisions focused on improving distractor quality, highlighting the need for human refinement. Generalizability and Decision studies confirmed consistency in expert ratings and recommended a minimum of seven experts for reliable evaluations. In conclusion, both AI tools demonstrated the capacity to generate psychometrically comparable items, highlighting their potential to support educators and test developers in test construction. The study concludes with practical recommendations for effectively incorporating AI into test development workflows. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2026 – Name: AN Label: Accession Number Group: ID Data: EJ1491386 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1491386 |
| RecordInfo | BibRecord: BibEntity: Languages: – Text: English PhysicalDescription: Pagination: PageCount: 25 StartPage: 1055 Subjects: – SubjectFull: Psychometrics Type: general – SubjectFull: Multiple Choice Tests Type: general – SubjectFull: Artificial Intelligence Type: general – SubjectFull: Natural Language Processing Type: general – SubjectFull: Test Items Type: general – SubjectFull: Item Analysis Type: general – SubjectFull: Content Validity Type: general – SubjectFull: Difficulty Level Type: general – SubjectFull: Statistical Significance Type: general – SubjectFull: Scores Type: general – SubjectFull: Undergraduate Students Type: general – SubjectFull: Schools of Education Type: general – SubjectFull: Technology Uses in Education Type: general – SubjectFull: Test Reliability Type: general – SubjectFull: Test Validity Type: general – SubjectFull: Test Construction Type: general – SubjectFull: Foreign Countries Type: general – SubjectFull: Expertise Type: general – SubjectFull: Feedback (Response) Type: general – SubjectFull: Turkey Type: general Titles: – TitleFull: ChatGPT vs. DeepSeek: A Comparative Psychometric Evaluation of AI Tools in Generating Multiple-Choice Questions Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Ceylan Gündeger Kilci IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 01 Type: published Y: 2025 Identifiers: – Type: issn-electronic Value: 2148-7456 Numbering: – Type: volume Value: 12 – Type: issue Value: 4 Titles: – TitleFull: International Journal of Assessment Tools in Education Type: main |
| ResultId | 1 |