Assessing AI's problem solving in physics: Analyzing reasoning, false positives and negatives through the force concept inventory.

Saved in:
Bibliographic Details
Title: Assessing AI's problem solving in physics: Analyzing reasoning, false positives and negatives through the force concept inventory.
Authors: Aldazharova, Salima1, Issayeva, Gulnara1, Maxutov, Samat2, Balta, Nuri2 baltanuri@gmail.com
Source: Contemporary Educational Technology. Oct2024, Vol. 16 Issue 4, p1-16. 16p.
Subject Terms: *Physics education, *Common misconceptions, *Educational outcomes, Newton's laws of motion, Generative pre-trained transformers
Abstract: This study investigates the performance of GPT-4, an advanced AI model developed by OpenAI, on the force concept inventory (FCI) to evaluate its accuracy, reasoning patterns, and the occurrence of false positives and false negatives. GPT-4 was tasked with answering the FCI questions across multiple sessions. Key findings include GPT-4's proficiency in several FCI items, particularly those related to Newton's third law, achieving perfect scores on many items. However, it struggled significantly with questions involving the interpretation of figures and spatial reasoning, resulting in a higher occurrence of false negatives where the reasoning was correct, but the answers were incorrect. Additionally, GPT-4 displayed several conceptual errors, such as misunderstanding the effect of friction and retaining the outdated impetus theory of motion. The study's findings emphasize the importance of refining AI-driven tools to make them more effective in educational settings. Addressing both AI limitations and common misconceptions in physics can lead to improved educational outcomes. [ABSTRACT FROM AUTHOR]
Copyright of Contemporary Educational Technology is the property of Bastas Publications and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Education Research Complete
Description
Abstract:This study investigates the performance of GPT-4, an advanced AI model developed by OpenAI, on the force concept inventory (FCI) to evaluate its accuracy, reasoning patterns, and the occurrence of false positives and false negatives. GPT-4 was tasked with answering the FCI questions across multiple sessions. Key findings include GPT-4's proficiency in several FCI items, particularly those related to Newton's third law, achieving perfect scores on many items. However, it struggled significantly with questions involving the interpretation of figures and spatial reasoning, resulting in a higher occurrence of false negatives where the reasoning was correct, but the answers were incorrect. Additionally, GPT-4 displayed several conceptual errors, such as misunderstanding the effect of friction and retaining the outdated impetus theory of motion. The study's findings emphasize the importance of refining AI-driven tools to make them more effective in educational settings. Addressing both AI limitations and common misconceptions in physics can lead to improved educational outcomes. [ABSTRACT FROM AUTHOR]
ISSN:1309517X
DOI:10.30935/cedtech/15592