How Do Physics Students Evaluate Artificial Intelligence Responses on Comprehension Questions? A Study on the Perceived Scientific Accuracy and Linguistic Quality of ChatGPT

Saved in:
Bibliographic Details
Title: How Do Physics Students Evaluate Artificial Intelligence Responses on Comprehension Questions? A Study on the Perceived Scientific Accuracy and Linguistic Quality of ChatGPT
Language: English
Authors: Dahlkemper, Merten Nikolay (ORCID 0000-0002-4453-056X), Lahme, Simon Zacharias (ORCID 0000-0002-7251-8832), Klein, Pascal (ORCID 0000-0003-3023-1478)
Source: Physical Review Physics Education Research. 2023 19(1).
Availability: American Physical Society. One Physics Ellipse 4th Floor, College Park, MD 20740-3844. Tel: 301-209-3200; Fax: 301-209-0865; e-mail: assocpub@aps.org; Web site: https://journals.aps.org/prper/
Peer Reviewed: Y
Page Count: 25
Publication Date: 2023
Document Type: Journal Articles
Reports - Research
Tests/Questionnaires
Education Level: Higher Education
Postsecondary Education
Descriptors: Physics, Science Instruction, Artificial Intelligence, Computer Software, Accuracy, Questioning Techniques, Mechanics (Physics), Difficulty Level, Undergraduate Students, Student Attitudes, Introductory Courses, Computational Linguistics, Misconceptions, Item Analysis, Comparative Analysis, Critical Thinking, Foreign Countries, German
Geographic Terms: Germany
DOI: 10.1103/PhysRevPhysEducRes.19.010142
ISSN: 2469-9896
Abstract: This study aimed at evaluating how students perceive the linguistic quality and scientific accuracy of ChatGPT responses to physics comprehension questions. A total of 102 first- and second-year physics students were confronted with three questions of progressing difficulty from introductory mechanics (rolling motion, waves, and fluid dynamics). Each question was presented with four different responses. All responses were attributed to ChatGPT, but in reality, one sample solution was created by the researchers. All ChatGPT responses obtained in this study were wrong, imprecise, incomplete, or misleading. We found little differences in the perceived linguistic quality between ChatGPT responses and the sample solution. However, the students rated the overall scientific accuracy of the responses significantly differently, with the sample solution being rated best for the questions of low and medium difficulty. The discrepancy between the sample solution and the ChatGPT responses increased with the level of self-assessed knowledge of the question content. For the question of highest difficulty (fluid dynamics) that was unknown to most students, a ChatGPT response was rated just as good as the sample solution. Thus, this study provides data on the students' perception of ChatGPT responses and the factors influencing their perception. The results highlight the need for careful evaluation of ChatGPT responses both by instructors and students, particularly regarding scientific accuracy. Therefore, future research could explore the potential of similar "spot the bot" activities in physics education to foster students' critical thinking skills.
Abstractor: As Provided
Entry Date: 2023
Accession Number: EJ1397189
Database: ERIC
FullText Text:
  Availability: 0
Header DbId: eric
DbLabel: ERIC
An: EJ1397189
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: How Do Physics Students Evaluate Artificial Intelligence Responses on Comprehension Questions? A Study on the Perceived Scientific Accuracy and Linguistic Quality of ChatGPT
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Dahlkemper%2C+Merten+Nikolay%22">Dahlkemper, Merten Nikolay</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-4453-056X">0000-0002-4453-056X</externalLink>)<br /><searchLink fieldCode="AR" term="%22Lahme%2C+Simon+Zacharias%22">Lahme, Simon Zacharias</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-7251-8832">0000-0002-7251-8832</externalLink>)<br /><searchLink fieldCode="AR" term="%22Klein%2C+Pascal%22">Klein, Pascal</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-3023-1478">0000-0003-3023-1478</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Physical+Review+Physics+Education+Research%22"><i>Physical Review Physics Education Research</i></searchLink>. 2023 19(1).
– Name: Avail
  Label: Availability
  Group: Avail
  Data: American Physical Society. One Physics Ellipse 4th Floor, College Park, MD 20740-3844. Tel: 301-209-3200; Fax: 301-209-0865; e-mail: assocpub@aps.org; Web site: https://journals.aps.org/prper/
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 25
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2023
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research<br />Tests/Questionnaires
– Name: Audience
  Label: Education Level
  Group: Audnce
  Data: <searchLink fieldCode="EL" term="%22Higher+Education%22">Higher Education</searchLink><br /><searchLink fieldCode="EL" term="%22Postsecondary+Education%22">Postsecondary Education</searchLink>
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Physics%22">Physics</searchLink><br /><searchLink fieldCode="DE" term="%22Science+Instruction%22">Science Instruction</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+Software%22">Computer Software</searchLink><br /><searchLink fieldCode="DE" term="%22Accuracy%22">Accuracy</searchLink><br /><searchLink fieldCode="DE" term="%22Questioning+Techniques%22">Questioning Techniques</searchLink><br /><searchLink fieldCode="DE" term="%22Mechanics+%28Physics%29%22">Mechanics (Physics)</searchLink><br /><searchLink fieldCode="DE" term="%22Difficulty+Level%22">Difficulty Level</searchLink><br /><searchLink fieldCode="DE" term="%22Undergraduate+Students%22">Undergraduate Students</searchLink><br /><searchLink fieldCode="DE" term="%22Student+Attitudes%22">Student Attitudes</searchLink><br /><searchLink fieldCode="DE" term="%22Introductory+Courses%22">Introductory Courses</searchLink><br /><searchLink fieldCode="DE" term="%22Computational+Linguistics%22">Computational Linguistics</searchLink><br /><searchLink fieldCode="DE" term="%22Misconceptions%22">Misconceptions</searchLink><br /><searchLink fieldCode="DE" term="%22Item+Analysis%22">Item Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Comparative+Analysis%22">Comparative Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Critical+Thinking%22">Critical Thinking</searchLink><br /><searchLink fieldCode="DE" term="%22Foreign+Countries%22">Foreign Countries</searchLink><br /><searchLink fieldCode="DE" term="%22German%22">German</searchLink>
– Name: Subject
  Label: Geographic Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Germany%22">Germany</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1103/PhysRevPhysEducRes.19.010142
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 2469-9896
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: This study aimed at evaluating how students perceive the linguistic quality and scientific accuracy of ChatGPT responses to physics comprehension questions. A total of 102 first- and second-year physics students were confronted with three questions of progressing difficulty from introductory mechanics (rolling motion, waves, and fluid dynamics). Each question was presented with four different responses. All responses were attributed to ChatGPT, but in reality, one sample solution was created by the researchers. All ChatGPT responses obtained in this study were wrong, imprecise, incomplete, or misleading. We found little differences in the perceived linguistic quality between ChatGPT responses and the sample solution. However, the students rated the overall scientific accuracy of the responses significantly differently, with the sample solution being rated best for the questions of low and medium difficulty. The discrepancy between the sample solution and the ChatGPT responses increased with the level of self-assessed knowledge of the question content. For the question of highest difficulty (fluid dynamics) that was unknown to most students, a ChatGPT response was rated just as good as the sample solution. Thus, this study provides data on the students' perception of ChatGPT responses and the factors influencing their perception. The results highlight the need for careful evaluation of ChatGPT responses both by instructors and students, particularly regarding scientific accuracy. Therefore, future research could explore the potential of similar "spot the bot" activities in physics education to foster students' critical thinking skills.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2023
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1397189
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1397189
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1103/PhysRevPhysEducRes.19.010142
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 25
    Subjects:
      – SubjectFull: Physics
        Type: general
      – SubjectFull: Science Instruction
        Type: general
      – SubjectFull: Artificial Intelligence
        Type: general
      – SubjectFull: Computer Software
        Type: general
      – SubjectFull: Accuracy
        Type: general
      – SubjectFull: Questioning Techniques
        Type: general
      – SubjectFull: Mechanics (Physics)
        Type: general
      – SubjectFull: Difficulty Level
        Type: general
      – SubjectFull: Undergraduate Students
        Type: general
      – SubjectFull: Student Attitudes
        Type: general
      – SubjectFull: Introductory Courses
        Type: general
      – SubjectFull: Computational Linguistics
        Type: general
      – SubjectFull: Misconceptions
        Type: general
      – SubjectFull: Item Analysis
        Type: general
      – SubjectFull: Comparative Analysis
        Type: general
      – SubjectFull: Critical Thinking
        Type: general
      – SubjectFull: Foreign Countries
        Type: general
      – SubjectFull: German
        Type: general
      – SubjectFull: Germany
        Type: general
    Titles:
      – TitleFull: How Do Physics Students Evaluate Artificial Intelligence Responses on Comprehension Questions? A Study on the Perceived Scientific Accuracy and Linguistic Quality of ChatGPT
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Dahlkemper, Merten Nikolay
      – PersonEntity:
          Name:
            NameFull: Lahme, Simon Zacharias
      – PersonEntity:
          Name:
            NameFull: Klein, Pascal
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2023
          Identifiers:
            – Type: issn-electronic
              Value: 2469-9896
          Numbering:
            – Type: volume
              Value: 19
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Physical Review Physics Education Research
              Type: main
ResultId 1