Automated Assessment in Math Education: A Comparative Analysis of LLMs for Open-Ended Responses

Saved in:
Bibliographic Details
Title: Automated Assessment in Math Education: A Comparative Analysis of LLMs for Open-Ended Responses
Language: English
Authors: Sami Baral, Eamon Worden, Wen-Chiang Lim, Zhuang Luo, Christopher Santorelli, Ashish Gurung, Neil Heffernan
Source: Grantee Submission. 2024Paper presented at the International Conference on Educational Data Mining (17th, Atlanta, GA, Jul 2024).
Peer Reviewed: Y
Page Count: 6
Publication Date: 2024
Sponsoring Agency: Institute of Education Sciences (ED)
National Science Foundation (NSF)
Office of Naval Research (ONR) (DOD)
National Institutes of Health (NIH) (DHHS)
Office of Postsecondary Education (ED)
Office of Elementary and Secondary Education (OESE) (ED), Education Innovation and Research (EIR)
Contract Number: R305N210049
R305D210031
R305A170137
R305A170243
R305A180401
R305A120125
R305R220012
R305T240029
2118725
2118904
1950683
1917808
1931523
1940236
1917713
1903304
1822830
1759229
1724889
1636782
1535428
N000141812768
R44GM146483
P200A120238
P200A180088
P200A150306
U411B190024
S411B210024
S411B220024
Document Type: Reports - Research
Speeches/Meeting Papers
Education Level: Junior High Schools
Middle Schools
Secondary Education
Descriptors: Automation, Scoring, Computer Assisted Testing, Natural Language Processing, Feedback (Response), Student Evaluation, Mathematics Education, Mathematics Tests, Scores, Middle School Students, Models, Algorithms
DOI: 10.5281/zenodo.12729932
Abstract: The effectiveness of feedback in enhancing learning outcomes is well documented within Educational Data Mining (EDM). Various prior research have explored methodologies to enhance the effectiveness of feedback to students in various ways. Recent developments in Large Language Models (LLMs) have extended their utility in enhancing automated feedback systems. This study aims to explore the potential of LLMs in facilitating automated feedback in math education in the form of numeric assessment scores. We examine the effectiveness of LLMs in evaluating student responses and scoring the responses by comparing 3 different models: Llama, SBERT-Canberra, and GPT4 model. The evaluation requires the model to provide a quantitative score on the student's responses to open-ended math problems. We employ Mistral, a version of Llama catered to math, and fine-tune this model for evaluating student responses by leveraging a dataset of student responses and teacher-provided scores for middle-school math problems. A similar approach was taken for training the SBERT-Canberra model, while the GPT4 model used a zero-shot learning approach. We evaluate and compare the models' performance in scoring accuracy. This study aims to further the ongoing development of automated assessment and feedback systems and outline potential future directions for leveraging generative LLMs in building automated feedback systems. [This paper was published in: "Proceedings of the 17th International Conference on Educational Data Mining," edited by B. Paaßen and C. D. Epp, International Educational Data Mining Society, 2024, pp. 732-737. Funding for this paper was provided by the U.S. Department of Education's Graduate Assistance in Areas of National Need (GAANN).]
Abstractor: As Provided
IES Funded: Yes
Entry Date: 2024
Accession Number: ED661970
Database: ERIC
FullText Text:
  Availability: 0
Header DbId: eric
DbLabel: ERIC
An: ED661970
AccessLevel: 3
PubType: Report
PubTypeId: report
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Automated Assessment in Math Education: A Comparative Analysis of LLMs for Open-Ended Responses
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Sami+Baral%22">Sami Baral</searchLink><br /><searchLink fieldCode="AR" term="%22Eamon+Worden%22">Eamon Worden</searchLink><br /><searchLink fieldCode="AR" term="%22Wen-Chiang+Lim%22">Wen-Chiang Lim</searchLink><br /><searchLink fieldCode="AR" term="%22Zhuang+Luo%22">Zhuang Luo</searchLink><br /><searchLink fieldCode="AR" term="%22Christopher+Santorelli%22">Christopher Santorelli</searchLink><br /><searchLink fieldCode="AR" term="%22Ashish+Gurung%22">Ashish Gurung</searchLink><br /><searchLink fieldCode="AR" term="%22Neil+Heffernan%22">Neil Heffernan</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Grantee+Submission%22"><i>Grantee Submission</i></searchLink>. 2024Paper presented at the International Conference on Educational Data Mining (17th, Atlanta, GA, Jul 2024).
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 6
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2024
– Name: SourceSuprt
  Label: Sponsoring Agency
  Group: SrcSuprt
  Data: Institute of Education Sciences (ED)<br />National Science Foundation (NSF)<br />Office of Naval Research (ONR) (DOD)<br />National Institutes of Health (NIH) (DHHS)<br />Office of Postsecondary Education (ED)<br />Office of Elementary and Secondary Education (OESE) (ED), Education Innovation and Research (EIR)
– Name: NumberContract
  Label: Contract Number
  Group: NumCntrct
  Data: R305N210049<br />R305D210031<br />R305A170137<br />R305A170243<br />R305A180401<br />R305A120125<br />R305R220012<br />R305T240029<br />2118725<br />2118904<br />1950683<br />1917808<br />1931523<br />1940236<br />1917713<br />1903304<br />1822830<br />1759229<br />1724889<br />1636782<br />1535428<br />N000141812768<br />R44GM146483<br />P200A120238<br />P200A180088<br />P200A150306<br />U411B190024<br />S411B210024<br />S411B220024
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Reports - Research<br />Speeches/Meeting Papers
– Name: Audience
  Label: Education Level
  Group: Audnce
  Data: <searchLink fieldCode="EL" term="%22Junior+High+Schools%22">Junior High Schools</searchLink><br /><searchLink fieldCode="EL" term="%22Middle+Schools%22">Middle Schools</searchLink><br /><searchLink fieldCode="EL" term="%22Secondary+Education%22">Secondary Education</searchLink>
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Automation%22">Automation</searchLink><br /><searchLink fieldCode="DE" term="%22Scoring%22">Scoring</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+Assisted+Testing%22">Computer Assisted Testing</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+Language+Processing%22">Natural Language Processing</searchLink><br /><searchLink fieldCode="DE" term="%22Feedback+%28Response%29%22">Feedback (Response)</searchLink><br /><searchLink fieldCode="DE" term="%22Student+Evaluation%22">Student Evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Mathematics+Education%22">Mathematics Education</searchLink><br /><searchLink fieldCode="DE" term="%22Mathematics+Tests%22">Mathematics Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Scores%22">Scores</searchLink><br /><searchLink fieldCode="DE" term="%22Middle+School+Students%22">Middle School Students</searchLink><br /><searchLink fieldCode="DE" term="%22Models%22">Models</searchLink><br /><searchLink fieldCode="DE" term="%22Algorithms%22">Algorithms</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.5281/zenodo.12729932
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: The effectiveness of feedback in enhancing learning outcomes is well documented within Educational Data Mining (EDM). Various prior research have explored methodologies to enhance the effectiveness of feedback to students in various ways. Recent developments in Large Language Models (LLMs) have extended their utility in enhancing automated feedback systems. This study aims to explore the potential of LLMs in facilitating automated feedback in math education in the form of numeric assessment scores. We examine the effectiveness of LLMs in evaluating student responses and scoring the responses by comparing 3 different models: Llama, SBERT-Canberra, and GPT4 model. The evaluation requires the model to provide a quantitative score on the student's responses to open-ended math problems. We employ Mistral, a version of Llama catered to math, and fine-tune this model for evaluating student responses by leveraging a dataset of student responses and teacher-provided scores for middle-school math problems. A similar approach was taken for training the SBERT-Canberra model, while the GPT4 model used a zero-shot learning approach. We evaluate and compare the models' performance in scoring accuracy. This study aims to further the ongoing development of automated assessment and feedback systems and outline potential future directions for leveraging generative LLMs in building automated feedback systems. [This paper was published in: "Proceedings of the 17th International Conference on Educational Data Mining," edited by B. Paaßen and C. D. Epp, International Educational Data Mining Society, 2024, pp. 732-737. Funding for this paper was provided by the U.S. Department of Education's Graduate Assistance in Areas of National Need (GAANN).]
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: CodeSource
  Label: IES Funded
  Group: SrcInfo
  Data: Yes
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2024
– Name: AN
  Label: Accession Number
  Group: ID
  Data: ED661970
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=ED661970
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.5281/zenodo.12729932
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 6
    Subjects:
      – SubjectFull: Automation
        Type: general
      – SubjectFull: Scoring
        Type: general
      – SubjectFull: Computer Assisted Testing
        Type: general
      – SubjectFull: Natural Language Processing
        Type: general
      – SubjectFull: Feedback (Response)
        Type: general
      – SubjectFull: Student Evaluation
        Type: general
      – SubjectFull: Mathematics Education
        Type: general
      – SubjectFull: Mathematics Tests
        Type: general
      – SubjectFull: Scores
        Type: general
      – SubjectFull: Middle School Students
        Type: general
      – SubjectFull: Models
        Type: general
      – SubjectFull: Algorithms
        Type: general
    Titles:
      – TitleFull: Automated Assessment in Math Education: A Comparative Analysis of LLMs for Open-Ended Responses
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Sami Baral
      – PersonEntity:
          Name:
            NameFull: Eamon Worden
      – PersonEntity:
          Name:
            NameFull: Wen-Chiang Lim
      – PersonEntity:
          Name:
            NameFull: Zhuang Luo
      – PersonEntity:
          Name:
            NameFull: Christopher Santorelli
      – PersonEntity:
          Name:
            NameFull: Ashish Gurung
      – PersonEntity:
          Name:
            NameFull: Neil Heffernan
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 07
              Type: published
              Y: 2024
          Titles:
            – TitleFull: Grantee Submission
              Type: main
ResultId 1