Nonstandard English and the Automated Scoring of Open-Ended Math Problems
Saved in:
| Title: | Nonstandard English and the Automated Scoring of Open-Ended Math Problems |
|---|---|
| Language: | English |
| Authors: | Abubakir Siedahm, Jaclyn Ocumpaugh, Zelda Ferris, Dinesh Kodwani, Eamon Worden, Neil Heffernan |
| Source: | International Educational Data Mining Society. 2025. |
| Availability: | International Educational Data Mining Society. e-mail: admin@educationaldatamining.org; Web site: https://educationaldatamining.org/conferences/ |
| Peer Reviewed: | Y |
| Page Count: | 11 |
| Publication Date: | 2025 |
| Document Type: | Speeches/Meeting Papers Reports - Research |
| Education Level: | Secondary Education |
| Descriptors: | Artificial Intelligence, Automation, Scoring, Mathematics Tests, Grading, Black Dialects, Nonstandard Dialects, Natural Language Processing, African American Students, Secondary School Mathematics |
| Abstract: | Recent advances in AI have opened the door for the automated scoring of open-ended math problems, which were previously much more difficult to assess at scale. However, we know that biases still remain in some of these algorithms. For example, recent research on the automated scoring of student essays has shown that certain varieties of English are more strongly penalized for nonstandard English than they are for other differences that reduce the quality of students' writing. This study examines that issue in a new domain, investigating the potential for large language models to accurately grade open-ended math problems produced by students who speak and write in non-standard English. Specifically, we look at four features of African American Vernacular English (AAVE), which range in the degree to which they are unique to AAVE or are common in other non-standard dialects. We then compare the scoring of answers that were produced by students using these dialect features to a control group of synthetic data--where we converted all non-standard dialect features to standard English. Results show that minor changes in the number of dialect features per student response do not impact GPTs automated scoring, but prompt engineering efforts did. [For the complete proceedings, see ED675583.] |
| Abstractor: | As Provided |
| Entry Date: | 2025 |
| Accession Number: | ED675646 |
| Database: | ERIC |
| FullText | Text: Availability: 0 CustomLinks: – Url: https://eric.ed.gov/contentdelivery/servlet/ERICServlet?accno=ED675646 Name: ERIC Full Text Category: fullText Text: Full Text from ERIC |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: ED675646 AccessLevel: 3 PubType: Conference PubTypeId: conference PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Nonstandard English and the Automated Scoring of Open-Ended Math Problems – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Abubakir+Siedahm%22">Abubakir Siedahm</searchLink><br /><searchLink fieldCode="AR" term="%22Jaclyn+Ocumpaugh%22">Jaclyn Ocumpaugh</searchLink><br /><searchLink fieldCode="AR" term="%22Zelda+Ferris%22">Zelda Ferris</searchLink><br /><searchLink fieldCode="AR" term="%22Dinesh+Kodwani%22">Dinesh Kodwani</searchLink><br /><searchLink fieldCode="AR" term="%22Eamon+Worden%22">Eamon Worden</searchLink><br /><searchLink fieldCode="AR" term="%22Neil+Heffernan%22">Neil Heffernan</searchLink> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22International+Educational+Data+Mining+Society%22"><i>International Educational Data Mining Society</i></searchLink>. 2025. – Name: Avail Label: Availability Group: Avail Data: International Educational Data Mining Society. e-mail: admin@educationaldatamining.org; Web site: https://educationaldatamining.org/conferences/ – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 11 – Name: DatePubCY Label: Publication Date Group: Date Data: 2025 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Speeches/Meeting Papers<br />Reports - Research – Name: Audience Label: Education Level Group: Audnce Data: <searchLink fieldCode="EL" term="%22Secondary+Education%22">Secondary Education</searchLink> – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Automation%22">Automation</searchLink><br /><searchLink fieldCode="DE" term="%22Scoring%22">Scoring</searchLink><br /><searchLink fieldCode="DE" term="%22Mathematics+Tests%22">Mathematics Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Grading%22">Grading</searchLink><br /><searchLink fieldCode="DE" term="%22Black+Dialects%22">Black Dialects</searchLink><br /><searchLink fieldCode="DE" term="%22Nonstandard+Dialects%22">Nonstandard Dialects</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+Language+Processing%22">Natural Language Processing</searchLink><br /><searchLink fieldCode="DE" term="%22African+American+Students%22">African American Students</searchLink><br /><searchLink fieldCode="DE" term="%22Secondary+School+Mathematics%22">Secondary School Mathematics</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Recent advances in AI have opened the door for the automated scoring of open-ended math problems, which were previously much more difficult to assess at scale. However, we know that biases still remain in some of these algorithms. For example, recent research on the automated scoring of student essays has shown that certain varieties of English are more strongly penalized for nonstandard English than they are for other differences that reduce the quality of students' writing. This study examines that issue in a new domain, investigating the potential for large language models to accurately grade open-ended math problems produced by students who speak and write in non-standard English. Specifically, we look at four features of African American Vernacular English (AAVE), which range in the degree to which they are unique to AAVE or are common in other non-standard dialects. We then compare the scoring of answers that were produced by students using these dialect features to a control group of synthetic data--where we converted all non-standard dialect features to standard English. Results show that minor changes in the number of dialect features per student response do not impact GPTs automated scoring, but prompt engineering efforts did. [For the complete proceedings, see ED675583.] – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2025 – Name: AN Label: Accession Number Group: ID Data: ED675646 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=ED675646 |
| RecordInfo | BibRecord: BibEntity: Languages: – Text: English PhysicalDescription: Pagination: PageCount: 11 Subjects: – SubjectFull: Artificial Intelligence Type: general – SubjectFull: Automation Type: general – SubjectFull: Scoring Type: general – SubjectFull: Mathematics Tests Type: general – SubjectFull: Grading Type: general – SubjectFull: Black Dialects Type: general – SubjectFull: Nonstandard Dialects Type: general – SubjectFull: Natural Language Processing Type: general – SubjectFull: African American Students Type: general – SubjectFull: Secondary School Mathematics Type: general Titles: – TitleFull: Nonstandard English and the Automated Scoring of Open-Ended Math Problems Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Abubakir Siedahm – PersonEntity: Name: NameFull: Jaclyn Ocumpaugh – PersonEntity: Name: NameFull: Zelda Ferris – PersonEntity: Name: NameFull: Dinesh Kodwani – PersonEntity: Name: NameFull: Eamon Worden – PersonEntity: Name: NameFull: Neil Heffernan IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 01 Type: published Y: 2025 Titles: – TitleFull: International Educational Data Mining Society Type: main |
| ResultId | 1 |