Nonstandard English and the Automated Scoring of Open-Ended Math Problems
Saved in:
| Title: | Nonstandard English and the Automated Scoring of Open-Ended Math Problems |
|---|---|
| Language: | English |
| Authors: | Abubakir Siedahm, Jaclyn Ocumpaugh, Zelda Ferris, Dinesh Kodwani, Eamon Worden, Neil Heffernan |
| Source: | International Educational Data Mining Society. 2025. |
| Availability: | International Educational Data Mining Society. e-mail: admin@educationaldatamining.org; Web site: https://educationaldatamining.org/conferences/ |
| Peer Reviewed: | Y |
| Page Count: | 11 |
| Publication Date: | 2025 |
| Document Type: | Speeches/Meeting Papers Reports - Research |
| Education Level: | Secondary Education |
| Descriptors: | Artificial Intelligence, Automation, Scoring, Mathematics Tests, Grading, Black Dialects, Nonstandard Dialects, Natural Language Processing, African American Students, Secondary School Mathematics |
| Abstract: | Recent advances in AI have opened the door for the automated scoring of open-ended math problems, which were previously much more difficult to assess at scale. However, we know that biases still remain in some of these algorithms. For example, recent research on the automated scoring of student essays has shown that certain varieties of English are more strongly penalized for nonstandard English than they are for other differences that reduce the quality of students' writing. This study examines that issue in a new domain, investigating the potential for large language models to accurately grade open-ended math problems produced by students who speak and write in non-standard English. Specifically, we look at four features of African American Vernacular English (AAVE), which range in the degree to which they are unique to AAVE or are common in other non-standard dialects. We then compare the scoring of answers that were produced by students using these dialect features to a control group of synthetic data--where we converted all non-standard dialect features to standard English. Results show that minor changes in the number of dialect features per student response do not impact GPTs automated scoring, but prompt engineering efforts did. [For the complete proceedings, see ED675583.] |
| Abstractor: | As Provided |
| Entry Date: | 2025 |
| Accession Number: | ED675646 |
| Database: | ERIC |
| Abstract: | Recent advances in AI have opened the door for the automated scoring of open-ended math problems, which were previously much more difficult to assess at scale. However, we know that biases still remain in some of these algorithms. For example, recent research on the automated scoring of student essays has shown that certain varieties of English are more strongly penalized for nonstandard English than they are for other differences that reduce the quality of students' writing. This study examines that issue in a new domain, investigating the potential for large language models to accurately grade open-ended math problems produced by students who speak and write in non-standard English. Specifically, we look at four features of African American Vernacular English (AAVE), which range in the degree to which they are unique to AAVE or are common in other non-standard dialects. We then compare the scoring of answers that were produced by students using these dialect features to a control group of synthetic data--where we converted all non-standard dialect features to standard English. Results show that minor changes in the number of dialect features per student response do not impact GPTs automated scoring, but prompt engineering efforts did. [For the complete proceedings, see ED675583.] |
|---|