Automating Self-Affirmation Essay Coding: Fine-Tuned BERT Performance Comparable to Human Coders and Comparison with GPT-4
Saved in:
| Title: | Automating Self-Affirmation Essay Coding: Fine-Tuned BERT Performance Comparable to Human Coders and Comparison with GPT-4 |
|---|---|
| Language: | English |
| Authors: | Cong Ye, Trisha H. Borman, Geoffrey D. Borman |
| Source: | Journal of Educational Data Mining. 2026 18(1):66-88. |
| Availability: | International Educational Data Mining. e-mail: jedm.editor@gmail.com; Web site: https://jedm.educationaldatamining.org/index.php/JEDM |
| Peer Reviewed: | Y |
| Page Count: | 23 |
| Publication Date: | 2026 |
| Sponsoring Agency: | Institute of Education Sciences (ED) |
| Contract Number: | R305A180230 |
| Document Type: | Journal Articles Reports - Research |
| Descriptors: | Automation, Artificial Intelligence, Technology Uses in Education, Coding, Essays, Accuracy, Interrater Reliability |
| ISSN: | 2157-2100 |
| Abstract: | Previous studies have demonstrated that a self-affirmation writing intervention, in which students reflect on personally important values, positively impacts students' school performance, and there is active research on this intervention. However, this research requires manual coding of students' writing exercises, and this manual coding has proved to be a time-consuming and expensive undertaking. To assist future selfaffirmation intervention studies or educators implementing the writing exercise, we employed our labeled data to fine-tune a pre-trained language model that achieves a comparable level of performance to that of human coders (Cohen's Kappa: 0.85 between machine coding and human coders as compared to 0.83 between human coders). To explore the potential of more advanced language models without requiring a large training dataset, we also evaluated OpenAI's GPT-4 in a zero-shot and few-shot classification setting. GPT-4's zeroshot predictions yield reasonable accuracy but do not reach the fine-tuned BERT model's performance or human-level agreement. Adding example essays (few-shot prompting) did not appreciably improve GPT-4's results. Our analysis also finds that the BERT model's performance is consistent across student subgroups, with minimal disparity between "stereotype-threatened" and "non-threatened" students, which are the focal groups for comparison in the self-affirmation intervention. We further demonstrate the generalizability of the fine-tuned model on an external dataset collected by a different research team: the model maintained a high agreement with human coders (Cohen's Kappa = 0.86) on this new sample. These results suggest that a finetuned transformer model can reliably code self-affirmation essays, thereby reducing the coding burden for future researchers and educators. We make the fine-tuned model publicly available to help the research community automate the burdensome task of coding at https://github.com/visortown/bert-self-affirm. |
| Abstractor: | As Provided |
| IES Funded: | Yes |
| Entry Date: | 2026 |
| Accession Number: | EJ1506380 |
| Database: | ERIC |
| FullText | Text: Availability: 0 CustomLinks: – Url: https://eric.ed.gov/contentdelivery/servlet/ERICServlet?accno=EJ1506380 Name: ERIC Full Text Category: fullText Text: Full Text from ERIC |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1506380 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Automating Self-Affirmation Essay Coding: Fine-Tuned BERT Performance Comparable to Human Coders and Comparison with GPT-4 – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Cong+Ye%22">Cong Ye</searchLink><br /><searchLink fieldCode="AR" term="%22Trisha+H%2E+Borman%22">Trisha H. Borman</searchLink><br /><searchLink fieldCode="AR" term="%22Geoffrey+D%2E+Borman%22">Geoffrey D. Borman</searchLink> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Journal+of+Educational+Data+Mining%22"><i>Journal of Educational Data Mining</i></searchLink>. 2026 18(1):66-88. – Name: Avail Label: Availability Group: Avail Data: International Educational Data Mining. e-mail: jedm.editor@gmail.com; Web site: https://jedm.educationaldatamining.org/index.php/JEDM – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 23 – Name: DatePubCY Label: Publication Date Group: Date Data: 2026 – Name: SourceSuprt Label: Sponsoring Agency Group: SrcSuprt Data: Institute of Education Sciences (ED) – Name: NumberContract Label: Contract Number Group: NumCntrct Data: R305A180230 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Automation%22">Automation</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Technology+Uses+in+Education%22">Technology Uses in Education</searchLink><br /><searchLink fieldCode="DE" term="%22Coding%22">Coding</searchLink><br /><searchLink fieldCode="DE" term="%22Essays%22">Essays</searchLink><br /><searchLink fieldCode="DE" term="%22Accuracy%22">Accuracy</searchLink><br /><searchLink fieldCode="DE" term="%22Interrater+Reliability%22">Interrater Reliability</searchLink> – Name: ISSN Label: ISSN Group: ISSN Data: 2157-2100 – Name: Abstract Label: Abstract Group: Ab Data: Previous studies have demonstrated that a self-affirmation writing intervention, in which students reflect on personally important values, positively impacts students' school performance, and there is active research on this intervention. However, this research requires manual coding of students' writing exercises, and this manual coding has proved to be a time-consuming and expensive undertaking. To assist future selfaffirmation intervention studies or educators implementing the writing exercise, we employed our labeled data to fine-tune a pre-trained language model that achieves a comparable level of performance to that of human coders (Cohen's Kappa: 0.85 between machine coding and human coders as compared to 0.83 between human coders). To explore the potential of more advanced language models without requiring a large training dataset, we also evaluated OpenAI's GPT-4 in a zero-shot and few-shot classification setting. GPT-4's zeroshot predictions yield reasonable accuracy but do not reach the fine-tuned BERT model's performance or human-level agreement. Adding example essays (few-shot prompting) did not appreciably improve GPT-4's results. Our analysis also finds that the BERT model's performance is consistent across student subgroups, with minimal disparity between "stereotype-threatened" and "non-threatened" students, which are the focal groups for comparison in the self-affirmation intervention. We further demonstrate the generalizability of the fine-tuned model on an external dataset collected by a different research team: the model maintained a high agreement with human coders (Cohen's Kappa = 0.86) on this new sample. These results suggest that a finetuned transformer model can reliably code self-affirmation essays, thereby reducing the coding burden for future researchers and educators. We make the fine-tuned model publicly available to help the research community automate the burdensome task of coding at https://github.com/visortown/bert-self-affirm. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: CodeSource Label: IES Funded Group: SrcInfo Data: Yes – Name: DateEntry Label: Entry Date Group: Date Data: 2026 – Name: AN Label: Accession Number Group: ID Data: EJ1506380 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1506380 |
| RecordInfo | BibRecord: BibEntity: Languages: – Text: English PhysicalDescription: Pagination: PageCount: 23 StartPage: 66 Subjects: – SubjectFull: Automation Type: general – SubjectFull: Artificial Intelligence Type: general – SubjectFull: Technology Uses in Education Type: general – SubjectFull: Coding Type: general – SubjectFull: Essays Type: general – SubjectFull: Accuracy Type: general – SubjectFull: Interrater Reliability Type: general Titles: – TitleFull: Automating Self-Affirmation Essay Coding: Fine-Tuned BERT Performance Comparable to Human Coders and Comparison with GPT-4 Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Cong Ye – PersonEntity: Name: NameFull: Trisha H. Borman – PersonEntity: Name: NameFull: Geoffrey D. Borman IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 01 Type: published Y: 2026 Identifiers: – Type: issn-electronic Value: 2157-2100 Numbering: – Type: volume Value: 18 – Type: issue Value: 1 Titles: – TitleFull: Journal of Educational Data Mining Type: main |
| ResultId | 1 |