Automating Self-Affirmation Essay Coding: Fine-Tuned BERT Performance Comparable to Human Coders and Comparison with GPT-4

Saved in:
Bibliographic Details
Title: Automating Self-Affirmation Essay Coding: Fine-Tuned BERT Performance Comparable to Human Coders and Comparison with GPT-4
Language: English
Authors: Cong Ye, Trisha H. Borman, Geoffrey D. Borman
Source: Journal of Educational Data Mining. 2026 18(1):66-88.
Availability: International Educational Data Mining. e-mail: jedm.editor@gmail.com; Web site: https://jedm.educationaldatamining.org/index.php/JEDM
Peer Reviewed: Y
Page Count: 23
Publication Date: 2026
Sponsoring Agency: Institute of Education Sciences (ED)
Contract Number: R305A180230
Document Type: Journal Articles
Reports - Research
Descriptors: Automation, Artificial Intelligence, Technology Uses in Education, Coding, Essays, Accuracy, Interrater Reliability
ISSN: 2157-2100
Abstract: Previous studies have demonstrated that a self-affirmation writing intervention, in which students reflect on personally important values, positively impacts students' school performance, and there is active research on this intervention. However, this research requires manual coding of students' writing exercises, and this manual coding has proved to be a time-consuming and expensive undertaking. To assist future selfaffirmation intervention studies or educators implementing the writing exercise, we employed our labeled data to fine-tune a pre-trained language model that achieves a comparable level of performance to that of human coders (Cohen's Kappa: 0.85 between machine coding and human coders as compared to 0.83 between human coders). To explore the potential of more advanced language models without requiring a large training dataset, we also evaluated OpenAI's GPT-4 in a zero-shot and few-shot classification setting. GPT-4's zeroshot predictions yield reasonable accuracy but do not reach the fine-tuned BERT model's performance or human-level agreement. Adding example essays (few-shot prompting) did not appreciably improve GPT-4's results. Our analysis also finds that the BERT model's performance is consistent across student subgroups, with minimal disparity between "stereotype-threatened" and "non-threatened" students, which are the focal groups for comparison in the self-affirmation intervention. We further demonstrate the generalizability of the fine-tuned model on an external dataset collected by a different research team: the model maintained a high agreement with human coders (Cohen's Kappa = 0.86) on this new sample. These results suggest that a finetuned transformer model can reliably code self-affirmation essays, thereby reducing the coding burden for future researchers and educators. We make the fine-tuned model publicly available to help the research community automate the burdensome task of coding at https://github.com/visortown/bert-self-affirm.
Abstractor: As Provided
IES Funded: Yes
Entry Date: 2026
Accession Number: EJ1506380
Database: ERIC
FullText Text:
  Availability: 0
CustomLinks:
  – Url: https://eric.ed.gov/contentdelivery/servlet/ERICServlet?accno=EJ1506380
    Name: ERIC Full Text
    Category: fullText
    Text: Full Text from ERIC
Header DbId: eric
DbLabel: ERIC
An: EJ1506380
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Automating Self-Affirmation Essay Coding: Fine-Tuned BERT Performance Comparable to Human Coders and Comparison with GPT-4
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Cong+Ye%22">Cong Ye</searchLink><br /><searchLink fieldCode="AR" term="%22Trisha+H%2E+Borman%22">Trisha H. Borman</searchLink><br /><searchLink fieldCode="AR" term="%22Geoffrey+D%2E+Borman%22">Geoffrey D. Borman</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Journal+of+Educational+Data+Mining%22"><i>Journal of Educational Data Mining</i></searchLink>. 2026 18(1):66-88.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: International Educational Data Mining. e-mail: jedm.editor@gmail.com; Web site: https://jedm.educationaldatamining.org/index.php/JEDM
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 23
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2026
– Name: SourceSuprt
  Label: Sponsoring Agency
  Group: SrcSuprt
  Data: Institute of Education Sciences (ED)
– Name: NumberContract
  Label: Contract Number
  Group: NumCntrct
  Data: R305A180230
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Automation%22">Automation</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Technology+Uses+in+Education%22">Technology Uses in Education</searchLink><br /><searchLink fieldCode="DE" term="%22Coding%22">Coding</searchLink><br /><searchLink fieldCode="DE" term="%22Essays%22">Essays</searchLink><br /><searchLink fieldCode="DE" term="%22Accuracy%22">Accuracy</searchLink><br /><searchLink fieldCode="DE" term="%22Interrater+Reliability%22">Interrater Reliability</searchLink>
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 2157-2100
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Previous studies have demonstrated that a self-affirmation writing intervention, in which students reflect on personally important values, positively impacts students' school performance, and there is active research on this intervention. However, this research requires manual coding of students' writing exercises, and this manual coding has proved to be a time-consuming and expensive undertaking. To assist future selfaffirmation intervention studies or educators implementing the writing exercise, we employed our labeled data to fine-tune a pre-trained language model that achieves a comparable level of performance to that of human coders (Cohen's Kappa: 0.85 between machine coding and human coders as compared to 0.83 between human coders). To explore the potential of more advanced language models without requiring a large training dataset, we also evaluated OpenAI's GPT-4 in a zero-shot and few-shot classification setting. GPT-4's zeroshot predictions yield reasonable accuracy but do not reach the fine-tuned BERT model's performance or human-level agreement. Adding example essays (few-shot prompting) did not appreciably improve GPT-4's results. Our analysis also finds that the BERT model's performance is consistent across student subgroups, with minimal disparity between "stereotype-threatened" and "non-threatened" students, which are the focal groups for comparison in the self-affirmation intervention. We further demonstrate the generalizability of the fine-tuned model on an external dataset collected by a different research team: the model maintained a high agreement with human coders (Cohen's Kappa = 0.86) on this new sample. These results suggest that a finetuned transformer model can reliably code self-affirmation essays, thereby reducing the coding burden for future researchers and educators. We make the fine-tuned model publicly available to help the research community automate the burdensome task of coding at https://github.com/visortown/bert-self-affirm.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: CodeSource
  Label: IES Funded
  Group: SrcInfo
  Data: Yes
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2026
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1506380
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1506380
RecordInfo BibRecord:
  BibEntity:
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 23
        StartPage: 66
    Subjects:
      – SubjectFull: Automation
        Type: general
      – SubjectFull: Artificial Intelligence
        Type: general
      – SubjectFull: Technology Uses in Education
        Type: general
      – SubjectFull: Coding
        Type: general
      – SubjectFull: Essays
        Type: general
      – SubjectFull: Accuracy
        Type: general
      – SubjectFull: Interrater Reliability
        Type: general
    Titles:
      – TitleFull: Automating Self-Affirmation Essay Coding: Fine-Tuned BERT Performance Comparable to Human Coders and Comparison with GPT-4
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Cong Ye
      – PersonEntity:
          Name:
            NameFull: Trisha H. Borman
      – PersonEntity:
          Name:
            NameFull: Geoffrey D. Borman
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-electronic
              Value: 2157-2100
          Numbering:
            – Type: volume
              Value: 18
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Journal of Educational Data Mining
              Type: main
ResultId 1