Leveraging AI-Generated Synthetic Data to Train Natural Language Processing Models for Qualitative Feedback Analysis

Saved in:
Bibliographic Details
Title: Leveraging AI-Generated Synthetic Data to Train Natural Language Processing Models for Qualitative Feedback Analysis
Language: English
Authors: Stephanie Fuchs (ORCID 0000-0001-9264-3094), Alexandra Werth (ORCID 0000-0003-0310-2654), Cristóbal Méndez (ORCID 0000-0002-1257-6707), Jonathan Butcher (ORCID 0000-0002-9309-6296)
Source: Journal of Engineering Education. 2025 114(4).
Availability: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 30
Publication Date: 2025
Contract Number: EF2222434
Document Type: Journal Articles
Reports - Research
Descriptors: Artificial Intelligence, Training, Data Analysis, Natural Language Processing, Feedback (Response), Student Evaluation, Engineering Education, Evaluation Methods, Models, Coding, Sentences, Accuracy, Classification
DOI: 10.1002/jee.70033
ISSN: 1069-4730
2168-9830
Abstract: Background: High-quality feedback is crucial for academic success, driving student motivation and engagement while research explores effective delivery and student interactions. Advances in artificial intelligence (AI), particularly natural language processing (NLP), offer innovative methods for analyzing complex qualitative data such as feedback interactions. Purpose: We developed a framework to train sentence transformers using generative AI--created synthetic data to categorize student-feedback interactions in engineering studios. We compared traditional thematic analysis with modern methods to evaluate the realism of synthetic datasets and their effectiveness in training NLP models by exploring how generative AI can aid qualitative coding. Methods: We deidentified and transcribed eight audio recordings from engineering studios. Synthetic feedback transcripts were generated using three locally hosted large language models: Llama 3.1, Gemma 2.0, and Mistral NeMo, adjusting parameters to produce datasets mimicking the real transcripts. We assessed the quality of synthetic transcripts using our framework and used a sentence transformer model (trained on both real and synthetic data) to compare changes in the model's percent accuracy when qualitatively coding feedback interactions. Results: Synthetic data improved the NLP model's performance in classifying feedback interactions, boosting the average accuracy from 68.4% to 81% with Llama 3.1. Although incorporating synthetic data improved classification, all models produced transcripts that occasionally included extraneous details and failed to capture instructor-dominant discourse. Conclusions: Synthetic data offers an opportunity to expand qualitative research, particularly in contexts where real data for NLP training is limited or hard to obtain; however, transparency in its use is paramount to maintain research integrity.
Abstractor: As Provided
Entry Date: 2025
Accession Number: EJ1487745
Database: ERIC
FullText Text:
  Availability: 0
Header DbId: eric
DbLabel: ERIC
An: EJ1487745
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Leveraging AI-Generated Synthetic Data to Train Natural Language Processing Models for Qualitative Feedback Analysis
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Stephanie+Fuchs%22">Stephanie Fuchs</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-9264-3094">0000-0001-9264-3094</externalLink>)<br /><searchLink fieldCode="AR" term="%22Alexandra+Werth%22">Alexandra Werth</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-0310-2654">0000-0003-0310-2654</externalLink>)<br /><searchLink fieldCode="AR" term="%22Cristóbal+Méndez%22">Cristóbal Méndez</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-1257-6707">0000-0002-1257-6707</externalLink>)<br /><searchLink fieldCode="AR" term="%22Jonathan+Butcher%22">Jonathan Butcher</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-9309-6296">0000-0002-9309-6296</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Journal+of+Engineering+Education%22"><i>Journal of Engineering Education</i></searchLink>. 2025 114(4).
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 30
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2025
– Name: NumberContract
  Label: Contract Number
  Group: NumCntrct
  Data: EF2222434
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Training%22">Training</searchLink><br /><searchLink fieldCode="DE" term="%22Data+Analysis%22">Data Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+Language+Processing%22">Natural Language Processing</searchLink><br /><searchLink fieldCode="DE" term="%22Feedback+%28Response%29%22">Feedback (Response)</searchLink><br /><searchLink fieldCode="DE" term="%22Student+Evaluation%22">Student Evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Engineering+Education%22">Engineering Education</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluation+Methods%22">Evaluation Methods</searchLink><br /><searchLink fieldCode="DE" term="%22Models%22">Models</searchLink><br /><searchLink fieldCode="DE" term="%22Coding%22">Coding</searchLink><br /><searchLink fieldCode="DE" term="%22Sentences%22">Sentences</searchLink><br /><searchLink fieldCode="DE" term="%22Accuracy%22">Accuracy</searchLink><br /><searchLink fieldCode="DE" term="%22Classification%22">Classification</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1002/jee.70033
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 1069-4730<br />2168-9830
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Background: High-quality feedback is crucial for academic success, driving student motivation and engagement while research explores effective delivery and student interactions. Advances in artificial intelligence (AI), particularly natural language processing (NLP), offer innovative methods for analyzing complex qualitative data such as feedback interactions. Purpose: We developed a framework to train sentence transformers using generative AI--created synthetic data to categorize student-feedback interactions in engineering studios. We compared traditional thematic analysis with modern methods to evaluate the realism of synthetic datasets and their effectiveness in training NLP models by exploring how generative AI can aid qualitative coding. Methods: We deidentified and transcribed eight audio recordings from engineering studios. Synthetic feedback transcripts were generated using three locally hosted large language models: Llama 3.1, Gemma 2.0, and Mistral NeMo, adjusting parameters to produce datasets mimicking the real transcripts. We assessed the quality of synthetic transcripts using our framework and used a sentence transformer model (trained on both real and synthetic data) to compare changes in the model's percent accuracy when qualitatively coding feedback interactions. Results: Synthetic data improved the NLP model's performance in classifying feedback interactions, boosting the average accuracy from 68.4% to 81% with Llama 3.1. Although incorporating synthetic data improved classification, all models produced transcripts that occasionally included extraneous details and failed to capture instructor-dominant discourse. Conclusions: Synthetic data offers an opportunity to expand qualitative research, particularly in contexts where real data for NLP training is limited or hard to obtain; however, transparency in its use is paramount to maintain research integrity.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2025
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1487745
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1487745
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1002/jee.70033
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 30
    Subjects:
      – SubjectFull: Artificial Intelligence
        Type: general
      – SubjectFull: Training
        Type: general
      – SubjectFull: Data Analysis
        Type: general
      – SubjectFull: Natural Language Processing
        Type: general
      – SubjectFull: Feedback (Response)
        Type: general
      – SubjectFull: Student Evaluation
        Type: general
      – SubjectFull: Engineering Education
        Type: general
      – SubjectFull: Evaluation Methods
        Type: general
      – SubjectFull: Models
        Type: general
      – SubjectFull: Coding
        Type: general
      – SubjectFull: Sentences
        Type: general
      – SubjectFull: Accuracy
        Type: general
      – SubjectFull: Classification
        Type: general
    Titles:
      – TitleFull: Leveraging AI-Generated Synthetic Data to Train Natural Language Processing Models for Qualitative Feedback Analysis
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Stephanie Fuchs
      – PersonEntity:
          Name:
            NameFull: Alexandra Werth
      – PersonEntity:
          Name:
            NameFull: Cristóbal Méndez
      – PersonEntity:
          Name:
            NameFull: Jonathan Butcher
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 10
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 1069-4730
            – Type: issn-electronic
              Value: 2168-9830
          Numbering:
            – Type: volume
              Value: 114
            – Type: issue
              Value: 4
          Titles:
            – TitleFull: Journal of Engineering Education
              Type: main
ResultId 1