Leveraging AI-Generated Synthetic Data to Train Natural Language Processing Models for Qualitative Feedback Analysis
Saved in:
| Title: | Leveraging AI-Generated Synthetic Data to Train Natural Language Processing Models for Qualitative Feedback Analysis |
|---|---|
| Language: | English |
| Authors: | Stephanie Fuchs (ORCID |
| Source: | Journal of Engineering Education. 2025 114(4). |
| Availability: | Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us |
| Peer Reviewed: | Y |
| Page Count: | 30 |
| Publication Date: | 2025 |
| Contract Number: | EF2222434 |
| Document Type: | Journal Articles Reports - Research |
| Descriptors: | Artificial Intelligence, Training, Data Analysis, Natural Language Processing, Feedback (Response), Student Evaluation, Engineering Education, Evaluation Methods, Models, Coding, Sentences, Accuracy, Classification |
| DOI: | 10.1002/jee.70033 |
| ISSN: | 1069-4730 2168-9830 |
| Abstract: | Background: High-quality feedback is crucial for academic success, driving student motivation and engagement while research explores effective delivery and student interactions. Advances in artificial intelligence (AI), particularly natural language processing (NLP), offer innovative methods for analyzing complex qualitative data such as feedback interactions. Purpose: We developed a framework to train sentence transformers using generative AI--created synthetic data to categorize student-feedback interactions in engineering studios. We compared traditional thematic analysis with modern methods to evaluate the realism of synthetic datasets and their effectiveness in training NLP models by exploring how generative AI can aid qualitative coding. Methods: We deidentified and transcribed eight audio recordings from engineering studios. Synthetic feedback transcripts were generated using three locally hosted large language models: Llama 3.1, Gemma 2.0, and Mistral NeMo, adjusting parameters to produce datasets mimicking the real transcripts. We assessed the quality of synthetic transcripts using our framework and used a sentence transformer model (trained on both real and synthetic data) to compare changes in the model's percent accuracy when qualitatively coding feedback interactions. Results: Synthetic data improved the NLP model's performance in classifying feedback interactions, boosting the average accuracy from 68.4% to 81% with Llama 3.1. Although incorporating synthetic data improved classification, all models produced transcripts that occasionally included extraneous details and failed to capture instructor-dominant discourse. Conclusions: Synthetic data offers an opportunity to expand qualitative research, particularly in contexts where real data for NLP training is limited or hard to obtain; however, transparency in its use is paramount to maintain research integrity. |
| Abstractor: | As Provided |
| Entry Date: | 2025 |
| Accession Number: | EJ1487745 |
| Database: | ERIC |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1487745 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Leveraging AI-Generated Synthetic Data to Train Natural Language Processing Models for Qualitative Feedback Analysis – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Stephanie+Fuchs%22">Stephanie Fuchs</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-9264-3094">0000-0001-9264-3094</externalLink>)<br /><searchLink fieldCode="AR" term="%22Alexandra+Werth%22">Alexandra Werth</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-0310-2654">0000-0003-0310-2654</externalLink>)<br /><searchLink fieldCode="AR" term="%22Cristóbal+Méndez%22">Cristóbal Méndez</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-1257-6707">0000-0002-1257-6707</externalLink>)<br /><searchLink fieldCode="AR" term="%22Jonathan+Butcher%22">Jonathan Butcher</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-9309-6296">0000-0002-9309-6296</externalLink>) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Journal+of+Engineering+Education%22"><i>Journal of Engineering Education</i></searchLink>. 2025 114(4). – Name: Avail Label: Availability Group: Avail Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 30 – Name: DatePubCY Label: Publication Date Group: Date Data: 2025 – Name: NumberContract Label: Contract Number Group: NumCntrct Data: EF2222434 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Training%22">Training</searchLink><br /><searchLink fieldCode="DE" term="%22Data+Analysis%22">Data Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+Language+Processing%22">Natural Language Processing</searchLink><br /><searchLink fieldCode="DE" term="%22Feedback+%28Response%29%22">Feedback (Response)</searchLink><br /><searchLink fieldCode="DE" term="%22Student+Evaluation%22">Student Evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Engineering+Education%22">Engineering Education</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluation+Methods%22">Evaluation Methods</searchLink><br /><searchLink fieldCode="DE" term="%22Models%22">Models</searchLink><br /><searchLink fieldCode="DE" term="%22Coding%22">Coding</searchLink><br /><searchLink fieldCode="DE" term="%22Sentences%22">Sentences</searchLink><br /><searchLink fieldCode="DE" term="%22Accuracy%22">Accuracy</searchLink><br /><searchLink fieldCode="DE" term="%22Classification%22">Classification</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1002/jee.70033 – Name: ISSN Label: ISSN Group: ISSN Data: 1069-4730<br />2168-9830 – Name: Abstract Label: Abstract Group: Ab Data: Background: High-quality feedback is crucial for academic success, driving student motivation and engagement while research explores effective delivery and student interactions. Advances in artificial intelligence (AI), particularly natural language processing (NLP), offer innovative methods for analyzing complex qualitative data such as feedback interactions. Purpose: We developed a framework to train sentence transformers using generative AI--created synthetic data to categorize student-feedback interactions in engineering studios. We compared traditional thematic analysis with modern methods to evaluate the realism of synthetic datasets and their effectiveness in training NLP models by exploring how generative AI can aid qualitative coding. Methods: We deidentified and transcribed eight audio recordings from engineering studios. Synthetic feedback transcripts were generated using three locally hosted large language models: Llama 3.1, Gemma 2.0, and Mistral NeMo, adjusting parameters to produce datasets mimicking the real transcripts. We assessed the quality of synthetic transcripts using our framework and used a sentence transformer model (trained on both real and synthetic data) to compare changes in the model's percent accuracy when qualitatively coding feedback interactions. Results: Synthetic data improved the NLP model's performance in classifying feedback interactions, boosting the average accuracy from 68.4% to 81% with Llama 3.1. Although incorporating synthetic data improved classification, all models produced transcripts that occasionally included extraneous details and failed to capture instructor-dominant discourse. Conclusions: Synthetic data offers an opportunity to expand qualitative research, particularly in contexts where real data for NLP training is limited or hard to obtain; however, transparency in its use is paramount to maintain research integrity. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2025 – Name: AN Label: Accession Number Group: ID Data: EJ1487745 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1487745 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1002/jee.70033 Languages: – Text: English PhysicalDescription: Pagination: PageCount: 30 Subjects: – SubjectFull: Artificial Intelligence Type: general – SubjectFull: Training Type: general – SubjectFull: Data Analysis Type: general – SubjectFull: Natural Language Processing Type: general – SubjectFull: Feedback (Response) Type: general – SubjectFull: Student Evaluation Type: general – SubjectFull: Engineering Education Type: general – SubjectFull: Evaluation Methods Type: general – SubjectFull: Models Type: general – SubjectFull: Coding Type: general – SubjectFull: Sentences Type: general – SubjectFull: Accuracy Type: general – SubjectFull: Classification Type: general Titles: – TitleFull: Leveraging AI-Generated Synthetic Data to Train Natural Language Processing Models for Qualitative Feedback Analysis Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Stephanie Fuchs – PersonEntity: Name: NameFull: Alexandra Werth – PersonEntity: Name: NameFull: Cristóbal Méndez – PersonEntity: Name: NameFull: Jonathan Butcher IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 10 Type: published Y: 2025 Identifiers: – Type: issn-print Value: 1069-4730 – Type: issn-electronic Value: 2168-9830 Numbering: – Type: volume Value: 114 – Type: issue Value: 4 Titles: – TitleFull: Journal of Engineering Education Type: main |
| ResultId | 1 |