Expanding Possibilities for Generative AI in Qualitative Analysis: Fostering Student Feedback Literacy through the Application of a Feedback Quality Rubric

Saved in:
Bibliographic Details
Title: Expanding Possibilities for Generative AI in Qualitative Analysis: Fostering Student Feedback Literacy through the Application of a Feedback Quality Rubric
Language: English
Authors: Katherine Drinkwater Gregg (ORCID 0009-0002-5998-9231), Olivia Ryan (ORCID 0009-0007-7981-6131), Andrew Katz (ORCID 0000-0002-3554-9015), Mark Huerta (ORCID 0000-0003-2962-0724), Susan Sajadi (ORCID 0000-0001-8511-7467)
Source: Journal of Engineering Education. 2025 114(3).
Availability: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 31
Publication Date: 2025
Document Type: Journal Articles
Reports - Research
Education Level: Higher Education
Postsecondary Education
Descriptors: Artificial Intelligence, Technology Uses in Education, Engineering Education, Student Evaluation, Peer Evaluation, Teamwork, Formative Evaluation, Feedback (Response), Natural Language Processing, Scoring Rubrics, College Freshmen, Computer Mediated Communication, Evaluators, Man Machine Systems, Interrater Reliability, Literacy
DOI: 10.1002/jee.70024
ISSN: 1069-4730
2168-9830
Abstract: Background: Courses in engineering often use peer evaluation to monitor teamwork behaviors and team dynamics. The qualitative peer comments written for peer evaluations hold potential as a valuable source of formative feedback for students, yet little is known about their content and quality. Purpose: This study uses a large language model (LLM) to apply a previously tested feedback quality rubric to peer feedback comments. Our research questions interrogate the reliability of LLMs for qualitative analysis with a rubric and use Bandura's self-regulated learning theory to assess peer feedback quality of first-year engineering students' comments. Method: An open-source, local LLM was used to score each comment according to four rubric criteria. Inter-rater reliability (IRR) with human raters using Cohen's quadratic weighted kappa was the primary metric of reliability. Our assessment of peer feedback quality utilized descriptive statistics. Results: The LLM achieved lower IRR than human raters, but the model's challenges mimic those of human raters. The model did achieve an excellent quadratic weighted kappa of 0.80 for one rubric criterion, which shows promise for LLM capability. For feedback quality, students generally wrote low- to medium-quality comments that were infrequently grounded in specific teamwork behaviors. We identified five types of peer feedback that inform how students perceive the feedback process. Conclusions: Our implementation of GAI suggests that LLMs can be helpful for rapid iteration of research designs, but consistent and reliable analysis with generative artificial intelligence (GAI) requires significant effort and testing. To develop feedback literacy, students must understand how to provide high-quality feedback.
Abstractor: As Provided
Entry Date: 2025
Accession Number: EJ1478628
Database: ERIC
Full text is not displayed to guests.
Description
Abstract:Background: Courses in engineering often use peer evaluation to monitor teamwork behaviors and team dynamics. The qualitative peer comments written for peer evaluations hold potential as a valuable source of formative feedback for students, yet little is known about their content and quality. Purpose: This study uses a large language model (LLM) to apply a previously tested feedback quality rubric to peer feedback comments. Our research questions interrogate the reliability of LLMs for qualitative analysis with a rubric and use Bandura's self-regulated learning theory to assess peer feedback quality of first-year engineering students' comments. Method: An open-source, local LLM was used to score each comment according to four rubric criteria. Inter-rater reliability (IRR) with human raters using Cohen's quadratic weighted kappa was the primary metric of reliability. Our assessment of peer feedback quality utilized descriptive statistics. Results: The LLM achieved lower IRR than human raters, but the model's challenges mimic those of human raters. The model did achieve an excellent quadratic weighted kappa of 0.80 for one rubric criterion, which shows promise for LLM capability. For feedback quality, students generally wrote low- to medium-quality comments that were infrequently grounded in specific teamwork behaviors. We identified five types of peer feedback that inform how students perceive the feedback process. Conclusions: Our implementation of GAI suggests that LLMs can be helpful for rapid iteration of research designs, but consistent and reliable analysis with generative artificial intelligence (GAI) requires significant effort and testing. To develop feedback literacy, students must understand how to provide high-quality feedback.
ISSN:1069-4730
2168-9830
DOI:10.1002/jee.70024