Expanding Possibilities for Generative AI in Qualitative Analysis: Fostering Student Feedback Literacy through the Application of a Feedback Quality Rubric

Saved in:
Bibliographic Details
Title: Expanding Possibilities for Generative AI in Qualitative Analysis: Fostering Student Feedback Literacy through the Application of a Feedback Quality Rubric
Language: English
Authors: Katherine Drinkwater Gregg (ORCID 0009-0002-5998-9231), Olivia Ryan (ORCID 0009-0007-7981-6131), Andrew Katz (ORCID 0000-0002-3554-9015), Mark Huerta (ORCID 0000-0003-2962-0724), Susan Sajadi (ORCID 0000-0001-8511-7467)
Source: Journal of Engineering Education. 2025 114(3).
Availability: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 31
Publication Date: 2025
Document Type: Journal Articles
Reports - Research
Education Level: Higher Education
Postsecondary Education
Descriptors: Artificial Intelligence, Technology Uses in Education, Engineering Education, Student Evaluation, Peer Evaluation, Teamwork, Formative Evaluation, Feedback (Response), Natural Language Processing, Scoring Rubrics, College Freshmen, Computer Mediated Communication, Evaluators, Man Machine Systems, Interrater Reliability, Literacy
DOI: 10.1002/jee.70024
ISSN: 1069-4730
2168-9830
Abstract: Background: Courses in engineering often use peer evaluation to monitor teamwork behaviors and team dynamics. The qualitative peer comments written for peer evaluations hold potential as a valuable source of formative feedback for students, yet little is known about their content and quality. Purpose: This study uses a large language model (LLM) to apply a previously tested feedback quality rubric to peer feedback comments. Our research questions interrogate the reliability of LLMs for qualitative analysis with a rubric and use Bandura's self-regulated learning theory to assess peer feedback quality of first-year engineering students' comments. Method: An open-source, local LLM was used to score each comment according to four rubric criteria. Inter-rater reliability (IRR) with human raters using Cohen's quadratic weighted kappa was the primary metric of reliability. Our assessment of peer feedback quality utilized descriptive statistics. Results: The LLM achieved lower IRR than human raters, but the model's challenges mimic those of human raters. The model did achieve an excellent quadratic weighted kappa of 0.80 for one rubric criterion, which shows promise for LLM capability. For feedback quality, students generally wrote low- to medium-quality comments that were infrequently grounded in specific teamwork behaviors. We identified five types of peer feedback that inform how students perceive the feedback process. Conclusions: Our implementation of GAI suggests that LLMs can be helpful for rapid iteration of research designs, but consistent and reliable analysis with generative artificial intelligence (GAI) requires significant effort and testing. To develop feedback literacy, students must understand how to provide high-quality feedback.
Abstractor: As Provided
Entry Date: 2025
Accession Number: EJ1478628
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwFcwxZYLTBym6IPdEjIbXXjAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDA3rWv803SK13HNc2wIBEICBm8PaQSYbNLn_KnaNbxk-jFGn0CWcS3btcrAeuhiRXFKlvWAtgxqH-NT3MpM-nQEoroMvWh46s5AI7mXVm6iPszkwG5rxp5vLqoqiPbm4cNqix0eawUvUeiQF5To9-vFY6iIl9PYZCfsGegyFD-GpdUPTVUiOdVCY4sLtR2MJjT6XlJSTqbl_VeP-A0SbBl7uDuYOg6vx9LVtV60C
Text:
  Availability: 1
  Value: <anid>AN0186992462;6m401jul.25;2025Jul31.06:28;v2.2.500</anid> <title id="AN0186992462-1">Expanding possibilities for generative AI in qualitative analysis: Fostering student feedback literacy through the application of a feedback quality rubric </title> <p>Background: Courses in engineering often use peer evaluation to monitor teamwork behaviors and team dynamics. The qualitative peer comments written for peer evaluations hold potential as a valuable source of formative feedback for students, yet little is known about their content and quality. Purpose: This study uses a large language model (LLM) to apply a previously tested feedback quality rubric to peer feedback comments. Our research questions interrogate the reliability of LLMs for qualitative analysis with a rubric and use Bandura's self‐regulated learning theory to assess peer feedback quality of first‐year engineering students' comments. Method: An open‐source, local LLM was used to score each comment according to four rubric criteria. Inter‐rater reliability (IRR) with human raters using Cohen's quadratic weighted kappa was the primary metric of reliability. Our assessment of peer feedback quality utilized descriptive statistics. Results: The LLM achieved lower IRR than human raters, but the model's challenges mimic those of human raters. The model did achieve an excellent quadratic weighted kappa of 0.80 for one rubric criterion, which shows promise for LLM capability. For feedback quality, students generally wrote low‐ to medium‐quality comments that were infrequently grounded in specific teamwork behaviors. We identified five types of peer feedback that inform how students perceive the feedback process. Conclusions: Our implementation of GAI suggests that LLMs can be helpful for rapid iteration of research designs, but consistent and reliable analysis with generative artificial intelligence (GAI) requires significant effort and testing. To develop feedback literacy, students must understand how to provide high‐quality feedback.</p> <p>Keywords: feedback; generative AI; rubric</p> <hd id="AN0186992462-2">INTRODUCTION</hd> <p>In recent years, most fields have rushed to understand how generative artificial intelligence (GAI) will impact future workflows. Researchers have proposed integrations of GAI in the qualitative research process from study design to data collection to analysis and beyond (Owoahene Acheampong & Nyaaba, [<reflink idref="bib47" id="ref1">47</reflink>]). Most of the AI tools used in research have been large language models (LLMs), AI models that generate textual information. LLMs have been used for code suggestions and thematic analysis (Gamieldien et al., [<reflink idref="bib24" id="ref2">24</reflink>]; Rietz & Maedche, [<reflink idref="bib49" id="ref3">49</reflink>]). However, procedures to determine the reliability of GAI tools when performing research tasks are largely undetermined (Holmes & Miao, [<reflink idref="bib30" id="ref4">30</reflink>]). Additional scholarship is needed to understand the capabilities of GAI and establish methods of verifying the quality of qualitative research.</p> <p>To understand the capabilities of LLMs in particular, this study focuses on its application to support engineering education research (EER). In education research, there is a capacity to collect large datasets (Abram et al., [<reflink idref="bib3" id="ref5">3</reflink>]) with methods such as publicly available data, surveys that include qualitative responses, and course materials. For large qualitative datasets, quality measures such as inter‐rater reliability require multiple coders, which can be time consuming and resource intensive (Cruz et al., [<reflink idref="bib16" id="ref6">16</reflink>]). The time required for quality analysis cannot keep up with the quantity of qualitative data researchers can collect, often leaving rich data unanalyzed. The recent emergence of easily accessible LLMs provides an opportunity to innovate qualitative analysis methods to reduce the necessary human resources and time; however, more research is required to understand their current capabilities and limitations for use in qualitative research. We identified an exemplary use case of assessing peer feedback quality to test the effectiveness of LLMs as a supplementary analysis mechanism for qualitative data in EER.</p> <p>Feedback literacy, or the capacity to make sense of feedback information and use it to enhance performance (Carless & Boud, [<reflink idref="bib10" id="ref7">10</reflink>]), is an essential competency for engineers (Coppens et al., [<reflink idref="bib14" id="ref8">14</reflink>]). Feedback comments from peer evaluation tools such as CATME, a web‐based platform that helps instructors form teams and facilitate peer feedback (Ohland et al., [<reflink idref="bib45" id="ref9">45</reflink>]), are suitable artifacts that can be qualitatively analyzed to measure how students provide feedback, which is typically assessed through reflections (Coppens et al., [<reflink idref="bib15" id="ref10">15</reflink>]) and survey instruments (Zhan, [<reflink idref="bib61" id="ref11">61</reflink>]). We have proposed a new method of assessing students' ability to provide feedback—an important part of feedback literacy (Dawson et al., [<reflink idref="bib18" id="ref12">18</reflink>])—by measuring improvements in the quality of written peer feedback comments using a rubric adapted from medical education (Drinkwater et al., [<reflink idref="bib20" id="ref13">20</reflink>]). Rubrics are a common method of analyzing qualitative data because they can guide raters to minimize subjectivity and can enable researchers to quantify qualitative data to systematically explore trends in their data (Cruz et al., [<reflink idref="bib16" id="ref14">16</reflink>]). Previous studies have suggested that providing LLMs with scoped coding tasks, such as applying a rubric, can yield more reliable results than inductive coding (Gamieldien et al., [<reflink idref="bib24" id="ref15">24</reflink>]), but we are unaware of any studies in EER that have used an LLM to apply a rubric to qualitative data of real student responses. Thus, this study incorporates a growing interest among education researchers in the importance of feedback literacy and an innovative use of LLMs to better understand how to optimize researchers' capacity with valid and replicable techniques for analyzing qualitative datasets.</p> <p>This study aims to explore the potential of open‐source LLMs in applying a rubric to assess qualitative data and highlight the feedback practices of first‐year engineering students. Our study is guided by two research questions to address our methodological and theoretical purposes:</p> <p></p> <ulist> <item> How reliable is an LLM in applying a rubric to assess the quality of engineering student peer feedback comments compared to researchers?</item> <p></p> <item> What do feedback quality measures reveal about first‐year engineering students' ability to provide feedback?</item> </ulist> <p>By addressing these research questions, we aim to (i) evaluate the effectiveness of using LLMs to apply a rubric to support research processes, and (ii) highlight aspects of first‐year engineering students' feedback literacy behaviors.</p> <hd id="AN0186992462-3">LITERATURE REVIEW</hd> <p>This literature review delves into previous research on LLMs in education research and peer feedback in engineering education.</p> <hd id="AN0186992462-4">Generative AI in education research</hd> <p>Recently, GAI has prompted questions from researchers about its use and role in research. Scholars have agreed that GAI can be used to accelerate and support research processes (Dwivedi et al., [<reflink idref="bib21" id="ref16">21</reflink>]; Liu & Jagadish, [<reflink idref="bib35" id="ref17">35</reflink>]; Morris, [<reflink idref="bib40" id="ref18">40</reflink>]). For example, LLMs can assist with data management, data analysis, and writing. Qualitative data analysis is a particularly ripe area of opportunity to apply LLMs in education research. Qualitative data analysis is a time‐consuming process when done effectively with reliability and validity measures. Therefore, there is an opportunity to utilize LLMs to assist researchers in the process. LLMs have recently been used in qualitative studies as a tool to provide code suggestions (Rietz & Maedche, [<reflink idref="bib49" id="ref19">49</reflink>]), contribute to grounded theory development (Sinha et al., [<reflink idref="bib55" id="ref20">55</reflink>]), and create a thematic analysis codebook (Gamieldien et al., [<reflink idref="bib24" id="ref21">24</reflink>]). However, these opportunities are accompanied by concerns regarding the ethics and quality of research assisted by AI. Qualitative data analysis is susceptible to demographic biases in AI models, and confidentiality and data privacy can be challenging to maintain (Fui‐Hoon Nah et al., [<reflink idref="bib23" id="ref22">23</reflink>]). The resources necessary to create and power large LLMs necessitate an ethical consideration of financial and environmental costs (Bender et al., [<reflink idref="bib6" id="ref23">6</reflink>]).</p> <p>Recent work in engineering education has shown that LLMs can be utilized to innovate the research process. For example, one paper showed how natural language processing and LLMs can be used to analyze and extract themes from a large dataset of student evaluations of teaching (Katz, Gerhardt, & Soledad, [<reflink idref="bib33" id="ref24">33</reflink>]). A very similar study utilized an open‐source GAI model to synthesize summaries of student feedback to instructors (Zhang et al., [<reflink idref="bib63" id="ref25">63</reflink>]). Other recent studies have used natural language processing to generate codebooks from a large corpus of text (Hingle et al., [<reflink idref="bib28" id="ref26">28</reflink>]) and shorter, simulated textual data (Katz, Fleming, & Main, [<reflink idref="bib32" id="ref27">32</reflink>]). These examples show that in EER, GAI can be used to analyze qualitative data thematically by summarizing it. However, little work has been done to show how GAI can be used to apply a rubric to assess qualitative data. Therefore, applying a rubric to assess qualitative comments is an area of opportunity for LLM use in qualitative research.</p> <hd id="AN0186992462-5">Feedback literacy</hd> <p>Feedback literacy is an important competency for engineering students (Coppens et al., [<reflink idref="bib14" id="ref28">14</reflink>]). Feedback literacy is defined as "the understandings, capacities and dispositions needed to make sense of information and use it to enhance work or learning strategies" (Carless & Boud, [<reflink idref="bib10" id="ref29">10</reflink>]). Once students enter the workplace, they will be required to give and receive feedback from their peers and managers, so developing feedback literacy skills during their education is critically important for their development. Feedback literacy is an emerging area of research, with recent publications suggesting definitions and development of instruments (Carless & Boud, [<reflink idref="bib10" id="ref30">10</reflink>]; Dawson et al., [<reflink idref="bib18" id="ref31">18</reflink>]; Molloy et al., [<reflink idref="bib39" id="ref32">39</reflink>]). Feedback literacy has four main features: appreciating feedback processes, making judgments, managing affect, and taking action. Appreciating feedback processes involves valuing feedback information and recognizing the student's active role in the feedback process. This step is crucial: without valuing feedback or understanding how to use it, students will lack motivation to take action. Making judgments relies on evaluative judgment; students need to assess their own performance and that of others to effectively give and receive feedback. Affect in the feature "managing affect" refers to feelings, emotions, and attitudes. Receiving constructive feedback can be challenging; students need to move beyond their initial reactions to recognize the value of the feedback they receive. The final feature of student feedback literacy is taking action. Once students interpret the feedback, they need to close the feedback loop by applying it to improve their work (Molloy & Boud, [<reflink idref="bib38" id="ref33">38</reflink>]).</p> <p>Developing students' feedback literacy skills enhances their ability to engage with feedback during their education and prepares them for the workplace, where feedback is consistently used in collaborative environments. For example, project‐based learning (PBL) courses are a common pedagogical approach for teaching engineering (Borrego et al., [<reflink idref="bib7" id="ref34">7</reflink>]) to help students develop essential professional skills such as communication (Beagon et al., [<reflink idref="bib5" id="ref35">5</reflink>]), conflict management (Ryan et al., [<reflink idref="bib50" id="ref36">50</reflink>]), and collaboration with diverse team members (L. Zhang & Ma, [<reflink idref="bib62" id="ref37">62</reflink>]). As a part of the collaborative work process, students are often required to rate and provide feedback to their teammates through a peer evaluation process. This process informs the instructor about team dynamics and helps teams to improve their dynamics and performance (Mentzer et al., [<reflink idref="bib37" id="ref38">37</reflink>]). Peer comments are a valuable opportunity for students to practice feedback literacy skills because they provide a wealth of information about team interactions, collaboration skills, and individual contributions. Further, applying tools such as generative AI to analyze peer comments can enable both researchers and instructors to gain insights into the feedback literacy of engineering students.</p> <p>Unfortunately, students tend not to have well‐developed feedback literacy skills, which is partly attributable to educators' lack of proficiency in creating optimal feedback engagement opportunities (Winstone et al., [<reflink idref="bib60" id="ref39">60</reflink>], p.18). The peer evaluation process in engineering PBL classes provides an opportunity for educators to engage students in feedback literacy behaviors; however, this is limited by the quality of feedback students provide and receive. Providing feedback is an essential feedback literacy behavior (Dawson et al., [<reflink idref="bib18" id="ref40">18</reflink>]). In the peer evaluation process in particular, providing quality feedback is critical to catalyzing further engagement in feedback literacy behaviors such as making sense of feedback and taking action. Unfortunately, it is rare for peer feedback comments to address teamwork behaviors in a detailed and actionable manner (Burgess et al., [<reflink idref="bib8" id="ref41">8</reflink>]) because engineering students, especially first‐year students, do not have training in writing effective feedback. With minimal guidance from instructors, peer feedback more often tends toward short phrases like "good teammate" or self level feedback comments (Hattie & Timperley, [<reflink idref="bib27" id="ref42">27</reflink>]) that do not relate to the team and are not actionable. These issues are worsened when feedback is summative and used for assessment, as students do not wish to negatively impact their teammates' grades (Sridharan et al., [<reflink idref="bib56" id="ref43">56</reflink>]). In order to help engineering students develop important feedback literacy skills, it is essential for educators and students to understand and apply the principles of quality feedback.</p> <p>Effective feedback has been shown to help students self‐assess their performance, which can improve team dynamics, develop self‐regulation behaviors, and support the development of other professional skills (Nicol et al., [<reflink idref="bib43" id="ref44">43</reflink>]). Studies have highlighted that, in order to attain these benefits of feedback, the feedback must be of high quality. Principles of quality feedback have been highlighted by several researchers (Hattie & Timperley, [<reflink idref="bib27" id="ref45">27</reflink>]; Shute, [<reflink idref="bib54" id="ref46">54</reflink>]). Shute ([<reflink idref="bib54" id="ref47">54</reflink>]) performed a literature review on formative feedback in education. They proposed that to enhance learning, formative feedback must concentrate on the task, be specific, be straightforward, and align with a learning objective. Similarly, Hattie and Timperley's model of feedback identifies four levels of feedback: task level, process level, self‐regulation level, and self level (Hattie & Timperley, [<reflink idref="bib27" id="ref48">27</reflink>]). The task level emphasizes the execution of the tasks being completed, while the process level outlines what is necessary to effectively carry out the tasks at hand. The self‐regulation level involves students monitoring and adjusting their actions, while the self level addresses personal evaluations by the learner.</p> <p>Improving engineering students' ability to provide high‐quality feedback would enable them to receive valuable formative feedback and have the opportunity to further develop their feedback literacy. Without high‐quality feedback, students are not given specific details on how to improve their teamwork behaviors and will not be able to take meaningful action. Additionally, without models of quality feedback, students cannot learn to effectively make or communicate judgments about the performance of others, which is a key component of feedback literacy (Carless & Boud, [<reflink idref="bib10" id="ref49">10</reflink>]). Our work addresses feedback quality as a key component of feedback literacy, a skill that is of growing interest to researchers broadly and can help engineering students develop their teamwork skills in preparation for the workplace.</p> <hd id="AN0186992462-6">Theoretical underpinnings</hd> <p>The concept of feedback literacy originated in assessment and feedback research (Chong, [<reflink idref="bib12" id="ref50">12</reflink>]). Despite theory being recognized as important to the field, empirical investigations often build on conceptualizations of assessment (Nieminen et al., [<reflink idref="bib44" id="ref51">44</reflink>]). In the case of feedback literacy, Chong ([<reflink idref="bib12" id="ref52">12</reflink>]) connects feedback literacy to sociocultural theory and learner agency to reconsider feedback literacy from an ecological perspective that includes a contextual, engagement, and individual dimension. Similarly, in our work, we focus on the "provision of feedback" component of feedback literacy (Dawson et al., [<reflink idref="bib18" id="ref53">18</reflink>]) and relate feedback literacy more broadly to learning using Bandura's theory of self‐regulation. Connecting these theoretical underpinnings to our research offers a valuable opportunity to explore specific learning environments in engineering education where feedback practices can be enhanced and make clear connections to our research questions and constructs used in our rubric.</p> <hd id="AN0186992462-7">Bandura's social cognitive theory of self‐regulation</hd> <p>Bandura's social cognitive theory of self‐regulation posits that individuals learn and regulate their behavior through an interplay between personal factors, environmental influences, and cognitive processes (Bandura, [<reflink idref="bib4" id="ref54">4</reflink>]). Self‐regulation, a central concept of this theory, involves goal‐setting, progress monitoring, and behavior adjustment to achieve desired outcomes. In the peer evaluation process, the quality of peer feedback provided by students can influence the self‐observation aspect of Bandura's self‐regulation theory. Quality peer feedback can provide students with knowledge of performance‐standard gaps or discrepancies between one's self‐perception and that of a comparator (peer) (Carver & Scheier, [<reflink idref="bib11" id="ref55">11</reflink>]). Such gaps emerge whenever there is a contrast between one's beliefs about their performance and the feedback provided by peers. However, this is possible only if peer feedback is specific and constructive; thus, to help students self‐regulate their learning and develop their teamwork skills, we must understand how students provide feedback to better understand feedback literacy. Access to quality feedback can stimulate self‐regulation, with more frequent attention to feedback‐standard gaps encouraging students to adjust their behavior and become more self‐aware (Zimmerman, [<reflink idref="bib64" id="ref56">64</reflink>]).</p> <hd id="AN0186992462-8">Quality feedback</hd> <p>Defining aspects of quality feedback is critical to assessing students' ability to provide feedback, a key element of feedback literacy. We considered identifying important elements of feedback quality as the first step to helping engineering students improve and regulate their teamwork behaviors and skills. This is another area where we leveraged Bandura's theory of self‐regulation as we evaluated existing rubrics.</p> <p>To evaluate the components of feedback quality, we sought to understand what constructive, effective peer feedback looks like and how peer feedback has been evaluated. We reviewed several frameworks or rubrics used to evaluate feedback. One major issue with other frameworks and rubrics used to evaluate feedback quality was their complexity, which was challenging to apply to relatively short peer comments, so we considered rubrics that approached evaluating feedback more generally. Gauthier et al. ([<reflink idref="bib25" id="ref57">25</reflink>]) from the medical education field developed a rubric using the Task, Gap, Action framework originally presented by Sadler. While Sadler's work was concerned with formative assessment in all educational fields (Sadler, [<reflink idref="bib51" id="ref58">51</reflink>]), the rubric created by Gautheir et al. evaluates peer feedback for medical residency training across three criteria: Task, Gap, and Action. Task describes the context in which the feedback was given, which can include the content of a student's participation or the value that their participation added to the team. Gap recognizes the differences between behaviors displayed and expected, and Action describes what can be done to improve (Gauthier et al., [<reflink idref="bib25" id="ref59">25</reflink>]). Other studies in medical education have used the Gauthier et al. rubric to evaluate undergraduate students (Abraham & Singaram, [<reflink idref="bib2" id="ref60">2</reflink>]) and medical students in a team environment (Burgess et al., [<reflink idref="bib8" id="ref61">8</reflink>]). Notably, our focus on task, gap, and action constructs aligns with observational learning, self‐reflection, and self‐efficacy as related to Bandura's work and the features of feedback literacy, as related to Carless and Boud. The description of a team member's tasks contributes to "appreciating feedback information" and "managing affect" by presenting a summary of how the team member's actions are positively contributing to the team. The gap construct helps with "making judgments" about performance, and action is clearly tied to the "taking action" feature of feedback literacy. Further details on the development of this rubric and its grounding in theory can be found in our prior work (Drinkwater et al., [<reflink idref="bib20" id="ref62">20</reflink>]). This framework became the basis of our rubric to evaluate feedback quality.</p> <hd id="AN0186992462-9">METHODS</hd> <p>This paper presents a deductive approach to the qualitative analysis of peer comments from first‐year engineering students enrolled in a required "Foundations of Engineering" course. Our approach utilized an open‐source LLM and researchers to create quantitative quality ratings from the peer comments. Matched mixed methods, as outlined by Hochwald et al. ([<reflink idref="bib29" id="ref63">29</reflink>]), analyze quantitative values derived from qualitative sources with both quantitative and qualitative methods to add robustness to the results. We adapt this method by using an LLM to determine the quantitative values of qualitative peer feedback comments using a feedback quality rubric. This section first describes the dataset, then discusses the two phases of rubric rating (researcher and LLM) and the analysis tools used to compare the phases.</p> <hd id="AN0186992462-10">Data collection and context</hd> <p>We collected data from a first‐year engineering PBL course at a large R1 university in the mid‐Atlantic region of the United States. The course is centered around a semester‐long design project where students are required to work in teams of 4–5 students. The teams are required to apply the engineering design process and utilize the tools they have learned in class to create a prototype of their design. The class size typically ranges from 60 to 72 students.</p> <p>The primary data source for this study is peer feedback comments from students enrolled in the first‐year course. Since students work in teams throughout the semester, peer feedback comments are collected periodically by the instructor to monitor team dynamics and provide formative feedback to students to help them improve their teamwork behaviors. The peer comments were collected through CATME, a web‐based tool that allows instructors to form teams based on specific criteria and facilitates a student peer feedback process through quantitative ratings on five dimensions of teamwork, along with qualitative self and peer feedback comments.</p> <p>CATME feedback was collected at three time points in a 16‐week semester (Week 7, Week 11, and Week 16). In the assignment description, students were tasked to write at least 2–3 specific comments for each team member (including themselves) regarding teamwork behaviors. Students were encouraged to provide both positive and constructive actionable feedback in their comments for Weeks 7 and 11. For Week 7, students were also instructed that these comments would not be directly released to their peers but instead used to inform personalized AI‐generated performance feedback reports for each student. For further details on how these AI reports work, see Sajadi et al. ([<reflink idref="bib52" id="ref64">52</reflink>]). For Week 11, students were instructed that their feedback comments would be directly released to their peers. For Week 16, students were instructed that their comments were not to be used as formative feedback but instead as summative comments for the instructors to use in assessment and, thus, would not be released to peers. This decision was made to encourage honesty and openness in student comments and minimize bias. Prior work has highlighted that the use of peer feedback scenarios for assessment purposes at the end of the semester tends to introduce increased bias (Sridharan et al., [<reflink idref="bib56" id="ref65">56</reflink>]). Our focus is on formative peer feedback, so only Week 7 and Week 11 comments were used for this study. The end‐of‐semester comments were not included in our dataset because these peer comments were summative in nature, while comments during the semester can be used to provide formative feedback that can inform teamwork behaviors. Table 1 summarizes the feedback schedule in the first‐year course.</p> <p>1 TABLE Peer feedback schedule and purpose.</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Aspect</th><th align="left">Week 7</th><th align="left">Week 11</th><th align="left">Week 16</th></tr></thead><tbody valign="top"><tr><td align="left">Purpose</td><td align="left">Provide formative feedback on teamwork</td><td align="left">Provide formative feedback on teamwork</td><td align="left">Used for summative assessment and grading purposes</td></tr><tr><td align="left">Data inclusion in study</td><td align="left">Included in study</td><td align="left">Included in study</td><td align="left">Not included in study</td></tr><tr><td align="left">Feedback visibility</td><td align="left">Not directly shared with peers; used for AI‐generated feedback reports</td><td align="left">Directly shared with peers</td><td align="left">Not shared with peers; only for instructor assessment</td></tr><tr><td align="left">Instructor use of feedback</td><td align="left">Used to inform formative feedback for teamwork improvement</td><td align="left">Used to provide direct peer feedback and improve teamwork behaviors</td><td align="left">Used for assessment and grading purposes, not formative feedback</td></tr></tbody></table> </ephtml> </p> <p>The dataset includes peer comments (self comments were not included) from the Fall 2023 semester. In Fall 2023, students participated in a 30‐min training about the CATME peer evaluation process before the first peer evaluation assignment was released. The training included an overview of the CATME dimensions, best practices for quantitative scoring and writing qualitative comments, and examples of high‐ and low‐quality feedback.</p> <hd id="AN0186992462-11">Participants</hd> <p>All procedures were approved by the IRB at the authors' institution. Students were asked to consent to the use of their peer comments for research purposes, and 118 students consented to their comments being used in the study. This led to a total of 295 comments used in this study. The demographic information for the participants can be found in Table 2.</p> <p>2 TABLE Demographic information.</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left" /><th align="left">Number of students</th></tr></thead><tbody valign="top"><tr><td align="left">Gender</td></tr><tr><td align="left">Male</td><td align="char" char=".">86</td></tr><tr><td align="left">Female</td><td align="char" char=".">26</td></tr><tr><td align="left">Other or prefer not to answer</td><td align="char" char=".">6</td></tr><tr><td align="left">Race/ethnicity</td></tr><tr><td align="left">Asian</td><td align="char" char=".">40</td></tr><tr><td align="left">Black or African American</td><td align="char" char=".">8</td></tr><tr><td align="left">Hispanic, Latino, or Spanish origin</td><td align="char" char=".">14</td></tr><tr><td align="left">Middle Eastern or North African</td><td align="char" char=".">2</td></tr><tr><td align="left">Native Hawaiian or other Pacific Islander</td><td align="char" char=".">0</td></tr><tr><td align="left">White</td><td align="char" char=".">40</td></tr><tr><td align="left">Other or prefer not to answer</td><td align="char" char=".">14</td></tr></tbody></table> </ephtml> </p> <hd id="AN0186992462-12">Phase 1: Researcher's rating</hd> <p>Before an LLM was integrated into the analysis phase of this study, the researchers iteratively developed a feedback quality rubric and quantitatively rated the peer comments. This section describes the theory‐guided development of the rubric and how the researchers approached the rating of the dataset.</p> <hd id="AN0186992462-13">First round coding</hd> <p>Informed by the principles of self‐regulated learning, we selected the Gauthier et al. ([<reflink idref="bib25" id="ref66">25</reflink>]) as the starting point for our study of peer feedback quality in engineering. Our prior work (Drinkwater et al., [<reflink idref="bib20" id="ref67">20</reflink>]) details how the inconsistency of the first round of coding resulted in changes made to the rubric. To summarize, during the first round of coding, the Gauthier et al. rubric was applied to each peer comment by two researchers, and inter‐rater reliability was calculated for each rubric criterion. This led to poor agreement in IRR with a quadratic weighted Cohen's kappa of 0.523 for Task, 0.054 for Gap, and 0.685 for Action. In our debrief after the first round of coding, we recognized that the Task criterion was ill‐defined, and we had different interpretations of the Gap criterion, which led to the low agreement. The low agreement in this first round of coding led us to revise the rubric.</p> <hd id="AN0186992462-14">Rubric iteration</hd> <p>Through our first round of coding, we found that some aspects of the engineering design context were not captured using the Gauthier et al. rubric, so we chose to revise it to align more closely with the parameters of the peer feedback we collected.</p> <p>Engineering students develop in both technical and professional domains. The Accreditation Board for Engineering and Technology (ABET) specifies technical outcomes for engineering students. The outcomes focus on (i) students' abilities to understand complex problems and apply design principles to solve problems, and (ii) professional outcomes, which focus on students' abilities to effectively communicate and work on teams (EAC Criteria, [<reflink idref="bib1" id="ref68">1</reflink>]). The Task criterion of the Gauthier et al. rubric was too broad for the engineering education context. For example, if a comment specifically focused on the professional skills rather than the technical tasks that students contributed to the project, it was unclear how to rate the comment in the Task criterion. Therefore, we extended the three criteria from the Gauthier et al. rubric to four quality measures that measure peer feedback quality <emph>in engineering</emph>: Contributions to Group Tasks, Behavior, Gap, and Action. The key difference is the addition of the Behavior criterion. Activities included in the Contribution to Group Tasks criterion include role fulfillment, task management, task execution, project contributions, and work output. The Behavior criterion considers actions that are not tied to the current project tasks, like interpersonal dynamics, team engagement, communication, or personal attributes. This change aligned with the technical and professional ABET outcomes and the overall purpose of team‐based activities in project‐based learning classes, where students are developing their technical abilities and simultaneously developing their teamwork abilities.</p> <p>The Gap criterion recognizes a negative difference between the ratee's performance and the expected standard of the team. Lastly, the Action criterion includes feedback that suggests how the ratee can improve their performance and remedy the identified gaps. These criteria were incorporated into a rubric with a scale ranging from 0 to 2, a change from the original Gauthier et al. rubric. A score of 0 indicates that the comment does not address the criterion. A score of 1 indicates that the comment provides a general or vague description of the criterion, while a score of 2 indicates that the comment provides a specific and detailed description related to that criterion. Table 3 shows the complete rubric.</p> <p>3 TABLE Feedback quality rubric.</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Feedback criterion</th><th align="left">0</th><th align="left">1</th><th align="left">2</th></tr></thead><tbody valign="top"><tr><td align="left">Contributions to group tasks (role fulfillment, task management, task execution, project contributions, work output)Exclude general references to teammate behaviors if they are not connected to a specific project task</td><td align="left">No task described</td><td align="left">General or vague description of how the team member contributes to group tasks, with minimal details</td><td align="left">Specific or detailed description of tasks that the team member contributes to that imply their value to the team</td></tr><tr><td align="left">Behavior (interpersonal dynamics, team engagement, communications, personal attributes)Exclude mentions of contributions to specific project assignments or tasks</td><td align="left">No behaviors described</td><td align="left">General or vague description of team member behaviors, with minimal details</td><td align="left">Specific or detailed description of team member behaviors that imply their value to the team</td></tr><tr><td align="left">Gap (recognition of a negative difference between the ratee's performance and an expected standard)</td><td align="left">No gap identified</td><td align="left">A gap is alluded to or briefly mentioned, but lacks specific details on how it compares to an expected standard</td><td align="left">A gap is discussed with specific details that highlight or easily imply how the gap compares to an expected standard</td></tr><tr><td align="left">Action (suggested steps to remedy gaps or improve performance)</td><td align="left">No actions identified</td><td align="left">General or vague description of an action with minimal details</td><td align="left">Specific or detailed description of how a team member should address a gap in their performance</td></tr></tbody></table> </ephtml> </p> <p>1 <emph>Note</emph>: The italicized text indicates guidance added to replicate the researchers' application rules for the LLM ratings.</p> <p>We also established rubric application rules for the researchers. First, we agreed that any portion of a peer comment could only contribute to a score in one criterion. For example, the same phrase or sentence of a comment could not count for a 1 in both Task and Behavior. Overall, our iterations included clarifying definitions of Gap, separating the Task domain into Contributions to group tasks and Behavior, using a smaller range of scores, and establishing rubric application rules. These iterations and rules were used for the second round of coding.</p> <hd id="AN0186992462-15">Second round coding</hd> <p>The rubric iteration provided clearer definitions and examples, resulting in greater consistency among coders in the second round of coding. The final rubric used in the second round is included in Table 3. The primary changes were expanding the Task domain to Contributions to Group Tasks (hereafter referred to as 'Task') and Behavior and using a smaller range of scores.</p> <p>We used the same process as the first round of coding, in which two researchers applied the new rubric to each peer comment. The new rubric led to an improved IRR score (Table 5 in Results), and we were satisfied with the level of agreement. We recognize that using a smaller range of scores could have improved the IRR, but we also expanded one of the criteria, which could have hurt the IRR. Overall, we believe the improved IRR can be attributed to the clearer definitions in the rubric, due to our discussions as a research group.</p> <hd id="AN0186992462-16">Phase 2: Large language model</hd> <p>Using an LLM to apply the rubric to the comments was a time‐consuming process because we tested multiple models and needed to iterate the prompt to get the most reliable and consistent results possible. In this section, we highlight the iterations and reasoning employed to utilize an LLM as a coder to rate feedback quality using the rubric.</p> <hd id="AN0186992462-17">Selecting large language models</hd> <p>We used open‐source generative text models for two parts of this study: comment de‐identification and rubric application. In our initial testing, we ran a small subset of 10 comments with five LLMs (llama‐3.1‐70b, mistral‐large‐123b, mistral‐8‐22b, mistral‐nemo‐12b, and qwen‐2.5‐32b). All of these models are open‐source, can be run locally, and have performed well on previous research tasks conducted by one of the researchers. We chose these five LLMs to represent a variety of parameter sizes (12–123 billion), developers, and training data.</p> <p>The de‐identification task used the mistral‐nemo‐12b model because it performs well for its size according to common leaderboard metrics and is an Apache‐2.0‐licensed model (Apache License, Version 2.0, [<reflink idref="bib57" id="ref69">57</reflink>]) (NB: the "b" in the model names, e.g., "12b" corresponds to the number of billions of parameters in the model). We set the temperature parameter to 0 for replicability purposes because a lower temperature setting makes the models more deterministic (Hingle et al., [<reflink idref="bib28" id="ref70">28</reflink>]).</p> <p>For the rubric application task, we used the qwen‐2.5‐32b model because of its Apache‐2.0 license and even better performance in our initial testing compared to other models. The models were evaluated based on the coherence of their responses and an average error calculation when compared to the researchers. The average error was calculated for each rubric criterion by averaging the absolute distance between the model's score and the researchers' average score. While the mistral‐large‐123b model had the lowest overall error (0.213 average point difference), we chose to use the qwen‐2.5‐32b for the whole dataset because the average error was similar (0.263 points), and qwen‐2.5‐32b is a much smaller model that requires fewer computational resources to run. Computation time and resources are key considerations in choosing an LLM, and our goal was to represent a process that is accessible to most researchers. We again set the temperature parameter to 0 for this model to make the results as consistent as possible.</p> <p>The process of determining which model was best was iterative. It required one of the authors to go through the results with a critical eye to determine if there were any instances where the model did not apply the rubric appropriately. For example, when we ran a subset of 10 comments for different models, the mistral‐nemo‐12b model had an instance where it scored a comment 3 for the Behavior criterion. The highest rubric score is 2, so this did not make sense, and the researchers both scored the comment 0 for Behavior, so there was a significant mismatch in interpretation. Despite recognizing that the score did not make sense in the final check, the model did not correct the score. The comment, rating, and justification are shown in Table 4. Illogical mistakes like the one exhibited here are a clue that the model is too small for the task, but careful review of pilot data is needed to identify the clues.</p> <p>4 TABLE Example of rubric error from mistral‐nemo‐12b.</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Comment</th><th align="left">Researcher rating – behavior</th><th align="left">Model rating – behavior</th><th align="left">Model justification</th></tr></thead><tbody valign="top"><tr><td align="left">Produces satisfactory work. Submits work well before deadlines. I do realize that this person is not that skilled in solidworks</td><td align="left">0</td><td align="left">3</td><td align="left">{"step_by_step_analysis": "The given written feedback mentions two behaviors of the teammate: 'Produces satisfactory work' and 'Submits work well before deadlines'. However, it also includes a negative comment about their skills in SolidWorks, which is not considered a behavior as per the rubric's definition. Therefore, we only consider the first two points for evaluation.", "justification": "'Produces satisfactory work' is a general description of the teammate's behavior but does not provide specific details (1 point). 'Submits work well before deadlines' is a specific and detailed description that implies their value to the team by showing responsibility and reliability, which are positive personal attributes (2 points).", "reflection": "The feedback could have been more comprehensive if it provided examples of how the teammate's work was satisfactory or the impact of submitting work early on the team's progress. The mention of SolidWorks skills was not considered in the evaluation as per the rubric's instructions." "score_and_confidence": {"score": 3, "confidence_level": "high"}, "final_check": "I have double‐checked my analysis and ensured that I have adhered to the rubric's definition of 'behavior' and its scoring guidelines. The total score is 3, which is not possible according to the rubric (maximum score is 2). Therefore, I need to correct my final score."}</td></tr></tbody></table> </ephtml> </p> <hd id="AN0186992462-18">Prompt development</hd> <p>We used a combination of suggested best practices to develop the prompts to instruct the model for both the de‐identification and rubric rating tasks. First, we assigned a persona to the model, which prior research has suggested can help model performance by shifting the distribution of its output into a particular outcome space (Olea et al., [<reflink idref="bib46" id="ref71">46</reflink>]). Second, we instructed the model to use a chain‐of‐thought approach to reasoning about the rubric application. Chain‐of‐thought prompting techniques encourage models to outline sequential steps before arriving at their conclusions (Wei et al., [<reflink idref="bib58" id="ref72">58</reflink>]). The third technique was the use of XML tags to add further structure to the prompt. The goal here is to assist the model in knowing which pieces of information to attend to. This is important because these language models are transformer‐based neural networks that use attention mechanisms to determine which pieces of context information are most relevant for determining the next token to generate. Using XML tags (or other structuring approaches) helps the model to identify the more salient information and improve the clarity of instruction.</p> <p>The process of developing the prompts was iterative. Similar to how we chose which model to use, we applied the prompt to a smaller subset of data, 100 comments, to determine if different prompts applied the rubric better. One of the researchers read through the results to determine which prompt applied the rubric better, which helped identify common issues. The researcher was one of the raters in Phase 1, so they were familiar with the data and rubric. They utilized this knowledge to assess the logical coherence of the models. Some models did not execute the rating procedure correctly, which made it a simple choice to remove them. The larger LLMs executed the task similarly, so the researcher qualitatively evaluated the justifications to choose the best prompt. More details on the errors identified and changes made to the prompts are included in the next two sections.</p> <hd id="AN0186992462-19">LLM comment de‐identification</hd> <p>A de‐identification process was used to remove identifying information from the peer comments. We wanted to remove names and gendered pronouns from the responses under the hypothesis that the LLM could mimic the biases found in industry that cause women to receive less actionable feedback than men (Doldor et al., [<reflink idref="bib19" id="ref73">19</reflink>]). Concern over inherent bias in LLMs has been widely acknowledged but is not easily measured or controlled (Fui‐Hoon Nah et al., [<reflink idref="bib23" id="ref74">23</reflink>]). To avoid any potential influence on the LLM from names or gendered pronouns, we chose to remove them from the data before applying the rubric. This involved replacing names and third‐person pronouns with the appropriate non‐gendered third‐person pronoun (e.g., "she" with "they", "his" with "their", "John" with "Person A"). The de‐identification process was conducted in two iterations. In the first iteration, we piloted the model with 100 comments, and a member of the research team then read these cleaned responses to confirm the process worked and to correct any errors where necessary. In our pilot iteration, we identified some issues with missed pronouns and incorrect deidentification when comments referenced multiple team members, so specific guidance for these issues was added to the prompt. The final version of the de‐identification prompt is included in Appendix A. While this process was proficient at identifying third‐person pronouns (i.e., he, she, they) and common English names, it occasionally struggled with possessive pronouns (i.e., his, her) at the end of comments and less common names. In the final iteration with all 295 comments, we did not identify any possessive pronouns that remained, but two names were included after de‐identification by the LLM. The names were removed manually by the researcher who reviewed the data after de‐identification.</p> <p>We recognize that we took the precaution to de‐identify the comments due to potential bias in LLM, but the researchers analyzed the comments without de‐identifying the data. We did not de‐identify the data for the researchers' analysis because this is not a common practice at the analysis stage of general qualitative analysis. Furthermore, the pronouns and names included in peer comments referred to the recipient, not the comment author whose feedback quality we were rating. So, the identities of the recipients should have had no bearing on the comment quality. However, we recognize that the inclusion of identifying information might impact our analysis, so we took several precautions to avoid bias. First, the instructor of the course did not analyze any of the comments because he knew the students in the course, and the researchers who analyzed the comments had no interaction with the participants. Additionally, the researchers recognized their positionality and practiced reflexivity to avoid making assumptions. Since the comments were short and we were evaluating feedback quality, we were focused exclusively on the content of the comments at large, which helped us to avoid bias. Although we took precautions to mitigate bias, we recognize that this is a limitation of our analysis.</p> <hd id="AN0186992462-20">LLM rubric rating</hd> <p>To apply the rubric and evaluate each de‐identified response, we broke the process into smaller steps. For example, we had to decide whether to apply all four rubric criteria at once (as seen in Table 3) or one criterion at a time, through four separate calls to the language model. We tested both options in our pilot of 10 peer comments. When all four criteria were rated together, the model's justifications were difficult for researchers to parse. There is also a risk of confusing the LLM by requesting too many pieces of information to attend to in the task. As such, we opted for the latter option—instructing the language model to apply one of the four criteria and provide a justification for the suggested score. Running one criterion at a time restricts the amount of information the model must process at one time. Additionally, no examples, training data, or dialogues were included in the prompting of the model, and a new conversation thread was started for each comment. Thus, it was as if the model was applying the rubric for the first time for every comment. We believe that four smaller tasks were more appropriate for the LLM and led to more accurate ratings, but we did not formally test this hypothesis. The complete prompt with XML tags used to rate the comments is included in Appendix B.</p> <hd id="AN0186992462-21">Data analysis</hd> <p>To compare the results from researchers and the LLM, inter‐rater reliability and a qualitative process analysis were used. To address our theoretical research question, descriptive statistics and a holistic qualitative analysis were conducted to analyze the feedback quality.</p> <hd id="AN0186992462-22">Inter‐rater reliability (IRR)</hd> <p>In order to answer our methodological research question to determine the reliability of the LLM as a "rater," we calculated inter‐rater reliability. We calculated the quadratic weighted Cohen's kappa statistic for each rubric category. Cohen's kappa is used to evaluate the reliability of inter‐rater agreement, adjusting for chance (Cohen, [<reflink idref="bib13" id="ref75">13</reflink>]). A quadratic weighted Cohen's kappa was considered because we felt the difference between 0 and 1 in the rubric is different than the difference between 1 and 2; a linear weighted Cohen's kappa assumes that the difference between each score is the same. The model's application of the rubric was compared to both researchers' application of the rubric using IRR.</p> <hd id="AN0186992462-23">Descriptive statistics</hd> <p>We used descriptive statistics to answer our theoretical research question. To determine the elements of feedback from the practices of first‐year engineering students, we calculated the mean, standard deviation, and interquartile range for each rubric category and disaggregated by researchers and model. This showed which feedback measures engineering students were generally better or worse at writing. Moreover, to visualize the distribution of scores for each rubric category, we used a stacked bar chart. This chart allowed us to see the difference in score distribution between the rubric categories and between the researchers and the model.</p> <hd id="AN0186992462-24">Feedback quality typology</hd> <p>While descriptive statistics are valuable in characterizing the distribution of each rubric criterion, they do not draw connections or patterns across the entire feedback quality rubric. Research question 2 addresses feedback quality as a whole, so it is important to analyze trends between the four feedback quality criteria. The researchers decided to represent these relevant patterns in a typology of peer feedback comments. After the second round of researcher coding, the researchers met to discuss themes of feedback quality. From these discussions and significant time spent with the raw dataset, the researchers identified five popular types of peer comments. To verify the frequency of these types, one researcher compared the distribution of total feedback quality scores (i.e., the sum of ratings for Task, Behavior, Gap, and Action) to the total scores that would be expected from each comment type in the typology. The Results section presents a description, example, range of total scores, and frequency in the dataset for each comment type.</p> <hd id="AN0186992462-25">Positionality</hd> <p>Acknowledging how our backgrounds and experiences shape qualitative research is essential to ensure an accurate representation of the data (Secules et al., [<reflink idref="bib53" id="ref76">53</reflink>]). We approached this study from a pragmatist philosophical stance, as our goal was to identify practical and reliable applications of LLMs in engineering education research. As engineering educators and researchers, we strive to pragmatically consider how research in feedback literacy can amplify the pedagogical impact of peer feedback in PBL contexts. Although most of our results in this paper are quantitative, our approach to this research and interpretation of peer comments was influenced by our experiences and interests in qualitative EER. Four of the authors have an interest in and commitment to improving the collaboration experiences of engineering students. Through our previous experiences, we recognize the importance of collaboration in engineering work and the necessity to help students to authentically practice teamwork skills. In prior or current roles as engineers, engineering students, instructors, and teaching assistants, we have observed that engineering students need feedback literacy to develop their teamwork skills.</p> <p>One of the authors has been utilizing GAI in his work over the past several years, and we identified an opportunity for collaboration by pooling the research team's interests. LLMs have many potential benefits for analyzing large datasets, and analyzing large datasets of peer comments can provide insight into the feedback practices of engineering students. Although wary of potential risks, we believe that tools like LLMs hold incredible potential for researchers when approached with tact and care. To minimize the risks involved in using AI, we were intentional about protecting students' identities and mitigating potential gender biases by removing names and pronouns from the student comments. This practice was motivated by several of the authors' experiences as women in engineering who have received biased and problematic feedback. Moreover, in order to make this approach more accessible, we used LLMs that balanced practicality and reliability.</p> <hd id="AN0186992462-26">Data limitations</hd> <p>There are limitations that arise from our data source. First‐year engineering students are likely new to writing peer feedback, and their ability to write quality feedback is limited. Many of the responses were relatively short and did not provide constructive feedback, so the dataset often lacked any reference to the Gap and Action rubric category. This limited our ability to test the language model's performance on the full range of comment quality.</p> <hd id="AN0186992462-27">RESULTS</hd> <p>The results of our parallel analysis of the methodological reliability of using an LLM as a rubric rater for qualitative data and the landscape of feedback quality for first‐year engineering students reveal mixed findings. The inter‐rater reliability of the LLM compared to researcher raters displayed similar patterns as the comparison between researchers. A review of the quality ratings from the dataset shows that first‐year students have a bias toward positive feedback over constructive feedback and tend to write comments that fit into one of five types.</p> <hd id="AN0186992462-28">Methodological results</hd> <p>Three methods of analysis were employed for the reliability of the LLM. The first is the Inter‐Rater Reliability. Table 5 displays the IRR results for each rubric criterion by comparing the GAI rating and two ratings from researchers. The researcher comparison (Rating1 vs. Rating2) IRR for Task is around 0.60, which is considered "substantial" reliability, and the Gap and Action values are above 0.8, which is "almost perfect" (Cohen, [<reflink idref="bib13" id="ref77">13</reflink>]). The IRR among researchers for Behavior is 0.63, which is still considered "substantial." On comparing the LLM rating to each researcher rating (Rating vs. Model), the IRR is worse than the researcher versus researcher IRR for all categories. The difference is particularly pronounced for Task and Behavior, where the IRR kappa values are below 0.40, which is considered only "fair" reliability. The difference is smaller for the Gap and Action categories, where the model still achieved "substantial" or "almost perfect" reliability compared to the researchers. Based on these results, the LLM performed better where researchers performed better, and both types of raters struggled with Task and Behavior.</p> <p>5 TABLE Inter‐rater reliability statistics.</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left" /><th align="left">Quadratic weighted kappa</th></tr><tr><th align="left">Comparison</th><th align="left">Task</th><th align="left">Behavior</th><th align="left">Gap</th><th align="left">Action</th></tr></thead><tbody valign="top"><tr><td align="left">Rating1 vs. Rating2</td><td align="left">0.58</td><td align="left">0.63</td><td align="left">0.88</td><td align="left">0.83</td></tr><tr><td align="left">Rating1 vs. Model</td><td align="left">0.40</td><td align="left">0.32</td><td align="left">0.79</td><td align="left">0.86</td></tr><tr><td align="left">Rating2 vs. Model</td><td align="left">0.34</td><td align="left">0.26</td><td align="left">0.78</td><td align="left">0.85</td></tr></tbody></table> </ephtml> </p> <p>To further examine the differences between human researchers and the GAI model, we qualitatively examined the reliability of the LLM when rating peer comments. Understanding the justification process used by the model is a key component of its reliability, and we did this by identifying cases of alignment and misalignment between the LLM and researchers.</p> <p>We first identified representative cases of high agreement and high disagreement between the researcher and the model. Our chain‐of‐thought prompt required the model to include a thorough explanation of its rating decisions, which enabled the researchers to compare the "thought processes" of the raters. Table 6 presents a comparison of researcher and model scores and justifications. The first example displays a case of total agreement between the model and researchers. This comment provides vague feedback on a peer's task completion and collaboration. The LLM was able to discern the lack of detail for Task and Behavior criteria and detect that no Gap or Action was included. The remaining four examples were selected to represent patterns of disagreement for each rubric criterion. For Task, the LLM was more likely to give a score of 1 when the researchers gave a 0. In the Behavior criterion, the LLM gave 2s for 16% more comments than the researchers, so the example shows a comment where the model gave a 2 for Behavior while the researchers gave a 1. The last two criteria had improved IRR, but there were still many discrepancies in scoring. For Gap, the researchers often gave higher scores than the model. So, in this case, we chose an example where the researchers gave a 2 for Gap while the model only scored a 1. Lastly, while both researchers and the model gave 0 for over 70% of Action scores, Table 6 shows an example where the model misinterpreted a comment and gave a higher Action score.</p> <p>6 TABLE Examples of score agreement and disagreement between LLM and researchers.</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Rubric criterion</th><th align="left">Peer comment</th><th align="left">LLM score</th><th align="left">LLM justification</th><th align="left">Researcher score</th><th align="left">Researchers' comments on discrepancy</th></tr></thead><tbody valign="top"><tr><td align="left">Agreement between LLM and researcher</td></tr><tr><td align="left">All</td><td align="left">"They are excellent at their job and complete everything on schedule without the need for reminders. Additionally, they are an excellent collaborator who never seeks to discredit others' views."</td><td align="left">Task: 1, Behavior:1, Gap:0, Action:0</td><td align="left">(For Task:) The feedback states "Person A is excellent at Person A's job and completes everything on schedule without the need for reminders." This indicates a general description but lacks specificity regarding particular tasks. Additionally, "Person A is an excellent collaborator who never seeks to discredit others' views" focuses more on behavior than specific task contributions.</td><td align="left">Task: 1, Behavior:1, Gap:0, Action:0</td><td align="left">All scores for this comment were in agreement. The LLM identified aspects that contributed to task, such as their task completion. The model correctly parsed that collaboration is more related to behavior than task management.</td></tr><tr><td align="left">Disagreement between LLM and researcher</td></tr><tr><td align="left">Task</td><td align="left">"I think that they need to work to be a bit more involved but overall have been more present in the classes."</td><td align="left">1</td><td align="left">The written feedback states "Person A needs to work to be a bit more involved but overall has been more present in the classes." This statement is general and vague, lacking any specific details about tasks or contributions. It does not provide enough information to justify a score higher than 1 point according to the rubric.</td><td align="left">0</td><td align="left">The AI incorrectly identifies a gap in the student's involvement as a task and does not provide justification for why the comment deserves a "1" instead of a "0"</td></tr><tr><td align="left">Behavior</td><td align="left">"They radiate positivity and enthusiasm, creating a vibrant atmosphere around them. However, I do wish they would express themselves more verbally, as their perspective and insights would undoubtedly add value to our discussions."</td><td align="left">2</td><td align="left">The written feedback describes Person A's positivity and enthusiasm as positive personal attributes. It also mentions a desire for more verbal expression from Person A, indicating an area of improvement in communication skills. These specific behaviors imply their value to the team, aligning with the criteria for 2 points under "Behavior."</td><td align="left">1</td><td align="left">The AI model correctly identified personal attributes that relate to behavior, but misidentified an area for improvement, or Gap, as a behavior‐related comment.</td></tr><tr><td align="left">Gap</td><td align="left">"They are a charismatic person that always brings life into the group, maintaining a positive vibe which results in increased productivity. They are also good at organization, as they asked if we could meet before the presentation to go over and practice our presentation. However, sometimes their work isn't of the highest quality, sometimes requiring review like the project scoping document. Still, they do put out good work overall. If they could improve, maybe they could spend a little more time making sure they meet all the requirements of the assignments."</td><td align="left">1</td><td align="left">The written feedback mentions that "sometimes Person A's work isn't of the highest quality, sometimes requiring review like the project scoping document." This indicates recognition of a gap but lacks detailed comparison to an expected standard. Therefore, it aligns more closely with scoring 1 point rather than 2 points.</td><td align="left">2</td><td align="left">The AI model felt that the work standard was not clear in this comment, but the researcher assumed that the accepted standard was to meet the assignment requirements. Additionally, the feedback mentioned an example of when their work was not of the highest quality—the project scoping document.</td></tr><tr><td align="left">Action</td><td align="left">"I think that they need to work to be a bit more involved but overall have been more present in the classes."</td><td align="left">1</td><td align="left">The feedback suggests an action ("needs to work to be a bit more involved") but does not provide detailed steps on how Person A should improve their involvement. This meets the criteria for 1 point as it is a general description of an action with minimal details.</td><td align="left">0</td><td align="left">The AI model misattributed the identified Gap as an Action. The sentence describes the Gap in the student's behavior without giving any details on what involvement would mean.</td></tr></tbody></table> </ephtml> </p> <p>While examining the justifications employed by the model, we identified some inconsistencies with the ultimate score given for similar comments. The model we chose, qwen‐2.5‐32b, had overall better justifications for scores and fewer hallucinations than some of the other models we considered, but there were some instances when it did not consistently apply the rubric. For example, two very similar comments are shown in Table 7. The researchers agreed that both these comments received a 0 for Task and Behavior, while the model rated one of them a 0 for Task and Behavior and the other received 1s for Task and Behavior. This shows some inconsistencies with the model's application of the rubric for very similar comments. Overall, we found that each model had inconsistencies, but we chose qwen‐2.5‐32b because it had the best balance of logical justification, score consistency, and reasonable model size.</p> <p>7 TABLE Example of model rating inconsistency from qwen‐2.5.32b.</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Comments</th><th align="left">Researcher ratings – task</th><th align="left">Model rating – task</th><th align="left">Model justification</th></tr></thead><tbody valign="top"><tr><td align="left">Good team member, does their part</td><td align="left">0; 0</td><td align="left">1</td><td align="left">"Step 1: I carefully read the rubric and the written feedback provided by the student.Step 2: Applying the rubric criteria step‐by‐step to the feedback, I noted that the feedback is very brief and lacks specific details about the teammate's contributions to group tasks. The phrase 'good team member' is vague and does not provide any concrete examples of how the teammate contributed to the project tasks.Step 3: Specific examples from the feedback include phrases like 'does their part,' which is general and does not specify what tasks were completed or how they contributed to the overall success of the group project. There are no specific details about role fulfillment, task management, task execution, project contributions, or work output.Step 4: The main challenge encountered during this evaluation was the lack of specificity in the feedback provided by the student. It is difficult to assign a higher score when there are no concrete examples given that align with the rubric's criteria for detailed descriptions of tasks and their value to the team.Step 5: Based on my analysis, I would assign a score of 1 point because the feedback provides a general description of how the teammate contributes but lacks specific details. The phrase 'does their part' implies some level of contribution without providing any concrete examples or context.Step 6: My confidence in this scoring decision is high because the rubric clearly defines what constitutes each score, and the feedback provided aligns most closely with the criteria for a score of 1 point.Step 7: I have double‐checked my analysis to ensure consistency and adherence to the rubric. The feedback does not meet the criteria for a higher score due to its lack of specific details about contributions to group tasks."</td></tr><tr><td align="left">Person A's work is satisfactory.</td><td align="left">0; 0</td><td align="left">0</td><td align="left">Step 1: I carefully read the rubric and the written feedback provided by the student.Step 2: Applying the rubric criteria step‐by‐step to the feedback:• The rubric requires specific or detailed descriptions of tasks that imply value to the team for a score of 2 points.• A general or vague description with minimal details would warrant 1 point.• No task described results in 0 points.Step 3: Justification based on examples from the feedback:• The written feedback states, 'Person A's work is satisfactory.' This statement does not describe any specific tasks performed by • Person A or their contributions to group tasks. It lacks detail and specificity required for a higher score.Step 4: Reflection on challenges or ambiguities encountered:• There are no significant challenges or ambiguities in applying the rubric to this feedback, as it clearly falls short of providing • any specific task details.Step 5: Assigning a score based on analysis:• Given that there is no description of tasks performed by Person A and only a general statement about their work being satisfactory, the appropriate score according to the rubric is 0 points.Step 6: Indicating confidence level in scoring decision:• My confidence level in this scoring decision is high because the feedback does not meet any criteria for higher scores as per the rubric.Step 7: Double‐checking analysis for consistency and adherence to the rubric:• The analysis aligns with the rubric's requirements, ensuring that no task description is present, which justifies a score of 0 points.</td></tr></tbody></table> </ephtml> </p> <hd id="AN0186992462-29">Feedback quality results</hd> <p>This section details quantitative descriptions of the distribution of feedback quality ratings and a typology of peer comments that synthesizes across quality criteria.</p> <hd id="AN0186992462-30">Descriptive statistics</hd> <p>To examine the peer feedback practices of first‐year engineering students, we calculated descriptive statistics of the rubric ratings for each category. Table 8 shows the central tendency and variation information, while Figure 1 displays the distribution of rubric scores for each feedback quality criterion.</p> <p>8 TABLE Descriptive statistics of feedback quality results.</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left" /><th align="left">Task</th><th align="left">Behavior</th><th align="left">Gap</th><th align="left">Action</th></tr></thead><tbody valign="top"><tr><td align="left">Model</td></tr><tr><td align="left">Mean</td><td align="left">1.05</td><td align="left">1.50</td><td align="left">0.45</td><td align="left">0.33</td></tr><tr><td align="left">SD</td><td align="left">0.43</td><td align="left">0.55</td><td align="left">0.55</td><td align="left">0.53</td></tr><tr><td align="left">IQR</td><td align="left">0</td><td align="left">1</td><td align="left">1</td><td align="left">1</td></tr><tr><td align="left">Researchers</td></tr><tr><td align="left">Mean</td><td align="left">0.99</td><td align="left">0.89</td><td align="left">0.56</td><td align="left">0.31</td></tr><tr><td align="left">SD</td><td align="left">0.76</td><td align="left">0.77</td><td align="left">0.73</td><td align="left">0.62</td></tr><tr><td align="left">IQR</td><td align="left">2</td><td align="left">1</td><td align="left">1</td><td align="left">0</td></tr></tbody></table> </ephtml> </p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/6M4/01jul25/jee70024-fig-0001.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jee70024-fig-0001.jpg" title="1 Comparison of Rubric Score Distributions for Feedback Quality by Researchers and LLM. The "AI‐" prefix denotes the distribution from the LLM, while "R‐" denotes researchers." /> </p> <p></p> <p>These figures together show that peer feedback quality in the first‐year PBL context is mostly in the lower ranges. Only the AI mean score for Behavior is substantially above a score of 1, and Figure 1 shows that researchers were much more conservative with their Behavior scores than the LLM. Another notable pattern is a difference in scores between the positive criteria (i.e., Task and Behavior) and the constructive criteria (i.e., Gap and Action). More than 60% of all comments were rated 0 in Gap, and around 70% of all comments were rated 0 in the Action criteria. Conversely, the maximum fraction of 0 scores for positive criteria was below 40% (36.1% for R‐Behavior).</p> <hd id="AN0186992462-32">Peer comment typology</hd> <p>While the distribution of scores from each rubric criterion is useful for identifying trends, it does not reveal how student comments perform over all four criteria. Based on patterns in our dataset of peer feedback comments, we created a typology of peer comments that combines all four rubric criteria. The typology consists of five types of peer comments that commonly appear in our data. The types are organized from low to high feedback quality.</p> <p>The first type in the feedback quality typology is <emph>Overall Low Quality</emph>. These comments are typically one sentence or less and receive a 0 or 1 score in every criterion. While some comments in this category are obvious, such as "none" or "N/A", others can initially seem reasonable. For example, the comment "Gets their portion of the work completed on time" may seem like a quality comment at first, but when compared to the rubric, the comment has limited elements that score a 1 at most in the Task criterion and 0 for all other criteria. A key element of Overall Low Quality comments is that there is nothing for the feedback recipient to learn from the comment. Based on researcher ratings, 20.8% of comments had a total score (sum of four criterion scores) of 0 or 1, which approximates the density of Overall Low Quality comments.</p> <p>Another common type of peer comment is <emph>Only Negative</emph>. A comment that is Only Negative does not score above a 0 in both Task and Behavior and only contains Gap‐ or Action‐related content. These comments typically come from teammates who are frustrated with the performance of their peers and use the peer evaluation to list grievances. An example of an Only Negative comment is</p> <p>They are good, and they would never do something to purposefully undermine the team, although they treat teamwork just as they do their own work. This would be okay if we didn't all have to plan around this and thus get work done last second. Although I believe this to be true, I can't say that I am much better because I understand it myself at times. We can both work on this.</p> <p>This comment scored a 2 for Gap, 1 for Action, and 0 for the other two criteria. Often, Only Negative comments contain only constructive or critical feedback, but this is not essential for a comment to be classified as Only Negative. A comment can have a short reference to a positive contribution or behavior, but it is not detailed enough to be a valuable addition. Furthermore, Only Negative comments sometimes recommend actions to remedy the identified gap(s), but this is rare.</p> <p>The analog to Only Negative comments is <emph>Only Positive</emph> comments. These comments score well on the Task and Behavior criteria but lack details on Gap and Action. These comments can have relatively high total scores (3–4) by scoring 2s in Task and Behavior, but they do not provide constructive feedback. After correcting for comments with zeros in every criterion, 54.9% of comments had zeros for both Gap and Action, indicating that they fall into the Only Positive type. This comment is a representative example of the Only Positive type:</p> <p>They have been a great help throughout this class/project. They are good at keeping everyone awake and on track. They are also good with CAD, so they have been a great help to me. I appreciate that they have been willing to work with me and I hope this continues.</p> <p>The fourth type of peer comment quality is a <emph>Descriptive</emph> comment. A comment that does a good job of describing the positive and negative aspects of a peer's performance but does not make any recommendations for how to improve is a Descriptive comment. In other words, this type of comment does not have the Action component of quality feedback. Descriptive comments are of overall good quality, but the lack of an action recommendation can hamper their utility for the recipient. For example, this comment scores well in all other criteria:</p> <p>They get their work done and do it well. They have missed some meetings, but even then, they stay informed when they aren't there and do their part for all of our team assignments. They always have a great attitude and great input.</p> <p>This comment contains specific details about the communication gap but does not suggest what the recipient should change. Descriptive comments have a nonzero score in Task, Behavior, and Gap, so their total scores can range from 3 to 6.</p> <p>The final type of comments performs well on all four quality criteria, so they are named <emph>Overall High Quality</emph>. These comments have nonzero scores in all criteria, usually with multiple 2s. Overall High Quality comments are typically much longer than other types, often bordering on a paragraph. A prime example of a comment that exemplifies high quality is this comment with a total score of 8 (researcher rating):</p> <p>Person A is present for every day of the class and has contributed heavily to group work. Person A has contributed to all team documents up to present day and has also taken initiative from my comments to assist other members. Person A sometimes misses discord notifications which causes some late work; however, Person A's output is a considerable benefit for the team, nonetheless. Person A also takes the initiative to reach out to the other members which is how we got crucial information about Person B's situation. Person A also frequently delivers critique about project goal reachability which is beneficial for increasing the foundation of the project. For instance, Person A has critiqued my design about autonomy, construction, and complexity which has given me the chance to rethink a couple topics and improve my design. I would like to see Person A improve a couple aspects; I would like Person A to be more aware of discord notifications as well as complete work on time. This is mainly due to teamwork being submitted with unified consensus.</p> <p>Less than 3% of comments in our dataset achieved a total score of 7 or above, so Overall High Quality comments are a small minority.</p> <hd id="AN0186992462-33">DISCUSSION</hd> <p>This study sought to answer two research questions, one methodological and one theoretical. Based on our analysis of the reliability of an LLM as a rater for a qualitative rubric, the results present an encouraging view of the capability of GAI in educational research. For our theoretical contribution to the study of peer feedback in first‐year engineering courses, we present feedback quality as an integral part of students' development of feedback literacy.</p> <hd id="AN0186992462-34">Large language model reliability</hd> <p>Our methodological research question is, "How reliable is a Large Language Model in applying a rubric to assess the quality of engineering student peer feedback comments compared to researchers?" This question is not as simple as the IRR results presented in Table 5. Reliability depends upon both the similarity of the ultimate rubric score between the LLM and human raters and the coherence of the justification for the score.</p> <hd id="AN0186992462-35">Ranking alignment</hd> <p>Considering the IRR values, the GAI model results mimicked the trends of the researchers. Researchers had lower IRR for Task and Behavior, and these criteria also had the lowest kappa values for the LLM. The model had lower IRR overall, particularly in the Behavior category, where the Rating2 versus Model IRR is 0.34 lower than the researcher comparison (Rating1 vs. Rating2).</p> <p>By studying the model's distribution of assigned scores, we can understand more about how the model applied the feedback quality rubric. The model tended to be more generous with subjective judgments about detail. In general, a 0 score had no detail, a 1 score on our rubric contained only vague detail, and a 2 score contained specific detail. The range of "vague" detail was much broader for the LLM than for the researchers, particularly in the Task and Behavior criteria. Figure 1 shows that the model assigned roughly 80% of comments a 1 score for Task, while the researchers only assigned 1s about 40% of the time. The broad application of 1s was less pronounced for the Gap and Action criteria. Overall, the dataset contained fewer references to performance gaps and actions, so the scores were primarily zeros for Gap and Action. If there is no mention of a Gap or Action, there is no subjective judgment required to assign a score. This lack of content could be the source of improved IRR results for Gap and Action.</p> <hd id="AN0186992462-36">Process considerations</hd> <p>Because of the process we used to run data through the GAI model, the model is most analogous to a novice rater. The prompt gives specific instructions on what to do with the peer comment, but there is little contextual information. The only feedback quality information given to the LLM is the rubric criteria definitions and score definitions (Table 3). Due to our methods, it was as if the model was applying the rubric for the first time for every comment. Conversely, the human raters were researchers who are familiar with feedback literature, created the rubric, and participated in prior rounds of coding with the data. Comparing the LLM to the researchers is akin to comparing novices to experts, which has obvious problems. When the researchers conducted their first round of rating, albeit on a slightly different rubric, the IRR values were closer to the IRR values from the model. For example, our initial IRR was lower for the Task than the Action criteria (Task κ = 0.52, Action κ = 0.69) (Drinkwater et al., [<reflink idref="bib20" id="ref78">20</reflink>]), which aligns with patterns in the model's IRR with Rating1 (Task κ = 0.39, Action κ = 0.86). While the model's IRR values are lower for Task, it is encouraging that the Action IRR achieved an "almost perfect" designation. As researchers, we had the ability to discuss these initial results for the Task criterion and identify discrepancies in our individual understanding of the rubric. With the LLM, we are able to read the model's justification only after the rating was complete. With the novice nature of the model in mind, the IRR values seem promising.</p> <p>Another contributing difference in the rating process between researchers and the LLM was double‐counting comment excerpts for Task and Behavior. When establishing the rubric application rules for researchers, we agreed that any part of a peer comment could contribute to a score in only one criterion. For example, the same phrase could not count for a 1 in both Task and Behavior. However, when running the LLM, each criterion was rated individually. This meant that the model evaluating a comment for Task had no idea what the same comment had scored for Behavior, Gap, or Action. Sometimes, this resulted in a comment phrase counting for scores in two criteria (i.e., double‐counting). After identifying this problem in our early iterations of prompt engineering, we added exclusionary criteria to the definitions for Task and Behavior. The italicized text in Table 3 is the language added to the rubric to help prevent double‐counting. The additions to the criteria definitions improved but did not eliminate the double‐counting of positive feedback for both Task and Behavior. Double‐counting could also potentially explain why the mean scores for Task and Behavior were higher with the model (AI‐Task <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:10694730:media:jee70024:jee70024-math-0001" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mover accent="true"><mi mathvariant="normal">x</mi><mo>¯</mo></mover></mrow><annotation encoding="application/x-tex">$$ \overline{\mathrm{x}} $$</annotation></semantics></math> </ephtml>  = 1.05, AI‐Behavior <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:10694730:media:jee70024:jee70024-math-0002" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mover accent="true"><mi mathvariant="normal">x</mi><mo>¯</mo></mover></mrow><annotation encoding="application/x-tex">$$ \overline{\mathrm{x}} $$</annotation></semantics></math> </ephtml>  = 1.50) than with the researchers (R‐Task <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:10694730:media:jee70024:jee70024-math-0003" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mover accent="true"><mi mathvariant="normal">x</mi><mo>¯</mo></mover></mrow><annotation encoding="application/x-tex">$$ \overline{\mathrm{x}} $$</annotation></semantics></math> </ephtml>  = 0.999, R‐Behavior <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:10694730:media:jee70024:jee70024-math-0004" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mover accent="true"><mi mathvariant="normal">x</mi><mo>¯</mo></mover></mrow><annotation encoding="application/x-tex">$$ \overline{\mathrm{x}} $$</annotation></semantics></math> </ephtml>  = 0.890).</p> <p>Overall, the LLM performed worse on IRR in comparison to researchers in the Task and Behavior criteria but had comparable reliability in Gap and Action. These trends mimic the researchers and replicate the results from the researchers' first round of coding (Drinkwater et al., [<reflink idref="bib20" id="ref79">20</reflink>]). These trends show the trade‐offs of methodological choices. In our case, we chose to run each criterion individually and use the LLM as a novice rater. While we believe these choices contributed to more coherent processing by the LLM, it did result in issues with double‐counting and novice rating mistakes. The following section examines the role of researchers in ensuring the reliability of data analysis when using GAI.</p> <hd id="AN0186992462-37">Ethical considerations</hd> <p>The results of our comparison between the ratings of the LLM and human researchers highlight important similarities and differences between the processes. As researchers consider how to ethically include the capabilities of GAI tools in their work, acknowledging the roles of human researchers and LLMs is crucial. The model applied the rubric similarly to novice researchers, and the justifications provided were coherent and logical. In areas where human researchers struggled with subjectivity, the model also struggled to consistently apply the rubric. This evidence suggests that LLMs, when tested appropriately, can be a reliable tool to apply the deductive rubric. Our results for a deductive task add a nuance to the encouraging results of other studies in EER that have used LLMs for inductive tasks, such as creating a codebook for qualitative data (e.g., Hingle et al., [<reflink idref="bib28" id="ref80">28</reflink>]; Katz, Fleming, & Main, [<reflink idref="bib32" id="ref81">32</reflink>]). The issues encountered with score discrepancies in our study suggest that LLMs could be most useful in the initial stages of the research process. Refining a codebook or rubric, for example, is a task that could use an LLM to understand unclear definitions or areas of subjectivity. If an LLM is going to be used as the rater in data analysis, our experience indicates that a highly refined rubric is necessary to ensure consistent and reliable results.</p> <p>Another limitation of the model that should be considered is the interpretation of grammatically incorrect comments. Researchers ignored small grammatical errors by assuming the desired meaning of comments, but the LLM processed comments the way they were written. For example, this comment uses imperative verb tense instead of declarative tense:</p> <p>Person A is a group organizer and is good at setting a schedule for the group members. Pay attention to the working status of each member. Provide input to team members when necessary. Listen to group members' ideas and revise the overall goal...</p> <p>Researchers assumed that the feedback writer meant to include a subject before the verb (e.g. "<emph>They</emph> provide input to team members when necessary..."), which means the last three sentences describe positive behaviors. The LLM took a different interpretation and counted these imperative statements as Action. It is difficult to tell what the writer originally meant, but the LLM's literal interpretation of comments could be a limitation for grammatically imperfect data, such as interview transcripts. Maintaining the voice of the participants is of utmost importance in qualitative research. It is up to the researcher to choose the methods with which to distill participant meaning. This simple grammatical disparity provides an example of how small differences between the researcher and LLM interpretation can change how the participant's meaning is interpreted.</p> <p>Many additional questions remain about the reasonable application of GAI in EER. Our challenges with double‐counting are only one example of how the process of entering data, running the model, and analyzing the data can greatly impact the results. Many tools at the disposal of researchers (e.g., group discussion, instantaneous clarification, contextual knowledge) remain difficult to implement with LLMs. The intellectual and collaborative capabilities of researchers highlight the ethical necessity of maintaining human direction in the research process (Davison et al., [<reflink idref="bib17" id="ref82">17</reflink>]). Currently, the implementation of GAI still requires a considerable amount of manual labor, but as the technology improves, it is essential to keep researchers at the center of decision making.</p> <hd id="AN0186992462-38">Feedback quality</hd> <p>The theoretical research question we sought to answer with the help of GAI is, "What do feedback quality measures reveal about first‐year engineering students' ability to provide feedback?" Both the analysis of scores from each criterion on the rubric and the comprehensive typology of common quality patterns help answer this question. By connecting the results back to frameworks of feedback literacy and self‐regulated learning theory, we reveal challenges, opportunities, and processes in first‐year engineering students' provision of peer feedback.</p> <hd id="AN0186992462-39">Challenges in providing feedback</hd> <p>The quantitative analysis of each criterion reveals the generally low quality of peer feedback comments. The average scores for every category were between 0 and 1 (based on researcher ratings), and the average score for Action was only 0.31. Furthermore, 50.1% of all the numerical ratings assigned by researchers across the four criteria were 0. Figure 1 shows that the scores for criteria that represent positive feedback (i.e., Task and Behavior) have overall higher scores than criteria that represent constructive feedback (i.e., Gap and Action). While both positive and constructive feedback are important, identifying performance gaps is crucial to self‐regulated reflection and improvement.</p> <p>Looking at these statistics alone does not paint an optimistic picture of the quality of peer feedback, which follows previous studies that have posited that peer feedback may not be eliciting constructive, specific feedback that can inform students' regulation of teamwork behaviors (Burgess et al., [<reflink idref="bib8" id="ref83">8</reflink>]). The overall lack of quality feedback provided by students can be explained from several angles. Winstone et al. ([<reflink idref="bib60" id="ref84">60</reflink>]) claim that instructors often lack the knowledge to explicitly teach students how to provide feedback. Students may also lack motivation to write quality feedback if they feel that it will have no effect on their team dynamics. Previous studies have indicated that feedback tends to be skimmed or ignored more than it is processed and acted upon (Gibbs & Simpson, [<reflink idref="bib26" id="ref85">26</reflink>]). If there is no benefit tied to writing quality feedback for teammates, students are disincentivized to spend time and energy providing feedback. Feedback tied to assessment also suffers from increased bias, as students are reluctant to negatively impact the grades of their peers (Sridharan et al., [<reflink idref="bib56" id="ref86">56</reflink>]). A related explanation for the lack of high‐quality feedback is a lack of students' ability to adequately judge the performance of their peers. Carless and Boud ([<reflink idref="bib10" id="ref87">10</reflink>]) list Making Judgments as one of the four key features of feedback literacy. First‐year students are often new to peer evaluation and feedback in an engineering context, so the development of their comparative judgment is nascent. Engaging in the full process of effective feedback is necessary for students to accurately judge the performance of their peers and themselves (Carless & Boud, [<reflink idref="bib10" id="ref88">10</reflink>]).</p> <hd id="AN0186992462-40">Opportunities for feedback literacy development</hd> <p>While the quantitative analysis of feedback quality scores displays an overall lack of quality, the typology of comments, however, provides more nuance and direction. There are certainly some students who write "No comment" or have Overall Low Quality comments, but less than 21% of comments fit in this category. A majority of students put effort into writing comments that will help their peers, which shows a desire to provide and receive good feedback. Appreciating and seeking feedback are important constructs of feedback literacy (Dawson et al., [<reflink idref="bib18" id="ref89">18</reflink>]), and with this initial buy‐in, there is an opportunity to guide students toward providing higher quality feedback.</p> <p>The typology of peer comments describes how students think about quality feedback. Overall High Quality comments include positive statements about performance and constructive criticism with recommended actions. Only Negative and Only Positive comments reveal that students often think that feedback is either negative or positive but not both. It is well known that both positive and constructive feedback are valuable for team and individual performance. The four criteria in our rubric assert that there is always something positive and something constructive that can (and should) be said about first‐year engineering students. In order to provide a high‐quality comment, the author must include positive and constructive elements. This aligns with existing feedback frameworks that highlight balance as the key to quality feedback (Camarata & Slieman, [<reflink idref="bib9" id="ref90">9</reflink>]; Hattie & Timperley, [<reflink idref="bib27" id="ref91">27</reflink>]). Furthermore, feedback on successes and areas of improvement is essential for effective self‐reflection, a key phase of self‐regulated learning (Zimmerman, [<reflink idref="bib64" id="ref92">64</reflink>]).</p> <p>The Descriptive comment type in the typology reveals a slightly different theme in student peer feedback; comments often place the goal‐setting phase of self‐regulation on the feedback recipient. Descriptive comments explain what a student is doing but do not suggest corrective action. It is then up to the feedback recipient to interpret what they should continue doing and what to change about their performance. This interpretation, however, may or may not align with what the feedback author intended. Environmental influences are important to the self‐observation element of self‐regulation because they allow the learner to collect data on their performance, but these influences are not helpful if they cannot be translated into a behavioral adjustment (Nicol & Macfarlane‐Dick, [<reflink idref="bib42" id="ref93">42</reflink>]). Encouraging students to identify Gap and Action in peer comments supports the development of self‐regulation and increases the likelihood of improved team dynamics. Furthermore, the practice of receiving constructive feedback can help students to prepare for common practices in the engineering industry, such as performance evaluations (Kremer & Burnette, [<reflink idref="bib34" id="ref94">34</reflink>]). Providing constructive feedback is clearly something that students need to practice, as more than 56% of comments in our data had zeros for both Gap and Action.</p> <p>Another opportunity to support feedback literacy development and build teamwork behaviors is to distinguish between Task and Behavior. This distinction was a point of difficulty for both the LLM and researchers in the analysis process, which begs the question of their true distinctiveness in the feedback quality rubric. Most frameworks of educational feedback focus only on relevant tasks because they are geared toward instructors giving feedback on specific assignments (e.g., Hattie & Timperley, [<reflink idref="bib27" id="ref95">27</reflink>]; Prins et al., [<reflink idref="bib48" id="ref96">48</reflink>]; Sadler, [<reflink idref="bib51" id="ref97">51</reflink>]). In an engineering education PBL setting, it is important to focus on the long‐term development of the student. Teamwork behaviors are developed over courses and often re‐developed in new courses and settings (Wei et al., [<reflink idref="bib59" id="ref98">59</reflink>]). Rating comments based only on the specific tasks that helped students gain points on a design assignment ignores key teamwork behaviors. Furthermore, students write about teamwork behaviors in peer evaluation more often than contributions to specific tasks. The rubric would miss a major component of peer comments and several ABET‐required outcomes if we only measured Task.</p> <p>By examining peer feedback comments through a typology, the perceptions of first‐year engineering students regarding what feedback should look like are illuminated. Most students put effort into providing peer comments but miss one or more quality criteria. The patterns of quality identify opportunities for further study of peer feedback processes and improvements in pedagogical interventions that teach students how to provide feedback.</p> <hd id="AN0186992462-41">Feedback as a process</hd> <p>Modern frameworks of feedback literacy are built upon the idea that feedback is a student‐centered process (Molloy & Boud, [<reflink idref="bib38" id="ref99">38</reflink>]). While grounded in higher education and assessment literature, current conceptions of feedback literacy require more grounding in learning theory (Chong, [<reflink idref="bib12" id="ref100">12</reflink>]). This paper provides Bandura's theory of self‐regulation to ground the learning process that students engage with during the feedback process. The focus of this analysis is feedback quality, which is recognized in the Provision of Feedback construct of Dawson et al.′s framework of feedback literacy ([<reflink idref="bib17" id="ref101">17</reflink>]). However, Dawson et al.′s and other frameworks do not describe what elements of quality are necessary for learners to develop feedback literacy. This paper contributes a rubric, typology, and process for characterizing peer feedback quality and links quality criteria to self‐regulated learning, which then informs the development of feedback literacy.</p> <p>Self‐regulated learning is a sociocultural perspective that relates environmental influences to cognitive processes through a process of self‐observation, goal setting, and behavioral adjustment (Bandura, [<reflink idref="bib4" id="ref102">4</reflink>]). The four criteria in our rubric follow this process. The Task and Behavior criteria describe what the learner is currently doing, which aids self‐observation by introducing data from peers. The Gap criterion informs goal setting by identifying an area of improvement, and Action is a behavioral adjustment. In short, the feedback quality criteria address the steps of self‐regulation. A comment rated as Overall High Quality through our rubric contains all the necessary information for a student to engage in the steps of self‐regulation. Through engagement in self‐regulation, students also engage in the development of feedback literacy by improving their comparative judgment and taking action (Carless & Boud, [<reflink idref="bib10" id="ref103">10</reflink>]).</p> <p>Our model of feedback quality also aligns with other emerging sociocultural perspectives of feedback literacy. Chong ([<reflink idref="bib12" id="ref104">12</reflink>]) outlined a socioecological framework of feedback literacy that identifies three dimensions of learning for feedback literacy: individual, engagement, and contextual. The individual dimension accounts for individual differences (e.g., beliefs, goals, abilities, previous experience) between learners that impact how they interact with feedback. Our results point to the importance of the individual dimension because, even though all students participated in the same in‐class feedback training, the peer comments still represented a wide variety of quality. The engagement dimension is an adapted version of Carless and Boud's framework and focuses on the learner's cognitive engagement, affective engagement, and behavioral engagement. Self‐regulation theory emphasizes reflection as a way for learners to engage cognitively and affectively with feedback (Nicol & Macfarlane‐Dick, [<reflink idref="bib42" id="ref105">42</reflink>]). Our grounding in self‐regulation theory aligns with Coppens et al. ([<reflink idref="bib15" id="ref106">15</reflink>]) and a forthcoming article from this study that assess the impact of reflection to support feedback literacy. The final dimension of Chong's framework is the contextual dimension, which addresses the manner in which the feedback process is conducted. Chong is primarily concerned with instructor feedback, so the contextual dimensions focus on the instructor–learner relationship. By focusing on peer feedback literacy, this study suggests team‐focused contextual factors. For example, the instructions given to students about peer feedback were specifically designed to encourage high‐quality feedback. Additionally, the timing of the evaluations (Week 7 and Week 11) was tailored to the timeline of the course project to provide timely formative feedback.</p> <p>Feedback literacy is a growing area of research in higher education (Chong, [<reflink idref="bib12" id="ref107">12</reflink>]) but is understudied in engineering. With the prevalence of peer evaluation in project‐based learning environments in engineering education, theory to support feedback processes is needed to inform pedagogical and research choices. This paper translates constructs of feedback quality from frameworks of feedback literacy to peer feedback in the engineering context. Self‐regulation theory is employed to operationalize the feedback process as a form of self‐regulated learning. Sociocultural theory is commonly employed in engineering education (Newsetter & Svinicki, [<reflink idref="bib41" id="ref108">41</reflink>]), so this connection makes frameworks of feedback literacy more accessible to the field. Finally, the tools to measure and characterize feedback quality that are presented in this study link the content included in a peer feedback comment to the greater iterative skill of feedback literacy.</p> <hd id="AN0186992462-42">Limitations</hd> <p>As with any research study, our results have limitations. Data limitations were mentioned in the Methods section but are also summarized here. Our dataset used peer comments from one semester of a first‐year engineering course. First‐year students are typically novices with peer evaluation and feedback, which could lead to lower quality comments. Conversely, the instructor of the course studies feedback and incorporates specific training and instruction for peer evaluation into the course. This is an uncommon practice, so students in this dataset may possess more knowledge about feedback than other first‐year students.</p> <p>Some limitations arose from the process used to incorporate the LLM, namely running each rubric criterion individually. Our initial tests suggested that breaking the model's task down into four segments would lead to higher model attention and higher performance as a result. A side effect of this choice was the double‐counting problem described in the discussion.</p> <p>It should also be acknowledged that this study used only one LLM. We initially tested five models, but only qwen‐2.5‐32b was used for all 295 comments. There is a possibility that another model could have performed better (or worse) with our dataset. However, this model was chosen because of its performance and size. Larger models take more processing time and computing power, which may be less accessible to researchers in engineering education. Consequently, we recommend running pilot tests with whatever model researchers choose to identify similar and/or different limitations.</p> <hd id="AN0186992462-43">IMPLICATIONS</hd> <p>The results of this study have several implications for stakeholders in educational research. We present implications related to our methodological findings regarding GAI and implications for those interested in peer feedback quality.</p> <hd id="AN0186992462-44">Using GAI for qualitative research</hd> <p>Considering the balanced results of our experimentation with the implementation of an LLM as a rater of qualitative comments, there are several implications for educational researchers who are using or wish to use LLMs in their work. The most salient takeaway from the numerous iterations of our process is that using LLMs for qualitative research is far from automatic. We iterated on the rubric three times, edited the prompts more than seven times, and piloted five models. Each of these changes was in response to a problem we identified by reviewing model output from a test run. While our iterations slowed the research process, they significantly improved our methods by increasing consistency and meeting our goals. We did not have the space to include every iteration of our process here, nor did we think they are relevant to other researchers. Each project has unique requirements, and, therefore, processes should be refined by the researchers familiar with the project. Our recommendation for researchers is to utilize LLMs as a sounding board and tester for analysis instruments. Whether refining a codebook, creating a rubric, or preparing to disseminate materials, LLMs can be a valuable tool for rapid iteration. An LLM can help researchers to identify discrepancies in understanding, areas of subjectivity, or unclear directions. The model acts as a novice researcher who does not have the background knowledge of experienced researchers. In actual analysis, this can be a weakness, but it is a strength in research design.</p> <p>For researchers planning to use LLMs to apply a rubric or for other deductive techniques, specificity is essential. Any subjectivity or ambiguity introduces volatility in the model's application of the rubric. We learned this through the model's broad definition of the word "vague." The complexity of the data is also an important factor when creating instruments and writing prompts. The longer the data, the more specific and deductive the rubric or codebook should be. The average length of comments in our dataset was only 43 words, and we had difficulty with ambiguity in the application of our rubric. Another key consideration is the technical logistics of choosing a model. We exclusively used open‐source, local models to ensure the security of the process and the data. When choosing a model, control over the back‐end processes, data privacy, and computational resources are all important considerations. We recommend running tests with a small subset of data to assess whether the process and results fit within the desired parameters.</p> <hd id="AN0186992462-45">Assessing peer feedback quality</hd> <p>Our investigation of peer feedback quality in the first‐year engineering PBL context creates several implications for engineering instructors and those concerned with feedback literacy. As with other professional skills, explicit instruction is an impactful way to support the development of feedback writing skills. Our results suggest that students need to be instructed on how to include effective constructive criticism along with positive feedback. This follows previous studies (Drinkwater et al., [<reflink idref="bib20" id="ref109">20</reflink>]; Huang et al., [<reflink idref="bib31" id="ref110">31</reflink>]; Loignon et al., [<reflink idref="bib36" id="ref111">36</reflink>]) that found a significant improvement in feedback quality after implementing a pedagogical intervention to teach students how to write feedback. Beyond telling students what constitutes good feedback, matching assessment mechanisms to quality metrics signals the importance of putting effort into writing high‐quality peer feedback. Adding an extrinsic motivator to write high‐quality feedback creates an additional reason for students to learn to write quality peer comments. Furthermore, a commonly cited reason for a lack of constructive feedback is that students are afraid to be thorough and honest in peer evaluations for fear of impacting their teammates' grades (Forsythe & Johnson, [<reflink idref="bib22" id="ref112">22</reflink>]). Assessing feedback <emph>quality</emph> instead of feedback <emph>content</emph> gives students more control over their teamwork grades, thus reducing pressure to inflate the ratings of their peers.</p> <p>Knowing how to write high‐quality feedback is only one element of feedback literacy. Recent frameworks in higher education (e.g., Chong, [<reflink idref="bib12" id="ref113">12</reflink>]; Dawson et al., [<reflink idref="bib18" id="ref114">18</reflink>]) demonstrate that, to maximize the benefit of the feedback process, students must be adept at writing, receiving, reflecting on, and responding to feedback. Most feedback research in engineering education has focused on one of these elements without considering the process holistically. Furthermore, most studies have limited grounding in educational theory. Investigating feedback literacy with a basis in learning theories, such as self‐regulation, is an important step toward understanding and realizing the impact of peer evaluation and feedback in engineering PBL settings.</p> <hd id="AN0186992462-46">CONCLUSIONS</hd> <p>This study conducted an intertwined analysis of the capability of an LLM to apply a rubric of feedback quality to peer comments and the landscape of feedback quality for first‐year engineering students in a project‐based course. The methodological vein of the paper utilized an existing rubric of peer feedback quality with four dimensions: Task, Behavior, Gap, and Action. Researchers and an LLM rated 295 peer comments with the rubric. Inter‐rater reliability between the model and researchers approached reasonable levels for the Gap and Action criteria but struggled to distinguish between Task and Behavior. Based on a review of the model's justifications for rubric scores, we believe the low reliability for Task and Behavior is caused by increased subjectivity in the criteria definitions, which impacted both the model and researcher raters. The theoretical vein sought to understand what feedback quality measures revealed about the feedback writing practices of first‐year engineering students through the lens of self‐regulation theory. Comments were more likely to have positive reinforcement than constructive criticism, and a majority of comments did not identify gaps in performance or suggest improvements. A typology of peer feedback quality was created based on patterns in the dataset. This study presents encouraging evidence for the capability of LLMs to accelerate research design and act as a valuable pilot test for codebook or rubric development, but caution must be taken to ensure that the process is verified before using LLMs for data analysis. Both the areas of AI‐supported qualitative analysis and peer feedback are ripe for innovation and further research. As this article (and this special issue) shows, the combination of new methodological approaches and theoretical questions can create synergistic results.</p> <hd id="AN0186992462-47">ACKNOWLEDGMENTS</hd> <p>The authors would like to acknowledge Marin Fisher Hale for her valuable contributions in rubric development and researcher coding of the data. We also thank all participants who agreed to share their peer comments.</p> <p>As part of a special issue for <emph>JEE</emph> focused on the research applications of GAI, this study incorporated the use of AI in several aspects of the research process. An open‐source local LLM was used to de‐identify data and apply a rubric to short excerpts of qualitative data. Iteration and researcher oversight were utilized at every stage of data analysis to ensure the reliability of results. No GAI tools were used to draft or edit this paper.</p> <hd id="AN0186992462-48">A APPENDIX De‐identification Prompt</hd> <p>"You are an expert text analyst." You specialize in following instructions perfectly. I need your help. I will provide you instructions in the <instructions> tag and examples in the <examples> tag. You should use the instructions and examples to help accomplish your task, which you will perform on the {data_type} in the <{data_type} > tag.</p> <p><instructions></p> <p>Your important task is to de‐identify the following {data_type} in the <data_type> tag by replacing all names and pronouns with aliases. The aliases should be "Person A", "Person B", et cetera. The goal of this de‐identification task is to have the {data_type} so that it is not possible to identify people referenced in the {data_type} other than the author of the {data_type} while retaining all other information exactly as it is written in the {data_type}.</p> <p>To accomplish your task, follow these steps:</p> <p></p> <ulist> <item> Read the {data_type} carefully.</item> <p></p> <item> Provide "Initial observations" where you identify names, pronouns, and any potential challenges in de‐identifying the text.</item> <p></p> <item> Copy the original {data_type} verbatim and replace names and pronouns that refer to a person in the {data_type} with aliases.</item> <p></p> <item> Do not change first person singular and plural pronouns that refer to the writer (e.g., I, me, we, etc.). These are not considered to be identifiable information and therefore should remain unchanged.</item> <p></p> <item> Do not change references to the team, the group, teammates (in a general sense), etc. These are not considered to be identifiable information, and therefore we do not need to de‐identify it.</item> </ulist> <p>Regarding formatting, your response should have two main sections:</p> <p></p> <ulist> <item> "My Initial expert observations:" followed by your notes on names, pronouns, and potential challenges.</item> <p></p> <item> "The de‐identified {data_type} is:" followed by the {data_type} that you have perfectly de‐identified by replacing the names and pronouns referencing other people with aliases that you have assigned to them.</item> </ulist> <p></instructions></p> <p>Now that you have studied these instructions, here are some examples to help you understand the task. Be aware that there are six examples. Each example has the original {data_type}, initial observations, the de‐identified version of the {data_type}, and an explanation about why the de‐identification is correct.</p> <p><examples></p> <p>Example 1. This example demonstrates the proper use of aliases (Person A, Person B) instead of generic pronouns, and shows how to handle multiple people in a single passage. Example 1 also shows how not to change a reference to the team.</p> <p>Original: "John and Sarah worked on the project together. He focused on coding while she handled the design. They made a great team."</p> <p>Initial observations: Names identified: John, Sarah. Pronouns to replace: He, she, They.</p> <p>De‐identified: "Person A and Person B worked on the project together. Person A focused on coding while Person B handled the design. Person A and Person B made a great team."</p> <p>Example 2. This example shows the preservation of first‐person pronouns referring to the writer ("We"), the correct handling of possessive pronouns ("her" becomes "Person A's"), multiple aliases (Person A, Person B), and not changing a reference to the team.</p> <p>Original: "The team met to discuss the project. Mary presented her ideas, and Tom took notes. We agreed on the next steps."</p> <p>Initial observations: Names identified: Mary, Tom. Pronouns to replace: her. "We" should be preserved as it refers to the writer. "The team" should not be changed.</p> <p>De‐identified: "The team met to discuss the project. Person A presented Person A's ideas, and Person B took notes. We agreed on the next steps."</p> <p>Example 3. This example illustrates the correct de‐identification of uncommon names (Adya and G‐man), ensuring that all names are replaced regardless of their familiarity.</p> <p>Original: "Adya and G‐man collaborated on the report. She wrote the introduction, while he analyzed the data."</p> <p>Initial observations: Names identified: Adya, G‐man (uncommon names). Pronouns to replace: She, he.</p> <p>De‐identified: "Person A and Person B collaborated on the report. Person A wrote the introduction, while Person B analyzed the data."</p> <p>Example 4. This example demonstrates the proper de‐identification when multiple names appear in a sentence, and shows how to replace "They" when it refers to specific individuals.</p> <p>Original: "I spoke with Katie and Olivia about the new initiative. They both had great insights."</p> <p>Initial observations: Names identified: Katie, Olivia. Pronouns to replace: They. "I" should be preserved as it refers to the writer.</p> <p>De‐identified: "I spoke with Person A and Person B about the new initiative. Person A and Person B both had great insights."</p> <p>Example 5. This example shows the preservation of collective nouns like "the group" and demonstrates how to handle pronouns referring to these collective nouns without introducing aliases.</p> <p>Original: "The group decided to postpone the meeting. They felt more preparation was needed."</p> <p>Initial observations: No specific names identified. "The group" should not be changed. "They" refers to the group and should be replaced with "The group".</p> <p>De‐identified: "The group decided to postpone the meeting. The group felt more preparation was needed."</p> <p>Example 6. This example demonstrates that when there are no explicit names or pronouns referring to individuals, the text should remain unchanged. It is crucial not to infer or add pronouns or aliases where none originally existed. Example 6 also shows how to maintain a generic reference to "the team".</p> <p>Original: "Communicates well, helps the team."</p> <p>Initial observations: No names or pronouns identified. No changes needed.</p> <p>De‐identified: "Communicates well, helps the team."</p> <p></examples></p> <p>Now that you have studied your instructions and six examples with their explanations, here is the actual {data_type} for you to de‐identify by replacing names and pronouns with aliases.</p> <p><{data_type}></p> <p>{text}</p> <p></{data_type}></p> <p>Begin your task now by first providing your initial observations, followed by the de‐identified {data_type}. Here is a full analysis template for you to use for your task.</p> <p>FULL ANALYSIS TEMPLATE</p> <p>My Initial expert observations:</p> <p>[your expert observations about which names and pronouns do require changing and which names and pronouns should not be changed].</p> <p>The de‐identified {data_type} is:</p> <p>[the {data_type} that you have perfectly de‐identified by replacing the names and pronouns referencing other people with aliases that you have assigned to them].</p> <hd id="AN0186992462-49">B APPENDIX Rubric Evaluation Prompt</hd> <p>You are an expert in educational assessment. You have famous skills in your ability to analyze {data_type}s and apply a grading rubric for evaluation purposes. You also have a reputation for being very strict in your interpretation and application of those grading rubrics. You have an important task I will describe for you in the <instructions> XML tag. Be aware that your instructions contain task‐specific instructions and output format instructions in their own respective XML tags. Before your instructions, you will be given the rubric in a < rubric> XML tag and the {data_type} to analyze using the rubric.</p> <p>First, here is the rubric for you to use.</p> <p><rubric></p> <p>{rubric}</p> <p></rubric></p> <p>Now that you have studied the rubric that you must use, here is the {data_type} in the <written_feedback> XML tag for you to analyze.</p> <p><written_feedback></p> <p>{text}</p> <p></written_feedback></p> <p>Now that you have studied the rubric and {data_type} to analyze, here are your instructions.</p> <p><instructions></p> <p><task_instructions></p> <p>You are evaluating written feedback from student teamwork settings. This evaluation is part of a larger assessment process aimed at improving students' ability to provide constructive feedback to their peers. Your task is to analyze the written feedback provided in the <written_feedback> XML tag using the rubric in the <rubric> XML tag. Here is the series of steps you should apply to accomplish this task.</p> <p>EVALUATION PROCESS:</p> <p>Step 1. Carefully read the rubric and the written feedback.</p> <p>Step 2. Apply the rubric criteria step‐by‐step to the feedback.</p> <p>Step 3. Provide specific examples from the feedback to justify your evaluation.</p> <p>Step 4. Reflect on any challenges or ambiguities you encountered during the evaluation.</p> <p>Step 5. Assign a score based on your analysis.</p> <p>Step 6. Indicate your confidence level in your scoring decision.</p> <p>Step 7. Double‐check your analysis for consistency and adherence to the rubric.</p> <p></task_instructions></p> <p><output_format_instructions></p> <p>You must structure your response as a JSON object with the following keys:</p> <p>‐ "step_by_step_analysis": A detailed, step‐by‐step reasoning of your analysis</p> <p>‐ "justification": Specific examples cited from the feedback to support your evaluation</p> <p>‐ "reflection": Discussion of any challenges or ambiguities encountered</p> <p>‐ "score_and_confidenceConfidence": An object containing:</p> <p>‐ "score": Your suggested score (as a number)</p> <p>‐ "confidence_level": Your confidence level (as a string: "low", "medium", or "high")</p> <p>‐ "final_check": A final check where you determine whether you need to make any corrections to your analysis</p> <p>Ensure that your response is a valid JSON object. Use proper JSON syntax, including quotes around keys and string values, and commas to separate elements.</p> <p>Here's an example of the expected JSON structure:</p> <p>{{</p> <p>"step_by_step_analysis": "[step‐by‐step analysis of your reasoning about the student feedback, rubric, and how to apply the rubric]",</p> <p>"justification": "[explanation of why you might assign a certain rubric score]",</p> <p>"reflection": "[your reflection on your response, identifying places of ambiguity and certainty]",</p> <p>"score_and_confidence": {{</p> <p>"score": "[rubric score]",</p> <p>"confidence_level": "low|medium|high".</p> <p>}},</p> <p>"final_check": "[check on weakest parts of your response]".</p> <p>}}</p> <p>Your response should follow this structure, but with more detailed content based on your analysis.</p> <p></output_format_instructions></p> <p></instructions></p> <p>Take a moment to carefully consider the rubric and the written feedback. Remember to be objective, consistent, and thorough in your evaluation. When you are ready, provide your analysis following the JSON structure outlined in the instructions. Your entire response should be a single, valid JSON object.</p> <ref id="AN0186992462-50"> <title> REFERENCES </title> <blist> <bibl id="bib1" idref="ref68" type="bt">1</bibl> <bibtext> ABET. (2023). Criteria for accrediting engineering programs, 2024–2025. ABET. https://<ulink href="http://www.abet.org/accreditation/accreditation-criteria/criteria-for-accrediting-engineering-programs-2024-2025/">www.abet.org/accreditation/accreditation-criteria/criteria-for-accrediting-engineering-programs-2024-2025/</ulink></bibtext> </blist> <blist> <bibl id="bib2" idref="ref60" type="bt">2</bibl> <bibtext> Abraham, R. M., & Singaram, V. S. (2019). Using deliberate practice framework to assess the quality of feedback in undergraduate clinical skills training. BMC Medical Education, 19 (1), 105. https://doi.org/10.1186/s12909-019-1547-5</bibtext> </blist> <blist> <bibl id="bib3" idref="ref5" type="bt">3</bibl> <bibtext> Abram, M. D., Mancini, K. T., & Parker, R. D. (2020). Methods to integrate natural language processing into qualitative research. International Journal of Qualitative Methods, 19, 1 – 6. https://doi.org/10.1177/1609406920984608</bibtext> </blist> <blist> <bibl id="bib4" idref="ref54" type="bt">4</bibl> <bibtext> Bandura, A. (1991). Social cognitive theory of self‐regulation. Organizational Behavior and Human Decision Processes, 50 (2), 248 – 287. https://doi.org/10.1016/0749-5978(91)90022-L</bibtext> </blist> <blist> <bibl id="bib5" idref="ref35" type="bt">5</bibl> <bibtext> Beagon, Ú., Niall, D., & Ní Fhloinn, E. (2019). Problem‐based learning: Student perceptions of its value in developing professional skills for engineering practice. European Journal of Engineering Education, 44 (6), 850 – 865. https://doi.org/10.1080/03043797.2018.1536114</bibtext> </blist> <blist> <bibl id="bib6" idref="ref23" type="bt">6</bibl> <bibtext> Bender, E. M., Gebru, T., McMillan‐Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models Be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency (pp. 610 – 623). Association for Computing Machine. https://doi.org/10.1145/3442188.3445922</bibtext> </blist> <blist> <bibl id="bib7" idref="ref34" type="bt">7</bibl> <bibtext> Borrego, M., Froyd, J. E., & Hall, T. S. (2010). Diffusion of engineering education innovations: A survey of awareness and adoption rates in U.S. engineering departments. Journal of Engineering Education, 99 (3), 185 – 207. https://doi.org/10.1002/j.2168-9830.2010.tb01056.x</bibtext> </blist> <blist> <bibl id="bib8" idref="ref41" type="bt">8</bibl> <bibtext> Burgess, A., Roberts, C., Lane, A. S., Haq, I., Clark, T., Kalman, E., Pappalardo, N., & Bleasel, J. (2021). Peer review in team‐based learning: Influencing feedback literacy. BMC Medical Education, 21 (1), 426. https://doi.org/10.1186/s12909-021-02821-6</bibtext> </blist> <blist> <bibl id="bib9" idref="ref90" type="bt">9</bibl> <bibtext> Camarata, T., & Slieman, T. A. (2020). Improving student feedback quality: A simple model using peer review and feedback rubrics. Journal of Medical Education and Curricular Development, 7, 2382120520936604. https://doi.org/10.1177/2382120520936604</bibtext> </blist> <blist> <bibtext> Carless, D., & Boud, D. (2018). The development of student feedback literacy: Enabling uptake of feedback. Assessment & Evaluation in Higher Education, 43 (8), 1315 – 1325. https://doi.org/10.1080/02602938.2018.1463354</bibtext> </blist> <blist> <bibtext> Carver, C. S., & Scheier, M. F. (1982). Control theory: A useful conceptual framework for personality–social, clinical, and health psychology. Psychological Bulletin, 92 (1), 111 – 135. https://doi.org/10.1037/0033-2909.92.1.111</bibtext> </blist> <blist> <bibtext> Chong, S. W. (2021). Reconsidering student feedback literacy from an ecological perspective. Assessment & Evaluation in Higher Education, 46 (1), 92 – 104. https://doi.org/10.1080/02602938.2020.1730765</bibtext> </blist> <blist> <bibtext> Cohen, J. (1968). Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit. Psychological Bulletin, 70 (4), 213 – 220. https://doi.org/10.1037/h0026256</bibtext> </blist> <blist> <bibtext> Coppens, K., Van den Broeck, L., Winstone, N., & Langie, G. (2023). Capturing student feedback literacy using reflective logs. European Journal of Engineering Education, 48 (4), 653 – 666. https://doi.org/10.1080/03043797.2023.2185501</bibtext> </blist> <blist> <bibtext> Coppens, K., Van den Broeck, L., Winstone, N., & Langie, G. (2024). A mixed method approach to exploring feedback literacy through student self‐reflection. Assessment & Evaluation in Higher Education, 50 (2), 173 – 186. https://doi.org/10.1080/02602938.2024.2373792</bibtext> </blist> <blist> <bibtext> Cruz, M. L., Saunders‐Smits, G. N., & Groen, P. (2020). Evaluation of competency methods in engineering education: A systematic review. European Journal of Engineering Education, 45 (5), 729 – 757. https://doi.org/10.1080/03043797.2019.1671810</bibtext> </blist> <blist> <bibtext> Davison, R. M., Chughtai, H., Nielsen, P., Marabelli, M., Iannacci, F., Van Offenbeek, M., Tarafdar, M., Trenz, M., Techatassanasoontorn, A. A., Díaz Andrade, A., & Panteli, N. (2024). The ethics of using generative AI for qualitative data analysis. Information Systems Journal, 34 (5), 1433 – 1439. https://doi.org/10.1111/isj.12504</bibtext> </blist> <blist> <bibtext> Dawson, P., Yan, Z., Lipnevich, A., Tai, J., Boud, D., & Mahoney, P. (2024). Measuring what learners do in feedback: The feedback literacy behaviour scale. Assessment & Evaluation in Higher Education, 49 (3), 348 – 362. https://doi.org/10.1080/02602938.2023.2240983</bibtext> </blist> <blist> <bibtext> Doldor, E., Wyatt, M., & Silvester, J. (2021). Research: Men get more actionable feedback than women. Harvard Business Review. https://hbr.org/2021/02/research-men-get-more-actionable-feedback-than-women</bibtext> </blist> <blist> <bibtext> Drinkwater, K., Ryan, O., Sajadi, S., Huerta, M., & Fisher, M. (2024). Improving peer feedback in project‐based learning contexts: An investigation into a first‐year engineering intervention. 2024 ASEE Annual Conference & Exposition. American Society of Engineering Education.</bibtext> </blist> <blist> <bibtext> Dwivedi, Y. K., Kshetri, N., Hughes, L., Slade, E. L., Jeyaraj, A., Kar, A. K., Baabdullah, A. M., Koohang, A., Raghavan, V., Ahuja, M., Albanna, H., Albashrawi, M. A., Al‐Busaidi, A. S., Balakrishnan, J., Barlette, Y., Basu, S., Bose, I., Brooks, L., Buhalis, D., ... Wright, R. (2023). Opinion paper: "So what if ChatGPT wrote it?" multidisciplinary perspectives on opportunities, challenges and implications of generative conversational AI for research, practice and policy. International Journal of Information Management, 71, 102642. https://doi.org/10.1016/j.ijinfomgt.2023.102642</bibtext> </blist> <blist> <bibtext> Forsythe, A., & Johnson, S. (2017). Thanks, but no‐thanks for the feedback. Assessment & Evaluation in Higher Education, 42 (6), 850 – 859. https://doi.org/10.1080/02602938.2016.1202190</bibtext> </blist> <blist> <bibtext> Fui‐Hoon Nah, F., Zheng, R., Cai, J., Siau, K., & Chen, L. (2023). Generative AI and ChatGPT: Applications, challenges, and AI‐human collaboration. Journal of Information Technology Case and Application Research, 25 (3), 277 – 304. https://doi.org/10.1080/15228053.2023.2233814</bibtext> </blist> <blist> <bibtext> Gamieldien, Y., Case, J. M., & Katz, A. (2023). Advancing qualitative analysis: An exploration of the potential of generative AI and NLP in thematic coding (SSRN scholarly paper 4487768). https://doi.org/10.2139/ssrn.4487768</bibtext> </blist> <blist> <bibtext> Gauthier, S., Cavalcanti, R., Goguen, J., & Sibbald, M. (2015). Deliberate practice as a framework for evaluating feedback in residency training. Medical Teacher, 37 (6), 551 – 557. https://doi.org/10.3109/0142159X.2014.956059</bibtext> </blist> <blist> <bibtext> Gibbs, G., & Simpson, C. (2004). Conditions under which assessment supports students' learning. Learning and Teaching in Higher Education, 1, 3 – 31.</bibtext> </blist> <blist> <bibtext> Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research, 77 (1), 81 – 112. https://doi.org/10.3102/003465430298487</bibtext> </blist> <blist> <bibtext> Hingle, A., Katz, A., & Johri, A. (2023). Exploring NLP‐based methods for generating engineering ethics assessment qualitative codebooks. 2023 IEEE Frontiers in Education Conference (FIE), 1–8. https://doi.org/10.1109/FIE58773.2023.10342985</bibtext> </blist> <blist> <bibtext> Hochwald, I. H., Green, G., Sela, Y., Radomyslsky, Z., Nissanholtz‐Gannot, R., & Hochwald, O. (2023). Converting qualitative data into quantitative values using a matched mixed‐methods design: A new methodological approach. Journal of Advanced Nursing, 79 (11), 4398 – 4410. https://doi.org/10.1111/jan.15649</bibtext> </blist> <blist> <bibtext> Holmes, W., & Miao, F. (2023). Guidance for generative AI in education and research. UNESCO Publishing.</bibtext> </blist> <blist> <bibtext> Huang, W., Wynkoop, R., Exter, M., & Berry, F. (2022). Feedback matters: Self‐and‐Peer assessment made better with instructional interventions. 2022 ASEE Annual Conference & Exposition. https://peer.asee.org/feedback-matters-self-and-peer-assessment-made-better-with-instructional-interventions</bibtext> </blist> <blist> <bibtext> Katz, A., Fleming, G. C., & Main, J. (2024). Thematic analysis with open‐source generative AI and machine learning: A new method for inductive qualitative codebook development (arXiv:2410.03721). arXiv. https://doi.org/10.48550/arXiv.2410.03721</bibtext> </blist> <blist> <bibtext> Katz, A., Gerhardt, M., & Soledad, M. (2024). Using generative text models to create qualitative codebooks for student evaluations of teaching (arXiv:2403.11984). arXiv. https://doi.org/10.48550/arXiv.2403.11984</bibtext> </blist> <blist> <bibtext> Kremer, G., & Burnette, D. (2008). Using performance reviews in capstone design courses for development and assessment of professional skills. 13.1349.1‐13.1349.11. https://peer.asee.org/using-performance-reviews-in-capstone-design-courses-for-development-and-assessment-of-professional-skills</bibtext> </blist> <blist> <bibtext> Liu, J., & Jagadish, H. V. (2024). Institutional efforts to help academic researchers implement generative AI in research. Harvard Data Science Review. https://doi.org/10.1162/99608f92.2c8e7e81</bibtext> </blist> <blist> <bibtext> Loignon, A. C., Woehr, D. J., Thomas, J. S., Loughry, M. L., Ohland, M. W., & Ferguson, D. M. (2017). Facilitating peer evaluation in team contexts: The impact of frame‐of‐reference rater training. Academy of Management Learning & Education, 16 (4), 562 – 578. https://doi.org/10.5465/amle.2016.0163</bibtext> </blist> <blist> <bibtext> Mentzer, N., Jackson, A., Richards, K. A., Zissimopoulos, A. N., & Laux, D. (2015). Student perceptions on the impact of formative peer team member effectiveness evaluation in an introductory design course. 26.1422.1‐26.1422.20. https://peer.asee.org/student-perceptions-on-the-impact-of-formative-peer-team-member-effectiveness-evaluation-in-an-introductory-design-course.</bibtext> </blist> <blist> <bibtext> Molloy, E., & Boud, D. (2012). Rethinking models of feedback for learning: The challenge of design. Assessment & Evaluation in Higher Education, 38 (6), 1 – 15. https://doi.org/10.1080/02602938.2012.691462</bibtext> </blist> <blist> <bibtext> Molloy, E., Boud, D., & Henderson, M. (2020). Developing a learning‐centred framework for feedback literacy. Assessment & Evaluation in Higher Education, 45 (4), 527 – 540. https://doi.org/10.1080/02602938.2019.1667955</bibtext> </blist> <blist> <bibtext> Morris, M. R. (2023). Scientists' perspectives on the potential for generative AI in their fields (arXiv:2304.01420). arXiv. https://doi.org/10.48550/arXiv.2304.01420</bibtext> </blist> <blist> <bibtext> Newsetter, W., & Svinicki, M. (2014). Learning theories for engineering education practice. In A. Johri & B. Olds (Eds.), Cambridge handbook of engineering education research (pp. 29 – 46). Cambridge University Press.</bibtext> </blist> <blist> <bibtext> Nicol, D., & Macfarlane‐Dick, D. (2006). Formative assessment and self‐regulated learning: A model and seven principles of good feedback practice. Studies in Higher Education, 31 (2), 199 – 218. https://doi.org/10.1080/03075070600572090</bibtext> </blist> <blist> <bibtext> Nicol, D., Thomson, A., & Breslin, C. (2014). Rethinking feedback practices in higher education: A peer review perspective. Assessment & Evaluation in Higher Education, 39 (1), 102 – 122. https://doi.org/10.1080/02602938.2013.795518</bibtext> </blist> <blist> <bibtext> Nieminen, J. H., Tai, J., Boud, D., & Henderson, M. (2022). Student agency in feedback: Beyond the individual. Assessment & Evaluation in Higher Education, 47 (1), 95 – 108. https://doi.org/10.1080/02602938.2021.1887080</bibtext> </blist> <blist> <bibtext> Ohland, M. W., Loughry, M. L., Woehr, D. J., Bullard, L. G., Felder, R. M., Finelli, C. J., Layton, R. A., Pomeranz, H. R., & Schmucker, D. G. (2012). The comprehensive assessment of team member effectiveness: Development of a behaviorally anchored rating scale for self‐ and peer evaluation. Academy of Management Learning & Education, 11 (4), 609 – 630. https://doi.org/10.5465/amle.2010.0177</bibtext> </blist> <blist> <bibtext> Olea, C., Tucker, H., Phelan, J., Pattison, C., Zhang, S., Lieb, M., Schmidt, D., & White, J. (2024). Evaluating persona prompting for question answering tasks, 63–81. https://doi.org/10.5121/csit.2024.141106</bibtext> </blist> <blist> <bibtext> Owoahene Acheampong, K., & Nyaaba, M. (2024). Review of qualitative research in the Era of generative artificial intelligence (SSRN Scholarly Paper 4686920). https://doi.org/10.2139/ssrn.4686920</bibtext> </blist> <blist> <bibtext> Prins, F. J., Sluijsmans, D. M. A., & Kirschner, P. A. (2006). Feedback for general practitioners in training: Quality, styles, and preferences. Advances in Health Sciences Education, 11 (3), 289 – 303. https://doi.org/10.1007/s10459-005-3250-z</bibtext> </blist> <blist> <bibtext> Rietz, T., & Maedche, A. (2021). Cody: An AI‐based system to semi‐automate coding for qualitative research. Proceedings of the 2021 CHI conference on human factors in computing systems, 1–14. https://doi.org/10.1145/3411764.3445591</bibtext> </blist> <blist> <bibtext> Ryan, O., Fisher, M. J., Schibelius, L., Huerta, M. V., & Sajadi, S. (2023). Using a scenario‐based learning approach with instructional technology to teach conflict management to engineering students. 2023 ASEE Annual Conference & Exposition. https://peer.asee.org/using-a-scenario-based-learning-approach-with-instructional-technology-to-teach-conflict-management-to-engineering-students</bibtext> </blist> <blist> <bibtext> Sadler, D. R. (1989). Formative assessment and the design of instructional systems. Instructional Science, 18 (2), 119 – 144. https://doi.org/10.1007/BF00117714</bibtext> </blist> <blist> <bibtext> Sajadi, S., Huerta, M. V., Ryan, O. J., & Drinkwater, K. (2024). Harnessing generative AI to enhance feedback quality in peer evaluations within project‐based learning contexts. International Journal of Engineering Education, 40 (5), 998 – 1012.</bibtext> </blist> <blist> <bibtext> Secules, S., McCall, C., Mejia, J. A., Beebe, C., Masters, A. S., Sánchez‐Peña, L. M., & Svyantek, M. (2021). Positionality practices and dimensions of impact on equity research: A collaborative inquiry and call to the community. Journal of Engineering Education, 110 (1), 19 – 43. https://doi.org/10.1002/jee.20377</bibtext> </blist> <blist> <bibtext> Shute, V. J. (2008). Focus on formative feedback. Review of Educational Research, 78 (1), 153 – 189. https://doi.org/10.3102/0034654307313795</bibtext> </blist> <blist> <bibtext> Sinha, R., Solola, I., Nguyen, H., Swanson, H., & Lawrence, L. (2024). The role of generative AI in qualitative research: GPT‐4's contributions to a grounded theory analysis. Proceedings of the 2024 Symposium on Learning, Design and Technology, 17–25. https://doi.org/10.1145/3663433.3663456</bibtext> </blist> <blist> <bibtext> Sridharan, B., Tai, J., & Boud, D. (2019). Does the use of summative peer assessment in collaborative group work inhibit good judgement? Higher Education, 77 (5), 853 – 870. https://doi.org/10.1007/s10734-018-0305-7</bibtext> </blist> <blist> <bibtext> The Apache Software Foundation. (2024). Apache License, version 2.0. The Apache Software Foundation. https://<ulink href="http://www.apache.org/licenses/LICENSE-2.0">www.apache.org/licenses/LICENSE-2.0</ulink></bibtext> </blist> <blist> <bibtext> Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E. H., Le, Q. V., & Zhou, D. (2024). Chain‐of‐thought prompting elicits reasoning in large language models. Proceedings of the 36th international conference on neural information processing systems, 24824–24837.</bibtext> </blist> <blist> <bibtext> Wei, S., Zhou, C., & Ohland, M. W. (2021). Longitudinal effects of team‐based training on students' peer rating quality. 2021 ASEE Virtual Annual Conference Content Access. https://peer.asee.org/longitudinal-effects-of-team-based-training-on-students-peer-rating-quality</bibtext> </blist> <blist> <bibtext> Winstone, N. E., Nash, R. A., Parker, M., & Rowntree, J. (2017). Supporting learners' agentic engagement with feedback: A systematic review and a taxonomy of Recipience processes. Educational Psychologist, 52 (1), 17 – 37. https://doi.org/10.1080/00461520.2016.1207538</bibtext> </blist> <blist> <bibtext> Zhan, Y. (2022). Developing and validating a student feedback literacy scale. Assessment & Evaluation in Higher Education, 47 (7), 1087 – 1100. https://doi.org/10.1080/02602938.2021.2001430</bibtext> </blist> <blist> <bibtext> Zhang, L., & Ma, Y. (2023). A study of the impact of project‐based learning on student learning effects: A meta‐analysis study. Frontiers in Psychology, 14, 1202728. https://doi.org/10.3389/fpsyg.2023.1202728</bibtext> </blist> <blist> <bibtext> Zhang, M., Lindsay, E. D., Thorbensen, F. B., Poulsen, D. B., & Bjerva, J. (2024). Leveraging large language models for actionable course evaluation student feedback to lecturers (arXiv:2407.01274). arXiv. https://doi.org/10.48550/arXiv.2407.01274</bibtext> </blist> <blist> <bibtext> Zimmerman, B. J. (2002). Becoming a self‐regulated learner: An overview. Theory into Practice, 41 (2), 64 – 70. https://doi.org/10.1207/s15430421tip4102_2</bibtext> </blist> </ref> <aug> <p>By Katherine Drinkwater Gregg; Olivia Ryan; Andrew Katz; Mark Huerta and Susan Sajadi</p> <p>Reported by Author; Author; Author; Author; Author</p> <p></p> <p>Katherine Drinkwater Gregg is a PhD candidate in the Department of Engineering Education at Virginia Tech, 635 Prices Fork Rd., Blacksburg, VA 24061, USA..</p> <p>Olivia Ryan is a PhD candidate in the Department of Engineering Education at Virginia Tech, 635 Prices Fork Rd., Blacksburg, VA 24061, USA..</p> <p>Andrew Katz is an Associate Professor in the Department of Engineering Education at Virginia Tech, 635 Prices Fork Rd., Blacksburg, VA 24061, USA..</p> <p>Mark Huerta is an Assistant Professor in the Department of Engineering Education at Virginia Tech, 635 Prices Fork Rd., Blacksburg, VA 24061, USA..</p> <p>Susan Sajadi is an Assistant Professor in the Department of Engineering Education at Virginia Tech, 635 Prices Fork Rd., Blacksburg, VA 24061, USA..</p> </aug> <nolink nlid="nl1" bibid="bib47" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib24" firstref="ref2"></nolink> <nolink nlid="nl3" bibid="bib49" firstref="ref3"></nolink> <nolink nlid="nl4" bibid="bib30" firstref="ref4"></nolink> <nolink nlid="nl5" bibid="bib16" firstref="ref6"></nolink> <nolink nlid="nl6" bibid="bib10" firstref="ref7"></nolink> <nolink nlid="nl7" bibid="bib14" firstref="ref8"></nolink> <nolink nlid="nl8" bibid="bib45" firstref="ref9"></nolink> <nolink nlid="nl9" bibid="bib15" firstref="ref10"></nolink> <nolink nlid="nl10" bibid="bib61" firstref="ref11"></nolink> <nolink nlid="nl11" bibid="bib18" firstref="ref12"></nolink> <nolink nlid="nl12" bibid="bib20" firstref="ref13"></nolink> <nolink nlid="nl13" bibid="bib21" firstref="ref16"></nolink> <nolink nlid="nl14" bibid="bib35" firstref="ref17"></nolink> <nolink nlid="nl15" bibid="bib40" firstref="ref18"></nolink> <nolink nlid="nl16" bibid="bib55" firstref="ref20"></nolink> <nolink nlid="nl17" bibid="bib23" firstref="ref22"></nolink> <nolink nlid="nl18" bibid="bib33" firstref="ref24"></nolink> <nolink nlid="nl19" bibid="bib63" firstref="ref25"></nolink> <nolink nlid="nl20" bibid="bib28" firstref="ref26"></nolink> <nolink nlid="nl21" bibid="bib32" firstref="ref27"></nolink> <nolink nlid="nl22" bibid="bib39" firstref="ref32"></nolink> <nolink nlid="nl23" bibid="bib38" firstref="ref33"></nolink> <nolink nlid="nl24" bibid="bib50" firstref="ref36"></nolink> <nolink nlid="nl25" bibid="bib62" firstref="ref37"></nolink> <nolink nlid="nl26" bibid="bib37" firstref="ref38"></nolink> <nolink nlid="nl27" bibid="bib60" firstref="ref39"></nolink> <nolink nlid="nl28" bibid="bib27" firstref="ref42"></nolink> <nolink nlid="nl29" bibid="bib56" firstref="ref43"></nolink> <nolink nlid="nl30" bibid="bib43" firstref="ref44"></nolink> <nolink nlid="nl31" bibid="bib54" firstref="ref46"></nolink> <nolink nlid="nl32" bibid="bib12" firstref="ref50"></nolink> <nolink nlid="nl33" bibid="bib44" firstref="ref51"></nolink> <nolink nlid="nl34" bibid="bib11" firstref="ref55"></nolink> <nolink nlid="nl35" bibid="bib64" firstref="ref56"></nolink> <nolink nlid="nl36" bibid="bib25" firstref="ref57"></nolink> <nolink nlid="nl37" bibid="bib51" firstref="ref58"></nolink> <nolink nlid="nl38" bibid="bib29" firstref="ref63"></nolink> <nolink nlid="nl39" bibid="bib52" firstref="ref64"></nolink> <nolink nlid="nl40" bibid="bib57" firstref="ref69"></nolink> <nolink nlid="nl41" bibid="bib46" firstref="ref71"></nolink> <nolink nlid="nl42" bibid="bib58" firstref="ref72"></nolink> <nolink nlid="nl43" bibid="bib19" firstref="ref73"></nolink> <nolink nlid="nl44" bibid="bib13" firstref="ref75"></nolink> <nolink nlid="nl45" bibid="bib53" firstref="ref76"></nolink> <nolink nlid="nl46" bibid="bib17" firstref="ref82"></nolink> <nolink nlid="nl47" bibid="bib26" firstref="ref85"></nolink> <nolink nlid="nl48" bibid="bib42" firstref="ref93"></nolink> <nolink nlid="nl49" bibid="bib34" firstref="ref94"></nolink> <nolink nlid="nl50" bibid="bib48" firstref="ref96"></nolink> <nolink nlid="nl51" bibid="bib59" firstref="ref98"></nolink> <nolink nlid="nl52" bibid="bib41" firstref="ref108"></nolink> <nolink nlid="nl53" bibid="bib31" firstref="ref110"></nolink> <nolink nlid="nl54" bibid="bib36" firstref="ref111"></nolink> <nolink nlid="nl55" bibid="bib22" firstref="ref112"></nolink>
Header DbId: eric
DbLabel: ERIC
An: EJ1478628
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Expanding Possibilities for Generative AI in Qualitative Analysis: Fostering Student Feedback Literacy through the Application of a Feedback Quality Rubric
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Katherine+Drinkwater+Gregg%22">Katherine Drinkwater Gregg</searchLink> (ORCID <externalLink term="https://orcid.org/0009-0002-5998-9231">0009-0002-5998-9231</externalLink>)<br /><searchLink fieldCode="AR" term="%22Olivia+Ryan%22">Olivia Ryan</searchLink> (ORCID <externalLink term="https://orcid.org/0009-0007-7981-6131">0009-0007-7981-6131</externalLink>)<br /><searchLink fieldCode="AR" term="%22Andrew+Katz%22">Andrew Katz</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-3554-9015">0000-0002-3554-9015</externalLink>)<br /><searchLink fieldCode="AR" term="%22Mark+Huerta%22">Mark Huerta</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-2962-0724">0000-0003-2962-0724</externalLink>)<br /><searchLink fieldCode="AR" term="%22Susan+Sajadi%22">Susan Sajadi</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-8511-7467">0000-0001-8511-7467</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Journal+of+Engineering+Education%22"><i>Journal of Engineering Education</i></searchLink>. 2025 114(3).
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 31
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2025
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Audience
  Label: Education Level
  Group: Audnce
  Data: <searchLink fieldCode="EL" term="%22Higher+Education%22">Higher Education</searchLink><br /><searchLink fieldCode="EL" term="%22Postsecondary+Education%22">Postsecondary Education</searchLink>
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Technology+Uses+in+Education%22">Technology Uses in Education</searchLink><br /><searchLink fieldCode="DE" term="%22Engineering+Education%22">Engineering Education</searchLink><br /><searchLink fieldCode="DE" term="%22Student+Evaluation%22">Student Evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Peer+Evaluation%22">Peer Evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Teamwork%22">Teamwork</searchLink><br /><searchLink fieldCode="DE" term="%22Formative+Evaluation%22">Formative Evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Feedback+%28Response%29%22">Feedback (Response)</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+Language+Processing%22">Natural Language Processing</searchLink><br /><searchLink fieldCode="DE" term="%22Scoring+Rubrics%22">Scoring Rubrics</searchLink><br /><searchLink fieldCode="DE" term="%22College+Freshmen%22">College Freshmen</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+Mediated+Communication%22">Computer Mediated Communication</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluators%22">Evaluators</searchLink><br /><searchLink fieldCode="DE" term="%22Man+Machine+Systems%22">Man Machine Systems</searchLink><br /><searchLink fieldCode="DE" term="%22Interrater+Reliability%22">Interrater Reliability</searchLink><br /><searchLink fieldCode="DE" term="%22Literacy%22">Literacy</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1002/jee.70024
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 1069-4730<br />2168-9830
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Background: Courses in engineering often use peer evaluation to monitor teamwork behaviors and team dynamics. The qualitative peer comments written for peer evaluations hold potential as a valuable source of formative feedback for students, yet little is known about their content and quality. Purpose: This study uses a large language model (LLM) to apply a previously tested feedback quality rubric to peer feedback comments. Our research questions interrogate the reliability of LLMs for qualitative analysis with a rubric and use Bandura's self-regulated learning theory to assess peer feedback quality of first-year engineering students' comments. Method: An open-source, local LLM was used to score each comment according to four rubric criteria. Inter-rater reliability (IRR) with human raters using Cohen's quadratic weighted kappa was the primary metric of reliability. Our assessment of peer feedback quality utilized descriptive statistics. Results: The LLM achieved lower IRR than human raters, but the model's challenges mimic those of human raters. The model did achieve an excellent quadratic weighted kappa of 0.80 for one rubric criterion, which shows promise for LLM capability. For feedback quality, students generally wrote low- to medium-quality comments that were infrequently grounded in specific teamwork behaviors. We identified five types of peer feedback that inform how students perceive the feedback process. Conclusions: Our implementation of GAI suggests that LLMs can be helpful for rapid iteration of research designs, but consistent and reliable analysis with generative artificial intelligence (GAI) requires significant effort and testing. To develop feedback literacy, students must understand how to provide high-quality feedback.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2025
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1478628
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1478628
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1002/jee.70024
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 31
    Subjects:
      – SubjectFull: Artificial Intelligence
        Type: general
      – SubjectFull: Technology Uses in Education
        Type: general
      – SubjectFull: Engineering Education
        Type: general
      – SubjectFull: Student Evaluation
        Type: general
      – SubjectFull: Peer Evaluation
        Type: general
      – SubjectFull: Teamwork
        Type: general
      – SubjectFull: Formative Evaluation
        Type: general
      – SubjectFull: Feedback (Response)
        Type: general
      – SubjectFull: Natural Language Processing
        Type: general
      – SubjectFull: Scoring Rubrics
        Type: general
      – SubjectFull: College Freshmen
        Type: general
      – SubjectFull: Computer Mediated Communication
        Type: general
      – SubjectFull: Evaluators
        Type: general
      – SubjectFull: Man Machine Systems
        Type: general
      – SubjectFull: Interrater Reliability
        Type: general
      – SubjectFull: Literacy
        Type: general
    Titles:
      – TitleFull: Expanding Possibilities for Generative AI in Qualitative Analysis: Fostering Student Feedback Literacy through the Application of a Feedback Quality Rubric
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Katherine Drinkwater Gregg
      – PersonEntity:
          Name:
            NameFull: Olivia Ryan
      – PersonEntity:
          Name:
            NameFull: Andrew Katz
      – PersonEntity:
          Name:
            NameFull: Mark Huerta
      – PersonEntity:
          Name:
            NameFull: Susan Sajadi
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 07
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 1069-4730
            – Type: issn-electronic
              Value: 2168-9830
          Numbering:
            – Type: volume
              Value: 114
            – Type: issue
              Value: 3
          Titles:
            – TitleFull: Journal of Engineering Education
              Type: main
ResultId 1