A Framework for Using SET When Evaluating Faculty

Saved in:
Bibliographic Details
Title: A Framework for Using SET When Evaluating Faculty
Language: English
Authors: Nina Zipser (ORCID 0000-0002-1431-1606), Lisa Mincieli (ORCID 0000-0001-7971-8338)
Source: Studies in Higher Education. 2025 50(1):168-182.
Availability: Routledge. Available from: Taylor & Francis, Ltd. 530 Walnut Street Suite 850, Philadelphia, PA 19106. Tel: 800-354-1420; Tel: 215-625-8900; Fax: 215-207-0050; Web site: http://www.tandf.co.uk/journals
Peer Reviewed: Y
Page Count: 15
Publication Date: 2025
Document Type: Journal Articles
Reports - Research
Education Level: Higher Education
Postsecondary Education
Descriptors: College Faculty, Faculty Evaluation, Student Evaluation of Teacher Performance, Evaluation Criteria, Evaluation Methods, Evaluation Problems, Evaluation Utilization, Data Interpretation
DOI: 10.1080/03075079.2024.2332412
ISSN: 0307-5079
1470-174X
Abstract: This paper presents a framework for utilizing Student Evaluations of Teaching (SET) in faculty evaluations. Recognizing the ongoing debate about the validity of SET as a measure of teaching effectiveness, the authors agree with scholars who propose viewing SET as a tool for gauging 'student perceptions of learning'. They present a method that combines post-estimation residuals from an Ordinary Least Squares (OLS) model, which accounts for key confounding factors, with simple mean SET scores. This approach provides a more nuanced perspective on each faculty member's individual ratings. The authors draw on SET data from courses taught by tenured faculty in the Faculty of Arts and Sciences at a large, private university in the Northeastern United States. Their findings suggest that the proposed residual framework can help address some of the concerns associated with the use of SET in personnel decision-making. However, they caution that this approach requires careful explanation and should be used in conjunction with a comprehensive review of a professor's full teaching portfolio. The study contributes to the ongoing discourse on the use of SET in faculty evaluations, offering a more equitable and nuanced approach to interpreting SET data.
Abstractor: As Provided
Entry Date: 2025
Accession Number: EJ1455581
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwHGgfrqlnqG_qNHt9zIiTtwAAAA4jCB3wYJKoZIhvcNAQcGoIHRMIHOAgEAMIHIBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDNEJiTornk-3aGhqCAIBEICBmjkmnseKJnmi2B8v5mUC0YTrPTYBXPDZckwl-MkBNXhzfPkpemAguqsurpER7lMTb0Zxib3pF0sWRe8Uh1MqwXo59PnepERNiOffJFFjsTdoeEGBDKnZGbmgfySl9Mu6lA8vmiuey_cCIRcRgbJdm5II8_2ylC26mn_ktb8wTD843PhqD-dXWqUl5GBjQoxHJ5SwhngAMljxmMk=
Text:
  Availability: 1
  Value: <anid>AN0181947092;she01jan.25;2025Jan01.02:52;v2.2.500</anid> <title id="AN0181947092-1">A framework for using SET when evaluating faculty </title> <p>This paper presents a framework for utilizing Student Evaluations of Teaching (SET) in faculty evaluations. Recognizing the ongoing debate about the validity of SET as a measure of teaching effectiveness, the authors agree with scholars who propose viewing SET as a tool for gauging 'student perceptions of learning'. They present a method that combines post-estimation residuals from an Ordinary Least Squares (OLS) model, which accounts for key confounding factors, with simple mean SET scores. This approach provides a more nuanced perspective on each faculty member's individual ratings. The authors draw on SET data from courses taught by tenured faculty in the Faculty of Arts and Sciences at a large, private university in the Northeastern United States. Their findings suggest that the proposed residual framework can help address some of the concerns associated with the use of SET in personnel decision-making. However, they caution that this approach requires careful explanation and should be used in conjunction with a comprehensive review of a professor's full teaching portfolio. The study contributes to the ongoing discourse on the use of SET in faculty evaluations, offering a more equitable and nuanced approach to interpreting SET data.</p> <p>Keywords: Student evaluations of teaching; postestimation predictions; faculty evaluation; bias</p> <hd id="AN0181947092-2">I. Introduction</hd> <p>Student Evaluations of Teaching (SET) frequently play a role in faculty assessments and the allocation of teaching awards. However, the academic community is divided over whether SET accurately gauges teaching effectiveness, and consequently, if they should influence personnel decisions (Kreitzer and Sweet-Cushman [<reflink idref="bib17" id="ref1">17</reflink>]). Given the widespread use of SET by institutions, we concur with Kreitzer and Sweet-Cushman ([<reflink idref="bib17" id="ref2">17</reflink>]) perspective that it is more precise and therefore, useful, to view SET as a tool for measuring 'student perceptions of learning' rather than a direct indicator of teaching effectiveness.</p> <p>SET allow universities to economically collect student feedback on their classroom learning experiences, providing faculty with insights into 'student perceptions of learning' and alerting administrators to potential issues in specific courses. Without this means of accountability, faculty may put less effort into their teaching and have less insight into what is working or not working for the students. Given the extensive reviews that need to be undertaken each academic term within universities, it is likely that SET will continue to be used as a metric for accountability in the classroom.</p> <p>While some scholars concur that Student Evaluations of Teaching (SET) can offer administrators valuable insights when included as a component of a comprehensive portfolio (Benton and Ryalls [<reflink idref="bib5" id="ref3">5</reflink>]; Esarey and Valdes [<reflink idref="bib10" id="ref4">10</reflink>]; Hammonds et al. [<reflink idref="bib11" id="ref5">11</reflink>]; Loeher et al. [<reflink idref="bib18" id="ref6">18</reflink>]), the majority of researchers recommend a cautious approach when utilizing SET. This caution is rooted in a large body of literature that investigates the existence and impact of measurement and equity biases in SET (Boring [<reflink idref="bib8" id="ref7">8</reflink>]; Esarey and Valdes [<reflink idref="bib10" id="ref8">10</reflink>]; Kreitzer and Sweet-Cushman [<reflink idref="bib17" id="ref9">17</reflink>]; MacNell, Driscoll, and Hunt [<reflink idref="bib19" id="ref10">19</reflink>]; Spooren, Brockx, and Mortelmans [<reflink idref="bib28" id="ref11">28</reflink>]; Uttl and Smibert [<reflink idref="bib32" id="ref12">32</reflink>]). We draw upon this research and the objectives of our university of study to guide the creation of a framework designed to reduce existing bias and the impact of confounding factors. This paper's framework emphasizes the use of residuals derived from a carefully constructed Ordinary Least Squares (OLS) model. We propose that these post-estimation residuals, which account for key confounding factors in student feedback, when combined with simple mean SET scores, provide more meaningful information than the mere use of simple means or even means juxtaposed with benchmark metrics.</p> <hd id="AN0181947092-3">II. Literature review</hd> <p>Numerous studies have debated the ability of SET scores to accurately measure teaching effectiveness, (Arroyo-Barriguete et al. [<reflink idref="bib2" id="ref13">2</reflink>]; Hornstein [<reflink idref="bib14" id="ref14">14</reflink>]; McPherson, Todd Jewell, and Kim [<reflink idref="bib21" id="ref15">21</reflink>]). Additionally, several have explored strategies to improve SET, both from a methodological perspective and to reduce measurement and equity bias (Kreitzer and Sweet-Cushman [<reflink idref="bib17" id="ref16">17</reflink>]). Over time, the role of SET has evolved from a tool primarily used for formative evaluation, benefiting the instructor, to a dual-purpose instrument serving both formative evaluation and summative faculty assessment (Hornstein [<reflink idref="bib14" id="ref17">14</reflink>]). As such, it is crucial to ensure that the data derived from SET is as free from bias and measurement error as possible (Hammonds et al. [<reflink idref="bib11" id="ref18">11</reflink>]). Hammonds et al. ([<reflink idref="bib11" id="ref19">11</reflink>]) offer a comprehensive overview of the concerns associated with SET and potential avenues for enhancement.</p> <p>In terms of enhancing data quality, studies have proposed that evaluations conducted before the final exam could yield higher quality data and a more accurate reflection of student attitudes (Vehovar and Štrlekar [<reflink idref="bib33" id="ref20">33</reflink>]). Furthermore, certain factors such as incentive structures, paper-based evaluations as opposed to online ones, emphasizing the importance of student feedback, and reminders to complete SET could boost response rates (Hammonds et al. [<reflink idref="bib11" id="ref21">11</reflink>]; Zipser and Mincieli [<reflink idref="bib36" id="ref22">36</reflink>]). This could result in larger and more representative datasets. Keeley, English, Irons, and Henslee ([<reflink idref="bib16" id="ref23">16</reflink>]) also suggest that educating students about the purpose and function of SET could further improve data quality. For instance, an experiment conducted by Peterson, Biederman, Andersen, Ditonto, and Roe ([<reflink idref="bib24" id="ref24">24</reflink>]) revealed a significant difference in the ratings provided by students who received evaluations with a description of bias, compared to those who received the standard evaluation form without any information on bias.</p> <p>An additional strategy that can complement the aforementioned recommendations involves adjusting evaluation scores to account for SET covariates that might confound 'students' perceptions of learning'. The literature has extensively documented that covariates such as academic discipline or the mandatory nature of a course can influence SET (Alquaraan, Alazzam, and Alkhateeb [<reflink idref="bib1" id="ref25">1</reflink>]; Benton and Cashin [<reflink idref="bib3" id="ref26">3</reflink>]; Rosen [<reflink idref="bib27" id="ref27">27</reflink>]; Spooren, Brockx, and Mortelmans [<reflink idref="bib28" id="ref28">28</reflink>]; Uttl and Smibert [<reflink idref="bib32" id="ref29">32</reflink>]; Zipser, Mincieli, and Kurochkin [<reflink idref="bib37" id="ref30">37</reflink>]). The covariates from the literature include student, course, and instructor characteristics (Spooren, Brockx, and Mortelmans [<reflink idref="bib28" id="ref31">28</reflink>]), which we discuss in detail in Section VI. Model Development.</p> <p>If institutions only consider average ratings without adjusting for these covariates, the insights they gain from student feedback on learning might be skewed. Similarly, if administrators or researchers identify equity bias in their university's SET scores, controlling for these factors can help mitigate such effects. Several articles propose potential methods to adjust scores, including controlling for confounding factors and generating an adjusted score (Benton et al. [<reflink idref="bib4" id="ref32">4</reflink>]) or using t-scores, ranges, or standard deviations instead of raw means (Benton and Young [<reflink idref="bib6" id="ref33">6</reflink>]).</p> <p>As previously discussed, there is ongoing debate about whether SET truly measures teaching effectiveness and whether students are suitable evaluators of faculty (Arroyo-Barriguete et al. [<reflink idref="bib2" id="ref34">2</reflink>]; Hammonds et al. [<reflink idref="bib11" id="ref35">11</reflink>]; Hornstein [<reflink idref="bib14" id="ref36">14</reflink>]; McPherson, Todd Jewell, and Kim [<reflink idref="bib21" id="ref37">21</reflink>]). As we noted in the introduction, we believe that SET should be utilized as a metric for 'student perceptions of learning'. Some studies have found correlations between SET and measures of student learning or their perceptions of learning. For instance, in their literature review, Spooren, Brockx, and Mortelmans ([<reflink idref="bib28" id="ref38">28</reflink>]) found varying evidence of convergent validity, or whether SET correlate with measures of student learning. Moreover, Remedios and Lieberman ([<reflink idref="bib26" id="ref39">26</reflink>]) found that student perceptions of teaching quality and engagement with course materials were more influential on SET scores than grades. Our current research, which analyzes over 200,000 student open-ended comments, reveals that the primary issues students mention relate to learning, and that quantitative SET scores increase when the prevalence of objectively positive comments about learning increases and decrease when the prevalence of a objectively negatively comments about learning increase (Zipser et al. [<reflink idref="bib35" id="ref40">35</reflink>]). Therefore, while SET may be imperfect measures of teaching effectiveness, they can serve as a measure of student perceptions of whether instructors' teaching methods are successful. This can then become an impetus for further review, by say faculty peers, of an instructors' teaching if problems are cited.</p> <hd id="AN0181947092-4">III. Conceptual framework</hd> <p>At our university of interest, SET have a longstanding history. Additionally, they have an overall response rate of greater than 80% (internal data). One factor contributing to this high response rate is that course evaluations are made public for students to utilize during the course selection process, thereby providing a strong incentive from them to complete the evaluations thoroughly and thoughtfully (Yu, Mincieli, and Zipser [<reflink idref="bib34" id="ref41">34</reflink>]). SET have been consistently used as one input for teaching awards and annual assessments. Faculty members receive benchmark data in the form of departmental and divisional averages along with their quantitative SET scores and qualitative student comments. Furthermore, any significant changes to the evaluation instrument or its administration require faculty approval, beginning with the Standing Committee on Undergraduate Educational Policy. Therefore, SET are a well-established tool used in the evaluation process.</p> <p>To mitigate potential biases and address confounding factors, we have devised a framework that utilizes adjusted SET scores, similar to the approach of Benton et al. ([<reflink idref="bib4" id="ref42">4</reflink>]), for administrative use. Our framework computes the residuals of an OLS model, which predicts instructor scores while controlling for endogenous factors that the institution may not wish the faculty to change (e.g. introducing enrollment caps to limit course size) or exogenous factors that the faculty member cannot change (e.g. academic discipline). This model enables administrators to look beyond average SET scores and identify faculty who may be successfully teaching courses that are inherently challenging for various reasons. This framework, which uses residuals in addition to simple means as one part of a teaching portfolio for reviewing faculty, offers institutions a method to use SET data to gain a more nuanced understanding of each faculty member's individual ratings. Our framework aligns with the method used by Arroyo-Barriguete et al. ([<reflink idref="bib2" id="ref43">2</reflink>]). They find that even after adjusting for confounding factors, potential biases remain (Arroyo-Barriguete et al. [<reflink idref="bib2" id="ref44">2</reflink>]). However, we posit that, despite its imperfections, this framework provides universities that continue to use SET in the evaluation process with a more refined, individualized supplement to simple mean analysis. Our framework also bears similarities to the method employed by the IDEA Center, an online evaluation system developed by Kansas State University and available to other universities (Campus Labs/Anthology).</p> <p>The difference in residual ratings and average SET ratings can be striking, with some faculty receiving high ratings in the residual framework, which takes control variables into account, compared to a framework based on of average SET ratings without any controls. In a residual framework, individual faculty are compared to an expected value derived from the characteristics of their courses. This method allows us to avoid direct comparison between faculty members, which can be advantageous for various reasons (Kreitzer and Sweet-Cushman [<reflink idref="bib17" id="ref45">17</reflink>]). Our approach involves of two steps:</p> <p></p> <ulist> <item> We estimate an OLS regression model, using multiple years of data, to predict expected instructor SET overall score, while controlling for factors that may confound SET ratings at our institution. The dependent variable, the average instructor SET overall score for a course, can take any real number between 1 and 5. We, therefore, treat this variable as continuous.</item> <p></p> <item> We use post-estimation residuals from the model to calculate the residual average for each faculty member over a specified period (we use five years, or 10 semesters).</item> </ulist> <p>These steps are elaborated in more depth in Section V. We recommend other institutions, which use SET, consider a similar technique, instead of solely relying on uncontrolled simple means.</p> <hd id="AN0181947092-5">IV. Data</hd> <p>We utilize SET data from undergraduate lecture and seminar courses[<reflink idref="bib1" id="ref46">1</reflink>] taught by tenured faculty in the Faculty of Arts and Sciences (FAS) at a large, private university located in the Northeastern United States. The sample includes 557 tenured faculty who were active as of 1 January 2022.[<reflink idref="bib2" id="ref47">2</reflink>],[<reflink idref="bib3" id="ref48">3</reflink>] The SET data ranges from Fall 2010 to Fall 2021.[<reflink idref="bib4" id="ref49">4</reflink>] For each course a faculty member has taught in the FAS, we create a 'faculty-course', which is a unique combination of semester, course catalog number and instructor ID (Zipser, Mincieli, and Kurochkin [<reflink idref="bib37" id="ref50">37</reflink>]). This method enables courses that are co-taught to be included in the dataset as discrete listings for each instructor. Our final dataset comprises 6,598 faculty-courses, corresponding to 218,338 evaluations completed by 29,263 individual students. Our study received approval from the Committee on the Use of Human Subjects (approval no. IRB15-3886). The requirement for informed consent was waived for the study.</p> <p>In our analysis, we focus on the 'instructor overall' score as our outcome of interest, leading us to exclude courses where instructor overall scores were not pertinent to our study.[<reflink idref="bib5" id="ref51">5</reflink>] This includes courses where enrollments were aggregated by the Registrar's Office, such as lectures where a course head is listed in the course catalog, but the actual instruction takes place in sections led by different instructors. We also exclude Reading/Research/Thesis writing courses where a student works one-on-one with a faculty member, as opposed to in a traditional classroom setting.</p> <hd id="AN0181947092-6">V. Method</hd> <p>To build an OLS regression model, we use as the dependent variable the faculty-course average SET instructor overall score. This is calculated by averaging each student's 'instructor overall' scores for a given faculty-course. We acknowledge the potential issues of treating an ordinal scale as interval data (Jamieson [<reflink idref="bib15" id="ref52">15</reflink>]), yet this practice is prevalent in SET literature (Park and Dooris [<reflink idref="bib23" id="ref53">23</reflink>]). Moreover, administrators often interpret SET as continuous when calculating means.</p> <p>When building a model, we advise institutions to utilize as many years of data as possible to enhance its predictive power. However, we caution institutions to examine trends in older data, as well as changes in evaluation administration, to ensure they are not including data that is no longer representative of the institution and its faculty.</p> <p>We estimate an OLS regression model to predict the faculty-course average for the SET 'instructor overall' score, controlling for the division of the course (i.e. Arts and Humanities, Social Science, Science, Engineering and Applied Science, General Education, and Freshman Seminars), the natural logarithm of enrollment for the course, the percent of students taking the course as an elective, and the semester in which the course was taught. The OLS regression model specified below is tailored to the Faculty of Arts and Sciences at our institution, but it is also theoretically grounded in existing research, which we will discuss further. Given the lack of consensus among researchers on the factors that might bias SET, and the likelihood that these factors differ from one institution to another (or even within the same institution) (Arroyo-Barriguete et al. [<reflink idref="bib2" id="ref54">2</reflink>]), it is important that each institution develop its own model. This model should consider the unique endogenous and exogenous factors that could confound SET at their respective institutions. We also recommend making the model as parsimonious as possible to increase its robustness as a tool that can be consistently applied over time. We present our model below:</p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mtable columnalign="right left" rowspacing=".5em" columnspacing="thickmathspace" displaystyle="true"><mtr><mtd /><mtd><mi>Faculty</mi><mspace width="thickmathspace" /><mi>course</mi><mspace width="thickmathspace" /><mi>average</mi></mtd></mtr><mtr><mtd /><mtd><mo>=</mo><mspace width="thickmathspace" /><msub><mtext fontfamily="times">β</mtext><mn>0</mn></msub><mo>+</mo><mspace width="thickmathspace" /><msub><mtext fontfamily="times">β</mtext><mn>1</mn></msub><mo>(</mo><mrow><mi>division</mi></mrow><mo>)</mo><mo>+</mo><mspace width="thickmathspace" /><msub><mtext fontfamily="times">β</mtext><mn>2</mn></msub><mi>ln</mi><mo>⁡</mo><mo>(</mo><mrow><mi>enrollment</mi></mrow><mo>)</mo><mo>+</mo><mspace width="thickmathspace" /></mtd></mtr><mtr><mtd /><mtd><msub><mtext fontfamily="times">β</mtext><mn>3</mn></msub><mo>(</mo><mrow><mspace width="thinmathspace" /><mi>fraction</mi><mspace width="thickmathspace" /><mi>taking</mi><mspace width="thickmathspace" /><mi>course</mi><mspace width="thickmathspace" /><mi>as</mi><mspace width="thickmathspace" /><mi>elective</mi></mrow><mo>)</mo><mo>+</mo><mspace width="thickmathspace" /><msub><mtext fontfamily="times">β</mtext><mn>4</mn></msub><mo>(</mo><mrow><mi>year</mi><mspace width="thickmathspace" /><mi mathvariant="normal">&</mi><mspace width="thickmathspace" /><mi>term</mi></mrow><mo>)</mo></mtd></mtr></mtable></math> </ephtml> </p> <p>The control variable <emph>year & term</emph> represents a set of twenty-two dummy variables each representing a term and year combination beginning with Fall 2010 and ending with Fall 2021.</p> <p>After running our model, we compute the post-estimation residuals of each faculty member for each of their faculty-course offerings.[<reflink idref="bib6" id="ref55">6</reflink>] The residual is the difference between an instructor's overall average score for a course, without any controls, and their expected score for the same course, with controls. We then average each faculty member's residuals from Fall 2016 to Fall 2021, a period encompassing 10 semesters' worth of scores instead of 11 because SET were not administered in Spring 2020 due to the Covid-19 pandemic. By using an average of 10 semesters' worth of residuals, we can more robustly identify students' perceptions of teaching for each faculty member. For example, we mitigate the risk of one or two poorly rated courses obscuring the average for an instructor who typically receives high evaluation scores. This data allows us to identify faculty members with higher-than-expected scores (residual averages greater than zero), lower than expected scores (residual averages less than zero), and scores that meet expectations (residual averages close to zero). The consideration of residuals controls for factors that could confound SET averages.</p> <hd id="AN0181947092-7">VI. Model development</hd> <p>As we discuss above, prior studies have indicated that SET scores can be influenced by a range of factors, including structural factors such as class size or discipline and individual factors such as gender or race/ethnicity. Through numerous conversations with university administrators who work with SET, we know that they mentally 'control' for many of these factors when evaluating faculty. The main factors that they control for are size of the course, discipline of the course, and whether a course is required. Combining what we have learned about administrator behavior and previous analytic research (Zipser, Mincieli, and Kurochkin [<reflink idref="bib37" id="ref56">37</reflink>]), including extensive literature reviews, we have determined the following factors are important potential confounds of SET at our institution: course size, course discipline, year/term, and course type (i.e. to what degree a course is required). We discuss each factor below. We then discuss other common factors cited in the literature for which we do not control but other universities may wish to, including gender, race/ethnicity, age, and student characteristics. In Appendix Table 1 we show the results when several of these other factors are included in the model.</p> <hd id="AN0181947092-8">1. Course size</hd> <p>In their review of potential confounds, Marsh and Roche ([<reflink idref="bib20" id="ref57">20</reflink>]) find that, in general, smaller courses receive higher ratings than larger courses, and they also suggest the relationship may be curvilinear. Benton and Cashin ([<reflink idref="bib3" id="ref58">3</reflink>]) find that, while some studies report a statistically significant correlation between SET scores and course size, it is often small. If small courses receive higher ratings because the students are learning more, as some studies have suggested, then course size would not be considered a confounding factor (Benton and Cashin [<reflink idref="bib3" id="ref59">3</reflink>]). In addition, courses may be large because they are effective and many students want to take them, which would lead to higher ratings. However, because we cannot test the relationship between course size and actual student learning, we control for course size as a potential confound in our analysis. We hypothesize that enrollment has a logarithmic relationship with SET scores, assuming that, for example, the score difference between enrollments of 10 and 20 is larger than the score difference between enrollments of 100 and 110. Graph 1, below, shows the average score at each level of enrollment, as well as the predicted logarithmic regression line Figure 1.</p> <p>Graph: Figure 1. Average instructor overall score by course enrollment.</p> <hd id="AN0181947092-9">2. Course discipline</hd> <p>Several articles have found that SET scores tend to be higher in courses in Arts and Humanities disciplines, as compared with Science or Engineering disciplines. In his study of almost 8,000,000 evaluations from RateMyProfessors.com, Rosen ([<reflink idref="bib27" id="ref60">27</reflink>]) finds that the top 10 disciplines rated for 'overall quality' are in Arts and Humanities or Social Science fields, while a majority of the bottom-10 rated fields are in Science or Engineering. Similarly, Uttl and Smibert ([<reflink idref="bib32" id="ref61">32</reflink>]) find that the mean scores for English courses at New York University are 0.61 of a point higher than the mean scores for math courses (this difference is greater than one standard deviation), and that almost three-quarters of English courses pass the 'overall mean' of 4.13, compared with only 21 percent of math courses. In their synthesis of SET literature, Benton and Cashin ([<reflink idref="bib3" id="ref62">3</reflink>]) posit that if courses that are more quantitative in nature are rated lower because students are less skilled in quantitative skills, then controlling for discipline is important when looking at SET scores. At our institution, we find that courses in the Arts and Humanities consistently receive higher average instructor overall scores than courses in Social Science, Science, or Engineering. Table 1 below shows the mean scores by discipline for our sample of tenured faculty. Note that we do not control for the discipline of the instructor, but rather the discipline of the course.</p> <p>Table 1. Average instructor overall score by course discipline, 2010–2022.</p> <p> <ephtml> <table><thead valign="bottom"><tr><td>Division</td><td>Mean</td><td>Std. Dev.</td></tr></thead><tbody><tr><td>Arts and Humanities</td><td char=".">4.59</td><td char=".">0.43</td></tr><tr><td>Engineering and Applied Science</td><td char=".">4.08</td><td char=".">0.62</td></tr><tr><td>Freshman Seminars</td><td char=".">4.55</td><td char=".">0.45</td></tr><tr><td>General Education</td><td char=".">4.24</td><td char=".">0.43</td></tr><tr><td>Science</td><td char=".">4.27</td><td char=".">0.54</td></tr><tr><td>Social Science</td><td char=".">4.37</td><td char=".">0.51</td></tr></tbody></table> </ephtml> </p> <hd id="AN0181947092-10">3. Year/Term</hd> <p>Changes to the administration of SET and the SET instrument at our institution between 2004 and 2013 decreased within-course instructor scores by, on average, 0.29 of a point (Zipser and Mincieli [<reflink idref="bib36" id="ref63">36</reflink>]). However, instructor overall scores across the board have gradually increased at our institution, even after controlling for instructor gender, age, rank, discipline, and course type (i.e. required versus elective) (Zipser, Mincieli, and Kurochkin [<reflink idref="bib37" id="ref64">37</reflink>]). Thus, based on evidence from our own institution, we find that SET are influenced by administrative and structural changes that occur, and are also increasing over the years studied. As such, we find it is important to control for year and term in our analysis.</p> <hd id="AN0181947092-11">4. Course type</hd> <p>Several studies find that courses where a larger proportion of students are taking the course as an elective are rated higher than required courses (Marsh and Roche [<reflink idref="bib20" id="ref65">20</reflink>]; Radchenko [<reflink idref="bib25" id="ref66">25</reflink>]; Ting [<reflink idref="bib30" id="ref67">30</reflink>]). Research from our own institution supports this finding (Zipser, Mincieli, and Kurochkin [<reflink idref="bib37" id="ref68">37</reflink>]), leading us to believe that courses where a large proportion of students are taking the course as a requirement are rated differently than popular elective courses. Since administrators do not want to penalize faculty teaching required courses, we include a control for the proportion of students who take a course as an elective in our model.</p> <p>We now discuss several of the notable potential covariates that are often cited in the literature, and we do not include in our model. As we explain below, there are multiple reasons that we do not include these variables in the paper. In Appendix Table 1, we show the results when several of them are included in our primary model as control variables. Note that we only discuss covariates that are often discussed as factors in the literature. See Spooren, Brockx, and Mortelmans ([<reflink idref="bib28" id="ref69">28</reflink>]) for a comprehensive list of research showing the relationships between SET scores and student, instructor, and course characteristics.</p> <hd id="AN0181947092-12">5. Gender</hd> <p>Two recent, large-scale syntheses of the literature have concluded that SET are disadvantageous to female academics (as well as other marginalized groups) (Heffernan [<reflink idref="bib13" id="ref70">13</reflink>]; Kreitzer and Sweet-Cushman [<reflink idref="bib17" id="ref71">17</reflink>]). However, in analyses of SET at our home institution, 32 combinations of instructor rank and discipline were tested and only two instances where scores differed significantly by gender were found; neither of which were observed among tenured faculty (Zipser, Mincieli, and Kurochkin [<reflink idref="bib37" id="ref72">37</reflink>]). This is not unheard of, as other researchers have also found no evidence of gender differences (Binderkrantz, Bisgaard, and Lassesen [<reflink idref="bib7" id="ref73">7</reflink>]), or else evidence of gender differences only among more junior faculty (Mengel, Sauermann, and Zölitz [<reflink idref="bib22" id="ref74">22</reflink>]). Because the technique described in this paper should be tailored to the individual needs of the institution, we do not control for instructor gender in our model, but other institutions may find it necessary to do so. In fact, because female instructors at our university receive similar or higher average scores than male instructors including a control variable for gender in the residual analysis could lower their predicted scores and be detrimental to them. (See Appendix Table 1 for models including gender.)</p> <hd id="AN0181947092-13">6. Race/ethnicity</hd> <p>Heffernan ([<reflink idref="bib13" id="ref75">13</reflink>]) also finds evidence of bias in SET for other marginalized groups, such as faculty from historically underrepresented races/ethnicities or faculty whose first language is not English. We do not control for this possibility because, unfortunately, we do not have enough tenured faculty from these historically underrepresented groups in our dataset to perform a robust statistical analysis. However, we recognize this is something we should continue to investigate and then update our model as appropriate.</p> <hd id="AN0181947092-14">7. Age</hd> <p>Similar to other course or demographic predictors, the inclusion of age (as a proxy for career stage) in our model is dependent upon whether we think there is bias against tenured faculty in certain career stages, or if we believe there could be differences in teaching effectiveness based on a faculty member's career stage that are not biases, per se, but rather effects that we want to capture. (For example, younger faculty with less experience may not be as effective as more experienced faculty, on average, or later career-stage faculty may not be as excited about teaching and may put less effort into their teaching, on average). If we believe there is statistical discrimination, we want to control for age. If there may be average career-stage effects that are not due to biases but to performance, then we would not control for age. We chose not to control for age. In their papers, McPherson, Todd Jewell, and Kim ([<reflink idref="bib21" id="ref76">21</reflink>]) and Stonebraker and Stone ([<reflink idref="bib29" id="ref77">29</reflink>]) find that age is negatively correlated with SET scores and thus this is a factor institutions should consider. (See Appendix Table 1 for models including age. Age is not a statistically significant predictor of SET in our model.)</p> <hd id="AN0181947092-15">8. Student characteristics</hd> <p>Finally, some papers include student characteristics in their models predicting overall instructor scores, as research has been found that students' demographic characteristics and backgrounds correlate with SET scores (Heffernan [<reflink idref="bib13" id="ref78">13</reflink>]). However, the focus of our research is to control for those factors administrators benchmark when making salary or other personnel decisions. At our institution, administrators do not know the demographics of the students in the course when reviewing SET. Thus, controlling for student characteristics in this context would not simulate the types of benchmarks administrators are using when reviewing SET. In addition, controlling for student demographics in this model would potentially have a deleterious effect, de-incentivizing faculty to strive to make their class inclusive for all students. At our institution, tenured faculty are expected to be able to teach successfully to all undergraduates and have a great deal of autonomy over the courses they teach and how they present material. Institutions should make their own evidence-backed decisions regarding whether to include student characteristics in their models. However, we are cognizant that many scholars are concerned that the fraction of female students in a class can affect the ratings of female faculty. In Appendix Table 1 we see that neither the gender of the instructor or the fraction of female students in the course are significant predictors of SET scores.</p> <hd id="AN0181947092-16">VII. Results</hd> <p>For our sample of tenured faculty in the Faculty of Arts and Sciences[<reflink idref="bib7" id="ref79">7</reflink>], the unadjusted mean instructor overall score is 4.38 on a five-point likert scale, with 5.0 representing the best score a faculty can receive. The standard deviation of the mean is 0.52.</p> <p>Our OLS regression model (hereafter referred to as Model 1) predicting the faculty-course average for each faculty member and controlling for the potential factors discussed above is shown below in Table 2.</p> <p>Table 2. OLS regression predicting instructor overall scores (Model 1).</p> <p> <ephtml> <table><thead valign="bottom"><tr><td /><td>Coef.</td><td>SE</td><td><italic>P</italic>-value</td></tr></thead><tbody><tr><td>Division (Arts & Humanities ref.)</td></tr><tr><td> Engineering and Applied Science</td><td char=".">−0.362</td><td char=".">0.024</td><td char=".">0.000</td></tr><tr><td> Freshman Seminars</td><td char=".">−0.201</td><td char=".">0.025</td><td char=".">0.000</td></tr><tr><td> General Education</td><td char=".">−0.114</td><td char=".">0.021</td><td char=".">0.000</td></tr><tr><td> Science</td><td char=".">−0.194</td><td char=".">0.017</td><td char=".">0.000</td></tr><tr><td> Social Science</td><td char=".">−0.118</td><td char=".">0.017</td><td char=".">0.000</td></tr><tr><td>Ln(Enrollment)</td><td char=".">−0.126</td><td char=".">0.006</td><td char=".">0.000</td></tr><tr><td>Fraction of Students Taking Course as Elective</td><td char=".">0.234</td><td char=".">0.025</td><td char=".">0.000</td></tr><tr><td>Year/Term (Fall 2010 ref.)</td></tr><tr><td> Spring 2011</td><td char=".">−0.050</td><td char=".">0.040</td><td char=".">0.209</td></tr><tr><td> Fall 2011</td><td char=".">0.007</td><td char=".">0.041</td><td char=".">0.865</td></tr><tr><td> Spring 2012</td><td char=".">0.009</td><td char=".">0.040</td><td char=".">0.831</td></tr><tr><td> Fall 2012</td><td char=".">0.037</td><td char=".">0.040</td><td char=".">0.355</td></tr><tr><td> Spring 2013</td><td char=".">−0.037</td><td char=".">0.040</td><td char=".">0.347</td></tr><tr><td> Fall 2013</td><td char=".">0.032</td><td char=".">0.040</td><td char=".">0.424</td></tr><tr><td> Spring 2014</td><td char=".">0.040</td><td char=".">0.039</td><td char=".">0.301</td></tr><tr><td> Fall 2014</td><td char=".">0.056</td><td char=".">0.041</td><td char=".">0.165</td></tr><tr><td> Spring 2015</td><td char=".">−0.001</td><td char=".">0.041</td><td char=".">0.975</td></tr><tr><td> Fall 2015</td><td char=".">0.071</td><td char=".">0.040</td><td char=".">0.073</td></tr><tr><td> Spring 2016</td><td char=".">0.071</td><td char=".">0.040</td><td char=".">0.072</td></tr><tr><td> Fall 2016</td><td char=".">0.117</td><td char=".">0.039</td><td char=".">0.003</td></tr><tr><td> Spring 2017</td><td char=".">0.091</td><td char=".">0.040</td><td char=".">0.021</td></tr><tr><td> Fall 2017</td><td char=".">0.087</td><td char=".">0.039</td><td char=".">0.027</td></tr><tr><td> Spring 2018</td><td char=".">0.135</td><td char=".">0.039</td><td char=".">0.001</td></tr><tr><td> Fall 2018</td><td char=".">0.077</td><td char=".">0.039</td><td char=".">0.049</td></tr><tr><td> Spring 2019</td><td char=".">0.110</td><td char=".">0.039</td><td char=".">0.004</td></tr><tr><td> Fall 2019</td><td char=".">0.227</td><td char=".">0.038</td><td char=".">0.000</td></tr><tr><td> Fall 2020</td><td char=".">0.282</td><td char=".">0.038</td><td char=".">0.000</td></tr><tr><td> Spring 2021</td><td char=".">0.296</td><td char=".">0.039</td><td char=".">0.000</td></tr><tr><td> Fall 2021</td><td char=".">0.212</td><td char=".">0.039</td><td char=".">0.000</td></tr><tr><td>Constant</td><td char=".">4.694</td><td char=".">0.037</td><td char=".">0.000</td></tr><tr><td>N</td><td>6598</td><td /><td /></tr><tr><td>R-Square</td><td char=".">0.2194</td><td /><td /></tr></tbody></table> </ephtml> </p> <p>As evident from Table 2, courses in the Arts and Humanities receive statistically significantly higher instructor overall scores, on average, than those in all other divisions, with the biggest difference in Engineering and Applied Science, where the average instructor scores are 0.37 of a point lower than in Arts and Humanities. The natural log of enrollment is also statistically significant, as is the fraction of students taking the course as an elective (e.g. students give lower scores, on average, to instructors teaching required courses). Using this model, we calculate the post-estimation residuals of each faculty member for each of their faculty-course offerings and average all their residuals from the past 10 semesters, creating for each faculty member a residual average score. We suggest universities assess this residual average scores together with simple means. In the section below, we provide examples showing the power of the residual framework.</p> <hd id="AN0181947092-17">Residual analysis</hd> <p>Let's consider an example to show how the residual framework can help elucidate the effectiveness of an instructor's teaching. Consider, for example, Professor A, who taught a 200-student, required course in Engineering and Applied Science and received an average score of 4.3. An administrator might ask, how should I think about this score? Using Model 1 to predict the professor's expected score, and then subtracting this expected score from the actual scores, we see that Professor A has a residual of +0.4 (an actual score of 4.3 minus an expected score of 3.9). This example exemplifies how teaching large, required courses in Engineering could be quite costly to faculty if their courses are not benchmarked by disciplines, course size, and type. The residual framework helps to ameliorate such a tax and appropriately recognize the efforts of the faculty member.</p> <p>To understand more broadly what information faculty and administrators can gain from the residual framework, which controls for covariates, we compare residuals, uncontrolled average SET ratings, and divisional average SET ratings for ten randomly chosen faculty from the Sciences. We summarize this data in Table 3.</p> <p>Table 3. Comparison of unadjusted SET averages to residual averages for randomly selected science faculty.</p> <p> <ephtml> <table><thead valign="bottom"><tr><td>Faculty Name</td><td>Division</td><td>SET Average</td><td>Residual Average</td><td>Divisional Average</td></tr></thead><tbody><tr><td>Fac 2</td><td>Science</td><td char=".">4.75</td><td char=".">0.17</td><td char=".">4.38</td></tr><tr><td>Fac 7</td><td char=".">Science</td><td char=".">4.46</td><td char=".">0.25</td><td char=".">4.38</td></tr><tr><td>Fac 9</td><td char=".">Science</td><td char=".">4.36</td><td char=".">0.14</td><td char=".">4.38</td></tr><tr><td>Fac 3</td><td char=".">Science</td><td char=".">4.30</td><td char=".">−0.27</td><td char=".">4.38</td></tr><tr><td>Fac 5</td><td char=".">Science</td><td char=".">4.29</td><td char=".">−0.03</td><td char=".">4.38</td></tr><tr><td>Fac 4</td><td char=".">Science</td><td char=".">4.16</td><td char=".">−0.22</td><td char=".">4.38</td></tr><tr><td>Fac 6</td><td char=".">Science</td><td char=".">3.98</td><td char=".">−0.47</td><td char=".">4.38</td></tr><tr><td>Fac 1</td><td char=".">Science</td><td char=".">3.94</td><td char=".">−0.21</td><td char=".">4.38</td></tr><tr><td>Fac 10</td><td char=".">Science</td><td char=".">3.89</td><td char=".">−0.35</td><td char=".">4.38</td></tr><tr><td>Fac 8</td><td char=".">Science</td><td char=".">3.66</td><td char=".">−0.79</td><td char=".">4.38</td></tr></tbody></table> </ephtml> </p> <p>This table illustrates that uncontrolled average SET ratings and divisional benchmarks lack important information that can help faculty and administrators understand how SET scores compare with similarly situated classes. As evident in the table, faculty with the highest SET averages are not necessarily those with the highest residual averages. Additionally, some faculty with SET averages close to the divisional mean may have substantially different residual averages. These examples also show that presenting benchmarks such as the average rating in the discipline is insufficient to handle multiple confounds.</p> <hd id="AN0181947092-18">VIII. Discussion and limitations</hd> <p>The residual framework discussed in this paper, which involves a multiyear analysis of SET scores, adjustment for potential confounding and biasing factors, and a comparison of instructors' average scores to their predicted ones, can address some of the concerns raised about the use of SET in personnel decision-making (Kreitzer and Sweet-Cushman [<reflink idref="bib17" id="ref80">17</reflink>]). Although this technique is not novel, there is little evidence of its application in evaluating SET for personnel decision-making. In this paper, we propose a straightforward framework for institutions to adopt. We argue that this framework minimizes biases as it is grounded in the literature on measurement and equity biases, and it formalizes what was previously an informal process of administrators mentally adjusting for confounds. We express concern that such an informal process can be inconsistent, potentially leading to unequal faculty evaluations and introducing further biases into the system. We are transparent, through publishing our methods, about the factors included in our model and identify which confounds are relevant at our institution. We believe that a key aspect of any such framework is its transparency, which, combined with the other factors, contributes to a more equitable and consistent faculty evaluation system.</p> <p>This framework also offers other benefits. For instance, it changes score interpretation from high versus low given the scale (3.5 looks low, 4.6 appears high) to high versus low given each faculty member's individual expected scores. Our use of multiple years of evaluations with a high response rate ensures that outliers do not significantly impact the analysis. Institutions should carefully consider what variables to control for. McPherson, Todd Jewell, and Kim ([<reflink idref="bib21" id="ref81">21</reflink>]) create a similar framework of producing predicted SET scores based off a random-effects model and conclude that institutions may want to use adjusted rankings when using SET for personnel decision-making.</p> <p>While many faculty dislike the use the SET in the evaluation process, many large universities will likely continue to use quantitative SET. If SET usage is discontinued, students might turn to other sources of information such as Rate My Professor instead of more thoughtfully designed institutional instruments. In universities where large numbers of faculty must be evaluated, SET data may provide accountability. While they should not be the only factor used in assessing teaching, as one piece of data in a larger teaching portfolio, they can provide information on 'student perceptions of learning'. Providing the context for the course, as we do in the paper, allows for a more nuanced understanding of how well an instructor is teaching, in the eyes of students, given the control variables we identify. For example, how should administrators view the scores of a tenure-track faculty member who is teaching a large, required course in the sciences? It doesn't make sense to compare this faculty member's SET score to that of a faculty member teaching a small seminar with a large fraction of the students taking the seminar as an elective. Using a residual approach, individualized to an institution's particular needs, helps reviewers make the adjustments for context, through control variables, that are now done intuitively, but most likely imprecisely, in peoples' heads.</p> <p>We caution institutions about using this approach if their SET have low response rates. Low response rates can lead to selection bias, which may obscure signals about teaching. SET at our institution have a response rate of over 80 percent, which is much higher than at many other institutions, where the average response rates for online SET are around 50 percent (Zipser and Mincieli [<reflink idref="bib36" id="ref82">36</reflink>]). Therefore, using SET and a residual framework may not be as effective at institutions with low response rates. He and Freeman ([<reflink idref="bib12" id="ref83">12</reflink>]) discuss the reliability of SET scores given response rates and provide some guidelines for their use.</p> <p>Another limitation of this method is that it requires careful explanation, unlike simple benchmark data that are often shared with faculty (e.g. the average SET score for the discipline of the course, the size of the course, and the number of students who responded to the survey). Replacing or augmenting simple benchmark data with residuals, which are an analytic summary of the cumulative effect of the benchmarking data, would require continual explanation and training. At our institution, most of the deans and administrators who use this information are also faculty members (academic administrators), and the SET instrument was developed by a committee of faculty and students. It is crucial that the use of SET for evaluation have faculty oversight. Thus far, academic administrators at our university have noted that having both residuals and simple means is of greater value to them than solely having access to simple SET means.</p> <p>Lastly, the residual framework should be supplemented by a discussion and contextualization of a professor's full portfolio of teaching, including additional information such as peer observations and teaching statements and student comments. We acknowledge and agree with other researchers who posit that there are many areas of teaching in which a student is not an appropriate evaluator.</p> <p>While there continue to be limitations with using SET in the faculty evaluation process, we believe our framework provides universities with a more nuanced and equitable way to approach SET data.</p> <hd id="AN0181947092-19">Disclosure statement</hd> <p>No potential conflict of interest was reported by the author(s).</p> <hd id="AN0181947092-20">Appendix: OLS regression models including additional covariates</hd> <p></p> <p> <ephtml> <table><thead valign="bottom"><tr><td>A.Our Model</td><td>B.Including Age + Age^2</td><td>C. Including Female Instructor dummy</td><td>D. Including Female instructor dummy and % Female students</td><td>E. Including Female Instructor dummy interacted with % of female student<xref ref-type="fn" rid="fn8" /></td><td>F. Including all</td></tr><tr><td>Coef.</td><td>SE</td><td><italic>P</italic>-value</td><td>Coef.</td><td>SE</td><td><italic>P</italic>-value</td><td>Coef.</td><td>SE</td><td><italic>P</italic>-value</td><td>Coef.</td><td>SE</td><td><italic>P</italic>-value</td><td>Coef.</td><td>SE</td><td><italic>P</italic>-value</td><td>Coef.</td><td>SE</td><td><italic>P</italic>-value</td></tr></thead><tbody><tr><td>Age</td><td char=".">−0.006</td><td char=".">0.004</td><td char=".">0.160</td><td char=".">−0.006</td><td char=".">0.004</td><td char=".">0.154</td></tr><tr><td>Age^2</td><td char=".">0.000</td><td char=".">0.000</td><td char=".">0.905</td><td char=".">0.000</td><td char=".">0.000</td><td char=".">0.912</td></tr><tr><td>Female instructor</td><td char=".">0.001</td><td char=".">0.013</td><td char=".">0.933</td><td char=".">0.002</td><td char=".">0.013</td><td char=".">0.898</td><td char=".">0.033</td><td char=".">2.610</td><td char=".">0.022</td><td char=".">−0.015</td><td char=".">0.013</td><td char=".">0.264</td></tr><tr><td>Fraction of female students</td><td char=".">−0.007</td><td char=".">0.027</td><td char=".">0.792</td><td char=".">0.035</td><td char=".">0.030</td><td char=".">0.254</td><td char=".">−0.004</td><td char=".">0.026</td><td char=".">0.885</td></tr><tr><td>Female inst.*female students</td><td char=".">−0.166</td><td char=".">0.060</td><td char=".">0.006</td></tr><tr><td>Division (Arts & Humanities ref.)</td></tr><tr><td> Engineering and Applied Science</td><td char=".">−0.362</td><td char=".">0.024</td><td char=".">0.000</td><td char=".">−0.403</td><td char=".">0.024</td><td char=".">0.000</td><td char=".">−0.361</td><td char=".">0.024</td><td char=".">0.000</td><td char=".">−0.362</td><td char=".">0.024</td><td char=".">0.000</td><td char=".">−0.362</td><td char=".">0.024</td><td char=".">0.000</td><td char=".">−0.406</td><td char=".">0.024</td><td char=".">0.000</td></tr><tr><td> Freshman Seminars</td><td char=".">−0.201</td><td char=".">0.025</td><td char=".">0.000</td><td char=".">−0.199</td><td char=".">0.025</td><td char=".">0.000</td><td char=".">−0.202</td><td char=".">0.025</td><td char=".">0.000</td><td char=".">−0.202</td><td char=".">0.025</td><td char=".">0.000</td><td char=".">−0.203</td><td char=".">0.025</td><td char=".">0.000</td><td char=".">−0.200</td><td char=".">0.025</td><td char=".">0.000</td></tr><tr><td> General Education</td><td char=".">−0.114</td><td char=".">0.021</td><td char=".">0.000</td><td char=".">−0.096</td><td char=".">0.021</td><td char=".">0.000</td><td char=".">−0.113</td><td char=".">0.021</td><td char=".">0.000</td><td char=".">−0.113</td><td char=".">0.021</td><td char=".">0.000</td><td char=".">−0.113</td><td char=".">0.021</td><td char=".">0.000</td><td char=".">−0.096</td><td char=".">0.021</td><td char=".">0.000</td></tr><tr><td> Science</td><td char=".">−0.194</td><td char=".">0.017</td><td char=".">0.000</td><td char=".">−0.214</td><td char=".">0.017</td><td char=".">0.000</td><td char=".">−0.191</td><td char=".">0.017</td><td char=".">0.000</td><td char=".">−0.191</td><td char=".">0.017</td><td char=".">0.000</td><td char=".">−0.191</td><td char=".">0.017</td><td char=".">0.000</td><td char=".">−0.216</td><td char=".">0.017</td><td char=".">0.000</td></tr><tr><td> Social Science</td><td char=".">−0.118</td><td char=".">0.017</td><td char=".">0.000</td><td char=".">−0.126</td><td char=".">0.017</td><td char=".">0.000</td><td char=".">−0.118</td><td char=".">0.017</td><td char=".">0.000</td><td char=".">−0.118</td><td char=".">0.017</td><td char=".">0.000</td><td char=".">−0.115</td><td char=".">0.017</td><td char=".">0.000</td><td char=".">−0.126</td><td char=".">0.017</td><td char=".">0.000</td></tr><tr><td>Ln(Enrollment)</td><td char=".">−0.126</td><td char=".">0.006</td><td char=".">0.000</td><td char=".">−0.129</td><td char=".">0.006</td><td char=".">0.000</td><td char=".">−0.126</td><td char=".">0.006</td><td char=".">0.000</td><td char=".">−0.126</td><td char=".">0.006</td><td char=".">0.000</td><td char=".">−0.127</td><td char=".">0.006</td><td char=".">0.000</td><td char=".">−0.129</td><td char=".">0.006</td><td char=".">0.000</td></tr><tr><td>Fraction of Students Taking Course as Elective</td><td char=".">0.234</td><td char=".">0.025</td><td char=".">0.000</td><td char=".">0.239</td><td char=".">0.025</td><td char=".">0.000</td><td char=".">0.237</td><td char=".">0.025</td><td char=".">0.000</td><td char=".">0.236</td><td char=".">0.025</td><td char=".">0.000</td><td char=".">0.237</td><td char=".">0.025</td><td char=".">0.000</td><td char=".">0.240</td><td char=".">0.025</td><td char=".">0.000</td></tr><tr><td>Year/Term (Fall 2010 ref.)</td></tr><tr><td> Spring 2011</td><td char=".">−0.050</td><td char=".">0.040</td><td char=".">0.209</td><td char=".">−0.050</td><td char=".">0.039</td><td char=".">0.206</td><td char=".">−0.047</td><td char=".">0.040</td><td char=".">0.236</td><td char=".">−0.047</td><td char=".">0.040</td><td char=".">0.236</td><td char=".">−0.048</td><td char=".">0.040</td><td char=".">0.227</td><td char=".">−0.050</td><td char=".">0.039</td><td char=".">0.204</td></tr><tr><td> Fall 2011</td><td char=".">0.007</td><td char=".">0.041</td><td char=".">0.865</td><td char=".">0.014</td><td char=".">0.040</td><td char=".">0.730</td><td char=".">0.005</td><td char=".">0.041</td><td char=".">0.905</td><td char=".">0.005</td><td char=".">0.041</td><td char=".">0.903</td><td char=".">0.004</td><td char=".">0.041</td><td char=".">0.926</td><td char=".">0.014</td><td char=".">0.040</td><td char=".">0.722</td></tr><tr><td> Spring 2012</td><td char=".">0.009</td><td char=".">0.040</td><td char=".">0.831</td><td char=".">0.023</td><td char=".">0.039</td><td char=".">0.558</td><td char=".">0.018</td><td char=".">0.040</td><td char=".">0.660</td><td char=".">0.018</td><td char=".">0.040</td><td char=".">0.655</td><td char=".">0.017</td><td char=".">0.040</td><td char=".">0.675</td><td char=".">0.023</td><td char=".">0.039</td><td char=".">0.556</td></tr><tr><td> Fall 2012</td><td char=".">0.037</td><td char=".">0.040</td><td char=".">0.355</td><td char=".">0.049</td><td char=".">0.040</td><td char=".">0.222</td><td char=".">0.037</td><td char=".">0.040</td><td char=".">0.356</td><td char=".">0.037</td><td char=".">0.040</td><td char=".">0.353</td><td char=".">0.039</td><td char=".">0.040</td><td char=".">0.337</td><td char=".">0.049</td><td char=".">0.040</td><td char=".">0.221</td></tr><tr><td> Spring 2013</td><td char=".">−0.037</td><td char=".">0.040</td><td char=".">0.347</td><td char=".">−0.032</td><td char=".">0.039</td><td char=".">0.416</td><td char=".">−0.036</td><td char=".">0.040</td><td char=".">0.368</td><td char=".">−0.036</td><td char=".">0.040</td><td char=".">0.369</td><td char=".">−0.037</td><td char=".">0.040</td><td char=".">0.356</td><td char=".">−0.031</td><td char=".">0.039</td><td char=".">0.424</td></tr><tr><td> Fall 2013</td><td char=".">0.032</td><td char=".">0.040</td><td char=".">0.424</td><td char=".">0.052</td><td char=".">0.039</td><td char=".">0.181</td><td char=".">0.031</td><td char=".">0.040</td><td char=".">0.426</td><td char=".">0.032</td><td char=".">0.040</td><td char=".">0.424</td><td char=".">0.031</td><td char=".">0.039</td><td char=".">0.427</td><td char=".">0.052</td><td char=".">0.039</td><td char=".">0.182</td></tr><tr><td> Spring 2014</td><td char=".">0.040</td><td char=".">0.039</td><td char=".">0.301</td><td char=".">0.055</td><td char=".">0.038</td><td char=".">0.150</td><td char=".">0.040</td><td char=".">0.039</td><td char=".">0.304</td><td char=".">0.040</td><td char=".">0.039</td><td char=".">0.303</td><td char=".">0.039</td><td char=".">0.039</td><td char=".">0.317</td><td char=".">0.056</td><td char=".">0.038</td><td char=".">0.148</td></tr><tr><td> Fall 2014</td><td char=".">0.056</td><td char=".">0.041</td><td char=".">0.165</td><td char=".">0.076</td><td char=".">0.040</td><td char=".">0.057</td><td char=".">0.056</td><td char=".">0.041</td><td char=".">0.165</td><td char=".">0.056</td><td char=".">0.041</td><td char=".">0.165</td><td char=".">0.054</td><td char=".">0.040</td><td char=".">0.180</td><td char=".">0.076</td><td char=".">0.040</td><td char=".">0.057</td></tr><tr><td> Spring 2015</td><td char=".">−0.001</td><td char=".">0.041</td><td char=".">0.975</td><td char=".">0.016</td><td char=".">0.040</td><td char=".">0.697</td><td char=".">−0.002</td><td char=".">0.040</td><td char=".">0.966</td><td char=".">−0.002</td><td char=".">0.040</td><td char=".">0.965</td><td char=".">−0.003</td><td char=".">0.040</td><td char=".">0.942</td><td char=".">0.016</td><td char=".">0.040</td><td char=".">0.689</td></tr><tr><td> Fall 2015</td><td char=".">0.071</td><td char=".">0.040</td><td char=".">0.073</td><td char=".">0.095</td><td char=".">0.039</td><td char=".">0.015</td><td char=".">0.071</td><td char=".">0.039</td><td char=".">0.073</td><td char=".">0.071</td><td char=".">0.039</td><td char=".">0.072</td><td char=".">0.071</td><td char=".">0.039</td><td char=".">0.073</td><td char=".">0.095</td><td char=".">0.039</td><td char=".">0.015</td></tr><tr><td> Spring 2016</td><td char=".">0.071</td><td char=".">0.040</td><td char=".">0.072</td><td char=".">0.098</td><td char=".">0.039</td><td char=".">0.013</td><td char=".">0.071</td><td char=".">0.040</td><td char=".">0.073</td><td char=".">0.071</td><td char=".">0.040</td><td char=".">0.073</td><td char=".">0.071</td><td char=".">0.040</td><td char=".">0.073</td><td char=".">0.098</td><td char=".">0.039</td><td char=".">0.012</td></tr><tr><td> Fall 2016</td><td char=".">0.117</td><td char=".">0.039</td><td char=".">0.003</td><td char=".">0.151</td><td char=".">0.039</td><td char=".">0.000</td><td char=".">0.120</td><td char=".">0.039</td><td char=".">0.002</td><td char=".">0.119</td><td char=".">0.039</td><td char=".">0.002</td><td char=".">0.118</td><td char=".">0.039</td><td char=".">0.002</td><td char=".">0.150</td><td char=".">0.039</td><td char=".">0.000</td></tr><tr><td> Spring 2017</td><td char=".">0.091</td><td char=".">0.040</td><td char=".">0.021</td><td char=".">0.120</td><td char=".">0.039</td><td char=".">0.002</td><td char=".">0.090</td><td char=".">0.040</td><td char=".">0.023</td><td char=".">0.090</td><td char=".">0.040</td><td char=".">0.023</td><td char=".">0.089</td><td char=".">0.040</td><td char=".">0.025</td><td char=".">0.120</td><td char=".">0.039</td><td char=".">0.002</td></tr><tr><td> Fall 2017</td><td char=".">0.087</td><td char=".">0.039</td><td char=".">0.027</td><td char=".">0.122</td><td char=".">0.039</td><td char=".">0.002</td><td char=".">0.086</td><td char=".">0.039</td><td char=".">0.028</td><td char=".">0.086</td><td char=".">0.039</td><td char=".">0.028</td><td char=".">0.086</td><td char=".">0.039</td><td char=".">0.029</td><td char=".">0.123</td><td char=".">0.039</td><td char=".">0.002</td></tr><tr><td> Spring 2018</td><td char=".">0.135</td><td char=".">0.039</td><td char=".">0.001</td><td char=".">0.176</td><td char=".">0.039</td><td char=".">0.000</td><td char=".">0.135</td><td char=".">0.039</td><td char=".">0.001</td><td char=".">0.135</td><td char=".">0.039</td><td char=".">0.001</td><td char=".">0.133</td><td char=".">0.039</td><td char=".">0.001</td><td char=".">0.176</td><td char=".">0.039</td><td char=".">0.000</td></tr><tr><td> Fall 2018</td><td char=".">0.077</td><td char=".">0.039</td><td char=".">0.049</td><td char=".">0.118</td><td char=".">0.038</td><td char=".">0.002</td><td char=".">0.076</td><td char=".">0.039</td><td char=".">0.049</td><td char=".">0.077</td><td char=".">0.039</td><td char=".">0.049</td><td char=".">0.076</td><td char=".">0.039</td><td char=".">0.049</td><td char=".">0.119</td><td char=".">0.038</td><td char=".">0.002</td></tr><tr><td> Spring 2019</td><td char=".">0.110</td><td char=".">0.039</td><td char=".">0.004</td><td char=".">0.148</td><td char=".">0.038</td><td char=".">0.000</td><td char=".">0.109</td><td char=".">0.038</td><td char=".">0.005</td><td char=".">0.109</td><td char=".">0.038</td><td char=".">0.004</td><td char=".">0.109</td><td char=".">0.038</td><td char=".">0.005</td><td char=".">0.148</td><td char=".">0.038</td><td char=".">0.000</td></tr><tr><td> Fall 2019</td><td char=".">0.227</td><td char=".">0.038</td><td char=".">0.000</td><td char=".">0.273</td><td char=".">0.038</td><td char=".">0.000</td><td char=".">0.227</td><td char=".">0.038</td><td char=".">0.000</td><td char=".">0.227</td><td char=".">0.038</td><td char=".">0.000</td><td char=".">0.226</td><td char=".">0.038</td><td char=".">0.000</td><td char=".">0.273</td><td char=".">0.038</td><td char=".">0.000</td></tr><tr><td> Fall 2020</td><td char=".">0.282</td><td char=".">0.038</td><td char=".">0.000</td><td char=".">0.350</td><td char=".">0.038</td><td char=".">0.000</td><td char=".">0.282</td><td char=".">0.038</td><td char=".">0.000</td><td char=".">0.282</td><td char=".">0.038</td><td char=".">0.000</td><td char=".">0.280</td><td char=".">0.038</td><td char=".">0.000</td><td char=".">0.350</td><td char=".">0.038</td><td char=".">0.000</td></tr><tr><td> Spring 2021</td><td char=".">0.296</td><td char=".">0.039</td><td char=".">0.000</td><td char=".">0.343</td><td char=".">0.039</td><td char=".">0.000</td><td char=".">0.295</td><td char=".">0.039</td><td char=".">0.000</td><td char=".">0.295</td><td char=".">0.039</td><td char=".">0.000</td><td char=".">0.294</td><td char=".">0.039</td><td char=".">0.000</td><td char=".">0.344</td><td char=".">0.039</td><td char=".">0.000</td></tr><tr><td> Fall 2021</td><td char=".">0.212</td><td char=".">0.039</td><td char=".">0.000</td><td char=".">0.268</td><td char=".">0.039</td><td char=".">0.000</td><td char=".">0.211</td><td char=".">0.039</td><td char=".">0.000</td><td char=".">0.211</td><td char=".">0.039</td><td char=".">0.000</td><td char=".">0.211</td><td char=".">0.039</td><td char=".">0.000</td><td char=".">0.269</td><td char=".">0.039</td><td char=".">0.000</td></tr><tr><td>Constant</td><td char=".">4.694</td><td char=".">0.037</td><td char=".">0.000</td><td char=".">5.028</td><td char=".">0.121</td><td char=".">0.000</td><td char=".">4.692</td><td char=".">0.037</td><td char=".">0.000</td><td char=".">4.695</td><td char=".">0.039</td><td char=".">0.000</td><td char=".">4.678</td><td char=".">0.039</td><td char=".">0.000</td><td char=".">5.039</td><td char=".">0.121</td><td char=".">0.000</td></tr><tr><td><italic>N</italic></td><td>6598</td><td>6566</td><td>6587</td><td>6587</td><td>6587</td><td>6566</td></tr><tr><td>R-Square</td><td>0.219</td><td>0.240</td><td>0.220</td><td>0.220</td><td>0.221</td><td>0.240</td></tr></tbody></table> </ephtml> </p> <ref id="AN0181947092-21"> <title> Notes </title> <blist> <bibl id="bib1" idref="ref25" type="bt">1</bibl> <bibtext> We include only courses classified as "Primarily for undergraduates" or "undergraduate/graduate" because our model is used as one piece of information in determining undergraduate teaching awards. However, institutions could also include graduate courses based on the needs of the institution.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref13" type="bt">2</bibl> <bibtext> An additional 21 tenured faculty who were active on January 1, 2022, have no undergraduate SET data between Fall 2010 and Fall 2021. A majority of these 21 faculty have only recently started their appointments.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref26" type="bt">3</bibl> <bibtext> Due to confidentiality concerns, the data used are not publicly available. However, aggregated data and code are available upon reasonable request.</bibtext> </blist> <blist> <bibl id="bib4" idref="ref32" type="bt">4</bibl> <bibtext> For data from Fall 2006 to Fall 2010, we found that female faculty members received statistically significantly lower scores than male faculty members, indicating potential bias against women. From data since Fall 2010, we have found no evidence of systematic gender bias against women in instructor overall effectiveness scores (Zipser, Mincieli, and Kurochkin [37]).</bibtext> </blist> <blist> <bibl id="bib5" idref="ref3" type="bt">5</bibl> <bibtext> The "instructor overall" question asks students to "Evaluate your instructor overall," with a scale of 1=unsatisfactory, 2=fair, 3=good, 4=very good, 5=excellent.</bibtext> </blist> <blist> <bibl id="bib6" idref="ref33" type="bt">6</bibl> <bibtext> This methodology is similar to that used in the educational value-added literature. Please see Topping and Sanders ([31]) for a description of this methodology.</bibtext> </blist> <blist> <bibl id="bib7" idref="ref73" type="bt">7</bibl> <bibtext> Due to confidentiality concerns, the data used are not publicly available. However, aggregated data and code are available upon request.</bibtext> </blist> <blist> <bibl id="bib8" idref="ref7" type="bt">8</bibl> <bibtext> Although neither the female instructor dummy variable nor the variable measuring the proportion of female students in the course is statistically significant in variation D of the model [or column 13], we also run a variation of the model (E) where we interact the gender of the instructor with the fraction of female students in the course. Using this model, we find that the expected score difference between female and male faculty at the mean proportion of female students in the course (47% female) is 0.01 of a point (holding the other variables in the model constant at the sample means). At one standard deviation below the mean (22% female), the difference is 0.05 of a point, and at one standard deviation above the mean (69% female), the difference is -0.02 of a point. In other words, as the fraction of <emph>male</emph> students increases, SET scores for female instructors increase, and as the fraction of <emph>female</emph> students increases, SET scores for female instructors decrease. This finding is the reverse of what is usually found in the literature, in which female instructors receive higher SET scores when more female students are enrolled in a course.</bibtext> </blist> </ref> <ref id="AN0181947092-22"> <title> References </title> <blist> <bibtext> Alquaraan, Mahmoud, Sulaf Alazzam, and Hakam Alkhateeb. 2023. " Using Measurement Invariance to Explore the Source of Variation in Basic Medical Science Students' Evaluation of Teaching Effectiveness." International Journal of Statistics in Medical Research 12 : 185 – 92. https://doi.org/10.6000/1929-6029.2023.12.23.</bibtext> </blist> <blist> <bibtext> Arroyo-Barriguete, J. L., C. Bada, L. Lazcano, J. Márquez, J. M. Ortiz-Lozano, and A. Rua-Vieites. 2023. " Is it Possible to Redress Noninstructional Biases in Student Evaluation of Teaching Surveys? Quantitative Analysis in Accounting and Finance Courses." Studies in Educational Evaluation 77 : 101263. https://doi.org/10.1016/j.stueduc.2023.101263.</bibtext> </blist> <blist> <bibtext> Benton, S. L., and William E. Cashin. 2012. "Student Ratings of Teaching: A Summary of Research and Literature." In Idea paper #50.</bibtext> </blist> <blist> <bibtext> Benton, Stephen L., Dan Li, Ron Brown, Meixi Guo, and Patricia Sullivan. 2015. "Revising the IDEA Student Ratings of Instruction Sysyem 2002-2011 Data." In IDEA Technical Report No. 18.</bibtext> </blist> <blist> <bibtext> Benton, S. L., and K. R. Ryalls. 2016. "Challenging Misconceptions about Student Ratings of Instruction." In IDEA Paper #58.</bibtext> </blist> <blist> <bibtext> Benton, Stephen L., and Suzanne Young. 2018. "Best Practices in the Evaluation of Teaching." In IDEA Paper #69.</bibtext> </blist> <blist> <bibtext> Binderkrantz, Anne Skorkjær, Mette Bisgaard, and Berit Lassesen. 2022. " Contradicting Findings of Gender Bias in Teaching Evaluations: Evidence from two Experiments in Denmark." Assessment & Evaluation in Higher Education. https://doi.org/10.1080/02602938.2022.2048355.</bibtext> </blist> <blist> <bibtext> Boring, Anne. 2017. " Gender Biases in Student Evaluations of Teaching." Journal of Public Economics 145 : 27 – 41. https://doi.org/10.1016/j.jpubeco.2016.11.006.</bibtext> </blist> <blist> <bibl id="bib9" type="bt">9</bibl> <bibtext> Campus Labs/Anthology. "About the Idea System." https://courseevaluationsupport.campuslabs.com/hc/en-us/articles/360038347493-About-IDEA#!.</bibtext> </blist> <blist> <bibtext> Esarey, Justin, and Natalie Valdes. 2020. " Unbiased, Reliable, and Valid Student Evaluations Can Still be Unfair." Assessment & Evaluation in Higher Education 45 (8): 1106 – 20. https://doi.org/10.1080/02602938.2020.1724875.</bibtext> </blist> <blist> <bibtext> Hammonds, Frank, Gina J. Mariano, Gracie Ammons, and Sheridan Chambers. 2017. " Student Evaluations of Teaching: Improving Teaching Quality in Higher Education." Perspectives: Policy and Practice in Higher Education 21 (1): 26 – 33. https://doi.org/10.1080/13603108.2016.1227388.</bibtext> </blist> <blist> <bibtext> He, Jun, and Lee A. Freeman. 2021. " Can we Trust Teaching Evaluations When Response Rates are not High? Implications from a Monte Carol Simulation." Studies in Higher Education 46 (9): 1934 – 48. https://doi.org/10.1080/03075079.2019.1711046.</bibtext> </blist> <blist> <bibtext> Heffernan, Troy. 2022. " Sexism, Racism, Prejudice, and Bias: A Literature Review and Synthesis of Research Surrounding Student Evaluations of Courses and Teaching." Assessment & Evaluation in Higher Education 47 (1): 144 – 54. https://doi.org/10.1080/02602938.2021.1888075.</bibtext> </blist> <blist> <bibtext> Hornstein, Henry A. 2017. " Student Evaluations of Teaching are an Inadequate Assessment Tool for Evaluating Faculty Performance." Cogent Education 4 (1): 1304016. https://doi.org/10.1080/2331186X.2017.1304016.</bibtext> </blist> <blist> <bibtext> Jamieson, Susan. 2004. " Likert scales: how to (ab)use them." Medical Education 38 (12). https://doi-org.ezp-prod1.hul.harvard.edu/ 10.1111/j.1365-2929.2004.02012.xopen_in_new.</bibtext> </blist> <blist> <bibtext> Keeley, Jared W., Taylor English, Jessica Irons, and Amber M. Henslee. 2013. " Investigating Halo and Ceiling Effects in Student Evaluations of Instruction." Educational and Psychological Measurement 440 – 457. https://doi-org.ezp-prod1.hul.harvard.edu/ 10.1177/0013164412475300.</bibtext> </blist> <blist> <bibtext> Kreitzer, Rebecca J., and Jennie Sweet-Cushman. 2022. " Evaluating Student Evaluations of Teaching: A Review of Measurement and Equity Bias in SETs and Recommendations for Ethical Reform." Journal of Acadmic Ethics 20 (1): 73 – 84. https://doi.org/10.1007/s10805-021-09400-w.</bibtext> </blist> <blist> <bibtext> Loeher, L.L., K. Haas, J. Valli-Marill, and J. Chang. 2006. " Guide to Evaluation of Instruction." Regents of the University of California.</bibtext> </blist> <blist> <bibtext> MacNell, Lillian, Adam Driscoll, and Andrea N. Hunt. 2015. " What's in a Name: Exposing Gender Bias in Student Ratings of Teaching." Innovative Higher Education 40 (4): 291 – 303. https://doi.org/10.1007/s10755-014-9313-4.</bibtext> </blist> <blist> <bibtext> Marsh, Herbert W., and Lawrence A. Roche. 1997. " Making Students' Evaluations of Teaching Effectiveness Effective." American Psychologist 52 (11): 1187 – 97. https://doi.org/10.1037/0003-066X.52.11.1187.</bibtext> </blist> <blist> <bibtext> McPherson, Michael A., R. Todd Jewell, and Myungsup Kim. 2009. " What Determines Student Evaluation Scores? A Random Effects Analysis of Undergraduate Economics Classes." Eastern Economic Journal 35 (1): 37 – 51. https://doi.org/10.1057/palgrave.eej.9050042.</bibtext> </blist> <blist> <bibtext> Mengel, Friederike, Jan Sauermann, and Ulf Zölitz. 2019. " Gender Bias in Teaching Evaluations." Journal of the European Economic Association 17 (2): 535 – 66. https://doi.org/10.1093/jeea/jvx057.</bibtext> </blist> <blist> <bibtext> Park, Eunkyoung, and John Dooris. 2020. " Predicting student evaluations of teaching using decision tree analysis." Assessment and Evaluation in Higher Education 45 (5): 776 – 793. https://doi-org.ezp-prod1.hul.harvard.edu/ 10.1080/02602938.2019.1697798.</bibtext> </blist> <blist> <bibtext> Peterson, D. A. M., L. A. Biederman, D. Andersen, T. M. Ditonto, and K. Roe. 2019. " Mitigating gender bias in student evaluations of teaching." PLoS One 14 (5). https://doi.org/10.1371/journal.pone.0216241.</bibtext> </blist> <blist> <bibtext> Radchenko, Natalia. 2020. " Student Evaluations of Teaching: Unidimensionality, Subjectivity, and Biases." Education Economics 28 (6): 549 – 66. https://doi.org/10.1080/09645292.2020.1814997.</bibtext> </blist> <blist> <bibtext> Remedios, Richard, and David A. Lieberman. 2008. " I Liked Your Course Because you Taught me Well: The Influence of Grades, Workload, Expectations and Goals on Students' Evaluations of Teaching." British Educational Research Journal 34 (1): 91 – 115. https://doi.org/10.1080/01411920701492043.</bibtext> </blist> <blist> <bibtext> Rosen, Andrew S. 2018. " Correlations, Trends and Potential Biases among Publically Accessible web-Based Student Evaluations of Teaching: A Large-Scale Study of RateMyProfessors.com Data." Assessment and Evaluation in Higher Education 43 (1): 31 – 44. https://doi.org/10.1080/02602938.2016.1276155.</bibtext> </blist> <blist> <bibtext> Spooren, Pieter, Bert Brockx, and Dimitri Mortelmans. 2013. " On the Validity of Student Evaluation of Teaching: The State of the Art." Review of Educational Research 83 (4): 598 – 642. https://doi.org/10.3102/0034654313496870.</bibtext> </blist> <blist> <bibtext> Stonebraker, Robert J., and Gary S. Stone. 2015. " Too Old to Teach? The Effect of Age on College and University Professors." Research in Higher Education 56 (8): 793 – 812. https://doi.org/10.1007/s11162-015-9374-y.</bibtext> </blist> <blist> <bibtext> Ting, Kwok-fai. 2000. " A Multilevel Perspective on Student Ratings of Instruction: Lessons from the Chinese Experience." Research in Higher Education 41 : 637. https://doi.org/10.1023/A:1007075516271.</bibtext> </blist> <blist> <bibtext> Topping, K. J., and W. L. Sanders. 2000. " Teacher Effectiveness and Computer Assessment of Reading." School Effectiveness and School Improvement 11 (3): 305 – 37. https://doi.org/10.1076/0924-3453(200009)11:3;1-G;FT305.</bibtext> </blist> <blist> <bibtext> Uttl, Bob, and Dylan Smibert. 2017. " Student Evaluations of Teaching: Teaching Quantitative Courses Can be Hazardous to One's Career." PeerJ 5 : e3299. https://doi.org/10.7717/peerj.3299.</bibtext> </blist> <blist> <bibtext> Vehovar, V., and L. Štrlekar. 2023. " When to conduct student evaluation of teaching surveys: before or after the final examination? " Assessment and Evaluation in Higher Education. https://doi.org/10.1080/02602938.2023.2298771.</bibtext> </blist> <blist> <bibtext> Yu, Kwok W., Lisa Mincieli, and Nina Zipser. 2021. " How Student Evaluations of Teaching Affect Course Enrollment." Assessment and Evaluation in Higher Education 46 (5): 779 – 92. https://doi.org/10.1080/02602938.2020.1808593.</bibtext> </blist> <blist> <bibtext> Zipser, Nina, Dmitry Kurochkin, Kwok Yu, and Lisa Mincieli. Under Review. "Gender Differences in Student Evaluations of Teaching: A Quantitative and Qualitative Analysis Using Structural Topic Modeling (Under Review).".</bibtext> </blist> <blist> <bibtext> Zipser, Nina, and Lisa Mincieli. 2018. " Administrative and Structural Changes in Student Evaluations of Teaching and Their Effects on Ovrall Instructor Scores." Assessment & Evaluation in Higher Education 43 (6): 995 – 1008. https://doi.org/10.1080/02602938.2018.1425368.</bibtext> </blist> <blist> <bibtext> Zipser, Nina, Lisa Mincieli, and Dmitry Kurochkin. 2021. " Are There Gender Differences in Quantitative Student Evaluations of Instructors? " Research in Higher Education 62 (7): 976 – 97. https://doi.org/10.1007/s11162-021-09628-w.</bibtext> </blist> </ref> <aug> <p>By Nina Zipser and Lisa Mincieli</p> <p>Reported by Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib17" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib10" firstref="ref4"></nolink> <nolink nlid="nl3" bibid="bib11" firstref="ref5"></nolink> <nolink nlid="nl4" bibid="bib18" firstref="ref6"></nolink> <nolink nlid="nl5" bibid="bib19" firstref="ref10"></nolink> <nolink nlid="nl6" bibid="bib28" firstref="ref11"></nolink> <nolink nlid="nl7" bibid="bib32" firstref="ref12"></nolink> <nolink nlid="nl8" bibid="bib14" firstref="ref14"></nolink> <nolink nlid="nl9" bibid="bib21" firstref="ref15"></nolink> <nolink nlid="nl10" bibid="bib33" firstref="ref20"></nolink> <nolink nlid="nl11" bibid="bib36" firstref="ref22"></nolink> <nolink nlid="nl12" bibid="bib16" firstref="ref23"></nolink> <nolink nlid="nl13" bibid="bib24" firstref="ref24"></nolink> <nolink nlid="nl14" bibid="bib27" firstref="ref27"></nolink> <nolink nlid="nl15" bibid="bib37" firstref="ref30"></nolink> <nolink nlid="nl16" bibid="bib26" firstref="ref39"></nolink> <nolink nlid="nl17" bibid="bib35" firstref="ref40"></nolink> <nolink nlid="nl18" bibid="bib34" firstref="ref41"></nolink> <nolink nlid="nl19" bibid="bib15" firstref="ref52"></nolink> <nolink nlid="nl20" bibid="bib23" firstref="ref53"></nolink> <nolink nlid="nl21" bibid="bib20" firstref="ref57"></nolink> <nolink nlid="nl22" bibid="bib25" firstref="ref66"></nolink> <nolink nlid="nl23" bibid="bib30" firstref="ref67"></nolink> <nolink nlid="nl24" bibid="bib13" firstref="ref70"></nolink> <nolink nlid="nl25" bibid="bib22" firstref="ref74"></nolink> <nolink nlid="nl26" bibid="bib29" firstref="ref77"></nolink> <nolink nlid="nl27" bibid="bib12" firstref="ref83"></nolink>
Header DbId: eric
DbLabel: ERIC
An: EJ1455581
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: A Framework for Using SET When Evaluating Faculty
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Nina+Zipser%22">Nina Zipser</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-1431-1606">0000-0002-1431-1606</externalLink>)<br /><searchLink fieldCode="AR" term="%22Lisa+Mincieli%22">Lisa Mincieli</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-7971-8338">0000-0001-7971-8338</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Studies+in+Higher+Education%22"><i>Studies in Higher Education</i></searchLink>. 2025 50(1):168-182.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Routledge. Available from: Taylor & Francis, Ltd. 530 Walnut Street Suite 850, Philadelphia, PA 19106. Tel: 800-354-1420; Tel: 215-625-8900; Fax: 215-207-0050; Web site: http://www.tandf.co.uk/journals
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 15
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2025
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Audience
  Label: Education Level
  Group: Audnce
  Data: <searchLink fieldCode="EL" term="%22Higher+Education%22">Higher Education</searchLink><br /><searchLink fieldCode="EL" term="%22Postsecondary+Education%22">Postsecondary Education</searchLink>
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22College+Faculty%22">College Faculty</searchLink><br /><searchLink fieldCode="DE" term="%22Faculty+Evaluation%22">Faculty Evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Student+Evaluation+of+Teacher+Performance%22">Student Evaluation of Teacher Performance</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluation+Criteria%22">Evaluation Criteria</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluation+Methods%22">Evaluation Methods</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluation+Problems%22">Evaluation Problems</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluation+Utilization%22">Evaluation Utilization</searchLink><br /><searchLink fieldCode="DE" term="%22Data+Interpretation%22">Data Interpretation</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1080/03075079.2024.2332412
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0307-5079<br />1470-174X
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: This paper presents a framework for utilizing Student Evaluations of Teaching (SET) in faculty evaluations. Recognizing the ongoing debate about the validity of SET as a measure of teaching effectiveness, the authors agree with scholars who propose viewing SET as a tool for gauging 'student perceptions of learning'. They present a method that combines post-estimation residuals from an Ordinary Least Squares (OLS) model, which accounts for key confounding factors, with simple mean SET scores. This approach provides a more nuanced perspective on each faculty member's individual ratings. The authors draw on SET data from courses taught by tenured faculty in the Faculty of Arts and Sciences at a large, private university in the Northeastern United States. Their findings suggest that the proposed residual framework can help address some of the concerns associated with the use of SET in personnel decision-making. However, they caution that this approach requires careful explanation and should be used in conjunction with a comprehensive review of a professor's full teaching portfolio. The study contributes to the ongoing discourse on the use of SET in faculty evaluations, offering a more equitable and nuanced approach to interpreting SET data.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2025
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1455581
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1455581
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1080/03075079.2024.2332412
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 15
        StartPage: 168
    Subjects:
      – SubjectFull: College Faculty
        Type: general
      – SubjectFull: Faculty Evaluation
        Type: general
      – SubjectFull: Student Evaluation of Teacher Performance
        Type: general
      – SubjectFull: Evaluation Criteria
        Type: general
      – SubjectFull: Evaluation Methods
        Type: general
      – SubjectFull: Evaluation Problems
        Type: general
      – SubjectFull: Evaluation Utilization
        Type: general
      – SubjectFull: Data Interpretation
        Type: general
    Titles:
      – TitleFull: A Framework for Using SET When Evaluating Faculty
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Nina Zipser
      – PersonEntity:
          Name:
            NameFull: Lisa Mincieli
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 0307-5079
            – Type: issn-electronic
              Value: 1470-174X
          Numbering:
            – Type: volume
              Value: 50
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Studies in Higher Education
              Type: main
ResultId 1