What's in a Score? A Decomposition of Observation Scores across Various Sources to Better Understand What Contributes to Scores

Saved in:
Bibliographic Details
Title: What's in a Score? A Decomposition of Observation Scores across Various Sources to Better Understand What Contributes to Scores
Language: English
Authors: Mark White
Source: Practical Assessment, Research & Evaluation. 2025 30.
Availability: University of Massachusetts Amherst Libraries. 154 Hicks Way, Amherst, MA 01003. e-mail: pare@umass.edu; Web site: https://openpublishing.library.umass.edu/pare/
Peer Reviewed: Y
Page Count: 34
Publication Date: 2025
Document Type: Journal Articles
Reports - Research
Education Level: Junior High Schools
Middle Schools
Secondary Education
Descriptors: Scores, Teacher Effectiveness, Teacher Evaluation, Observation, Accuracy, Interrater Reliability, Teaching Conditions, Student Characteristics, Teacher Characteristics, English Teachers, Middle School Teachers, Middle School Students, Classroom Environment
Assessment and Survey Identifiers: Classroom Assessment Scoring System
ISSN: 1531-7714
Abstract: Systematized, observational approaches to measuring teaching quality are an important tool in research and practice. Termed observation systems, these approaches include a rubric that operationalizes a set of teaching quality constructs and structures to support rater training and monitoring. Scores from observation systems, through their interpretation as capturing the intended teaching quality constructs, are used to develop theoretical understandings of teaching quality. This paper explores what factors contribute to observation scores in a secondary analysis of the Understanding Teaching Quality project. Leveraging calibration data, I combine mixed-effects regression analyses of calibration data that examine rater accuracy (i.e., deviations from master scores) with analyses of operational data to explore the extent to which raters, students, teachers, and the teaching context contribute to scores. These analyses highlight that (1) some rater error may be invisible in typical analyses examining rater agreement; (2) rater error is largely systematic; and (3) differences in student composition across schools largely explain between-school differences in scores. These results highlight potential biases in estimates of score reliability and validity coefficients that might exist in studies that model rater agreement rather than rater accuracy and/or that fail to consider differences in between-teacher and between-school variation in observation scores.
Abstractor: As Provided
Entry Date: 2025
Accession Number: EJ1472538
Database: ERIC
Description
Abstract:Systematized, observational approaches to measuring teaching quality are an important tool in research and practice. Termed observation systems, these approaches include a rubric that operationalizes a set of teaching quality constructs and structures to support rater training and monitoring. Scores from observation systems, through their interpretation as capturing the intended teaching quality constructs, are used to develop theoretical understandings of teaching quality. This paper explores what factors contribute to observation scores in a secondary analysis of the Understanding Teaching Quality project. Leveraging calibration data, I combine mixed-effects regression analyses of calibration data that examine rater accuracy (i.e., deviations from master scores) with analyses of operational data to explore the extent to which raters, students, teachers, and the teaching context contribute to scores. These analyses highlight that (1) some rater error may be invisible in typical analyses examining rater agreement; (2) rater error is largely systematic; and (3) differences in student composition across schools largely explain between-school differences in scores. These results highlight potential biases in estimates of score reliability and validity coefficients that might exist in studies that model rater agreement rather than rater accuracy and/or that fail to consider differences in between-teacher and between-school variation in observation scores.
ISSN:1531-7714