Examining the Dual Purpose Use of Student Learning Objectives for Classroom Assessment and Teacher Evaluation
Saved in:
| Title: | Examining the Dual Purpose Use of Student Learning Objectives for Classroom Assessment and Teacher Evaluation |
|---|---|
| Language: | English |
| Authors: | Briggs, Derek C., Chattergoon, Rajendra, Burkhardt, Amy |
| Source: | Journal of Educational Measurement. Win 2019 56(4):686-714. |
| Availability: | Wiley-Blackwell. 350 Main Street, Malden, MA 02148. Tel: 800-835-6770; Tel: 781-388-8598; Fax: 781-388-8232; e-mail: cs-journals@wiley.com; Web site: http://www.wiley.com/WileyCDA |
| Peer Reviewed: | Y |
| Page Count: | 29 |
| Publication Date: | 2019 |
| Document Type: | Journal Articles Reports - Evaluative |
| Descriptors: | Student Educational Objectives, Student Evaluation, Teacher Evaluation, Scores, Urban Schools, School District Size, Growth Models, Measurement, High Stakes Tests, Inferences |
| DOI: | 10.1111/jedm.12233 |
| ISSN: | 0022-0655 |
| Abstract: | The process of setting and evaluating student learning objectives (SLOs) has become increasingly popular as an example where classroom assessment is intended to fulfill the dual purpose use of informing instruction and holding teachers accountable. A concern is that the high-stakes purpose may lead to distortions in the inferences about students and teachers that SLOs can support. This concern is explored in the present study by contrasting student SLO scores in a large urban school district to performance on a common objective external criterion. This external criterion is used to evaluate the extent to which student growth scores appear to be inflated. Using 2 years of data, growth comparisons are also made at the teacher level for teachers who submit SLOs and have students that take the state-administered large-scale assessment. Although they do show similar relationships with demographic covariates and have the same degree of stability across years, the two different measures of growth are weakly correlated. |
| Abstractor: | As Provided |
| Entry Date: | 2019 |
| Accession Number: | EJ1236254 |
| Database: | ERIC |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwHweiSfketWbPQWEdpo4lwkAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDEwecBp1AEGjvlMGtgIBEICBmzA0N6uplBGsrRoBjnA6clDj3Qcedzfo3S5Gbw9uMePmQkmz5guyFfLmopfFWX5NfaB2KN9rH7uZh_LzsbuV5UIeb6HYJzh7oqX6G4X0QQ24hkjLOp2vMPgAyfg-7eDAPQL98w-YpDnIz2nonoYKlJqQU0Vm1wa5w0foVcX6leFJzA6-4DZRDQAWHkdIfJlU4FaJdH85p31Lv48C Text: Availability: 1 Value: <anid>AN0140159251;mea01dec.19;2019Dec07.04:40;v2.2.500</anid> <title id="AN0140159251-1">Examining the Dual Purpose Use of Student Learning Objectives for Classroom Assessment and Teacher Evaluation </title> <p>The process of setting and evaluating student learning objectives (SLOs) has become increasingly popular as an example where classroom assessment is intended to fulfill the dual purpose use of informing instruction and holding teachers accountable. A concern is that the high‐stakes purpose may lead to distortions in the inferences about students and teachers that SLOs can support. This concern is explored in the present study by contrasting student SLO scores in a large urban school district to performance on a common objective external criterion. This external criterion is used to evaluate the extent to which student growth scores appear to be inflated. Using 2 years of data, growth comparisons are also made at the teacher level for teachers who submit SLOs and have students that take the state‐administered large‐scale assessment. Although they do show similar relationships with demographic covariates and have the same degree of stability across years, the two different measures of growth are weakly correlated.</p> <p>Classroom assessment and large‐scale assessment are generally understood to serve different purposes. To the extent that there is a conventional distinction to be drawn between the two, it is that classroom assessments are optimal as a means of conveying feedback about student learning (i.e., assessment <emph>for</emph> learning), whereas large‐scale assessments are optimal as a means of conveying information that can be used to monitor and evaluate student learning (i.e., assessment <emph>of</emph> learning). This distinction seems sensible because classroom assessments tend to be locally developed, their administration is in the control of teachers, and they are generally proximal to the teacher and students' enacted curriculum (cf., Ruiz‐Primo, Shavelson, Hamilton, &amp; Klein, [<reflink idref="bib22" id="ref1">22</reflink>]). Furthermore, the feedback students and teachers receive from these assessments is often immediate or follows from a short time lag. In contrast, large‐scale assessments are externally developed, their administration is outside the control of teachers, and the results are typically not available for many months. However, because of their standardization, they provide for an efficient and reliable way to compare students across classrooms, schools, and districts.</p> <p>Although these conventional distinctions in the purposes served by classroom and large‐scale assessments might appear self‐evident, an increase in educational accountability policies that emphasize teacher evaluation has led to some blurring of boundaries. In many instances, large‐scale assessments have been increasingly expected to not only provide information relevant to systems of educational accountability, but to also provide information that is formatively useful to teachers, students, and parents. One reason for this is the potential for computerized adaptive assessments to provide teachers with much more timely (even immediate) feedback on student performance relative to traditional paper‐and‐pencil exams.[<reflink idref="bib1" id="ref2">1</reflink>] This boundary blurring between assessments used for both high‐ and low‐stakes purposes is also becoming more prevalent coming from the opposite direction. That is, in some states and school districts, locally developed assessments, previously used solely to inform instruction or support student grading within classrooms, are being used for comparative purposes <emph>across</emph> classrooms. One of the highest profile examples of this can be found in the process commonly used to set and evaluate <emph>student learning objectives</emph> (SLOs).</p> <p>Our focus in the present study is on the use of SLOs in support of teacher evaluation <emph>and</emph> formative assessment practices in Denver Public Schools (DPS). On the one hand, there is a clear desire among DPS assessment and evaluation staff for SLOs to be used formatively to inform instructional practice. On the other hand, DPS teachers are well aware that SLOs also provide the district with aggregate measures that can count both toward a teacher's monetary compensation and their annual evaluation. The motivating question we explore in this study is whether there is evidence that the use of classroom assessment for high‐stakes purposes vis‐à‐vis their role in SLOs is associated with distortions in the information the assessments convey about student growth.</p> <hd id="AN0140159251-2">Student Learning Objectives</hd> <p>Over the past decade, a growing number of states and school districts in the United States have sought to incorporate evidence of student growth into formal evaluations of teachers and schools under the auspices of educational accountability (cf. Doherty &amp; Jacobs, [<reflink idref="bib8" id="ref3">8</reflink>]). Although a great deal of research and debate has surrounded the use of statistical models for this purpose, only about one‐third of classroom teachers teach students for whom state‐administered standardized tests are available as inputs (Hall, Gagnon, Schneider, Marion, &amp; Thompson, [<reflink idref="bib12" id="ref4">12</reflink>]). Hence, for a majority of teachers other evidence is needed to support inferences about student growth. In more than 30 states, this evidence has been, and/or continues to be, gathered through the development and evaluation of student growth using SLOs (Lacireno‐Paquet, Morgan, &amp; Mello, [<reflink idref="bib18" id="ref5">18</reflink>]). Indeed, in many places, SLOs are now being developed and submitted by teachers in both state‐tested and non–state‐tested subjects and grades (Hall et al., [<reflink idref="bib12" id="ref6">12</reflink>]). Common elements of the SLO process typically include the following:</p> <p></p> <ulist> <item> specification of a goal as part of an "objective statement" that defines the content and skills students are expected to learn;</item> <p></p> <item> specification of the interval of instruction over which the learning is expected to occur (e.g., semester or year‐long);</item> <p></p> <item> identification of assessments to be given to students at the beginning and end of the instructional interval;</item> <p></p> <item> specification of growth targets set for individual students based on performance achieved on identified assessments; and</item> <p></p> <item> designation of a teacher rating based on growth achieved by students.</item> </ulist> <p>The implementation of these common SLO elements varies across states and districts depending upon the degree of centralization and comparability desired in the set of learning goals used across all teachers, the set of data sources used to support the process, and the methodology used to compute teacher ratings (Lachlan‐Haché, Cushing, &amp; Bivona, [<reflink idref="bib17" id="ref7">17</reflink>]).</p> <p>In theory, SLOs can provide a means for teachers to establish learning goals, monitor students' progress toward these goals, and then evaluate the degree to which students achieve these goals using relevant measures. In this sense, SLOs merely represent the formalization of a process that should be part of good teaching. What makes them unique is that beyond their instructional use, SLOs are intended to provide a quantitative source of evidence that can be used to differentiate effective and ineffective teachers. As a quantitative indicator, this brings SLOs squarely into the realm of what has been referred to as Campbell's Law (Briggs, [<reflink idref="bib3" id="ref8">3</reflink>]; Campbell, [<reflink idref="bib5" id="ref9">5</reflink>]):</p> <p>The more any quantitative social indicator (or even some qualitative indicator) is used for social decision‐making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor. (Campbell, [<reflink idref="bib5" id="ref10">5</reflink>], p. 49)</p> <p>SLOs are especially susceptible to corruption because it will often be the case (as it is in the data context in this study) that many, or even all, of the elements that figure into the quantification and aggregation of student growth—the choice of learning objective, the choice of assessments to administer, the scoring of these assessments, and the classification of students into mastery levels—are directly influenced by teachers' judgments.</p> <hd id="AN0140159251-3">Student Learning Objectives in Denver Public Schools</hd> <p>Denver Public Schools was one of the first school districts in the country to establish a formal process for using locally developed assessments to serve both instructional and evaluative purposes (Gonring, Teske, &amp; Jupp, [<reflink idref="bib11" id="ref11">11</reflink>]; Hershberg &amp; Robertson‐Kraft, [<reflink idref="bib13" id="ref12">13</reflink>]). A pilot version of SLOs (called "Student Growth Objectives") was being used by subsets of DPS as far back as 2001, and by 2006 all DPS teachers received small salary bonuses or salary increases of up to $376 if some designated percentage of students could be shown to have met one or two of their growth objectives. More recently, as of the 2015–2016 school year, results from SLOs have been factored into teacher evaluation. It can be argued that the dual use concept for SLOs was pioneered by DPS, and that the SLO process taken in many other districts and states across the country represents an emulation of the DPS approach.</p> <hd id="AN0140159251-4">The Role SLOs Play in the DPS Accountability Context</hd> <p>By state law,[<reflink idref="bib2" id="ref13">2</reflink>] public school teachers in Colorado must be evaluated annually, and 50% of this evaluation must be based on evidence of student growth in academic achievement. In DPS, 50% of a teacher's overall leading effective academic practice (LEAP)[<reflink idref="bib3" id="ref14">3</reflink>] evaluation score comes from three sources: a collective measure of school growth, student growth on state‐administered achievement tests, and student growth as indicated by SLO performance. Among these three sources, a measure of school‐level growth contributes 10%, and another 10% comes from the mean of student growth percentiles (Betebenner, [<reflink idref="bib2" id="ref15">2</reflink>]) computed for students who have taken the state‐administered tests (i.e., the Partnership for Assessment of Readiness in College and Career Consortium [PARCC tests]) in English Language Arts (ELA) and/or Math two years in a row. For these same teachers, SLOs contribute the remaining 30% of growth evidence toward a LEAP rating. For teachers that do not have students who take state‐administered tests, SLOs contribute 40% to their overall LEAP rating. Hence, for all teachers, whether their students take state‐administered large‐scale assessments or not, student performance on SLOs is by far the predominant source of information used to satisfy the growth component of the evaluation system under Colorado state law.</p> <hd id="AN0140159251-5">Components of SLOs at Denver Public Schools</hd> <p>In this study, we focus on DPS SLO data from the 2015–2016 and 2016–2017 school years. DPS's Department of Accountability, Research and Evaluation offered the following description of SLOs in the 2016 edition of its <emph>SLO Teacher Handbook</emph> under the heading, "What are Student Learning Objectives?":</p> <p>Effective teachers have learning goals for their students and use assessments to measure progress toward these goals. They have a deep understanding of where students are at the beginning of a course, and what they can achieve by the end. Effective teachers analyze standards, select and administer rigorous assessments aligned to those standards, and measure how their students grow during the school year. They use this data to drive their instruction and are constantly reflecting on and refining their craft. Student Learning Objectives (SLOs) embody these effective pedagogical practices by helping DPS educators focus on high impact standards, set ambitious learning goals, and measure students' progress toward attaining them. This process will yield greater student growth on critical learning outcomes by allowing teachers to plan backward from an end vision of student success, ensuring that every minute of instruction is geared toward our district vision that <emph>Every Child Succeeds</emph>.</p> <p>This description makes the intended formative use of SLOs fairly clear. Beyond the evaluative role they play as an input into a teacher's LEAP rating, SLOs are expected to provide a tool for teachers to use to facilitate student achievement and to improve upon their pedagogical practices.</p> <p>Any given SLO in DPS contains three key components: an objective statement, performance criteria, and a "learning progression" rubric. Although teachers can create SLOs on their own, and many do, DPS staff provide two "template" SLOs in math and ELA for each grade level from Kindergarten through the 12th grade. To provide a concrete example, consider a district‐developed Grade 2 math SLO template for understanding the concept of place value. The objective statement of this SLO is as follows:</p> <p>All students will be able to model and explain the meaning of the digits in three‐digit numbers as the amount of hundreds, tens, and ones both verbally and in writing. Students will apply this understanding to compare two three‐digit numbers, use the symbols &gt;, &lt;, and = to record the comparison results, and justify the comparison both verbally and in writing.</p> <p>The four performance criteria used to operationalize this objective are</p> <p></p> <ulist> <item> Students demonstrate understanding of the three digits of a three‐digit number by independently modeling a three‐digit number with a visual representation.</item> <p></p> <item> Students compare two three‐digit numbers and record the results using symbols &gt;, &lt;, and =.</item> <p></p> <item> Students independently express a three‐digit number in expanded notation. Students represent the quantity in terms of hundreds, tens, and ones.</item> <p></p> <item> Students use grade‐level academic and content language to explain their representations and justify comparisons orally and in writing based on the place value system.</item> </ulist> <p>Finally, a rubric is used to score/categorize students with respect to each of these performance criteria on a criterion‐referenced scale from 1 to 4. Along with these three components, each DPS template comes with one or more performance‐based assessments that teachers are encouraged to administer near the end of the SLO instructional period. Two of the six items for a Grade 2 place value assessment are depicted in Figure . This assessment was designed under the leadership of content experts in the district's curriculum and instruction division, and was written to align to its associated SLO objective statements and performance criteria, which are themselves aligned to the Common Core State Standards.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/MEA/01dec19/jedm12233-fig-0001.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jedm12233-fig-0001.jpg" title="District‐developed assessment tasks for place value SLO." /> </p> <p></p> <hd id="AN0140159251-7">Monitoring and Scoring of Students on SLOs</hd> <p>At the student level, there are three important SLO‐related variables: (a) preparedness levels, (b) end‐of‐year command levels, and (c) growth points earned. DPS teachers assign students to preparedness levels at the outset of the instructional period for an SLO (typically the beginning of the academic school year in September). They are able to choose from one of five preparedness levels:</p> <p></p> <ulist> <item> <bold> Significantly Underprepared [1] </bold> : Students who enter the course/grade with particularly minimal mastery of the prerequisite knowledge and skills for the course/grade.</item> <p></p> <item> <bold> Underprepared [2] </bold> : Students who enter the course/grade with minimal mastery of the prerequisite knowledge and skills for the course/grade.</item> <p></p> <item> <bold> Somewhat Prepared [3] </bold> : Students who enter the course/grade has some, but not all, prerequisite knowledge and skills for the course/grade</item> <p></p> <item> <bold> Prepared [4] </bold> : Students who enter the course/grade with sufficient prerequisite knowledge and skills for the course/grade. Students are academically prepared to engage in the content area of the SLO.</item> <p></p> <item> <bold> Ahead [5] </bold> : Students who enter the course/grade with a deep command of the prerequisite knowledge and skills for the course/grade. These students are able to apply previous learning to a variety of contexts.</item> </ulist> <p>At the end of a school year, teachers rate students' command of the SLO. They can choose from one of five command levels: below limited command, limited command, moderate command, strong command, or distinguished command.[<reflink idref="bib4" id="ref16">4</reflink>] In between the beginning and end of the school year, teachers are expected to monitor their students' progress, and to this end teachers can use one or more of the assessments provided by the district (if using a district template SLO, see example in Figure ) and/or they can administer their own assessment tasks. Figure  provides an example of two assessment items for place value developed by a group of Grade 2 teachers at a specific DPS elementary school. These particular teachers met regularly as part of a professional learning community and were encouraged to share examples of their students' responses to progress monitoring tasks such as the ones illustrated in Figures  and . The intent was for discussions of these assessment results to prompt suggestions and strategies that could help teachers improve their instructional practices (for details, see Briggs et al., [<reflink idref="bib4" id="ref17">4</reflink>]). However, these activities are not mandated or standardized, so the extent to which they occur will vary from school to school.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/MEA/01dec19/jedm12233-fig-0002.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jedm12233-fig-0002.jpg" title="Example of a locally developed assessment for monitoring understanding of place value. [Color figure can be viewed at wileyonlinelibrary.com]" /> </p> <p></p> <p>In classifying their students both in terms of preparedness for the SLO at the start of the year and level of command by the end of the year, teachers are instructed to use their professional judgment on the basis of the "body of evidence" collected for each student. Teachers are given considerable autonomy in classifying students into preparedness and command categories. For preparedness classifications, these measures are assembled primarily from student performance in prior year classes and assessments. For end‐of‐year command classifications, a body of evidence is defined formally by DPS as "data derived from a variety of assessment tools that measure the degree to which students are progressing toward each Performance Criterion, and more broadly the Objective Statement." For students who use one of the district template SLOs, results from performance‐based items (such as those illustrated in Figure ) would be expected to play a prominent role, but these could also be supplemented with other teacher‐developed assessments (such as the one illustrated in Figure ). With regard to both classification decisions, teachers are expected to be able to provide a "strong, clear and thorough rationale" for the evidence used to support the classification, and each teacher's SLO must be approved by their principal. However, beyond these guiding principles there is little standardization in the specific process that teachers use to make SLO classifications.</p> <p>Student growth scores are computed according to the relationship between a student's preparedness level and end‐of‐year command levels. With a few exceptions, these growth scores are computed automatically by an online SLO application using a series of decision rules. The version of the decision rules used during the 2015–2016 school year is shown in Figure . (The gray boxes in Figure  reflect the decision points where manual entries are made by the evaluator to finalize a teacher rating.)</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/MEA/01dec19/jedm12233-fig-0003.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jedm12233-fig-0003.jpg" title="DPS SLO Scoring Matrix.Note: Cells with gray shading reflect the decision points where manual entries are made by the evaluator to finalize a teacher rating.* For students starting a course Prepared or Ahead, Below Limited Command represents less mastery than they begin the course with. This is based on Limited Command representing a level of mastery expected around the beginning of a course. Thus, Below Limited Command is not an option for these students. Limited Command is the lowest level of mastery they can attain." /> </p> <p></p> <p>Finally, after SLO growth points have been computed for each student, a teacher‐level SLO score is computed as the percent of maximum possible growth points earned by a teacher's students, where this maximum represents the number of students in the teacher's class multiplied by 3 (because each student can theoretically earn up to 3 points for their end of course level of command, regardless of their rated level of preparedness).</p> <hd id="AN0140159251-10">Some Threats to the Validity of SLOs as an Indicator of Growth</hd> <p>The SLO scoring matrix depicted in Figure  represents an example of a <emph>value table</emph> (Castellano &amp; Ho, [<reflink idref="bib6" id="ref18">6</reflink>]) in which a student's end of course achievement status is differentially valued as a function of his or her incoming starting point or preparedness level. Because of this, a student that starts the year "significantly underprepared" or "underprepared" but ends at a "moderate" level of command for the SLO, can earn the same three points for growth as a student that starts the year "prepared" and ends with a "distinguished" level of command. The intent is to give both students and their teacher credit for the progress made during the school year, as opposed to potentially penalizing them for a lack of opportunity in the past.</p> <p>In this study, we examine a fundamental threat to the validity of SLOs as teacher‐level measures of student growth: the potential for distortion caused by Campbell's Law. Unlike student growth percentiles (SGPs), which are based on objective and standardized measures of student achievement, SLOs are based upon teachers' holistic classifications of student achievement. Because these scores can represent 30%–40% of a teacher's LEAP rating, this could create an incentive for teachers to inflate student growth either by classifying students as less "prepared" for an SLO than they actually are at the outset of an instructional period, or by classifying students as having a higher level of command than they have actually attained at the end of an instructional period. We look for correlational evidence that this kind of score distortion is occurring by using information about prior and current year student performance on Colorado's state‐administered assessment (i.e., PARCC tests in ELA and Math at the time of this study) as a criterion against which the SLO preparedness of students can be compared. We examine whether students who scored in the "met expectations" or "exceeds expectations" performance levels on a PARCC test in ELA or Math for their previous grade are classified as "prepared" or "ahead" for their same subject SLO in the following grade. We consider a mismatch between a PARCC classification and an SLO preparedness classification as evidence of a classification that appears to be inconsistent. We then more formally examine whether certain student characteristics are associated with the probability of a higher preparedness classification after controlling for prior year PARCC performance. We take a parallel approach with respect to end‐of‐year SLO command, but in this case we look for evidence that these scores have been inflated. Finally, we use a regression model to examine the extent to which a teacher‐level SLO measure is associated with the aggregate characteristics of a teacher's students. We then contrast the use of a teacher‐level measure of student growth based on SLOs to one based on to the use of SGPs (the mean SGP, or MGP).</p> <hd id="AN0140159251-11">Data and Descriptive Statistics</hd> <p>Our populations in 2015–2016 and 2016–2017 consist of all traditional district schools and teachers within those schools that implemented the SLO process. Alternative schools are excluded along with individual students who were not assigned SLO growth ratings due to factors such as attendance. One important distinction between the data collected in these two years is that 2015–2016 represented the first year in which SLOs were fully implemented as a component of teachers' LEAP evaluation throughout all schools in DPS. However, in this first year of full implementation teachers were only required to submit one SLO. In the second year of full implementation, 2016–2017, teachers were required to submit two SLOs. Another important distinction between the first and second year of this study is in the availability of prior grade PARCC scores to inform teacher preparedness ratings. Due to a delay in the release of these scores from the 2014–2015 administration, prior grade scores were not available to DPS teachers in the fall of 2015 when they were making preparedness classifications for the 2015–2016 school year. In contrast, prior grade PARCC scores were available for teachers to consult in the fall of 2016. This distinction likely explains one of the findings we present later, and we will return to it then.</p> <p>As shown in Table , there were a total of 132 schools, 4,187 teachers, and 55,479 unique students with SLO ratings in 2015–2016. The following year, in 2016–2017, there were a total of 140 schools, 4,383 teachers, and 61,023 students. We further restrict the population in each year to only those teachers who chose SLOs that came from district‐developed templates in the subjects of ELA and Math. We do this for two reasons. First, it allows for a contrast within the same subject domain relative to student performance on the PARCC assessment. Second, it reduces some of the variance in SLO outcomes from the inclusion of SLOs that may vary tremendously in quality. Filtering on the use of district‐developed SLOs throughout ensures at least some minimum level of comparability with respect to the components of the SLO process (objective statement, performance criteria, rubric, and performance‐based tasks), but it also makes the resulting analyses something of a best case scenario.[<reflink idref="bib5" id="ref19">5</reflink>]</p> <p>Analytic Populations for Analyses</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th /&gt;&lt;th /&gt;&lt;th /&gt;&lt;th align="center"&gt;District Templates&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th /&gt;&lt;th align="center"&gt;All SLOs&lt;/th&gt;&lt;th align="center"&gt;ELA SLOs&lt;/th&gt;&lt;th align="center"&gt;Math SLOs&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th /&gt;&lt;th align="center"&gt;2015&amp;#8211;2016&lt;/th&gt;&lt;th align="center"&gt;2016&amp;#8211;2017&lt;/th&gt;&lt;th align="center"&gt;2015&amp;#8211;2016&lt;/th&gt;&lt;th align="center"&gt;2016&amp;#8211;2017&lt;/th&gt;&lt;th align="center"&gt;2015&amp;#8211;2016&lt;/th&gt;&lt;th align="center"&gt;2016&amp;#8211;2017&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Students&lt;/td&gt;&lt;td&gt;56,295&lt;/td&gt;&lt;td&gt;61,023&lt;/td&gt;&lt;td&gt;25,115&lt;/td&gt;&lt;td&gt;27,463&lt;/td&gt;&lt;td&gt;18,238&lt;/td&gt;&lt;td&gt;28,671&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Teachers&lt;/td&gt;&lt;td&gt;4,267&lt;/td&gt;&lt;td&gt;4,383&lt;/td&gt;&lt;td&gt;1,415&lt;/td&gt;&lt;td&gt;1,444&lt;/td&gt;&lt;td&gt;872&lt;/td&gt;&lt;td&gt;1,316&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Schools&lt;/td&gt;&lt;td&gt;135&lt;/td&gt;&lt;td&gt;140&lt;/td&gt;&lt;td&gt;126&lt;/td&gt;&lt;td&gt;133&lt;/td&gt;&lt;td&gt;120&lt;/td&gt;&lt;td&gt;132&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Figure  presents frequency distributions of SLO preparedness, end‐of‐year command, and growth points across all students in both academic school years. Interestingly, these distributions are almost identical by year and by subject domain. In both years, the average DPS student was classified as "somewhat prepared" for his or her SLO at the start of the school year (mean between 2.6 and 2.7), and as having "moderate command" of the SLO by the end of the instructional period (mean between 3.2 and 3.3). Most students received 2 or 3 growth points out of a maximum possible of 3 points (after applying the scoring rules shown in Figure ). This suggests that, according to their teachers, the vast majority of DPS students showed evidence of "moderate" to "high" growth.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/MEA/01dec19/jedm12233-fig-0004.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jedm12233-fig-0004.jpg" title="Distributions of SLO outcomes for ELA and math students by year." /> </p> <p></p> <p>The SLO variable of interest at the teacher level is the percent of SLO points earned by each teacher. On the basis of their percent of SLO points earned (which we will refer to henceforth as "SLO Growth") and cut scores set by DPS leadership, teachers fell into one of four effectiveness categories. For example, out of 1,415 teachers with ELA SLOs in 2015–2016, 0.3% were rated "ineffective," 8.6% of teachers were rated "approaching," 60.5% were rated "effective," and 30.6% were rated "distinguished." One year later, in 2016–2017, out of 1,444 teachers with ELA SLOs, 0.3% were rated "ineffective," 15.4% were rated "approaching," 72.5% were rated "effective," and 11.7% were rated "distinguished." Similar changes occurred for Math. Figure  displays the distribution of this variable by SLO subject and year with the cut points for LEAP effectiveness categories from 2015–2016 superimposed.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/MEA/01dec19/jedm12233-fig-0005.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jedm12233-fig-0005.jpg" title="Distribution of percent of SLO growth points earned by a teacher by subject and year. [Color figure can be viewed at wileyonlinelibrary.com]" /> </p> <p></p> <p>The distribution of the SLO Growth variable has a strong skew toward high values. To put this in perspective, one‐third of teachers with ELA SLOs in 2015–2016 (486 out of 1,415) and about a quarter in 2016–2017 (315 out of 1,444) earned 90% or more of possible SLO growth points. At the extremes, 14% and 6% in each year earned 100% of possible SLO growth points. This is indicative of a ceiling effect on teacher growth scores, although the effect appears to have diminished to a considerable extent from the first to second year of full SLO implementation. The skew toward high values may be the result of SLOs that were not sufficiently ambitious (i.e., it was very easy for students to demonstrate strong or distinguished command) or of student growth scores that were inflated by the way teachers classified students with respect to preparedness and end of year command. We explore this latter hypothesis in the next section.</p> <hd id="AN0140159251-14">Validity of SLO Preparedness and Command Classifications</hd> <p></p> <hd id="AN0140159251-15">Preparedness Classifications</hd> <p>To examine the validity of teachers' SLO preparedness classifications, we use as an external indicator students' ELA and Math PARCC scores from the previous school year (either the spring of 2015, or the spring of 2016) as a criterion to characterize the consistency of teachers' decisions. An important limitation of this analysis is that it can only be performed for students with ELA and/or Math SLO classifications who have PARCC scores from the preceding year. For example, using 2015–2016 data, this restriction reduces our student sample from 25,115 to 9,760 in ELA, and from 18,238 to 8,642 in Math.</p> <p>Table  shows the cross‐tabulation of these two student classifications using 2015–2016 data (we created and examined a similar cross‐tabulation with 2016–2017 data). There are a number of different ways that one could characterize preparedness classifications as consistent with PARCC test score performance. The strictest rule would be a one‐to‐one agreement between PARCC performance levels and SLO preparedness levels such that only cells along the main diagonal of each cross‐tabulation are considered consistent. Using this criterion with 2015–2016 data, we would find that 39.9% of ELA and 41.5% of Math SLO preparedness classifications were consistent. Using 2016–2017 data, these values increase to 47.3% and 53.3%, respectively, likely because prior grade PARCC scores were available for teachers to consult when making preparedness ratings in 2016–2017, but not in 2015–2016.</p> <p>SLO Preparedness Levels by PARCC Performance Levels: 2015–2016 Data</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;PARCC Performance Level&lt;/th&gt;&lt;th align="center"&gt;SLO Preparedness Level Fall 2015&lt;/th&gt;&lt;th /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th /&gt;&lt;th align="center"&gt;Sig. Up.&lt;/th&gt;&lt;th align="center"&gt;Up.&lt;/th&gt;&lt;th align="center"&gt;S. Prep.&lt;/th&gt;&lt;th align="center"&gt;Prepared&lt;/th&gt;&lt;th align="center"&gt;Ahead&lt;/th&gt;&lt;th align="center"&gt;Totals&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th&gt;Spring 2015&lt;/th&gt;&lt;th align="center"&gt;[1]&lt;/th&gt;&lt;th align="center"&gt;[2]&lt;/th&gt;&lt;th align="center"&gt;[3]&lt;/th&gt;&lt;th align="center"&gt;[4]&lt;/th&gt;&lt;th align="center"&gt;[5]&lt;/th&gt;&lt;th /&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;English Language Arts&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Did Not Yet Meet Expectations [1]&lt;/td&gt;&lt;td&gt;1,086&lt;/td&gt;&lt;td&gt;680&lt;/td&gt;&lt;td&gt;378&lt;/td&gt;&lt;td&gt;45&lt;/td&gt;&lt;td&gt;2&lt;/td&gt;&lt;td&gt;2,191&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Partially Met Expectations [2]&lt;/td&gt;&lt;td&gt;477&lt;/td&gt;&lt;td&gt;815&lt;/td&gt;&lt;td&gt;761&lt;/td&gt;&lt;td&gt;160&lt;/td&gt;&lt;td&gt;3&lt;/td&gt;&lt;td&gt;2,216&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Approached Expectations [3]&lt;/td&gt;&lt;td&gt;211&lt;/td&gt;&lt;td&gt;603&lt;/td&gt;&lt;td&gt;985&lt;/td&gt;&lt;td&gt;486&lt;/td&gt;&lt;td&gt;36&lt;/td&gt;&lt;td&gt;2,321&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Met Expectations [4]&lt;/td&gt;&lt;td&gt;73&lt;/td&gt;&lt;td&gt;321&lt;/td&gt;&lt;td&gt;984&lt;/td&gt;&lt;td&gt;912&lt;/td&gt;&lt;td&gt;190&lt;/td&gt;&lt;td&gt;2,480&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Exceeded Expectations [5]&lt;/td&gt;&lt;td&gt;6&lt;/td&gt;&lt;td&gt;34&lt;/td&gt;&lt;td&gt;171&lt;/td&gt;&lt;td&gt;245&lt;/td&gt;&lt;td&gt;96&lt;/td&gt;&lt;td&gt;552&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Total&lt;/td&gt;&lt;td&gt;1,853&lt;/td&gt;&lt;td&gt;2,453&lt;/td&gt;&lt;td&gt;3,279&lt;/td&gt;&lt;td&gt;1,848&lt;/td&gt;&lt;td&gt;327&lt;/td&gt;&lt;td&gt;9,760&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Mathematics&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Did Not Yet Meet Expectations [1]&lt;/td&gt;&lt;td&gt;686&lt;/td&gt;&lt;td&gt;534&lt;/td&gt;&lt;td&gt;230&lt;/td&gt;&lt;td&gt;27&lt;/td&gt;&lt;td&gt;1&lt;/td&gt;&lt;td&gt;1,478&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Partially Met Expectations [2]&lt;/td&gt;&lt;td&gt;554&lt;/td&gt;&lt;td&gt;982&lt;/td&gt;&lt;td&gt;818&lt;/td&gt;&lt;td&gt;134&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;2,492&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Approached Expectations [3]&lt;/td&gt;&lt;td&gt;184&lt;/td&gt;&lt;td&gt;626&lt;/td&gt;&lt;td&gt;965&lt;/td&gt;&lt;td&gt;481&lt;/td&gt;&lt;td&gt;20&lt;/td&gt;&lt;td&gt;2,276&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Met Expectations [4]&lt;/td&gt;&lt;td&gt;59&lt;/td&gt;&lt;td&gt;205&lt;/td&gt;&lt;td&gt;762&lt;/td&gt;&lt;td&gt;853&lt;/td&gt;&lt;td&gt;184&lt;/td&gt;&lt;td&gt;2,063&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Exceeded Expectations [5]&lt;/td&gt;&lt;td&gt;1&lt;/td&gt;&lt;td&gt;3&lt;/td&gt;&lt;td&gt;73&lt;/td&gt;&lt;td&gt;157&lt;/td&gt;&lt;td&gt;99&lt;/td&gt;&lt;td&gt;333&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Total&lt;/td&gt;&lt;td&gt;1,484&lt;/td&gt;&lt;td&gt;2,350&lt;/td&gt;&lt;td&gt;2,848&lt;/td&gt;&lt;td&gt;1,652&lt;/td&gt;&lt;td&gt;308&lt;/td&gt;&lt;td&gt;8,642&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>1 <emph>Notes</emph>: Cells with gray shading are considered "consistent SLO classifications" and those without shading are considered to be "inconsistent."</p> <p>However, there is no particular reason to require or expect a one to one relationship between PARCC performance levels and SLO preparedness levels. An alternative approach is to regard the PARCC performance levels of 3 ("approached expectations") and 4 ("met expectations") as a key dividing line. That is, if a student fell in a PARCC performance level of 4 or 5, we consider a preparedness level of 4 or 5 to be a consistent classification for that student (cells shaded gray in Table ), and a preparedness level of 1, 2 or 3 to be an inconsistent classification (cells that are not shaded). For a student in PARCC performance level of 3 or lower, so long as the student's preparedness level is also 3 or lower, we consider it a consistent classification (cells shaded gray), whereas a preparedness level of 4 or 5 is considered an inconsistent classification (cells that are not shaded). In summary, using this dichotomous classification criterion, we would find that 76.2% of ELA and 79.5% of Math SLO preparedness classifications were consistent in 2015–2016, with these rates increasing to 80.6% and 84.5% in 2016–2017.</p> <p>A closer look at Table  suggests that at least some teachers may have had a tendency to underestimate the preparedness levels of their students. This is seen most clearly by looking at the 3,032 and 2,396 students who either met or exceeded expectations on the PARCC tests for ELA and Math, respectively (performance level 4 or 5). Only 48% and 54% of these students were classified consistently—in other words, about 52% and 46% of these students were classified as significantly underprepared, underprepared, or only somewhat prepared. In contrast, consider the 6,728 and 6,246 students who were classified as approaching expectations or lower on the PARCC tests for ELA and Math (performance level 1, 2, or 3). Only about 10%–11% of students in these groups were classified inconsistently in an upward direction as prepared or ahead for their SLO. This demonstrates an asymmetry in inconsistent preparedness classifications—teachers were more likely to underestimate a student's preparedness relative to PARCC performance than they were to overestimate it. The top half of Table  summarizes these different types of SLO preparedness classification rates by subject domain and year.</p> <p>Summary of Consistency With PARCC Performance for SLO Preparedness and End‐of‐Year Command Classifications</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th /&gt;&lt;th align="center"&gt;ELA&lt;/th&gt;&lt;th align="center"&gt;Math&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th /&gt;&lt;th align="center"&gt;2015&amp;#8211;2016&lt;/th&gt;&lt;th align="center"&gt;2016&amp;#8211;2017&lt;/th&gt;&lt;th align="center"&gt;2015&amp;#8211;2016&lt;/th&gt;&lt;th align="center"&gt;2016&amp;#8211;2017&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;SLO Year Preparedness&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Strict Consistency %&lt;/td&gt;&lt;td char="."&gt;39.9&lt;/td&gt;&lt;td char="."&gt;47.3&lt;/td&gt;&lt;td char="."&gt;41.5&lt;/td&gt;&lt;td char="."&gt;53.3&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Dichotomous Consistency %&lt;/td&gt;&lt;td char="."&gt;76.2&lt;/td&gt;&lt;td char="."&gt;80.6&lt;/td&gt;&lt;td char="."&gt;79.5&lt;/td&gt;&lt;td char="."&gt;84.5&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Underestimation %&lt;/td&gt;&lt;td char="."&gt;52.4&lt;/td&gt;&lt;td char="."&gt;39.1&lt;/td&gt;&lt;td char="."&gt;46.0&lt;/td&gt;&lt;td char="."&gt;35.8&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Overestimation %&lt;/td&gt;&lt;td char="."&gt;10.9&lt;/td&gt;&lt;td char="."&gt;8.8&lt;/td&gt;&lt;td char="."&gt;10.7&lt;/td&gt;&lt;td char="."&gt;7.0&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;SLO End&amp;#8208;of&amp;#8208;Year Command&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Strict Consistency %&lt;/td&gt;&lt;td char="."&gt;40.4&lt;/td&gt;&lt;td char="."&gt;39.6&lt;/td&gt;&lt;td char="."&gt;42.7&lt;/td&gt;&lt;td char="."&gt;39&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Dichotomous Consistency&lt;/td&gt;&lt;td char="."&gt;76.6&lt;/td&gt;&lt;td char="."&gt;78.7&lt;/td&gt;&lt;td char="."&gt;79.1&lt;/td&gt;&lt;td char="."&gt;80.5&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Underestimation %&lt;/td&gt;&lt;td char="."&gt;25.5&lt;/td&gt;&lt;td char="."&gt;20.7&lt;/td&gt;&lt;td char="."&gt;18.0&lt;/td&gt;&lt;td char="."&gt;11.8&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Overestimation %&lt;/td&gt;&lt;td char="."&gt;22.4&lt;/td&gt;&lt;td char="."&gt;21.7&lt;/td&gt;&lt;td char="."&gt;22.3&lt;/td&gt;&lt;td char="."&gt;23.2&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>It is important to appreciate that an inconsistent classification as defined here (i.e., when a student's SLO preparedness appears to be underestimated or overestimated), is not necessarily an <emph>inaccurate</emph> classification. It is entirely possible that a student who performed at a high level on the prior grade PARCC test might not show evidence of being "prepared" for an SLO at the beginning of the next school year. Explanations for inconsistent classifications could be plausibly rooted in differences between the content of the prior grade PARCC test and the current year SLO, or might be attributable to summer learning loss.</p> <hd id="AN0140159251-16">End‐of‐Year Command Classifications</hd> <p>Here we follow an approach that is somewhat parallel to the one above, only this time we compare SLO end‐of‐year command classifications of students (as of spring 2016 or spring 2017) to the performance of these students on the same subject PARCC test (also taken in the spring of 2016 or the spring of 2017). The bottom half of Table  summarizes these results. With respect to the overall consistency of end‐of‐year classifications, the results are almost identical to what was found for preparedness classifications. That is, taking a strict approach to consistent classification (agreement along the main diagonal), only about 40% and 43% of student were consistently classified for ELA and math SLOs respectively. Using a dichotomous classification approach, the values jump to 77% and 79%. These values stayed about the same in 2016–2017.</p> <p>If teachers have an incentive to inflate their growth scores by giving higher end‐of‐year command ratings than their students merit, a place to look for evidence of this would be the students who were in the bottom three performance levels of the PARCC ELA and Math tests. With both 2015–2016 and 2016–2017 data, we find that in both subject domains about 22% of students appear to have had their end‐of‐year rating overestimated. These are students who did not "meet expectations" on the PARCC tests taken in March but were classified as "strong" or "distinguished" on their SLO in May. This is primarily driven by students placed into the strong command categories—fewer than 1% of students in the bottom three PARCC performance levels were placed into the "distinguished" SLO category.</p> <hd id="AN0140159251-17">Is the Preparedness of Certain Kinds of Students Underestimated?</hd> <p>Next, we examine whether certain kinds of students were more likely than others to have their preparedness underestimated. We first do this by conducting ordered logit (i.e., ordinal) regressions (e.g., Long, [<reflink idref="bib19" id="ref20">19</reflink>]) with students as the units of analysis. The dependent variable in these regressions is the preparedness level of the student from 1 to 5. The independent variables consist of within‐grade standardized PARCC scale scores from the previous year and student‐specific demographic dummy variables that take on values of 1 if a student is female ("Female"), nonwhite ("Non‐White"), eligible for free or reduced price lunches ("FRL"), an English Language Learner ("ELL"), receiving special education services through an individualized education plan ("IEP"), or part of the gifted and talented ("GT") program, and a value of 0 otherwise. We run this regression four times, once for each SLO subject domain (math and ELA), and once for each academic school year (2015–2016 and 2017–2018). The respective student sample sizes for ELA and Math SLOs were 9,749 and 8,630 in 2015–2016; 10,463 and 12,017 in 2016–2017. Of primary interest is whether the demographic covariates are predictive of SLO preparedness levels, even after controlling for prior year PARCC scale scores.</p> <p>Columns 2 through 5 of Table  presents the results from these ordinal regressions, with coefficients expressed in an odds ratio metric. In both subjects across years, as one would expect, the odds of a student being in a higher preparedness category or categories relative to the lower category or categories increases dramatically for a 1 SD increase in prior PARCC scores. Here we see again that the coefficients for the odds ratios increase significantly from 2015–2016 to 2016–2017, consistent with the results shown earlier, and the availability of prior grade PARCC scores in 2016–2017, but not in 2015–2016. We also see that students with an IEP are much more likely to be in a lower preparedness category, whereas GT students have greater odds of being in a higher preparedness category. Neither of these results is surprising, because each variable represents sources of information teachers would be expected to consult when making a classification along with prior year test performance.</p> <p>Student‐level Ordered Logit Regressions (Odds Ratios)</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th /&gt;&lt;th align="center"&gt;&lt;italic&gt;Dependent Variable&lt;/italic&gt;:&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th /&gt;&lt;th align="center"&gt;Ordinal Categories 1&amp;#8211;5&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th /&gt;&lt;th align="center"&gt;SLO Preparedness&lt;/th&gt;&lt;th align="center"&gt;SLO End&amp;#8208;of&amp;#8208;Year Command&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th /&gt;&lt;th align="center"&gt;2015&amp;#8211;2016&lt;/th&gt;&lt;th align="center"&gt;2016&amp;#8211;2017&lt;/th&gt;&lt;th align="center"&gt;2015&amp;#8211;2016&lt;/th&gt;&lt;th align="center"&gt;2016&amp;#8211;2017&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th /&gt;&lt;th align="center"&gt;ELA&lt;/th&gt;&lt;th align="center"&gt;Math&lt;/th&gt;&lt;th align="center"&gt;ELA&lt;/th&gt;&lt;th align="center"&gt;Math&lt;/th&gt;&lt;th align="center"&gt;ELA&lt;/th&gt;&lt;th align="center"&gt;Math&lt;/th&gt;&lt;th align="center"&gt;ELA&lt;/th&gt;&lt;th align="center"&gt;Math&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;PARCC.SS&lt;/td&gt;&lt;td char="."&gt;3.320001&lt;/td&gt;&lt;td char="."&gt;3.250001&lt;/td&gt;&lt;td char="."&gt;5.440001&lt;/td&gt;&lt;td char="."&gt;6.890001&lt;/td&gt;&lt;td char="."&gt;4.780001&lt;/td&gt;&lt;td char="."&gt;6.200001&lt;/td&gt;&lt;td char="."&gt;3.600001&lt;/td&gt;&lt;td char="."&gt;4.520001&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Female&lt;/td&gt;&lt;td char="."&gt;1.07&lt;/td&gt;&lt;td char="."&gt;1.04&lt;/td&gt;&lt;td char="."&gt;1.00&lt;/td&gt;&lt;td char="."&gt;0.94&lt;/td&gt;&lt;td char="."&gt;1.130001&lt;/td&gt;&lt;td char="."&gt;1.110001&lt;/td&gt;&lt;td char="."&gt;1.170001&lt;/td&gt;&lt;td char="."&gt;1.150001&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Non&amp;#8208;White&lt;/td&gt;&lt;td char="."&gt;0.92&lt;/td&gt;&lt;td char="."&gt;0.780001&lt;/td&gt;&lt;td char="."&gt;0.99&lt;/td&gt;&lt;td char="."&gt;0.890001&lt;/td&gt;&lt;td char="."&gt;0.96&lt;/td&gt;&lt;td char="."&gt;0.96&lt;/td&gt;&lt;td char="."&gt;0.900001&lt;/td&gt;&lt;td char="."&gt;0.900001&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;FRL&lt;/td&gt;&lt;td char="."&gt;0.870001&lt;/td&gt;&lt;td char="."&gt;0.760001&lt;/td&gt;&lt;td char="."&gt;0.840001&lt;/td&gt;&lt;td char="."&gt;0.830001&lt;/td&gt;&lt;td char="."&gt;0.93&lt;/td&gt;&lt;td char="."&gt;1.03&lt;/td&gt;&lt;td char="."&gt;0.87&lt;/td&gt;&lt;td char="."&gt;0.80&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;ELL&lt;/td&gt;&lt;td char="."&gt;1.01&lt;/td&gt;&lt;td char="."&gt;0.860001&lt;/td&gt;&lt;td char="."&gt;1.110001&lt;/td&gt;&lt;td char="."&gt;0.920001&lt;/td&gt;&lt;td char="."&gt;0.910001&lt;/td&gt;&lt;td char="."&gt;1.05&lt;/td&gt;&lt;td char="."&gt;1.090001&lt;/td&gt;&lt;td char="."&gt;1.080001&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;IEP&lt;/td&gt;&lt;td char="."&gt;0.330001&lt;/td&gt;&lt;td char="."&gt;0.400001&lt;/td&gt;&lt;td char="."&gt;0.340001&lt;/td&gt;&lt;td char="."&gt;0.420001&lt;/td&gt;&lt;td char="."&gt;0.430001&lt;/td&gt;&lt;td char="."&gt;0.530001&lt;/td&gt;&lt;td char="."&gt;0.790001&lt;/td&gt;&lt;td char="."&gt;0.720001&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;GT&lt;/td&gt;&lt;td char="."&gt;1.300001&lt;/td&gt;&lt;td char="."&gt;1.520001&lt;/td&gt;&lt;td char="."&gt;1.07&lt;/td&gt;&lt;td char="."&gt;0.94&lt;/td&gt;&lt;td char="."&gt;1.410001&lt;/td&gt;&lt;td char="."&gt;1.350001&lt;/td&gt;&lt;td char="."&gt;1.580001&lt;/td&gt;&lt;td char="."&gt;1.440001&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Students&lt;/td&gt;&lt;td align="center"&gt;9,749&lt;/td&gt;&lt;td align="center"&gt;8,630&lt;/td&gt;&lt;td align="center"&gt;10,463&lt;/td&gt;&lt;td align="center"&gt;12,017&lt;/td&gt;&lt;td align="center"&gt;12,146&lt;/td&gt;&lt;td align="center"&gt;9,802&lt;/td&gt;&lt;td align="center"&gt;13,303&lt;/td&gt;&lt;td align="center"&gt;15,956&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>2 <emph>Notes</emph>:*<emph>p</emph> &lt; 0.1; ** <emph>p</emph> &lt; 0.05; *** <emph>p</emph> &lt; 0.01. The results in each column are instances of an ordered logit regression model with four thresholds identified by fixing the constant at 0. Threshold estimates are not included here but are available upon request.</p> <p>Of concern is the significant finding across subjects and years that poorer students (those eligible for free and reduced price lunches), have a slightly higher odds of being in lower preparedness categories. Although the magnitude of this association is small, the fact that it is statistically significant even after controlling for prior PARCC performance is worrisome, because this represents information that teachers should not be using to judge SLO preparedness. Similarly, there is some evidence that the odds of landing in a higher preparedness category is related to a student's status as an English language learner and for Non‐White students on Math SLOs (though again, the magnitudes tend to be small, and for these variables the direction and statistical significance are not always consistent).</p> <p>The ordinal regression results presented in Table  make a proportional odds or parallel regression assumption. That is, it is assumed that the odds of shifting from one category to the next are the same across all categories. A closer inspection of this assumption suggests that it may not hold for the covariates Non‐White and FRL for the two highest SLO preparedness categories. To dig into this more deeply, we restricted our focus to the subset of students who scored in performance levels 4 or 5 on the PARCC ELA and Math tests. We then specify logistic regressions where a value of 1 for the outcome variable indicates that a student was inconsistently classified in a manner that may have <emph>underestimated</emph> the student's level of preparedness, and a value of 0 indicates that the student was consistently classified. This logistic regression includes the same set of demographic variables as covariates that were in the ordinal regression described above, but does not include prior year PARCC scale scores since prior year PARCC performance levels are being used directly to both restrict the sample and define the outcome variable. In this approach, we again look for evidence that students with certain demographic characteristics have their preparedness underestimated, but we do so by focusing on a subset of students (those who met expectations on PARCC the previous spring), and a specific preparedness threshold (the three lowest preparedness levels vs. the two highest).</p> <p>In taking this approach, we find some evidence of a larger magnitude of underestimation for students who are FRL‐eligible and/or non‐White. Using 2016–2017 SLO data for ELA as an example, a non‐White male who is eligible for free and reduced lunch, a native English speaker, with no IEP and not identified as Gifted and Talented had a probability of 51.7% of being classified as Somewhat Underprepared, Underprepared, or Significantly Underprepared—even though the student had scored in one of the top two PARCC ELA performance levels about 4–5 months earlier. In contrast, if the student differed only by being White and not FRL‐eligible, the probability of a low preparedness classification would decrease by almost 17 points to 35%. The marginal change in probability is similarly large for math SLOs, where the race/ethnicity and FRL indicators increase the probability of underestimation about 13 points (from 28.5% to 41.6%). These marginal changes in probability within SLO subject domain were fairly similar for other reference students with different combinations of the demographic covariates.</p> <hd id="AN0140159251-18">Is the End‐of‐Year Command of Certain Kinds of Students Overestimated?</hd> <p>We performed a parallel set of analyses with the end‐of‐year SLO command classifications as the outcome variable of interest. Columns 6 through 9 of Table  present the results from the ordinal regressions, with coefficients expressed in an odds ratio metric. Once again, we similar relationships between PARCC scores, IEP, and GT indicators, and classification in higher end‐of‐year command category as was found for preparedness classifications. However, we see no association with race, poverty, or English learner status. Instead, females tend to be slightly more likely than males to be in a higher end‐of‐year command category.</p> <hd id="AN0140159251-19">Relationship of Teacher SLO Growth Points to Other Teacher‐Level Variables</hd> <p>In this section, we shift from examining the validity of student‐level preparedness and end‐of‐year command classifications to the validity and reliability of using SLO.Growth scores as a measure of teacher effectiveness. We do so by adopting a framework that has been used in the past to evaluate the properties of value‐added models in the context of teachers who teach students for whom state‐administered standardized test scores are available across adjacent years (e.g., Ehlert, Koedel, Parsons, &amp; Podgursky, [<reflink idref="bib9" id="ref21">9</reflink>]). When averaged over students and attached to teachers, an SLO measure can be cast as a crude version of a value‐added model in that it attempts to distinguish teachers on the basis of differences in student achievement conditional on students' starting points designated during the fall. In a typical value‐added model for teachers in tested subjects, preparedness and mastery would be established objectively on the basis of prior and current grade standardized tests; on SLOs, both preparedness and mastery are determined with some degree of subjectivity by teachers. Because value‐added models take preexisting differences in student achievement (and often other variables) into account, in theory at least, they provide a fairer basis for comparing teachers relative to comparisons based solely on students' end‐of‐year achievement. With this in mind, we can examine to what extent a teacher‐level growth measure based on SLOs appears to "level the playing field" in a manner analogous to a value‐added model when contrasted to the use of a teacher‐level status measure based on SLOs.</p> <p>We examine teachers with ELA and Math SLOs separately in this analysis, and also impose a restriction that each teacher must have at least 15 students. This reduces our teacher samples shown in Table  by about 29%–35% in 2015–2016, and 18%–25% in 2016–2017. In what follows, we contrast two different teacher‐level SLO outcome variables. The first, "SLO.Status," is the average of a students' end‐of‐year SLO mastery classifications (the average of ordinal values on a scale from 1 to 5). The second, "SLO.Growth," represents the percent of possible growth points earned by a teacher (on a scale from 0% to 100%). Table  shows correlations between these two variables as well as between other available and relevant teacher‐level variables for two different teachers samples in 2015–2016: the 920 teachers that submitted ELA SLOs (upper right triangle above the main diagonal) and the 652 that submitted Math SLOs (lower left triangle below the main diagonal). The main diagonal provides an estimate of the SD of each variable, computed as the average of the ELA and Math teacher samples. A variable of particular interest is each teacher's professional practice rating ("PP.Pts"). This score represents a combination of classroom observation ratings, a professionalism rating, and student perception survey ratings. The remaining covariates are all created by computing teacher‐level means for the students associated with each teacher. These include the mean subject‐specific PARCC scale scores at the teacher level for two different years ("PARCC.15" and "PARCC.16"), the mean of current year student growth percentiles ("MGP.16") and all the aggregate demographic variables used in the student‐level logistic regressions previously presented. Because the PARCC tests are not on the same scale across grades, all scores were standardized across the full population of students in the district by grade and subject for a given year. We only present this correlation matrix for 2015–2016 data; the results for 2016–2017 data were very similar (within about 0.05 on a scale from 0 to 1).</p> <p>Relationships Among Teacher‐Level Variables (2015–2016)</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr valign="bottom"&gt;&lt;th /&gt;&lt;th align="center"&gt;SLO. Status&lt;/th&gt;&lt;th align="center"&gt;SLO. Growth&lt;/th&gt;&lt;th align="center"&gt;PP.Pts&lt;/th&gt;&lt;th align="center"&gt;MSS.15&lt;/th&gt;&lt;th align="center"&gt;MSS.16&lt;/th&gt;&lt;th align="center"&gt;MGP.16&lt;/th&gt;&lt;th align="center"&gt;%Fem&lt;/th&gt;&lt;th align="center"&gt;%Non&amp;#8208;White&lt;/th&gt;&lt;th align="center"&gt;%FRL&lt;/th&gt;&lt;th align="center"&gt;%ELL&lt;/th&gt;&lt;th align="center"&gt;%IEP&lt;/th&gt;&lt;th align="center"&gt;%GT&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;SLO.Status&lt;/td&gt;&lt;td&gt;(0.60)&lt;/td&gt;&lt;td&gt;0.49&lt;/td&gt;&lt;td&gt;0.29&lt;/td&gt;&lt;td&gt;0.54&lt;/td&gt;&lt;td&gt;0.56&lt;/td&gt;&lt;td&gt;0.22&lt;/td&gt;&lt;td&gt;&amp;#8722;0.01&lt;/td&gt;&lt;td&gt;&amp;#8722;0.48&lt;/td&gt;&lt;td&gt;&amp;#8722;0.49&lt;/td&gt;&lt;td&gt;&amp;#8722;0.45&lt;/td&gt;&lt;td&gt;&amp;#8722;0.23&lt;/td&gt;&lt;td&gt;0.09&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;SLO.Growth&lt;/td&gt;&lt;td&gt;0.46&lt;/td&gt;&lt;td&gt;(0.13)&lt;/td&gt;&lt;td&gt;0.21&lt;/td&gt;&lt;td&gt;0.06&lt;/td&gt;&lt;td&gt;0.11&lt;/td&gt;&lt;td&gt;0.13&lt;/td&gt;&lt;td&gt;0.04&lt;/td&gt;&lt;td&gt;&amp;#8722;0.04&lt;/td&gt;&lt;td&gt;&amp;#8722;0.04&lt;/td&gt;&lt;td&gt;&amp;#8722;0.11&lt;/td&gt;&lt;td&gt;&amp;#8722;0.08&lt;/td&gt;&lt;td&gt;&amp;#8722;0.07&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;PP.Pts&lt;/td&gt;&lt;td&gt;0.33&lt;/td&gt;&lt;td&gt;0.25&lt;/td&gt;&lt;td&gt;(4.49)&lt;/td&gt;&lt;td&gt;0.24&lt;/td&gt;&lt;td&gt;0.30&lt;/td&gt;&lt;td&gt;0.26&lt;/td&gt;&lt;td&gt;0.06&lt;/td&gt;&lt;td&gt;&amp;#8722;0.28&lt;/td&gt;&lt;td&gt;&amp;#8722;0.30&lt;/td&gt;&lt;td&gt;&amp;#8722;0.20&lt;/td&gt;&lt;td&gt;&amp;#8722;0.15&lt;/td&gt;&lt;td&gt;0.14&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;PARCC.15&lt;/td&gt;&lt;td&gt;0.64&lt;/td&gt;&lt;td&gt;0.07&lt;/td&gt;&lt;td&gt;0.28&lt;/td&gt;&lt;td&gt;(0.62)&lt;/td&gt;&lt;td&gt;0.93&lt;/td&gt;&lt;td&gt;0.31&lt;/td&gt;&lt;td&gt;0.14&lt;/td&gt;&lt;td&gt;&amp;#8722;0.81&lt;/td&gt;&lt;td&gt;&amp;#8722;0.85&lt;/td&gt;&lt;td&gt;&amp;#8722;0.66&lt;/td&gt;&lt;td&gt;&amp;#8722;0.37&lt;/td&gt;&lt;td&gt;0.72&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;PARCC.16&lt;/td&gt;&lt;td&gt;0.63&lt;/td&gt;&lt;td&gt;0.08&lt;/td&gt;&lt;td&gt;0.36&lt;/td&gt;&lt;td&gt;0.90&lt;/td&gt;&lt;td&gt;(0.59)&lt;/td&gt;&lt;td&gt;0.62&lt;/td&gt;&lt;td&gt;0.12&lt;/td&gt;&lt;td&gt;&amp;#8722;0.80&lt;/td&gt;&lt;td&gt;&amp;#8722;0.83&lt;/td&gt;&lt;td&gt;&amp;#8722;0.60&lt;/td&gt;&lt;td&gt;&amp;#8722;0.25&lt;/td&gt;&lt;td&gt;0.61&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;MGP.16&lt;/td&gt;&lt;td&gt;0.29&lt;/td&gt;&lt;td&gt;0.12&lt;/td&gt;&lt;td&gt;0.34&lt;/td&gt;&lt;td&gt;0.27&lt;/td&gt;&lt;td&gt;0.60&lt;/td&gt;&lt;td&gt;(13.26)&lt;/td&gt;&lt;td&gt;0.10&lt;/td&gt;&lt;td&gt;&amp;#8722;0.31&lt;/td&gt;&lt;td&gt;&amp;#8722;0.30&lt;/td&gt;&lt;td&gt;&amp;#8722;0.22&lt;/td&gt;&lt;td&gt;&amp;#8722;0.11&lt;/td&gt;&lt;td&gt;0.29&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;%Fem&lt;/td&gt;&lt;td&gt;0.00&lt;/td&gt;&lt;td&gt;0.02&lt;/td&gt;&lt;td&gt;0.04&lt;/td&gt;&lt;td&gt;0.02&lt;/td&gt;&lt;td&gt;0.01&lt;/td&gt;&lt;td&gt;&amp;#8722;0.03&lt;/td&gt;&lt;td&gt;(0.10)&lt;/td&gt;&lt;td&gt;&amp;#8722;0.01&lt;/td&gt;&lt;td&gt;&amp;#8722;0.02&lt;/td&gt;&lt;td&gt;&amp;#8722;0.01&lt;/td&gt;&lt;td&gt;&amp;#8722;0.14&lt;/td&gt;&lt;td&gt;0.12&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;%Non&amp;#8208;White&lt;/td&gt;&lt;td&gt;&amp;#8722;0.53&lt;/td&gt;&lt;td&gt;&amp;#8722;0.04&lt;/td&gt;&lt;td&gt;&amp;#8722;0.31&lt;/td&gt;&lt;td&gt;&amp;#8722;0.79&lt;/td&gt;&lt;td&gt;&amp;#8722;0.81&lt;/td&gt;&lt;td&gt;&amp;#8722;0.30&lt;/td&gt;&lt;td&gt;0.03&lt;/td&gt;&lt;td&gt;(0.30)&lt;/td&gt;&lt;td&gt;0.94&lt;/td&gt;&lt;td&gt;0.65&lt;/td&gt;&lt;td&gt;0.18&lt;/td&gt;&lt;td&gt;&amp;#8722;0.31&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;%FRL&lt;/td&gt;&lt;td&gt;&amp;#8722;0.54&lt;/td&gt;&lt;td&gt;&amp;#8722;0.04&lt;/td&gt;&lt;td&gt;&amp;#8722;0.30&lt;/td&gt;&lt;td&gt;&amp;#8722;0.81&lt;/td&gt;&lt;td&gt;&amp;#8722;0.84&lt;/td&gt;&lt;td&gt;&amp;#8722;0.33&lt;/td&gt;&lt;td&gt;0.01&lt;/td&gt;&lt;td&gt;0.95&lt;/td&gt;&lt;td&gt;(0.33)&lt;/td&gt;&lt;td&gt;0.65&lt;/td&gt;&lt;td&gt;0.20&lt;/td&gt;&lt;td&gt;&amp;#8722;0.35&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;%ELL&lt;/td&gt;&lt;td&gt;&amp;#8722;0.46&lt;/td&gt;&lt;td&gt;&amp;#8722;0.10&lt;/td&gt;&lt;td&gt;&amp;#8722;0.25&lt;/td&gt;&lt;td&gt;&amp;#8722;0.58&lt;/td&gt;&lt;td&gt;&amp;#8722;0.61&lt;/td&gt;&lt;td&gt;&amp;#8722;0.27&lt;/td&gt;&lt;td&gt;&amp;#8722;0.02&lt;/td&gt;&lt;td&gt;0.68&lt;/td&gt;&lt;td&gt;0.67&lt;/td&gt;&lt;td&gt;(0.33)&lt;/td&gt;&lt;td&gt;0.07&lt;/td&gt;&lt;td&gt;&amp;#8722;0.12&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;%IEP&lt;/td&gt;&lt;td&gt;&amp;#8722;0.17&lt;/td&gt;&lt;td&gt;0.04&lt;/td&gt;&lt;td&gt;&amp;#8722;0.05&lt;/td&gt;&lt;td&gt;&amp;#8722;0.36&lt;/td&gt;&lt;td&gt;&amp;#8722;0.24&lt;/td&gt;&lt;td&gt;&amp;#8722;0.14&lt;/td&gt;&lt;td&gt;&amp;#8722;0.08&lt;/td&gt;&lt;td&gt;0.16&lt;/td&gt;&lt;td&gt;0.19&lt;/td&gt;&lt;td&gt;0.08&lt;/td&gt;&lt;td&gt;(0.08)&lt;/td&gt;&lt;td&gt;&amp;#8722;0.17&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;%GT&lt;/td&gt;&lt;td&gt;0.14&lt;/td&gt;&lt;td&gt;&amp;#8722;0.07&lt;/td&gt;&lt;td&gt;0.06&lt;/td&gt;&lt;td&gt;0.72&lt;/td&gt;&lt;td&gt;0.63&lt;/td&gt;&lt;td&gt;0.27&lt;/td&gt;&lt;td&gt;0.07&lt;/td&gt;&lt;td&gt;&amp;#8722;0.27&lt;/td&gt;&lt;td&gt;&amp;#8722;0.32&lt;/td&gt;&lt;td&gt;&amp;#8722;0.11&lt;/td&gt;&lt;td&gt;&amp;#8722;0.17&lt;/td&gt;&lt;td&gt;(0.17)&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>3 <emph>Note</emph>: Correlations for ELA teacher sample in the upper triangle, correlations for math teacher sample in the lower triangle. Correlations were computed using pairwise complete observations. The values in parentheses along the main diagonal represent the average standard deviation of each variable for the two samples.</p> <p>Note in Table  that SLO.Status tends to be more strongly correlated with aggregated demographic and achievement variables than SLO.Growth. In particular, while the teacher‐level mean of prior year PARCC scale scores in ELA and Math have correlations of about 0.54 and 0.64 with the teacher‐level mean of SLO.Status, the two variables are essentially uncorrelated with SLO.Growth.</p> <p>We explore these relationships more formally by regressing (with teachers as the unit of analysis) aggregate ELA and Math SLO outcome variables on a set of variables that capture aggregate characteristics of the students in each teacher's class, the grade level being taught, and teacher scores for professional practice:[<reflink idref="bib6" id="ref22">6</reflink>]</p> <p> <ephtml> &lt;math display="block" altimg="urn:x-wiley:00220655:media:jedm12233:jedm12233-math-0001" xmlns="http://www.w3.org/1998/Math/MathML"&gt;Tch.SLO.Outcome=&amp;#946;0+&amp;#946;1PP.Pts+&amp;#946;2PARCC.lagged+&amp;#946;3%Female+&amp;#946;4%FRL+&amp;#946;5%ELL+&amp;#946;6%IEP+&amp;#946;7%GT+&amp;#946;8Middle+&amp;#946;9High+&amp;#946;10OtherSchool+&amp;#949;.&lt;/math&gt; </ephtml> </p> <p>In the model above, <emph>Tch.SLO.Outcome</emph> is a teacher's average of student's end‐of‐year SLO mastery classification (SLO.Status) or a teacher's percent of SLO points earned (SLO.Growth). The variable <emph>PP.Pts</emph> represents the professional practice points a teacher earned in the school year in consideration, and <emph>PARCC.lagged</emph> represents the mean standardized student PARCC scores from either spring 2015 (for 2015–2016 SLO results) or spring 2016 (for 2016–2017 SLO results). The variables <emph>Middle</emph>, <emph>High</emph>, and <emph>Other School</emph>[<reflink idref="bib7" id="ref23">7</reflink>] capture school level differences with Elementary as the omitted reference category. Finally, %<emph>Female</emph>, %<emph>FRL</emph>, %<emph>ELL</emph>, %<emph>IEP</emph>, and %<emph>GT</emph> are all teacher‐level variables that were created by aggregating student demographic information for a particular teacher using just the students a teacher tracked on SLOs. We run separate regressions for each year, subject domain, and SLO outcome. To facilitate comparisons across models, we standardize all regression variables with the exception of our school level indicators for each subject domain and year combination. The estimated coefficients can be found in Table  and are interpretable as the number of SDs the dependent variable of interest would be predicted to increase for a 1 SD increase in the independent variable of interest, holding other covariates constant.</p> <p>Regression Models for Teachers' SLO Outcomes</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th /&gt;&lt;th align="center"&gt;ELA 2015&amp;#8211;2016&lt;/th&gt;&lt;th align="center"&gt;Math 2015&amp;#8211;2016&lt;/th&gt;&lt;th align="center"&gt;ELA 2016&amp;#8211;2017&lt;/th&gt;&lt;th align="center"&gt;Math 2016&amp;#8211;2017&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th /&gt;&lt;th align="center"&gt;SLO.&lt;/th&gt;&lt;th align="center"&gt;SLO.&lt;/th&gt;&lt;th align="center"&gt;SLO.&lt;/th&gt;&lt;th align="center"&gt;SLO.&lt;/th&gt;&lt;th align="center"&gt;SLO.&lt;/th&gt;&lt;th align="center"&gt;SLO.&lt;/th&gt;&lt;th align="center"&gt;SLO.&lt;/th&gt;&lt;th align="center"&gt;SLO.&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th /&gt;&lt;th align="center"&gt;Status&lt;/th&gt;&lt;th align="center"&gt;Growth&lt;/th&gt;&lt;th align="center"&gt;Status&lt;/th&gt;&lt;th align="center"&gt;Growth&lt;/th&gt;&lt;th align="center"&gt;Status&lt;/th&gt;&lt;th align="center"&gt;Growth&lt;/th&gt;&lt;th align="center"&gt;Status&lt;/th&gt;&lt;th align="center"&gt;Growth&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;PP.Pts&lt;/td&gt;&lt;td&gt;0.150001&lt;/td&gt;&lt;td&gt;0.200001&lt;/td&gt;&lt;td&gt;0.150001&lt;/td&gt;&lt;td&gt;0.200001&lt;/td&gt;&lt;td&gt;0.140001&lt;/td&gt;&lt;td&gt;0.330001&lt;/td&gt;&lt;td&gt;0.140001&lt;/td&gt;&lt;td&gt;0.270001&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;PARCC.lagged&lt;/td&gt;&lt;td&gt;0.480001&lt;/td&gt;&lt;td&gt;0.17&lt;/td&gt;&lt;td&gt;0.350001&lt;/td&gt;&lt;td&gt;0.03&lt;/td&gt;&lt;td&gt;0.280001&lt;/td&gt;&lt;td&gt;&amp;#8722;0.05&lt;/td&gt;&lt;td&gt;0.420001&lt;/td&gt;&lt;td&gt;0.07&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;%Fem&lt;/td&gt;&lt;td&gt;&amp;#8722;0.03&lt;/td&gt;&lt;td&gt;0.02&lt;/td&gt;&lt;td&gt;0.03&lt;/td&gt;&lt;td&gt;&amp;#8722;0.05&lt;/td&gt;&lt;td&gt;&amp;#8722;0.05&lt;/td&gt;&lt;td&gt;0.03&lt;/td&gt;&lt;td&gt;0.02&lt;/td&gt;&lt;td&gt;0.08&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;%FRL&lt;/td&gt;&lt;td&gt;&amp;#8722;0.04&lt;/td&gt;&lt;td&gt;0.10&lt;/td&gt;&lt;td&gt;&amp;#8722;0.16&lt;/td&gt;&lt;td&gt;0.11&lt;/td&gt;&lt;td&gt;&amp;#8722;0.210001&lt;/td&gt;&lt;td&gt;&amp;#8722;0.002&lt;/td&gt;&lt;td&gt;&amp;#8722;0.210001&lt;/td&gt;&lt;td&gt;&amp;#8722;0.05&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;%ELL&lt;/td&gt;&lt;td&gt;&amp;#8722;0.03&lt;/td&gt;&lt;td&gt;0.02&lt;/td&gt;&lt;td&gt;0.04&lt;/td&gt;&lt;td&gt;0.15&lt;/td&gt;&lt;td&gt;0.03&lt;/td&gt;&lt;td&gt;&amp;#8722;0.005&lt;/td&gt;&lt;td&gt;&amp;#8722;0.003&lt;/td&gt;&lt;td&gt;0.10&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;%IEP&lt;/td&gt;&lt;td&gt;&amp;#8722;0.08&lt;/td&gt;&lt;td&gt;0.01&lt;/td&gt;&lt;td&gt;&amp;#8722;0.04&lt;/td&gt;&lt;td&gt;0.002&lt;/td&gt;&lt;td&gt;&amp;#8722;0.03&lt;/td&gt;&lt;td&gt;0.01&lt;/td&gt;&lt;td&gt;&amp;#8722;0.03&lt;/td&gt;&lt;td&gt;0.003&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;%GT&lt;/td&gt;&lt;td&gt;&amp;#8722;0.05&lt;/td&gt;&lt;td&gt;&amp;#8722;0.08&lt;/td&gt;&lt;td&gt;0.150001&lt;/td&gt;&lt;td&gt;0.200001&lt;/td&gt;&lt;td&gt;0.230001&lt;/td&gt;&lt;td&gt;&amp;#8722;0.001&lt;/td&gt;&lt;td&gt;0.130001&lt;/td&gt;&lt;td&gt;0.01&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Middle&lt;/td&gt;&lt;td&gt;0.04&lt;/td&gt;&lt;td&gt;&amp;#8722;0.23&lt;/td&gt;&lt;td&gt;&amp;#8722;0.19&lt;/td&gt;&lt;td&gt;&amp;#8722;0.06&lt;/td&gt;&lt;td&gt;&amp;#8722;0.310001&lt;/td&gt;&lt;td&gt;&amp;#8722;0.14&lt;/td&gt;&lt;td&gt;&amp;#8722;0.370001&lt;/td&gt;&lt;td&gt;0.08&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;&amp;#8722;0.370001&lt;/td&gt;&lt;td&gt;&amp;#8722;0.06&lt;/td&gt;&lt;td&gt;&amp;#8722;0.380001&lt;/td&gt;&lt;td&gt;&amp;#8722;0.330001&lt;/td&gt;&lt;td&gt;&amp;#8722;0.540001&lt;/td&gt;&lt;td&gt;&amp;#8722;0.14&lt;/td&gt;&lt;td&gt;&amp;#8722;0.520001&lt;/td&gt;&lt;td&gt;0.12&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Other School&lt;/td&gt;&lt;td&gt;&amp;#8722;0.350001&lt;/td&gt;&lt;td&gt;&amp;#8722;0.16&lt;/td&gt;&lt;td&gt;&amp;#8722;0.04&lt;/td&gt;&lt;td&gt;0.02&lt;/td&gt;&lt;td&gt;&amp;#8722;0.17&lt;/td&gt;&lt;td&gt;&amp;#8722;0.430001&lt;/td&gt;&lt;td&gt;&amp;#8722;0.05&lt;/td&gt;&lt;td&gt;&amp;#8722;0.16&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Constant&lt;/td&gt;&lt;td&gt;0.11&lt;/td&gt;&lt;td&gt;0.07&lt;/td&gt;&lt;td&gt;0.140001&lt;/td&gt;&lt;td&gt;0.09&lt;/td&gt;&lt;td&gt;0.170001&lt;/td&gt;&lt;td&gt;0.11&lt;/td&gt;&lt;td&gt;0.140001&lt;/td&gt;&lt;td&gt;&amp;#8722;0.004&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Observations&lt;/td&gt;&lt;td align="center"&gt;343&lt;/td&gt;&lt;td align="center"&gt;343&lt;/td&gt;&lt;td align="center"&gt;301&lt;/td&gt;&lt;td align="center"&gt;301&lt;/td&gt;&lt;td align="center"&gt;380&lt;/td&gt;&lt;td align="center"&gt;380&lt;/td&gt;&lt;td align="center"&gt;434&lt;/td&gt;&lt;td align="center"&gt;434&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;R&lt;sup&gt;2&lt;/sup&gt;&lt;/td&gt;&lt;td&gt;0.35&lt;/td&gt;&lt;td&gt;0.05&lt;/td&gt;&lt;td&gt;0.46&lt;/td&gt;&lt;td&gt;0.10&lt;/td&gt;&lt;td&gt;0.47&lt;/td&gt;&lt;td&gt;0.12&lt;/td&gt;&lt;td&gt;0.59&lt;/td&gt;&lt;td&gt;0.10&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>4 <emph>Note</emph>: *<emph>p</emph> &lt; 0.05, **<emph>p</emph> &lt; 0.01, and ***<emph>p</emph> &lt; 0.001.</p> <p>Two patterns in these results stand out. First, covariates such as %GT, %IEP, %FRL, and school‐level indicators, which frequently have a significant partial association with SLO.Status for a given subject domain and year, seldom have a significant partial association with SLO.Growth. Second, the only covariate that retains a significant association with both SLO.Status and SLO.Growth outcomes is a teacher's professional practice score (PP.Pts). In fact, though the differences tend to be relatively small, the regression coefficient for the PP.Pts covariate is always larger in magnitude when regressed on the SLO.Growth outcome relative to the SLO.Status outcome. For example, for 2015–2016 data, a 1 SD increase in a teacher's professional practice score is associated with about a 0.15 SD increase on SLO.Status scores. In contrast, a 1 SD increase is associated with a about a 0.20 SD increase on SLO.Growth scores.</p> <hd id="AN0140159251-20">Comparisons with Mean Student Growth Percentiles</hd> <p>In principle, both an SGP and SLO are supposed to convey something about student growth that, when aggregated to a teacher, can be used to discern that some teachers have been more effective than others in their academic instruction. An important question is whether a growth indicator based on SLOs (i.e., SLO.Growth) is an adequate substitute for a measure based on SGPs (i.e., MGP). To address this question, we examine the slightly smaller subset of teachers from the regression analysis above for whom MGP scores are also available. Figure  displays the scatterplots of SLO.Growth and MGP by subject domain and year. Although both SLO.Growth and MGP are intended to be teacher‐level indicators of student growth, with correlations ranging between 0.12 and 0.28, they tend to not be much more strongly correlated with one another than they are with professional practice ratings, even though the latter are intended to capture a different dimension of teacher effectiveness.[<reflink idref="bib8" id="ref24">8</reflink>]</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/MEA/01dec19/jedm12233-fig-0006.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jedm12233-fig-0006.jpg" title="Relationship between SLO points and mean of student growth percentiles. [Color figure can be viewed at wileyonlinelibrary.com]" /> </p> <p></p> <p>Table  replicates the format of Table , but adds a column in which we replace the SLO‐based aggregate outcome (SLO.Status or SLO.Growth) with an SGP‐based aggregate outcome (i.e., a teacher's MGP). There are again two interesting and interpretable findings.[<reflink idref="bib9" id="ref25">9</reflink>] First, the <emph>R</emph><sups>2</sups> for regressions with MGP as the outcome are always higher than the <emph>R</emph><sups>2</sups> for regressions with SLO.Growth as the outcome. More specifically, the covariate %GT maintains a significant partial association with MGP for each subject domain and year contribution; a similar association is not evident for the SLO.Growth outcome. Second, the professional practice score has about the same partial association with the MGP outcome (ranging from a low of 0.21 to a high of 0.30) as it does for the SLO.Growth outcome (ranging from a low of 0.20 to a high of 0.34).</p> <p>Regression Models for Teachers with MGPs</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th /&gt;&lt;th align="center"&gt;ELA 2015&amp;#8208;16&lt;/th&gt;&lt;th align="center"&gt;Math 2015&amp;#8208;16&lt;/th&gt;&lt;th align="center"&gt;ELA 2016&amp;#8208;17&lt;/th&gt;&lt;th align="center"&gt;Math 2016&amp;#8208;17&lt;/th&gt;&lt;/tr&gt;&lt;tr valign="bottom"&gt;&lt;th /&gt;&lt;th align="center"&gt;SLO. Status&lt;/th&gt;&lt;th align="center"&gt;SLO. Growth&lt;/th&gt;&lt;th align="center"&gt;MGP&lt;/th&gt;&lt;th align="center"&gt;SLO. Status&lt;/th&gt;&lt;th align="center"&gt;SLO. Growth&lt;/th&gt;&lt;th align="center"&gt;MGP&lt;/th&gt;&lt;th align="center"&gt;SLO. Status&lt;/th&gt;&lt;th align="center"&gt;SLO. Growth&lt;/th&gt;&lt;th align="center"&gt;MGP&lt;/th&gt;&lt;th align="center"&gt;SLO. Status&lt;/th&gt;&lt;th align="center"&gt;SLO. Growth&lt;/th&gt;&lt;th align="center"&gt;MGP&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;PP.Pts&lt;/td&gt;&lt;td&gt;0.100001&lt;/td&gt;&lt;td&gt;0.200001&lt;/td&gt;&lt;td&gt;0.210001&lt;/td&gt;&lt;td&gt;0.110001&lt;/td&gt;&lt;td&gt;0.210001&lt;/td&gt;&lt;td&gt;0.290001&lt;/td&gt;&lt;td&gt;0.130001&lt;/td&gt;&lt;td&gt;0.340001&lt;/td&gt;&lt;td&gt;0.300001&lt;/td&gt;&lt;td&gt;0.140001&lt;/td&gt;&lt;td&gt;0.270001&lt;/td&gt;&lt;td&gt;0.290001&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;PARCC.lagged&lt;/td&gt;&lt;td&gt;0.380001&lt;/td&gt;&lt;td&gt;0.16&lt;/td&gt;&lt;td&gt;0.01&lt;/td&gt;&lt;td&gt;0.410001&lt;/td&gt;&lt;td&gt;0.06&lt;/td&gt;&lt;td&gt;&amp;#8722;0.310001&lt;/td&gt;&lt;td&gt;0.290001&lt;/td&gt;&lt;td&gt;&amp;#8722;0.01&lt;/td&gt;&lt;td&gt;0.03&lt;/td&gt;&lt;td&gt;0.450001&lt;/td&gt;&lt;td&gt;0.05&lt;/td&gt;&lt;td&gt;&amp;#8722;0.02&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;%Fem&lt;/td&gt;&lt;td&gt;&amp;#8722;0.02&lt;/td&gt;&lt;td&gt;0.01&lt;/td&gt;&lt;td&gt;0.08&lt;/td&gt;&lt;td&gt;0.02&lt;/td&gt;&lt;td&gt;&amp;#8722;0.07&lt;/td&gt;&lt;td&gt;&amp;#8722;0.03&lt;/td&gt;&lt;td&gt;&amp;#8722;0.03&lt;/td&gt;&lt;td&gt;0.04&lt;/td&gt;&lt;td&gt;&amp;#8722;0.05&lt;/td&gt;&lt;td&gt;0.01&lt;/td&gt;&lt;td&gt;0.07&lt;/td&gt;&lt;td&gt;&amp;#8722;0.04&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;%FRL&lt;/td&gt;&lt;td&gt;&amp;#8722;0.08&lt;/td&gt;&lt;td&gt;0.14&lt;/td&gt;&lt;td&gt;&amp;#8722;0.11&lt;/td&gt;&lt;td&gt;&amp;#8722;0.17&lt;/td&gt;&lt;td&gt;0.12&lt;/td&gt;&lt;td&gt;&amp;#8722;0.380001&lt;/td&gt;&lt;td&gt;&amp;#8722;0.220001&lt;/td&gt;&lt;td&gt;0.05&lt;/td&gt;&lt;td&gt;&amp;#8722;0.09&lt;/td&gt;&lt;td&gt;&amp;#8722;0.170001&lt;/td&gt;&lt;td&gt;&amp;#8722;0.06&lt;/td&gt;&lt;td&gt;&amp;#8722;0.14&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;%ELL&lt;/td&gt;&lt;td&gt;&amp;#8722;0.10&lt;/td&gt;&lt;td&gt;&amp;#8722;0.02&lt;/td&gt;&lt;td&gt;&amp;#8722;0.05&lt;/td&gt;&lt;td&gt;0.01&lt;/td&gt;&lt;td&gt;0.16&lt;/td&gt;&lt;td&gt;0.02&lt;/td&gt;&lt;td&gt;0.05&lt;/td&gt;&lt;td&gt;&amp;#8722;0.03&lt;/td&gt;&lt;td&gt;0.01&lt;/td&gt;&lt;td&gt;&amp;#8722;0.02&lt;/td&gt;&lt;td&gt;0.10&lt;/td&gt;&lt;td&gt;0.08&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;%IEP&lt;/td&gt;&lt;td&gt;&amp;#8722;0.09&lt;/td&gt;&lt;td&gt;0.02&lt;/td&gt;&lt;td&gt;0.02&lt;/td&gt;&lt;td&gt;&amp;#8722;0.03&lt;/td&gt;&lt;td&gt;&amp;#8722;0.01&lt;/td&gt;&lt;td&gt;&amp;#8722;0.120001&lt;/td&gt;&lt;td&gt;&amp;#8722;0.03&lt;/td&gt;&lt;td&gt;0.01&lt;/td&gt;&lt;td&gt;&amp;#8722;0.07&lt;/td&gt;&lt;td&gt;&amp;#8722;0.03&lt;/td&gt;&lt;td&gt;&amp;#8722;0.02&lt;/td&gt;&lt;td&gt;0.03&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;%GT&lt;/td&gt;&lt;td&gt;0.06&lt;/td&gt;&lt;td&gt;&amp;#8722;0.03&lt;/td&gt;&lt;td&gt;0.190001&lt;/td&gt;&lt;td&gt;0.09&lt;/td&gt;&lt;td&gt;0.17&lt;/td&gt;&lt;td&gt;0.220001&lt;/td&gt;&lt;td&gt;0.240001&lt;/td&gt;&lt;td&gt;0.01&lt;/td&gt;&lt;td&gt;0.180001&lt;/td&gt;&lt;td&gt;0.130001&lt;/td&gt;&lt;td&gt;0.01&lt;/td&gt;&lt;td&gt;0.160001&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Middle&lt;/td&gt;&lt;td&gt;0.06&lt;/td&gt;&lt;td&gt;&amp;#8722;0.22&lt;/td&gt;&lt;td&gt;&amp;#8722;0.29&lt;/td&gt;&lt;td&gt;&amp;#8722;0.15&lt;/td&gt;&lt;td&gt;&amp;#8722;0.03&lt;/td&gt;&lt;td&gt;0.04&lt;/td&gt;&lt;td&gt;&amp;#8722;0.310001&lt;/td&gt;&lt;td&gt;&amp;#8722;0.14&lt;/td&gt;&lt;td&gt;&amp;#8722;0.01&lt;/td&gt;&lt;td&gt;&amp;#8722;0.370001&lt;/td&gt;&lt;td&gt;0.07&lt;/td&gt;&lt;td&gt;&amp;#8722;0.320001&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;&amp;#8722;0.610001&lt;/td&gt;&lt;td&gt;0.02&lt;/td&gt;&lt;td&gt;0.04&lt;/td&gt;&lt;td&gt;&amp;#8722;0.17&lt;/td&gt;&lt;td&gt;0.002&lt;/td&gt;&lt;td&gt;0.17&lt;/td&gt;&lt;td&gt;&amp;#8722;0.570001&lt;/td&gt;&lt;td&gt;&amp;#8722;0.500001&lt;/td&gt;&lt;td&gt;0.10&lt;/td&gt;&lt;td&gt;&amp;#8722;0.420001&lt;/td&gt;&lt;td&gt;0.13&lt;/td&gt;&lt;td&gt;&amp;#8722;0.18&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Other School&lt;/td&gt;&lt;td&gt;&amp;#8722;0.330001&lt;/td&gt;&lt;td&gt;&amp;#8722;0.15&lt;/td&gt;&lt;td&gt;&amp;#8722;0.08&lt;/td&gt;&lt;td&gt;&amp;#8722;0.01&lt;/td&gt;&lt;td&gt;0.02&lt;/td&gt;&lt;td&gt;0.13&lt;/td&gt;&lt;td&gt;&amp;#8722;0.17&lt;/td&gt;&lt;td&gt;&amp;#8722;0.420001&lt;/td&gt;&lt;td&gt;0.09&lt;/td&gt;&lt;td&gt;&amp;#8722;0.05&lt;/td&gt;&lt;td&gt;&amp;#8722;0.16&lt;/td&gt;&lt;td&gt;&amp;#8722;0.22&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Constant&lt;/td&gt;&lt;td&gt;0.09&lt;/td&gt;&lt;td&gt;0.06&lt;/td&gt;&lt;td&gt;0.06&lt;/td&gt;&lt;td&gt;0.05&lt;/td&gt;&lt;td&gt;0.004&lt;/td&gt;&lt;td&gt;&amp;#8722;0.05&lt;/td&gt;&lt;td&gt;0.140001&lt;/td&gt;&lt;td&gt;0.14&lt;/td&gt;&lt;td&gt;&amp;#8722;0.02&lt;/td&gt;&lt;td&gt;0.110001&lt;/td&gt;&lt;td&gt;0.004&lt;/td&gt;&lt;td&gt;0.11&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Observations&lt;/td&gt;&lt;td align="center"&gt;298&lt;/td&gt;&lt;td align="center"&gt;298&lt;/td&gt;&lt;td align="center"&gt;298&lt;/td&gt;&lt;td align="center"&gt;249&lt;/td&gt;&lt;td align="center"&gt;249&lt;/td&gt;&lt;td align="center"&gt;249&lt;/td&gt;&lt;td align="center"&gt;347&lt;/td&gt;&lt;td align="center"&gt;347&lt;/td&gt;&lt;td align="center"&gt;347&lt;/td&gt;&lt;td align="center"&gt;407&lt;/td&gt;&lt;td align="center"&gt;407&lt;/td&gt;&lt;td align="center"&gt;407&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;R&lt;sup&gt;2&lt;/sup&gt;&lt;/td&gt;&lt;td&gt;0.40&lt;/td&gt;&lt;td&gt;0.05&lt;/td&gt;&lt;td&gt;0.17&lt;/td&gt;&lt;td&gt;0.48&lt;/td&gt;&lt;td&gt;0.08&lt;/td&gt;&lt;td&gt;0.23&lt;/td&gt;&lt;td&gt;0.50&lt;/td&gt;&lt;td&gt;0.14&lt;/td&gt;&lt;td&gt;0.22&lt;/td&gt;&lt;td&gt;0.61&lt;/td&gt;&lt;td&gt;0.10&lt;/td&gt;&lt;td&gt;0.15&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>5 <emph>Note</emph>: *<emph>p</emph> &lt; 0.05, **<emph>p</emph> &lt; 0.01, and ***<emph>p</emph> &lt; 0.001.</p> <p>If one criterion for the validity of a classroom growth indicator for use in teacher evaluation is that it should "level the playing" field such that teachers are not rewarded or penalized for characteristics of students that are outside of their control, then an argument could be made based on these results that the SLO.Growth indicator fulfills this criterion as well or better than the MGP indicator. On the other hand, one might argue that for a growth indicator to be valid that it <emph>should</emph> retain at least a small positive correlation with a variable such as %GT, under the premise that students flagged with gifted and talented status have typically demonstrated more growth each year than students not flagged with this status. However, while the results are open to slightly different interpretations, the bottom line is the regression relationships tend to be remarkably similar whether one uses SLO.Growth or MGP as the teacher‐level dependent variable. This suggests the need for caution when using this as the basis for a validity argument, and we return to discuss the implications of these comparisons in the final section of the article.</p> <hd id="AN0140159251-22">Stability</hd> <p>Beyond concerns about their validity as measures of teacher effectiveness, a key drawback to any growth‐based statistic is that measures of growth tend to be more volatile than measures of status. In the literature on value‐added models, the year‐to‐year correlation of teacher growth statistics has been found to be weak to moderate, ranging from about 0.2 to 0.6 (Goldhaber &amp; Hansen, [<reflink idref="bib10" id="ref26">10</reflink>]; McCaffrey, Sass, Lockwood, &amp; Mihaly, [<reflink idref="bib20" id="ref27">20</reflink>]). Kane and Staiger ([<reflink idref="bib15" id="ref28">15</reflink>], [<reflink idref="bib16" id="ref29">16</reflink>]) and McCaffrey et al. ([<reflink idref="bib20" id="ref30">20</reflink>]) have argued that under certain assumptions, such intertemporal correlations can be interpreted as an estimate of stability, in which case any intertemporal correlation less than 0.5 would imply that more than half of the variability in value‐added can be explained by chance factors unrelated to characteristics of a teacher that persist over time.</p> <p>The scatterplots in Figure  depict the relative stability of teacher‐level SLO and SGP growth statistics when focusing on the subset of teachers for whom both statistics were available across two years of data. The intertemporal correlations tend to be small to moderate, ranging from about 0.30 to 0.47. The stability of growth based on SLOs is lower than the stability of growth based on SGPs. For ELA, the stability is just slightly higher for MGPs (<emph>r</emph> = 0.38 instead of <emph>r</emph> = 0.32), but for Math it is considerably higher (<emph>r</emph> = 0.47 instead of <emph>r</emph> = 0.30). Nonetheless, the lower correlation for SLOs is still well within the range of what has been typically observed for most growth‐based statistics (McCaffrey et al., [<reflink idref="bib20" id="ref31">20</reflink>]).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/MEA/01dec19/jedm12233-fig-0007.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jedm12233-fig-0007.jpg" title="Intertemporal stability of teacher‐level growth: SLO versus SGP. [Color figure can be viewed at wileyonlinelibrary.com]" /> </p> <p></p> <p>These results need to be interpreted with some caution because, to the extent that each of these teacher‐level statistics is biased, the correlations may be overstated to an unknown degree. For example, MGPs make no attempt to disentangle influences on student achievement beyond those captured by a student's individual prior test achievement profile. Assume that a student's classroom peers have a distinct effect on student achievement. If some teachers tend to always work with students in classrooms with positive peer effects, and others tend to work with students in classrooms with negative peer effects, then the former teachers will always be more likely to have an MGP that is above average in any given year, and the latter teachers will always be more likely to have an MGP that is below average in any given year. This bias could manifest itself through inflated intertemporal correlations. The potential for bias is just as great (if not greater) for measures of growth based on SLOs. In this context, one might imagine two different sources of bias, one due to omitted variables associated with both classroom contexts and student achievement, and another due to the fact that teachers are the ones classifying the preparedness and end‐of‐year command of their own students. If certain teachers are more likely to underestimate the preparedness of their students every year and others are more likely to overestimate it, this could also inflate the observed stability of SLO growth indicators.</p> <hd id="AN0140159251-24">Discussion</hd> <p>In the way that they are implemented in Denver Public Schools, SLOs are premised on a dual purpose use for classroom assessment. That is, teachers are expected to use classroom assessments that they choose, administer, and score so that they can track and facilitate student learning over the course of the year. At the same time, the result of this process also includes numeric ratings that get aggregated into a teacher‐level growth indicator, which in turn figures prominently in teachers' annual performance evaluations. The <emph>Standards for Educational and Psychological Testing</emph> offer a general warning when it comes to the use of educational assessments to meet multiple purposes: "Most educational tests will serve one purpose better than others; and the more purposes an educational test is purported to serve, the less likely it is to serve any of those purposes effectively" (American Educational Research Association, American Psychological Association, &amp; National Council on Measurement in Education, [<reflink idref="bib1" id="ref32">1</reflink>], p. 206)</p> <p>In the present study, we do not formally evaluate the claim that SLOs are a successful vehicle for formative assessment. However, the results from focus group interviews and surveys that we have conducted (but do not present here) do not paint an encouraging picture in this regard (Diaz‐Bilello, Chattergoon, &amp; Briggs, [<reflink idref="bib7" id="ref33">7</reflink>]). For example, in one survey administered in early 2016, only about 60% of a sample of 1,712 DPS teachers agreed or strongly agreed with the prompt "SLOs provide me valuable information about what my students know and can do." Although these findings have somewhat limited generalizability because of self‐selection among survey (and focus group) respondents and our inability to link these respondents to the data considered in this article, they suggest at a minimum that many DPS teachers have not bought into the dual purpose use of SLOs. Instead, they tend to view SLOs almost exclusively as something required of them to meet district accountability demands.</p> <p>What we have looked for in this study is evidence, through available data over a 2‐year time period, that student and teacher‐level growth statistics based on the SLO performances of students are being distorted as one might predict from Campbell's Law. We have also compared the use of an SLO‐based growth statistic for teacher evaluation to the use of an SGP‐based growth statistic. When using PARCC test performance as an external criterion, we find evidence to suggest that teachers may have a tendency to underestimate some students' SLO preparedness and/or overestimate some students' end‐of‐year level of SLO command, two practices that would lead to an inflation of SLO growth points. The underestimation of preparedness appears to be a bigger concern than the overestimation of end‐of‐year command. Using 2015–2016 data restricted to the population of students that were in one of PARCC's two highest performance levels in the Spring of 2015, we found that roughly 50% of these students had SLO preparedness classifications that seemed too low. In contrast, when using the same data and restricting to the population of students that scored in one of PARCC's three lowest performance levels near the end of the school year, we found that only about 10%–11% of students had SLO end‐of‐year command classifications that seemed too high.</p> <p>The general finding that many students had preparedness classifications that were lower than what was suggested by their PARCC performance has at least two possible explanations. One explanation is that many teachers had the (mistaken) impression that SLO preparedness and end‐of‐year command exist on the same dimension. A teacher might look past the labels used for the preparedness and command categories and instead focus just on the numbers. From this perspective, it might be hard to imagine that a student at the beginning of the year could be at level 4 on a scale that runs from 1 to 5. Another explanation, more consistent with Campbell's Law, would be that at least some teachers, aware that student growth would count toward their LEAP evaluation ratings, responded to this incentive by placing students into lower preparedness levels to maximize their growth ratings. These two explanations are not mutually exclusive.</p> <p>A more troubling finding is that certain characteristics of students—in particular gender, race/ethnicity, and poverty status—are often predictive of having preparedness underestimated or end‐of‐year command overestimated. Although the magnitude of this underestimation or overestimation tended to be small overall, it raises concerns about equity. Holding constant prior academic achievement in the form of PARCC performance, students in poverty are still significantly more likely to have their preparedness underestimated relative to students not in poverty.</p> <p>SLO growth points are typically aggregated to create a teacher‐level growth measure. For teachers that have both SLO‐based growth measures and SGP‐based growth measures, the correlation between the two is weak. Somewhat surprisingly, with respect to its relationship to other teacher‐level variables and with respect to its intertemporal stability, SLO.Growth is difficult to distinguish from an MGP. Both SLO.Growth and MGPs have the same low (but statistically significant) correlation with professional practice ratings, and this correlation (not corrected for measurement error) is of about the same magnitude as that found between value‐added estimates and teacher professional practice ratings in at least one other prominent empirical context (Kane, McCaffrey, Miller, &amp; Staiger, [<reflink idref="bib14" id="ref34">14</reflink>]; Kane &amp; Staiger, [<reflink idref="bib16" id="ref35">16</reflink>]).</p> <p>If an advantage of measures of growth over status is that they are less correlated (or uncorrelated) with aggregate measures of student status, then one could argue that SLO.Growth displays this advantage as well or better than MGPs. Both SLO.Growth and MGPs have a weak to moderate degree of intertemporal stability. In general, the finding of low intertemporal stability for teacher or school‐level measures of growth has been a key to the argument that such measures should be employed with great caution for high‐stakes purposes (Morganstein &amp; Wasserstein, [<reflink idref="bib21" id="ref36">21</reflink>]). This argument certainly applies in this context as well, and it was a greater concern for SLO.Growth than it was for MGP, especially in Math.</p> <p>There are limitations to the generalizability of findings in this study. In our analyses, we focus on the nonrandom sample of DPS teachers who submitted district‐created SLOs in Math and ELA, and who also teach students for whom large‐scale assessment results were available. Recall that one of the biggest motivations for SLOs is that they provide growth evidence for teachers who teach in subjects for which state‐administered large‐scale assessments are <emph>not</emph> available. If the SLOs for these teachers tend to be of lesser quality or involve more distorted practices than those analyzed in this study, our findings may well present an overly optimistic evaluation. Another limitation is that we cannot claim that the use of SLOs for high stakes has <emph>caused</emph> a distortion in SLO scores. It is entirely possible that the evidence we have found with respect to under‐ and overestimation would be evident even in the absence of the use of SLOs for accountability. It would be interesting to replicate our approach in a comparable school district or an entire state that uses SLOs, but only for formative purposes.</p> <p>To what extent do the analyses in the second half of this article suggest that SLO.Growth can be validly used as not just an alternative to MGPs for accountability decisions, but as a replacement? Although some of our findings might imply support to this notion, we would caution readers against it. It is true that SLO.Growth looks very similar to MGPs when it comes to relationships with certain teacher‐level covariates and in terms of intertemporal stability, but they clearly do not convey the same information about student growth. Even when focusing on student achievement results in the same subject domain, the observed correlation between SLO.Growth and MGPs is weak—the highest was <emph>r</emph> = 0.13 in ELA and <emph>r</emph> = 0.29 in Math. It is important to appreciate that SLO.Growth represents an attempt at a criterion‐referenced growth measure, whereas an MGP is purely norm‐referenced. Although one might expect to see that when students in a classroom all score higher on a test than peers across the district with similar prior year scores, that these students will have also shown growth in a criterion‐referenced sense, this need not be the case. Conversely, if all students in a classroom are showing the same evidence of criterion‐referenced growth on their SLOs, a teacher's MGP might still be close to the district average if it is generally the case that students across the district are all showing similar amounts of criterion‐referenced growth. The results from this study suggest that when both SLO.Growth and an MGP are available they should not be regarded as interchangeable. Whether it is defensible to use SLO.Growth as an alternative growth indicator when an MGP is not available is still an open question given the restrictions we placed on our analytic samples.</p> <p>Ultimately, the validity of SLOs and the indicators that derive from them will depend to a great extent upon the validity of the assessments that underlie SLO preparedness and end‐of‐year command levels. What constructs do these assessments measure? What is the underlying theory of student cognition? What design principles were used to write tasks and items? When performance tasks are used, to what extent are the scores generalizable? And so on. There is considerable evidence that can be brought to the table for PARCC (and other large‐scale assessments) in this regard. The validity of the many different assessments that factor into SLO classifications is in this sense both a black box and an open question. In other words, although we would not argue that the assessments currently used in places like DPS for purposes of SLO classifications are necessarily invalid (indeed, the open‐ended tasks shown in Figure  have some very promising features), very little systematic or formal evidence has been gathered and made publicly available in this regard. As such validity evidence is critical, the limited capacity of school district staff to collect and present such evidence represents a major impediment to the long‐term dual purpose use of SLOs that has been envisioned.</p> <ref id="AN0140159251-25"> <title> Footnotes </title> <blist> <bibl id="bib1" idref="ref2" type="bt">1</bibl> <bibtext> The popularity of computer‐based "interim" assessments in many school districts across the country suggests that many educators believe that these products are formatively useful, even though the assessments are standardized and externally developed. Yet because these sorts of assessments have generally not been used as inputs for high‐stakes accountability decisions, their "dual use" potential is an open question.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref13" type="bt">2</bibl> <bibtext> https://<ulink href="http://www.cde.state.co.us/educatoreffectiveness/senatebill10191rulesdocument">www.cde.state.co.us/educatoreffectiveness/senatebill10191rulesdocument</ulink></bibtext> </blist> <blist> <bibl id="bib3" idref="ref8" type="bt">3</bibl> <bibtext> <ulink href="http://careers.dpsk12.org/wp-content/uploads/2016/09/2017-LEAP-Teacher-Handbook-lo-res.pdf">http://careers.dpsk12.org/wp-content/uploads/2016/09/2017-LEAP-Teacher-Handbook-lo-res.pdf</ulink> </bibtext> </blist> <blist> <bibl id="bib4" idref="ref16" type="bt">4</bibl> <bibtext> These were the descriptors used during the 2015–2016 school year to align with the PARCC performance descriptors. For the 2016–2017 school year, these were changed to stay consistent with the change made to PARCC performance descriptors in the same year: (a) Did Not Yet Meet Expectations, (b) Partially Met Expectations, (c) Approached Expectations, (d) Met Expectations, and (e) Exceeded Expectations.</bibtext> </blist> <blist> <bibl id="bib5" idref="ref9" type="bt">5</bibl> <bibtext> The teachers in each subject domain were generally distinct in 2015–2016 because in this year teachers were only required to submit a single SLO, though some did submit two. In 2015–2016, there were a total of 1,974 unique teachers and 313 who submitted an SLO for both ELA and math. In 2016–2017 (when teachers were required to submit two SLOs), there were a total of 2,025 unique teachers and 735 who submitted in both subject domains. The overlap across subject domain samples for students was 9,067 and 19,030 for 2015–2016 and 2016–2017, respectively. Importantly, however, for both years when conditioning on subject domain, sample sizes represent unique teachers and students. In a relatively small number of cases, the same student was present for the SLO of more than one teacher within the math and ELA subject domains. In these instances, to ensure that the same student was only present once in our analyses, we randomly assigned the student to a single teacher.</bibtext> </blist> <blist> <bibl id="bib6" idref="ref18" type="bt">6</bibl> <bibtext> As shown in Table 5, once aggregated to the teacher level, the variables representing the percentage of non‐White students (%Non‐White) and percentage of students eligible for free and reduced price lunches (%FRL) have a correlation of about 0.94. To avoid problems with a severe form of multicollinearity, we only retain %FRL as a covariate in our regression models.</bibtext> </blist> <blist> <bibl id="bib7" idref="ref23" type="bt">7</bibl> <bibtext> "Other School" represents schools with the atypical grade configurations of K‐8, 6‐12, or K‐12.</bibtext> </blist> <blist> <bibl id="bib8" idref="ref3" type="bt">8</bibl> <bibtext> In both subject domains, the correlation between SLO.Growth and MGP increased from 2015–2016 to 2016–2017. Again, this finding may well be attributable to the impact of having prior year PARCC scores available when making preparedness ratings in 2016–2017, but not having them available in 2015–2016.</bibtext> </blist> <blist> <bibl id="bib9" idref="ref21" type="bt">9</bibl> <bibtext> A puzzling finding is the statistically significant regression coefficients of −0.38 and −0.31 for the covariates %FRL and PARCC.lagged when math MGP is the outcome using 2015–2016 data. After ruling out coding errors as an explanation and performing a variety of regression diagnostics, our conclusion is that this difference relative to the results in the other three models is most likely being caused by some interaction between sample size (this year and subject domain had the smallest teacher sample) and multicollinearity. That is, all four of the regressions shown in Table 7 have the same issues with multicollinearity among the independent variables induced by the aggregation of student‐level variables. But only in this one case does it manifest itself with a shift in the magnitude and sign of a regression coefficient that does not cohere with theory or findings from previous studies.</bibtext> </blist> </ref> <ref id="AN0140159251-26"> <title> References </title> <blist> <bibtext> American Educational Research Association, American Psychological Association, &amp; National Council on Measurement in Education. (2014). Standards for educational and psychological testing. Washington, DC : American Educational Research Association.</bibtext> </blist> <blist> <bibtext> Betebenner, D. (2009). Norm‐ and criterion‐referenced student growth. Educational Measurement: Issues and Practice, 28 (4), 42 – 51.</bibtext> </blist> <blist> <bibtext> Briggs, D. C. (2016). Can Campbell's Law be mitigated? In H. Braun (Ed.), Meeting the challenges to measurement in an era of accountability (pp. 168 – 179). New York, NY : Routledge.</bibtext> </blist> <blist> <bibtext> Briggs, D. C., Diaz‐Bilello, E., Peck, F., Alzen, J., Chattergoon, R., &amp; McClelland, A. (2014). Learning progression project: Documentation of pilot work and lessons in the 2013–2014 school year. Boulder, CO : Center for Assessment, Design, Research and Evaluation (CADRE) and National Center for the Improvement of Educational Assessment.</bibtext> </blist> <blist> <bibtext> Campbell, D. T. (1976). Assessing the impact of planned social change. Hanover, NH : Public Affairs Center, Dartmouth College.</bibtext> </blist> <blist> <bibtext> Castellano, K. E., &amp; Ho, A. D. (2013). A practitioner's guide to growth models. Washington, DC : Council of Chief State School Officers.</bibtext> </blist> <blist> <bibtext> Diaz‐Bilello, E., Chattergoon, R., &amp; Briggs, D. (2016). District‐wide perceptions about the student learning objectives process in 2015–2016. Brief commissioned by the Denver Public Schools. Boulder, CO : Center for Assessment Design Research and Evaluation (CADRE).</bibtext> </blist> <blist> <bibtext> Doherty, K. M., &amp; Jacobs, S. (2013). State of the states 2013: Connect the dots: Using evaluations of teacher effectiveness to inform policy and practice. Washington, DC : National Council on Teacher Quality. Retrieved from <ulink href="http://www.nctq.org/dmsView/State%5fof%5fthe%5fStates%5f2013%5fUsing%5fTeacher%5fEvaluations%5fNCTQ%5fReport">http://www.nctq.org/dmsView/State%5fof%5fthe%5fStates%5f2013%5fUsing%5fTeacher%5fEvaluations%5fNCTQ%5fReport</ulink></bibtext> </blist> <blist> <bibtext> Ehlert, M., Koedel, C., Parsons, E., &amp; Podgursky, M. J. (2014). The sensitivity of value‐added estimates to specification adjustments: Evidence from school‐ and teacher‐level models in missouri. Statistics and Public Policy, 1 (1), 19 – 27. https://doi.org/10.1080/2330443X.2013.856152</bibtext> </blist> <blist> <bibtext> Goldhaber, D., &amp; Hansen, M. (2009). National board certification and teachers' career paths: Does NBPTS certification influence how long teachers remain in the profession and where they teach? Education Finance and Policy, 4, 229 – 262.</bibtext> </blist> <blist> <bibtext> Gonring, P., Teske, P., &amp; Jupp, B. (2007). Pay‐for‐performance teacher compensation: An inside view of Denver's ProComp Plan. Cambridge, MA : Harvard Education Press.</bibtext> </blist> <blist> <bibtext> Hall, E., Gagnon, D., Schneider, M. C., Marion, S., &amp; Thompson, J. (2014). State practices related to the use of student achievement measures in the evaluation of teachers in non‐tested subjects and grades. Unpublished manuscript. Retrieved from <ulink href="http://www.nciea.org/publication%5fPDFs/Gates%20NTGS%5fHall%20082614.pdf">http://www.nciea.org/publication%5fPDFs/Gates%20NTGS%5fHall%20082614.pdf</ulink></bibtext> </blist> <blist> <bibtext> Hershberg, T., &amp; Robertson‐Kraft, C. (2012). A grand bargain for education reform: New rewards and supports for new accountability. Cambridge, MA : Harvard Education Press.</bibtext> </blist> <blist> <bibtext> Kane, T. J., McCaffrey, D. F., Miller, T., &amp; Staiger, D. O. (2013). Have we identified effective teachers? Validating measures of effective teaching using random assignment. Research Paper, MET Project. Seattle, WA : Bill &amp; Melinda Gates Foundation.</bibtext> </blist> <blist> <bibtext> Kane, T. J., &amp; Staiger, D. O. (2002). Volatility in school test scores: Implications for test‐based accountability systems. Brookings Papers on Education Policy. Washington, DC : Brookings Institution Press.</bibtext> </blist> <blist> <bibtext> Kane, T. J., &amp; Staiger, D. O. (2012). Gathering feedback for teaching: Combining high‐quality observations with student surveys and achievement gains. Research Paper, MET Project. Seattle, WA : Bill &amp; Melinda Gates Foundation.</bibtext> </blist> <blist> <bibtext> Lachlan‐Haché, L., Cushing, E., &amp; Bivona, L. (2012). Student learning objectives as measures of educator effectiveness: The basics. Washington, DC : American Institutes for Research.</bibtext> </blist> <blist> <bibtext> Lacireno‐Paquet, N., Morgan, C., &amp; Mello, D. (2014). How states use student learning objectives in teacher evaluation systems: A review of state websites (REL 2014–013). Washington, DC : U.S. Department of Education, Institute of Education Sciences, National Center for Education Evaluation and Regional Assistance, Regional Educational Laboratory Northeast &amp; Islands. Retrieved from <ulink href="http://ies.ed.gov/ncee/edlabs">http://ies.ed.gov/ncee/edlabs</ulink></bibtext> </blist> <blist> <bibtext> Long, J. S. (1997). Regression models for categorical and limited dependent variables. Thousand Oaks, CA : Sage.</bibtext> </blist> <blist> <bibtext> McCaffrey, D. F., Sass, T. R., Lockwood, J. R., &amp; Mihaly, K. (2009). The intertemporal variability of teacher effect estimates. Education Finance and Policy, 4, 572 – 606.</bibtext> </blist> <blist> <bibtext> Morganstein, D., &amp; Wasserstein, R. (2014). ASA statement on value‐added models. Statistics and Public Policy, 1 (1), 108 – 110. https://doi.org/10.1080/2330443X.2014.956906</bibtext> </blist> <blist> <bibtext> Ruiz‐Primo, M. A., Shavelson, R. J., Hamilton, L., &amp; Klein, S. (2002). On the evaluation of systemic science education reform: Searching for instructional sensitivity. Journal of Research in Science Teaching, 39, 369 – 393.</bibtext> </blist> </ref> <aug> <p>By Derek C. Briggs; Rajendra Chattergoon and Amy Burkhardt</p> <p>Reported by Author; Author; Author</p> <p></p> <p>DEREK C. BRIGGS is Professor at the School of Education, University of Colorado Boulder, and Director of the Center for Assessment, Research and Evaluation;. His research interests focus on issues of measurement and psychometrics in research on growth and student learning.</p> <p>RAJENDRA CHATTERGOON is PhD student in the Research and Evaluation Methodology program at the University of Colorado Boulder. His research interests include issues related to measurement, assessment, and learning progressions.</p> <p>AMY BURKHARDT is a PhD student in the Research and Evaluation Methodology program at the University of Colorado Boulder;. Her research interests focus on issues related to measurement and automated scoring.</p> </aug> <nolink nlid="nl1" bibid="bib22" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib12" firstref="ref4"></nolink> <nolink nlid="nl3" bibid="bib18" firstref="ref5"></nolink> <nolink nlid="nl4" bibid="bib17" firstref="ref7"></nolink> <nolink nlid="nl5" bibid="bib11" firstref="ref11"></nolink> <nolink nlid="nl6" bibid="bib13" firstref="ref12"></nolink> <nolink nlid="nl7" bibid="bib19" firstref="ref20"></nolink> <nolink nlid="nl8" bibid="bib10" firstref="ref26"></nolink> <nolink nlid="nl9" bibid="bib20" firstref="ref27"></nolink> <nolink nlid="nl10" bibid="bib15" firstref="ref28"></nolink> <nolink nlid="nl11" bibid="bib16" firstref="ref29"></nolink> <nolink nlid="nl12" bibid="bib14" firstref="ref34"></nolink> <nolink nlid="nl13" bibid="bib21" firstref="ref36"></nolink> |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1236254 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Examining the Dual Purpose Use of Student Learning Objectives for Classroom Assessment and Teacher Evaluation – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Briggs%2C+Derek+C%2E%22">Briggs, Derek C.</searchLink><br /><searchLink fieldCode="AR" term="%22Chattergoon%2C+Rajendra%22">Chattergoon, Rajendra</searchLink><br /><searchLink fieldCode="AR" term="%22Burkhardt%2C+Amy%22">Burkhardt, Amy</searchLink> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Journal+of+Educational+Measurement%22"><i>Journal of Educational Measurement</i></searchLink>. Win 2019 56(4):686-714. – Name: Avail Label: Availability Group: Avail Data: Wiley-Blackwell. 350 Main Street, Malden, MA 02148. Tel: 800-835-6770; Tel: 781-388-8598; Fax: 781-388-8232; e-mail: cs-journals@wiley.com; Web site: http://www.wiley.com/WileyCDA – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 29 – Name: DatePubCY Label: Publication Date Group: Date Data: 2019 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Evaluative – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Student+Educational+Objectives%22">Student Educational Objectives</searchLink><br /><searchLink fieldCode="DE" term="%22Student+Evaluation%22">Student Evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Teacher+Evaluation%22">Teacher Evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Scores%22">Scores</searchLink><br /><searchLink fieldCode="DE" term="%22Urban+Schools%22">Urban Schools</searchLink><br /><searchLink fieldCode="DE" term="%22School+District+Size%22">School District Size</searchLink><br /><searchLink fieldCode="DE" term="%22Growth+Models%22">Growth Models</searchLink><br /><searchLink fieldCode="DE" term="%22Measurement%22">Measurement</searchLink><br /><searchLink fieldCode="DE" term="%22High+Stakes+Tests%22">High Stakes Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Inferences%22">Inferences</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1111/jedm.12233 – Name: ISSN Label: ISSN Group: ISSN Data: 0022-0655 – Name: Abstract Label: Abstract Group: Ab Data: The process of setting and evaluating student learning objectives (SLOs) has become increasingly popular as an example where classroom assessment is intended to fulfill the dual purpose use of informing instruction and holding teachers accountable. A concern is that the high-stakes purpose may lead to distortions in the inferences about students and teachers that SLOs can support. This concern is explored in the present study by contrasting student SLO scores in a large urban school district to performance on a common objective external criterion. This external criterion is used to evaluate the extent to which student growth scores appear to be inflated. Using 2 years of data, growth comparisons are also made at the teacher level for teachers who submit SLOs and have students that take the state-administered large-scale assessment. Although they do show similar relationships with demographic covariates and have the same degree of stability across years, the two different measures of growth are weakly correlated. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2019 – Name: AN Label: Accession Number Group: ID Data: EJ1236254 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1236254 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1111/jedm.12233 Languages: – Text: English PhysicalDescription: Pagination: PageCount: 29 StartPage: 686 Subjects: – SubjectFull: Student Educational Objectives Type: general – SubjectFull: Student Evaluation Type: general – SubjectFull: Teacher Evaluation Type: general – SubjectFull: Scores Type: general – SubjectFull: Urban Schools Type: general – SubjectFull: School District Size Type: general – SubjectFull: Growth Models Type: general – SubjectFull: Measurement Type: general – SubjectFull: High Stakes Tests Type: general – SubjectFull: Inferences Type: general Titles: – TitleFull: Examining the Dual Purpose Use of Student Learning Objectives for Classroom Assessment and Teacher Evaluation Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Briggs, Derek C. – PersonEntity: Name: NameFull: Chattergoon, Rajendra – PersonEntity: Name: NameFull: Burkhardt, Amy IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 01 Type: published Y: 2019 Identifiers: – Type: issn-print Value: 0022-0655 Numbering: – Type: volume Value: 56 – Type: issue Value: 4 Titles: – TitleFull: Journal of Educational Measurement Type: main |
| ResultId | 1 |