Measuring Returns to Experience Using Supervisor Ratings of Observed Performance: The Case of Classroom Teachers
Saved in:
| Title: | Measuring Returns to Experience Using Supervisor Ratings of Observed Performance: The Case of Classroom Teachers |
|---|---|
| Language: | English |
| Authors: | Courtney Bell, Jessalynn James (ORCID |
| Source: | Journal of Policy Analysis and Management. 2025 44(1):12-44. |
| Availability: | Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us |
| Peer Reviewed: | Y |
| Page Count: | 33 |
| Publication Date: | 2025 |
| Document Type: | Journal Articles Reports - Evaluative |
| Descriptors: | Lesson Observation Criteria, Teaching Experience, Teacher Evaluation, Supervisors, Supervisory Methods, Teachers, Statistical Bias, Test Reliability, Context Effect, Robustness (Statistics) |
| Geographic Terms: | Tennessee, District of Columbia |
| DOI: | 10.1002/pam.22584 |
| ISSN: | 0276-8739 1520-6688 |
| Abstract: | We study the returns to experience in teaching, estimated using supervisor ratings from classroom observations. We describe the assumptions required to interpret changes in observation ratings over time as the causal effect of experience on performance. We compare two difference-in-differences strategies: the two-way fixed effects estimator common in the literature, and an alternative which avoids potential bias arising from effect heterogeneity. Using data from Tennessee and Washington, DC, we show empirical tests relevant to assessing the identifying assumptions and substantive threats--e.g., leniency bias, manipulation, changes in incentives or job assignments--and find our estimates are robust to several threats. |
| Abstractor: | As Provided |
| Entry Date: | 2025 |
| Accession Number: | EJ1456394 |
| Database: | ERIC |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwHFF-CVMpGsEzxtuRBw4hSMAAAA6DCB5QYJKoZIhvcNAQcGoIHXMIHUAgEAMIHOBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDEVudRmenGH_VMiJggIBEICBoMBhwQLsYHqG8vJ9pMVYlx261feBX5ytxDgtGuDBFjun9Nfjz4hfSigTyHevPtlNfFKO8k1_iwTDfr1-1CDM-EgtX90Mu1Y9fvKdiPvSJ4aOO9kmIruX2N-rX8D95Cz3-eS1ijUHLHcq3-gx_8tYvk6S6lcRGeuiKkTuGnIkB9ueXzovXH3zGUo1zbgRkysYG1iN-pWsfvsFFo3RJBi2Z2g= Text: Availability: 1 Value: <anid>AN0184198818;jpa01jan.25;2025Apr04.05:31;v2.2.500</anid> <title id="AN0184198818-1">Measuring returns to experience using supervisor ratings of observed performance: The case of classroom teachers </title> <sbt id="AN0184198818-2">INTRODUCTION</sbt> <p>We study the returns to experience in teaching, estimated using supervisor ratings from classroom observations. We describe the assumptions required to interpret changes in observation ratings over time as the causal effect of experience on performance. We compare two difference‐in‐differences strategies: the two‐way fixed effects estimator common in the literature, and an alternative which avoids potential bias arising from effect heterogeneity. Using data from Tennessee and Washington, DC, we show empirical tests relevant to assessing the identifying assumptions and substantive threats—e.g., leniency bias, manipulation, changes in incentives or job assignments—and find our estimates are robust to several threats.</p> <p>Monitoring employee job performance is a fundamental task in personnel management. In particular, understanding how performance improves with experience—the "returns to experience"—is critical to decisions about hiring and turnover, investments in employee training, and others. Consider the choice between retaining a current employee or replacing that employee with a novice new hire; the optimal choice depends not simply on the current performance of the two individuals, but rather on each person's expected future performance over time. However, isolating the causal effects of experience is complicated by imperfect and incomplete performance measures, and selection on performance through hiring and turnover decisions.</p> <p>Supervisor ratings of observed performance—a ubiquitous job performance measure—present a particular challenge when measuring returns to experience. For example, the relative subjectivity of supervisor ratings creates scope for leniency bias (Prendergast, [<reflink idref="bib51" id="ref1">51</reflink>]), and supervisors' leniency bias may itself depend on the employee's years of experience. Subjective supervisor ratings are quite common in public‐sector jobs where organizational objectives are diffuse and often difficult to quantify. We examine the case of classroom teachers and the most common performance measure for public‐school teachers: ratings by the school principal based on classroom observations.</p> <p>Teacher personnel management is central to education policy. Teacher salaries dwarf other public‐school expenses, consuming 3 out of 5 dollars, and teachers' contributions to student academic and social development dwarf other contributions by schools.</p> <p>Understanding the causal effects of experience on teaching—as measured by observation ratings—can improve teacher policy in two ways. First, as an input to developing new policy options. Over the last 2 decades, scholars have produced an expansive literature on how to measure teacher performance at a given point in time. But there remains comparatively little evidence on how teaching skills improve. Consequently, policymakers have struggled to develop policies or practices which consistently yield improvements in teaching effectiveness. Second, as an input to implementing existing policies. Consider, for example, school systems where tenure and salary decisions are now linked to observation ratings. The optimal rules for granting tenure or setting salaries depend critically on expected performance over time, not just at a single point in time.</p> <p>Our empirical focus in this paper is estimating the returns to experience in teaching using classroom observation ratings. We define "returns to experience" as the causal effect of 1 additional year of teaching experience on teacher performance, estimating returns separately for the first year of experience, second year, third year, etc. We define experience broadly to include whatever professional experiences occur over the course of a teacher's first year (or second year, etc.). Our primary objective is evaluating claims about returns to experience for (a) performance of the teaching practice inputs which the observation rubrics are designed to measure. But we also consider inferences about returns to experience on (b) broader output‐based measures of teacher performance, like teachers' value‐added contributions to student achievement scores. The extent to which experience affects (a) and (b) differently partly motivates our work, because input‐based measures are much more common in schools than output‐based measures.</p> <p>We use a difference‐in‐differences framework to make explicit the causal inference features of the returns‐to‐experience estimates. Our preferred estimates come from applying a diff‐in‐diff strategy proposed by de Chaisemartin and D'Haultfœuille ([<reflink idref="bib16" id="ref2">16</reflink>], 2023a). Briefly, the first difference is the observed change in a teacher's observation rating, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0001" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$({{{{\bar{s}}}&amp;#95;{jt}} - {{{\bar{s}}}&amp;#95;{j,t - 1}}})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , from year <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0002" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$({t - 1})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> to <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0003" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$t$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> when the teacher's experience changes from <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0004" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$({e - 1})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> to <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0005" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$e$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . The second difference is between early‐career (treated) teachers and veteran (comparison) teachers. These estimates are the solid lines in Figure 1 using data from Tennessee and the Washington, DC, Public Schools (DCPS). The identifying assumptions require: First, that, on average, veteran (comparison) teachers no longer experience returns to an additional year of experience. Second, that the process, explicit or implicit, that maps true performance to ratings does not depend on a teacher's years of experience.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/JPA/01jan25/pam22584-fig-0001.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="pam22584-fig-0001.jpg" title="1 Returns to experience measured in classroom observation ratings.Notes: The solid line reports estimates using our preferred diff‐in‐diff strategy described in the section &quot;Estimation Methods.&quot; The dashed line reports estimates using the conventional two‐way fixed effects approach described in the section &quot;Alternative Estimation Methods.&quot; The vertical lines mark the 95% confidence intervals which are corrected for clustering (teacher). In both cases the outcome variable is teacher j$j$'s classroom observation score, s¯jt${{\bar{s}}_{jt}}$, which is an average of several item‐level ratings recorded during a given school year t$t$. Observation scores are standardized (M = 0, SD = 1) by school year using the distribution of all teachers in the jurisdiction, Tennessee or DCPS, respectively. The solid line estimates are the difference between two means: (a) The average first‐difference, (s¯jt−s¯j,t−1)$({{{{\bar{s}}}_{jt}} - {{{\bar{s}}}_{j,t - 1}}})$, among &quot;treated&quot; teachers—those with e$e$ years of prior experience (x‐axis) in school year t$t$, and e−1$e - 1$ years in school year t−1$t - 1$. (b) The average first‐difference, (s¯jt−s¯j,t−1)$({{{{\bar{s}}}_{jt}} - {{{\bar{s}}}_{j,t - 1}}})$, among &quot;comparison&quot; teachers—those with ≥$ \ge$9 years of prior experience in both year t$t$ and t−1$t - 1$. The (a) minus (b) second‐difference is calculated separately for each unique combination of e$e$ and t$t$ in the data. Then the plotted points are the weighted average across t$t$ for a given e$e$, where the weights are the number of &quot;treated&quot; teachers. For the dashed line estimates we fit a single two‐way fixed effects regression, with teacher j$j$ and school year t$t$ fixed effects. The specification includes indicators for years of prior experience 0 through 8 individually, with ≥$ \ge$9 years the omitted category, but no other controls. The plotted points are the coefficients on the experience indicators. The sample size for the dashed line in Tennessee is 375,072 teacher‐by‐year observations for 81,847 unique teachers; and similarly 349,920 and 66,156 for solid line Tennessee; 33,484 and 7,268 for dashed line DCPS; and 33,040 and 7,201 for solid line DCPS." /> </p> <p></p> <p>We find that teacher job performance—as measured by classroom observation ratings—improves by roughly 1 standard deviation over the first 10 years of a teaching career, as shown in Figure 1. One‐third of the growth comes in just the first year. The improvements measured by observation ratings are somewhat larger in magnitude compared to improvements measured by teachers' contributions to student achievement scores. In our sample, value‐added improves by 0.10 student standard deviations over the first 10 years, as shown in Figure 2. One teacher standard deviation in value‐added is between 0.10 and 0.20 student standard deviations. Still, this pair of estimates of the returns to experience—1 standard deviation in ratings and 0.10 in value‐added—is consistent with prior cross‐sectional estimates of the ratings‐to‐value‐added relationship (Araujo et al., [<reflink idref="bib5" id="ref3">5</reflink>]; Burgess et al., [<reflink idref="bib9" id="ref4">9</reflink>]; Kane et al., [[<reflink idref="bib39" id="ref5">39</reflink>], [<reflink idref="bib37" id="ref6">37</reflink>]]).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/JPA/01jan25/pam22584-fig-0002.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="pam22584-fig-0002.jpg" title="2 Returns to experience measured in value‐added contributions to student achievement.Notes: The solid line reports estimates using our preferred diff‐in‐diff strategy described in the section &quot;Estimation Methods.&quot; The dashed line reports estimates using the conventional two‐way fixed effects approach described in the section &quot;Alternative Estimation Methods.&quot; The vertical lines mark the 95% confidence intervals which are corrected for clustering (teacher). In both cases the outcome variable is student i$i$'s test score, Aijst${{A}_{ijst}}$, in subject s$s$ and school year t$t$. Test scores are standardized (M = 0, SD = 1) within each grade‐by‐subject‐by‐year cell using the distribution for all students in the jurisdiction, Tennessee or DCPS respectively. For the dashed line estimates we fit a single two‐way fixed effects regression, with teacher j$j$ and school year t$t$ fixed effects. The specification includes indicators for years of prior experience 0 through 8 individually, with ≥$ \ge$9 years the omitted category. Additional controls are a quadratic in prior‐year test score, where the parameters are allowed to differ across grade‐by‐subject‐by‐year cells, b(Ais(t−1))$b({{{A}_{is({t - 1})}}})$. The plotted points are the coefficients on the experience indicators. For the solid line estimates, we begin by estimating teacher contributions to student test scores, μ̂jt${{\hat{\mu }}_{jt}}$. We fit a regression of student scores Aijst${{A}_{ijst}}$ on the same prior score controls, b(Ais(t−1))$b({{{A}_{is({t - 1})}}})$, and teacher fixed effects; and then obtain the residuals Aijst−b̂(Ais(t−1))${{A}_{ijst}} - \hat{b}({{{A}_{is({t - 1})}}})$. Our estimate μ̂jt${{\hat{\mu }}_{jt}}$ is the average residual for teacher j$j$ in year t$t$. The dashed line estimates are the difference between two means: (a) The average first‐difference, (μ̂jt−μ̂j,t−1)$({{{{\hat{\mu }}}_{jt}} - {{{\hat{\mu }}}_{j,t - 1}}})$, among &quot;treated&quot; teachers—those with e$e$ years of prior experience (x‐axis) in school year t$t$, and e−1$e - 1$ years in school year t−1$t - 1$. (b) The average first‐difference, (μ̂jt−μ̂j,t−1)$({{{{\hat{\mu }}}_{jt}} - {{{\hat{\mu }}}_{j,t - 1}}})$, among &quot;comparison&quot; teachers—those with ≥$ \ge$9 years of prior experience in both year t$t$ and t−1$t - 1$. The (a) minus (b) second‐difference is calculated separately for each unique combination of e$e$ and t$t$ in the data. Then the plotted points are the weighted average across t$t$ for a given e$e$, where the weights are the number of &quot;treated&quot; teachers. The sample size for the dashed line in Tennessee is 4,222,939 student‐by‐subject‐by‐year observations and 92,403 teacher‐by‐year observations for 34,395 unique teachers; and similarly 71,474 and 20,954 for solid line Tennessee; 247,005, 5,413 and 2,268 for dashed line DCPS; and 4,249 and 1,280 for solid line DCPS." /> </p> <p></p> <p>The paper goes on to evaluate several threats to the two identifying assumptions. Most threats are reasons why observation ratings might rise (or fall) over time even if a teacher's true performance is unchanged. One simple example is when changes are made to the scoring rubric, as happened in DCPS in 2017. As we discuss in detail, changes to the rubric (or to rater training, or to rater‐to‐teacher matching rules) do not necessarily threaten causal inferences about returns to experience estimates. Veteran teachers—the diff‐in‐diff comparison group—provide an estimate of the effect of such changes under the first assumption above, and that estimate is a reasonable counterfactual for early‐career teachers under the second assumption above. We use similar reasoning, combined with empirical evidence where available, to address other threats: rater leniency bias, raters using information from outside the observation, changes in incentives that distort teacher effort, manipulation behaviors by teachers which raise scores but not performance, the effect of job changes, and others. We find little evidence that these potential threats compromise a causal interpretation of our estimates.</p> <p>Our preferred estimation method is new to the literature on teacher returns to experience. Thus, we compare our estimates to estimates which use the conventional strategy. That strategy is also a difference‐in‐differences strategy using a two‐way fixed effects estimator, and both strategies require the same two main identifying assumptions. However, the conventional two‐way FE strategy requires additional assumptions.</p> <p>This is the first paper, to our knowledge, that studies the causal returns to experience reflected in supervisor ratings of observed performance. That contribution to the literature comes from combining explicit causal inference reasoning with panel data on performance ratings. Many prior papers have contributed causal estimates of the returns to experience using other measures of performance, for example, wages (e.g., Altonji &amp; Williams, [<reflink idref="bib2" id="ref7">2</reflink>]; Angrist, [<reflink idref="bib4" id="ref8">4</reflink>]; Grogger, [<reflink idref="bib27" id="ref9">27</reflink>]) or outputs like teacher value‐added to student achievement scores (e.g., Ost, [<reflink idref="bib47" id="ref10">47</reflink>]; Rivkin et al., [<reflink idref="bib52" id="ref11">52</reflink>]; Rockoff, [<reflink idref="bib53" id="ref12">53</reflink>]).</p> <p>Our focus is teachers and there is already a large literature on the returns to experience in teaching (see Taylor, [<reflink idref="bib58" id="ref13">58</reflink>], for a review). Still, we contribute to that literature in three ways. First, our paper provides a thorough discussion of causal inference considerations—identification strategies, assumptions, and threats—and a new estimation strategy. Claims of having estimated "returns to experience" are causal claims, but the contemporary tools of causal inference are mostly implicit (or entirely absent) in prior papers. We make explicit the difference‐in‐differences nature of returns to experience estimates, which clarifies identifying assumptions and threats. We further highlight how the typical estimator is a two‐way FE estimator, which will be biased if the returns to experience change over time or the distribution of experience changes over time. Finally, we demonstrate a new estimation strategy for the returns to experience in teaching. That new strategy incorporates recent developments in difference‐in‐differences methods (for reviews see de Chaisemartin &amp; D'Haultfœuille, [<reflink idref="bib18" id="ref14">18</reflink>], and Roth et al., [<reflink idref="bib55" id="ref15">55</reflink>]), and avoids the potential bias of the common two‐way‐FE strategy.</p> <p>Second, we contribute estimates of the returns to experience using a novel teacher performance measure: classroom observation ratings. Existing estimates of the returns to experience in teaching nearly all use value‐added measures of performance. Our estimates of how observation ratings (inputs) change with experience complement estimates of how value‐added scores (outputs) change, in part by contributing to efforts to understand the mechanisms behind teachers' improvements in value added (Atteberry et al., [<reflink idref="bib6" id="ref16">6</reflink>]; Kraft &amp; Papay, [<reflink idref="bib42" id="ref17">42</reflink>]; Ost, [<reflink idref="bib47" id="ref18">47</reflink>]). Two other papers also estimate returns to experience using observation ratings: Kraft et al. ([<reflink idref="bib43" id="ref19">43</reflink>]) and work concurrent to ours by Laski and Papay ([<reflink idref="bib44" id="ref20">44</reflink>]). Those two papers focused on understanding how the returns to experience vary across teachers, schools, etc. By contrast, our paper focuses on whether (or under what circumstances) the estimates should be interpreted as the causal effects of experience.</p> <p>Third, we explain and examine additional identifying assumptions and threats specific to the observation ratings case. These additional considerations are not relevant to the test‐score value‐added case. Our examination includes several novel empirical tests relevant to identification threats; tests which can be repeated in other settings.</p> <p>A final, broader contribution of our paper is to current policy debates involving teacher classroom observations ratings. First, most of our empirical results are framed by causal inference threats. However, those same results are also relevant to broader concerns about measurement in classroom observations, and in some cases our tests provide new evidence. For example, rater leniency bias (Kraft &amp; Gilmour, [<reflink idref="bib41" id="ref21">41</reflink>]; Steinberg &amp; Kraft, [<reflink idref="bib57" id="ref22">57</reflink>]), the influence of the students in the classroom (Campbell &amp; Ronfeldt, [<reflink idref="bib10" id="ref23">10</reflink>]), and unintended effects of teacher‐rater pairings (Chi, [<reflink idref="bib12" id="ref24">12</reflink>]), among other concerns (Cohen &amp; Goldhaber, [<reflink idref="bib13" id="ref25">13</reflink>]; Grissom &amp; Bartanen, [<reflink idref="bib26" id="ref26">26</reflink>]). Second, effective policy decisions depend on establishing a causal relationship between a (proposed) policy and a valued outcome. This paper provides an important foundation for policymakers and policy researchers who employ observed performance ratings to better understand teacher development policy.</p> <hd id="AN0184198818-5">DATA AND SETTING</hd> <p>Both DCPS and Tennessee maintain panel data on teachers, including ratings from classroom observations over several years. In DCPS, the panel begins with the start of its current evaluation system, IMPACT, in 2009/2010, and we use data through 2018/201919. Tennessee's current evaluation system began in 2011/2012, and our data run from that start date through 2018/2019. In both cases the data include item‐level ratings for several specific teaching tasks evaluated in a given observation visit. Teachers in tested grades and subjects can be linked to their students and achievement scores. Characteristics of the teachers and their students in our data are summarized in Table 1.</p> <p>1 TABLE Characteristics of the two samples.</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th /&gt;&lt;th align="center"&gt;Tennessee&lt;/th&gt;&lt;th align="center"&gt;DCPS&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th /&gt;&lt;th align="center"&gt;(1)&lt;/th&gt;&lt;th align="center"&gt;(2)&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align="center"&gt;(A) Students&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;At or above proficiency on NAEP&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Math, grade 4&lt;/td&gt;&lt;td&gt;0.39&lt;/td&gt;&lt;td&gt;0.31&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Math, grade 8&lt;/td&gt;&lt;td&gt;0.30&lt;/td&gt;&lt;td&gt;0.18&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Reading, grade 4&lt;/td&gt;&lt;td&gt;0.34&lt;/td&gt;&lt;td&gt;0.27&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Reading, grade 8&lt;/td&gt;&lt;td&gt;0.32&lt;/td&gt;&lt;td&gt;0.20&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Race/ethnicity&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Black&lt;/td&gt;&lt;td&gt;0.22&lt;/td&gt;&lt;td&gt;0.64&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Hispanic&lt;/td&gt;&lt;td&gt;0.09&lt;/td&gt;&lt;td&gt;0.18&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;White&lt;/td&gt;&lt;td&gt;0.64&lt;/td&gt;&lt;td&gt;0.13&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Other or multiple race or ethnicity&lt;/td&gt;&lt;td&gt;0.05&lt;/td&gt;&lt;td&gt;0.04&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Urbanicity&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;City&lt;/td&gt;&lt;td&gt;0.34&lt;/td&gt;&lt;td&gt;1.00&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Suburb&lt;/td&gt;&lt;td&gt;0.25&lt;/td&gt;&lt;td&gt;0.00&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Town&lt;/td&gt;&lt;td&gt;0.14&lt;/td&gt;&lt;td&gt;0.00&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Rural&lt;/td&gt;&lt;td&gt;0.27&lt;/td&gt;&lt;td&gt;0.00&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Share of school&amp;#8208;age population in poverty&lt;/td&gt;&lt;td&gt;0.22&lt;/td&gt;&lt;td&gt;0.28&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;English language learner&lt;/td&gt;&lt;td&gt;0.04&lt;/td&gt;&lt;td&gt;0.10&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Special education&lt;/td&gt;&lt;td&gt;0.13&lt;/td&gt;&lt;td&gt;0.17&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="center"&gt;(B) Teachers&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Observation score (original units)&lt;/td&gt;&lt;td&gt;3.94&lt;/td&gt;&lt;td&gt;3.17&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;(0.57)&lt;/td&gt;&lt;td&gt;(0.47)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Observation score, administrators&lt;/td&gt;&lt;td&gt;3.94&lt;/td&gt;&lt;td&gt;3.22&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;(0.57)&lt;/td&gt;&lt;td&gt;(0.49)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Observation score, master educators&lt;/td&gt;&lt;td /&gt;&lt;td&gt;3.02&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td /&gt;&lt;td /&gt;&lt;td&gt;(0.53)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;In student test score sample&lt;/td&gt;&lt;td&gt;0.23&lt;/td&gt;&lt;td&gt;0.15&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Female&lt;/td&gt;&lt;td&gt;0.79&lt;/td&gt;&lt;td&gt;0.74&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Race/ethnicity&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Black&lt;/td&gt;&lt;td&gt;0.06&lt;/td&gt;&lt;td&gt;0.51&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Hispanic&lt;/td&gt;&lt;td&gt;0.00&lt;/td&gt;&lt;td&gt;0.05&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;White&lt;/td&gt;&lt;td&gt;0.86&lt;/td&gt;&lt;td&gt;0.32&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Other or multiple race or ethnicity&lt;/td&gt;&lt;td&gt;0.08&lt;/td&gt;&lt;td&gt;0.04&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Graduate degree&lt;/td&gt;&lt;td&gt;0.55&lt;/td&gt;&lt;td&gt;0.69&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Years of experience&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Mean&lt;/td&gt;&lt;td&gt;11.83&lt;/td&gt;&lt;td&gt;10.86&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Standard deviation&lt;/td&gt;&lt;td&gt;(9.61)&lt;/td&gt;&lt;td&gt;(8.25)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Categorical&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;1st year teaching&lt;/td&gt;&lt;td&gt;0.06&lt;/td&gt;&lt;td&gt;0.07&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;2nd&lt;/td&gt;&lt;td&gt;0.06&lt;/td&gt;&lt;td&gt;0.07&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;3rd&lt;/td&gt;&lt;td&gt;0.06&lt;/td&gt;&lt;td&gt;0.07&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;4th&lt;/td&gt;&lt;td&gt;0.05&lt;/td&gt;&lt;td&gt;0.06&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;5th&lt;/td&gt;&lt;td&gt;0.05&lt;/td&gt;&lt;td&gt;0.06&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;6th&lt;/td&gt;&lt;td&gt;0.05&lt;/td&gt;&lt;td&gt;0.05&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;7th&lt;/td&gt;&lt;td&gt;0.04&lt;/td&gt;&lt;td&gt;0.05&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;8th&lt;/td&gt;&lt;td&gt;0.04&lt;/td&gt;&lt;td&gt;0.04&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;9th&lt;/td&gt;&lt;td&gt;0.04&lt;/td&gt;&lt;td&gt;0.04&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;10th or more&lt;/td&gt;&lt;td&gt;0.55&lt;/td&gt;&lt;td&gt;0.48&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>1 <emph>Notes</emph>: Panel A: National Assessment of Educational Progress (NAEP) scores are the simple mean of NAEP tests which occurred during the years in our analysis sample.</p> <p>2 Descriptive statistics for students are form the from National Center for Education Statistics' Common Core of Data. The exception is the "in poverty" statistic which comes from US Census Bureau Small Area Income and Poverty Estimates. Panel B: Authors' calculations using administrative data.</p> <hd id="AN0184198818-6">Features common to both settings</hd> <p>The DCPS and Tennessee settings share many features. In both locations, all teachers, regardless of experience level, are evaluated every school year by trained observers. The resulting observation ratings are a highly‐weighted component, among a larger set of evaluation measures including value‐added scores which measure teacher contributions to student achievement. The larger evaluation systems are used to identify exemplary teachers, those in need of additional support or training, or individuals who will be dismissed. During most of the period we study, teachers in DCPS were observed five times per year. After a change in the rubric in 2017, teachers were observed up to three times per year depending on experience and performance. In Tennessee, the number of evaluations per year varies according to teachers' prior performance and licensure status, but teachers are typically evaluated multiple times per year. The median novice teacher in Tennessee receives 2.5 formal observations and the median novice teacher in DCPS receives five formal observations.</p> <p>While the two systems use different observation rubrics, both rubrics assess similar tasks and teaching practices, and both rubrics have roots in Danielson's ([<reflink idref="bib15" id="ref27">15</reflink>]) Framework for Teaching. Tennessee uses the Tennessee Educator Acceleration Model (TEAM) evaluation rubric. The TEAM rubric's 19 items are divided into three categories of skills: instruction, planning, and environment. Each category is comprised of multiple items for teaching tasks. Ratings for each item range from 1 to 5 (5 = <emph>significantly above expectations</emph>, 1 = <emph>significantly below expectations</emph>). During most of the period of our analysis, DCPS used an observation rubric called the Teaching and Learning Framework (TLF). The TLF rubric has a 1 to 4 rating scale (4 = <emph>highly effective</emph>, 1 = <emph>ineffective</emph>) for items measuring nine teaching tasks. In 2017, DCPS transitioned to the Essential Practices (EP) observation rubric, which covers similar skills to the TLF, but with more concise definitions for each related task and explicit alignment to the Common Core State Standards.</p> <p>One frequent, but potentially misleading, criticism of such classroom observation systems is that the scores produced have little variation, with most teachers scoring in one or two top categories (Kraft &amp; Gilmour, [<reflink idref="bib41" id="ref28">41</reflink>]; Weisberg et al., [<reflink idref="bib60" id="ref29">60</reflink>]). This criticism, and most policy discussions, focus on the final end‐of‐year "summative" scores that end up in a teacher's personnel file. These final scores lack variation in part because final scores are rounded off to integer values. In this paper we use observation scores that average across many item ratings (several items and several observations of a given item), and those scores vary meaningfully, with a relatively Gaussian density (as shown in Appendix Figure A1).</p> <hd id="AN0184198818-7">Differences between the two settings</hd> <p>While both evaluation systems share many features, there are a number of useful differences. First, both places use trained school administrators as raters (e.g., principals and assistant principals, or other instructional leaders). However, until a change in 2017, in DCPS teachers were also observed and rated by "master educators"—specialized observers external to the school with subject‐ and grade‐specific expertise. Two of each teacher's annual observations were conducted by a master educator.</p> <p>Second, the two systems have different incentives and consequences associated with teachers' performance scores. While both DCPS and Tennessee might be considered high‐stakes evaluation systems, DCPS's has notably higher stakes. In DCPS, teachers with low performance (a final annual score below effective) are subject to involuntary dismissal. Prior work documents that these incentives influence teachers' behavior at work and their decisions about remaining at DCPS (Dee et al., [<reflink idref="bib19" id="ref30">19</reflink>]; Dee &amp; Wyckoff, [<reflink idref="bib20" id="ref31">20</reflink>]). There are also rewards in DCPS for high performance. Teachers who demonstrate exceptional performance (a final annual score of highly effective) are eligible for substantial bonuses and, if they continue to perform well, large base pay increases. In Tennessee, to earn tenure a teacher must receive a final composite score of "above expectations" or higher (roughly the top two‐thirds of teachers) for 2 consecutive years, after working at least 5 years total. Tenure can be revoked based on evaluation scores but that is rare: a teacher must score "below expectations" or lower (roughly the bottom 5% of teachers) for 2 consecutive years, and this rule does not apply to teachers who were tenured before 2011/2012.</p> <p>Finally, in addition to the specifics of their evaluation systems, DCPS and Tennessee differ from each other in size and many other characteristics. TEAM is used by nearly the entire state of Tennessee, and therefore includes teachers and schools across a range of settings and demographics. Each year the Tennessee data include roughly 84,000 teachers, of whom 5,500 are in their first year teaching, with 450,000 students at 1,350 schools. DCPS, on the other hand, is an urban majority‐minority and low‐income district, with approximately 3,500 teachers (290 novice) at 125 schools serving 46,000 students each year.</p> <hd id="AN0184198818-8">Additional data</hd> <p>In addition to classroom observation ratings, we have access to other data for teachers and students. For DCPS and Tennessee teachers, we know when they entered teaching, their experience in teaching, and other demographic characteristics. We have information regarding the observation raters and timing of the observation visits. In both settings we have the usual information regarding each teacher's students, for tested subjects and grades, including eligibility for free or reduced‐price lunch, race and ethnicity, and standardized achievement scores.</p> <p>Additionally, DCPS began using student surveys in 2016/2017 as teacher performance measures. This measure is adapted from the Tripod survey (Ferguson &amp; Danielson, [<reflink idref="bib21" id="ref32">21</reflink>]), which asks students questions about their teachers' practice. An example question is: "When explaining new ideas or skills in class, my teacher tells us about common mistakes that students might make."</p> <hd id="AN0184198818-9">RETURNS TO EXPERIENCE ESTIMATES</hd> <p></p> <hd id="AN0184198818-10">Estimation methods</hd> <p>We estimate the "returns to experience"—the improvement in performance caused by additional experience—using a difference‐in‐differences strategy. Our measure of performance, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0053" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\bar{s}}&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , is the classroom observation score for teacher <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0054" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$j$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> in school year <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0055" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$t$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . At the start of year <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0056" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$t$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , teacher <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0057" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$j$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> has <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0058" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$exp{{r}&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> years of prior teaching experience. Using these inputs, we apply the diff‐in‐diff estimator proposed by de Chaisemartin and D'Haultfœuille ([<reflink idref="bib16" id="ref33">16</reflink>], 2023a). Later, in "Alternative Estimation Methods," we compare this strategy to the more‐common estimation approach in the literature.</p> <p>Let <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0059" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\delta }&amp;#95;e}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> be the improvement in performance caused by gaining the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0060" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$e$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> th year of teaching experience. Our estimate of <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0061" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\delta }&amp;#95;e}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is: 1 <ephtml> &lt;math display="block" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0062" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&amp;#948;&amp;#770;e=&amp;#8721;tNet&amp;#948;&amp;#770;et&amp;#8721;tNet&amp;#948;&amp;#770;et=1Net&amp;#8721;j:exprj,t=e,exprj,t&amp;#8722;1=e&amp;#8722;1s&amp;#175;j,t&amp;#8722;s&amp;#175;j,t&amp;#8722;1&amp;#8722;1Met&amp;#8721;j:exprj,t&amp;#8722;1&amp;#8805;e&amp;#175;s&amp;#175;j,t&amp;#8722;s&amp;#175;j,t&amp;#8722;1,&lt;annotation encoding="application/x-tex"&gt;$$\begin{eqnarray}{{\hat{\delta }}&amp;#95;e} &amp;=&amp; \frac{{\mathop \sum \nolimits&amp;#95;t {{N}&amp;#95;{et}}{{{\hat{\delta }}}&amp;#95;{et}}}}{{\mathop \sum \nolimits&amp;#95;t {{N}&amp;#95;{et}}}}\nonumber\\ {{\hat{\delta }}&amp;#95;{et}} &amp;=&amp; \left[ {\frac{1}{{{{N}&amp;#95;{et}}}}\mathop \sum \limits&amp;#95;{ \def\eqcellsep{&amp;}\begin{array}{@{}*{1}{c}@{}} {j:exp{{r}&amp;#95;{j,t}} = e,}\\ {exp{{r}&amp;#95;{j,t - 1}} = e - 1} \end{array} } \left({{{{\bar{s}}}&amp;#95;{j,t}} - {{{\bar{s}}}&amp;#95;{j,t - 1}}} \right)} \right] - \left[ {\frac{1}{{{{M}&amp;#95;{et}}}}\mathop \sum \limits&amp;#95;{j:exp{{r}&amp;#95;{j,t - 1}} \ge \bar{e}} \left({{{{\bar{s}}}&amp;#95;{j,t}} - {{{\bar{s}}}&amp;#95;{j,t - 1}}} \right)} \right],\end{eqnarray}$$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> where <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0063" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\delta }&amp;#95;{et}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is simply the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0064" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\delta }&amp;#95;e}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> effect for a specific school year <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0065" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$t$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . The number of treated teachers is <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0066" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{N}&amp;#95;{et}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> and comparison teachers is <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0067" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;M&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{M}&amp;#95;{et}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . Because a teacher, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0068" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$j$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , may contribute to several <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0069" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\hat{\delta }}&amp;#95;e}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , our standard error estimates correct for clustering at the teacher level.</p> <p>The treatment is gaining the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0086" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$e$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> th year of experience. Thus, teacher <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0087" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$j$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is in the treated sample if they had <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0088" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$e$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> years of experience at the start of school year <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0089" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$t$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , but only <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0090" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$({e - 1})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> years of experience at the start of school year <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0091" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$({t - 1})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . Inside the brackets on the left is the average first‐difference in observation score, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0092" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\bar{s}}&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , for the sample of treated teachers. That first‐difference is the observed change in a teacher <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0093" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$j$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 's score between year <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0094" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$({t - 1})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> to <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0095" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$t$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> when that teacher's experience changes from <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0096" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$({e - 1})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> to <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0097" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$e$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . Still, a given teacher's scores may change over time for reasons unrelated to their own experience, which motivates the second difference between treated and comparison teachers.</p> <p>Our comparison sample is veteran teachers—teachers who have at least <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0098" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;annotation encoding="application/x-tex"&gt;$\bar{e}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> years of experience. One identifying assumption, which we formalize below, is that past <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0099" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;annotation encoding="application/x-tex"&gt;$\bar{e}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> years of teaching experience, there are no longer any returns to experience for the average teacher. Inside the brackets on the right is the average first‐difference for the veteran comparison teachers. Any observed change from <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0100" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$({t - 1})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> to <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0101" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$t$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> among veterans is, by assumption, unrelated to experience and differenced out. Our main estimates set <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0102" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;9&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$\bar{e} = 9$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , but our estimates are robust to higher values of <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0103" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;annotation encoding="application/x-tex"&gt;$\bar{e}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> as we show later.</p> <p>Each <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0104" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\hat{\delta }}&amp;#95;{et}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> estimate uses data from just two school years: one treated year, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0105" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$t$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , and one pre year, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0106" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$({t - 1})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . A teacher's performance in year <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0107" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$t$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is affected by that teacher's <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0108" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$e$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> th year of experience. A teacher's performance in year <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0109" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$({t + 1})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is also affected by their <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0110" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$e$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> th year of experience, but also affected by their <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0111" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$({e + 1})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> th year of experience. Thus, we observe the marginal effect of the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0112" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$e$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> th year of experience for one school year, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0113" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$t$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , after which the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0114" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$e$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> th year is confounded with further experience gains.</p> <p>The outcome variable, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0115" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\bar{s}}&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , is teacher <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0116" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$j$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 's classroom observation score for school year <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0117" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;.&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$t.$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> More precisely, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0118" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\bar{s}}&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is the average of several task‐specific scores, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0119" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mfrac&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mi&gt;K&lt;/mi&gt;&lt;/mfrac&gt;&lt;msubsup&gt;&lt;mo&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;K&lt;/mi&gt;&lt;/msubsup&gt;&lt;mspace width="-0.16em" /&gt;&lt;msub&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;${{\bar{s}}&amp;#95;{jt}} = \frac{1}{K}\mathop \sum \nolimits&amp;#95;{k=1}^{K}\!{{s}&amp;#95;{kjt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . The Tennessee rubric includes <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0120" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;K&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;19&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$K= 19$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> items and DCPS includes <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0121" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;K&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;9&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$K= 9$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . Our focus on the average observation score is motivated by an empirical constraint: while the tasks being scored are distinct—for example "teacher content knowledge" and "managing student behavior"—in practice the scores across tasks are highly correlated. In our Tennessee data, the mean correlation between items is 0.53 with a standard deviation of 0.05; in a factor analysis the first factor explains 95% of the variation in item scores. This correlation of items is common in classroom observation rubric scores (e.g., Kane et al., [<reflink idref="bib39" id="ref34">39</reflink>]). The <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0122" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\bar{s}}&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> scores are scaled in teacher standard deviation units, within jurisdiction (Tennessee or DC) by year cells.</p> <hd id="AN0184198818-11">Main results</hd> <p>Teacher performance measured in classroom observations improves with experience. In Figure 1 the solid line plots our returns to experience estimates from the difference‐in‐differences strategy in equation (<reflink idref="bib1" id="ref35">1</reflink>). Observation scores are scaled in standard deviation units, and, by construction, the zero line on the y‐axis is the average score among veteran teachers, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0126" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8805;&lt;/mo&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;9&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$exp{{r}&amp;#95;{jt}} \ge \bar{e} = 9$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . The vertical lines mark cluster‐corrected 95% confidence intervals.</p> <p>Just 1 year of teaching experience improves performance by one‐quarter to one‐third of a standard deviation. Over the first 10 years of a teaching career, performance in observations improves 1 standard deviation. The patterns in DC and Tennessee are quite similar. What are the educational or economic consequences of these gains? Improvements in teaching inputs contribute to outputs, including student learning which we can measure with teachers' value added to student test scores. There is evidence that teachers' test‐score value‐added contributions translate to better longer‐run outcomes, including college going and labor market success (Chetty et al., [<reflink idref="bib11" id="ref36">11</reflink>]). However, there are currently no estimates linking observation scores to longer‐run student outcomes.</p> <p>The pattern in Figure 1, using classroom observation ratings, is similar to the pattern of returns to experience for teacher value added to student achievement scores. In Figure 2 the solid line plots estimates where the performance measure is a teacher's value‐added contribution to student test scores. We first obtain value‐added scores, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0127" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#956;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\hat{\mu }}&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , then apply the estimator in equation (<reflink idref="bib1" id="ref37">1</reflink>) substituting <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0128" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#956;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\hat{\mu }}&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> for <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0129" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\bar{s}}&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . In Figure 2, the y‐axis, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0130" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#956;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\hat{\mu }}&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , is measured in student standard deviation units, and the sample is limited to teachers of grades 4 through 8 in math and English language arts. The pattern for Tennessee in Figure 2 matches estimates from several other places (see Taylor, [<reflink idref="bib58" id="ref38">58</reflink>], for a review). The DC estimates are much nosier but consistent with the typical pattern. In both cases teacher value‐added improves by about 0.10 student test‐score standard deviations over the first 10 years of teaching.</p> <p>The improvements measured by observation ratings are somewhat larger in magnitude compared to improvements measured by value added. One teacher standard deviation in value‐added is between 0.10 and 0.20 student test‐score standard deviations. There are many potential explanations for the difference, given the variety of teacher skills and tasks not captured by observation scores or by test‐score value‐added. Still, there are several existing cross‐sectional estimates of the relationship between observation ratings and value‐added (Araujo et al., [<reflink idref="bib5" id="ref39">5</reflink>]; Burgess et al., [<reflink idref="bib9" id="ref40">9</reflink>]; Kane et al., [[<reflink idref="bib39" id="ref41">39</reflink>], [<reflink idref="bib37" id="ref42">37</reflink>]]). In those estimates a 1 standard deviation increase in observation ratings predicts a 0.05 to 0.11 increase in value‐added. Our pair of estimates of the returns to experience, Figures 1 and 2, are quite consistent with those prior cross‐sectional estimates.</p> <hd id="AN0184198818-12">Causal inference</hd> <p>The difference‐in‐differences setup provides a familiar framework for evaluating causal claims about the estimates in Figure 1. Stated in general terms, the identifying assumption in this case is: any change over time we observe in veteran (comparison) teachers' scores is the same change we would see in early‐career (treated) teachers' scores if there were no returns to experience. We can clarify the identifying assumption further with the help of a simple conceptual framework.</p> <hd id="AN0184198818-13">Observation scores and true performance</hd> <p>A teacher's job involves many tasks: learning content, planning lessons, asking questions in class, responding to misbehavior, grading, communicating with parents, any many more. Each of those tasks produces some input to the production of student achievement or other goals of schooling. Let <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0131" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\theta }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> measure true performance of task <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0132" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$k$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . Higher performance is synonymous with producing more or higher‐quality task <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0133" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$k$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> inputs.</p> <p>Classroom observation rubrics are designed to measure task performance, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0134" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\theta }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , at least for some subset of a teacher's tasks. Rubrics are not designed to measure outcomes like student achievement. For example, observers are asked to score the nature and frequency of questions teachers ask students, but observers are not asked to assess whether these questions generated student learning. Observation scores are also sometimes described as measures of a teacher's skills. But an observation score is a function of both skills and effort, thus we prefer describing those scores as measures of performance.</p> <p>Still, classroom observations are an imperfect way to measure performance. An observation score, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0135" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{s}&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , is inevitably some combination of true performance, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0136" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\theta }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , and other factors unrelated to performance, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0137" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . For exposition we assume: 2 <ephtml> &lt;math display="block" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0138" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;mo linebreak="badbreak"&gt;=&lt;/mo&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mfenced separators="" open="(" close=")"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mo linebreak="goodbreak"&gt;=&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;mo linebreak="goodbreak"&gt;+&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;mspace width="0.33em" /&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$$\begin{equation}{{s}&amp;#95;k} = g\left({{{\theta }&amp;#95;k},{{\nu }&amp;#95;k}} \right) = {{\theta }&amp;#95;k} + {{\nu }&amp;#95;k}\ \end{equation}$$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml></p> <p>Those other factors, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0139" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , include much more than just classical measurement error. Even as the number of observations grows, features of the evaluation process will create some difference between <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0140" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;mo&gt;[&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;]&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$E[ {{{s}&amp;#95;k}} ]$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> and <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0141" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;mo&gt;[&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;]&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$E[ {{{\theta }&amp;#95;k}} ]$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . First, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0142" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> includes explicit features of the evaluation process, for example, the rubric itself, how evaluators are trained, how evaluators are assigned to teachers, incentives attached to scores. Such explicit features are (mostly) controllable by those designing and implementing the evaluation. But <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0143" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> also includes less‐explicit less‐controllable features, for example, the behaviors teachers or evaluators choose in response to the explicit features.</p> <hd id="AN0184198818-14">Identifying assumptions</hd> <p>Interpreting Figure 1 as the returns to experience—the causal effect of teaching experience on true task performance—requires two identifying assumptions. <emph>Assumption 1</emph>: Factors which contribute to observation scores but are unrelated to performance, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0144" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> in equation (<reflink idref="bib2" id="ref43">2</reflink>), do not depend on teaching experience. Specifically, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0145" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;[&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;]&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;[&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mo&gt;]&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$E[ {{{\nu }&amp;#95;{kjt}}|k,t,exp{{r}&amp;#95;{jt}}} ] = E[ {{{\nu }&amp;#95;{kjt}}|k,t} ]$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . This assumption requires that if an early‐career and a veteran teacher both have the same true task performance, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0146" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\theta }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , they will have the same observation score, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0147" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{s}&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . <emph>Assumption 2</emph>: True performance is not changing over time, on average, in the comparison group of teachers. Specifically, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0148" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;mo&gt;[&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8805;&lt;/mo&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;/mrow&gt;&lt;mo&gt;]&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$E[ {{{\theta }&amp;#95;{kjt}} - {{\theta }&amp;#95;{kj({t - 1})}}|exp{{r}&amp;#95;{jt}} \ge \bar{e}} ] = 0$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> .</p> <p>The importance of a comparison group is shown by stating the assumption that would replace Assumption 2 in the absence of a comparison group. <emph>Assumption 3</emph>: The <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0149" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> factors do not change over time. Specifically, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0150" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;[&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;]&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;[&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mo&gt;]&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$E[ {{{\nu }&amp;#95;{kjt}}|k,t,exp{{r}&amp;#95;{jt}}} ] = E[ {{{\nu }&amp;#95;{kjt}}|k} ]$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . If we used only early‐career teachers' data, we could not separate the returns to experience from changes in <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0151" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> over time, because <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0152" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$exp{{r}&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> and <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0153" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$t$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> are colinear within teacher. Later, in "Alternative Explanations and Threats to Casual Inference," we discuss several different substantive threats to these identifying assumptions, but some of the quite‐plausible threats are known changes in <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0154" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> over time.</p> <p>These are the assumptions required for claims about performance of the teaching tasks which classroom observations are designed to measure. We might also be interested in claims about other aspects of teacher performance, like teachers' value‐added to student achievement scores. Imagine a production process for student achievement; some of the inputs will be the teaching tasks described in an observation rubric. However, to make any inference from observation scores to value‐added would require a much better understanding of that production process than currently exists. Later we provide some new empirical evidence relevant to that broader inference.</p> <hd id="AN0184198818-15">Alternative estimation methods</hd> <p>Our estimation methods, described in "Estimation Methods," are new to the literature on returns to experience in teaching. Here we compare our estimation strategy to the conventional estimation strategy—the strategy which, to date, has been most common in that literature (see Taylor, [<reflink idref="bib58" id="ref44">58</reflink>], for a review). The conventional strategy is also a difference‐in‐differences strategy, but using a two‐way fixed effects estimator, though it is not often described in those terms. Both strategies require the same core set of identifying assumptions, but the conventional strategy requires additional assumptions about effect heterogeneity.</p> <p>In the conventional approach, estimates of the returns to experience come from a least‐squares regression. The basic specification is: 3 <ephtml> &lt;math display="block" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0160" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo linebreak="badbreak"&gt;=&lt;/mo&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mfenced separators="" open="(" close=")"&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mo linebreak="goodbreak"&gt;+&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#955;&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;/msub&gt;&lt;mo linebreak="goodbreak"&gt;+&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#960;&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;mo linebreak="goodbreak"&gt;+&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#949;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$$\begin{equation}{{\bar{s}}&amp;#95;{jt}} = h\left({exp{{r}&amp;#95;{jt}}} \right) + {{\lambda }&amp;#95;j} + {{\pi }&amp;#95;t} + {{\varepsilon }&amp;#95;{jt}},\end{equation}$$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> where the outcome is a measure of teacher performance, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0161" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\bar{s}}&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> in our case.</p> <p>Selective attrition is a fundamental threat to any returns‐to‐experience estimate. Attrition from teaching is likely negatively correlated with performance. In response to that threat, nearly all estimation strategies focus on variation within individual teachers over time. Our main strategy uses only within‐teacher variation by first differences. The conventional approach uses teacher fixed effects (e.g., Rockoff, [<reflink idref="bib53" id="ref45">53</reflink>]).</p> <p>However, for a given teacher, years of experience, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0162" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$exp{{r}&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , is colinear with school year, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0163" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$t$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , unless the teacher takes a leave of absence. Specification 3 includes both teacher fixed effects, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0164" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#955;&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\lambda }&amp;#95;j}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , and school year fixed effects, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0165" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#960;&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\pi }&amp;#95;t}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , and thus requires some restriction on <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0166" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$h({exp{{r}&amp;#95;{jt}}})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> to avoid the age‐period‐cohort problem. The typical restriction is to assume no returns to experience after some number of years, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0167" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;annotation encoding="application/x-tex"&gt;$\bar{e}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . Then <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0168" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$h({exp{{r}&amp;#95;{jt}}})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is a series of indicator variables for years of experience up to <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0169" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;annotation encoding="application/x-tex"&gt;$\bar{e}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> : 4 <ephtml> &lt;math display="block" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0170" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;hexprjt=&amp;#8721;e=0e&amp;#175;&amp;#8722;1&amp;#946;e&amp;#215;1exprjt=eand&amp;#948;e=&amp;#946;e&amp;#8722;&amp;#946;e&amp;#8722;1.&lt;annotation encoding="application/x-tex"&gt;$$\begin{eqnarray} &amp;&amp;h\left({exp{{r}&amp;#95;{jt}}} \right) = \sum\limits&amp;#95;{e = 0}^{\bar{e} - 1} {{\beta }&amp;#95;e} \times {\bf 1}\left\{ {exp{{r}&amp;#95;{jt}} = e} \right\}\nonumber \\ &amp;&amp;\hbox{and }{{\delta }&amp;#95;e} = {{\beta }&amp;#95;e} - {{\beta }&amp;#95;{e - 1}}.\end{eqnarray}$$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml></p> <p>The omitted category is veterans, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0171" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;{&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8805;&lt;/mo&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;/mrow&gt;&lt;mo&gt;}&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$1\{ {exp{{r}&amp;#95;{jt}} \ge \bar{e}} \}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . This restriction maps to identifying Assumption 2, as stated earlier. That required assumption is well known in the literature on returns to experience in teaching; the assumption is sometimes stated explicitly (e.g., Rockoff, [<reflink idref="bib53" id="ref46">53</reflink>]) and sometimes criticized (e.g., Papay &amp; Kraft, [<reflink idref="bib48" id="ref47">48</reflink>]).</p> <p>This conventional estimation strategy uses a two‐way fixed effects estimator. Notice in Specification 3 the characteristic group and period fixed effects, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0181" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#955;&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\lambda }&amp;#95;j}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> and <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0182" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#960;&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\pi }&amp;#95;t}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , and a series of treatment indicators, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0183" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$h({exp{{r}&amp;#95;{jt}}})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . A recent, growing literature clarifies several properties of two‐way FE estimators; in particular, how those estimators can produce biased estimates when treatment effects are heterogeneous (for reviews see de Chaisemartin and D'Haultfœuille, [<reflink idref="bib18" id="ref48">18</reflink>], and Roth et al., [<reflink idref="bib55" id="ref49">55</reflink>]).</p> <p>Three types of heterogeneity can create bias in two‐way FE estimates. The first two types, and the resulting bias, are now regularly discussed in papers using difference‐in‐differences methods. The third type is specific to settings with multiple treatments, including the returns to experience estimates which produce <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0184" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\hat{\delta }}&amp;#95;e}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> for several <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0185" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$e$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> .</p> <p>The first source of potential bias would arise if the effects of the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0186" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$e$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> th year of experience, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0187" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\delta }&amp;#95;e}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , differ across cohorts of teachers. In other words, variation over time in <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0188" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\delta }&amp;#95;e}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , given the link between cohort, time, and experience. Why might <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0199" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\delta }&amp;#95;e}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> change over time? First, selection into (or out of) teaching may change over time. For example, schools may get better (or worse) over time at selecting hires on potential job performance (e.g., Jacob et al., [<reflink idref="bib34" id="ref50">34</reflink>]), or self‐selection by (prospective) teachers may change in response to compensation for potential performance inside or outside schools (e.g., Leaver et al., [<reflink idref="bib45" id="ref51">45</reflink>]; Nagler et al., [<reflink idref="bib46" id="ref52">46</reflink>]). Second, the nature of treatment itself may change over time. For example, if schools devote more resources to mentoring for early‐career teachers (e.g., Kraft et al., [<reflink idref="bib40" id="ref53">40</reflink>]; Rockoff, [<reflink idref="bib54" id="ref54">54</reflink>]).</p> <p>The second source of potential bias would, theoretically, arise if the effects of the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0200" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$e$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> th year of experience increase (or decrease) over time within a cohort of teachers. However, in practice, this second bias is not a concern in the conventional returns to experience estimates. In Specification 3, with <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0209" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$h({exp{{r}&amp;#95;{jt}}})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> in equation 4, treated teachers (early‐career teachers) are not part of the comparison group until they are beyond <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0210" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;annotation encoding="application/x-tex"&gt;$\bar{e}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> (until they become veteran teachers).</p> <p>The third source of potential bias arises when there are multiple treatments, as detailed in de Chaisemartin and D'Haultfœuille ([<reflink idref="bib17" id="ref55">17</reflink>]). In the current setting, first, the estimate for the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0215" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$e$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> th year, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0216" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\hat{\delta }}&amp;#95;e}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , can be biased if the effects in other years, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0217" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8712;&lt;/mo&gt;&lt;mrow&gt;&lt;mo&gt;{&lt;/mo&gt;&lt;mrow&gt;&lt;mtext&gt;...&lt;/mtext&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mtext&gt;...&lt;/mtext&gt;&lt;/mrow&gt;&lt;mo&gt;}&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;${{\delta }&amp;#95;{({ - e})}} \in \{ { \ldots ,{{\delta }&amp;#95;{e - 2}},{{\delta }&amp;#95;{e - 1}},{{\delta }&amp;#95;{e + 1}},{{\delta }&amp;#95;{e + 2}}, \ldots } \}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , differ across cohorts. Second, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0218" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\hat{\delta }}&amp;#95;e}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> can be biased if the distribution of teacher experience is changing over time; in other words, if the probability of the other treatments is changing over time. Even if all the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0219" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\delta }&amp;#95;e}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> are not changing over cohorts, this second bias threat remains. The distribution of experience may be fairly constant for an entire state, like Tennessee, but could change over time for a district, like DCPS, with changes in hiring and retention strategies.</p> <p>Do these potential biases affect our estimates? The dashed line in Figure 1 shows our estimates from the common two‐way FE strategy, alongside our preferred strategy. In Tennessee the two lines are nearly identical, suggesting little change from cohort to cohort in the returns to experience. Estimated growth over 10 years is just 2% of a standard deviation smaller with the two‐way FE approach compared to our preferred estimates. By contrast, there is some difference for DCPS, suggesting the two‐way FE strategy underestimates the steepness of returns to experience. Over 10 years the accumulated difference is about 17% of a standard deviation. Estimated growth in any 1 year, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0232" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$({e - 1})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> to <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0233" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$e$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , differs by 2% of a standard deviation on average. Up through the fifth or sixth year, the year‐to‐year estimated changes are nearly identical.</p> <p>The differences in DCPS estimates are largely explained by changes in the distribution of teacher experience in DCPS over time. Appendix Figure A3 shows that the distribution of experience shifted away from early‐career teachers over time but became more stable from 2014/2015 on. If we restrict out analysis to this more‐stable more‐recent period, the standard and alternative approaches are quite similar, as shown in Appendix Figure A4. The observable changes in the distribution of experience in DCPS may also be correlated with changes in the returns to experience among DCPS teachers.</p> <p>The two strategies also yield similar estimates when the performance measure is a teacher's value‐added contribution to student test scores, as shown in Figure 2. The value‐added returns‐to‐experience estimates are much noisier, given the much smaller samples. For Tennessee, estimated growth over ten years is one‐quarter of a standard deviation smaller with the two‐way FE approach compared to our preferred estimates. Estimates are more similar for DCPS. However, in both settings, we cannot reject the null hypothesis of no difference between the two strategies.</p> <hd id="AN0184198818-16">ALTERNATIVE EXPLANATIONS AND THREATS TO CAUSAL INFERENCE</hd> <p>Observation ratings may improve (or decline) over time for reasons unrelated to a teacher's gains from experience. In this section we describe several alternative explanations for changing ratings, and whether an alternative explanation threatens a causal "returns to experience" interpretation of Figure 1. We focus specifically on interpreting changes in observation ratings as the causal effect of experience on performance of the tasks which the rubric is designed to measure.</p> <hd id="AN0184198818-17">General evidence</hd> <p>Before taking up specific alternative explanations, we begin with some general evidence relevant to the plausibility of identifying Assumptions 1 and 2. First, consider Assumption 1 which requires: factors which contribute to observation scores but are unrelated to performance, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0234" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , do not depend on teacher experience. We cannot test this assumption directly. However, if Assumption 1 is true, we would predict that <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0235" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;v&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#956;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$cov({{{\nu }&amp;#95;{jt}},{{{\hat{\mu }}}&amp;#95;{jt}}})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> should not depend on experience, and thus <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0236" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;v&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#956;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$cov({{{{\bar{s}}}&amp;#95;{jt}},{{{\hat{\mu }}}&amp;#95;{jt}}})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> should also not depend on experience. Here <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0243" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#956;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\hat{\mu }}&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is the teacher's value‐added contribution to student achievement test scores. We can test this prediction which is relevant to judging Assumption 1.</p> <p>Why might <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0244" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;v&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#956;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$cov({{{\nu }&amp;#95;{jt}},{{{\hat{\mu }}}&amp;#95;{jt}}})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> depend on experience? Raters may give greater scrutiny to early‐career teachers, perhaps with the specific goal of increasing the correlation between observation scores and value added. Alternatively, raters may be more lenient with early‐career teachers, reducing <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0245" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;v&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#956;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$cov({{{\nu }&amp;#95;{jt}},{{{\hat{\mu }}}&amp;#95;{jt}}})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . We discuss other possibilities in "The Evaluation System" and "Behavior of the Raters." However, Assumption 1 could be violated without affecting <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0246" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;v&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#956;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$cov({{{\nu }&amp;#95;{jt}},{{{\hat{\mu }}}&amp;#95;{jt}}})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . For example, raters might simply add 1 point to all scores for early‐career teachers, which, barring any ceiling effects, would not affect the correlation between observation scores and value added conditional on experience.</p> <p>Figure 3 provides information on the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0247" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;v&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#956;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$cov({{{{\bar{s}}}&amp;#95;{jt}},{{{\hat{\mu }}}&amp;#95;{jt}}})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , whether that relationship depends on teacher experience, and thus a test of the prediction outlined above. The x‐axis is years of prior experience. The y‐axis is the predicted increase in value added, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0248" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#956;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\hat{\mu }}&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , if a teacher's observation score, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0249" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\bar{s}}&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , increases by 1 standard deviation. As shown in Figure 3, the estimated relationship between observation scores and value added is largely unrelated to experience. There is no clear trend over experience, and we cannot reject the null hypothesis that each point estimate is equal to the average of the series it belongs to, though the DCPS estimates are quite noisy. The one exception is the earliest years in Tennessee using only within‐teacher variation (solid line series). Those estimates suggest the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0250" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;v&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#956;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$cov({{{{\bar{s}}}&amp;#95;{jt}},{{{\hat{\mu }}}&amp;#95;{jt}}})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> may be higher in a teacher's first year of employment. Some of the specific threats described below could be a mechanism behind the first‐year correlation. In summary, the lack of a relationship between <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0254" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;v&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#956;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$cov({{{{\bar{s}}}&amp;#95;{jt}},{{{\hat{\mu }}}&amp;#95;{jt}}})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> and experience in Figure 3 is consistent with Assumption 1, through the prediction outlined above.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/JPA/01jan25/pam22584-fig-0003.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="pam22584-fig-0003.jpg" title="3 Predicting student test scores with teacher observation scores by years of teacher experience.Notes: The solid and dashed lines each report estimates from a separate linear regression. The vertical lines mark the 95% confidence intervals which are corrected for clustering (teacher). In both cases the outcome variable is student i$i$'s test score, Aijst${{A}_{ijst}}$, in subject s$s$ (maths or English language arts pooled) and school year t$t$. Test scores are standardized (M = 0, SD = 1) within each grade‐by‐subject‐by‐year cell using the distribution for all students in the jurisdiction, Tennessee or DCPS respectively. In both cases the specification includes (a) indicators for years of prior experience 0 through 8 individually, with ≥$ \ge$9 years the omitted category; (b) classroom observation score, s¯jt${{\bar{s}}_{jt}}$; and (c) the interactions of (a) and (b). Each plotted point is sum of the coefficient on the (a)*(b) interaction for e$e$ years of prior experience (x‐axis) plus the main‐effect coefficient on (b). Additional controls are a quadratic in prior‐year test score, where the parameters are allowed to differ across grade‐by‐subject‐by‐year cells, b(Ais(t−1))$b({{{A}_{is({t - 1})}}})$. The solid line specification includes year and teacher fixed effects. The dashed line includes only year fixed effects, omitting the teacher fixed effects. The sample size the same for the two lines; in Tennessee 4,222,939 student‐by‐subject‐by‐year observations and 92,403 teacher‐by‐year observations for 34,395 unique teachers, and similarly in DCPS 252,400, 5,429, and 2,274." /> </p> <p></p> <p>We can also partially test identifying Assumption 2. That assumption requires that, on average, true performance, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0263" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\theta }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , is not changing over time among the comparison group of veteran teachers, i.e., <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0264" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;mo&gt;[&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8805;&lt;/mo&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;/mrow&gt;&lt;mo&gt;]&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$E[ {{{\theta }&amp;#95;{kjt}} - {{\theta }&amp;#95;{kj({t - 1})}}|exp{{r}&amp;#95;{jt}} \ge \bar{e}} ] = 0$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . Our main estimates in Figure 1 set <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0265" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;9&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$\bar{e} = 9$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> to define the veteran group. If Assumption 2 holds, then our estimates for returns at <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0266" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mspace width="0.33em" /&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$e\ = $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 0–8 should be robust to setting <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0267" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;annotation encoding="application/x-tex"&gt;$\bar{e}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> above 9.</p> <p>Our returns to experience estimates are quite robust to changes in <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0268" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;annotation encoding="application/x-tex"&gt;$\bar{e}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . The solid line in Figure 4 simply repeats the solid line in Figure 1 for convenient comparison, with <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0269" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;9&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$\bar{e} = 9$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . The two dashed lines show estimates where <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0270" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;14&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$\bar{e} = 14$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> and <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0271" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;19&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$\bar{e} = 19$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . The three lines have different intercepts; the intercept in this case is the average performance among veteran teachers with <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0272" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8805;&lt;/mo&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$exp{{r}&amp;#95;{jt}} \ge \bar{e}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . Still, the slopes of the lines are quite similar across estimates over the range of 0 to 9 years of prior experience. For example, in Tennessee, the slope between <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0273" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$({e - 1}) = 0$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> and <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0274" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$e= 1$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is 0.346 standard deviations in the estimates with <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0275" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;9&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$\bar{e} = 9$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> (solid line) and 0.352 when <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0276" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;14&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$\bar{e} = 14$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> (long‐dash line), a difference of 0.006. Indeed, that same difference in slope estimates is 0.006‐0.007 for all of the pairwise <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0277" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$({e - 1})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> to <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0278" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$e$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> slopes. The accumulated difference over the first ten years is 0.05. Those slopes are the returns to experience we want to estimate, and those estimated changes are robust to the choice of <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0279" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;annotation encoding="application/x-tex"&gt;$\bar{e}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> .</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/JPA/01jan25/pam22584-fig-0004.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="pam22584-fig-0004.jpg" title="4 Estimates by definition of comparison group.Notes: Each of the three lines reports estimates using our preferred diff‐in‐diff strategy described in the section &quot;Estimation Methods.&quot; The vertical lines mark the 95% confidence intervals which are corrected for clustering (teacher). The solid line is identical to the solid line in Figure 1. For the two dashed lines, the details of estimation are identical to the solid with one exception. For the solid line, the comparison group is teachers with ≥$ \ge $9 years of experience, e¯=9$\bar{e} = 9$. The two dashed lines show e¯=14$\bar{e} = 14$ and e¯=19$\bar{e} = 19$ respectively. The sample size the same for all three lines; in Tennessee 375,072 teacher‐by‐year observations for 81,847 unique teachers, and similarly in DCPS 33,484 and 7,267." /> </p> <p></p> <p>Additionally, while we cannot observe <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0287" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#916;&lt;/mi&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mspace width="0.33em" /&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mspace width="0.33em" /&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;mo&gt;[&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8805;&lt;/mo&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;/mrow&gt;&lt;mo&gt;]&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;${{\Delta}}\theta \ = \ E[ {{{\theta }&amp;#95;{kjt}} - {{\theta }&amp;#95;{kj({t - 1})}}|exp{{r}&amp;#95;{jt}} \ge \bar{e}} ]$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> directly, we can observe <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0288" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#916;&lt;/mi&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;[&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8805;&lt;/mo&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;/mrow&gt;&lt;mo&gt;]&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;${{\Delta}}\bar{s} = E[ {{{{\bar{s}}}&amp;#95;{kjt}} - {{{\bar{s}}}&amp;#95;{kj({t - 1})}}|exp{{r}&amp;#95;{jt}} \ge \bar{e}} ]$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . Among veteran teachers, the mean first‐difference in observation scores is 0.004 standard deviations (st.err. 0.002) in Tennessee and <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0289" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$ - $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 0.073 standard deviations (st.err. 0.006) in DCPS. Under what conditions would <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0291" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#916;&lt;/mi&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mo&gt;&amp;#8773;&lt;/mo&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;${{\Delta}}\bar{s} \cong 0$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> but <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0292" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#916;&lt;/mi&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mo&gt;&amp;#8800;&lt;/mo&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;${{\Delta}}\theta \ne 0$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> ? Only in the knife‐edge case where any change in true performance, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0293" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\theta }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , is just offset by a change in the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0294" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> component of scores.</p> <hd id="AN0184198818-20">The evaluation system</hd> <p>Changes in observation ratings over time may be caused by changes to the evaluation system's tools and procedures. Key features of an evaluation system include the scoring rubric, the training provided to raters, and the rules for assigning teachers to raters. Even if a teacher's performance, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0295" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\theta }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , remains constant, the rating assigned to that performance, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0296" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{s}&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , may go up or down if the system's processes change. In other words, the evaluation system's tools and procedures are key features of <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0297" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> in equation (<reflink idref="bib2" id="ref56">2</reflink>) where <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0298" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;${{s}&amp;#95;k} = {{\theta }&amp;#95;k} + {{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . (The incentives or consequences attached to performance measures are also a key feature of an evaluation system, and we discuss those incentives below.)</p> <p>The most straightforward example of a change in <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0299" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is a change in the scoring rubric. In 2017 DCPS switched from the Teaching and Learning Framework (TLF) rubric to an entirely new Essential Practices (EP) rubric. The new EP rubric did not measure exactly the same set of tasks, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0300" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$k$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , as the old TLF rubric. Changes in other settings might be smaller, like word choices, even if the tasks scored remain the same. Still, large or small rubric changes would not necessarily threaten our identifying assumptions, as long as the rubric changes affect early‐career (treatment) and veteran (comparison) teachers equally.</p> <p>The DCPS changes allow us to compare estimates from different rubrics. In Figure 5 the short dash line shows estimates of returns to experience using only ratings generated by the TLF, while the long dash blue line uses only EP ratings. Both dashed lines are limited to scores from school administrators. For both rubrics the average first‐year teacher's rating is much lower than the average veteran's rating, but that starting gap is smaller with the EP rubric. Using TLF rubric data suggests the average teacher improves by 1.3 standard deviations over the first 10 years, compared to 1.2 standard deviations using the EP rubric. Though the differences are not statistically significant. The differences suggest a potential threat to Assumption 1—that <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0301" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> does not depend on experience—at the time of the change in rubrics in DCPS. However, the difference between the dashed (TLF) and long‐dashed (EP) estimates could be a compositional change. Starting in 2011, and thus concurrent with our data, DCPS became more selective in both hiring and retention decisions, with selection strategies based explicitly on performance measure (Dee &amp; Wyckoff, [<reflink idref="bib20" id="ref57">20</reflink>]; Jacob et al., [<reflink idref="bib34" id="ref58">34</reflink>]). There were noticeably fewer early‐career teachers by 2017 (Appendix Figure A3). Thus, in Figure 5, the higher scores with the EP rubric may reflect true higher performance because of selection.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/JPA/01jan25/pam22584-fig-0005.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="pam22584-fig-0005.jpg" title="5 Estimates using different rubrics and rater types (DCPS).Notes: Each of the three lines reports estimates using our preferred diff‐in‐diff strategy described in the section &quot;Estimation Methods.&quot; The vertical lines mark the 95% confidence intervals which are corrected for clustering (teacher). The details of estimation are identical to the solid line in Figure 1 with the following exceptions. First, the estimation sample is limited by the type of rater: external &quot;Master Educators&quot; for the solid line, and school administrators for the dashed and long dashed lines. Second, the estimation sample is limited by the rubric used: TLF from 2010 to 2016 and EP from 2017 to 2019. The sample size for the solid line is 18,715 teacher‐by‐year observations for 5,118 unique teachers; and similarly 21,080 and 5,380 for dashed line, and 10,190 and 3,726 for the long dash line." /> </p> <p></p> <p>Choosing raters is also a key evaluation design decision, and a decision which itself may change over time. Figure 5 also compares estimates by rater type for DCPS. The solid red line uses only ratings from the master educator raters, who specialize in rating and are external to the school, while the dashed red line uses only ratings from school administrators. Both lines are limited to scores generated by the TLF rubric, and there is no composition concern since each teacher was rated by both a school administrator and master educator each year. The slopes of the two TLF lines are quite similar, especially over the first 5 years of a teacher's career. Using master educator ratings suggests improvement of over 1.6 standard deviations over the 10 years, compared to 1.3 using principal ratings. Though, again, the differences are not statistically significant. Figure 5 does obscure one important difference between master educator scores and school administrator scores: School administrators give higher average scores on the 1 to 4 scale; in other words, the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0302" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> component in equation (<reflink idref="bib2" id="ref59">2</reflink>) does depend on rater type. However, the difference in scores between the rater types is the same for all teachers regardless of experience; thus, the rater type difference in <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0303" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> does not violate Assumption 1.</p> <p>In general, changes to the evaluation system are changes to the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0304" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> component in equation (<reflink idref="bib2" id="ref60">2</reflink>). Interpreting Figure 1 as the causal returns to experience does not require that <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0305" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> remain unchanged over time. The only restriction on <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0306" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is that <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0307" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> not depend on experience. This applies to obvious changes in <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0308" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , like the rubric or types of evaluators, and to changes which are more difficult (for the researcher) to observe. Thus, while the tests in Figure 5 address some threats, they do not rule out all threats from changes in the evaluation system. One potentially difficult to observe change is changes to the training of raters. Imagine that system administrators determine, at a given point in time, that raters need to be re‐trained on some aspect of scoring. That re‐training might be in fact be motivated by administrators' belief that scores, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0309" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{s}&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , are not reflecting performance, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0310" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\theta }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , as they should. A second example is a change to the rules for assigning teachers to raters. Chi ([<reflink idref="bib12" id="ref61">12</reflink>]), among others, has documented teacher‐rater match effects on observation scores; for example, when a teacher and rater share a gender or race, the teacher's scores are higher. Imagine the evaluation system administrators decide, at some point, to make gender or race an explicit factor in the rules for making assignments.</p> <hd id="AN0184198818-22">Behavior of the raters</hd> <p>Changes in ratings over time may reflect changes in the behavior of the raters. Raters have some discretion within any evaluation system's designed procedures. Rubric‐based classroom observation ratings fall somewhere in between the theoretical poles of truly objective evaluation and purely subjective evaluation. Moreover, raters may also take actions which violate the designed procedures they were trained to follow. The behavior of raters, whether intended or unintended in the system design, is part of the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0311" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> component in equation (<reflink idref="bib2" id="ref62">2</reflink>).</p> <p>One behavior that is frequently cited, given rater discretion, is leniency bias—the tendency for raters to give scores which are higher than warranted. Histograms of observation ratings (Appendix Figure A1) are consistent with systematic leniency bias in both Tennessee and DCPS, although such bias is less evident for ratings assigned by the master educators in DCPS. The skew in the ratings distribution could also accurately reflect teacher performance using a rubric with ceiling effects. Leniency bias is often cited as a concern in classroom observation scores by both researchers and in public debate (Anderson, [<reflink idref="bib3" id="ref63">3</reflink>]; Kraft &amp; Gilmour, [<reflink idref="bib41" id="ref64">41</reflink>]), but leniency bias is common in many occupations beyond teaching (Prendergast, [<reflink idref="bib51" id="ref65">51</reflink>]).</p> <p>However, leniency bias does not necessarily threaten our interpretation of Figure 1 as the causal returns to experience. To violate Assumption 1— <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0312" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> does not depend on experience—rater leniency would need to be correlated with teacher experience. For example, imagine that raters are less lenient with a first‐year teacher compared to their rating of the same teacher in the teacher's second year; then Figure 1 would over‐state the returns to the first year of teaching. Such a change in leniency might be a mechanism behind the decline in Figure 3 after the first year, for teachers in Tennessee. However, if it is not correlated with experience, leniency bias will be differenced out in the same way as rubric changes or other evaluation system features.</p> <p>Another potential mechanism is that raters may use information learned outside an official observation visit. Consider the case of a teacher rated by their school principal. A few brief classroom observations are a small fraction of the interactions a teacher and principal will have in a school year; the principal likely learns much about the teacher's performance outside of official observations. Ho and Kane ([<reflink idref="bib29" id="ref66">29</reflink>]) showed evidence that a teacher's own principal scores a video of the teacher's classroom differently than a principal from another school in the district scores the same video, perhaps because the teacher's own principal begins the scoring with a prior on the teacher's performance. Additionally, because the rubric covers only some teaching tasks, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0313" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$k$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , a principal may raise (or lower) observation scores to reflect the principal's beliefs about the teacher's performance of tasks not covered by the rubric. A principal using outside information is a potentially rational behavior if the observation ratings are used for personnel decisions and the principal cares much less about observation scores than about student outcomes and teacher value‐added to those outcomes.</p> <p>This outside information explanation may threaten Assumption 1— <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0314" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> does not depend on experience—but only if raters both have and use different outside information depending on a teacher's years of experience. The number of years a teacher–principal pair has worked together likely will be correlated with the teacher's years of experience, but it does not need to be strongly correlated if school principals switch schools frequently. A high correlation would suggest principal raters might have different outside information on early‐career and veteran teachers. Empirically the correlation between years‐worked‐together and experience is 0.17 in the DCPS data and 0.15 in the Tennessee data.</p> <p>One test relevant to this outside‐information question is the event study of ratings in Figure 6. Event time is relative to a change in the school principal, with year zero the new principal's first year, and we allow the time series to differ for early‐career and veteran teachers as shown by the two plotted lines. If principals learn about a teacher's performance outside of formal classroom observations, we might expect observation scores to rise or fall. However, scores do not change on average as a principal and teacher work together longer. This pattern holds for both early‐career and veteran teachers. In Tennessee there is some evidence that principals give slightly lower scores in their first year in a new school (about 5% of a standard deviation lower).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/JPA/01jan25/pam22584-fig-0006.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="pam22584-fig-0006.jpg" title="6 Event study of a change in school principal.Notes: All estimates are from a single linear regression. The vertical lines mark the 95% confidence intervals which are corrected for clustering (teacher). The dependent variable is teacher j$j$'s classroom observation score, s¯jt${{\bar{s}}_{jt}}$, which is an average of several item‐level scores recorded during a given school year t$t$. Observation scores are standardized (M = 0, SD = 1) by school year using the distribution of all teachers in the jurisdiction, Tennessee or DCPS respectively. The specification includes (a) indicators for year relative to a change in school principal; (b) an indicator =1$ = 1$ if teacher j$j$ has ≤$ \le$4 years of prior experience, and =0$ =\!0$ if teacher j$j$ has ≥$ \ge$9 years; and the interaction of (a) and (b). The new principal's first year, x‐axis =0$ =\! 0$, is omitted for both groups defined by (b). The specification also includes indicators for years of prior experience, with ≥$ \ge$9 years omitted, plus teacher and year fixed effects. If a teacher experiences two (or more) principal changes, we stack the data to include each teacher‐by‐event‐study case in the data. DCPS observation scores in panel B represent administrator‐assigned scores only, but can include multiple administrators (i.e., principals and assistant principals) within a given teacher‐year. The sample size for the solid line in Tennessee is 72,850 teacher‐by‐year observations for 29,193 unique teachers; and similarly 136,443 and 32,244 for dashed line Tennessee; 6,927 and 2,511 for solid line DCPS; and 9,597 and 2,406 for dashed line DCPS." /> </p> <p></p> <p>Figure 6 alone cannot exclude the threat. While Figure 6 suggests principals do not use outside information, that interpretation assumes the information principals have is increasing year over year. Perhaps principals get to know teachers very well in just one year, and then information increases very little after the first year. The outside information gained in year 1 could affect observation scores in future years, but we would not see change over time in Figure 6. We emphasize these points as a reminder that the empirical tests in this section are intended to be informative, about both causal inference considerations and policy debates, but these tests are not dispositive.</p> <hd id="AN0184198818-24">Incentives and distortion of effort</hd> <p>Changes in ratings may reflect changes in the incentives attached to those ratings. Those incentives might be explicitly linked to observation ratings, like monetary bonuses or the threat of dismissal, or less‐explicit career concerns incentives. Still, a change in incentives alone does not threaten inferences about true performance, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0328" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\theta }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , for tasks covered by the rubric. A new or stronger incentive attached to task <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0329" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$k$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 's score, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0330" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{s}&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , can induce a teacher to raise their performance of that task, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0331" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\theta }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , through more effort for task <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0332" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$k$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> or investing in skills for task <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0333" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mo&gt;.&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$k.$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> Thus, inferences about true performance, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0334" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\theta }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , of tasks covered by the rubric are not necessarily threatened by a change in incentives attached to ratings, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0335" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{s}&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> .</p> <p>However, an increase in effort for tasks covered by the rubric, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0336" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$k$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , may come at the expense of teacher performance in other tasks not covered, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0337" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$ - k$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . This asymmetry between scored tasks and unscored tasks suggests scope for the well‐known multitask distortion problem (Holmstrom &amp; Milgrom, [<reflink idref="bib30" id="ref67">30</reflink>]). Given that potential distortion, a change in incentives attached to rubric ratings can threaten inferences about teacher performance beyond the scope of what is covered by the rubric. Recall that the rubric tasks are inputs to the broader education production responsibilities of teachers, including improving student math achievement, social skills, earnings as an adult, etc.</p> <p>Still, using ratings and incentives to shift teacher effort away from some tasks and toward other tasks will not necessarily lead to distortion. There is (quasi‐)experimental evidence that rubric‐based classroom observations can improve teachers' contributions to student test scores, even when teachers are not evaluated based on those test scores (Briole &amp; Maurin, [<reflink idref="bib7" id="ref68">7</reflink>]; Burgess et al., [<reflink idref="bib8" id="ref69">8</reflink>]; Taylor &amp; Tyler, [<reflink idref="bib59" id="ref70">59</reflink>]). In DCPS specifically, teacher performance improves more when the teacher spends more of the year anticipating an unannounced rater visit (Phipps, [<reflink idref="bib49" id="ref71">49</reflink>]; Phipps &amp; Wiseman, [<reflink idref="bib50" id="ref72">50</reflink>]).</p> <p>While incentives do not necessarily threaten our causal interpretation of Figure 1 as the returns to experience, changes in incentives may be a mechanism behind the improvements seen in Figure 1. The simplest example is tenure rules. In Tennessee, teachers can earn tenure after 5 years, but tenure requires sufficiently high observation ratings in years 4 and 5. Thus, teachers have somewhat more incentive to focus effort on the rubric‐measured tasks in years 4 and 5 compared to years 1, 2, and 3, which might contribute to the pattern in Figure 1. Still, it seems unlikely a teacher concerned about tenure would wait until year 4 to pay attention to the rubric, and the slope from years 3 to 4 in Figure 1 is not obviously a departure from the trend suggested by the other year‐to‐year slopes.</p> <p>Unlike Tennessee, the evaluation incentives in DCPS were not explicitly a function of years of experience but could have been correlated with experience. DCPS teachers are dismissed if rated "Minimally Effective" (the second‐lowest rating) in 2 consecutive years or if they fail to exceed a "Developing" rating (the third‐lowest rating) within 3 consecutive years. Before fall 2012, teachers could receive permanent salary increases after 2 consecutive years of being rated "Highly Effective" (the top rating). Figure 7 shows the proportion of teachers in each rating category by years of experience, suggesting the incentives are not strongly correlated with experience.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/JPA/01jan25/pam22584-fig-0007.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="pam22584-fig-0007.jpg" title="7 Incidence of consequential performance ratings (DCPS).Notes: Each plotted series reports the percentage of teachers scoring at the relevant consequential rating level. In DCPS, teachers who receive their first Minimally Effective rating must improve the following year or risk dismissal. Beginning in 2012/2013, teachers who have earned a second consecutive Developing rating are likewise subject to dismissal if they fail to improve. Through spring 2012, Highly Effective teachers were conversely eligible for large financial rewards. The share of teachers facing each performance incentive are estimated only within the respective years in which the incentive was in place. The sample for the solid line includes 35,672 teachers‐by‐year and 9,455 unique teachers; and similarly for the dashed line 22,344 and 6,936, and for the long dashed line 10,004 and 4,755." /> </p> <p></p> <hd id="AN0184198818-26">Manipulation of ratings</hd> <p>Observation ratings may reflect changes in teachers' actions unrelated to their job performance. Teachers, like professionals in any other occupation, may adopt behaviors or actions which do raise their ratings, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0338" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{s}&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , but do not raise their true job performance, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0339" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\theta }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . In the literature on job performance evaluation these actions are known as manipulation. This manipulation of observation ratings might occur, for example, because classroom observations are infrequent and brief; thus, a teacher could prepare a special lesson or even rehearse the lesson with his students in advance of the rater's visit. By contrast, if the evaluation process or incentives prompted a teacher to improve their lessons on all (or many of) the days the rater would not be present, that would be an improvement in performance and not manipulation.</p> <p>Manipulation plausibly threatens our casual returns‐to‐experience interpretation of Figure 1. In our framework, teacher manipulation results from the evaluation system's procedures and incentives, and is part of the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0340" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> component in equation (<reflink idref="bib2" id="ref73">2</reflink>). A teacher's awareness of how to manipulate likely grows as he gains experience with the evaluation system. That suggests a plausible correlation between potential for manipulation and general teaching experience, which threatens Assumption 1 that <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0341" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is invariant to experience. However, that correlation might be weakened if more‐experienced teachers share their manipulation strategies with newly‐hired teachers. If the manipulation component of observation scores is unrelated to general experience, then manipulation will be differenced out in Figure 1.</p> <p>The decline in correlation after year 1 in Tennessee in Figure 3 may be explained by increasing manipulation over the first few years of a teacher's career. However, we cannot rule out other mechanisms, such as, for example, raters becoming more lenient as a teacher moves from the first to the second year. And there are other limitations to the test in Figure 3, as discussed above. On the other hand, while underpowered, the evidence from DCPS in Figure 3 does not indicate a decline in the relationship between classroom observation scores and student achievement over experience. In addition, the relatively stable correlation between classroom observation ratings and student survey scores across levels of teaching experience in DCPS (Appendix Figure A5) provide evidence against the presence of manipulation, unless teachers were similarly able to manipulate scores on both measures across levels of experience.</p> <p>Dee and Wyckoff ([<reflink idref="bib20" id="ref74">20</reflink>]) examined whether DCPS school administers manipulate observation scores, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0342" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{s}&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , in the face of increased incentives. Consider the teachers who received their first Minimally Effective rating in 2010/2011, and thus were under a significant threat of dismissal during 2011/2012. Observation ratings did improve in 2011/2012 for these teachers, on average. However, master educators also scored these teachers as having improved, and the increase in observation scores was similar across both types of raters. Additionally, these teachers under dismissal threat also improved on their test‐score value added. Taken together, these results suggest that the dismissal threat did not improve observation ratings through manipulation alone.</p> <hd id="AN0184198818-27">Changes in job assignments</hd> <p>Changes in a teacher's ratings may reflect changes in their job assignment. A teacher's observation ratings, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0343" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{s}&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , might decline (or improve) after a job change for either of two reasons: First, the teacher's actual performance, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0344" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\theta }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , could decline (or improve) because of the job change. Using teacher value‐added to student test score, Ost ([<reflink idref="bib47" id="ref75">47</reflink>]) provided evidence that teaching skills and experience are not fully transferable across grade levels. Switching from 3rd to 5th grade, for example, likely requires some adjusting of questioning techniques, or shifting effort to new lesson plans at the expense of in‐class performance.</p> <p>Let <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0345" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$a$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> and <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0346" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msup&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mo&gt;&amp;#8242;&lt;/mo&gt;&lt;/msup&gt;&lt;annotation encoding="application/x-tex"&gt;$a^{\prime}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> be two different job assignments; <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0347" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\theta }&amp;#95;{kjta}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is the actual performance of teacher <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0348" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$j$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> in task <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0349" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$k$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> during school year <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0350" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$t$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> with job assignment <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0351" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$a$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . We can write: 5 <ephtml> &lt;math display="block" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0352" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;mfenced separators="" open="[" close="]"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mfenced separators="" open="(" close=")"&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mo linebreak="badbreak"&gt;=&lt;/mo&gt;&lt;munder&gt;&lt;munder accentunder="true"&gt;&lt;mrow&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;mfenced separators="" open="[" close="]"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mfenced separators="" open="(" close=")"&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;mo&gt;&amp;#65080;&lt;/mo&gt;&lt;/munder&gt;&lt;msup&gt;&lt;mi mathvariant="normal"&gt;&amp;#916;&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msup&gt;&lt;/munder&gt;&lt;mo linebreak="goodbreak"&gt;+&lt;/mo&gt;&lt;munder&gt;&lt;munder accentunder="true"&gt;&lt;mrow&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;mfenced separators="" open="[" close="]"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mfenced separators="" open="(" close=")"&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mfenced separators="" open="(" close=")"&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;msup&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mo&gt;&amp;#8242;&lt;/mo&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;mo&gt;&amp;#65080;&lt;/mo&gt;&lt;/munder&gt;&lt;msup&gt;&lt;mi mathvariant="normal"&gt;&amp;#916;&lt;/mi&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;/msup&gt;&lt;/munder&gt;&lt;mspace width="0.33em" /&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$$\begin{equation}E\left[ {{{\theta }&amp;#95;{kjt}} - {{\theta }&amp;#95;{kj\left({t - 1} \right)}}} \right] = \underbrace{{E\left[ {{{\theta }&amp;#95;{kjta}} - {{\theta }&amp;#95;{kj\left({t - 1} \right)a}}} \right]}}&amp;#95;{{{{{{\Delta}}}^t}}} + \underbrace{{pE\left[ {{{\theta }&amp;#95;{kj\left({t - 1} \right)a}} - {{\theta }&amp;#95;{kj\left({t - 1} \right)a^{\prime}}}} \right]}}&amp;#95;{{{{{{\Delta}}}^a}}}\ \end{equation}$$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> where <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0353" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$p$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is the probability of switching from job <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0354" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msup&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mo&gt;&amp;#8242;&lt;/mo&gt;&lt;/msup&gt;&lt;annotation encoding="application/x-tex"&gt;$a^{\prime}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> to <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0355" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$a$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> .</p> <p>The intuitive notion of returns to experience implies that the job is constant and experience increases, which matches <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0356" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msup&gt;&lt;mi mathvariant="normal"&gt;&amp;#916;&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msup&gt;&lt;annotation encoding="application/x-tex"&gt;${{{{\Delta}}}^t}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> in Expression 5. If identifying Assumption 2 holds—no returns to additional experience for veterans—then Figure 1 reports estimates of <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0357" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mi mathvariant="normal"&gt;&amp;#916;&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msup&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msup&gt;&lt;mi mathvariant="normal"&gt;&amp;#916;&lt;/mi&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$({{{{{\Delta}}}^t} + p{{{{\Delta}}}^a}})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . Assuming further that job changes reduce performance, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0358" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mi mathvariant="normal"&gt;&amp;#916;&lt;/mi&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;/msup&gt;&lt;mo&gt;&amp;#60;&lt;/mo&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;${{{{\Delta}}}^a} &amp;#60; 0$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , then Figure 1 underestimates the intuitive <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0359" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msup&gt;&lt;mi mathvariant="normal"&gt;&amp;#916;&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msup&gt;&lt;annotation encoding="application/x-tex"&gt;${{{{\Delta}}}^t}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . Alternatively, some researchers or policymakers may be interested <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0360" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mi mathvariant="normal"&gt;&amp;#916;&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msup&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msup&gt;&lt;mi mathvariant="normal"&gt;&amp;#916;&lt;/mi&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$({{{{{\Delta}}}^t} + p{{{{\Delta}}}^a}})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , which we could describe as the "returns to experience including job changes typical of early‐career teachers."</p> <p>Job changes do threaten identifying Assumption 2, which requires that <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0361" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;mo&gt;[&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mo&gt;&amp;#8805;&lt;/mo&gt;&lt;/mrow&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;/mrow&gt;&lt;mo&gt;]&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$E[ {{{\theta }&amp;#95;{kjt}} - {{\theta }&amp;#95;{kj({t - 1})}}|expr \ge \bar{e}} ] = 0$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> in our comparison group of veteran teachers. A veteran's performance might change because of a job change, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0362" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mi mathvariant="normal"&gt;&amp;#916;&lt;/mi&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;/msup&gt;&lt;mo&gt;&amp;#8800;&lt;/mo&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;${{{{\Delta}}}^a} \ne 0$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , even if that veteran's performance would not have otherwise changed, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0363" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mi mathvariant="normal"&gt;&amp;#916;&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msup&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;${{{{\Delta}}}^t} = 0$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . If job changes do reduce veteran (comparison) teacher performance, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0364" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mi mathvariant="normal"&gt;&amp;#916;&lt;/mi&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;/msup&gt;&lt;mo&gt;&amp;#60;&lt;/mo&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;${{{{\Delta}}}^a} &amp;#60; 0$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , then the estimates in Figure 1 overstate the intuitive <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0365" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msup&gt;&lt;mi mathvariant="normal"&gt;&amp;#916;&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msup&gt;&lt;annotation encoding="application/x-tex"&gt;${{{{\Delta}}}^t}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> for novices. This bias is positive, and the bias described in the prior paragraph is negative, but the two would only cancel each other out under the assumption that <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0366" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$p$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> and <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0367" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msup&gt;&lt;mi mathvariant="normal"&gt;&amp;#916;&lt;/mi&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;/msup&gt;&lt;annotation encoding="application/x-tex"&gt;${{{{\Delta}}}^a}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> do not depend on experience.</p> <p>The second reason scores might change is that the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0369" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> component in equation (<reflink idref="bib2" id="ref76">2</reflink>) might differ across jobs. For example, typically the same rubric is used for all teachers, leaving any adaptation to grade‐level or subject circumstances up to the rater or training process. More subtly, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0370" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> might depend on the students in the classroom (Campbell &amp; Ronfeldt, [<reflink idref="bib10" id="ref77">10</reflink>]). Students are themselves an important feature of a teacher's job assignment, and a feature which can change even if grade level or subject do not. The threat to identification parallels other features of <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0371" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> discussed above. As long as job‐specific differences in <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0372" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> are unrelated to experience, this second reason is not a serious threat to identification. A job‐specific difference might be, for example, if raters are more lenient with novices after a job change than they are with veterans.</p> <p>In Figure 8 we test the robustness of Figure 1 to changes in the students a teacher is assigned. Using data from Tennessee and DCPS, we plot returns‐to‐experience estimates with and without controls for students prior‐year test scores. Accounting for changes in students assigned does not affect our estimates. The similarity of all the estimates in Figures 1 and 8 is partly because they all use only within‐teacher variation. The <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0375" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> component might well depend on the students in the classroom (Campbell &amp; Ronfeldt, [<reflink idref="bib10" id="ref78">10</reflink>]), but most of the variation in students assigned is between teachers or schools, not within teachers over time.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/JPA/01jan25/pam22584-fig-0008.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="pam22584-fig-0008.jpg" title="8 Estimates controlling for student baseline test scores.Notes: Both the solid and dashed lines report estimates using our preferred diff‐in‐diff strategy described in the section &quot;Estimation Methods.&quot; Both use the same identical sample of teacher‐by‐year observations. The vertical lines mark the 95% confidence intervals, which are corrected for clustering (teacher). For the solid line &quot;teachers with baseline test scores&quot; estimates, the details of estimation are identical to the solid line in Figure 1 except that we restrict the estimation sample. The solid line sample includes only teacher‐by‐year observations where we have both an average observation rating, s¯jt${{\bar{s}}_{jt}}$, and baseline test scores, Ai(t−1)${{A}_{i({t - 1})}}$, for the students i$i$ assigned to teacher j$j$ in year t$t$. For the dashed line &quot;controlling for baseline test scores&quot; estimates, the details of estimation are identical to the solid line except that we first residualize the outcome, s¯jt${{\bar{s}}_{jt}}$, using the mean baseline test score, Ai(t−1)${{A}_{i({t - 1})}}$, among teacher j$j$'s students. The sample size the same for the two lines; in Tennessee 3,076,946 student‐by‐subject‐by‐year observations and 65,750 teacher‐by‐year observations for 25,017 unique teachers, and similarly in DCPS 250,377, 5,369 and 2,258." /> </p> <p></p> <hd id="AN0184198818-29">Performance improvements among veteran teachers</hd> <p>The true performance of veteran (comparison group) teachers may change over time—violating Assumption 2—even if there are no returns to experience for veterans. For example, veterans may increase their effort in response to incentives. How would interpretation change if Assumption 2 was violated in this way, but Assumption 1 held? If the veteran gains were only among veterans, then the estimates in Figure 1 would likely understate the true returns to experience for early‐career teachers. The veterans' improvements would be subtracted off any improvements for early‐career teachers.</p> <hd id="AN0184198818-30">Turnover</hd> <p>One final consideration in interpreting Figure 1 is turnover or attrition from our estimation sample. The estimates in Figure 1 use only within‐teacher variation in observation scores. This feature addresses a first‐order potential bias: average observation ratings might rise with experience, even if each individual teacher's scores remain constant, if lower‐rated teachers are more likely to leave teaching (or at least leave the district or state).</p> <p>Still, even using only within‐teacher variation, Figure 1 is still partly determined by turnover. In Figure 1 the slope between year 1 and year 2 is an average of <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0384" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{N}&amp;#95;{1,2}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> different individual teacher slopes, where <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0385" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{N}&amp;#95;{1,2}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is the sample of individuals who are observed in year 1 and year 2 (and perhaps future years). Similarly, the slope between year 4 and year 5 uses only the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0386" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;4&lt;/mn&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mn&gt;5&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{N}&amp;#95;{4,5}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> sample. However, these are not the same samples: <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0387" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;4&lt;/mn&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mn&gt;5&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8800;&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;${{N}&amp;#95;{4,5}} \ne {{N}&amp;#95;{1,2}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . First, for any given cohort of novice hires, attrition from the profession over time will make <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0388" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;4&lt;/mn&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mn&gt;5&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8834;&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;${{N}&amp;#95;{4,5}} \subset {{N}&amp;#95;{1,2}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . Second, experienced teachers who transfer into the system from elsewhere may contribute to <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0389" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;4&lt;/mn&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mn&gt;5&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{N}&amp;#95;{4,5}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> even if they do not contribute to <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0390" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{N}&amp;#95;{1,2}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . The slope from year 1 to year 2 in Figure 1 might be different if we could estimate it with the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0391" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;4&lt;/mn&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mn&gt;5&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{N}&amp;#95;{4,5}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> sample.</p> <p>Empirically, however, our Figure 1 estimates are not strongly influenced by this second‐order composition concern. Figure 9 shows our returns‐to‐experience estimates using subsamples defined by when the teacher leaves teaching in Tennessee or DCPS. The changes from year 1 to 2, 2 to 3, etc. are quite similar across samples. The exception is that the trajectory appears to change in a teacher's final year before leaving teaching in Tennessee or DCPS.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/JPA/01jan25/pam22584-fig-0009.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="pam22584-fig-0009.jpg" title="9 Estimates by year of exit.Notes: All lines report estimates using our preferred diff‐in‐diff strategy described in the section &quot;Estimation Methods.&quot; The vertical lines mark the 95% confidence intervals which are corrected for clustering (teacher). For each line in the figure, the details of estimation are identical to the solid line in Figure 1 except for the estimation sample. The sample for each of the four solid lines is defined by how many years the teacher taught in the jurisdiction (Tennessee or DC). Each teacher is observed for exactly 2, 3, 4, or 5 consecutive years and then not observed in the data subsequently. The dashed line includes teachers observed for 6 or more consecutive years. The sample size the same for the two series; in Tennessee 27,853 teacher‐by‐year observations for 6,613 unique teachers, and similarly in DCPS 31,785 and 8,931." /> </p> <p></p> <hd id="AN0184198818-32">CONCLUSION</hd> <p>The typical estimates of returns to experience, applied to observation ratings, can reasonably be interpreted as the causal effect of additional experience on teachers' job performance—specifically, performance of the input tasks covered by the rubric. The estimates are difference‐in‐differences estimates, where veteran teachers are the comparison group. Veterans provide a plausible counterfactual estimate for several often‐stated threats, including for example, leniency bias from raters, manipulation by teachers, changes in the evaluation system, and changes in teachers' job assignments. Our estimates are robust to changes in the rubric, different rater types, and controlling for student baseline achievement, among other things. Still, these tests are not dispositive; there are reasons to remain cautious about a causal interpretation. We find, in one setting, a weakening correlation between teacher observation scores and student test scores in the very first years of teaching. That weakening is consistent with some threats to the identifying assumptions, but it would also be consistent with changes in optimal teaching strategies as experience increases.</p> <p>Our analyses should be interpreted carefully. First, we focus on the performance of the input tasks covered by classroom observation rubrics. Stronger assumptions are required when using observation ratings to make inferences about teacher performance measured by contributions to student outcomes. Second, taking differences in scores over time addresses many concerns. But several of those concerns would remain when making claims based on score levels at a single point in time. Third, the various tests in "Alternative Explanations and Threats to Casual Inference" are intended to be informative about potential threats to casual interpretation, but those tests are not dispositive. Partly because some test results have competing interpretations, as with Figure 6, and partly because we have data from specific settings. We hope future work in other settings will repeat these tests and add new ones.</p> <p>A final note of caution is that our estimates using data from Tennessee and DCPS may differ from estimates in other settings employing teacher observations. Our identification strategy—using differences in scores over time and between early career and veteran teachers—can be applied to other settings. However, the implementation of observations in other settings may open those systems to violations of the identifying assumptions explained and explored here.</p> <p>Our own prediction is that our results will generalize well to other settings. The Tennessee and DCPS settings are quite similar to other settings on many design features: the detail and content of the observation rubric, the frequency and duration of observation visits, the use of school principals as observers, and others (Kraft &amp; Gilmour, [<reflink idref="bib41" id="ref79">41</reflink>]; Steinberg &amp; Donaldson, [<reflink idref="bib56" id="ref80">56</reflink>]). Though other features differ: the formal incentives attached to teachers' observation ratings, DCPS's use of master educators as raters, and others. Comparisons of the actual data produced by different systems—including variation in scores and reliability—are scarce and quite limited. In general, teacher observation ratings have moderate to low reliability (Ho &amp; Kane, [<reflink idref="bib29" id="ref81">29</reflink>]; Kane &amp; Staiger, [<reflink idref="bib38" id="ref82">38</reflink>]; Taylor, [<reflink idref="bib58" id="ref83">58</reflink>]), but there is no evidence that Tennessee and DCPS systems produce scores which are more (or less) reliable than other systems. Tennessee and DCPS were early adopters of the redesigned teacher evaluation systems of the last decade, and our estimates may be more relevant to well established systems. The DCPS system provided strong and atypical financial incentives, which may have motivated teachers to improve their skills more quickly.</p> <p>In general, other settings may have their own new sources of error and bias in observation scores—the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0392" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> component—when those scores are used to measure teacher performance. But estimates of the returns to experience will be robust to those sources of error and bias under Assumption 1, that <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0393" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\nu }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> does not vary with teacher experience. One example from our study illustrates this point. It is plausible that the DCPS master educators might produce more reliable scores than school principals, but our estimates of the returns to experience were similar for both types of raters.</p> <p>Finally, we see two ways for our work to benefit teacher policy. First, a better understanding of how experience causes improvements in teaching provides a necessary foundation for several questions relevant to policy design. For example, does early‐career development depend on formal training, either in teacher certification programs or professional development for new teachers? Or does early‐career development come though learning by doing? Do teachers improve differentially across the various tasks of teaching, like managing student behavior, planning, or instruction? Is some minimal level of expertise in certain tasks a prerequisite for the development of other tasks? Answering these questions could provide useful insights to school managers and policymakers developing policies and practices intended to help teachers develop new skills. However, without credible evidence that measured performance improvements reflect true performance improvements, the resulting policy insights disappear.</p> <p>Second, a better understanding of how experience causes improvements in teaching is a critical input to implementing existing policies. For example, many states and districts now use classroom observation ratings to inform tenure decisions. Optimizing performance‐based tenure decisions requires understanding how teacher performance improves with experience. That improvement trajectory is relevant to deciding when to make tenure decisions. And that same improvement trajectory is also relevant to understanding the key cost of denying tenure: a dismissed teacher will be replaced by a novice new hire. In the first year, the novice new hire will likely under‐preform the dismissed teacher the novice replaced, but the costs and benefits accrue over an entire career not just 1 year. Similarly, how teacher performance changes over time is relevant to designing teacher compensation, including bonuses or salary increases linked to observation ratings.</p> <p>Whether and how teachers learn new skills is central to education policy decisions about selecting and investing in teachers. Our focus in this paper is estimating the returns to experience—improvements in performance caused by teaching experience. But that focus is motivated by the underlying teacher policy decisions and debates. We provide evidence that the gains in classroom observation ratings, for Tennessee and Washington, DC, teachers, do reflect stronger teaching performance resulting from early career teaching experience.</p> <hd id="AN0184198818-33">ACKNOWLEDGMENTS</hd> <p>We thank the District of Columbia Public Schools, Tennessee Department of Education, and Tennessee Education Research Alliance. Generous financial support was provided by the Spencer Foundation. The authors have no relevant conflict of interest.</p> <p>GRAPH: Online Appendix</p> <ref id="AN0184198818-34"> <title> Footnotes </title> <blist> <bibl id="bib1" idref="ref35" type="bt">1</bibl> <bibtext> For reviews of the related research literature see Garrett et al. ([24]), Goe et al. ([25]), Jackson et al. ([31]), James and Wyckoff ([35]), Kane et al. ([36]), and Taylor ([58]).</bibtext> </blist> <blist> <bibl id="bib2" idref="ref7" type="bt">2</bibl> <bibtext> Some of the estimates in Kraft et al. ([43]) and Laski and Papay ([44]) are also relevant to identification threats, though neither paper discussed causal inference explicitly.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref63" type="bt">3</bibl> <bibtext> The best exception is Kraft and Gilmour ([41]), who found that scores from Tennessee may have less variation than other settings. However, the comparison is limited to the percentage of teachers who score at or above expectations on the end‐of‐year "summative" ratings. Another exception is Weisberg et al. ([60]) where the data proceed current designs, and indeed helped to spur those designs.</bibtext> </blist> <blist> <bibl id="bib4" idref="ref8" type="bt">4</bibl> <bibtext> Additionally, the causes of observation scores (un)reliability are not well understood. For example, Ho and Kane ([29]) found that a teacher's own principal (direct supervisor) produces a more reliable observation score for them than does a principal from another school. Burgess et al. ([9]) found that peer observers who received very little training produce scores as reliable as found in research using highly trained researchers as observers.</bibtext> </blist> <blist> <bibl id="bib5" idref="ref3" type="bt">5</bibl> <bibtext> The literature does include many estimates of the correlation between observation scores and teacher value‐added, which is typically much less than 0.50. In our data that correlation is 0.38 for Tennessee and 0.30 for DCPS. Appendix Table A1 reports on these estimates in detail. However, 0.38 and 0.30 are likely to underestimate the true correlation. First, there is the common attenuation because of measurement error. Second, the simple mean <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0155" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\bar{s}}&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> gives equal weight to each task <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0156" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$k$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , but it seems unlikely the elasticity of value‐added, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0157" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#956;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\mu $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , with respect to <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0158" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\theta }&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is equal for all <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0159" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$k$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . If we knew the production function for student achievement, we would likely choose un‐equal weights.</bibtext> </blist> <blist> <bibl id="bib6" idref="ref16" type="bt">6</bibl> <bibtext> Though the conventional strategy is common, in nearly all prior papers the performance measure is teachers' value‐added contributions to student test scores.</bibtext> </blist> <blist> <bibl id="bib7" idref="ref68" type="bt">7</bibl> <bibtext> It is more common in the literature to make the first year of teaching, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0172" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$exp{{r}&amp;#95;{jt}} = 0$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , the omitted category in <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0173" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$h({exp{{r}&amp;#95;{jt}}})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . We prefer to omit veterans in part to make the comparison with our main estimates easier. Nevertheless, the choice of omitted category does not change the estimates of interest obtained from fitting the two‐way FE specification in equations 3 and 4. If we reproduced the dashed line in Figure 1 from scratch, but instead with the omitted category <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0174" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$exp{{r}&amp;#95;{jt}} = 0$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , the only thing that would change is the y‐intercept value. All of the slopes between points would remain exactly as they are in Figure 1.</bibtext> </blist> <blist> <bibl id="bib8" idref="ref69" type="bt">8</bibl> <bibtext> There are alternative specifications of <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0175" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$h({exp{{r}&amp;#95;{jt}}})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> in the literature: (i) Specifying <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0176" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$h$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> as cubic in <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0177" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$exp{{r}&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , or other higher‐order polynomial, though often still with <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0178" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$exp{{r}&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> top‐coded at some point (e.g., Rockoff, [53]). (ii) Dividing <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0179" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$exp{{r}&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> into bins, e.g., 1–2, 3–4, 5–9, 10–14, 15–24, and 25+ (e.g., Harris &amp; Sass, [28]). (iii) Using the non‐standard age‐experience progressions, e.g., leaves of absence, to estimate Specification 1 without restrictions on <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0180" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$h$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> (e.g., Wiswall, [61]).</bibtext> </blist> <blist> <bibl id="bib9" idref="ref4" type="bt">9</bibl> <bibtext> This potential bias arises in part through the implicit weights in the two‐way FE estimator (de Chaisemartin &amp; D'Haultfœuille, [18]). When estimating <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0189" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\hat{\delta }}&amp;#95;e}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , the weight given to <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0190" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\hat{\delta }}&amp;#95;{et}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is increasing in <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0191" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{N}&amp;#95;{et}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0192" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;M&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{M}&amp;#95;{et}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , and <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0193" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;v&lt;/mi&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo stretchy="false"&gt;{&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;}&lt;/mo&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$var({1\{ {exp{{r}&amp;#95;{jt}} = e} \}})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . By contrast, our preferred estimates only weight by <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0194" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{N}&amp;#95;{et}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . In our current setting, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0195" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{N}&amp;#95;{et}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> and <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0196" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;M&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{M}&amp;#95;{et}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> will likely covary over time. However, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0197" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;v&lt;/mi&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo stretchy="false"&gt;{&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;}&lt;/mo&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$var({1\{ {exp{{r}&amp;#95;{jt}} = e} \}})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> will be constant across cohorts given the specification of <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0198" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$h({exp{{r}&amp;#95;{jt}}})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> .</bibtext> </blist> <blist> <bibtext> Conceptually, experience gained in the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0201" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$e$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> th year of teaching can affect performance in year <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0202" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$({e + 1})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0203" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$({e + 2})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0204" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;3&lt;/mn&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$({e + 3})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , and on into the future. Empirically, however, the effects of the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0205" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$e$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> th year on <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0206" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$({e + 2})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> performance cannot be separately identified from the effects of the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0207" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$({e + 1})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> th year on <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0208" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$({e + 2})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> performance.</bibtext> </blist> <blist> <bibtext> The specification of <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0211" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$h({exp{{r}&amp;#95;{jt}}})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is analogous to the common event study specification, which is not subject to this second source of potential bias. In other words, the conventional strategy only uses treated‐teacher data from years <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0212" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$e$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> and <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0213" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$({e - 1})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> to estimate the effect of the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0214" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$e$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> th year of experience.</bibtext> </blist> <blist> <bibtext> Briefly, first, when estimating <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0220" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\delta }&amp;#95;{et}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> for a given <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0221" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$e$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , two‐way FE uses observations from both veterans and early‐career teachers (as long as <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0222" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8800;&lt;/mo&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;{&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;}&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$exp{{r}&amp;#95;{jt}} \ne \{ {e,e - 1} \}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> ) to estimate the year effects, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0223" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#960;&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\pi }&amp;#95;t}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . But the estimation of <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0224" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#960;&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\pi }&amp;#95;t}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> must account for the fact that there are other treatments occurring—specifically, the early‐career teachers are also gaining <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0225" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\delta }&amp;#95;{({ - e})t}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . The two‐way FE estimator uses the pooled estimate <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0226" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\hat{\delta }}&amp;#95;{({ - e})}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> instead of the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0227" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$t$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> ‐specific estimate <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0228" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\hat{\delta }}&amp;#95;{({ - e})t}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . Second, accounting for other treatments also requires knowing the probability of those other treatments in the sample at time <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0229" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$t$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . In the current setting, the probability of other treatments is the proportion of teachers at each <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0230" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#8712;&lt;/mo&gt;&lt;mo stretchy="false"&gt;{&lt;/mo&gt;&lt;mrow&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mtext&gt;...&lt;/mtext&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;}&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$e \in \{ {0,1, \ldots ,\bar{e}} \}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> —the distribution of teacher experience. The two‐way FE estimator uses the pooled estimate of the probability instead of the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0231" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$t$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> ‐specific estimate.</bibtext> </blist> <blist> <bibtext> Our DCPS data begin in 2009/2010 and thus the early years coincide with the slow labor recovery following the recession. We do not see the same pattern in Tennessee where the experience distribution has been stable over the years we study.</bibtext> </blist> <blist> <bibtext> Appendix B provides details of our value‐added estimation methods.</bibtext> </blist> <blist> <bibtext> In DCPS classroom observations account for 75% of overall IMPACT scores for the more than 80% of teachers without a value‐added score. For teacher with value added as part of their evaluation, observations account for between 30% and 40% depending on the year. In Tennessee, classroom observations are 50% and 85% of the overall TEAM score for teachers with and without value‐added scores, respectively.</bibtext> </blist> <blist> <bibtext> Not all Tennessee districts use the TEAM rubric, but our analysis in this paper uses only data from the TEAM rubric.</bibtext> </blist> <blist> <bibtext> The first seven tasks align generally with the domain of instruction, while the final two align with the domains of classroom management and environment.</bibtext> </blist> <blist> <bibtext> All appendices are available at the end of this article as it appears in JPAM online. Go to the publisher's website and use the search engine to locate the article at <ulink href="http://onlinelibrary.wiley.com">http://onlinelibrary.wiley.com</ulink>.</bibtext> </blist> <blist> <bibtext> In practice, to facilitate standard error estimation, we estimate the many individual <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0070" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\hat{\delta }}&amp;#95;{et}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> terms simultaneously in one system of regressions, stacking together one regression for each <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0071" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\hat{\delta }}&amp;#95;{et}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . The stack includes one regression (synonymously, one layer) for each unique combination of <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0072" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$e$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> and <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0073" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$t$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , where <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0074" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#8712;&lt;/mo&gt;&lt;mo stretchy="false"&gt;{&lt;/mo&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mtext&gt;...&lt;/mtext&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;}&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$e \in \{ {1,2, \ldots ,\bar{e}} \}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> and <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0075" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$t$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is a school year in our data. Each regression in the stack has the same simple specification, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0076" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#916;&lt;/mi&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;&amp;#945;&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#949;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;${{\Delta}}{{\bar{s}}&amp;#95;{jt}} = \alpha + \delta {{T}&amp;#95;{jt}} + {{\epsilon }&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , but the estimation sample is limited to only teacher‐by‐year observations, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0077" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$jt$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , where either (i) <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0078" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$exp{{r}&amp;#95;{j,t}} = e$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> and <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0079" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$exp{{r}&amp;#95;{j,t - 1}} = e - 1$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , or (ii) <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0080" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8805;&lt;/mo&gt;&lt;mover accent="true"&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$exp{{r}&amp;#95;{j,t - 1}} \ge \bar{e}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . The indicator variable <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0081" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;D&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;${{D}&amp;#95;{jt}} = 1$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> for group (i) and <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0082" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$ = 0$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> for group (ii). When stacked together, the specification becomes: <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0083" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#916;&lt;/mi&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#945;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;msub&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#949;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;${{\Delta}}{{\bar{s}}&amp;#95;{jt}} = {{\alpha }&amp;#95;{et}} + {{\delta }&amp;#95;{et}}{{T}&amp;#95;{jt}} + {{\epsilon }&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . To obtain <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0084" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\hat{\delta }}&amp;#95;e}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> we take the weighted average of <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0085" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#948;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\hat{\delta }}&amp;#95;{et}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> terms as shown in equation (1). This stacked approach allows us to correct for clusters (teachers) across regressions.</bibtext> </blist> <blist> <bibtext> We begin with the item‐by‐observation‐visit data recorded by observers in the original rubric units (integer scores 1 to 4 in DCPS and 1 to 5 in Tennessee). Separately for DCPS and Tennessee: (i) We standardize the item‐by‐visit ratings so that, by school year, each item is mean 0, standard deviation 1. (ii) For each teacher <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0123" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$j$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> by item by school year, we calculate the school‐year average of the standardized item‐by‐visit ratings. We then re‐standardize the item‐average scores. (iii) For each teacher <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0124" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$j$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> by school year, we average the item‐average scores to create the overall average score, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0125" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\bar{s}}&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . Finally, we again standardize the overall average scores by year.</bibtext> </blist> <blist> <bibtext> Appendix B provides details of our value‐added estimation methods.</bibtext> </blist> <blist> <bibtext> Additionally, Appendix Figure A2 reports results using student survey measures of teacher performance, and again the pattern of returns to experience is quite similar.</bibtext> </blist> <blist> <bibtext> Using equation (2) we can write: <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0237" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;v&lt;/mi&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#956;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;|&lt;/mo&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;v&lt;/mi&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#956;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;|&lt;/mo&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;v&lt;/mi&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#956;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;|&lt;/mo&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$cov({{{{\bar{s}}}&amp;#95;{jt}},{{{\hat{\mu }}}&amp;#95;{jt}}|exp{{r}&amp;#95;{jt}}}) = cov({{{\theta }&amp;#95;{jt}},{{{\hat{\mu }}}&amp;#95;{jt}}|exp{{r}&amp;#95;{jt}}}) + cov({{{\nu }&amp;#95;{jt}},{{{\hat{\mu }}}&amp;#95;{jt}}|exp{{r}&amp;#95;{jt}}})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . Assume quite plausibly that <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0238" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;v&lt;/mi&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#956;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;|&lt;/mo&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;v&lt;/mi&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#956;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$cov({{{\theta }&amp;#95;{jt}},{{{\hat{\mu }}}&amp;#95;{jt}}|exp{{r}&amp;#95;{jt}}}) = cov({{{\theta }&amp;#95;{jt}},{{{\hat{\mu }}}&amp;#95;{jt}}})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . In plain language, this assumption requires that the education production process—which turns teaching input tasks, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0239" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\theta }&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , into a teacher's value‐added contributions, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0240" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#956;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\hat{\mu }}&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> —does not depend on experience. Under this assumption, the prediction that <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0241" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;v&lt;/mi&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#956;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;|&lt;/mo&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;v&lt;/mi&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#957;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#956;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$cov({{{\nu }&amp;#95;{jt}},{{{\hat{\mu }}}&amp;#95;{jt}}|exp{{r}&amp;#95;{jt}}}) = cov({{{\nu }&amp;#95;{jt}},{{{\hat{\mu }}}&amp;#95;{jt}}})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> implies the further prediction that <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0242" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;v&lt;/mi&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#956;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;|&lt;/mo&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;v&lt;/mi&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#956;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$cov({{{{\bar{s}}}&amp;#95;{jt}},{{{\hat{\mu }}}&amp;#95;{jt}}|exp{{r}&amp;#95;{jt}}}) = cov({{{{\bar{s}}}&amp;#95;{jt}},{{{\hat{\mu }}}&amp;#95;{jt}}})$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> .</bibtext> </blist> <blist> <bibtext> The estimation details for Figure 3 are summarized in its note and described in Appendix B. The solid line uses only within‐teacher over‐time variation (by including teacher FE in the estimation), and the dashed line uses both within‐ and between‐teacher variation (by omitting teacher FE). To get a sense of the correlation between observation scores and value added, multiply the y‐axis by about 5 for Tennessee and 3 for DCPS. Additionally, Appendix Table A1 provides complementary evidence. That table reports the average relationship between observation scores and value added, but that average relationship does not change when we control for experience.</bibtext> </blist> <blist> <bibtext> An alternative explanation is the following: We assumed that the production process which turns teaching tasks, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0251" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\theta }&amp;#95;{jk}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , into value added output, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0252" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#956;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\hat{\mu }}&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , does not depend on experience. Even if the production process is constant, the way in which a teacher chooses to optimize that production process may depend on experience. For example, perhaps as early‐career teachers gain experience they shift more effort to tasks which are not measured by the observation rubric, or more subtly shift effort across tasks in a way not well captured by the simple average of ratings, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0253" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\bar{s}}&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> .</bibtext> </blist> <blist> <bibtext> Figure 4 suggests there may be continued returns to experience beyond a teacher's first decade (beyond <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0280" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;9&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$e = 9$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> ). For example, in Tennessee, the average observation score at 11 years of prior experience is statistically significantly greater than at 9 years but less than at 14 or 19 years. In DCPS the estimates are much nosier. Such continued returns would, strictly speaking, violate Assumption 1. However, first, the continued returns would generate downward bias in our estimates, leading us to understate the returns to experience. Second, the magnitude of bias is empirically small, for example, 0.346 versus 0.352. The bias is small for two reasons: (i) The gains after the first decade are quite small in relative terms. Teacher scores improve more than 1 full standard deviation in the first decade, but only another 10% of a standard deviation the next 5 years. (ii) Observations with <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0281" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$e$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 10–14 are a small share of all observations with <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0282" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;&amp;#62;&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$e &amp;#62; $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 9.</bibtext> </blist> <blist> <bibtext> In DCPS, compositional differences in the teaching force over time (Dee et al., [19]; Dee &amp; Wyckoff, [20]; James &amp; Wyckoff, [35]) could make it appear, with our preferred within‐year standardization process, as if experienced teachers were declining over time as the average performance of incoming teachers improves. However, relying on alternative standardization approaches, including standardizing relative to veteran teachers within year and standardizing scores across years, do not change the slopes shown in Figure 1. Differences in point estimates across standardization approaches never exceed 0.037, with an average difference in point estimates across approaches and levels of experiences of 0.005. In rubric units, the average first difference for veteran teachers is also quite small, at <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0290" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$ - $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 0.014 (st.err. 0.003).</bibtext> </blist> <blist> <bibtext> Our language and examples in this discussion mainly imply the evaluation systems designed or used by schools, districts, or states. The features and reasoning also apply to scores collected by researchers or for other purposes.</bibtext> </blist> <blist> <bibtext> On additional note on rater behavior. As described in the section "Estimation Methods," the item level observation scores for specific tasks <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0326" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{s}&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> are strongly correlated, in these data and most teacher observation data. This fact is sometimes interpreted as evidence that raters do not actually differentiate between tasks, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0327" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$k$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , but instead score teachers on some single general dimension of teaching performance. This seems unlikely given that the item level correlations are not equal to one. A more plausible explanation is that the rubrics define tasks where true performance is in fact strongly correlated. Whatever the explanation, this issue is not central to our analysis in this paper which focuses on the average score. This issue does limit our ability to make conclusions about how experience may affect tasks differentially.</bibtext> </blist> <blist> <bibtext> More precisely, tenure requires being rated "4. Effective" or "5. Highly Effective" on the 1 to 5 integer scale. While only one input to that overall final rating, classroom observation scores get a weight of 50% to 85% for the teachers.</bibtext> </blist> <blist> <bibtext> Also studying DCPS, Adnot ([1]) reported evidence that teachers facing the 2‐consecutive‐years‐minimally‐effective dismissal threat shift effort across tasks within the rubric toward tasks which are more likely raise their scores. This is a sort of distortion within measured tasks but suggests that teachers are aware of this margin.</bibtext> </blist> <blist> <bibtext> Empirical examples of manipulation by teachers include cheating on student tests (Jacob &amp; Levitt, [33]) and intentionally excluding low‐scoring students from high‐stakes tests (Cullen &amp; Reback, [14]; Figlio, [22]; Figlio &amp; Getzler, [23]; Jacob, [32]).</bibtext> </blist> <blist> <bibtext> This assumption is sufficient but not strictly necessary. We only require that the product <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0368" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;msup&gt;&lt;mi mathvariant="normal"&gt;&amp;#916;&lt;/mi&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$p{{{{\Delta}}}^a}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> not depend on experience, which should be a weaker assumption.</bibtext> </blist> <blist> <bibtext> The estimation for Figure 8 is identical to our preferred strategy used in Figure 1 with two exceptions. First, we limit the sample to teacher‐by‐year, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0373" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$jt$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , observations where we have prior‐year test scores for students assigned to the teacher, grades 4 through 8 math and language classes. Second, for the dashed line, the outcome variable is the residual from a regression of observation score, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:02768739:media:pam22584:pam22584-math-0374" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\bar{s}}&amp;#95;{jt}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , on the average prior‐year test score for students assigned to the teacher.</bibtext> </blist> <blist> <bibtext> This subtraction might be desirable in specific cases. Imagine, for example, that veterans improved because of some new training, and that training was given to all teachers, early‐career and veteran. If, roughly, the effect of the training was similar for all teachers, then the subtraction makes the Figure 1 estimates returns to experience controlling for any general training effects.</bibtext> </blist> </ref> <ref id="AN0184198818-35"> <title> REFERENCES </title> <blist> <bibtext> Adnot, M. (2016). Teacher evaluation, instructional practice, and student achievement: Evidence from the District of Columbia Public Schools and the Measures of Effective Teaching Project [ Doctoral dissertation ]. University of Virginia. ProQuest Dissertations Publishing. https://doi.org/10.18130/V31C72</bibtext> </blist> <blist> <bibtext> Altonji, J. G., &amp; Williams, N. (1992). The effects of labor market experience, job seniority, and job mobility on wage growth [Working paper no. 4133]. National Bureau of Economic Research. https://doi.org/10.3386/w4133</bibtext> </blist> <blist> <bibtext> Anderson, J. (2013, March 30). Curious grade for teachers: Nearly all pass. New York Times. https://<ulink href="http://www.nytimes.com/2013/03/31/education/curious&amp;#8208;grade&amp;#8208;for&amp;#8208;teachers&amp;#8208;nearly&amp;#8208;all&amp;#8208;pass.html">www.nytimes.com/2013/03/31/education/curious&amp;#8208;grade&amp;#8208;for&amp;#8208;teachers&amp;#8208;nearly&amp;#8208;all&amp;#8208;pass.html</ulink></bibtext> </blist> <blist> <bibtext> Angrist, J. D. (1990). Lifetime earnings and the Vietnam era draft lottery: Evidence from social security administrative records. American Economic Review, 80 (3), 313 – 336.</bibtext> </blist> <blist> <bibtext> Araujo, M. C., Carneiro, P., Cruz‐Aguayo, Y., &amp; Schady, N. (2016). Teacher quality and learning outcomes in kindergarten. Quarterly Journal of Economics, 131 (3), 1415 – 1453. https://doi.org/10.1093/qje/qjw016</bibtext> </blist> <blist> <bibtext> Atteberry, A., Loeb, S., &amp; Wyckoff, J. (2015). Do first impressions matter? Predicting early career teacher effectiveness. AERA Open, 1 (4), 1 – 23. https://doi.org/10.1177/2332858415607834</bibtext> </blist> <blist> <bibtext> Briole, S., &amp; Maurin, E. (2022). There's always room for improvement: The persistent benefits of repeated teacher evaluations. Journal of Human Resources, Advance online publication. https://doi.org/10.3368/jhr.1220‐11370R1</bibtext> </blist> <blist> <bibtext> Burgess, S., Rawal, S., &amp; Taylor, E. S. (2021). Teacher peer observation and student test scores: Evidence from a field experiment in English secondary schools. Journal of Labor Economics, 39 (4), 1155 – 1186. https://doi.org/10.1086/712997</bibtext> </blist> <blist> <bibtext> Burgess, S., Rawal, S., &amp; Taylor, E. S. (2023). Teachers' use of class time and student achievement. Economics of Education Review, 94, 102405. https://doi.org/10.1016/j.econedurev.2023.102405</bibtext> </blist> <blist> <bibtext> Campbell, S. L., &amp; Ronfeldt, M. (2018). Observational evaluation of teachers: Measuring more than we bargained for? American Educational Research Journal, 55 (6), 1233 – 1267. https://doi.org/10.3102/0002831218776216</bibtext> </blist> <blist> <bibtext> Chetty, R., Friedman, J. N., &amp; Rockoff, J. E. (2014b). Measuring the impacts of teachers II: Teacher value‐added and student outcomes in adulthood. American Economic Review, 104 (9), 2633 – 2679. https://doi.org/10.1257/aer.104.9.2633</bibtext> </blist> <blist> <bibtext> Chi, O. L. (2023). A classroom observer like me: The effects of race‐congruence and gender‐congruence between teachers and raters on observation scores. Education Finance and Policy, 18 (3), 442 – 466. https://doi.org/10.1162/edfp_a_00367</bibtext> </blist> <blist> <bibtext> Cohen, J., &amp; Goldhaber, D. (2016). Building a more complete understanding of teacher evaluation using classroom observations. Educational Researcher, 45 (6), 378 – 387. https://doi.org/10.3102/0013189X16659442</bibtext> </blist> <blist> <bibtext> Cullen, J. B., &amp; Reback, R. (2006). Tinkering toward accolades: School gaming under a performance accountability system. In T. J. Gronberg &amp; D. W. Jansen (Eds.), Advances in applied microeconomics (Vol. 14, pp. 1 – 34). Emerald Group Publishing Limited. https://doi.org/10.1016/S0278‐0984(06)14001‐8</bibtext> </blist> <blist> <bibtext> Danielson, C. (1997). Enhancing professional practice: A framework for teaching. Association for Supervision and Curriculum Development.</bibtext> </blist> <blist> <bibtext> de Chaisemartin, C., &amp; D'Haultfœuille, X. (2020). Two‐way fixed effects estimators with heterogeneous treatment effects. American Economic Review, 110 (9), 2964 – 2996. https://doi.org/10.1257/aer.20181169</bibtext> </blist> <blist> <bibtext> de Chaisemartin, C., &amp; D'Haultfœuille, X. (2023a). Two‐way fixed effects and difference‐in‐differences estimators with several treatments. Journal of Econometrics, 236 (2), 105480. https://doi.org/10.1016/j.jeconom.2023.105480</bibtext> </blist> <blist> <bibtext> de Chaisemartin, C., &amp; D'Haultfœuille, X. (2023b). Two‐way fixed effects and differences‐in‐differences with heterogeneous treatment effects: A survey. The Econometrics Journal, 26 (3), C1 – C30. https://doi.org/10.1093/ectj/utac017</bibtext> </blist> <blist> <bibtext> Dee, T. S., James, J., &amp; Wyckoff, J. (2021). Is effective teacher evaluation sustainable? Evidence from DCPS. Education Finance and Policy, 16 (2), 313 – 346. https://doi.org/10.1162/edfp_a_00303</bibtext> </blist> <blist> <bibtext> Dee, T. S., &amp; Wyckoff, J. (2015). Incentives, selection, and teacher performance: Evidence from impact. Journal of Policy Analysis and Management, 34 (2), 267 – 297. https://doi.org/10.1002/pam.21818</bibtext> </blist> <blist> <bibtext> Ferguson, R. F., &amp; Danielson, C. (2015). How Framework for Teaching and Tripod 7Cs evidence distinguish key components of effective teaching. In T. J. Kane, K. A. Kerr, &amp; R. C. Pianta (Eds.), Designing teacher evaluation systems: New guidance from the Measures of Effective Teaching Project (pp. 98 – 143). Jossey‐Bass. https://doi.org/10.1002/9781119210856.ch4</bibtext> </blist> <blist> <bibtext> Figlio, D. N. (2006). Testing, crime and punishment. Journal of Public Economics, 90 (4–5), 837 – 851. https://doi.org/10.1016/j.jpubeco.2005.01.003</bibtext> </blist> <blist> <bibtext> Figlio, D. N., &amp; Getzler, L. (2006). Accountability, ability, and disability: Gaming the system? In T. J. Gronberg &amp; D. W. Jansen (Eds.), Advances in applied microeconomics (Vol. 14, pp. 35 – 49). Emerald Group Publishing Limited. https://doi.org/10.1016/S0278‐0984(06)14002‐X</bibtext> </blist> <blist> <bibtext> Garrett, R., Citkowicz, M., &amp; Williams, R. (2019). How responsive is a teacher's classroom practice to intervention? A meta‐analysis of randomized field studies. Review of Research in Education, 43 (1), 106 – 137. https://doi.org/10.3102/0091732X19830634</bibtext> </blist> <blist> <bibtext> Goe, L., Bell, C., &amp; Little, O. (2008). Approaches to evaluating teacher effectiveness: A research synthesis. National Comprehensive Center for Teacher Quality, ETS.</bibtext> </blist> <blist> <bibtext> Grissom, J. A., &amp; Bartanen, B. (2019). Strategic retention: Principal effectiveness and teacher turnover in multiple‐measure teacher evaluation systems. American Educational Research Journal, 56 (2), 514 – 555. https://doi.org/10.3102/0002831218797931</bibtext> </blist> <blist> <bibtext> Grogger, J. (2009). Welfare reform, returns to experience, and wages: Using reservation wages to account for sample selection bias. Review of Economics and Statistics, 91 (3), 490 – 502. https://doi.org/10.1162/rest.91.3.490</bibtext> </blist> <blist> <bibtext> Harris, D. N., &amp; Sass, T. R. (2011). Teacher training, teacher quality and student achievement. Journal of Public Economics, 95 (7–8), 798 – 812. https://doi.org/10.1016/j.jpubeco.2010.11.009</bibtext> </blist> <blist> <bibtext> Ho, A. D., &amp; Kane, T. J. (2013). The reliability of classroom observations by school personnel. Bill &amp; Melinda Gates Foundation.</bibtext> </blist> <blist> <bibtext> Holmstrom, B., &amp; Milgrom, P. (1991). Multitask principal‐agent analyses: Incentive contracts, asset ownership, and job design. Journal of Law, Economics, &amp; Organization, 7 (1), 24 – 52. https://doi.org/10.1093/jleo/7.special_issue.24</bibtext> </blist> <blist> <bibtext> Jackson, C. K., Rockoff, J. E., &amp; Staiger, D. O. (2014). Teacher effects and teacher‐related policies. Annual Review of Economics, 6 (1), 801 – 825. https://doi.org/10.1146/annurev‐economics‐080213‐040845</bibtext> </blist> <blist> <bibtext> Jacob, B. A. (2005). Accountability, incentives and behavior: The impact of high‐stakes testing in the Chicago Public Schools. Journal of Public Economics, 89 (5–6), 761 – 796. https://doi.org/10.1016/j.jpubeco.2004.08.004</bibtext> </blist> <blist> <bibtext> Jacob, B. A., &amp; Levitt, S. D. (2003). Rotten apples: An investigation of the prevalence and predictors of teacher cheating. Quarterly Journal of Economics, 118 (3), 843 – 877. https://doi.org/10.1162/00335530360698441</bibtext> </blist> <blist> <bibtext> Jacob, B. A., Rockoff, J. E., Taylor, E. S., Lindy, B., &amp; Rosen, R. (2018). Teacher applicant hiring and teacher performance: Evidence from DC Public Schools. Journal of Public Economics, 166, 81 – 97. https://doi.org/10.1016/j.jpubeco.2018.08.011</bibtext> </blist> <blist> <bibtext> James, J., &amp; Wyckoff, J. (2020). Teacher labor markets: An overview. In S. Bradley &amp; C. Green (Eds.), The economics of education (2nd ed., pp. 355 – 370). Elsevier.</bibtext> </blist> <blist> <bibtext> Kane, T. J., Kerr, K., &amp; Pianta, R. (2014). Designing teacher evaluation systems: New guidance from the Measures of Effective Teaching Project. Jossey‐Bass. https://doi.org/10.1002/9781119210856</bibtext> </blist> <blist> <bibtext> Kane, T. J., McCaffrey, D. F., Miller, T., &amp; Staiger, D. O. (2013). Have we identified effective teachers? Validating measures of effective teaching using random assignment. Bill &amp; Melinda Gates Foundation.</bibtext> </blist> <blist> <bibtext> Kane, T. J., &amp; Staiger, D. O. (2012). Gathering feedback for teaching: Combining high‐quality observations with student surveys and achievement gains. Bill &amp; Melinda Gates Foundation.</bibtext> </blist> <blist> <bibtext> Kane, T. J., Taylor, E. S., Tyler, J. H., &amp; Wooten, A. L. (2011). Identifying effective classroom practices using student achievement data. Journal of Human Resources, 46 (3), 587 – 613. https://doi.org/10.3368/jhr.46.3.587</bibtext> </blist> <blist> <bibtext> Kraft, M. A., Blazar, D., &amp; Hogan, D. (2018). The effect of teacher coaching on instruction and achievement: A meta‐analysis of the causal evidence. Review of Educational Research, 88 (4), 547 – 588. https://doi.org/10.3102/0034654318759268</bibtext> </blist> <blist> <bibtext> Kraft, M. A., &amp; Gilmour, A. F. (2017). Revisiting the widget effect: Teacher evaluation reforms and the distribution of teacher effectiveness. Educational Researcher, 46 (5), 234 – 249. https://doi.org/10.3102/0013189X17718797</bibtext> </blist> <blist> <bibtext> Kraft, M. A., &amp; Papay, J. P. (2014). Can professional environments in schools promote teacher development? Explaining heterogeneity in returns to teaching experience. Educational Evaluation and Policy Analysis, 36 (4), 476 – 500. https://doi.org/10.3102/0162373713519496</bibtext> </blist> <blist> <bibtext> Kraft, M. A., Papay, J. P., &amp; Chi, O. L. (2020). Teacher skill development: Evidence from performance ratings by principals. Journal of Policy Analysis and Management, 39 (2), 315 – 347. https://doi.org/10.1002/pam.22193</bibtext> </blist> <blist> <bibtext> Laski, M., &amp; Papay, J. (2020). Understanding the dynamics of teacher productivity development: Evidence on teacher improvement in Tennessee. Unpublished manuscript.</bibtext> </blist> <blist> <bibtext> Leaver, C., Ozier, O., Serneels, P., &amp; Zeitlin, A. (2021). Recruitment, effort, and retention effects of performance contracts for civil servants: Experimental evidence from Rwandan primary schools. American Economic Review, 111 (7), 2213 – 2246. https://doi.org/10.1257/aer.20191972</bibtext> </blist> <blist> <bibtext> Nagler, M., Piopiunik, M., &amp; West, M. R. (2020). Weak markets, strong teachers: Recession at career start and teacher effectiveness. Journal of Labor Economics, 38 (2), 453 – 500. https://doi.org/10.1086/705883</bibtext> </blist> <blist> <bibtext> Ost, B. (2014). How do teachers improve? The relative importance of specific and general human capital. American Economic Journal: Applied Economics, 6 (2), 127 – 151. https://doi.org/10.1257/app.6.2.127</bibtext> </blist> <blist> <bibtext> Papay, J. P., &amp; Kraft, M. A. (2015). Productivity returns to experience in the teacher labor market: Methodological challenges and new evidence on long‐term career improvement. Journal of Public Economics, 130, 105 – 119. https://doi.org/10.1016/j.jpubeco.2015.02.008</bibtext> </blist> <blist> <bibtext> Phipps, A. R. (2018). Incentive contracts in complex environments: Theory and evidence on effective teacher performance incentives [ Doctoral dissertation ]. University of Virginia. ProQuest Dissertations Publishing. https://doi.org/10.18130/V3610VR6M</bibtext> </blist> <blist> <bibtext> Phipps, A. R., &amp; Wiseman, E. A. (2021). Enacting the rubric: Teacher improvements in windows of high‐stakes observation. Education Finance and Policy, 16 (2), 283 – 312. https://doi.org/10.1162/edfp_a_00295</bibtext> </blist> <blist> <bibtext> Prendergast, C. (1999). The provision of incentives in firms. Journal of Economic Literature, 37 (1), 7 – 63. https://doi.org/10.1257/jel.37.1.7</bibtext> </blist> <blist> <bibtext> Rivkin, S. G., Hanushek, E. A., &amp; Kain, J. F. (2005). Teachers, schools, and academic achievement. Econometrica, 73 (2), 417 – 458. https://doi.org/10.1111/j.1468‐0262.2005.00584.x</bibtext> </blist> <blist> <bibtext> Rockoff, J. E. (2004). The impact of individual teachers on student achievement: Evidence from panel data. American Economic Review, 94 (2), 247 – 252. https://doi.org/10.1257/0002828041302244</bibtext> </blist> <blist> <bibtext> Rockoff, J. E. (2008). Does mentoring reduce turnover and improve skills of new employees? Evidence from teachers in New York City [ Working paper no. 13868 ]. National Bureau of Economic Research. https://doi.org/10.3386/w13868</bibtext> </blist> <blist> <bibtext> Roth, J., Sant'Anna, P. H. C., Bilinski, A., &amp; Poe, J. (2023). What's trending in difference‐in‐differences? A synthesis of the recent econometrics literature. Journal of Econometrics, 235 (2), 2218 – 2244. https://doi.org/10.1016/j.jeconom.2023.03.008</bibtext> </blist> <blist> <bibtext> Steinberg, M. P., &amp; Donaldson, M. L. (2016). The new educational accountability: Understanding the landscape of teacher evaluation in the post‐NCLB era. Education Finance and Policy, 11 (3), 340 – 359. https://doi.org/10.1162/EDFP_a_00186</bibtext> </blist> <blist> <bibtext> Steinberg, M. P., &amp; Kraft, M. A. (2017). The sensitivity of teacher performance ratings to the design of teacher evaluation systems. Educational Researcher, 46 (7), 378 – 396. https://doi.org/10.3102/0013189X17726752</bibtext> </blist> <blist> <bibtext> Taylor, E. (2023). Teacher evaluation and training. In E. A. Hanushek, S. Machin, &amp; L. Woessmann (Eds.), The handbook of the economics of education (Vol. 7, pp. 61 – 141). Elsevier. https://doi.org/10.1016/bs.hesedu.2023.03.002</bibtext> </blist> <blist> <bibtext> Taylor, E., &amp; Tyler, J. (2012). The effect of evaluation on teacher performance. American Economic Review, 102 (7), 3628 – 3651. https://doi.org/10.1257/aer.102.7.3628</bibtext> </blist> <blist> <bibtext> Weisberg, D., Sexton, S., Mulhern, J., Keeling, D., Schunck, J., Palcisco, A., &amp; Morgan, K. (2009). The widget effect: Our national failure to acknowledge and act on differences in teacher effectiveness. The New Teacher Project.</bibtext> </blist> <blist> <bibtext> Wiswall, M. (2013). The dynamics of teacher quality. Journal of Public Economics, 100, 61 – 78. https://doi.org/10.1016/j.jpubeco.2013.01.006</bibtext> </blist> </ref> <aug> <p>By Courtney Bell; Jessalynn James; Eric S. Taylor and James Wyckoff</p> <p>Reported by Author; Author; Author; Author</p> <p></p> <p>Courtney Bell is Professor of Learning Sciences and Director of the Wisconsin Center for Education Research, University of Wisconsin–Madison, 1025 West Johnson Street, Madison, WI 53706‐1706 (email: Courtney.bell@wisc.edu).</p> <p>Jessalynn James is Analytics Director at TNTP, 500 7th Avenue, 8th Floor, New York, NY 10018 (email: jessalynn.james@tntp.org).</p> <p>Eric S. Taylor is an Associate Professor at Harvard University, and a Faculty Research Fellow at NBER, Gutman Library 469, 6 Appian Way, Cambridge, MA 02138 (email: eric_taylor@gse.harvard.edu).</p> <p>James Wyckoff is Memorial Professor of Education and Policy at University of Virginia, PO Box 400265, Charlottesville, VA 22904 (email: wyckoff@virginia.edu).</p> </aug> <nolink nlid="nl1" bibid="bib51" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib16" firstref="ref2"></nolink> <nolink nlid="nl3" bibid="bib39" firstref="ref5"></nolink> <nolink nlid="nl4" bibid="bib37" firstref="ref6"></nolink> <nolink nlid="nl5" bibid="bib27" firstref="ref9"></nolink> <nolink nlid="nl6" bibid="bib47" firstref="ref10"></nolink> <nolink nlid="nl7" bibid="bib52" firstref="ref11"></nolink> <nolink nlid="nl8" bibid="bib53" firstref="ref12"></nolink> <nolink nlid="nl9" bibid="bib58" firstref="ref13"></nolink> <nolink nlid="nl10" bibid="bib18" firstref="ref14"></nolink> <nolink nlid="nl11" bibid="bib55" firstref="ref15"></nolink> <nolink nlid="nl12" bibid="bib42" firstref="ref17"></nolink> <nolink nlid="nl13" bibid="bib43" firstref="ref19"></nolink> <nolink nlid="nl14" bibid="bib44" firstref="ref20"></nolink> <nolink nlid="nl15" bibid="bib41" firstref="ref21"></nolink> <nolink nlid="nl16" bibid="bib57" firstref="ref22"></nolink> <nolink nlid="nl17" bibid="bib10" firstref="ref23"></nolink> <nolink nlid="nl18" bibid="bib12" firstref="ref24"></nolink> <nolink nlid="nl19" bibid="bib13" firstref="ref25"></nolink> <nolink nlid="nl20" bibid="bib26" firstref="ref26"></nolink> <nolink nlid="nl21" bibid="bib15" firstref="ref27"></nolink> <nolink nlid="nl22" bibid="bib60" firstref="ref29"></nolink> <nolink nlid="nl23" bibid="bib19" firstref="ref30"></nolink> <nolink nlid="nl24" bibid="bib20" firstref="ref31"></nolink> <nolink nlid="nl25" bibid="bib21" firstref="ref32"></nolink> <nolink nlid="nl26" bibid="bib11" firstref="ref36"></nolink> <nolink nlid="nl27" bibid="bib48" firstref="ref47"></nolink> <nolink nlid="nl28" bibid="bib34" firstref="ref50"></nolink> <nolink nlid="nl29" bibid="bib45" firstref="ref51"></nolink> <nolink nlid="nl30" bibid="bib46" firstref="ref52"></nolink> <nolink nlid="nl31" bibid="bib40" firstref="ref53"></nolink> <nolink nlid="nl32" bibid="bib54" firstref="ref54"></nolink> <nolink nlid="nl33" bibid="bib17" firstref="ref55"></nolink> <nolink nlid="nl34" bibid="bib29" firstref="ref66"></nolink> <nolink nlid="nl35" bibid="bib30" firstref="ref67"></nolink> <nolink nlid="nl36" bibid="bib59" firstref="ref70"></nolink> <nolink nlid="nl37" bibid="bib49" firstref="ref71"></nolink> <nolink nlid="nl38" bibid="bib50" firstref="ref72"></nolink> <nolink nlid="nl39" bibid="bib56" firstref="ref80"></nolink> <nolink nlid="nl40" bibid="bib38" firstref="ref82"></nolink> |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1456394 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Measuring Returns to Experience Using Supervisor Ratings of Observed Performance: The Case of Classroom Teachers – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Courtney+Bell%22">Courtney Bell</searchLink><br /><searchLink fieldCode="AR" term="%22Jessalynn+James%22">Jessalynn James</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-6765-2642">0000-0001-6765-2642</externalLink>)<br /><searchLink fieldCode="AR" term="%22Eric+S%2E+Taylor%22">Eric S. Taylor</searchLink><br /><searchLink fieldCode="AR" term="%22James+Wyckoff%22">James Wyckoff</searchLink> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Journal+of+Policy+Analysis+and+Management%22"><i>Journal of Policy Analysis and Management</i></searchLink>. 2025 44(1):12-44. – Name: Avail Label: Availability Group: Avail Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 33 – Name: DatePubCY Label: Publication Date Group: Date Data: 2025 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Evaluative – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Lesson+Observation+Criteria%22">Lesson Observation Criteria</searchLink><br /><searchLink fieldCode="DE" term="%22Teaching+Experience%22">Teaching Experience</searchLink><br /><searchLink fieldCode="DE" term="%22Teacher+Evaluation%22">Teacher Evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Supervisors%22">Supervisors</searchLink><br /><searchLink fieldCode="DE" term="%22Supervisory+Methods%22">Supervisory Methods</searchLink><br /><searchLink fieldCode="DE" term="%22Teachers%22">Teachers</searchLink><br /><searchLink fieldCode="DE" term="%22Statistical+Bias%22">Statistical Bias</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Reliability%22">Test Reliability</searchLink><br /><searchLink fieldCode="DE" term="%22Context+Effect%22">Context Effect</searchLink><br /><searchLink fieldCode="DE" term="%22Robustness+%28Statistics%29%22">Robustness (Statistics)</searchLink> – Name: Subject Label: Geographic Terms Group: Su Data: <searchLink fieldCode="DE" term="%22Tennessee%22">Tennessee</searchLink><br /><searchLink fieldCode="DE" term="%22District+of+Columbia%22">District of Columbia</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1002/pam.22584 – Name: ISSN Label: ISSN Group: ISSN Data: 0276-8739<br />1520-6688 – Name: Abstract Label: Abstract Group: Ab Data: We study the returns to experience in teaching, estimated using supervisor ratings from classroom observations. We describe the assumptions required to interpret changes in observation ratings over time as the causal effect of experience on performance. We compare two difference-in-differences strategies: the two-way fixed effects estimator common in the literature, and an alternative which avoids potential bias arising from effect heterogeneity. Using data from Tennessee and Washington, DC, we show empirical tests relevant to assessing the identifying assumptions and substantive threats--e.g., leniency bias, manipulation, changes in incentives or job assignments--and find our estimates are robust to several threats. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2025 – Name: AN Label: Accession Number Group: ID Data: EJ1456394 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1456394 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1002/pam.22584 Languages: – Text: English PhysicalDescription: Pagination: PageCount: 33 StartPage: 12 Subjects: – SubjectFull: Lesson Observation Criteria Type: general – SubjectFull: Teaching Experience Type: general – SubjectFull: Teacher Evaluation Type: general – SubjectFull: Supervisors Type: general – SubjectFull: Supervisory Methods Type: general – SubjectFull: Teachers Type: general – SubjectFull: Statistical Bias Type: general – SubjectFull: Test Reliability Type: general – SubjectFull: Context Effect Type: general – SubjectFull: Robustness (Statistics) Type: general – SubjectFull: Tennessee Type: general – SubjectFull: District of Columbia Type: general Titles: – TitleFull: Measuring Returns to Experience Using Supervisor Ratings of Observed Performance: The Case of Classroom Teachers Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Courtney Bell – PersonEntity: Name: NameFull: Jessalynn James – PersonEntity: Name: NameFull: Eric S. Taylor – PersonEntity: Name: NameFull: James Wyckoff IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 01 Type: published Y: 2025 Identifiers: – Type: issn-print Value: 0276-8739 – Type: issn-electronic Value: 1520-6688 Numbering: – Type: volume Value: 44 – Type: issue Value: 1 Titles: – TitleFull: Journal of Policy Analysis and Management Type: main |
| ResultId | 1 |