Growth across Grades and Common Item Grade Alignment in Vertical Scaling Using the Rasch Model
Saved in:
| Title: | Growth across Grades and Common Item Grade Alignment in Vertical Scaling Using the Rasch Model |
|---|---|
| Language: | English |
| Authors: | Sanford R. Student (ORCID |
| Source: | Educational Measurement: Issues and Practice. 2025 44(1):84-95. |
| Availability: | Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us |
| Peer Reviewed: | Y |
| Page Count: | 12 |
| Publication Date: | 2025 |
| Document Type: | Journal Articles Reports - Research |
| Education Level: | Elementary Education Junior High Schools Middle Schools Secondary Education |
| Descriptors: | Elementary School Mathematics, Elementary School Students, Middle School Mathematics, Middle School Students, Growth Models, Achievement Gains, Scaling, Rating Scales, Error of Measurement, Mathematics Achievement, Test Reliability, Test Validity, Grade Level Differences |
| DOI: | 10.1111/emip.12639 |
| ISSN: | 0731-1745 1745-3992 |
| Abstract: | Vertical scales are frequently developed using common item nonequivalent group linking. In this design, one can use upper-grade, lower-grade, or mixed-grade common items to estimate the linking constants that underlie the absolute measurement of growth. Using the Rasch model and a dataset from Curriculum Associates' i-Ready Diagnostic in math in grades 3-7, we demonstrate how grade-to-grade mean differences in mathematics proficiency appear much larger when upper-grade linking items are used instead of lower-grade items, with linkings based on a mixture of items falling in between. We then consider salient properties of the three calibrated scales including invariance of the different sets of common items to student grade and item difficulty reversals. These exploratory analyses suggest that upper-grade common items in vertical scaling are more subject to threats to score comparability across grades, even though these items also tend to imply the most growth. |
| Abstractor: | As Provided |
| Entry Date: | 2025 |
| Accession Number: | EJ1460458 |
| Database: | ERIC |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwGK_SEIhP6lAmS2Jnz1Guu4AAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDNa3Ob1IaWvWVJWXbgIBEICBmxGF0hA7KD3X1DU3jtp_OX8UdGXzz-pLLap1wnLst5EV5YGOKx2wHDLQaCRNVz9mnJwVG2znA-4SaX89M_oYYK1hGbvdE5VS7WRoqgeaneoAGYGkMyXAdm00ZGltqCXRGlmJBuhxiusQ2BvNTgBNiC6Tc1xpGwoUTVgup7R5hU6VWos8SCpqbnCmoGIolzqk_CbNKIFI_i7tWkxQ Text: Availability: 1 Value: <anid>AN0183983945;ems01mar.25;2025Mar26.05:51;v2.2.500</anid> <title id="AN0183983945-1">Growth across Grades and Common Item Grade Alignment in Vertical Scaling Using the Rasch Model </title> <p>Vertical scales are frequently developed using common item nonequivalent group linking. In this design, one can use upper‐grade, lower‐grade, or mixed‐grade common items to estimate the linking constants that underlie the absolute measurement of growth. Using the Rasch model and a dataset from Curriculum Associates' i‐Ready Diagnostic in math in grades 3–7, we demonstrate how grade‐to‐grade mean differences in mathematics proficiency appear much larger when upper‐grade linking items are used instead of lower‐grade items, with linkings based on a mixture of items falling in between. We then consider salient properties of the three calibrated scales including invariance of the different sets of common items to student grade and item difficulty reversals. These exploratory analyses suggest that upper‐grade common items in vertical scaling are more subject to threats to score comparability across grades, even though these items also tend to imply the most growth.</p> <p>Keywords: invariance; linking</p> <p>When constructing a vertical scale to measure students' growth over time, numerous decisions must be made about how to construct the scale. Prior studies have investigated how many of these decisions can lead to different conclusions about the differences in proficiency between students in different grades (e.g., Briggs &amp; Weeks, [<reflink idref="bib10" id="ref1">10</reflink>]; Tong &amp; Kolen, [<reflink idref="bib44" id="ref2">44</reflink>]; Yen, [<reflink idref="bib53" id="ref3">53</reflink>]; Yen &amp; Burket, [<reflink idref="bib54" id="ref4">54</reflink>], among many others; for an unpublished but comprehensive summary of much of this research, see Reckase, [<reflink idref="bib37" id="ref5">37</reflink>]). In this study, we investigate the effect that the grade level of the items used to link grade‐specific tests has on mean score differences between students in adjacent grades. While there appears to be an understanding in the field that grade‐to‐grade score differences, colloquially referred to as "growth" (although the differences are in fact cross‐sectional), are likely to change with different common item grade alignments (Kolen &amp; Brennan, [<reflink idref="bib26" id="ref6">26</reflink>], p. 429), the little publicly available research involving this distinction provides mixed results. With this study, we provide evidence that absent procedures that run the risk of limiting the construct representation of the linking items, common item grade alignment can produce very different estimates of grade‐to‐grade growth when applied to the exact same underlying population of examinees. This finding is based on primary analyses using the Rasch model (Rasch, [<reflink idref="bib36" id="ref7">36</reflink>]) to estimate across‐grade linking constants. We then consider the properties of, and tradeoffs between, the designs we consider, asking: what insights does this study provide about the "right way" to build a vertical scale?</p> <hd id="AN0183983945-2">Research Question</hd> <p>The fundamental purpose of a vertical scale is to take test forms with distinct, developmentally appropriate content in a common domain (such as standards‐aligned math content for an upper and a lower grade) and place the scores from those distinct forms onto a common scale to facilitate the measurement of growth. To accomplish this, overlapping sets of items—"common items"—are typically embedded in adjacent test forms so that their relative difficulty on a common scale can be estimated. This is the common item, nonequivalent group (CING) data collection design for parameter linking (Tong &amp; Kolen, [<reflink idref="bib44" id="ref8">44</reflink>]; van der Linden &amp; Barrett, [<reflink idref="bib45" id="ref9">45</reflink>]), and while it is not the only data collection design available for vertical scaling (e.g., Petersen et al., [<reflink idref="bib34" id="ref10">34</reflink>]), it currently appears to be the most popular, at least for large‐scale assessment programs in the United States (Student, [<reflink idref="bib41" id="ref11">41</reflink>]).</p> <p>The task of selecting common items introduces a unique tension into data collection for vertical scaling that is not present in other linking and equating applications such as year‐to‐year "horizontal" equating of test forms measuring a common construct, even though the mathematical models used to link vertical and horizontal scales are largely the same. The tension is this: the purpose of vertical scaling is to create a common scale for test forms that usually imply different conceptualizations of the target construct (either in terms of breadth of content or the cognitive complexity of the content). However, the common items used to calibrate a vertical scale are by definition the same. It is therefore not possible for the common items to align perfectly with the construct motivating each form, especially when this construct is being defined operationally from a grade‐specific blueprint of content by process specifications, as is typical in an achievement testing context. If it were possible, the task of selecting common items would be identical to common item selection for horizontal equating. A critical decision point for an achievement test developer is to specify the grade level alignment for the common items, and the ramifications of this decision should not be neglected. After all, those common items are the basis for figuring out the relative difficulty of different test forms, and those differences in difficulty form the basis of measuring growth in an absolute sense along a calibrated measuring scale.</p> <p>In this study, we consider three potential common item grade level alignments and their implications for interpretations of growth on a vertical scale. Growth between two time points might be a matter of improvement on material one has already had the chance to learn by that first time point; improvement on material one has had the chance to learn only after the first time point; or improvement on both types of material. When the time points are defined by students' grade level, the first conceptualization implies the use of common items aligned to the lower grade's content standards; the second conceptualization implies the use of upper‐grade items; and the third implies the use of a mixture of items with both alignments. We refer to these as lower‐grade, upper‐grade, and mixed common item alignment designs.</p> <p>Despite the consistent attention to vertical scales in research spanning the last forty years, and the fact that the measurement of growth has been of interest for almost as long as the field of psychometrics has existed (Thurstone, [<reflink idref="bib43" id="ref12">43</reflink>]), empirical studies of the grade level alignment of common items used in constructing a vertical scale have been rare. This study addresses this gap in the context of data from an existing testing program–Curriculum Associates' i‐Ready Diagnostic Assessment in math–with a vertical scale calibrated using the Rasch model. The field testing design by which new items are added to the i‐Ready assessment item bank provides an opportunity to make direct comparisons between different vertical scales constructed using sets of common items with different grade alignments.</p> <p>During the Winter 2019 administration of the i‐Ready Diagnostic mathematics assessment, Curriculum Associates administered hundreds of Embedded Field Test (EFT) items to students in each grade in the study using a design that ensures administration of every item across three grades–the grade to which the item is aligned, as well as the grades above and below–with random administration within each grade. We use these EFT items in three distinct common item nonequivalent group linking designs that use different common item grade level alignments–upper, lower, and mixed. These analyses address the question: <emph>How do the grade‐level score distributions that define a vertical scale for grade 3–7 math compare when using upper‐grade, lower‐grade or mixed‐grade common items to link the scale with the Rasch model using the same item response data?</emph> As we will show, there is evidence–at least in the content domain of mathematics–that growth will appear largest when employing upper grade common items, and smallest when employing lower grade common items. The mixed design falls predictably in between.</p> <p>Despite large differences in apparent growth, the results of this analysis alone provide no evidence for which choice of common item grade alignment is the "right" one because each reflects a distinct operationalization of the idea of "growth." The remainder of the paper considers the implications of this main finding for additional criteria by which vertical scales are often evaluated. While we do not claim that our results are generalizable to all vertical scales (or even all vertical scales covering the domain of math in grades 3–7), the questions we ask in this extended discussion are representative of the types of questions that one should ask when developing a new vertical scale. We focus on issues surrounding the extent to which item responses appear compatible with a hypothesis of year‐over‐year growth in the target domain and the importance of content expertise in deciding which common item grade alignment to use as the foundation of one's vertical scale using CING.</p> <hd id="AN0183983945-3">Background</hd> <p></p> <hd id="AN0183983945-4">Common Item Grade Level Alignment in Vertical Scaling</hd> <p></p> <hd id="AN0183983945-5">Common item choice bears directly upon growth interpretations</hd> <p>Most common vertical scaling designs invoke a "grade‐to‐grade" conceptualization of growth (Kolen &amp; Brennan, [<reflink idref="bib26" id="ref13">26</reflink>]), in which growth from, for example, grade 3 to 4 can reflect different content than growth from grade 4 to 5. This is also the conceptualization that forms the basis for the present study. The CING design can be used to operationalize a grade‐to‐grade conceptualization of growth. As briefly noted above, this design involves administering overlapping sets of items to students in adjacent grades, and using differences in their performance to link each grade‐specific scale so that it can be expressed relative to a common unit across grades.</p> <p>This entails a linear adjustment involving two "linking constants," an additive <emph>B</emph> constant and a multiplicative <emph>A</emph> constant. There are different approaches that have been developed and applied to estimate these constants (Haebara, [<reflink idref="bib19" id="ref14">19</reflink>]; Loyd &amp; Hoover, [<reflink idref="bib28" id="ref15">28</reflink>]; Marco, [<reflink idref="bib29" id="ref16">29</reflink>]; Stocking &amp; Lord, [<reflink idref="bib39" id="ref17">39</reflink>]; for equivalent expressions of the linking constants using either item or person parameters, see Kolen &amp; Brennan, [<reflink idref="bib26" id="ref18">26</reflink>], p. 180). Because the <emph>B</emph> and <emph>A</emph> linking constants are the basis for placing different test forms onto a common scale, they form the basis for growth‐related interpretations at both aggregate cross‐sectional and individual longitudinal levels. At the individual level, longitudinal measurement of growth on a vertical scale is implicitly based on the linking constants, and as such, different linking constant values will inevitably produce different conclusions about the amount an individual student has grown over time. In the Methods section, we describe how we approached linking in this study.</p> <p>Various criteria can be used to select the common items used to calibrate a vertical scale. This study focuses on items' grade level alignment. Other criteria have been proposed—for example, Kapoor ([<reflink idref="bib25" id="ref19">25</reflink>]) suggests selecting items based on their apparent sensitivity to growth. This method, however, is not compatible with the CING design, and we follow other prior studies in focusing on grade level as the main criterion for common item selection (Briggs &amp; Dadey, [<reflink idref="bib8" id="ref20">8</reflink>]; [<reflink idref="bib38" id="ref21">38</reflink>]).</p> <hd id="AN0183983945-6">Empirical work on common item grade alignment and relevant learning theory</hd> <p>To our knowledge, only one peer‐reviewed study–Briggs and Dadey ([<reflink idref="bib8" id="ref22">8</reflink>])–has explicitly considered the role of common item grade alignment in vertical scaling. This study introduced the notions of "forward" and "backward" linking designs, corresponding respectively to upper‐ and lower‐grade common item alignments, for the purposes of investigating the phenomenon in which a common item appears more difficult for upper grade students than it does for lower grade students (i.e., a proportion‐correct "reversal"). Although Briggs and Dadey compared difficulty estimate distributions for items reflecting an upper‐grade or lower‐grade alignment, their analysis was limited to comparing item proportions correct, and did not involve the calibration of IRT models, nor did it involve linking a vertical scale (as we do in this study). Still, their work provides motivation for ours in that Briggs and Dadey did discover patterns in these difficulty estimates corresponding to different amounts of growth depending on choice of common items. In their study, growth appeared largest in most grade pairs when using upper‐grade common items; however, this pattern was not always consistent as grade levels increased. Our study takes up this issue in greater detail with a primary focus on comparing growth magnitudes based upon calibrated vertical scales in the domain of math.</p> <p>It bears mentioning that although the grade level alignment of common items has not been a focus in the published research literature on vertical scaling, this may not imply a lack of awareness of its importance among test developers. Kolen and Brennan ([<reflink idref="bib26" id="ref23">26</reflink>]), for example, provide explicit guidance on this issue. However, the research base underlying such guidance is limited for at least two reasons: (<reflink idref="bib1" id="ref24">1</reflink>) the lack of availability of real data with the structure needed to study this issue empirically, and (<reflink idref="bib2" id="ref25">2</reflink>) the complexity of trying to speculatively create plausible data‐generating mechanisms to study this issue via simulation. The first point boils down to the fact that in nearly all applications of vertical scaling with which we are familiar, the common item grade alignment appears to be chosen a priori, so data representing different grade level alignments are simply not available. Briefly touching upon the second point, much can be accomplished in psychometric research, even psychometric research closely related to the present study, via simulation. Using simulation, one can study important and complex issues such as the impact of unmodeled guessing on vertical scaling using the Rasch model (Waterbury &amp; DeMars, [<reflink idref="bib46" id="ref26">46</reflink>]), or how vertical linking methods perform under violations of the Rasch model assumption of common discrimination (Fischer et al., [<reflink idref="bib17" id="ref27">17</reflink>]). Notably, such studies generate data from IRT models (e.g., the three‐parameter logistic [3PL] and two‐parameter logistic [2PL], respectively; Birnbaum, [<reflink idref="bib11" id="ref28">11</reflink>]), and are somewhat agnostic as to the reasons why these might be more plausible data‐generating models than the Rasch model (although the use of the 3PL does rely on an assumption of guessing, of course, and the Rasch model is widely considered a restrictive model compared to IRT models that include item‐level discrimination parameters). There is, to our knowledge, no coherent theory about how common item grade level alignment corresponds to Rasch model misfit, at least not to the extent that one could specify a simulation study with any confidence that the data‐generating model resembles "truth." The relationship between grade level common item alignment could plausibly involve such disparate issues as dimensionality (Hansen &amp; Monroe, [<reflink idref="bib20" id="ref29">20</reflink>]; Li &amp; Lissitz, [<reflink idref="bib27" id="ref30">27</reflink>]; Martineau, [<reflink idref="bib30" id="ref31">30</reflink>]), heterogeneous item‐level effects of a year of schooling (Gilbert et al., [<reflink idref="bib18" id="ref32">18</reflink>]), (lack of) opportunity to learn new material and/or maintain performance on previously learned material, and more. Given the lack of clarity around how to simulate data to study the impact of common item grade level on vertical scaling, we aim to provide empirical findings that might facilitate this type of work in the future.</p> <p>Common item grade alignment has likely received more attention from test developers working to construct new vertical scales than it has from researchers. As such analyses related to this topic are likely documented as technical reports that may or may not be publicly available. One publicly available example comes from a report on the development of Texas's state achievement test's vertical scales in math and reading ([<reflink idref="bib38" id="ref33">38</reflink>]). The report details a study in which more common items than necessary were administered to students across Texas, and linking constants were compared across scales calibrated with lower‐, upper‐, and mixed‐grade designs using item response data gathered from the same population of students. The generalizability of the findings is limited because, as detailed in the report, common items were filtered so that upper‐grade items were always reflective of content that had already been taught in the lower grade. This reduces the representativeness of the common item pool relative to the upper grade test form, and while this choice is endorsed in, for example, Kolen and Brennan ([<reflink idref="bib26" id="ref34">26</reflink>]), it also means that the upper‐grade items in the Texas study do not necessarily reflect the full breadth of upper‐grade content. The study was conducted by calculating separate linking constants using pools of items representing upper, lower, and mixed designs, and then comparing the magnitudes of gradewise differences in logits by alignment. The study found the differences to be small (&lt;0.1 logits larger mean difference using an upper‐grade design across the whole span of grades 3–8) and therefore negligible in both reading and math for grades 3–8. For some grade pairs, a <emph>lower‐grade</emph> design actually appeared to produce slightly larger growth estimates (in reading, in this case). The study also does not report any information about design‐based differences in the variance of the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0001" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> distribution over grades.</p> <p>Learning theory may well suggest scenarios in which growth within certain subdomains of math could appear greater when lower‐grade common items are being used. For example, "backward transfer" (Hohensee, [<reflink idref="bib22" id="ref35">22</reflink>]; Hohensee et al., [<reflink idref="bib23" id="ref36">23</reflink>]) refers to the phenomenon of students applying their more sophisticated conceptualizations of mathematics to problems that could already be solved using less sophisticated conceptualizations (for example, using insights from instruction on quadratic functions to become more capable of solving problems involving linear functions). Yet, in contrast, there are some grade pairs where major shifts in curriculum mean that the content on which students receive instruction during the year will represent a major departure from what they had the chance to learn in years prior (e.g., grade 6 in math; Briggs &amp; Peck, [<reflink idref="bib9" id="ref37">9</reflink>]). The phenomenon of backward transfer could theoretically lead to lower‐grade items showing more evidence of growth than upper‐grade items; curricular shifts could potentially produce the opposite pattern. This further motivates this study, which both details methods by which growth can be compared across different sets of common items and includes some speculation as to the possible reasons why different sets of common items might lead to different estimates of grade‐to‐grade differences in ability—the foundation for the measurement of growth.</p> <hd id="AN0183983945-7">Data: The i‐Ready Diagnostic</hd> <p>The data for this study come from the winter 2019 administration of the i‐Ready Diagnostic interim assessment in math (Curriculum Associates, [<reflink idref="bib14" id="ref38">14</reflink>]). For all analyses, we start with a dataset consisting of all item responses to the winter 2019 i‐Ready Diagnostic nationally. All item responses used in our analyses are for EFT items, which were not operational at the time and administered at random to students in three adjacent grades centered at the grade to whose standards the item is aligned. We limited our sample to students in grades 3 through 7, as there were too few EFT items to support our analyses for grades K–2 or 8–12. The EFT items are generally representative of the operational item bank in terms of content alignment; however, because our dataset included all EFT items from this period, it may include items that were eventually not included as part of the operational item bank due to poor item statistics (e.g., excessively easy or difficult items that contribute very little test information) or poor fit to the Rasch model, a point to which we attend before calibrating the vertical scales for this study. Nearly every EFT item has 500+ responses, with some items coming in at just under 500 after the application of Curriculum Associates' internal data cleaning rules.</p> <p>Figure 1 presents the item difficulty distributions in a proportion correct metric, broken out by student and item grade. We can see that average item difficulty is monotonically decreasing in student grade and increasing in item grade, holding the other constant: higher‐grade students tended to perform better on a given set of items than lower‐grade students, and higher‐grade items tended to be harder for a given set of students than lower‐grade items.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/EMS/01mar25/emip12639-fig-0001.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="emip12639-fig-0001.jpg" title="1 Distributions of Item Difficulty by Item and Student Grade" /> </p> <p></p> <hd id="AN0183983945-9">Methods</hd> <p></p> <hd id="AN0183983945-10">Creating a Vertical Scale Using EFT Items</hd> <p>To address the primary aim of this study, we calculate three sets of four linking constants (for grades 3–4, 4–5, 5–6 and 6–7) using responses from math EFT items. The sets of linking constants correspond to upper, lower, and mixed grade level alignments of the common items. We use the Rasch model to calibrate our vertical scales, in line with the operational model used for i‐Ready and many other vertically scaled tests.</p> <hd id="AN0183983945-11">Initial item filtering</hd> <p>Because the Rasch model is a more constrained model compared to the 2PL and 3PL IRT models, one runs the risk of items that misfit the model to a greater extent than one does using more flexible models (Fischer et al., [<reflink idref="bib17" id="ref39">17</reflink>]; Waterbury &amp; DeMars, [<reflink idref="bib46" id="ref40">46</reflink>]). Given this possibility, we applied an initial, fairly restrictive item fit criterion for the purposes of this study. We began by calibrating the Rasch model with five different sets of item responses, one per student grade, to all applicable items (so, for example, the Rasch model was calibrated using grade 4 student responses to grade 3–5 items, using grade 5 student responses to grade 4–6 items, etc.). Each model was identified by setting the mean of the <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0002" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> distribution to zero. Because each student in the study answered only a handful of items, the model‐based <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0003" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> estimates are very imprecise; however, because these items were answered as part of an operational testing occasion, we have a much more reliable <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0004" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> estimate in the form of students' operational scores. For each student grade, we rescaled operational scores to the same scale as the newly calibrated models by setting the mean of the estimates to 0 and the <emph>SD</emph> to the population <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0005" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml><emph>SD</emph> estimated for the relevant model. We then used these rescaled <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0006" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> estimates to get estimates of Rasch outfit, the unweighted mean‐square standardized residual of responses to the item given each individual's <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0007" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> estimate (Wright &amp; Masters, [<reflink idref="bib49" id="ref41">49</reflink>]). Values of outfit below 1 indicate that an item over‐fits the Rasch model (i.e., its empirical item characteristic curve [ICC] is steeper than the model‐implied curve), while values above 1 indicate under‐fit (i.e., a less steep empirical ICC; Wu &amp; Adams, [<reflink idref="bib51" id="ref42">51</reflink>]). In line with Wu and Adams' findings on the sensitivity of outfit to sample size, we apply a conservative criterion for item fit, removing any item with outfit &gt; 1.2 for any student grade from the pool of linking items for the main analysis. Table 1 below summarizes the number of items showing misfit for each student grade. In general, the items fit the Rasch model well. Between 0.6% and 11.1% of candidate items were flagged as misfitting using the outfit &gt; 1.2 criterion. These items were removed from all subsequent analyses to preclude the possibility of item misfit biasing our findings.</p> <p>1 Table Distributions of Rasch Outfit Statistic by Student Grade</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left"&gt;Grade&lt;/th&gt;&lt;th align="center"&gt;&lt;italic&gt;N&lt;/italic&gt;&lt;/th&gt;&lt;th align="center"&gt;&lt;italic&gt;N&lt;/italic&gt; &amp;#62; 1.20&lt;/th&gt;&lt;th align="center"&gt;&lt;italic&gt;P&lt;/italic&gt; &amp;#62; 1.20&lt;/th&gt;&lt;th align="center"&gt;Mean Outfit&lt;/th&gt;&lt;th align="center"&gt;&lt;italic&gt;SD&lt;/italic&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;3&lt;/td&gt;&lt;td&gt;341&lt;/td&gt;&lt;td&gt;2&lt;/td&gt;&lt;td&gt;0.006&lt;/td&gt;&lt;td&gt;0.990&lt;/td&gt;&lt;td&gt;0.064&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;455&lt;/td&gt;&lt;td&gt;20&lt;/td&gt;&lt;td&gt;0.044&lt;/td&gt;&lt;td&gt;0.975&lt;/td&gt;&lt;td&gt;0.131&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;5&lt;/td&gt;&lt;td&gt;360&lt;/td&gt;&lt;td&gt;20&lt;/td&gt;&lt;td&gt;0.056&lt;/td&gt;&lt;td&gt;0.980&lt;/td&gt;&lt;td&gt;0.107&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;6&lt;/td&gt;&lt;td&gt;340&lt;/td&gt;&lt;td&gt;21&lt;/td&gt;&lt;td&gt;0.062&lt;/td&gt;&lt;td&gt;0.992&lt;/td&gt;&lt;td&gt;0.121&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;7&lt;/td&gt;&lt;td&gt;226&lt;/td&gt;&lt;td&gt;25&lt;/td&gt;&lt;td&gt;0.111&lt;/td&gt;&lt;td&gt;0.982&lt;/td&gt;&lt;td&gt;0.146&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>1 <emph>Note. N</emph> = number of items. <emph>N</emph> &gt; 1.20 = number of items with outfit above 1.20. <emph>P</emph> &gt; 1.20 = proportion of items with outfit above 1.20.</p> <hd id="AN0183983945-12">Linking constants</hd> <p>To construct the three vertical scales used in our main analysis, we use chain‐linking with separate calibrations. Here, we detail the specific procedures we use for chain‐linking. We start the process by calibrating the Rasch model separately for each student grade. The Rasch model specifies the item response function 1 <ephtml> &lt;math display="block" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0008" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;mspace width="0.33em" /&gt;&lt;mfenced separators="" open="(" close=")"&gt;&lt;mrow&gt;&lt;mspace width="0.33em" /&gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo linebreak="goodbreak"&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mo linebreak="badbreak"&gt;=&lt;/mo&gt;&lt;mfrac&gt;&lt;msup&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;msup&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;/mfrac&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mspace width="0.33em" /&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$$\begin{equation}P\ \left({\ {{X}&amp;#95;{ij}} = 1{\mathrm{|}}{{\theta }&amp;#95;i},{{b}&amp;#95;j}} \right) = \frac{{{{e}^{{{\theta }&amp;#95;i} - {{b}&amp;#95;j}}}}}{{1 + {{e}^{{{\theta }&amp;#95;i} - {{b}&amp;#95;j}}}}},\ \end{equation}$$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> where the model is commonly anchored by fixing the mean of <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0009" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> to 0 or by fixing the mean of item difficulty to 0. No variance needs to be specified to identify the model; a population <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0010" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> variance <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0011" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mi&gt;&amp;#963;&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;${{\sigma }^2}(\theta)$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is estimated.</p> <p>We place items onto a common scale using a "mean‐mean" linking methodology (Loyd &amp; Hoover, [<reflink idref="bib28" id="ref43">28</reflink>]). Broadly speaking, IRT linking involves an additive linking constant <emph>B</emph> and a multiplicative linking constant <emph>A</emph>. In the mean‐mean design, two adjacent grades <emph>g</emph> and <emph>g</emph> + 1 are linked as follows (Fischer et al., [<reflink idref="bib16" id="ref44">16</reflink>]). First, the multiplicative linking constant <emph>A</emph> is calculated as 2 <ephtml> &lt;math display="block" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0012" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mspace width="0.33em" /&gt;&lt;msub&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo linebreak="badbreak"&gt;=&lt;/mo&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;msubsup&gt;&lt;mo&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/msubsup&gt;&lt;msub&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;jg&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/mfrac&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;msubsup&gt;&lt;mo&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/msubsup&gt;&lt;msub&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mi&gt;jg&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/mfrac&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$$\begin{equation} \ {A}&amp;#95;{g,g+1}=\frac{{\sum}&amp;#95;{j=1}^{n}{a}&amp;#95;{\textit{jg}+1}}{n}\frac{{\sum}&amp;#95;{j=1}^{n}{a}&amp;#95;{\textit{jg}}}{n}, \end{equation}$$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> where <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0013" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{a}&amp;#95;{j - }}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is a discrimination parameter for item <emph>j</emph> in grade <emph>g</emph> or <emph>g</emph> + 1. The additive constant <emph>B</emph> calculated as 3 <ephtml> &lt;math display="block" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0014" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mspace width="0.33em" /&gt;&lt;msub&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo linebreak="badbreak"&gt;=&lt;/mo&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;msubsup&gt;&lt;mo&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/msubsup&gt;&lt;msub&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/mfrac&gt;&lt;mspace width="0.33em" /&gt;&lt;mo linebreak="goodbreak"&gt;&amp;#8722;&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;msubsup&gt;&lt;mo&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/msubsup&gt;&lt;msub&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/mfrac&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$$\begin{equation}\ {{B}&amp;#95;{g,g + 1}} = \frac{{\sum&amp;#95;{j = 1}^n {{b}&amp;#95;{jg}}}}{n}\ - {{A}&amp;#95;{g,g + 1}}\frac{{\sum&amp;#95;{j = 1}^n {{b}&amp;#95;{jg + 1}}}}{n},\end{equation}$$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> where <emph>n</emph> is the number of common items and <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0015" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{b}&amp;#95;{jg}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is the difficulty of item <emph>j</emph> for students in grade <emph>g</emph>. In Equation (<reflink idref="bib2" id="ref45">2</reflink>), <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0016" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{a}&amp;#95;j}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is a discrimination parameter that varies by item, but in the Rasch model specification of Equation (<reflink idref="bib1" id="ref46">1</reflink>) this parameter is fixed to 1 for all items. Therefore, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0017" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{A}&amp;#95;{g,g + 1}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is equal to 1 and <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0018" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{B}&amp;#95;{g,g + 1}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> takes the simplified form 4 <ephtml> &lt;math display="block" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0019" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mspace width="0.33em" /&gt;&lt;msub&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo linebreak="badbreak"&gt;=&lt;/mo&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;msubsup&gt;&lt;mo&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/msubsup&gt;&lt;msub&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/mfrac&gt;&lt;mspace width="0.33em" /&gt;&lt;mo linebreak="goodbreak"&gt;&amp;#8722;&lt;/mo&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;msubsup&gt;&lt;mo&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/msubsup&gt;&lt;msub&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/mfrac&gt;&lt;mo&gt;.&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$$\begin{equation}\ {{B}&amp;#95;{g,g + 1}} = \frac{{\sum&amp;#95;{j = 1}^n {{b}&amp;#95;{jg}}}}{n}\ - \frac{{\sum&amp;#95;{j = 1}^n {{b}&amp;#95;{jg + 1}}}}{n}.\end{equation}$$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml></p> <p>That is, the linking constant for grade <emph>g</emph> and grade <emph>g</emph> + 1 is the mean difference between the item difficulty estimates for the common items for examinees in grade <emph>g</emph> and the estimates for examinees in grade <emph>g</emph> + 1.</p> <p>To link the scale, we start with grade 3 as the base scale. To put grade 4 items onto the grade 3 scale, we add the corresponding <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0020" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;3&lt;/mn&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mn&gt;4&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{B}&amp;#95;{3,4}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> to their difficulties. To then place grade 5 items onto that same grade 3 scale, we add both <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0021" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;3&lt;/mn&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mn&gt;4&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{B}&amp;#95;{3,4}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> and <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0022" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;4&lt;/mn&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mn&gt;5&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{B}&amp;#95;{4,5}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> to their difficulties. This "chained" approach is also applied to grades 6 and 7 in our study, producing a scale where all item parameters end up on the grade 3 scale. The term "chain‐linking" reflects the fact that linking is done separately for each grade pair but carries through the prior linkages.</p> <p>In all cases, the item responses come from a random sample of the same student population. All analyses were conducted using R (R Core Team, [<reflink idref="bib35" id="ref47">35</reflink>]). We used mirt (Chalmers, [<reflink idref="bib13" id="ref48">13</reflink>]) for model calibration with full‐information marginal maximum likelihood estimation (Bock &amp; Aitkin, [<reflink idref="bib5" id="ref49">5</reflink>]) and we anchor each grade specific item calibration by fixing the mean of <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0023" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> to 0 since this is the mirt default.</p> <hd id="AN0183983945-13">Item pools</hd> <p>The upper‐, lower‐, and mixed‐grade designs are distinguished solely by which items are used as linking items for which grade pair(s). For the lower‐grade design, the common items used to link grade <emph>g</emph> and grade <emph>g</emph> + 1 are aligned to grade <emph>g</emph>. For the upper‐grade design, they are aligned to grade <emph>g</emph> + 1. For the mixed design, they are aligned to either. For each design, we calibrate one model per student grade, in which we include all on‐grade items as well as the off‐grade items needed for the given linking design.</p> <hd id="AN0183983945-14">Linking constants as measures of gradewise differences in θ${\bm{\theta }}$</hd> <p>To assess the grade‐to‐grade differences in performance implied by linking constants, we compare linking constants within each grade pair as well as patterns in those differences across all grade pairs in the study. We do so in logits, the <emph>SD</emph> of the base grade, and adjacent grade pooled <emph>SD</emph> units. Recall that the Rasch model produces a fixed discrimination of 1 and a freely estimated <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0025" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> variance. Each additive linking constant <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0026" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{B}&amp;#95;{g,g + 1}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> expresses the mean <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0027" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> difference across the two linked grades in logit units. There are two ways to give differences expressed in logits an interpretable frame of reference. The first follows the ethos of the Rasch measurement paradigm to use the locations of items, and the distances (or variability) between them, to interpret the practical significance of a mean <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0028" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> difference. The approach, which leverages the use of person‐item maps (i.e., "Wright Maps"), has much to recommend it, but is outside the scope of the present study (for details on how this could be done, see Briggs, [<reflink idref="bib7" id="ref50">7</reflink>] and Wilson, [<reflink idref="bib48" id="ref51">48</reflink>]). The second way is to express mean logit differences relative to the <emph>SD</emph> of one or more <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0029" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> distributions. One example we present is to divide each adjacent grade mean <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0030" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> difference with respect to the <emph>SD</emph> of the grade 3 population distribution. Yen ([<reflink idref="bib53" id="ref52">53</reflink>]) suggested the use of a pooled <emph>SD</emph> across adjacent grades, which turns mean <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0031" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> differences into expressions of the overlap in the adjacent <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0032" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> distributions. For each grade pair <emph>g</emph>, <emph>g</emph> + 1, we also report the mean difference in <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0033" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> between the grades as 5 <ephtml> &lt;math display="block" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0034" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo linebreak="badbreak"&gt;=&lt;/mo&gt;&lt;mfrac&gt;&lt;msub&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;msqrt&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mi&gt;&amp;#963;&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;mfenced separators="" open="(" close=")"&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;/msub&gt;&lt;/mfenced&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;msup&gt;&lt;mi&gt;&amp;#963;&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;mfenced separators="" open="(" close=")"&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/mfrac&gt;&lt;/msqrt&gt;&lt;/mfrac&gt;&lt;mo&gt;.&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$$\begin{equation} E{S}&amp;#95;{g,g+1}=\frac{{B}&amp;#95;{g,g+1}}{\sqrt{\frac{{\sigma}^{2}\left({\theta}&amp;#95;{g}\right)+{\sigma}^{2}\left({\theta}&amp;#95;{g+1}\right)}{2}}}. \end{equation}$$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml></p> <p>Mean differences in effect size units have been used in previous studies of population‐level cross‐sectional growth, or "years of learning" (Bloom et al., [<reflink idref="bib4" id="ref53">4</reflink>]; Dadey &amp; Briggs, [<reflink idref="bib15" id="ref54">15</reflink>]; Hill et al., [<reflink idref="bib21" id="ref55">21</reflink>]; Student, [<reflink idref="bib41" id="ref56">41</reflink>]). Note that <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0035" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$E{{S}&amp;#95;{g,g + 1}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , by definition, does not have the same unit of measurement for each adjacent grade pairing, and will have an ambiguous interpretation if the population <emph>SD</emph>s in different vertical scaling contexts stay constant across grades, grow ("scale expansion") or decrease ("scale shrinkage"). For more on scale expansion versus shrinkage, see Camilli et al. ([<reflink idref="bib12" id="ref57">12</reflink>]) and Yen ([<reflink idref="bib52" id="ref58">52</reflink>]).</p> <hd id="AN0183983945-15">Results</hd> <p>Table 2 reports the results of the separate Rasch calibrations, before linking is applied. We report the population <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0036" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml><emph>SD</emph> and the mean item difficulty for each model. Notice that irrespective of design, we see evidence of increasing variability as grade level increases from 3 to 7.</p> <p>2 Table Grade‐Specific θ$\theta $ Variances and Item Difficulty Means</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left" /&gt;&lt;th align="center" /&gt;&lt;th align="center"&gt;Lower Design&lt;/th&gt;&lt;th align="center"&gt;Mixed Design&lt;/th&gt;&lt;th align="center"&gt;Upper Design&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align="left"&gt;Grades linked&lt;/td&gt;&lt;td align="center"&gt;Student grade&lt;/td&gt;&lt;td align="center"&gt;&lt;p&gt;&lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0038" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics xmlns=""&gt;&lt;mi&gt;&amp;#963;&lt;/mi&gt;2&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;${{\sigma }^2}(\theta)$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="center"&gt;&lt;p&gt;&lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0039" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics xmlns=""&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$\bar{b}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="center"&gt;&lt;p&gt;&lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0040" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics xmlns=""&gt;&lt;mi&gt;&amp;#963;&lt;/mi&gt;2&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;${{\sigma }^2}(\theta)$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="center"&gt;&lt;p&gt;&lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0041" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics xmlns=""&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$\bar{b}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="center"&gt;&lt;p&gt;&lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0042" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics xmlns=""&gt;&lt;mi&gt;&amp;#963;&lt;/mi&gt;2&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;${{\sigma }^2}(\theta)$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="center"&gt;&lt;p&gt;&lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0043" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics xmlns=""&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$\bar{b}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;3&amp;#8211;4&lt;/td&gt;&lt;td&gt;3&lt;/td&gt;&lt;td&gt;0.73&lt;/td&gt;&lt;td&gt;0.07&lt;/td&gt;&lt;td&gt;0.51&lt;/td&gt;&lt;td&gt;0.38&lt;/td&gt;&lt;td&gt;0.52&lt;/td&gt;&lt;td&gt;0.85&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" /&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;0.80&lt;/td&gt;&lt;td&gt;&amp;#8722;0.47&lt;/td&gt;&lt;td&gt;0.68&lt;/td&gt;&lt;td&gt;&amp;#8722;0.24&lt;/td&gt;&lt;td&gt;0.57&lt;/td&gt;&lt;td&gt;0.08&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;4&amp;#8722;5&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;0.80&lt;/td&gt;&lt;td&gt;0.09&lt;/td&gt;&lt;td&gt;0.68&lt;/td&gt;&lt;td&gt;0.41&lt;/td&gt;&lt;td&gt;0.57&lt;/td&gt;&lt;td&gt;0.85&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" /&gt;&lt;td&gt;5&lt;/td&gt;&lt;td&gt;0.87&lt;/td&gt;&lt;td&gt;&amp;#8722;0.35&lt;/td&gt;&lt;td&gt;0.69&lt;/td&gt;&lt;td&gt;&amp;#8722;0.17&lt;/td&gt;&lt;td&gt;0.60&lt;/td&gt;&lt;td&gt;0.07&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;5&amp;#8722;6&lt;/td&gt;&lt;td&gt;5&lt;/td&gt;&lt;td&gt;0.87&lt;/td&gt;&lt;td&gt;0.07&lt;/td&gt;&lt;td&gt;0.69&lt;/td&gt;&lt;td&gt;0.57&lt;/td&gt;&lt;td&gt;0.60&lt;/td&gt;&lt;td&gt;1.15&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" /&gt;&lt;td&gt;6&lt;/td&gt;&lt;td&gt;0.93&lt;/td&gt;&lt;td&gt;&amp;#8722;0.24&lt;/td&gt;&lt;td&gt;0.72&lt;/td&gt;&lt;td&gt;0.14&lt;/td&gt;&lt;td&gt;0.75&lt;/td&gt;&lt;td&gt;0.58&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;6&amp;#8722;7&lt;/td&gt;&lt;td&gt;6&lt;/td&gt;&lt;td&gt;0.93&lt;/td&gt;&lt;td&gt;0.59&lt;/td&gt;&lt;td&gt;0.72&lt;/td&gt;&lt;td&gt;0.99&lt;/td&gt;&lt;td&gt;0.75&lt;/td&gt;&lt;td&gt;1.35&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" /&gt;&lt;td&gt;7&lt;/td&gt;&lt;td&gt;0.98&lt;/td&gt;&lt;td&gt;0.29&lt;/td&gt;&lt;td&gt;0.95&lt;/td&gt;&lt;td&gt;0.58&lt;/td&gt;&lt;td&gt;0.95&lt;/td&gt;&lt;td&gt;0.83&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>2 <emph>Note</emph>. <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0044" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mi&gt;&amp;#963;&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;${{\sigma }^2}(\theta)$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = population variance of <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0045" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> ; <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0046" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mover accent="true"&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;annotation encoding="application/x-tex"&gt;$\bar{b}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = average item difficulty. Mean <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0047" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is 0 in all cases to identify model.</p> <p>Table 3 then presents the results of computing linking constants based upon the results in Table 2. The main comparisons of interest in Table 3 are between the linking constants for a given grade pair across the three grade level alignments; these comparisons tell us to what extent each design produced larger or smaller estimates of growth than the others for that grade pair. As noted above, we provide this comparison in the logit metric (i.e., the measurement unit of the Rasch model), a base <emph>SD</emph> metric, and a pooled adjacent grade <emph>SD</emph> metric (see Equation 5).</p> <p>3 Table Grade‐to‐Grade Linking Constants in Logits and SD Units</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left"&gt;Design&lt;/th&gt;&lt;th align="center"&gt;Grades&lt;/th&gt;&lt;th align="center"&gt;&lt;italic&gt;B&lt;/italic&gt;&lt;/th&gt;&lt;th align="center"&gt;&lt;p&gt;&lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0048" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics xmlns=""&gt;&lt;mrow&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;mo&gt;/&lt;/mo&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;D&lt;/mi&gt;&lt;mn&gt;3&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$B/S{{D}&amp;#95;3}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="center"&gt;&lt;italic&gt;PSD&lt;/italic&gt;&lt;/th&gt;&lt;th align="center"&gt;&lt;italic&gt;B/PSD&lt;/italic&gt;&lt;/th&gt;&lt;th align="center"&gt;Cumulative&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align="left"&gt;Lower&lt;/td&gt;&lt;td&gt;3&amp;#8722;4&lt;/td&gt;&lt;td&gt;0.54&lt;/td&gt;&lt;td&gt;0.63&lt;/td&gt;&lt;td&gt;0.88&lt;/td&gt;&lt;td&gt;0.62&lt;/td&gt;&lt;td&gt;1.59&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" /&gt;&lt;td&gt;4&amp;#8722;5&lt;/td&gt;&lt;td&gt;0.44&lt;/td&gt;&lt;td&gt;0.51&lt;/td&gt;&lt;td&gt;0.91&lt;/td&gt;&lt;td&gt;0.48&lt;/td&gt;&lt;td align="left" /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" /&gt;&lt;td&gt;5&amp;#8722;6&lt;/td&gt;&lt;td&gt;0.31&lt;/td&gt;&lt;td&gt;0.36&lt;/td&gt;&lt;td&gt;0.95&lt;/td&gt;&lt;td&gt;0.32&lt;/td&gt;&lt;td align="left" /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" /&gt;&lt;td&gt;6&amp;#8722;7&lt;/td&gt;&lt;td&gt;0.30&lt;/td&gt;&lt;td&gt;0.34&lt;/td&gt;&lt;td&gt;0.98&lt;/td&gt;&lt;td&gt;0.31&lt;/td&gt;&lt;td align="left" /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Mixed&lt;/td&gt;&lt;td&gt;3&amp;#8722;4&lt;/td&gt;&lt;td&gt;0.62&lt;/td&gt;&lt;td&gt;0.87&lt;/td&gt;&lt;td&gt;0.77&lt;/td&gt;&lt;td&gt;0.81&lt;/td&gt;&lt;td&gt;2.04&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" /&gt;&lt;td&gt;4&amp;#8722;5&lt;/td&gt;&lt;td&gt;0.58&lt;/td&gt;&lt;td&gt;0.82&lt;/td&gt;&lt;td&gt;0.83&lt;/td&gt;&lt;td&gt;0.71&lt;/td&gt;&lt;td align="left" /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" /&gt;&lt;td&gt;5&amp;#8722;6&lt;/td&gt;&lt;td&gt;0.43&lt;/td&gt;&lt;td&gt;0.61&lt;/td&gt;&lt;td&gt;0.84&lt;/td&gt;&lt;td&gt;0.52&lt;/td&gt;&lt;td align="left" /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" /&gt;&lt;td&gt;6&amp;#8722;7&lt;/td&gt;&lt;td&gt;0.41&lt;/td&gt;&lt;td&gt;0.58&lt;/td&gt;&lt;td&gt;0.91&lt;/td&gt;&lt;td&gt;0.45&lt;/td&gt;&lt;td align="left" /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Upper&lt;/td&gt;&lt;td&gt;3&amp;#8722;4&lt;/td&gt;&lt;td&gt;0.77&lt;/td&gt;&lt;td&gt;1.07&lt;/td&gt;&lt;td&gt;0.74&lt;/td&gt;&lt;td&gt;1.04&lt;/td&gt;&lt;td&gt;2.65&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" /&gt;&lt;td&gt;4&amp;#8722;5&lt;/td&gt;&lt;td&gt;0.78&lt;/td&gt;&lt;td&gt;1.08&lt;/td&gt;&lt;td&gt;0.77&lt;/td&gt;&lt;td&gt;1.02&lt;/td&gt;&lt;td align="left" /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" /&gt;&lt;td&gt;5&amp;#8722;6&lt;/td&gt;&lt;td&gt;0.57&lt;/td&gt;&lt;td&gt;0.79&lt;/td&gt;&lt;td&gt;0.82&lt;/td&gt;&lt;td&gt;0.70&lt;/td&gt;&lt;td align="left" /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" /&gt;&lt;td&gt;6&amp;#8722;7&lt;/td&gt;&lt;td&gt;0.53&lt;/td&gt;&lt;td&gt;0.73&lt;/td&gt;&lt;td&gt;0.92&lt;/td&gt;&lt;td&gt;0.57&lt;/td&gt;&lt;td align="left" /&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>3 <emph>Note. B</emph> = additive linking constant in logits; = linking constant divided by standard deviation in base grade 3; <emph>PSD</emph> = pooled standard deviation of <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0049" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> for grade pair; <emph>B/PSD</emph> = linking constant divided by PSD (see Equation 5); Cumulative = total grade 3–7 <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0050" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> mean increase in logits.</p> <p>In all three metrics, the main big picture finding is clear: within each grade pair, grade‐to‐grade differences are largest in all cases using upper‐grade common items and smallest in all cases using lower‐grade common items. In the logit metric, grade‐to‐grade mean differences (i.e., <emph>B</emph> linking constants) appear roughly 1.5×–2× larger for upper‐grade items than for lower‐grade items, with mixed items producing linking constants roughly in between. This naturally carries through to total mean differences in logits across the fully linked scales (column "Total"). Pooled standard deviations (see "PSD" column) are consistent with the "scale expansion" in each design shown in Table 2, but the grade‐specific variances are largest using the lower‐grade design and smallest using the upper‐grade design. In the based grade <emph>SD</emph> and pooled <emph>SD</emph> metrics, growth under the upper‐grade design is around 2x greater than growth under the lower‐grade design. This contrast is shown visually in Figure 2 below.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/EMS/01mar25/emip12639-fig-0002.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="emip12639-fig-0002.jpg" title="2 Growth on a Vertical Scale Using Three Designs and Three Metrics" /> </p> <p></p> <p>One can also look within each design to assess grade‐to‐grade patterns; in all cases we find evidence of growth deceleration (Bloom et al., [<reflink idref="bib4" id="ref59">4</reflink>]), whether one uses the logit metric or either standardized metric; again, this is readily apparent in Figure 2.</p> <hd id="AN0183983945-17">Discussion</hd> <p>Our main finding should not surprise readers familiar with guidance on vertical scaling (Kolen &amp; Brennan, [<reflink idref="bib26" id="ref60">26</reflink>]), but this is nonetheless the first full‐fledged demonstration of this phenomenon in the peer‐reviewed literature. However, the generalizability and implications of our findings are not yet clear. Our main finding alone provides no evidence regarding the question of whether any of the three designs we compare is "better" than the others, but prior practice in vertical scaling has suggested empirical analyses that can be described as investigations of the "reasonableness" of a vertical scale. A study comparing vertical scales would be incomplete without touching upon these and discussing their implications for practice going forward. Here, we report the motivation for and results of follow‐up analyses before turning to this study's overall implications, limitations, and next steps.</p> <hd id="AN0183983945-18">Common Item Grade Alignment and Vertical Scale "Reasonableness"</hd> <p>The practical literature on vertical scaling does point to several expectations one can justifiably have of a vertical scale, as outlined by Patz ([<reflink idref="bib33" id="ref61">33</reflink>]). The Texas technical report cited above refers to these as indications of a scale's "reasonableness" ([<reflink idref="bib38" id="ref62">38</reflink>], p. 36). One is that the items will show growth (that is, they will be easier for upper‐grade students than for lower‐grade students). Violations of this assumption are known as "reversals" (per Briggs &amp; Dadey, [<reflink idref="bib8" id="ref63">8</reflink>], who investigated this issue qualitatively and in the <emph>p</emph>‐value metric). Another is that the items will function similarly for upper‐ and lower‐grade students—that is, their difficulty ordering should be similar no matter which students are taking a given set of items. Both expectations combine the IRT property of parameter invariance with an implied theory of growth: that growth is roughly uniform across different items potentially representing different content sub‐domains. The two expectations are closely related: when growth from a lower grade to an upper grade is positive, it is not possible for items to show reversals unless parameter invariance has been violated to some extent (or the item difficulty standard errors are large, which is not the case in this study given our sample sizes).</p> <p>Here, we consider the extent to which each vertical scale in this study (based on upper, lower, and mixed grade common item alignments) meets the "reasonableness" expectations that items will function similarly across grades and show growth. In prior work, researchers and measurement practitioners have approached these issues in different ways. We begin by considering parameter invariance, as this has received attention in a variety of studies and practitioner guides. As described by Patz ([<reflink idref="bib33" id="ref64">33</reflink>]) and implemented in the Texas STAAR vertical scale technical report cited in this study, one approach is to simply inspect the correlations between the two sets of parameter estimates for each set of common items. Other potential approaches include (<reflink idref="bib1" id="ref65">1</reflink>) using the calibrated scale to predict item <emph>p</emph>‐values for different grade levels (Najarian et al., [<reflink idref="bib32" id="ref66">32</reflink>]) or (<reflink idref="bib2" id="ref67">2</reflink>) conducting formal significance tests of differential item functioning using for example Mantel–Haenszel techniques (Holland &amp; Thayer, [<reflink idref="bib24" id="ref68">24</reflink>]) for all potential common items. We did not conduct DIF tests because of the extreme sparseness of our data. Rather, we calibrated each scale with all available items that fit the Rasch model. Here, we now consider the extent to which each scale may introduce concerns about parameter invariance.</p> <p>First, we report the Pearson's correlation between the two sets of parameter estimates for each set of common items from each design, as suggested by Patz ([<reflink idref="bib33" id="ref69">33</reflink>]). Second, we report the proportion of items that have violations of Rasch parameter invariance (i.e., their difficulties are not the same for upper‐ and lower‐grade students). To assess parameter invariance, we follow the "quality control" methods introduced by Wright and Stone ([<reflink idref="bib50" id="ref70">50</reflink>], p. 95). For a given item, its lower‐grade difficulty estimate is compared to its upper‐grade difficulty estimate, adjusted by the <emph>B</emph> linking constant so that they would be identical under perfect parameter invariance. This difference is compared to an "error unit," which is the square root of the summed squared standard errors of the two estimates. If the difference between the two estimates exceeds 1.96 error units, this is a statistically significant invariance violation with a nominal Type I error rate of 0.05. Table 4 presents the results of these analyses.</p> <p>4 Table Summary of Invariance Analysis</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left" /&gt;&lt;th align="center"&gt;Lower&lt;/th&gt;&lt;th align="center"&gt;Mixed&lt;/th&gt;&lt;th align="center"&gt;Upper&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Grades&lt;/td&gt;&lt;td&gt;r&lt;/td&gt;&lt;td&gt;N&lt;/td&gt;&lt;td&gt;%&lt;/td&gt;&lt;td&gt;r&lt;/td&gt;&lt;td&gt;N&lt;/td&gt;&lt;td&gt;%&lt;/td&gt;&lt;td&gt;r&lt;/td&gt;&lt;td&gt;N&lt;/td&gt;&lt;td&gt;%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;3&amp;#8211;4&lt;/td&gt;&lt;td&gt;0.95&lt;/td&gt;&lt;td&gt;200&lt;/td&gt;&lt;td&gt;31.0&lt;/td&gt;&lt;td&gt;0.91&lt;/td&gt;&lt;td&gt;334&lt;/td&gt;&lt;td&gt;43.4&lt;/td&gt;&lt;td&gt;0.77&lt;/td&gt;&lt;td&gt;134&lt;/td&gt;&lt;td&gt;52.2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;4&amp;#8211;5&lt;/td&gt;&lt;td&gt;0.94&lt;/td&gt;&lt;td&gt;134&lt;/td&gt;&lt;td&gt;29.1&lt;/td&gt;&lt;td&gt;0.90&lt;/td&gt;&lt;td&gt;235&lt;/td&gt;&lt;td&gt;46.0&lt;/td&gt;&lt;td&gt;0.89&lt;/td&gt;&lt;td&gt;100&lt;/td&gt;&lt;td&gt;42.0&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;5&amp;#8211;6&lt;/td&gt;&lt;td&gt;0.94&lt;/td&gt;&lt;td&gt;100&lt;/td&gt;&lt;td&gt;45.0&lt;/td&gt;&lt;td&gt;0.93&lt;/td&gt;&lt;td&gt;202&lt;/td&gt;&lt;td&gt;47.5&lt;/td&gt;&lt;td&gt;0.91&lt;/td&gt;&lt;td&gt;85&lt;/td&gt;&lt;td&gt;45.9&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;6&amp;#8211;7&lt;/td&gt;&lt;td&gt;0.98&lt;/td&gt;&lt;td&gt;85&lt;/td&gt;&lt;td&gt;20.0&lt;/td&gt;&lt;td&gt;0.97&lt;/td&gt;&lt;td&gt;194&lt;/td&gt;&lt;td&gt;25.3&lt;/td&gt;&lt;td&gt;0.97&lt;/td&gt;&lt;td&gt;102&lt;/td&gt;&lt;td&gt;25.5&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>4 <emph>Note</emph>: <emph>r</emph> = Pearson's correlation; <emph>N</emph> = total common items; % = percentage of total common items with difficulty estimates falling outside 95% "quality control" interval, that is, invariance violations.</p> <p>We highlight two key results in Table 4. First, for most grade pairs, there is substantially weaker evidence in support of parameter invariance using upper‐grade common items, in that the correlations are generally lowest for the upper‐grade design within each grade pair (this is less of a concern for grade 6–7, and most stark for grades 3–4). While Patz ([<reflink idref="bib33" id="ref71">33</reflink>]) does not provide any guidance on what a "good enough" correlation would be, nor is there any hypothesis test associated with these correlation coefficients, we note that the operational example cited throughout this manuscript ([<reflink idref="bib38" id="ref72">38</reflink>]) reported correlations between 0.90 and 0.98. All but two of the correlations in Table 4 fall within this range, which can be considered <emph>de facto</emph> sufficient for operational use, at least compared to what little prior works exists. However, we also see invariance violations well above their nominal 5% Type I error rate in all cases. These are generally—but not always—highest under the upper‐grade design, again suggesting that parameter invariance is a particular concern when using upper‐grade common items.</p> <p>We also report the distribution of item difficulty reversals by student grade pair and common item grade alignment. An item has a reversal for two grades being linked if its pre‐linking item difficulty is lower in the lower grade than in the upper grade. That is, because the Rasch calibrations for individual student grades all are anchored by a mean <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0051" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> of 0, a lower difficulty in the lower grade means that the item was easier for lower‐grade students than for upper‐grade students. Table 5 reports our results. We can see that reversals are most common under a lower‐grade design. This is almost mechanically guaranteed once one establishes that grade‐to‐grade mean differences are smaller under this design, as we found above, since a distribution of difficulty differences with a given variance will always overlap with 0 more as the mean of that distribution gets closer to 0. However, the within‐design differences are also striking; in every design, the most reversals occur when comparing grade 5 and grade 6 students' performance, and 23% of lower‐grade items show reversals. This is notable because prior work on vertical scaling in elementary and middle school math (Briggs &amp; Peck, [<reflink idref="bib9" id="ref73">9</reflink>]) has noted that this grade pair represents a major shift in content emphasis, meaning that U.S. standards for grade 6 emphasize instruction on new material with less overlap with the prior grade's material than is found in the standards of earlier grades. That is, our findings are consistent with a general finding regarding math standards, making it likely that they are not entirely unique to this one set of common items.</p> <p>5 Table Summary of Reversals by Common Item Grade Alignment</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left" /&gt;&lt;th align="center"&gt;Lower&lt;/th&gt;&lt;th align="center"&gt;Mixed&lt;/th&gt;&lt;th align="center"&gt;Upper&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Grades&lt;/td&gt;&lt;td align="center"&gt;Total&lt;/td&gt;&lt;td align="center"&gt;Rev.&lt;/td&gt;&lt;td align="center"&gt;% rev.&lt;/td&gt;&lt;td align="center"&gt;Total&lt;/td&gt;&lt;td align="center"&gt;Rev.&lt;/td&gt;&lt;td align="center"&gt;% rev.&lt;/td&gt;&lt;td align="center"&gt;Total&lt;/td&gt;&lt;td align="center"&gt;Rev.&lt;/td&gt;&lt;td align="center"&gt;% rev.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;3&amp;#8211;4&lt;/td&gt;&lt;td&gt;200&lt;/td&gt;&lt;td&gt;9&lt;/td&gt;&lt;td&gt;4.5&lt;/td&gt;&lt;td&gt;334&lt;/td&gt;&lt;td&gt;10&lt;/td&gt;&lt;td&gt;3.0&lt;/td&gt;&lt;td&gt;134&lt;/td&gt;&lt;td&gt;2&lt;/td&gt;&lt;td&gt;1.5&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;4&amp;#8211;5&lt;/td&gt;&lt;td&gt;134&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;3.0&lt;/td&gt;&lt;td&gt;235&lt;/td&gt;&lt;td&gt;7&lt;/td&gt;&lt;td&gt;3.0&lt;/td&gt;&lt;td&gt;100&lt;/td&gt;&lt;td&gt;2&lt;/td&gt;&lt;td&gt;2.0&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;5&amp;#8211;6&lt;/td&gt;&lt;td&gt;100&lt;/td&gt;&lt;td&gt;23&lt;/td&gt;&lt;td&gt;23.0&lt;/td&gt;&lt;td&gt;202&lt;/td&gt;&lt;td&gt;33&lt;/td&gt;&lt;td&gt;16.3&lt;/td&gt;&lt;td&gt;85&lt;/td&gt;&lt;td&gt;6&lt;/td&gt;&lt;td&gt;7.1&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;6&amp;#8211;7&lt;/td&gt;&lt;td&gt;85&lt;/td&gt;&lt;td&gt;7&lt;/td&gt;&lt;td&gt;8.2&lt;/td&gt;&lt;td&gt;194&lt;/td&gt;&lt;td&gt;16&lt;/td&gt;&lt;td&gt;8.2&lt;/td&gt;&lt;td&gt;102&lt;/td&gt;&lt;td&gt;5&lt;/td&gt;&lt;td&gt;4.9&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>5 <emph>Note</emph>: Total = total # of linking items; Rev. = # of reversals; % rev. = percentage of reversals.</p> <p>Taken together, we take two key points away from our findings in this section. First, our findings on parameter invariance suggest that invariance violations at the item level can be quite common, even when correlations between the difficulty estimates for a given set of items for two grade levels are reasonably strong. It may not be surprising that 53.7% of items appear non‐invariant when the correlation between the estimates is 0.78 (grade 3–4, upper grade design), but the percentage for a correlation of 0.91 (mixed design grades 3–4, lower design grades 5–6) still exceeds 40%. Notably, these analyses of invariance also are conducted <emph>after</emph> linking constants are applied, so they show that item‐specific grade‐to‐grade performances are variable, but they do not provide any evidence on the potential effects of DIF in the common items on the linking constants themselves. The full impact of potential common item DIF on linking constants might be better understood via future studies that take up this issue in greater detail, for example by using recently‐introduced DIF tests that attempt to circumvent the non‐identifiability of individual item difficulties in the Rasch model (Bechger &amp; Maris, [<reflink idref="bib2" id="ref74">2</reflink>], as suggested by Fischer et al., [<reflink idref="bib16" id="ref75">16</reflink>]) or that model DIF and grade‐to‐grade <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0052" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> distributions simultaneously and include them in the same overall measurement model (Bauer &amp; Hussong, [<reflink idref="bib1" id="ref76">1</reflink>]).</p> <p>Second, we assert that our findings underscore the importance of content expertise in vertical scale common item selection and refinement. We start by emphasizing that a reversal does not automatically mean that the item is poorly written or fails to measure the targeted content accurately; reversals can occur for legitimate, content‐based reasons such as instruction deemphasizing earlier‐grade content, which can manifest as DIF favoring lower‐grade students (Tao &amp; Mix, [<reflink idref="bib42" id="ref77">42</reflink>]). Reversals may also, as noted above, be a natural consequence of using linking items that produce smaller overall linking constants combined with natural variability in item‐by‐item gradewise differences in performance. That is, we do not advocate rejecting a lower‐grade alignment just because it produces more reversals. This contrasts with the perspective that showing growth from one grade to the next should be an empirical criterion for selecting common items (Kapoor, [<reflink idref="bib25" id="ref78">25</reflink>]). Rather, the presence of large numbers of reversals points to the limitations of purely quantitative analysis of vertical scale common items. The role of content expertise in making sense of the findings in this section is crucial. As described by Briggs and Dadey ([<reflink idref="bib8" id="ref79">8</reflink>]), it can be true both that a reversal reflects legitimate trends in student learning and that a large number of reversals might call into question aspects of the theory of teaching and learning that motivated the design of the vertical scale. Content review of common items can help shed light on whether items with reversals are appropriate for inclusion or not, based on the extent to which they reflect curricular content that is expected to be instructionally sensitive. If items meet such content criteria, then reversals are indicative of a real empirical phenomenon that a vertical scale likely should attempt to capture. Of course, this also requires a well‐articulated learning theory that underlies the scale, something that in our experience often receives limited attention in operational vertical scaling. Similarly, when it comes to parameter invariance issues, content experts should be consulted to determine, for example, if the items appear likely to favor upper‐ or lower‐grade students in the aggregate, or if they represent a roughly equal mix of items that favor either group, as well as the extent to which items with potential DIF are necessary for content representation. If the former is the case, analytic methods to account for the potential bias in the linking constants introduced by DIF may be warranted (Bauer &amp; Hussong, [<reflink idref="bib1" id="ref80">1</reflink>]), but future work needs to investigate this issue in greater detail.</p> <hd id="AN0183983945-19">Implications</hd> <p>This study demonstrates the extent to which different approaches to choosing vertical scale common items–using item responses from the same population of examinees–can lead to different conclusions about the magnitude of gradewise mean differences in <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0053" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , the foundation of vertical scaling with a CING linking design and a grade‐to‐grade conceptualization of growth. We begin by noting that we found the same pattern–largest mean differences using upper‐grade items, smallest differences using lower‐grade items–in every grade pair we were able to include in this study. This suggests that a major component of how much students appear to grow on a vertical scale for mathematics is opportunity to learn relative to the common items, in that the lower‐grade design uses items that students had the opportunity to learn in the lower grade, while some of the upper grade items reflect new content where opportunity to learn first occurs during the upper grade (Briggs &amp; Dadey, [<reflink idref="bib8" id="ref81">8</reflink>]). Although there is some precedent for the perspective that vertical scale common items are better if they show more growth (Kapoor, [<reflink idref="bib25" id="ref82">25</reflink>]; Reckase, [<reflink idref="bib37" id="ref83">37</reflink>]), the use of upper‐grade items as vertical scale common items appears to require particular care to ensure that the items function at least roughly equivalently across grades. Using a large pool of items aligned to upper‐grade standards, we found weaker evidence of parameter invariance for most grade pairs compared to using lower‐grade items. This suggests that while the linking constants for the scale with upper‐grade items are largest, they may also be the most fraught, with any type of cross‐grade comparison weakened.</p> <p>However, the solution may not be as simple as removing individual problematic upper‐grade common items: we note that if opportunity to learn drives larger apparent growth on a scale linked using upper‐grade items, and opportunity to learn also drives DIF in upper‐grade items, then removing noninvariant upper‐grade items from the linking pool could certainly reduce the magnitude of apparent growth, even though this would run counter to the justification for using upper‐grade items in the first place. The large differences in growth between upper‐grade and lower‐grade items also might complicate the validation and interpretation of a mixed design, in that this design appears to involve mixing together two pools of items that give very different answers about growth and treating them as a homogeneous set. To justify the use of the mixed design, one would have to provide evidence that growth is similar on upper‐ and lower‐grade items (as in e.g., [<reflink idref="bib38" id="ref84">38</reflink>]), but again, achieving this might involve having to remove a large number of potential common items from the pool, reducing the extent to which the common items represent the content of the grade‐specific forms, a crucial issue in vertical scaling (Kolen &amp; Brennan, [<reflink idref="bib26" id="ref85">26</reflink>]). In short, there are conceptual tensions in the mixed design, and it certainly seems that using upper‐grade designs runs a greater risk of producing a scale where comparability of scores across grades is threatened. Yet, upper‐grade and mixed designs were used on every U.S. scale for which a recent study of growth on vertical scales includes data (Student, [<reflink idref="bib41" id="ref86">41</reflink>]). Perhaps not surprisingly given our findings, that study documents the fact that growth is systematically larger on scales for grade 3–8 math linked using upper‐grade designs than on scales linked using mixed designs.</p> <p>Ultimately, one cannot decide which common item alignment to use without deciding what growth on that test is supposed to mean. Thus, test developers should not be agnostic to the role of common item alignment when constructing new vertical scales; rather, common item alignment is a design feature that should be explicitly decided upon just like other features of the scale that directly affect apparent growth from one grade to the next (e.g., the measurement model or calibration method; Briggs &amp; Weeks, [<reflink idref="bib10" id="ref87">10</reflink>]). The question of how growth is to be conceptualized and–when adopting a grade‐to‐grade approach–operationalized requires explicit consideration early in the test development process. At a minimum, publicly available technical documentation should be clear about <emph>why</emph> the common item alignment used for that scale was chosen, and the design should be connected to the conceptualization of growth that it implies (i.e., what is being measured when score gains are compared across individuals or populations?). More ideally, for designs that appear likelier to run into issues of parameter noninvariance/lack of cross‐grade comparability, it should be standard practice for test developers to collect the data necessary to reckon honestly with the question of whether empirical data support the use of the desired linking design.</p> <p>Our findings are also highly relevant to researchers who use large scale assessment scores to assess the impact of programs or interventions. As noted, there is already evidence that growth on operational vertical scales is larger when those scales use upper‐grade common items, especially in math (Student, [<reflink idref="bib41" id="ref88">41</reflink>]). Our findings further underscore that growth on a scale linked with upper‐grade items may not be comparable with growth on a scale linked using mixed or lower‐grade items, because the type of growth they measure is not the same. Researchers should thus ideally be aware of the linking design used to calibrate a vertically scaled outcome measure when inferences premised on magnitudes of growth are of interest. For example, this can be especially important when benchmarking causal effects against growth estimates from vertically scaled tests (Bloom et al., [<reflink idref="bib4" id="ref89">4</reflink>]). More broadly, comparing growth across scales linked using different common item grade alignments in a substantive manner (such as longitudinal studies of student growth in different contexts) may be questionable, even after translating growth into effect size units (May et al., [<reflink idref="bib31" id="ref90">31</reflink>]).</p> <hd id="AN0183983945-20">Limitations and Future Directions</hd> <p>To conclude, we note both some limitations of this study and corresponding future directions for research. A future direction for research is to explore the relationship between common item grade alignment and linking constants using more flexible IRT models, such as the 2PL or 3PL. The data used in this study are too sparse for stable estimation of item‐specific discrimination or guessing parameters, but the methods we used for this study could be applied just as easily for more flexible IRT models if a suitable dataset were available. Of course, the choice to use a non‐Rasch measurement model for vertical scaling would mean dispensing with the possibility of an invariant item‐side reference unit, which makes reporting the meaning of growth along the scale in absolute terms much more tenuous (Briggs, [[<reflink idref="bib6" id="ref91">6</reflink>]]).</p> <p>We also found larger differences in linking constants by common item alignment compared to a similar, state‐specific study ([<reflink idref="bib38" id="ref92">38</reflink>]), and while we speculated that this was due to up‐front item filtering in that study, the fact remains that our findings are not consistent. This underscores the likelihood that a study similar to ours, applied to field test data for a different vertical scale, could produce somewhat different results (just as different vertical scales in the same subject tend to produce different estimates of grade‐to‐grade mean differences; Dadey &amp; Briggs, [<reflink idref="bib15" id="ref93">15</reflink>]). The vertical scaling research base would benefit from more empirical studies in this area, especially studies using responses to items specifically selected for vertical scale linking. Additionally, it would be worthwhile to conduct a study similar to this one in ELA/reading, given Student ([<reflink idref="bib41" id="ref94">41</reflink>])'s finding that growth on vertical scales in ELA/reading is quite variable, but not strongly predicted by common item grade alignment to the same extent as in math.</p> <p>We also note that even in grade pairs where parameter invariance seemed to hold (e.g., 6–7), and after accounting for item‐model fit, we still found that an upper‐grade common item alignment produced the largest grade‐to‐grade difference in <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0054" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> . Further exploration of the empirical drivers of grade‐to‐grade differences in <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12639:emip12639-math-0055" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\theta $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> by common item alignment are surely warranted. One salient area we could not investigate is "construct shift." This refers to violations of the assumption of common item unidimensionality; such violations invite the inference that the interpretation of the construct is changing from one location on the scale to another. Several methods have been proposed both to conceptualize the empirical consequences of construct shift and to account for construct shift in measurement models (e.g., Hansen &amp; Monroe, [<reflink idref="bib20" id="ref95">20</reflink>]; Li &amp; Lissitz, [<reflink idref="bib27" id="ref96">27</reflink>]; Martineau, [<reflink idref="bib30" id="ref97">30</reflink>]; Strachan et al., [<reflink idref="bib40" id="ref98">40</reflink>]; Weeks, [<reflink idref="bib47" id="ref99">47</reflink>]). However, the sparseness of the data used in this study precluded any investigation of dimensionality, and as such, we could not investigate construct shift.</p> <p>More broadly, this study moves the field toward the possibility of empirically‐grounded simulation studies of common item grade level alignment, in that it points to parameter invariance differences and heterogeneous item‐level effects of a year of learning as two things that appear to contribute to differences in how much a given population of students will appear to grow. While this study focused on unidimensional scales, perhaps studies in the world of multidimensional vertical scaling (e.g., Hansen &amp; Monroe, [<reflink idref="bib20" id="ref100">20</reflink>]; Strachan et al., [<reflink idref="bib40" id="ref101">40</reflink>]; Weeks, [<reflink idref="bib47" id="ref102">47</reflink>]) can provide a starting point for simulating the types of multidimensionality that could theoretically also manifest as construct shift. Taken together, these would provide researchers with a much stronger basis for simulation studies than they have had previously.</p> <p>Next, this study is specific to the CING design using a grade‐to‐grade definition of growth (Kolen &amp; Brennan, [<reflink idref="bib26" id="ref103">26</reflink>]). Future work should explore the relationship between common item selection and grade‐to‐grade score differences under different linking methods (Béguin &amp; Wools, [<reflink idref="bib3" id="ref104">3</reflink>]; Petersen et al., [<reflink idref="bib34" id="ref105">34</reflink>]). Although our focus in this study was on the results from a grade‐to‐grade conceptualization of growth, and this appears to reflect the approach commonly taken to calibrate most operational vertical scales in the United States (Student, [<reflink idref="bib41" id="ref106">41</reflink>]), we also note that there are at least two other ways to conceive of growth in the design of a vertical scale, the "domain" approach (Kolen &amp; Brennan, [<reflink idref="bib26" id="ref107">26</reflink>]), and the "learning progression" approach (Briggs &amp; Peck, [<reflink idref="bib9" id="ref108">9</reflink>]). It would be especially worthwhile to replicate this study in the context of a learning progression approach to vertical scaling (Briggs &amp; Peck, [<reflink idref="bib9" id="ref109">9</reflink>]). In this context, when linking constants differ by common item grade, this can be taken as evidence that the learning theory underlying the construct of measurement requires further refinement, or that the items themselves do not properly reflect the theory. Future work on linking in vertical scaling will need to address conceptual issues like these in addition to the empirical ones described above.</p> <ref id="AN0183983945-21"> <title> Footnotes </title> <blist> <bibl id="bib1" idref="ref24" type="bt">1</bibl> <bibtext> Note that while the window for the Winter administration can be quite wide (November through February) to accommodate different school district testing schedules, data collection for EFT items was only active for a couple of weeks during February 2019. As such, and given the shift in the latter part of the spring semester in most schools to focus on preparation for state testing, we believe that these data reflect a substantial amount of the expected student growth from within the current school year.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref25" type="bt">2</bibl> <bibtext> We thank four anonymous peer reviewers for suggestions relating to this section's framing and analyses.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref104" type="bt">3</bibl> <bibtext> Although vertical scales can also be created using a "scaling test" that is common to students in all grades on the scale (Petersen et al., [34]), and this approach enables additional analysis of common items (Béguin &amp; Wools, [3]), the scaling test design does not seem to have gained much traction, based on our review of technical manuals for operational vertical scales.</bibtext> </blist> <blist> <bibl id="bib4" idref="ref53" type="bt">4</bibl> <bibtext> The term "full‐information" refers to the fact that the likelihood function is maximized relative to all individual response patterns rather than the item response covariance matrix, a necessity given the EFT item administration design. Full‐information maximum likelihood assumes that missing data are missing completely at random, in line with the EFT administration design.</bibtext> </blist> </ref> <ref id="AN0183983945-22"> <title> References </title> <blist> <bibtext> Bauer, D. J., &amp; Hussong, A. M. (2009). Psychometric approaches for developing commensurate measures across independent studies: Traditional and new models. Psychological Methods, 14 (2), 101 – 125. https://doi.org/10.1037/a0015583</bibtext> </blist> <blist> <bibtext> Bechger, T. M., &amp; Maris, G. (2015). A statistical test for differential item pair functioning. Psychometrika, 80 (2), 317 – 340. https://doi.org/10.1007/s11336‐014‐9408‐y</bibtext> </blist> <blist> <bibtext> Béguin, A. A., &amp; Wools, S. (2015). Vertical comparison using reference sets. In R. E. Millsap, D. M. Bolt, L. A. VanDer Ark, &amp; W.‐C. Wang (Eds.), Quantitative Psychology Research: The 78th Annual Meeting of the Psychometric Society. (Vol., 89, pp. 195 – 211). Springer International Publishing, https://doi.org/10.1007/978‐3‐319‐07503‐7</bibtext> </blist> <blist> <bibtext> Bloom, H. S., Hill, C. J., Black, A. R., &amp; Lipsey, M. W. (2008). Performance trajectories and performance gaps as achievement effect‐size benchmarks for educational interventions. Journal of Research on Educational Effectiveness, 1 (4), 289 – 328. https://doi.org/10.1080/19345740802400072</bibtext> </blist> <blist> <bibl id="bib5" idref="ref49" type="bt">5</bibl> <bibtext> Bock, R. D., &amp; Aitkin, M. (1981). Marginal maximum likelihood estimation of item parameters: Application of an EM algorithm. Psychometrika, 46 (4), 443 – 459. https://doi.org/10.1007/BF02293801</bibtext> </blist> <blist> <bibl id="bib6" idref="ref91" type="bt">6</bibl> <bibtext> Briggs, D. C. (2013). Measuring growth with vertical scales. Journal of Educational Measurement, 50 (2), 204 – 226. https://doi.org/10.1111/jedm.12011</bibtext> </blist> <blist> <bibl id="bib7" idref="ref50" type="bt">7</bibl> <bibtext> Briggs, D. C. (2019). Interpreting and visualizing the unit of measurement in the Rasch Model. Measurement, 146, 961 – 971. https://doi.org/10.1016/j.measurement.2019.07.035</bibtext> </blist> <blist> <bibl id="bib8" idref="ref20" type="bt">8</bibl> <bibtext> Briggs, D. C., &amp; Dadey, N. (2015). Making sense of common test items that do not get easier over time: Implications for vertical scale designs. Educational Assessment, 20 (1), 1 – 22. https://doi.org/10.1080/10627197.2014.995165</bibtext> </blist> <blist> <bibl id="bib9" idref="ref37" type="bt">9</bibl> <bibtext> Briggs, D. C., &amp; Peck, F. A. (2015). Using learning progressions to design vertical scales that support coherent inferences about student growth. Measurement: Interdisciplinary Research and Perspectives, 13 (2), 75 – 99. https://doi.org/10.1080/15366367.2015.1042814</bibtext> </blist> <blist> <bibtext> Briggs, D. C., &amp; Weeks, J. P. (2009). The impact of vertical scaling decisions on growth interpretations. Educational Measurement: Issues and Practice, 28 (4), 3 – 14. https://doi.org/10.1111/j.1745‐3992.2009.00158.x</bibtext> </blist> <blist> <bibtext> Birnbaum, A. (1968). Some latent trait models and their use in inferring an examinee's ability. In F. M. Lord &amp; M. R. Novick (Eds.), Statistical theories of mental test scores (pp. 395 – 479). Addison‐Wesley.</bibtext> </blist> <blist> <bibtext> Camilli, G., Yamamoto, K., &amp; Wang, M. (1993). Scale shrinkage in vertical equating. Applied Psychological Measurement, 17 (4), 379 – 388. https://doi.org/10.1177/014662169301700407</bibtext> </blist> <blist> <bibtext> Chalmers, R. P. (2012). mirt: A multidimensional item response theory package for the R environment. Journal of Statistical Software, 48 (6), 1 – 29. https://doi.org/10.18637/jss.v048.i06</bibtext> </blist> <blist> <bibtext> Curriculum Associates. (2019). I‐Ready assessments technical manual. Curriculum Associates.</bibtext> </blist> <blist> <bibtext> Dadey, N., &amp; Briggs, D. C. (2012). A meta‐analysis of growth trends from vertically scaled assessments. Practical Assessment, Research and Evaluation, 17. https://doi.org/10.7275/F2BM‐6R59</bibtext> </blist> <blist> <bibtext> Fischer, L., Gnambs, T., Rohm, T., &amp; Carstensen, C. H. (2019). Longitudinal linking of Rasch‐model‐scaled competence tests in large‐scale assessments: A comparison and evaluation of different linking methods and anchoring designs based on two tests on mathematical competence administered in grades 5 and 7. Psychological Test and Assessment Modeling, 61 (1), 37 – 64.</bibtext> </blist> <blist> <bibtext> Fischer, L., Rohm, T., Carstensen, C. H., &amp; Gnambs, T. (2021). Linking of Rasch‐scaled tests: Consequences of limited item pools and model misfit. Frontiers in Psychology, 12, 633896. https://doi.org/10.3389/fpsyg.2021.633896</bibtext> </blist> <blist> <bibtext> Gilbert, J. B., Kim, J. S., &amp; Miratrix, L. W. (2023). Modeling item‐level heterogenous treatment effect with the Explanatory Item Response Model: Leveraging large‐scale online assessments to pinpoint the impact of educational interventions. Journal of Educational and Behavioral Statistics, 48 (6), 889 – 913. https://doi.org/10.3102/10769986231171710</bibtext> </blist> <blist> <bibtext> Haebara, T. (1980). Equating logistic ability scales by a weighted least squares method. Japanese Psychological Research, 22 (3), 144 – 149. https://doi.org/10.4992/psycholres1954.22.144</bibtext> </blist> <blist> <bibtext> Hansen, M., &amp; Monroe, S. (2018). Linking not‐quite‐vertical scales through Multidimensional Item Response Theory. Measurement: Interdisciplinary Research and Perspectives, 16 (3), 155 – 167. https://doi.org/10.1080/15366367.2018.1497350</bibtext> </blist> <blist> <bibtext> Hill, C. J., Bloom, H. S., Black, A. R., &amp; Lipsey, M. W. (2008). Empirical benchmarks for interpreting effect sizes in research. Child Development Perspectives, 2 (3), 172 – 177. https://doi.org/10.1111/j.1750‐8606.2008.00061.x</bibtext> </blist> <blist> <bibtext> Hohensee, C. (2014). Backward Transfer: An investigation of the influence of quadratic functions instruction on students' prior ways of reasoning about linear functions. Mathematical Thinking and Learning, 16 (2), 135 – 174. https://doi.org/10.1080/10986065.2014.889503</bibtext> </blist> <blist> <bibtext> Hohensee, C., Willoughby, L., &amp; Gartland, S. (2024). Backward transfer, the relationship between new learning and prior ways of reasoning, and action versus process views of linear functions. Mathematical Thinking and Learning, 26 (1), 71 – 89. https://doi.org/10.1080/10986065.2022.2037043</bibtext> </blist> <blist> <bibtext> Holland, P. W., &amp; Thayer, D. T. (1986). Differential item functioning and the Mantel‐Haenszel procedure. ETS Research Report Series, 1986 (2), 1 – 24. https://doi.org/10.1002/j.2330‐8516.1986.tb00186.x</bibtext> </blist> <blist> <bibtext> Kapoor, S. (2014). Growth sensitivity and standardized assessments: New evidence on the relationship [ Doctor of Philosophy, University of Iowa ]. https://doi.org/10.17077/etd.gyavem36</bibtext> </blist> <blist> <bibtext> Kolen, M. J., &amp; Brennan, R. L. (2014). Test equating, scaling, and linking (3rd ed.). New York : Springer. https://doi.org/10.1007/978‐1‐4939‐0317‐7</bibtext> </blist> <blist> <bibtext> Li, Y., &amp; Lissitz, R. W. (2012). Exploring the full‐information bifactor model in vertical scaling with construct shift. Applied Psychological Measurement, 36 (1), 3 – 20. https://doi.org/10.1177/0146621611432864</bibtext> </blist> <blist> <bibtext> Loyd, B. H., &amp; Hoover, H. D. (1980). Vertical equating using the Rasch model. Journal of Educational Measurement, 17 (3), 179 – 193. https://doi.org/10.1111/j.1745‐3984.1980.tb00825.x</bibtext> </blist> <blist> <bibtext> Marco, G. L. (1977). Item characteristic curve solutions to three intractable testing problems (ETS Research Bulletin Series). https://onlinelibrary.wiley.com/doi/10.1002/j.2333‐8504.1977.tb01136.x</bibtext> </blist> <blist> <bibtext> Martineau, J. A. (2006). Distorting value added: The use of longitudinal, vertically scaled student achievement data for growth‐based, value‐added accountability. Journal of Educational and Behavioral Statistics, 31 (1), 35 – 62. https://doi.org/10.3102/10769986031001035</bibtext> </blist> <blist> <bibtext> May, H., Perez‐Johnson, I., Haimson, J., Sattar, S., &amp; Gleason, P. (2009). Using state tests in education experiments: A discussion of the issues (Technical Methods Report No. NCEE 2009013). National Center for Education Evaluation.</bibtext> </blist> <blist> <bibtext> Najarian, M., Pollack, J. M., &amp; Sorongon, A. G. (2009). Early childhood longitudinal study, kindergarten class of 1998–99 (ECLS‐K): Psychometric report for the eighth grade. National Center for Education Statistics. https://doi.org/10.1037/e600812011‐001</bibtext> </blist> <blist> <bibtext> Patz, R. J. (2007). Vertical scaling in standards‐based educational assessment and accountability systems. Council of Chief State School Officers.</bibtext> </blist> <blist> <bibtext> Petersen, N. S., Kolen, M. J., &amp; Hoover, H. D. (1989). Scaling, norming and equating. In R. L. Linn (Ed.), Educational measurement (3rd ed., pp. 221 – 262). Macmillan Publishing Company.</bibtext> </blist> <blist> <bibtext> Core Team R. (2023). R: A language and environment for statistical computing [Computer software]. R Foundation for Statistical Computing. https://www.R‐project.org/</bibtext> </blist> <blist> <bibtext> Rasch, G. (1960). Probabilistic models for some intelligence and attainment tests. Danish Institute for Educational Research.</bibtext> </blist> <blist> <bibtext> Reckase, M. D. (2010). Study of best practices for vertical scaling and standard setting with recommendations for FCAT 2.0. https://<ulink href="http://www.fldoe.org/core/fileparse.php/5663/urlt/0086369&amp;#8208;studybestpracticesverticalscalingstandardsetting.pdf">www.fldoe.org/core/fileparse.php/5663/urlt/0086369&amp;#8208;studybestpracticesverticalscalingstandardsetting.pdf</ulink></bibtext> </blist> <blist> <bibtext> State of Texas Assessments of Academic Readiness (STAAR) vertical scale technical report. (2013). https://tea.texas.gov/sites/default/files/2013‐STAARVerticalScaleTechReport.pdf</bibtext> </blist> <blist> <bibtext> Stocking, M. L., &amp; Lord, F. M. (1983). Developing a common metric in item response theory. Applied Psychological Measurement, 7 (2), 201 – 210. https://doi.org/10.1177/014662168300700208</bibtext> </blist> <blist> <bibtext> Strachan, T., Cho, U. H., Kim, K. Y., Willse, J. T., Chen, S., Ip, E. H., Ackerman, T. A., &amp; Weeks, J. P. (2021). Using a projection IRT method for vertical scaling when construct shift is present. Journal of Educational Measurement, 58 (2), 211 – 235. https://doi.org/10.1111/jedm.12278</bibtext> </blist> <blist> <bibtext> Student, S. R. (2024). Growth on 2019 state achievement tests: Empirical benchmarks and the role of scale choice. Journal of Research on Educational Effectiveness, 1 – 27. https://doi.org/10.1080/19345747.2024.2360534</bibtext> </blist> <blist> <bibtext> Tao, S., &amp; Mix, D. (2016). Assessing the impacts of using off‐grade items in adaptive testing—A differential item functioning approach. Annual Meeting of the National Council on Measurement in Education, Washington, D.C.</bibtext> </blist> <blist> <bibtext> Thurstone, L. L. (1927). The unit of measurement in educational scales. Journal of Educational Psychology, 18 (8), 505 – 524. https://doi.org/10.1037/h0075524</bibtext> </blist> <blist> <bibtext> Tong, Y., &amp; Kolen, M. J. (2007). Comparisons of methodologies and results in vertical scaling for educational achievement tests. Applied Measurement in Education, 20 (2), 227 – 253. https://doi.org/10.1080/08957340701301207</bibtext> </blist> <blist> <bibtext> van der Linden, W. J., &amp; Barrett, M. D. (2016). Linking item response model parameters. Psychometrika, 81 (3), 650 – 673. https://doi.org/10.1007/s11336‐015‐9469‐6</bibtext> </blist> <blist> <bibtext> Waterbury, G. T., &amp; DeMars, C. E. (2021). Anchors aweigh: How the choice of anchor items affects the vertical scaling of 3PL data with the Rasch Model. Educational Assessment, 26 (3), 175 – 197. https://doi.org/10.1080/10627197.2020.1858782</bibtext> </blist> <blist> <bibtext> Weeks, J. P. (2018). An application of multidimensional vertical scaling. Measurement: Interdisciplinary Research and Perspectives, 16 (3), 139 – 154. https://doi.org/10.1080/15366367.2018.1502005</bibtext> </blist> <blist> <bibtext> Wilson, M. (2023). Constructing measures: An item response modeling approach (2nd ed.). Routledge.</bibtext> </blist> <blist> <bibtext> Wright, B. D., &amp; Masters, G. N. (1990). Computation of OUTFIT and INFIT statistics. Rasch Measurement Transactions, 3 (4), 84 – 85.</bibtext> </blist> <blist> <bibtext> Wright, B. D., &amp; Stone, M. H. (1979). Best test design. MESA Press.</bibtext> </blist> <blist> <bibtext> Wu, M., &amp; Adams, R. J. (2013). Properties of Rasch residual fit statistics. Journal of Applied Measurement, 14 (4), 339 – 355.</bibtext> </blist> <blist> <bibtext> Yen, W. M. (1985). Increasing item complexity: A possible cause of scale shrinkage for unidimensional item response theory. Psychometrika, 50 (4), 399 – 410. https://doi.org/10.1007/BF02296259</bibtext> </blist> <blist> <bibtext> Yen, W. M. (1986). The choice of scale for educational measurement: An IRT perspective. Journal of Educational Measurement, 23 (4), 299 – 325. https://doi.org/10.1111/j.1745‐3984.1986.tb00252.x</bibtext> </blist> <blist> <bibtext> Yen, W. M., &amp; Burket, G. R. (1997). Comparison of Item Response Theory and Thurstone methods of vertical scaling. Journal of Educational Measurement, 34 (4), 293 – 313. https://doi.org/10.1111/j.1745‐3984.1997.tb00520.x</bibtext> </blist> </ref> <aug> <p>By Sanford R. Student; Derek C. Briggs and Laurie Davis</p> <p>Reported by Author; Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib10" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib44" firstref="ref2"></nolink> <nolink nlid="nl3" bibid="bib53" firstref="ref3"></nolink> <nolink nlid="nl4" bibid="bib54" firstref="ref4"></nolink> <nolink nlid="nl5" bibid="bib37" firstref="ref5"></nolink> <nolink nlid="nl6" bibid="bib26" firstref="ref6"></nolink> <nolink nlid="nl7" bibid="bib36" firstref="ref7"></nolink> <nolink nlid="nl8" bibid="bib45" firstref="ref9"></nolink> <nolink nlid="nl9" bibid="bib34" firstref="ref10"></nolink> <nolink nlid="nl10" bibid="bib41" firstref="ref11"></nolink> <nolink nlid="nl11" bibid="bib43" firstref="ref12"></nolink> <nolink nlid="nl12" bibid="bib19" firstref="ref14"></nolink> <nolink nlid="nl13" bibid="bib28" firstref="ref15"></nolink> <nolink nlid="nl14" bibid="bib29" firstref="ref16"></nolink> <nolink nlid="nl15" bibid="bib39" firstref="ref17"></nolink> <nolink nlid="nl16" bibid="bib25" firstref="ref19"></nolink> <nolink nlid="nl17" bibid="bib38" firstref="ref21"></nolink> <nolink nlid="nl18" bibid="bib46" firstref="ref26"></nolink> <nolink nlid="nl19" bibid="bib17" firstref="ref27"></nolink> <nolink nlid="nl20" bibid="bib11" firstref="ref28"></nolink> <nolink nlid="nl21" bibid="bib20" firstref="ref29"></nolink> <nolink nlid="nl22" bibid="bib27" firstref="ref30"></nolink> <nolink nlid="nl23" bibid="bib30" firstref="ref31"></nolink> <nolink nlid="nl24" bibid="bib18" firstref="ref32"></nolink> <nolink nlid="nl25" bibid="bib22" firstref="ref35"></nolink> <nolink nlid="nl26" bibid="bib23" firstref="ref36"></nolink> <nolink nlid="nl27" bibid="bib14" firstref="ref38"></nolink> <nolink nlid="nl28" bibid="bib49" firstref="ref41"></nolink> <nolink nlid="nl29" bibid="bib51" firstref="ref42"></nolink> <nolink nlid="nl30" bibid="bib16" firstref="ref44"></nolink> <nolink nlid="nl31" bibid="bib35" firstref="ref47"></nolink> <nolink nlid="nl32" bibid="bib13" firstref="ref48"></nolink> <nolink nlid="nl33" bibid="bib48" firstref="ref51"></nolink> <nolink nlid="nl34" bibid="bib15" firstref="ref54"></nolink> <nolink nlid="nl35" bibid="bib21" firstref="ref55"></nolink> <nolink nlid="nl36" bibid="bib12" firstref="ref57"></nolink> <nolink nlid="nl37" bibid="bib52" firstref="ref58"></nolink> <nolink nlid="nl38" bibid="bib33" firstref="ref61"></nolink> <nolink nlid="nl39" bibid="bib32" firstref="ref66"></nolink> <nolink nlid="nl40" bibid="bib24" firstref="ref68"></nolink> <nolink nlid="nl41" bibid="bib50" firstref="ref70"></nolink> <nolink nlid="nl42" bibid="bib42" firstref="ref77"></nolink> <nolink nlid="nl43" bibid="bib31" firstref="ref90"></nolink> <nolink nlid="nl44" bibid="bib40" firstref="ref98"></nolink> <nolink nlid="nl45" bibid="bib47" firstref="ref99"></nolink> |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1460458 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Growth across Grades and Common Item Grade Alignment in Vertical Scaling Using the Rasch Model – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Sanford+R%2E+Student%22">Sanford R. Student</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-7419-2437">0000-0001-7419-2437</externalLink>)<br /><searchLink fieldCode="AR" term="%22Derek+C%2E+Briggs%22">Derek C. Briggs</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-1628-4661">0000-0003-1628-4661</externalLink>)<br /><searchLink fieldCode="AR" term="%22Laurie+Davis%22">Laurie Davis</searchLink> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Educational+Measurement%3A+Issues+and+Practice%22"><i>Educational Measurement: Issues and Practice</i></searchLink>. 2025 44(1):84-95. – Name: Avail Label: Availability Group: Avail Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 12 – Name: DatePubCY Label: Publication Date Group: Date Data: 2025 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Audience Label: Education Level Group: Audnce Data: <searchLink fieldCode="EL" term="%22Elementary+Education%22">Elementary Education</searchLink><br /><searchLink fieldCode="EL" term="%22Junior+High+Schools%22">Junior High Schools</searchLink><br /><searchLink fieldCode="EL" term="%22Middle+Schools%22">Middle Schools</searchLink><br /><searchLink fieldCode="EL" term="%22Secondary+Education%22">Secondary Education</searchLink> – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Elementary+School+Mathematics%22">Elementary School Mathematics</searchLink><br /><searchLink fieldCode="DE" term="%22Elementary+School+Students%22">Elementary School Students</searchLink><br /><searchLink fieldCode="DE" term="%22Middle+School+Mathematics%22">Middle School Mathematics</searchLink><br /><searchLink fieldCode="DE" term="%22Middle+School+Students%22">Middle School Students</searchLink><br /><searchLink fieldCode="DE" term="%22Growth+Models%22">Growth Models</searchLink><br /><searchLink fieldCode="DE" term="%22Achievement+Gains%22">Achievement Gains</searchLink><br /><searchLink fieldCode="DE" term="%22Scaling%22">Scaling</searchLink><br /><searchLink fieldCode="DE" term="%22Rating+Scales%22">Rating Scales</searchLink><br /><searchLink fieldCode="DE" term="%22Error+of+Measurement%22">Error of Measurement</searchLink><br /><searchLink fieldCode="DE" term="%22Mathematics+Achievement%22">Mathematics Achievement</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Reliability%22">Test Reliability</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Validity%22">Test Validity</searchLink><br /><searchLink fieldCode="DE" term="%22Grade+Level+Differences%22">Grade Level Differences</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1111/emip.12639 – Name: ISSN Label: ISSN Group: ISSN Data: 0731-1745<br />1745-3992 – Name: Abstract Label: Abstract Group: Ab Data: Vertical scales are frequently developed using common item nonequivalent group linking. In this design, one can use upper-grade, lower-grade, or mixed-grade common items to estimate the linking constants that underlie the absolute measurement of growth. Using the Rasch model and a dataset from Curriculum Associates' i-Ready Diagnostic in math in grades 3-7, we demonstrate how grade-to-grade mean differences in mathematics proficiency appear much larger when upper-grade linking items are used instead of lower-grade items, with linkings based on a mixture of items falling in between. We then consider salient properties of the three calibrated scales including invariance of the different sets of common items to student grade and item difficulty reversals. These exploratory analyses suggest that upper-grade common items in vertical scaling are more subject to threats to score comparability across grades, even though these items also tend to imply the most growth. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2025 – Name: AN Label: Accession Number Group: ID Data: EJ1460458 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1460458 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1111/emip.12639 Languages: – Text: English PhysicalDescription: Pagination: PageCount: 12 StartPage: 84 Subjects: – SubjectFull: Elementary School Mathematics Type: general – SubjectFull: Elementary School Students Type: general – SubjectFull: Middle School Mathematics Type: general – SubjectFull: Middle School Students Type: general – SubjectFull: Growth Models Type: general – SubjectFull: Achievement Gains Type: general – SubjectFull: Scaling Type: general – SubjectFull: Rating Scales Type: general – SubjectFull: Error of Measurement Type: general – SubjectFull: Mathematics Achievement Type: general – SubjectFull: Test Reliability Type: general – SubjectFull: Test Validity Type: general – SubjectFull: Grade Level Differences Type: general Titles: – TitleFull: Growth across Grades and Common Item Grade Alignment in Vertical Scaling Using the Rasch Model Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Sanford R. Student – PersonEntity: Name: NameFull: Derek C. Briggs – PersonEntity: Name: NameFull: Laurie Davis IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 03 Type: published Y: 2025 Identifiers: – Type: issn-print Value: 0731-1745 – Type: issn-electronic Value: 1745-3992 Numbering: – Type: volume Value: 44 – Type: issue Value: 1 Titles: – TitleFull: Educational Measurement: Issues and Practice Type: main |
| ResultId | 1 |