Are Online and Paper Tests Comparable? Evidence from Statewide K-12 Tests

Saved in:
Bibliographic Details
Title: Are Online and Paper Tests Comparable? Evidence from Statewide K-12 Tests
Language: English
Authors: Ben Backes, James Cowan
Source: Grantee Submission. 2024.
Peer Reviewed: Y
Page Count: 30
Publication Date: 2024
Sponsoring Agency: Institute of Education Sciences (ED)
National Center for Analysis of Longitudinal Data in Education Research (CALDER) at American Institutes for Research (AIR)
Contract Number: R305A170119
Document Type: Reports - Research
Education Level: Elementary Secondary Education
Descriptors: Computer Assisted Testing, Test Format, Differences, Academic Achievement, State Programs, Standardized Tests, Elementary Secondary Education, Grade Prediction
Geographic Terms: Massachusetts
Assessment and Survey Identifiers: Massachusetts Comprehensive Assessment System
DOI: 10.1080/08957347.2024.2311933
Abstract: We investigate two research questions using a recent statewide transition from paper to computer-based testing: first, the extent to which test mode effects found in prior studies can be eliminated in large-scale administration; and second, the degree to which online and paper assessments offer different information about underlying student ability. In contrast to the first test transition in Massachusetts, we find very small mode effects for a more recent transition, which may be attributable to an additional step matching on observable characteristics in the equating process. Second, we investigate the predictive evidence of validity for paper and online tests for predictions of future test scores and grades. We generally find minimal differences for the extent to which scores on paper tests can differentially predict future online versus paper test scores. Finally, online and paper test scores are similarly predictive of future grade point average. We conclude that the online test penalty can vary substantially by test and that extreme care should be taken when administering online tests to some students and paper tests to others.
Abstractor: As Provided
IES Funded: Yes
Entry Date: 2024
Accession Number: ED647304
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwFgcn1XnLwp1C4zCLElayfLAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDANEsCoFLjvXtafk9AIBEICBm6jNJMy0w7LRIW4ZraKPwrOysVWSyA_WSTCN_L4BlMSNXdRR9iBo_y0g_KtiJjaIOSKc0wkLB9i9VEMLADNXqaPUOlXYHYITExQ331HOkFYmISdVEj_Guhh9-LrsCT7yJ3mDYKmnMq8x6BEMrO0BoRvCm5PuGuw_sixt6ZjdCz5oVO9Q251CvsoLWPjCwfLZTlzIqn1VWIUVRRE8
Text:
  Availability: 1
  Value: <anid>AN0175570130;7lg01jan.24;2024Feb23.05:30;v2.2.500</anid> <title id="AN0175570130-1">Are Online and Paper Tests Comparable? Evidence from Statewide K-12 Tests </title> <p>We investigate two research questions using a recent statewide transition from paper to computer-based testing: first, the extent to which test mode effects found in prior studies can be eliminated; and second, the degree to which online and paper assessments offer different information about underlying student ability. We first find very small mode effects for a more recent transition in Massachusetts. Second, we investigate the predictive evidence of validity for paper and online tests for predictions of future test scores and grades. We generally find minimal differences for the extent to which scores on paper tests can differentially predict future online versus paper test scores. Finally, online and paper test scores are similarly predictive of future grade point average. We conclude that the online test penalty can vary substantially by test and that extreme care should be taken when administering online tests to some students and paper tests to others.</p> <p>Testing occupies a central role in the measurement of student learning. In many states, the results of standardized assessments support teacher evaluation, school accountability determinations, student graduation, or the distribution of school resources. Recent years have seen a rapid transition to online testing, and the vast majority of states' test administrations are partially or fully online (Olson, [<reflink idref="bib15" id="ref1">15</reflink>]). In addition, as schools across the country transitioned to partial or full remote instruction during the COVID-19 pandemic, online assessments were used to assess student learning more than ever before. In particular, much of what was known about the extent of learning loss suffered early in the pandemic came from online assessments (Brody & Koh, [<reflink idref="bib4" id="ref2">4</reflink>]; Kuhfeld, Tarasawa, Johnson, Ruzek, & Lewis, [<reflink idref="bib11" id="ref3">11</reflink>]).</p> <p>With the prevalence of online testing rising over time, it is important to understand how the administration of tests on computers affects measurement of student learning. To offer guidance, we provide evidence on two important questions related to recent findings spanning multiple states that have uniformly found significant online test penalties: that is, students scoring lower when taking a test on a computer than similar students taking the same test on paper (Backes & Cowan, [<reflink idref="bib1" id="ref4">1</reflink>]; Gordanier, Ozturk, & Zhan, [<reflink idref="bib8" id="ref5">8</reflink>]). The first portion of this paper provides additional evidence about the effects of test mode on measured student achievement using the transition to a new assessment administered to some students online and others on paper.</p> <p>The second question concerns the degree to which online and paper assessments offer different information about student ability when administered at scale. In this setting, there is little evidence about the extent to which online and paper tests measure the same components of student learning. For example, online tests could measure, in part, students' familiarity with computers as these tests often require students to navigate computer programs and use editing tools to respond to open-ended questions. White, Kim, Chen, and Liu ([<reflink idref="bib23" id="ref6">23</reflink>]) and Sandene et al. ([<reflink idref="bib18" id="ref7">18</reflink>]) have found that prior computer experience predicted student performance in pilots of the National Assessment of Educational Progress in mathematics and English Language Arts (ELA). Goldberg and Pedulla ([<reflink idref="bib7" id="ref8">7</reflink>]) documented similar patterns for some online versions of the Graduate Record Examinations.</p> <p>In this paper, we make two primary contributions. The first is to examine a statewide administration of a K-12 assessment in which the equating process between online and paper scores included a step that matched students based on prior test scores to ensure comparability of the equating samples. In contrast to prior studies finding large online test mode penalties (Backes & Cowan, [<reflink idref="bib1" id="ref9">1</reflink>]; Gordanier, Ozturk, & Zhan, [<reflink idref="bib8" id="ref10">8</reflink>]), we find that this equating process resulted in minimal mode effects. As discussed in more detail below, we attribute the diminished mode effects relative to the assessment administered by the Partnership for Assessment of Readiness for College and Careers (PARCC) to post hoc adjustments to the testing scale through this more complete matching based on prior student ability. Second, we examine the relative predictive evidence of validity of online and paper test scores over outcomes such as future test scores and future course grades.</p> <hd id="AN0175570130-2">1. Background</hd> <p></p> <hd id="AN0175570130-3">1.1. Prior Literature</hd> <p>There is an existing literature that examines several national tests when they were converted from pencil and paper to online. Powers ([<reflink idref="bib17" id="ref11">17</reflink>]) examined the GRE scores of paper and online test takers, with minimal differences across test mode. However, the online test was computer adaptive – and thus not directly comparable – and the study participants had already selected their type of mode. In a random assignment study, Schaeffer et al. ([<reflink idref="bib19" id="ref12">19</reflink>]) also found that examinees who answered all GRE questions scored similarly on the paper and online versions of the test. In another random assignment study, Bennett et al. ([<reflink idref="bib3" id="ref13">3</reflink>]) compared paper and online NAEP math scores, finding an online penalty of 0.14 standard deviations. Bennett et al. ([<reflink idref="bib3" id="ref14">3</reflink>]) found that the differences are largely driven by items which were adapted to online and paper differently (i.e., answering in the online version had additional steps). For NAEP writing scores, there was generally no difference found between online and paper (Horkay, Bennett, Allen, Kaplan, & Yan, [<reflink idref="bib9" id="ref15">9</reflink>]). The ACT is a notable exception to the pattern of students generally scoring lower online than on paper. In another random assignment study, Steedle, Pashley, and Cho ([<reflink idref="bib21" id="ref16">21</reflink>]) found large positive online effects for reading, English, and writing, which motivated ACT to equate scores across modes.</p> <p>In a meta-analysis of computerized tests at the K–12 level, Wang, Jiao, Young, Brooks, and Olson ([<reflink idref="bib22" id="ref17">22</reflink>]) concluded that the average study finds that students taking a paper test score about 10% of a standard deviation higher than those taking a CBT. In addition, effects tend to be larger on linear (as opposed to computer-adaptive) tests, such as the ones under consideration in this study.</p> <p>A limitation of the studies in Wang, Jiao, Young, Brooks, and Olson ([<reflink idref="bib22" id="ref18">22</reflink>]) meta analysis and the GRE, NAEP, and ACT results discussed above is that the samples tended to be smaller. The largest studies cited in Wang, Jiao, Young, Brooks, and Olson ([<reflink idref="bib22" id="ref19">22</reflink>]) had fewer than 7,000 students. We are aware of two studies that have examined a statewide test transition with much larger sample sizes. Backes and Cowan ([<reflink idref="bib1" id="ref20">1</reflink>]) and Gordanier, Ozturk, and Zhan ([<reflink idref="bib8" id="ref21">8</reflink>]) each used recent test transitions and found large negative effects associated with taking exams online relative to on paper.</p> <hd id="AN0175570130-4">1.2. State Context</hd> <p>This study covers seven years (2012 through 2018) in which Massachusetts administered three different statewide standardized tests. Massachusetts adopted new state curriculum frameworks incorporating the Common Core State Standards in 2011, with implementation beginning in the 2012–13 school year (from here forward, we will refer to school years by the spring of the year under discussion).[<reflink idref="bib1" id="ref22">1</reflink>] Until 2014, all districts used the Massachusetts Comprehensive Assessment System (MCAS), which was administered on paper and referred to as MCAS 1.0 throughout. In 2015 and 2016, districts chose between MCAS 1.0 and the new PARCC assessment, with three districts having a mix of online and paper tests.[<reflink idref="bib2" id="ref23">2</reflink>] About 72% of elementary or middle schools in our sample administered PARCC in either 2015 or 2016. PARCC districts had the additional option of offering the test online or on paper. Of those schools administering PARCC in either 2015 or 2016, 57% administered the test online at least once. Finally, in 2017 and 2018, all schools switched to MCAS 2.0, which was initially a mix of online and paper, with the online version rolled out over the first three years. In 2017, the 4th- and 8th-grade tests were online only and the remaining tests were online optional. The 5th- and 7th-grade tests became online only in 2018, and the 3rd- and 6th-grade tests followed in 2019.</p> <p>The online versions of both PARCC and MCAS 2.0 are linear (i.e., not computer adaptive) and are meant to be similar to the paper versions of the respective assessments. For PARCC, the paper versions were adapted from the online forms and used a similar set of items. That said, the online versions of the test included some interactive questions not present in the paper version, meaning that the paper and online versions were not exactly equivalent. In addition, there are other differences between test formats. For example, Figure 1, displays reading passages from the sample PARCC assessment's paper and online formats. The paper version of the test (Figure 1a) displays reading passages across multiple pages in the test booklet, while the online version (Figure 1b) displays the full passage in a box embedded in a single page with multiple-choice questions. Other differences include dropdown menus in the online format (Figure 2) and the ability to type essay responses (Figure 3).[<reflink idref="bib3" id="ref24">3</reflink>]</p> <p>Graph: Figure 1. Reading passage display formats on online and paper assessments.</p> <p>Graph: Figure 2. Multiple-choice question display formats on online and paper assessments.</p> <p>Graph: Figure 3. Free-response question formats on online and paper assessments.</p> <p>Prior work in Massachusetts has compared changes in achievement in schools adopting online tests to those administering tests by paper and found strong evidence of a negative causal online mode effect in both math and ELA scaled scores (Backes & Cowan, [<reflink idref="bib1" id="ref25">1</reflink>]). Of the two PARCC years in the study, the largest test mode penalties were found in the first year; however, mode effects continued into the second year. This penalty did not appear to be explained by differences in school context – the schools that switched to online testing were disproportionately high achieving – and the analysis passed a series of placebo tests. For example, schools that switched to the online PARCC assessment and saw a drop in math and ELA scores did not see a corresponding drop in grades 4 and 8 science, which was administered on paper to all schools throughout the time period.</p> <hd id="AN0175570130-5">2. Data and Methods</hd> <p>We constructed a panel of test scores for students in grades 3–8 and 10 who took the PARCC, MCAS 1.0, or MCAS 2.0 tests between 2012 and 2018.[<reflink idref="bib4" id="ref26">4</reflink>] Massachusetts offered MCAS 1.0 from 2011 to 2016, PARCC in 2015 and 2016, and MCAS 2.0 in 2017 and 2018. Generally speaking, school districts selected a single test and mode for both of the PARCC years, but some districts switched between tests and modes following the 2015 school year. In addition, as discussed above, three large school districts were allowed to choose the mode separately for each school. The high school MCAS is a state graduation requirement and Massachusetts did not administer the PARCC high school assessment online. Massachusetts switched to an online version of the 10<sups>th</sups> grade MCAS assessment in 2019 (after the data used in this study).</p> <hd id="AN0175570130-6">2.1. Test Mode Equating Procedures</hd> <p>Neither PARCC nor MCAS 2.0 had access to randomly assigned pools of online and paper test takers to use for performing test mode equating or for studying mode comparability. For PARCC, both modes included a subset of linked items common across the tests to facilitate the reporting of student scores on a common scale (Educational Testing Service, Pearson, & Measured Progress, [<reflink idref="bib6" id="ref27">6</reflink>]; Pearson, [<reflink idref="bib16" id="ref28">16</reflink>]). Following the administration of the test, PARCC scored the tests for each mode separately and then transformed results from the paper tests onto the online scale, using results from the common set of linked items. The scores were therefore intended to be comparable across modes, and a mode comparability study found that "When comparing the performance on the common items, the effect sizes ranged from negligible to small for most of the tests evaluated (Educational Testing Service, Pearson, & Measured Progress, [<reflink idref="bib6" id="ref29">6</reflink>], p. 141)." However, in the one state for which the study obtained prior performance data, the study found that the paper and online samples were not comparable in terms of prior achievement.</p> <p>The Massachusetts Department of Elementary and Secondary Education also used a set of linked items for the equating process in MCAS 2.0, but in addition, subsequently adjusted the scale scores using a sample of test takers matched on observable characteristics, including prior test scores on a given test. For example, a student who took online PARCC in 2016 and online MCAS 2.0 in 2017 would be matched to a student who took paper MCAS 2.0 in 2017 with similar online PARCC scores in 2016 in order to ensure comparability of prior-year scores when equating. The creation of the paper forms involved using the computer-based test as a starting point and then replacing all technology-based items with items that could be administered on paper, with examples of such items being dropdown menus and "drag-and-drop" tasks. These replacement items were intended to be aligned to a similar content standard as the online-only items (DESE, [<reflink idref="bib13" id="ref30">13</reflink>]).</p> <p>We used the scaled scores for this analysis, which were the only scores available for both MCAS and PARCC assessments. The Massachusetts data included both the PARCC scale scores and a conversion of the PARCC scale scores to the MCAS scale.[<reflink idref="bib5" id="ref31">5</reflink>] We used the converted scale scores for the main analyses, but the results were not sensitive to using the original PARCC scales instead. We then standardized test scores by grade and year. In some years, Massachusetts applied a nonlinear transformation to the student theta scores to obtain the MCAS scales, so we also tested the sensitivity of the results to using a normal curve equivalent transformation of the scale scores (Jacob & Rothstein, [<reflink idref="bib10" id="ref32">10</reflink>]).</p> <p>For each grade and school in the sample, we constructed a panel of data spanning 2012 through 2018 with (<reflink idref="bib1" id="ref33">1</reflink>) average current test scores in math and ELA, (<reflink idref="bib2" id="ref34">2</reflink>) the average two-year gain score (i.e., the difference between current test scores and twice-lagged test scores), and (<reflink idref="bib3" id="ref35">3</reflink>) the correlation between students' current scores and twice-lagged test scores. We used the first two variables to assess mode effects and the final variable to assess whether online tests measure, in part, students' online test-taking ability. We used the twice-lagged scores in the construction of (<reflink idref="bib2" id="ref36">2</reflink>) and (<reflink idref="bib3" id="ref37">3</reflink>) so that current test results during each transition are compared to tests using the previous version of the test.[<reflink idref="bib6" id="ref38">6</reflink>]</p> <hd id="AN0175570130-7">2.2. Empirical Strategy for Test Mode Effects</hd> <p>We begin by describing our empirical strategy for estimating test mode effects on MCAS 2.0. Our research design relied on the rollout of the online PARCC and online MCAS 2.0. We used the entire panel of test scores from 2012 to 2018 and estimated models comparing test scores in grades and schools switching to online tests to those remaining on paper tests. Formally, we estimated</p> <p>(<reflink idref="bib1" id="ref39">1</reflink>)</p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>Y</mi><mrow><mi mathvariant="italic">igst</mi></mrow></msub></mrow><mo>=</mo><mrow><msub><mi>X</mi><mrow><mi mathvariant="italic">igst</mi></mrow></msub></mrow><mi mathvariant="italic">β</mi><mo>+</mo><mi mathvariant="italic">Onlin</mi><mrow><msub><mi>e</mi><mrow><mi mathvariant="italic">gst</mi></mrow></msub></mrow><mi mathvariant="italic">δ</mi><mo>+</mo><mrow><msub><mi>α</mi><mrow><mi mathvariant="italic">gs</mi></mrow></msub></mrow><mo>+</mo><mrow><msub><mi>λ</mi><mi>t</mi></msub></mrow><mo>+</mo><mrow><msub><mtext mathcolor="red">\isin</mtext><mrow><mi mathvariant="italic">igst</mi></mrow></msub></mrow><mo>,</mo></math> </ephtml> </p> <p>a model with grade-by-school fixed effects (</p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>α</mi><mrow><mi mathvariant="italic">gs</mi></mrow></msub></mrow></math> </ephtml> ), test-by-year fixed effects (</p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>λ</mi><mi>t</mi></msub></mrow></math> </ephtml> ), student characteristics (</p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>X</mi><mrow><mi mathvariant="italic">igst</mi></mrow></msub></mrow></math> </ephtml> ), and an error term</p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext mathcolor="red">\isin</mtext><mrow><mi mathvariant="italic">igst</mi></mrow></msub></mrow></math> </ephtml> . In addition, the term</p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">Onlin</mi><mrow><msub><mi>e</mi><mrow><mi mathvariant="italic">gst</mi></mrow></msub></mrow></math> </ephtml> indicates whether the student took the test online, and</p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>Y</mi><mrow><mi mathvariant="italic">igst</mi></mrow></msub></mrow></math> </ephtml> is some outcome variable for student <emph>i</emph> in year <emph>t</emph>. The research design incorporated two sources of variation in test mode to estimate the coefficient of interest,</p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">δ</mi></math> </ephtml> : differences across grades in the application of the PARCC or MCAS 2.0 online testing mandate and differences within grades between schools voluntarily switching to the online format.[<reflink idref="bib7" id="ref40">7</reflink>]</p> <p>The difference-in-differences design assumes that "online schools" would have followed the same achievement trajectory as "paper schools" in the absence of the test mode switch. Backes and Cowan ([<reflink idref="bib1" id="ref41">1</reflink>]) assessed the plausibility of this assumption in the context of the initial switch from MCAS 1.0 to PARCC. They found no evidence of differential pretreatment trends among PARCC online and PARCC paper schools, and found no effect of the mode switch on a science test that was administered on paper throughout the PARCC testing era in Massachusetts. We further investigated the plausibility of the identifying assumptions for measuring the mode effect under MCAS 2.0 in Figure 4. We estimated a version of Eq. (<reflink idref="bib1" id="ref42">1</reflink>) that replaced the online MCAS indicator with indicators for year relative to the first administration of the online MCAS (at the school-by-grade level). We plot the effects of online MCAS 2.0 testing relative to the year prior to implementation in Figure 4. We do not find any evidence of differential pretreatment trends among the online testing cells. Offering a preview of the results, the magnitudes of test mode effects after the transition (x-axis values of 0 and 1) are very small for both ELA and math relative to prior studies examining statewide transitions (Backes & Cowan, [<reflink idref="bib1" id="ref43">1</reflink>]; Gordanier, Ozturk, & Zhan, [<reflink idref="bib8" id="ref44">8</reflink>]).</p> <p>Graph: Figure 4. Test mode effects by year relative to implementation.</p> <p>We additionally estimated models that isolated each of these distinct sources of variation in test mode. If the online testing mandates were associated with other changes to the test across grades, then we may conflate test difficulty with mode effects. We therefore replaced the year effects with year-by-grade effects. Because these models remove variation within grades and years in online testing status, they compare changes in test scores among schools and grades voluntarily switching to the online test format to those remaining on the paper test. We then restrict the sample to schools that never switched to PARCC. This sample of schools never administered online testing prior to MCAS 2.0 and thus isolates a sample without school-level familiarity with online testing. These estimated effects may be a more realistic expectation for schools switching to online testing for the first time than the combined PARCC-MCAS 1.0 sample.</p> <hd id="AN0175570130-8">2.3. Empirical Strategy for Mode-Specific Relationship Between Current and Future Scores</hd> <p>We then considered whether online tests differentially measure students' academic performance. We assessed differential measurement by comparing the extent to which current scores differed from twice-lagged scores (as measured by Euclidean distance) for online and paper schools. We also constructed a difference measure between test scores and course grades by taking the Euclidean difference between GPA (standardized to be mean 0, standard deviation 1 as with test scores) and twice-lagged test scores. Because several factors can influence the measurement error in the test (e.g., conditions during the test administration, students' academic proficiency), we again relied on a difference-in-differences design to compare changes in distance among students in schools that adopt the online form to those adopting the paper form.</p> <p>We measured test score distances across years and in particular, the extent to which these distances depend on whether year pairs are of the same test mode. The design used test scores two years apart to make comparisons across assessment types (i.e., MCAS 2.0 versus PARCC; and MCAS 2.0 versus MCAS 1.0). We also used twice-lagged data because we used high school tests in some specifications and the MCAS was only administered in grades 8 and 10.</p> <p>Our basic research design relied on test type transitions that also generated variation in test mode between the current year administration and the administration two years prior. The first transition used data form MCAS 1.0 and PARCC. In 2015 and 2016, students in PARCC schools had taken a paper version of the MCAS 1.0 two years prior. We compared differences between current and prior test scores for the "Paper MCAS 1.0 → Paper PARCC" group to changes in the "Paper MCAS 1.0 → Online PARCC" group. Because we made comparisons only among PARCC schools, this transition implicitly controlled for the fact that correlations in scores among different tests may differ (Backes et al., [<reflink idref="bib2" id="ref45">2</reflink>]).</p> <p>The second transition used data from MCAS 1.0 and MCAS 2.0; i.e., schools that never switched to PARCC. As noted above, all students took MCAS 1.0 on paper. For MCAS 2.0, the switch to online testing for MCAS 2.0 was staggered over grades: Grades 4 and 8 in 2017; Grades 5 and 7 in 2018; and Grades 3 and 6 in 2019 (the final not considered in this paper). For this transition, we obtained a sample of students who took MCAS 1.0 on paper and compared those who took MCAS 2.0 online to those who took MCAS 2.0 on paper. Like the first design, this began with a sample of paper test-takers and examined how their test scores changed two years later for future online takers relative to future paper takers.</p> <p>The third source of variation we considered was between the 8<sups>th</sups> grade PARCC scores and the 10<sups>th</sups> grade MCAS scores for students in schools that administered the PARCC assessment in 2015 and 2016. While test mode varied by district for the 8<sups>th</sups> grade PARCC test, the 10<sups>th</sups> grade MCAS was administered on paper forms throughout the state. We used data on the 10<sups>th</sups> grade MCAS and compared changes in the distances between 8<sups>th</sups> and 10<sups>th</sups> grade MCAS scores among students who took the PARCC online test relative to those who took to the PARCC paper test.</p> <p>In each case, we can recover the effect of a mode switch on differences between current and prior test scores (or current GPA and prior test scores) using estimating equations similar to the following:</p> <p>(<reflink idref="bib2" id="ref46">2</reflink>)</p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>D</mi><mrow><mi mathvariant="italic">igs</mi><mo>,</mo><mi mathvariant="italic">t</mi><mo>,</mo><mi mathvariant="italic">t</mi><mo>−</mo><mn>2</mn></mrow></msub></mrow><mo>=</mo><mrow><msub><mi>X</mi><mrow><mi mathvariant="italic">igst</mi></mrow></msub></mrow><mi mathvariant="italic">β</mi><mo>+</mo><mi mathvariant="italic">Mode</mi><mi mathvariant="normal">_switc</mi><mrow><msub><mi>h</mi><mrow><mi mathvariant="italic">gst</mi></mrow></msub></mrow><mi mathvariant="italic">δ</mi><mo>+</mo><mrow><msub><mi>α</mi><mrow><mi mathvariant="italic">gs</mi></mrow></msub></mrow><mo>+</mo><mrow><msub><mi>λ</mi><mi>t</mi></msub></mrow><mo>+</mo><mrow><msub><mtext mathcolor="red">\isin</mtext><mrow><mi mathvariant="italic">igst</mi></mrow></msub></mrow><mo>,</mo></math> </ephtml> </p> <p>where</p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>D</mi><mrow><mi mathvariant="italic">igs</mi><mo>,</mo><mi mathvariant="italic">t</mi><mo>,</mo><mi mathvariant="italic">t</mi><mo>−</mo><mn>2</mn></mrow></msub></mrow></math> </ephtml> represents the Euclidean distance between current and twice-lagged scores, and the model includes student characteristics</p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>X</mi><mrow><mi mathvariant="italic">igst</mi></mrow></msub></mrow></math> </ephtml> , grade-by-school fixed effects (</p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>α</mi><mrow><mi mathvariant="italic">gs</mi></mrow></msub></mrow></math> </ephtml> ), year fixed effects (</p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>λ</mi><mi>t</mi></msub></mrow></math> </ephtml> ), and an error term</p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext mathcolor="red">\isin</mtext><mrow><mi mathvariant="italic">gst</mi></mrow></msub></mrow></math> </ephtml> . In addition,</p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">Mode</mi><mi mathvariant="normal">_</mi><mi mathvariant="italic">switc</mi><mrow><msub><mi>h</mi><mrow><mi mathvariant="italic">gst</mi></mrow></msub></mrow></math> </ephtml> represents whether a student took the test with different modes in time period <emph>t</emph> and in time period <emph>t-2</emph>. All models cluster standard errors at the school level.</p> <p>In Eq. (<reflink idref="bib2" id="ref47">2</reflink>) above, the coefficient</p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">δ</mi></math> </ephtml> represents how the difference between time <emph>t</emph> and time <emph>t − 2</emph> scores differed in years that switched test modes between <emph>t</emph> and <emph>t − 2</emph> and those that did not. For example, when the difference is measured for the grade 10 test as in the third test above – always administered on paper – then</p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">δ</mi></math> </ephtml> measures the difference in Euclidean distance for the "online → paper" sample relative to "paper → paper" sample. In this case, a positive value of</p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">δ</mi></math> </ephtml> would suggest that the relationship between current scores and past scores is weaker (i.e., the scores are further apart) for students whose past scores were obtained from an online assessment.</p> <hd id="AN0175570130-9">3. Results</hd> <p></p> <hd id="AN0175570130-10">3.1. MCAS 2.0 Mode Effects</hd> <p>In this section, we use the more recent test transition in Massachusetts as a piece of evidence about whether scale score penalties from online tests are always large, as they were for the initial years of PARCC and in another state test (Gordanier, Ozturk, & Zhan, [<reflink idref="bib8" id="ref48">8</reflink>]). As described above, in 2017 and 2018, Massachusetts switched to MCAS 2.0. The math and ELA MCAS were administered online in grades 3 through 8 beginning in 2017, with optional paper administrations in some grades.[<reflink idref="bib8" id="ref49">8</reflink>] In 2017, the state required the grades 4 and 8 MCAS to be offered online, and 95% of fourth and eighth graders took the online MCAS. In addition, 43% of primary school students in other grades took the test online. In 2018, Massachusetts required online administration of the grades 5 and 7 MCAS. Nearly 90% of students took the test online that year, with grade 3 (62%) and grade 6 (78%) lagging behind the other grades.</p> <p>Main results for the mode effects investigation are shown in Table 1. Using student-level controls including twice-lagged test scores, we estimated mode effects of −0.02 in math and −0.01 in ELA for MCAS 2.0 (column 3). The confidence intervals are precise enough to rule out mode effects of about −0.05 standard deviations in math and −0.03 standard deviations in ELA, suggesting that the mode effects on MCAS 2.0, if any, were substantially smaller than on the PARCC assessment. In particular, the results PARCC in the same specification were −0.096 in math and −0.226 in ELA. In column 2, we included grade-by-year effects to focus on effect of online testing on measured student performance for schools voluntarily switching to the online format. The estimates were nearly identical to those in column 3. We repeated these analyses in columns 4 through 6 using schools that administered the MCAS 1.0 in both 2015 and 2016. The point estimates were not statistically significant for either test in column 6, although they are quite similar in magnitude on both tests. Notably, because these schools did not offer online testing prior to 2017 (MCAS 1.0 was entirely on paper), these findings indicate that the reduction in the online penalty from PARCC to MCAS 2.0 could not be driven by prior familiarity with online testing.</p> <p>Table 1. Online test penalties for PARCC and MCAS 2.0.</p> <p> <ephtml> <table><thead><tr><td /><td>(1)</td><td>(2)</td><td>(3)</td><td>(4)</td><td>(5)</td><td>(6)</td></tr></thead><tbody><tr><td><italic>Panel A. Math</italic></td><td /><td /><td /><td /><td /><td /></tr><tr><td>Online PARCC</td><td>−0.095***</td><td>−0.096***</td><td>−0.096***</td><td /><td /><td /></tr><tr><td /><td>(0.008)</td><td>(0.008)</td><td>(0.008)</td><td /><td /><td /></tr><tr><td>Online MCAS</td><td>−0.018**</td><td>−0.032***</td><td>−0.022***</td><td>−0.005</td><td>−0.030**</td><td>−0.010</td></tr><tr><td /><td>(0.007)</td><td>(0.009)</td><td>(0.007)</td><td>(0.011)</td><td>(0.013)</td><td>(0.011)</td></tr><tr><td><italic>Panel B. ELA</italic></td><td /><td /><td /><td /><td /><td /></tr><tr><td>Online PARCC</td><td>−0.225***</td><td>−0.226***</td><td>−0.226***</td><td /><td /><td /></tr><tr><td /><td>(0.008)</td><td>(0.008)</td><td>(0.008)</td><td /><td /><td /></tr><tr><td>Online MCAS</td><td>−0.008</td><td>−0.015</td><td>−0.014*</td><td>−0.008</td><td>−0.029*</td><td>−0.014</td></tr><tr><td /><td>(0.008)</td><td>(0.010)</td><td>(0.008)</td><td>(0.012)</td><td>(0.016)</td><td>(0.011)</td></tr><tr><td><italic>N</italic></td><td>2,318,252</td><td>2,318,252</td><td>2,318,252</td><td>669,141</td><td>669,141</td><td>669,141</td></tr><tr><td>Student controls</td><td /><td /><td>Y</td><td /><td /><td>Y</td></tr><tr><td>Grade-year FE</td><td /><td>Y</td><td /><td /><td>Y</td><td /></tr><tr><td>MCAS only</td><td /><td /><td>Y</td><td>Y</td><td>Y</td></tr></tbody></table> </ephtml> </p> <p>1 Notes: Each number represents the coefficient from a regression of standardized (mean zero, standard deviation one) test scores on test mode. The coefficients represent estimated online test mode penalties measured in test score standard deviations. Each regression includes school-by-grade fixed effects and year indicators. Student controls include student race/ethnicity, gender, free- or reduced- price lunch status, special education status, and limited English proficiency status. "PARCC" denotes the assessment created by the Partnership for Assessment of Readiness for College and Careers, while "MCAS" stands for Massachusetts Comprehensive Assessment System. * significant at <emph>p</emph> <.10, ** significant at <emph>p</emph> <.05, *** significant at <emph>p</emph> <.01.</p> <p>Overall, online mode effects appear to be substantially smaller on the MCAS 2.0 than for the PARCC assessments. We estimated mode effects of about −0.014 standard deviations in ELA, and between −0.01 and −0.03 in math. There appear to be two possible explanations for why there is a substantial online penalty for PARCC and not for MCAS 2.0. The first possibility is that the online version of PARCC was genuinely harder due to item design. Although we cannot assess this possibility directly, we note that the MCAS 2.0 did license some items from the PARCC item bank, meaning that there was some overlap between the two assessments (Massachusetts Department of Elementary and Secondary Education, [<reflink idref="bib13" id="ref50">13</reflink>], [<reflink idref="bib14" id="ref51">14</reflink>]). The second is that, as noted above, the equating between the online and paper versions of the tests was conducted differently in the two assessment systems. The Massachusetts Department of Elementary and Secondary Education used a set of linked items as did PARCC, but in addition, performed a subsequent step of adjusting the scale scores using a sample of test takers matched on observable characteristics, including prior test scores on a given test. Notably, the mode effects <emph>after</emph> equating based on linked items but <emph>prior to</emph> the adjustment based on matched comparisons were similar in magnitude to estimates from the PARCC assessments (Backes & Cowan, [<reflink idref="bib1" id="ref52">1</reflink>]; Massachusetts Department of Elementary and Secondary Education, [<reflink idref="bib13" id="ref53">13</reflink>], [<reflink idref="bib14" id="ref54">14</reflink>]), suggesting the need to perform an additional equating step that accounts for prior achievement when the underlying sample of paper and online test takers is dissimilar.</p> <hd id="AN0175570130-11">3.2. Predictive Evidence of Validity of Online and Paper Tests</hd> <p>To provide evidence on how the predictive evidence of validity of online and of paper scores varies by test mode, we examined the predictive evidence of validity of test scores by mode in the three samples of students described above. Results are shown in Table 2.</p> <p>Table 2. Predictive evidence of validity of online and paper tests.</p> <p> <ephtml> <table><thead><tr><td /><td /><td>Math</td><td>ELA</td></tr><tr><td /><td /><td>(1)</td><td>(2)</td><td>(3)</td><td>(4)</td></tr></thead><tbody><tr><td><italic>Panel A. MCAS 1.0 to PARCC</italic></td><td /><td /><td /><td /></tr><tr><td>Paper –> Online</td><td>5-8th Grade Tests</td><td>0.002</td><td>0.001</td><td>−0.008</td><td>−0.008</td></tr><tr><td /><td /><td>(0.005)</td><td>(0.005)</td><td>(0.005)</td><td>(0.005)</td></tr><tr><td>N</td><td /><td>465,606</td><td>465,606</td><td>465,606</td><td>465,606</td></tr><tr><td><italic>Panel B. MCAS 1.0 to MCAS 2.0</italic></td><td /><td /></tr><tr><td>Paper –> Online</td><td>5-8th Grade Tests</td><td>0.010*</td><td>0.010*</td><td>0.012**</td><td>0.011**</td></tr><tr><td /><td /><td>(0.005)</td><td>(0.005)</td><td>(0.005)</td><td>(0.005)</td></tr><tr><td>N</td><td /><td>437,050</td><td>437,050</td><td>437,050</td><td>437,050</td></tr><tr><td><italic>Panel C. 10th Grade Outcomes</italic></td><td /><td /></tr><tr><td>Online –> Paper</td><td>Paper MCAS 1.0</td><td>0.017**</td><td>0.017**</td><td>0.008</td><td>0.008</td></tr><tr><td /><td /><td>(0.008)</td><td>(0.008)</td><td>(0.006)</td><td>(0.006)</td></tr><tr><td>N</td><td /><td>231,905</td><td>231,905</td><td>231,905</td><td>231,905</td></tr><tr><td>GPA</td><td /><td>0.006</td><td>0.006</td><td>0.013</td><td>0.014</td></tr><tr><td /><td /><td>(0.018)</td><td>(0.018)</td><td>(0.018)</td><td>(0.019)</td></tr><tr><td>N</td><td /><td>230,207</td><td>230,207</td><td>230,207</td><td>230,207</td></tr><tr><td>Student controls</td><td /><td /><td>Y</td><td /><td>Y</td></tr></tbody></table> </ephtml> </p> <p>2 Notes: Each number represents the estimated difference in Euclidean distance between standardized test scores in a given year (or, in the bottom panel, standardized grade point average [GPA] in a given year) and test scores from two years prior for students that experienced test mode switches relative to those that did not. Each regression includes school-by-grade fixed effects and year indicators. Additional student-level controls include student race/ethnicity, free- or reduced- price lunch status, special education status, and limited English proficiency status. "PARCC" denotes the assessment created by the Partnership for Assessment of Readiness for College and Careers, while "MCAS" stands for Massachusetts Comprehensive Assessment System. *significant at <emph>p</emph><.10, **significant at <emph>p</emph> <.05, ***significant at <emph>p</emph> <.01.</p> <p>The first sample contains students who took the MCAS 1.0 on paper and switched to paper or online PARCC. Results are shown in panel A. Results for this sample were small and not significant, indicating that the difference in test scores between PARCC scores and MCAS 1.0 scores two years prior was similar for students who had taken online PARCC and students who had taken paper PARCC.</p> <p>The second sample we examined was students in schools who never transitioned to PARCC and thus took the paper MCAS 1.0 in 2015 and 2016. These students took the MCAS 1.0 on paper and then MCAS 2.0 on paper or online. There are two potential sources of variation in test mode for this sample. In 2017 and 2018, only some grades were required to administer the MCAS online. There was therefore some variation across grades and within schools in the timing of the transition to online testing. Second, as with the PARCC, schools had the option of administering the MCAS online in other grades. We therefore observed some variation across schools in the test mode within the same grade and school year. Results for this sample are shown in panel B. The coefficient on transitioning from paper to online is positive in all specifications. However, the practical effect was very small and represents an increase in Euclidean distance of about 0.01 relative to a mean of about 0.50 for the entire sample, a difference of about 2%. Together, the results from Panels A and B, which collectively examine switches from paper tests to either online or paper tests in another assessment, suggest that the ability of test scores on paper to predict future online test scores is very similar to the ability of paper scores to predict future paper scores.</p> <p>Finally, we examined the 10th-grade MCAS assessment among a sample of students whose schools switched to PARCC for 2015 and 2016. The 10th-grade MCAS 2.0 was administered on paper to all students in 2017 and 2018, but the test scores from two years prior (2015 and 2016) were obtained from online PARCC for some students and paper PARCC for other students, depending on if their school administered PARCC online or on paper in a given year. We therefore tested for changes in the difference between 8th and 10th grade scores for students that switched to the online PARCC in 2015 relative to those whose tests remained on paper. Results are shown in panel C. Coefficients were larger and significant in math, giving some suggestive evidence that, relative to paper tests in math, online tests in math are less predictive of performance on future math paper tests.</p> <p>Thus far, results have focused on how past test scores predict future test scores. Panel C conducts an additional test of the relative predictive power of online and paper test scores and measures whether the relationship between test scores in grade 8 and grade point average in grade 10 differs by whether a student took PARCC online or on paper in grade 8. Results are shown in the bottom of panel C. Coefficients were very small and not significant, suggesting that the predictive power of PARCC test scores in grade 8 over students' grades in grade 10 was similar whether the student took the test online or on paper, despite the very large test mode effects for PARCC.</p> <hd id="AN0175570130-12">4. Lessons Learned</hd> <p>Online tests are perceived as an improvement over paper tests, offering "reduced testing time and cost, quicker results, greater access for English language learners and students with disabilities, individualized questions, automated scoring, and technology-enhanced performance tasks that can assess more complex skills" (Olson, [<reflink idref="bib15" id="ref55">15</reflink>]). In this study, we found modest evidence that online tests measure distinct skills relative to paper tests in math: in our sample, paper tests were a better predictor of future paper tests than online tests were. However, these differences were small, did not show up when examining transitions from paper to online testing, and did not appear to carry over to other outcomes like course grades. The findings suggest that the skills measured by online tests, such as the ability to compose documents on a computer, do not predict grades or other more complex academic outcomes. In addition, we found that the online testing penalty found in prior studies can be mitigated in large-scale testing systems. Although both the PARCC and Next-Generation MCAS exhibited mode effects after equating scores using a sample of matched items, the mode effects on the Next-Generation MCAS were significantly reduced through matching procedures.</p> <p>In order for measurable mode effects to appear in our analysis of PARCC, the following four conditions had to be met: (<reflink idref="bib1" id="ref56">1</reflink>) an absence of random assignment to draw upon for test mode equating, (<reflink idref="bib2" id="ref57">2</reflink>) incomparability in the prior achievement of the online and paper samples, (<reflink idref="bib3" id="ref58">3</reflink>) prior achievement not being used in the equating procedure, and (<reflink idref="bib4" id="ref59">4</reflink>) the same test being offered online to some students and on paper to others in the same year. In a sense, it is thus possible that the mode effects found here offer limited generalizability because all four conditions may not be met frequently. On the other hand, because schools and districts typically have input into test mode, obtaining random samples that can be used for test mode equating at scale can be very difficult in practice. In the initial version of the mode comparability study, PARCC began with an intention to perform a random assignment study but encountered practical difficulties, including schools lacking the technical infrastructure to implement online testing and other schools that had already switched to online testing that "did not want to take a step back" (Brown et al., [<reflink idref="bib5" id="ref60">5</reflink>], p. 12). To the extent that schools that have the resources in place to offer online testing are systematically different than those that do not, we might expect conditions (<reflink idref="bib1" id="ref61">1</reflink>) and (<reflink idref="bib2" id="ref62">2</reflink>) to be met in other settings in the future. In particular, we might expect a correlation between a school's technical infrastructure challenges and the socioeconomic status of its student body. This was the case in Massachusetts in the mid-2010s, with the paper PARCC sample being disproportionately likely to consist of free- or reduced-price lunch eligible students in addition to having lower prior test scores in math and ELA (Backes & Cowan, [<reflink idref="bib1" id="ref63">1</reflink>]). One of the main lessons of this study is thus that in the absence of random assignment, it is important to obtain students' prior achievement in order to measure whether online and paper samples are comparable.</p> <p>The experience of Massachusetts' transition to online testing offers two main takeaways. First, it is possible for online and paper assessments to offer very different estimates of student learning. Despite largely similar item design, students who took PARCC online scored substantially worse than similar students who took PARCC on paper. Thus, extreme care should be taken when administering online tests to some students and paper tests to others. This is now especially relevant for schools that are offering a mix of test types by in-person offering. In addition, test vendors should ensure that test scores are equated across modes in a manner that eliminates test mode penalties. One way to maximize the likelihood of the equating process being conducted in a way that yields comparable scores across test modes is to ensure that the samples of online and paper students have similar baseline characteristics and prior test scores measured on a common test (Massachusetts Department of Elementary and Secondary Education, [<reflink idref="bib12" id="ref64">12</reflink>]). Second, even when test mode penalties are present, online and paper tests appear to offer similar value as a diagnostic tool in terms of predicting future student outcomes. Even across test modes and different tests, ELA scores are still highly correlated with future ELA scores, and math scores are highly correlated with future math scores.</p> <hd id="AN0175570130-13">Disclosure statement</hd> <p>No potential conflict of interest was reported by the author(s).</p> <ref id="AN0175570130-14"> <title> References </title> <blist> <bibl id="bib1" idref="ref4" type="bt">1</bibl> <bibtext> Backes, B., & Cowan, J. (2019). Is the pen mightier than the keyboard? The effect of online testing on measured student achievement. Economics of Education Review, 68, 89 – 103. doi: 10.1016/j.econedurev.2018.12.007</bibtext> </blist> <blist> <bibl id="bib2" idref="ref23" type="bt">2</bibl> <bibtext> Backes, B., Cowan, J., Goldhaber, D., Koedel, C., Miller, L. C., & Xu, Z. (2018). The common core conundrum: To what extent should we worry that changes to assessments will affect test-based measures of teacher performance? Economics of Education Review, 62, 48 – 65. doi: 10.1016/j.econedurev.2017.10.004</bibtext> </blist> <blist> <bibl id="bib3" idref="ref13" type="bt">3</bibl> <bibtext> Bennett, R. E., Braswell, J., Oranje, A., Sandene, B., Kaplan, B., & Yan, F. (2008). Does it matter if I take my mathematics test on computer? A second empirical study of mode effects in NAEP. Journal of Technology, Learning, and Assessment, 6 (9).</bibtext> </blist> <blist> <bibl id="bib4" idref="ref2" type="bt">4</bibl> <bibtext> Brody, L., & Koh, Y. (2020, Nov 21). Student test scores drop in math since covid-19 pandemic. The Wall Street Journal. https://<ulink href="http://www.wsj.com/articles/student-test-scores-drop-in-math-since-covid-19-pandemic-11605974400">www.wsj.com/articles/student-test-scores-drop-in-math-since-covid-19-pandemic-11605974400</ulink>.</bibtext> </blist> <blist> <bibl id="bib5" idref="ref31" type="bt">5</bibl> <bibtext> Brown, T., Chen, J., Ali, U., Costanzo, K., Chun, S., & Ling, G. (2015). Mode comparability study based on spring 2014 field test data. Partnership for Assessment of Readiness for College and Careers.</bibtext> </blist> <blist> <bibl id="bib6" idref="ref27" type="bt">6</bibl> <bibtext> Educational Testing Service, Pearson, & Measured Progress. (2016). Final technical report for 2015 administration.</bibtext> </blist> <blist> <bibl id="bib7" idref="ref8" type="bt">7</bibl> <bibtext> Goldberg, A. L., & Pedulla, J. J. (2002). Performance differences according to test mode and computer familiarity on a practice graduate record exam. Educational and Psychological Measurement, 62 (6), 1053 – 1067. doi: 10.1177/0013164402238092</bibtext> </blist> <blist> <bibl id="bib8" idref="ref5" type="bt">8</bibl> <bibtext> Gordanier, J., Ozturk, O., & Zhan, C. (2023). Pencils down? Computerized testing and student achievement. Education Finance and Policy, 18 (2), 232 – 252.</bibtext> </blist> <blist> <bibl id="bib9" idref="ref15" type="bt">9</bibl> <bibtext> Horkay, N., Bennett, R. E., Allen, N., Kaplan, B., & Yan, F. (2006). Does it matter if I take my writing test on computer? An empirical study of mode effects in NAEP. Journal of Technology, Learning, and Assessment, 5 (2). Retrieved July 29, 2007, from <ulink href="http://escholarship.bc.edu/jtla/vol5/2/">http://escholarship.bc.edu/jtla/vol5/2/</ulink></bibtext> </blist> <blist> <bibtext> Jacob, B., & Rothstein, J. (2016). The measurement of student ability in modern assessment systems. Journal of Economic Perspectives, 30 (3), 85 – 108. doi: 10.1257/jep.30.3.85</bibtext> </blist> <blist> <bibtext> Kuhfeld, M., Tarasawa, B., Johnson, A., Ruzek, E., & Lewis, K. (2020). Learning during COVID-19: Initial findings on students' reading and math achievement and growth. Portland, OR : NWEA Research.</bibtext> </blist> <blist> <bibtext> Massachusetts Department of Elementary and Secondary Education. (2016). Representative samples and PARCC to MCAS concordance studies. Malden, MA : Massachusetts Department of Elementary and Secondary Education.</bibtext> </blist> <blist> <bibtext> Massachusetts Department of Elementary and Secondary Education. (2017). 2017 next-generation MCAS and MCAS-Alt technical report. Malden, MA : Massachusetts Department of Elementary and Secondary Education.</bibtext> </blist> <blist> <bibtext> Massachusetts Department of Elementary and Secondary Education. (2018). 2018 next-generation MCAS and MCAS-Alt technical report. Malden, MA : Massachusetts Department of Elementary and Secondary Education.</bibtext> </blist> <blist> <bibtext> Olson, L. (2019). The new testing landscape. FutureEd. https://<ulink href="http://www.future-ed.org/wp-content/uploads/2019/09/FutureEdTestingLandscapeReport.pdf">www.future-ed.org/wp-content/uploads/2019/09/FutureEdTestingLandscapeReport.pdf</ulink></bibtext> </blist> <blist> <bibtext> Pearson. (2017). Final technical report for 2016 administration.</bibtext> </blist> <blist> <bibtext> Powers, D. E. (1999). Test anxiety and test performance: Comparing paper‐based and computer‐adaptive versions of the GRE general test. ETS Research Report Series, 1999 (2), i – 32. doi: 10.1002/j.2333-8504.1999.tb01813.x</bibtext> </blist> <blist> <bibtext> Sandene, B., Horkay, N., Bennett, R. E., Allen, N., Braswell, J., Kaplan, B., & Oranje, A. (2005). Online assessment in mathematics and writing: Reports from the NAEP technology-based assessment project, research and development series. (NCES 2005–457). National Center for Education Statistics.</bibtext> </blist> <blist> <bibtext> Schaeffer, G. A., Bridgeman, B., Golub‐Smith, M. L., Lewis, C., Potenza, M. T., & Steffen, M. (1998). Comparability of paper‐and‐pencil and computer adaptive test scores on the GRE® general test. ETS Research Report Series, 1998 (2), i – 25. doi: 10.1002/j.2333-8504.1998.tb01787.x</bibtext> </blist> <blist> <bibtext> StataCorp. (2021). Stata statistical software: Release 17. College Station, TX : StataCorp LLC.</bibtext> </blist> <blist> <bibtext> Steedle, J., Pashley, P., & Cho, Y. (2020). Three studies of comparability between paper-based and computer-based testing for the ACT. ACT Research & Policy. Research Report. ACT, Inc.</bibtext> </blist> <blist> <bibtext> Wang, S., Jiao, H., Young, M. J., Brooks, T., & Olson, J. (2007). A meta-analysis of testing mode effects in grade K-12 mathematics tests. Educational and Psychological Measurement, 67 (2), 219 – 238. doi: 10.1177/0013164406288166</bibtext> </blist> <blist> <bibtext> White, S., Kim, Y. Y., Chen, J., & Liu, F. (2015). Performance of fourth-grade students in the 2012 NAEP computer-based writing pilot assessment: Scores, text length, and use of editing tools. Working Paper Series. NCES 2015–119. National Center for Education Statistics.</bibtext> </blist> </ref> <ref id="AN0175570130-15"> <title> Footnotes </title> <blist> <bibtext> Previous work spanning several states, including Massachusetts, suggests that test scores across different standards or assessment regimes provide similar information about students' underlying ability (Backes et al., [2]).</bibtext> </blist> <blist> <bibtext> Boston, Worcester, and Springfield had the option of assigning individual schools to the online or paper format. Otherwise, districts selected a single test administration for the entire district. In November 2015, the Massachusetts State Board of Education voted to discontinue the PARCC assessment and implement a redeveloped version of the MCAS in all schools beginning in 2017.</bibtext> </blist> <blist> <bibtext> Figures 1–3 originally displayed in Backes and Cowan ([1]). These sample tests may be found at https://dc.mypearsonsupport.com/practice-tests/.</bibtext> </blist> <blist> <bibtext> As noted below, because some specifications use twice-lagged test scores, the actual underlying data used are 2010 through 2018.</bibtext> </blist> <blist> <bibtext> This conversion was performed by the Massachusetts Department of Elementary and Secondary Education in order to ensure that PARCC scores could be used for accountability purposes along with MCAS scores. Briefly, this process was conducted in two steps. First, comparable representative samples of MCAS and PARCC students were selected by ensuring that the prior performance and demographic variables for each sample looked similar to the overall mean from the final year that each student took the MCAS; i.e., 2014. Second, an equipercentile method was used to identify comparable test scores across the two different tests using student achievement percentiles for each test. For the full technical description, please see Massachusetts Department of Elementary and Secondary Education ([12]).</bibtext> </blist> <blist> <bibtext> For instance, test results under PARCC (2015 and 2016) are compared to test results using the MCAS 1.0 (2013 and 2014); similarly, test results under MCAS 2.0 (2017 and 2018) are compared to results using the PARCC (2015 and 2016).</bibtext> </blist> <blist> <bibtext> All estimates were generated using StataCorp ([20]).</bibtext> </blist> <blist> <bibtext> Schools could obtain a waiver to administer the MCAS tests on paper, but waivers were granted only for cases in which online testing was not possible at a school (e.g., due to lack of technology required) or if every student in the school would take a paper-based test due to testing accommodations. For more information, see <ulink href="http://www.doe.mass.edu/news/news.aspx?id=25169">http://www.doe.mass.edu/news/news.aspx?id=25169</ulink>.</bibtext> </blist> </ref> <aug> <p>By Ben Backes and James Cowan</p> <p>Reported by Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib15" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib11" firstref="ref3"></nolink> <nolink nlid="nl3" bibid="bib23" firstref="ref6"></nolink> <nolink nlid="nl4" bibid="bib18" firstref="ref7"></nolink> <nolink nlid="nl5" bibid="bib17" firstref="ref11"></nolink> <nolink nlid="nl6" bibid="bib19" firstref="ref12"></nolink> <nolink nlid="nl7" bibid="bib21" firstref="ref16"></nolink> <nolink nlid="nl8" bibid="bib22" firstref="ref17"></nolink> <nolink nlid="nl9" bibid="bib16" firstref="ref28"></nolink> <nolink nlid="nl10" bibid="bib13" firstref="ref30"></nolink> <nolink nlid="nl11" bibid="bib10" firstref="ref32"></nolink> <nolink nlid="nl12" bibid="bib14" firstref="ref51"></nolink> <nolink nlid="nl13" bibid="bib12" firstref="ref64"></nolink>
Header DbId: eric
DbLabel: ERIC
An: ED647304
AccessLevel: 3
PubType: Report
PubTypeId: report
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Are Online and Paper Tests Comparable? Evidence from Statewide K-12 Tests
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Ben+Backes%22">Ben Backes</searchLink><br /><searchLink fieldCode="AR" term="%22James+Cowan%22">James Cowan</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Grantee+Submission%22"><i>Grantee Submission</i></searchLink>. 2024.
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 30
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2024
– Name: SourceSuprt
  Label: Sponsoring Agency
  Group: SrcSuprt
  Data: Institute of Education Sciences (ED)<br />National Center for Analysis of Longitudinal Data in Education Research (CALDER) at American Institutes for Research (AIR)
– Name: NumberContract
  Label: Contract Number
  Group: NumCntrct
  Data: R305A170119
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Reports - Research
– Name: Audience
  Label: Education Level
  Group: Audnce
  Data: <searchLink fieldCode="EL" term="%22Elementary+Secondary+Education%22">Elementary Secondary Education</searchLink>
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Computer+Assisted+Testing%22">Computer Assisted Testing</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Format%22">Test Format</searchLink><br /><searchLink fieldCode="DE" term="%22Differences%22">Differences</searchLink><br /><searchLink fieldCode="DE" term="%22Academic+Achievement%22">Academic Achievement</searchLink><br /><searchLink fieldCode="DE" term="%22State+Programs%22">State Programs</searchLink><br /><searchLink fieldCode="DE" term="%22Standardized+Tests%22">Standardized Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Elementary+Secondary+Education%22">Elementary Secondary Education</searchLink><br /><searchLink fieldCode="DE" term="%22Grade+Prediction%22">Grade Prediction</searchLink>
– Name: Subject
  Label: Geographic Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Massachusetts%22">Massachusetts</searchLink>
– Name: SubjectThesaurus
  Label: Assessment and Survey Identifiers
  Group: Su
  Data: <searchLink fieldCode="SU" term="%22Massachusetts+Comprehensive+Assessment+System%22">Massachusetts Comprehensive Assessment System</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1080/08957347.2024.2311933
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: We investigate two research questions using a recent statewide transition from paper to computer-based testing: first, the extent to which test mode effects found in prior studies can be eliminated in large-scale administration; and second, the degree to which online and paper assessments offer different information about underlying student ability. In contrast to the first test transition in Massachusetts, we find very small mode effects for a more recent transition, which may be attributable to an additional step matching on observable characteristics in the equating process. Second, we investigate the predictive evidence of validity for paper and online tests for predictions of future test scores and grades. We generally find minimal differences for the extent to which scores on paper tests can differentially predict future online versus paper test scores. Finally, online and paper test scores are similarly predictive of future grade point average. We conclude that the online test penalty can vary substantially by test and that extreme care should be taken when administering online tests to some students and paper tests to others.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: CodeSource
  Label: IES Funded
  Group: SrcInfo
  Data: Yes
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2024
– Name: AN
  Label: Accession Number
  Group: ID
  Data: ED647304
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=ED647304
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1080/08957347.2024.2311933
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 30
    Subjects:
      – SubjectFull: Computer Assisted Testing
        Type: general
      – SubjectFull: Test Format
        Type: general
      – SubjectFull: Differences
        Type: general
      – SubjectFull: Academic Achievement
        Type: general
      – SubjectFull: State Programs
        Type: general
      – SubjectFull: Standardized Tests
        Type: general
      – SubjectFull: Elementary Secondary Education
        Type: general
      – SubjectFull: Grade Prediction
        Type: general
      – SubjectFull: Massachusetts
        Type: general
      – SubjectFull: Massachusetts Comprehensive Assessment System
        Type: general
    Titles:
      – TitleFull: Are Online and Paper Tests Comparable? Evidence from Statewide K-12 Tests
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Ben Backes
      – PersonEntity:
          Name:
            NameFull: James Cowan
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 02
              Type: published
              Y: 2024
          Titles:
            – TitleFull: Grantee Submission
              Type: main
ResultId 1