Synergy and Tension between Large-Scale and Classroom Assessment: International Trends

Saved in:
Bibliographic Details
Title: Synergy and Tension between Large-Scale and Classroom Assessment: International Trends
Language: English
Authors: Volante, Louis, DeLuca, Christopher, Adie, Lenore, Baker, Eva, Harju-Luukkainen, Heidi, Heritage, Margaret, Schneider, Christoph, Stobart, Gordon, Tan, Kelvin, Wyatt-Smith, Claire
Source: Educational Measurement: Issues and Practice. Win 2020 39(4):21-29.
Availability: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 9
Publication Date: 2020
Document Type: Journal Articles
Reports - Evaluative
Descriptors: Educational Trends, Trend Analysis, Measurement, Teaching Methods, Educational Assessment, Learning Processes, Cross Cultural Studies, Foreign Countries, Educational Policy, Test Results, Correlation
Geographic Terms: United States, Canada, Finland, Singapore, Australia, Germany, United Kingdom (England)
DOI: 10.1111/emip.12382
ISSN: 0731-1745
Abstract: The synergy, or lack thereof, between large-scale and classroom assessment has been fiercely debated in both academic and policy spheres for decades around the world. This paper seeks to explicate how different countries are utilizing large-scale testing and test results at the classroom level. Through country profiles, this paper analyzes contemporary developments on the tensions and synergies between large-scale assessment and classroom teaching, learning, and assessment observed across seven international jurisdictions: United States, Canada, Australia, England, Germany, Finland, and Singapore. The paper concludes with an analysis of international trends leading to a synthesis of root causes contributing to the current limited uptake of large-scale assessment results at classroom levels.
Abstractor: As Provided
Entry Date: 2020
Accession Number: EJ1276712
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwH8WxllSMSsDaiGgj8GUav7AAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDC9h0AopLc9X2F0lzAIBEICBm9p8PAT8qvwjX7MdSH-hnXvcjaBOlKfdBK5Kji4KKXK_IsCxcDD24BBp1Vgj9-2GQmBLoruI5UVSWxNo0kQvgUEfBo9-NVAqJJLXeaupiSQ78nkktTeAym-epzPu014H9vGVHUnvIbThtmYbbtezsFHLj36JWXUnBBn3nLGDV6Q74oJiNfAFRAHzHCC4_wxfnu_z-M19azGBnjoZ
Text:
  Availability: 1
  Value: <anid>AN0147461954;ems01dec.20;2020Dec09.04:50;v2.2.500</anid> <title id="AN0147461954-1">Synergy and Tension between Large‐Scale and Classroom Assessment: International Trends </title> <p>The synergy, or lack thereof, between large‐scale and classroom assessment has been fiercely debated in both academic and policy spheres for decades around the world. This paper seeks to explicate how different countries are utilizing large‐scale testing and test results at the classroom level. Through country profiles, this paper analyzes contemporary developments on the tensions and synergies between large‐scale assessment and classroom teaching, learning, and assessment observed across seven international jurisdictions: United States, Canada, Australia, England, Germany, Finland, and Singapore. The paper concludes with an analysis of international trends leading to a synthesis of root causes contributing to the current limited uptake of large‐scale assessment results at classroom levels.</p> <p>Keywords: classroom assessment; education policy; large‐scale testing; student assessment</p> <p>The synergy, or lack thereof, between large‐scale and classroom assessment has been fiercely debated in both academic and policy spheres for decades. Researchers around the world have regularly documented the pronounced influence of external testing measures on teachers' instructional practices—with both intended positive and unintended negative consequences and effects on teachers and students alike (Volante, 2012). Indeed, one of the earliest references to the "encroaching power" of high‐stakes testing comes from an American publication where the author documents how tests influence the "prevalent views of life and work among young men and how it affects parents, teachers, the writers of educational books, and the notions of the public about education" (Latham, 1886, p. 2). More than a century later, the international community is still grappling with the challenges posed by test‐based accountability systems and how results from large‐scale tests are used in classrooms around the world (Stobart, 2008).</p> <p>This paper offers an analysis of contemporary developments and tensions observed across seven international jurisdictions: United States, Canada, Australia, England, Germany, Finland, and Singapore. Although these countries certainly do not encapsulate the full range of global developments in this field, they have been purposefully selected to represent a cross‐section of nations that vary by (a) accountability structures, (b) large‐scale and classroom assessment systems, (c) influence/"stakes" associated with large‐scale assessment, and (d) results on international achievement tests. As a result of these criteria, the selected jurisdictions reflect industrialized high‐income countries.</p> <p>Our paper seeks to explicate how different countries are utilizing large‐scale testing and test results at the classroom level. Our interest is primarily on the linkage between country‐specific large‐scale tests and classroom assessments rather than on exploring the influence of various international assessments (e.g., the Programme for International Student Assessment [PISA]; Trends in International Mathematics and Science Study [TIMSS]; Progress in International Reading Literacy Study [PIRLS]) on education systems (for reviews of international assessments, see Smith, 2014; Volante, 2016). In this article, we present country profiles written by leading academics from each country that provide an international review of developments on the tensions and synergies between large‐scale assessment and classroom teaching, learning, and assessment.</p> <p>By necessity, the country profiles are brief and cannot provide a comprehensive historical overview of changes in policies and uses of assessment in each context; rather, their intent is to provide a high‐level articulation of current (i.e., 2019) state‐of‐the‐field in relation to tensions and synergies between country‐based large‐scale assessments and classroom contexts that have been noted in the available empirical research. Much of the evidence to support key claims stems from many of the authors' engagement in national‐level assessment research and policy work; nevertheless, the authors also draw upon the most salient national policy developments and other relevant empirical findings to provide complementary evidence in compositing the national profiles. The paper concludes with an analysis of international trends leading to a synthesis of root causes contributing to the current limited uptake of large‐scale assessment results at classroom levels. These findings will be of particular interest to researchers and policymakers tasked with developing and implementing assessment systems that effectively integrate large‐scale assessments with classroom practices.</p> <hd id="AN0147461954-2">Australia</hd> <p>Currently, Australia has a complex assemblage of nation‐wide, state, district, and school‐level assessments intended to align with a national curriculum and accompanying year‐level achievement standards for Foundation to Grade 10 (4–15 years). The National Assessment Program–Literacy and Numeracy (NAPLAN) is a census test administered in all Australian schools at years 3, 5, 7, and 9. It is the country's only large‐scale, longitudinal assessment, now in its 13th year of implementation. NAPLAN involves the application of a 1–10 national achievement scale to classify student performance into bands that represent the national minimum standard for each year level. Historically, NAPLAN has been a pencil and paper test, moving to staggered implementation of online testing from 2018–2020, which encountered some implementation difficulties. The decision to move to full online implementation of the test remains contested at education policy level.</p> <p>In addition to participating in international assessments (i.e., PISA, TIMSS, PIRLS), Australia has 3‐yearly sample assessments in civics and citizenship, information and communication technology (ICT) literacy, and science literacy tested on a rolling basis.</p> <p>Reports on NAPLAN have been used primarily by governments at the national and state levels for accountability purposes, while international assessments are used for benchmarking purposes. Referring to PISA, reports of declining comparative performance have received considerable media coverage and policy attention with the Australian Government installing various policy levers to improve international rankings in reading, mathematical, and science literacy. This move has impacted state education policy priorities and classroom assessment practices (Thomas, Emery, Prain, Papageorgiou, & McKendrick, 2019) and, in part, explains why reading has been given significantly more intense policy attention than writing. This has been the case even though results in writing have declined (Wyatt‐Smith & Jackson, 2016).</p> <p>NAPLAN results also routinely receive intense media coverage and policy scrutiny with comparisons made at state and school levels (Cumming, Van Der Kleij, & Adie, 2019). There has been reported narrowing of curriculum with target goals for improvement, including for identified target populations; intensified test preparation; and a focus on data for comparative purposes (Lingard, Thompson, & Sellar, 2017). NAPLAN has widely been reported as impacting on school and teacher decision‐making regarding curriculum priorities and implementation practices with some state governments moving to develop teacher and classroom materials at scale. A notable example is <emph>Curriculum to Classroom</emph> (https://education.qld.gov.au/curriculum/school-curriculum/C2C) developed in Queensland to model pedagogic practice and assessment tasks including accompanying assessment criteria. A review of NAPLAN is underway (Wyatt‐Smith & Jackson, 2016), triggered by increasing recognition of the test's unintended consequences on teachers, schools, student learning, and well‐being. This is occurring against a backdrop of increasing interest in ways of personalizing assessment and in strengthening teachers' practice in formative assessment (Gonski et al., 2018).</p> <p>Currently, three directions are clear in Australia's assessment profile: first, a growing interest in school performance and evidence of student growth and achievement (Gonski et al., 2018); second, the customization of learning and assessment enabled through digital technologies; and third, learning analytics for longitudinal tracking of performance patterns at cohort, subgroup, and individual levels.</p> <hd id="AN0147461954-3">Canada</hd> <p>Each of Canada's 10 provinces and 3 territories governs its own education system; however, there are striking similarities in curricular priorities and pedagogical trends across the country as reinforced by overarching agencies such as the Council of Ministers of Education, Canada. With respect to international testing programs, each province participates as its own jurisdiction, on TIMSS, PISA, and PIRLS international assessments. Each province and territory also maintains its own provincial large‐scale testing program. Tests within provincial programs vary in their purpose with some serving gatekeeping functions (e.g., British Columbia's Graduation Provincial Examinations), monitoring functions (e.g., Newfoundland and Labrador's Criterion Reference Tests), instructional diagnosis functions (e.g., Manitoba's Middle Years Assessment), and accountability functions (e.g., Ontario's Grades 3, 6, and 9 and Secondary Literacy Test) (Klinger, DeLuca, & Miller, 2008). Testing occurs in various grades across provinces ranging from Grade 2 to Grade 12, and across subjects, most notably in mathematics, literacy, and sciences. Given the highly diverse student population within most provincial education systems in Canada, provisions for accommodation across ability and language groups or exclusion are common.</p> <p>Through a current systematic review of classroom‐level grading policies and practices, DeLuca, Braund, Valiquette, and Cheng (2017) documented that provincial tests hold varying weights ranging from 10% to 50% on student's final course grades at upper secondary levels, with a recent trend in some provinces toward increasing the value of provincial test results on students' final grades. Various classroom assessment policy documents across provinces provide general guidance on the relationship between the provincial testing program and classroom‐level assessment but offer limited guidance to teachers on how provincial tests should be used productively in classrooms to align or inform their teaching and assessments. There are few deliberate policy structures linking provincial testing processes to classroom‐level assessments in Canada. Instead, provincial priorities in professional learning and resource allotment tend to follow achievement trends on provincial results. For example, declining mathematics scores on Ontario's Education Quality and Accountability Office (EQAO) tests has sparked a priority toward stimulating mathematics teaching and learning initiatives. Further, unlike other countries that use large‐scale assessment results to narrow funding to underperforming schools (leading to school closures or restaffing), provincial governments in Canada take the opposite approach: to support public schools with additional professional development and resources to ensure that they begin to perform better over time and to ensure greater consistency across schools throughout a province.</p> <p>Despite the real or perceive "stake" of the assessment for students, reports across provinces often document the washback effects of provincial testing on teacher pedagogy including instances of curriculum narrowing, test drills, and teaching test‐taking skills, particularly around testing periods and tested subject domains (Volante & Ben Jaafar, 2008). However, under a framework of public accountability, provincial tests are more often used at school levels—by teachers and administrators—to help establish school targets, professional learning goals, and school improvement plans. In this way, while large‐scale assessment in Canada may not systematically drive classroom assessment practices, they do hold influence on the direction for school improvement initiatives and professional development, which ultimately shapes classroom teaching, learning, and assessment practices.</p> <hd id="AN0147461954-4">England</hd> <p>The English education system is dominated by its high‐stakes accountability processes. Central to these are the external assessments by which schools are judged. These come at the end of primary school (grade 6, age 11), when pupils sit national "key stage 2" tests in mathematics, reading and English grammar, punctuation and spelling, and at age 16 when students sit subject‐specific General Certificate in Secondary Education (GCSE). There are also national assessments at age 7 and results at advanced‐level, single‐subject selective examinations in year 13, which are also collected and used to evaluate schools.</p> <p>Each school has its pupils' results aggregated and is expected to meet government‐determined threshold levels of performance and to make year‐on‐year progress. Failure to do this can trigger an inspection from the Office of Standards in Education (Ofsted) and schools being put in "special measures." There is concern about schools "gaming" the system to optimize results, and Ofsted is currently addressing "off‐rolling" by which schools suspend, expel, or transfer students so that their results will not be included (Staufenberg, 2017).</p> <p>This assessment landscape is against the backdrop of constantly shifting political education and assessment policies. The last 5 years have seen changes to the national curriculum to increase "challenge," the narrowing of key stage 2 tests, revised GCSE specifications, and a changed GCSE grading system (Cife, 2018). Within‐course modular examinations have been replaced by linear "final" examinations and teacher assessment no longer contributes to grades in most mainstream subjects. A current proposal is to test children on entry to school aged 4 and to store the results until they leave at age 11 so that the school's "value added" can be measured.</p> <p>England takes part in sample‐based international large‐scale assessments such as PIRLS, TIMMS, and PISA. However, the findings are rarely accessed by, or utilized in, schools, though they are often high stakes for politicians (is the country improving or slipping comparatively?) (Hopfenbeck & Gorgen, 2017). The more informal day‐to‐day assessment processes linked to Assessment <emph>for</emph> Learning (A<emph>f</emph>L) are familiar to most teachers. A<emph>f</emph>L implementation will depend on the context, and teachers teaching examination classes are more likely to focus on past papers and "mock exams." The synergy between external assessments and classroom preparation is clear, if restricting. In nonexamination classes, there may be greater freedom, though teachers are required to "track" pupil progress, often using computer packages, as part of school accountability (Dann, 2018).</p> <hd id="AN0147461954-5">Finland</hd> <p>According to Vainikainen and Harju‐Luukhainen (in press), very little attention has been paid to the model of the Finnish educational assessment system and the lack of standardized measurement and control. Thus, these factors in large part are contributing to the overall functioning of the system. <emph>School‐Based Assessments (SBAs)</emph> are the most common form of assessments in Finland and conducted by the teachers at the classroom level. The teacher evaluates formatively how student performance is improving over time on all grade levels. This focus on formative assessment allows the teacher to adjust teaching methods to the individual student's specific needs and to improve the learning outcome. The national core curriculum for basic education determines the learning objectives for each school subject. Also grading guidelines are given. This approach serves as a baseline while grading students. The objectiveness is important here since, for instance, the received grades at the end of the compulsory education determine students' next steps in the education system (see also Harju‐Luukkainen, Vettenranta, Oukrim‐Soivio, & Bernelius, 2016).</p> <p> <emph>National Assessments (NAs)</emph> provide valuable information for educational authorities. All of the national assessments, conducted by the Finnish Education Evaluation Centre (FINEEC), are sample‐based (around 5–10% of the age group). These assess how well students have reached the learning objectives of the national curriculum in different subject areas. Basic education is expected to secure equal educational opportunities for all students, and therefore, the equity of learning outcomes is especially evaluated. Schools selected for national assessments receive feedback in the form of reference data and use this feedback to develop their instruction in different subjects. However, there are no nation‐wide examinations during or at the end of basic education. Further, a school's participation in NA is infrequent (see also Vainikainen et al., 2017).</p> <p>Interestingly, <emph>International Large‐Scale Assessments (ILSAs)</emph> are gaining more momentum in Finland. Data from these assessments are used to develop the overall education system. A declining trend in international assessments since 2006 has been observed (Harju‐Luukkainen, Nissinen, Sulkunen, & Suni, 2014; Hautamäki, Kupiainen, Marjanen, Vainikainen, & Hotulainen, 2013). This trend has been taken seriously and programs have been launched to address this decline, including thematic assessments of the education system and additional funding for municipalities and schools to improve their practices.</p> <p>As a summary, assessments on different levels are used to ensure equality of education on all levels of the education system. Lack of standardized examinations gives more freedom to schools and teachers to implement curriculum in a purposeful way. According to Vainikainen and Harju‐Luukkainen (in press), the declining trend in international assessments and increasing segregation call for a slightly more detailed monitoring system across the country.</p> <hd id="AN0147461954-6">Germany</hd> <p>In Germany, a strong focus on the monitoring of school outcomes emerged in the early 2000s, when unexpected mediocre PISA outcomes raised very wide public attention. In reaction to this infamous "PISA shock," school policy undertook vast reform efforts. Most significantly, many of Germany's 16 Länder (federal states) implemented structural changes for tracking practices in lower secondary education, namely, replacing the traditional highly selective three‐track system (i.e., academic vs. intermediate vs. basic track) with a more permeable two‐track system (academic vs. nonacademic track). Furthermore, the Standing Conference of the Cultural Ministers of the Länder (KMK) expedited the use of outcomes in LSAs to inform education policy. Its strategy for educational monitoring through examinations and assessments included (<reflink idref="bib1" id="ref1">1</reflink>) participation in international LSAs, (<reflink idref="bib2" id="ref2">2</reflink>) regular national evaluation of outcomes in primary, lower, and higher secondary education in relation to educational standards and competence models, and (<reflink idref="bib3" id="ref3">3</reflink>) assessments and interventions designed to enhance educational quality at the school level (KMK, 2015). While the first two have little impact on the school or classroom level, low‐stakes <emph>school inspections</emph> (SIs) and <emph>comparative tests</emph> (CTs) pertaining to the third are decidedly aimed at providing feedback to schools.</p> <p>SI practice, focusing on quality of instruction and processes inside schools, varies somewhat in methodology, variables assessed, and modes of feedback across the Länder. Yet, Germany's overall approach to SI builds strongly on insights from school‐level stakeholders rather than on promoting competition among schools or on imposing consequences for schools with insufficient outcomes (e.g., Rürup, 2008). Feedback is exclusively directed at individual schools, and the outcomes lead to agreements on objectives for school improvement. Studies on the effects of SI revealed that it is well accepted and perceived as helpful by school principals (Böhm‐Kasper, Selders, & Lambrecht, 2016), but teachers found it difficult to use the feedback they receive for improving their instructional practice (Böttcher, Hense, & Keune, 2013). There is limited evidence of impact on students' competence tests outcomes (e.g., Pietsch, Janke, & Mohr, 2014). Recently, some Länder have opted to suspend their external SI and transfer responsibility directly to schools.</p> <p>CTs ("Vergleichsarbeiten [VERA]") are conducted nationwide in grades 3 and 8 to generate feedback on students' competences in a number of domains. Reports including individual student outcomes, class‐ and school‐aggregated data, and reference group data are provided to schools and their teachers, who are expected to deduce relevant information for enhancing the quality of their work. Empirically, using CT reports for purposes such as developing instructional materials and assignments, competency orientation, and individualization was fairly common among teachers, but teachers' self‐reported intensity in using CT reports had no impact on students' learning outcomes (Wurster, Richter, & Lenski, 2017). Only half of school principals reported using CT feedback to inform school improvement practices (Bach, Wurster, Thillmann, Pant Anand, & Thiel, 2014). In summary, a lack of consistent effects of SI and CT on outcome variables has raised critiques that while there is an extensive national strategy on monitoring and significant data are continuously collected, an adjacent change management strategy on how to make use of this data is widely missing (e.g., Böttcher, 2014).</p> <hd id="AN0147461954-7">Singapore</hd> <p>Singapore has an examination‐oriented education system, and has enjoyed international success in international assessments, including TIMSS and PISA. The confluence of success in international assessments and the use of meritocracy as a grand narrative for social order and ranking (Tan & Deneen, 2015) entrenches "assessment" as an institutional authority of standards, teaching (performativity), and classroom learning (Leong & Tan, 2014).</p> <p>Admission to each level of education is guarded almost solely by examinations, rendering them high stakes. Students sit for national examinations at the end of primary six, secondary four, and the second year of preuniversity. These highly competitive examinations serve as school placement assessments for secondary school, preuniversity, and university schools/course, respectively. Large‐scale assessments influence school‐based assessment at all levels by focusing students and teachers on preparation for national examinations (Tan, 2011). As early as primary one, students are taught how to "scaffold" their language compositions with memorized phrases and suggested formats in adherence to scoring rubrics. Such is the reliance on examinations that the ministry's recent announcement of removing school mid‐term examinations in odd numbered years led to parents requesting tuition agencies to enact mid‐year examinations for their children (https://<ulink href="http://www.channelnewsasia.com/news/singapore/exams-tuition-centres-rush-in-to-fill-gap-soothe-anxious-parents-10827734">www.channelnewsasia.com/news/singapore/exams-tuition-centres-rush-in-to-fill-gap-soothe-anxious-parents-10827734</ulink>).</p> <p>Strenuous efforts have been made to balance high‐stakes summative assessment with more formative uses of assessment in schools. In 2008, the Ministry of Education (MOE) introduced "Holistic Assessment" for primary schools to place greater emphasis on skills development and to ease assessment stress on students. Primary schools were encouraged to provide more qualitative feedback for improvement and to design "bite‐sized" assessment as an alternative to one‐off examination(s) in order to reduce stress. However, teachers report difficulty in couching feedback in language that could be understood and acted upon by students. Likewise, converting an assessment task into bite‐sized episodes seemed to create more frequent occurrences of assessment, and consequently more stress for students and teachers (Ratnam‐Lim & Tan, 2015).</p> <p>Attempts to introduce A<emph>f</emph>L in secondary schools and preuniversities remain mixed. Teachers value A<emph>f</emph>L but perceive a lack of assessment literacy and opportunities to practice it (Deneen et al., 2019). In contrast, summative assessment is valued less than formative assessment, but teachers claimed to be more proficient in it and use it more than formative assessment. It would seem that Singaporean teachers are still struggling to balance formative uses of assessment in schools and preparation for high‐stakes summative assessment.</p> <p>As previously noted, the MOE recently announced the removal of mid‐year examinations in odd numbered years in primary and secondary schools (MOE, 2018). Educators in Singapore remain hopeful that ensuing opportunities for more curriculum time and less frequent examinations would enable more formative uses of assessment in the classroom. Collectively, these reforms are designed to "nurture lifelong learners" and "help students discover more joy and develop stronger intrinsic motivation in learning" (see https://<ulink href="http://www.moe.gov.sg/news/press-releases/-learn-for-life—preparing-our-students-to-excel-beyond-exam-results">www.moe.gov.sg/news/press-releases/-learn-for-life—preparing-our-students-to-excel-beyond-exam-results</ulink>).</p> <hd id="AN0147461954-8">United States</hd> <p>The United States has a long history of testing. Since 1969, the National Assessment of Educational Progress (NAEP) has been used to measure student learning across the nation, states, and in some urban districts resulting in what is known as The Nation's Report Card. Congressionally mandated and administered by the National Center for Education Statistics, NAEP involves sample assessments across the country with results reported based on groups of students (similar in demographic characteristics) and to the public, states, and districts. These assessments, while not directly used to drive curriculum and instruction, do provide public accountability information on the "state of education" in America.</p> <p>Accountability measures have been heightened in the United States over the past three decades and came to dominate when the No Child Left Behind Act (NCLB), the nation's major law governing public schools, was enacted in January 2002. NCLB mandated annual testing against state standards for students in grades 3–8 in reading and mathematics, and intensified the consequences for schools that did not make "adequate yearly progress." The axiom "what gets tested gets taught" was realized in schools, resulting in a narrowing of curriculum and instruction (Dee & Jacobs, 2010).</p> <p>The testing requirements of NCLB were retained in NCLB's replacement, Every Student Succeeds Act (https://<ulink href="http://www.ed.gov/essa?src=rn">www.ed.gov/essa?src=rn</ulink>). Large‐scale assessments enshrined in NCLB and ESSA are primarily informed by research focused on psychometric models and represent content domains as the end points of learning (Shepard, 2019). Such assessments are not designed for instructional purposes, and so have little utility for informing ongoing teaching and learning in the classroom (Perie, Marion, & Gong, 2009).</p> <p>In an effort to redress the balance between classroom and large‐scale assessment advocated by the seminal report <emph>Knowing What Students Know</emph> (<emph>KWSK</emph>) (NRC, 2001), many states proposed "balanced" assessment systems, which generally included end‐of the‐year state assessments and interim assessments administered several times each year to monitor progress toward standards and formative assessment. Mostly, though, such systems did not correspond to the vision of a coordinated system of multiple assessments working together to collectively support a common set of learning goals, which <emph>KWSK</emph> advocated.</p> <p>More ambitious learning standards were adopted by states beginning in 2010, at which time the federal government funded two consortia to develop new assessments aligned to them. No doubt, the consortia assessments were an improvement on previous end‐of‐year assessments. Nonetheless, because their purpose is to measure learning against grade‐level standards at the end of the year, they have had limited utility for informing teaching and learning day by day.</p> <p>ESSA made provision for a limited number of states to develop innovative assessment systems that support student learning, and some states are currently attempting to do this. It remains to be seen if these innovations can lead to assessments that better balance value to classroom practice and accountability. Various curricular and assessment alignment methodologies as well as computerized assessment practices and learning analytics appear to be pushing forward innovations in this area (Wilson & Scalise, 2016). Formative assessment, which has the potential to improve classroom learning, although not without its challenges, is still facing strong headwinds after years of accountability pressures with a number of states making considerable efforts to support this practice in schools (Birenbaum et al., 2015; Brookhart, 2019).</p> <hd id="AN0147461954-9">International Trends</hd> <p>In this section, we engage in an analysis of international trends related to the influence of large‐scale tests on classroom assessments based on the seven country profiles. As evident across contexts, with the exception of Finland, large‐scale assessments operate as highly visible public policy instruments that function to report on and, in some instances, improve teaching and learning in educational systems (Mazzeo, 2001). Most contexts have a variety of testing programs operating at various grades (ranging from grade 2 to grade 12) and in relation to diverse content areas (traditional subject domains plus contemporary domains such as information communication technology, citizenship, and scientific literacy). The growing preoccupation with large‐scale assessments across geographies has endorsed a thickening testing culture in schools globally (https://eric.ed.gov/?id=ED563487) focusing attention on the measurable outcomes of teacher practice and student learning (Verger & Parcerisa, 2017); thus, in these contexts, teachers and students are both users and subjects of national and state tests (Holloway, Sørensen, & Verger, 2017). Such a positioning calls our attention to consider the ways these large‐scale tests are taken up in classrooms and potentially constrain best practices.</p> <p>Across several countries (Australia, England, United States, and Singapore), more often than not, the presence of large‐scale assessments leads to tensions in practice: observed narrowing of classroom teaching and assessment (e.g., mock exams, prioritizing content) to more worrying activities including gaming the system, reclassifying students for exemptions, and school closures, staff removal, and funding redistribution. Additionally, a growing source of incongruence in many systems is the struggle to mount a robust and well‐endorsed A<emph>f</emph>L culture amid the dominating influence of national and state testing programs. Singapore offers a concrete case where efforts toward holistic assessment and A<emph>f</emph>L, while endorsed in policy, may be perceived to run counter to the known effects of the exam‐oriented system. This incongruence presents a further potential conflict for teacher and students, the concurrent valuing of two diverse forms of evidence for student learning—the scaled and numerical form produced through external tests and the qualitative, ongoing feedback of evidence produced through formative and more holistic assessments. These diverse evidences potentially signal different values for learning with challenges in their alignment and interpretation. While formative assessment can provoke gains on tests, the fundamental challenge is not to chase the grade but value the feedback.</p> <p>Accountability featured as an anchoring commitment across the majority of contexts as justification for large‐scale assessments. However, as observed, each country has taken up the mantel of accountability differently: tests vary in terms of their "stakes" (i.e., ranging from nonexistent consequences to high levels of gatekeeping based on results), sampling of students differs (i.e., ranging from sporadic stratified sampling to census testing of every student), and reporting and feedback practices diverge (i.e., ranging from public reporting of school results to internal feedback). This finding is unsurprising as it reflects what Hopmann (2008) called the "multiple realities of accountability" (p. 426) noting that "accountability concepts change over time and are different in different places" and that "accountability theories provide a wide array of possible causes and implications" (p. 422). As evident across profiles, accountability either lacks attention and is satisfied by international assessments (e.g., PISA, TIMSS) as in the case of Finland or serves as a linchpin of public education through a variety of high‐stakes examinations as seen in the England, United States, Singapore, and NAPLAN in Australia. Germany and Canada offer alternative test‐based accountability measures and conceptions, diminishing the stakes and gatekeeping functions of tests for students. In both contexts, provincial and state tests are used for monitoring and feedback processes, with Germany coupling these tests with school inspections. Importantly, even in contexts where tests are intended to provide feedback, productive uptake by teachers appears limited, for example, in Canada; rather test results tend to be valued by broader school and school district‐level administrators as information to direct system priorities, resources, and professional learning plans.</p> <p>Considering both within and between countries, what emerges is a complicated assemblage of assessment practices governed by nations, states, districts, schools, and teachers. Such an assemblage leads, in a positive view, to a significant systemic capacity to monitor progress, plan from program improvements, and identity system inequities. On the other hand, and as noted in some country profiles, there has been a need to "balance" and "coordinate" the variety of tests and assessments, and their resulting data, in an effort to effectively inform systems, schools, and students. While data are undoubtedly a fixture of the accountability movement, there appears to be few instances where large‐scale assessment data are effectively driving classroom instruction. Instead, what has emerged are more elaborate testing architectures, such as the balanced assessment systems developed in some states in the United States, where interim state and district tests are intended to align with end‐of‐year assessments and national assessments resulting in a significant number of external tests during instructional time. Part of the potential backlash of these more elaborate testing assemblages is a valuing of external assessments in lieu of teacher assessment, as emerging in the England.</p> <hd id="AN0147461954-10">Discussion</hd> <p>In the restricted space available, this paper provides a top‐level description of seven countries' testing practices. However, it is clear in most cases (excepting Singapore on one end and Finland on the other end) that desires to integrate policy‐relevant large‐scale assessment with teachers' daily assessments have been largely unachieved and may be futile. Based on the country profiles, the authors agree that there are three potential fundamental explanations why assessments are largely unintegrated.</p> <p>The first explanation is that separate levels of authority promote different uses of tests, with varied procurement and far different expectations for data use. A good case of this explanation is evident in the United States where there exists national‐, state‐, district‐, and school‐level assessments of various forms and stakes, leading to a system of assessment with varied priorities and approaches. In classrooms, assessment data are expected to guide the teachers' practices in identifying and giving feedback on student weaknesses or strengths (Shepard, 2019). Data use depends upon teachers' internalizing classroom assessments (if they are externally provided), interpreting findings, finding and applying differentiated pedagogical options, and conducting these procedures seamlessly in a dynamic classroom. Given pacing that restricts flexible instructional time, teachers' own predilections and skills, and limited access to high‐quality support materials, the ideals of classroom assessment and their synergy with larger scale assessments will likely remain distant goals. On the other hand, large‐scale assessment can occur with sparse sampling of content and rare validity studies relative to the policy uses intended (Baker, 2016). As a result, these exams may neither give full descriptions nor provide comprehensive guidance for specific policy options, but their regular use may convince the authorities that they are upholding managerial principles of guidance, evidence, and accountability. Their true cost is whether teachers can legitimately affect these scores, which has given rise to value‐added models and extensive tracking practices as evident in the United States, England, and Singapore, and to a lesser extent in Germany, Canada, and Australia.</p> <p>A second reason for the lack of integration derives from how tests are conceived, designed, and developed. As noted in country profiles, and in the previous research, most national‐ or state‐/provincial‐level large‐scale assessments are created and used largely for one purpose these days—accountability (Baker, 2016; Stobart, 2008). However, while accountability is a core feature of national or state/provincial assessments across nearly all systems discussed in this paper (Finland is the exception), the assessment practices, responses, and relative "stake" for students, teachers, and schools ascribed to accountability mandates differ substantially. In some cases (e.g., United States, England), trends toward declining performance have led to increased external evaluation or redistribution of staff and students, whereas in other contexts (e.g., Canada), declining performance promoted increases in financial and human resource supports for schools.</p> <p>Consistent across countries, the guidelines for most school tests tend to be standards or general statements of content and skills expected of students. The central challenge is that assessment design and alignment at each level of the system—teacher, school, district, state, national, and international—‐vary substantially based on skill, referencing standards, alignment approach, and the use of diverse (or nonexistent) learning frameworks. Most assessments discussed in the country profiles were explicitly curriculum‐referenced, yet these curricula change over time and level, and may not align with international assessments. Moreover, the emergence of assessing (a) broader competencies such as civics and citizenship and information communication technology, as seen in Australia, (b) learning progressions that extend over multiple years, and (c) wider learning capacities (e.g., creativity, collaboration, critical thinking) as developing in international assessments (e.g., PISA 2021), extend beyond traditional narrowly defined content domains further provoking this challenge.</p> <p>To create integrated systems of assessment, the design of assessments should share more than general standards and broad constructs. Instead, the framework guiding these assessments (as well as curriculum) should focus on clear descriptions of cognitive requirements, e.g., what kinds of problem‐solving at grade 3 and grade 8, bounded content (what is fair for instruction and testing), and what criteria define adequacy operationally. One could easily create such conceptual and operational frameworks to apply across purposes, with approaches that vary sampling, domain, time, scoring, and modeling appropriate for particular uses. These approaches are largely absent because of inertia, profit motive, and lack of insight, and as a result, the two worlds of classroom and large‐scale testing spin in largely different orbits. As observed by Pellegrino and Chudowsky (2003) in their summary article of the National Research Council report, <emph>Knowing What Students Know: The Science and Design of Educational Assessment</emph> (2001), large‐scale "tests do not focus on many aspects of cognition that research indicates are important, such as students' organization of knowledge, problem representations, use of strategies, self‐monitoring skills, and individual contributions to group problem solving" (p. 107). Accordingly, they advocate for a more explicit and robust linking of assessment design to contemporary learning frameworks. Such a call is resonant with other scholars who recognize that learning sciences have evolved significantly to better understanding the underpinning cognitive mechanisms supporting student learning across levels and contexts; developing assessments on these newer models, and in conjunction with one another, would more accurately measure student learning (Baird, Andrich, Hopfenbeck, & Stobart, 2017; Pellegrino, DiBello, & Goldman, 2016).</p> <p>Finally, large‐scale and classroom assessment remain separated because there are no active advocates for their integration. We call out educational researchers and measurement people who have acquiesced to current practice. Our focus has been on using technology to be more innovative, efficient, and interesting without a thought to integrating test purposes. We usually offer technology (thoughtfully constructed) as an answer to many educational problems, but in this case, we have played at the margin without trying to generate a multipurpose model, with domain independent and dependent elements. We certainly have the technology to do so. Where is our motivation? Where are the validity studies that demonstrate required action? Where is the science or practical evidence about integrated systems? If we are serious about the value of large‐scale testing for classroom teaching and learning, then in addition to documenting the sad state of integrated assessment, we need to take action.</p> <p>Overall, our profiles underscored the significant tensions that currently exist between classroom and large‐scale assessment in the majority of countries examined. Although synergy is possible, it remains elusive in systems that rely on test‐based accountability models. The latter is somewhat perplexing given the substantial empirical evidence that has consistently documented the positive effects of formative assessment, the deleterious effects of high‐stakes external testing, and the potential for intelligent accountability models (Stobart, 2008). If anything, the gulf between assessment evidence and policy in contemporary education systems suggests that we have a long way to go in terms of promoting evidenced‐based policies.</p> <p>Interestingly, the International Association for the Evaluation of Educational Achievement recently established an international digital literacy assessment and the Organisation for Economic Cooperation and Development (OECD) is also currently developing newer assessments in areas such as digital learning and creative thinking (OECD, 2019). This expansion will undoubtedly spur policymakers to consider a broader array of "essential skills" in students which may provoke a rethinking of the priorities of assessment—at both large scale and classroom levels (Volante, 2019). Certainly, recent cross‐cultural research suggests that national assessment systems—such as those in Germany and Canada—are increasingly being modeling to align with PISA (Volante, 2018). This convergence of assessment purposes (and models) has largely characterized standards‐based reform contexts for the last century (Baker, 2016). Thus, we are left with the prospect that an expansion of tested subject matter and domains may expand teacher's and schools' priorities—albeit in predictable the ways that produce persistent washback effects.</p> <ref id="AN0147461954-11"> <title> References </title> <blist> <bibl id="bib1" idref="ref1" type="bt">1</bibl> <bibtext> Bach, A., Wurster, S., Thillmann, K., Pant Anand, H., & Thiel, F. (2014). Vergleichsarbeiten und schulische Personalentwicklung ‐ Ausmaß und Voraussetzungen der Datennutzung [Comparative tests and teacher development: Extent of data use and requirements for this use]. Zeitschrift für Erziehungswissenschaft, 17, 61 – 84.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref2" type="bt">2</bibl> <bibtext> Baird, J., Andrich, D., Hopfenbeck, T. N., & Stobart, G. (2017). Assessment and learning: Fields apart? Assessment in Education: Principles, Policy & Practice, 24 (3), 317 – 350.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref3" type="bt">3</bibl> <bibtext> Baker, E. L., Chung, G. K. W. K., & Cai, L. (2016). Assessment gaze, refraction, and blur: The course of achievement testing in the past 100 years. Review of Research in Education, 40 (1), 94 – 142. https://doi.org/10.3102/0091732X16679806</bibtext> </blist> <blist> <bibl id="bib4" type="bt">4</bibl> <bibtext> Birenbaum, M., DeLuca, C., Earl, L., Heritage, M., Klenowski, V., Looney, A., ... Wyatt‐Smith, C. (2015). International trends in the implementation of assessment for learning: Implications for policy and practice. Policy Futures in Education, 13 (1), 117 – 140.</bibtext> </blist> <blist> <bibl id="bib5" type="bt">5</bibl> <bibtext> Böhm‐Kasper, O., Selders, O., & Lambrecht, M. (2016). Schulinspektion und Schulentwicklung ‐ Ergebnisse der quantitativen Schulleitungsbefragung [School inspection and school development – Results from a quantitative survey amongst school principals]. In Arbeitsgruppe Schulinspektion (Ed.), Schulinspektion als Steuerungsimpuls? Ergebnisse aus Forschungsprojekten (pp. 1 – 50). Wiesbaden : Springer.</bibtext> </blist> <blist> <bibl id="bib6" type="bt">6</bibl> <bibtext> Böttcher, W. (2014). Steuerung durch Bildungspolitik. Wie Politik die Schulen steuert ‐ und was sie vernachlässigt [Educational policy. How politics governs schools – And what it neglects]. Pädagogik, 66 (5), 44 – 47.</bibtext> </blist> <blist> <bibl id="bib7" type="bt">7</bibl> <bibtext> Böttcher, W., Hense, J., & Keune, M. (2013). Schulinspektion als eine Form externer Evaluation ‐ ein Forschungsüberblick [School inspection as a form of external evaluation – A review]. In J. Hense, S. Rädiker, W. Böttcher, & T. Widmer (Eds.), Forschung über Evaluation. Bedingungen, Prozesse und Wirkungen (pp. 231 – 250). Münster u.a : Waxmann.</bibtext> </blist> <blist> <bibl id="bib8" type="bt">8</bibl> <bibtext> Brookhart, S. M. (2019). Feedback and measurement. In S. M. Brookhart & J. H. McMillan (Eds.), Classroom assessment and educational measurement. (pp. 63 – 78). New York, NY : Routledge.</bibtext> </blist> <blist> <bibl id="bib9" type="bt">9</bibl> <bibtext> Cife. (2018). The new A level and GCSE exams. Retrieved from https://<ulink href="http://www.cife.org.uk/article/the-new-a-level-and-gcse-exams/">www.cife.org.uk/article/the-new-a-level-and-gcse-exams/</ulink></bibtext> </blist> <blist> <bibtext> Cumming, J. J., Van Der Kleij, F. M., & Adie, L. (2019). Contesting educational assessment policies in Australia. Journal of Education Policy, 34 (6), 836 – 857. https://doi.org/10.1080/02680939.2019.1608375</bibtext> </blist> <blist> <bibtext> Dann, R. (2018). Developing feedback for pupil learning. New York, NY : Routledge.</bibtext> </blist> <blist> <bibtext> Dee, T., & Jacob, B. A. (2010). The impact of no child left behind on students, teachers, and schools. Brookings Papers on Economic Activity (pp. 149–207).</bibtext> </blist> <blist> <bibtext> DeLuca, C., Braund, H., Valiquette, A., & Cheng, L. (2017). Grading policies and practices in Canada: A landscape study. Canadian Journal of Educational Administration and Policy, 184, 4 – 22.</bibtext> </blist> <blist> <bibtext> Deneen, C., Fulmer, G. W., Brown, G. T. L., Tan, K., Leong, W. S., & Tay, H. Y. (2019). Value, practice and proficiency: Teachers' complex relationship with assessment for learning. Teacher and Teacher Education, 80, 39 – 47.</bibtext> </blist> <blist> <bibtext> Gonski, D., Arcus, T., Boston, K., Gould, V., Johnson, W., O'Brien, L., ... Roberts, M. (2018). Through growth to achievement: Report of the review to achieve educational excellence in Australian schools. Department of Education and Training. Retrieved from https://<ulink href="http://www.appa.asn.au/wp-content/uploads/2018/04/20180430-Through-Growth-to-Achievement%5fText.pdf">www.appa.asn.au/wp-content/uploads/2018/04/20180430-Through-Growth-to-Achievement%5fText.pdf</ulink></bibtext> </blist> <blist> <bibtext> Harju‐Luukkainen, H., Nissinen, K., Sulkunen, S., & Suni, M. (2014). Avaimet osaamiseen ja tulevaisuuteen: Selvitys maahanmuuttajataustaisten nuorten osaamisen tasosta ja siihen liittyvistä taustatekijöistä PISA 2012 – tutkimuksessa [Keys to competence and future: A report on PISA 2012 results and related underlying factors for students with an immigrant background]. Finnish Institute for Educational Research.</bibtext> </blist> <blist> <bibtext> Harju‐Luukkainen, H., Vettenranta, J., Oukrim‐Soivio, N., & Bernelius, V. (2016). Differences between PISA reading literacy scores and grading for mother tongue and literature at school: A geostatistical analysis of the Finnish PISA 2009 data. Education Inquiry, 7 (4), 463 – 479.</bibtext> </blist> <blist> <bibtext> Hautamäki, J., Kupiainen, S., Marjanen, J., Vainikainen, M.‐P., & Hotulainen, R. (2013). Oppimaan oppiminen peruskoulun päättövaiheessa: Tilanne vuonna 2012 ja muutos vuodesta 2001 [Learning to learn at the end of basic education: Situation in 2012 and change from 2001]. Department of Teacher Education Research Report 347. University of Helsinki, Unigrafia.</bibtext> </blist> <blist> <bibtext> Holloway, J., Sørensen, T. B., & Verger, A. (2017). Global perspectives on high‐stakes teacher accountability policies: An introduction. Education Policy Analysis Archives, 25 (85), 1 – 18. https://doi.org/10.14507/epaa.25.3325</bibtext> </blist> <blist> <bibtext> Hopfenbeck, T., & Gorgen, K. (2017). The politics of PISA: The media, policy and public responses in Norway and England, European Journal of Education, 52 (2), 192 – 205.</bibtext> </blist> <blist> <bibtext> Hopmann, S. T. (2008). No child, no school, no state left behind: Schooling in the age of accountability. Journal of Curriculum Studies, 40 (4), 417 – 456.</bibtext> </blist> <blist> <bibtext> Klinger, D., DeLuca, C., & Miller, T. (2008). The evolving culture of large‐scale assessments in Canadian education. Canadian Journal of Educational Administration and Policy, 76, 1 – 34.</bibtext> </blist> <blist> <bibtext> KMK (Ständige Konferenz der Kultusminister der Länder in der Bundesrepublik Deutschland). (2015). Gesamtstrategie der Kultusministerkonferenz zum Bildungsmonitoring [Comprehensive strategy of the Standing Conference of Cultural Ministers for educational monitoring]. Retrieved from https://<ulink href="http://www.kmk.org/fileadmin/Dateien/veroeffentlichungen%5fbeschluesse/2015/2015%5f06%5f11-Gesamtstrategie-Bildungsmonitoring.pdf">www.kmk.org/fileadmin/Dateien/veroeffentlichungen%5fbeschluesse/2015/2015%5f06%5f11-Gesamtstrategie-Bildungsmonitoring.pdf</ulink></bibtext> </blist> <blist> <bibtext> Latham, H. (1886). On the action of examinations: Considered as a means of selection. Willard Small. Retrieved from https://books.google.ca/books?hl=en&lr=&id=mvUBAAAAYAAJ&oi=fnd&pg=PR1&ots=N-paw-67-W&sig=b7rluv4q-K-OWre_2zI8Afqwiv4&redir_esc=y#v=onepage&q&f=false</bibtext> </blist> <blist> <bibtext> Leong, W. S., & Tan, K. (2014). What (more) can, and should, assessment do for learning? Observations from 'successful learning context' in Singapore. The Curriculum Journal, 25 (4), 593 – 619.</bibtext> </blist> <blist> <bibtext> Lingard, B., Thompson, G., & Sellar, S. (2017). National testing from an Australian perspective. In B. Lingard, G. Thompson, & S. Sellar (Eds.), National testing in schools: An Australian assessment (pp. 1 – 17). New York, NY : Routledge.</bibtext> </blist> <blist> <bibtext> Mazzeo, C. (2001). Frameworks of state: Assessment policy in historical perspective. Teachers College Record, 103, 367 – 397.</bibtext> </blist> <blist> <bibtext> Ministry of Education. (2018). 'Learn for life' – Preparing our students to excel beyond exam results. Retrieved from https://<ulink href="http://www.moe.gov.sg/news/press-releases/-learn-for-life—preparing-our-students-to-excel-beyond-exam-results">www.moe.gov.sg/news/press-releases/-learn-for-life—preparing-our-students-to-excel-beyond-exam-results</ulink></bibtext> </blist> <blist> <bibtext> National Research Council. (2001). Knowing what students know: The science of design and educational assessment. Washington, DC : National Academy Press.</bibtext> </blist> <blist> <bibtext> Organisation for Economic Cooperation and Development. (2019). PISA 2021 Creative thinking framework. Retrieved from https://<ulink href="http://www.oecd.org/pisa/publications/PISA-2021-creative-thinking-framework.pdf">www.oecd.org/pisa/publications/PISA-2021-creative-thinking-framework.pdf</ulink></bibtext> </blist> <blist> <bibtext> Pellegrino, J. W., & Chudowsky, N. (2003). The foundations of assessment. Measurement: Interdisciplinary Research and Perspectives, 1 (2), 103 – 148.</bibtext> </blist> <blist> <bibtext> Pellegrino, J. W., DiBello, L. V., & Goldman, S. R. (2016). A framework for conceptualizing and evaluating the validity of instructionally relevant assessments. Educational Psychologist, 51 (1), 59 – 81.</bibtext> </blist> <blist> <bibtext> Perie, M., Marion, S., & Gong, B. (2009). Moving toward a comprehensive assessment system: A framework for considering interim assessments. Educational Measurement: Issues and Practice, 28 (3), 5 – 13.</bibtext> </blist> <blist> <bibtext> Pietsch, M., Janke, N., & Mohr, I. (2014). Führt Schulinspektion zu besseren Schülerleistungen? Difference‐in‐Differences‐Studien zu Effekten der Schulinspektion Hamburg auf Lernzuwächse und Leistungstrends [Does school inspection result in higher competence of students? Difference‐in‐difference‐studies on the effects of Hamburg's school inspection on learning outcomes and trends in achievement]. Zeitschrift für Pädagogik, 60 (3), 446 – 470.</bibtext> </blist> <blist> <bibtext> Ratnam‐Lim, C., & Tan, K. H. (2015). Large‐scale implementation of formative assessment practices in an examination oriented culture. Assessment in Education, 22 (1), 61 – 78.</bibtext> </blist> <blist> <bibtext> Rürup, M. (2008). Typen der Schulinspektion in den deutschen Bundesländern [Types of school inspection in the federal states of Germany]. Die Deutsche Schule, 100 (4), 467 – 477.</bibtext> </blist> <blist> <bibtext> Shepard, L. A. (2019). Classroom assessment to support teaching and learning. In A. Berman, M. J. Feuer, & J. W. Pellegrino (Eds.), The ANNALS of the American Academy of Political and Social Science (pp. 183 – 200). Thousand Oaks, CA : Sage Publications.</bibtext> </blist> <blist> <bibtext> Smith, W. C. (2014). The global transformation toward testing for accountability. Education Policy Analysis Archives, 22 (116), 1 – 34. https://doi.org/10.14507/epaa.v22.1571.</bibtext> </blist> <blist> <bibtext> Staufenberg, J. (2017). Ofsted inspectors urged to crack down on schools 'off‐rolling' pupils. Retrieved from https://schoolsweek.co.uk/ofsted-inspectors-urged-to-crack-down-on-schools-off-rolling-pupils/</bibtext> </blist> <blist> <bibtext> Stobart, G. (2008). Testing times: The uses and abuses of assessment. New York, NY : Routledge.</bibtext> </blist> <blist> <bibtext> Tan, K. (2011). Assessment for learning reform in Singapore ‐ enduring, sustainable or threshold? Assessment Reform in Education. In R. Berry & B. Adamson (Eds.), Assessment in education (pp. 75 – 88). Cham, Switzerland : Springer.</bibtext> </blist> <blist> <bibtext> Tan, K., & Deneen, C. C. (2015). Aligning and sustaining meritocracy, curriculum and assessment validity in Singapore. Assessment Matters, 8, 31 – 52.</bibtext> </blist> <blist> <bibtext> Thomas, D. P., Emery, S., Prain, V., Papageorgiou, J., & McKendrick, A.‐M. (2019). Influences on local curriculum innovation in times of change: A literacy case study. Australian Educational Researcher, 46 (3), 469 – 487.</bibtext> </blist> <blist> <bibtext> Vainikainen, M. P., Thuneberg, H., Marjanen, J., Hautamäki, J., Kupiainen, S., & Hotulainen, R. (2017). How do Finns know? Educational monitoring without inspection and standard‐setting. In S. Blömeke & J.‐E. Gustafsson (Eds.), Standard setting: International state of research and practices in the Nordic countries (pp. 243 – 259). Cham : Springer. Retrieved from 10.1007/978-3-319-50856-6_14.</bibtext> </blist> <blist> <bibtext> Vainikainen, M.‐P., & Harju‐Luukkainen, H. (in press). Educational assessment in Finland. In H. Harju‐Luukkainen, N. McElvany, & J. Stang (Eds.), Monitoring of student achievement in the 21st century. European policy perspectives and assessment strategies. Cham : Springer.</bibtext> </blist> <blist> <bibtext> Verger, A., & Parcerisa, L. (2017). A difficult relationship. Accountability policies and teachers: International evidence and key premises for future research. In M. Akiba & G. LeTendre (Eds.), International handbook of teacher quality and policy (pp. 241 – 254). New York, NY : Routledge.</bibtext> </blist> <blist> <bibtext> Volante, L. (Ed.). (2012). School Leadership in the context of standards‐based reform: International perspectives. Cham : Springer.</bibtext> </blist> <blist> <bibtext> Volante, L. (Ed.). (2016). The intersection of international achievement testing and educational policy: Global perspectives on large‐scale reform. New York; London : Routledge.</bibtext> </blist> <blist> <bibtext> Volante, L. (Ed.). (2018). The PISA effect on global educational governance. New York; London : Routledge.</bibtext> </blist> <blist> <bibtext> Volante, L., Schnepf, S., Jerrim, J., & Klinger, D. (Eds.). (2019). Socioeconomic inequality and student outcomes: Cross‐national trends, policies, and practices. Singapore : Springer.</bibtext> </blist> <blist> <bibtext> Volante, L., & Ben Jaafar, S. (2008). Educational assessment in Canada. Assessment in Education: Principles, Policy & Practice, 15 (2), 201 – 210.</bibtext> </blist> <blist> <bibtext> Wilson, M., & Scalise, K. (2016). Learning analytics: Negotiating the intersection of measurement technology and information technology. In M. J. Spector, B. B. Lockee, & M. D. Childress (Eds.), Learning, design, and technology: An international compendium of theory, research, practice, and policy. (pp. 1 – 23). Cham : Springer.</bibtext> </blist> <blist> <bibtext> Wurster, S., Richter, D., & Lenski, A. E. (2017). Datenbasierte Unterrichtsentwicklung und ihr Zusammenhang zur Schülerleistung [Teachers' use of evaluation data to improve instruction and its relationship to student achievement]. Zeitschrift für Erziehungswissenschaft, 20, 628 – 650.</bibtext> </blist> <blist> <bibtext> Wyatt‐Smith, C., & Jackson, C. (2016). A picture of accelerating negative change. The Australian Journal of Language and Literacy, 39 (3), 233 – 244.</bibtext> </blist> </ref> <aug> <p>By Louis Volante; Christopher DeLuca; Lenore Adie; Eva Baker; Heidi Harju‐Luukkainen; Margaret Heritage; Christoph Schneider; Gordon Stobart; Kelvin Tan and Claire Wyatt‐Smith</p> <p>Reported by Author; Author; Author; Author; Author; Author; Author; Author; Author; Author</p> </aug>
Header DbId: eric
DbLabel: ERIC
An: EJ1276712
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Synergy and Tension between Large-Scale and Classroom Assessment: International Trends
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Volante%2C+Louis%22">Volante, Louis</searchLink><br /><searchLink fieldCode="AR" term="%22DeLuca%2C+Christopher%22">DeLuca, Christopher</searchLink><br /><searchLink fieldCode="AR" term="%22Adie%2C+Lenore%22">Adie, Lenore</searchLink><br /><searchLink fieldCode="AR" term="%22Baker%2C+Eva%22">Baker, Eva</searchLink><br /><searchLink fieldCode="AR" term="%22Harju-Luukkainen%2C+Heidi%22">Harju-Luukkainen, Heidi</searchLink><br /><searchLink fieldCode="AR" term="%22Heritage%2C+Margaret%22">Heritage, Margaret</searchLink><br /><searchLink fieldCode="AR" term="%22Schneider%2C+Christoph%22">Schneider, Christoph</searchLink><br /><searchLink fieldCode="AR" term="%22Stobart%2C+Gordon%22">Stobart, Gordon</searchLink><br /><searchLink fieldCode="AR" term="%22Tan%2C+Kelvin%22">Tan, Kelvin</searchLink><br /><searchLink fieldCode="AR" term="%22Wyatt-Smith%2C+Claire%22">Wyatt-Smith, Claire</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Educational+Measurement%3A+Issues+and+Practice%22"><i>Educational Measurement: Issues and Practice</i></searchLink>. Win 2020 39(4):21-29.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 9
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2020
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Evaluative
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Educational+Trends%22">Educational Trends</searchLink><br /><searchLink fieldCode="DE" term="%22Trend+Analysis%22">Trend Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Measurement%22">Measurement</searchLink><br /><searchLink fieldCode="DE" term="%22Teaching+Methods%22">Teaching Methods</searchLink><br /><searchLink fieldCode="DE" term="%22Educational+Assessment%22">Educational Assessment</searchLink><br /><searchLink fieldCode="DE" term="%22Learning+Processes%22">Learning Processes</searchLink><br /><searchLink fieldCode="DE" term="%22Cross+Cultural+Studies%22">Cross Cultural Studies</searchLink><br /><searchLink fieldCode="DE" term="%22Foreign+Countries%22">Foreign Countries</searchLink><br /><searchLink fieldCode="DE" term="%22Educational+Policy%22">Educational Policy</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Results%22">Test Results</searchLink><br /><searchLink fieldCode="DE" term="%22Correlation%22">Correlation</searchLink>
– Name: Subject
  Label: Geographic Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22United+States%22">United States</searchLink><br /><searchLink fieldCode="DE" term="%22Canada%22">Canada</searchLink><br /><searchLink fieldCode="DE" term="%22Finland%22">Finland</searchLink><br /><searchLink fieldCode="DE" term="%22Singapore%22">Singapore</searchLink><br /><searchLink fieldCode="DE" term="%22Australia%22">Australia</searchLink><br /><searchLink fieldCode="DE" term="%22Germany%22">Germany</searchLink><br /><searchLink fieldCode="DE" term="%22United+Kingdom+%28England%29%22">United Kingdom (England)</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1111/emip.12382
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0731-1745
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: The synergy, or lack thereof, between large-scale and classroom assessment has been fiercely debated in both academic and policy spheres for decades around the world. This paper seeks to explicate how different countries are utilizing large-scale testing and test results at the classroom level. Through country profiles, this paper analyzes contemporary developments on the tensions and synergies between large-scale assessment and classroom teaching, learning, and assessment observed across seven international jurisdictions: United States, Canada, Australia, England, Germany, Finland, and Singapore. The paper concludes with an analysis of international trends leading to a synthesis of root causes contributing to the current limited uptake of large-scale assessment results at classroom levels.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2020
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1276712
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1276712
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1111/emip.12382
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 9
        StartPage: 21
    Subjects:
      – SubjectFull: Educational Trends
        Type: general
      – SubjectFull: Trend Analysis
        Type: general
      – SubjectFull: Measurement
        Type: general
      – SubjectFull: Teaching Methods
        Type: general
      – SubjectFull: Educational Assessment
        Type: general
      – SubjectFull: Learning Processes
        Type: general
      – SubjectFull: Cross Cultural Studies
        Type: general
      – SubjectFull: Foreign Countries
        Type: general
      – SubjectFull: Educational Policy
        Type: general
      – SubjectFull: Test Results
        Type: general
      – SubjectFull: Correlation
        Type: general
      – SubjectFull: United States
        Type: general
      – SubjectFull: Canada
        Type: general
      – SubjectFull: Finland
        Type: general
      – SubjectFull: Singapore
        Type: general
      – SubjectFull: Australia
        Type: general
      – SubjectFull: Germany
        Type: general
      – SubjectFull: United Kingdom (England)
        Type: general
    Titles:
      – TitleFull: Synergy and Tension between Large-Scale and Classroom Assessment: International Trends
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Volante, Louis
      – PersonEntity:
          Name:
            NameFull: DeLuca, Christopher
      – PersonEntity:
          Name:
            NameFull: Adie, Lenore
      – PersonEntity:
          Name:
            NameFull: Baker, Eva
      – PersonEntity:
          Name:
            NameFull: Harju-Luukkainen, Heidi
      – PersonEntity:
          Name:
            NameFull: Heritage, Margaret
      – PersonEntity:
          Name:
            NameFull: Schneider, Christoph
      – PersonEntity:
          Name:
            NameFull: Stobart, Gordon
      – PersonEntity:
          Name:
            NameFull: Tan, Kelvin
      – PersonEntity:
          Name:
            NameFull: Wyatt-Smith, Claire
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2020
          Identifiers:
            – Type: issn-print
              Value: 0731-1745
          Numbering:
            – Type: volume
              Value: 39
            – Type: issue
              Value: 4
          Titles:
            – TitleFull: Educational Measurement: Issues and Practice
              Type: main
ResultId 1