A Two-Stage Method for Classroom Assessments of Essay Writing

Saved in:
Bibliographic Details
Title: A Two-Stage Method for Classroom Assessments of Essay Writing
Language: English
Authors: Humphry, Stephen Mark, Heldsinger, Sandy
Source: Journal of Educational Measurement. Fall 2019 56(3):505-520.
Availability: Wiley-Blackwell. 350 Main Street, Malden, MA 02148. Tel: 800-835-6770; Tel: 781-388-8598; Fax: 781-388-8232; e-mail: cs-journals@wiley.com; Web site: http://www.wiley.com/WileyCDA
Peer Reviewed: Y
Page Count: 16
Publication Date: 2019
Document Type: Journal Articles
Reports - Evaluative
Education Level: Elementary Education
Descriptors: Essays, Elementary School Students, Writing (Composition), Writing Evaluation, Elementary School Teachers, Persuasive Discourse, Achievement Rating, Interrater Reliability, Expertise
DOI: 10.1111/jedm.12223
ISSN: 0022-0655
Abstract: To capitalize on professional expertise in educational assessment, it is desirable to develop and test methods of rater-mediated assessment that enable classroom teachers to make reliable and informative judgments. Accordingly, this article investigates the reliability of a two-stage method used by classroom teachers to assess primary school students' persuasive writing. Stage 1 involves pairwise comparisons and stage 2 involves rating against calibrated exemplars from stage 1 plus performance descriptors. A high level of interrater reliability among teachers was obtained. This is consistent with previous evidence that the two-stage method is a viable classroom assessment method without extensive training and moderation. Implications for assessment practices in education are discussed with a focus on the widely expressed desire to value professional expertise.
Abstractor: As Provided
Entry Date: 2019
Accession Number: EJ1227713
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwFARx8ZpCKmR-CWcO-HddPiAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDI1yv0HLKAYQF8LSawIBEICBmy5T5Mgjz8kFp6KdfpHoZjkGZ8lt0FtgMKtSymdWBK4gZvri-G3DfUQdIDDhBltqrEOkc16XLVoGOGyG3cRsOhIIZgMk_MYIz0dxwRKtgm2dnO6fpqQUS8OMAmyFnmnV3fxWfVTGy9XltJK4Zzn7_sN982elN3W7Mj25WagKWUcjO-Qv6ofmwWLHCmkvyDsdcWdrdUdUC9Hauyro
Text:
  Availability: 1
  Value: <anid>AN0138541329;mea01sep.19;2019Sep12.03:10;v2.2.500</anid> <title id="AN0138541329-1">A Two‐Stage Method for Classroom Assessments of Essay Writing </title> <p>To capitalize on professional expertise in educational assessment, it is desirable to develop and test methods of rater‐mediated assessment that enable classroom teachers to make reliable and informative judgments. Accordingly, this article investigates the reliability of a two‐stage method used by classroom teachers to assess primary school students' persuasive writing. Stage 1 involves pairwise comparisons and stage 2 involves rating against calibrated exemplars from stage 1 plus performance descriptors. A high level of interrater reliability among teachers was obtained. This is consistent with previous evidence that the two‐stage method is a viable classroom assessment method without extensive training and moderation. Implications for assessment practices in education are discussed with a focus on the widely expressed desire to value professional expertise.</p> <p>There has been considerable work internationally in recent decades to support teachers in their use of assessments to improve student learning. Assessment is inherent to effective teaching and a teacher's professional judgment influences various aspects of their work (e.g., Allal, [<reflink idref="bib1" id="ref1">1</reflink>]; Black & Wiliam, [<reflink idref="bib7" id="ref2">7</reflink>]; Du Four, [<reflink idref="bib16" id="ref3">16</reflink>]; Hattie & Timperley, [<reflink idref="bib19" id="ref4">19</reflink>]). Although there is a widespread preference for teacher judgment as the primary means of obtaining summative assessments, there are challenges associated with obtaining reliable teacher judgments (Brookhart, [<reflink idref="bib12" id="ref5">12</reflink>]; Harlen, [<reflink idref="bib17" id="ref6">17</reflink>]; Johnson, [<reflink idref="bib23" id="ref7">23</reflink>]).</p> <p>A growing body of research is showing that teachers obtain reliable judgments using the method of pairwise comparison in a range of learning domains (Steedle & Ferrara, [<reflink idref="bib35" id="ref8">35</reflink>]). Direct pairwise comparisons leaves little or no scope for differences in rater harshness (Andrich, [<reflink idref="bib2" id="ref9">2</reflink>]). However, pairwise comparisons are time‐consuming (Bramley, Bell, & Pollitt, [<reflink idref="bib10" id="ref10">10</reflink>]) and do not directly provide information that can be used formatively. In order to capitalize on the high levels of reliability typically obtained using the method of pairwise comparisons, while enabling a more time‐effective process, a two‐stage method is adopted in this research. The two‐stage method is designed to enable reliable and comparable school‐based assessments with minimal training or professional development. The aim of the research is to ascertain the degree to which this two‐stage method of assessment can be used by classroom teachers to reliably assess student performances on persuasive essay tasks.</p> <p>In the first stage, a relatively large sample of performances is calibrated using the method of pairwise comparison. In the second stage teachers score their students' work by placing them on the calibrated scale, using implicit comparisons as well as descriptors derived from the calibrated exemplars. The reliability of judgments using this method of assessment was previously examined in the context of recount writing obtained from 4–7‐year‐olds (Heldsinger & Humphry, [<reflink idref="bib21" id="ref11">21</reflink>]). The aim of the study is to investigate the reliability of the method in the context of persuasive writing obtained from students aged 5–12.</p> <p>The article is set out as follows. First, background is provided about teacher judgments and the desire to place more value on the judgments made by teachers as raters. A brief overview is then presented of the literature focusing on the reliability of classroom assessments. Following this, the article provides an overview of the use of pairwise comparisons in educational contexts, and the rationale for the two‐stage method is outlined. The design of the empirical study is detailed and results are presented. Lastly, the article discusses the results in light of the objectives of achieving reliable classroom assessments and placing greater emphasis on classroom assessments.</p> <hd id="AN0138541329-2">Background</hd> <p></p> <hd id="AN0138541329-3">Teacher judgments</hd> <p>Johnson ([<reflink idref="bib23" id="ref12">23</reflink>]) observed a widely expressed desire in education to value teachers' professional expertise and, at least in the United Kingdom, an associated belief that classroom summative assessment may counteract the perceived issues of externally imposed testing. In 2015, The Commission for Assessment Without Levels found that too great a reliance was being placed by government on external tests, particularly for school accountability purposes (Standards and Testing Agency, [<reflink idref="bib34" id="ref13">34</reflink>]). There is little indication internationally that genuine gains have been made in establishing parity between classroom assessments and external standardized tests. Possible reasons are indicated in the two reviews of the research evidence of the reliability and validity of classroom assessment referred to earlier: the U.K. review conducted by Harlen ([<reflink idref="bib17" id="ref14">17</reflink>]) and the long‐term review by Brookhart ([<reflink idref="bib12" id="ref15">12</reflink>]).</p> <p>In the review focusing on the United Kingdom it was concluded that "the findings ... by no means constitute a ringing endorsement of teachers' assessment; there was evidence of low reliability and bias in teachers' judgments made in certain circumstances" (Harlen, [<reflink idref="bib18" id="ref16">18</reflink>], p. 245). In a long‐term review focusing on the United States, Brookhart ([<reflink idref="bib12" id="ref17">12</reflink>]) concluded that for public accountability, more trust is placed in standardized tests than teacher judgments, and commented that until the quality of teacher judgements in summative assessment is addressed in both research and practice, students and schools in the USA will probably continue to be addressed in a two‐level manner. Teachers' grades will have immediate, localised effects on students, but for the more politically charged judgements of school accountability, standardised tests will garner more trust. (p. 86)</p> <p>The general lack of value and trust in teachers' professional judgments is unfortunate given that classroom assessments influence various aspects of teaching from lesson planning to decision making (Allal, [<reflink idref="bib1" id="ref18">1</reflink>]). Classroom assessments include a range of activities such as a teacher creating an assessment rubric, using a preexisting marking key, making judgments for reporting grades, and making an on‐balance judgment of a performance against criteria. Such assessment activities may be undertaken alone or with other teachers, and with or without explicit exemplification of performance levels.</p> <p>The breadth of activity covered by the concept of summative assessment by teachers is a confounding issue when evaluating the reliability of teacher summative judgment. The purpose of this article is not to examine reliability of various forms of assessment. However, the breadth of activities is relevant to the present research because it is seldom acknowledged that task design and judgment process are likely to impact the reliability of teacher judgments. The purpose of the present research is to examine the reliability of an assessment method specifically designed to be accessible to school teachers.</p> <p>In reviewing the evidence for the reliability and validity of teacher judgment, Harlen ([<reflink idref="bib18" id="ref19">18</reflink>]) identified 12 studies that investigated the reliability of classroom assessment. The studies focused on remarking or moderation of the data from the assessment process, or the influence of school or student variables. Only three of these studies employed experimental control or manipulation in the research. The rest of the studies examined teacher judgment in naturally occurring contexts or looked at the relationship between variables. The review observed low reliability and bias, and recommended amongst other things that "there needs to be research into the effectiveness of different approaches to improving the dependability of teachers' summative assessment including moderation procedures" (Harlen, [<reflink idref="bib18" id="ref20">18</reflink>], p. 268).</p> <p>Brookhart ([<reflink idref="bib12" id="ref21">12</reflink>]) found that research generally could be divided into two broad categories: teachers' summative grading practices, and examinations of the relation between teacher judgments and large‐scale summative assessment, most of which were standardized tests. The review found that the quality of teacher judgments was variable.</p> <hd id="AN0138541329-4">Reliability of teacher judgments</hd> <p>There has been a growing focus on the reliability and consistency of teachers' judgments, which has increased over the past decades at least partly in response to the increase and perceived significance of high‐stakes testing and reporting of results at a school level (Connolly, Klenowski, & Wyatt‐Smith, [<reflink idref="bib15" id="ref22">15</reflink>]). Although an increasing number of research studies provide evidence that teacher judgments are not reliable, a review indicated that there are few studies reporting investigations of reliability of classroom assessment in the literature (Johnson, [<reflink idref="bib23" id="ref23">23</reflink>], p. 94).</p> <p>In designing assessments that enable higher levels of reliability, it is helpful to consider factors that influence the quality of classroom assessments. One factor that can affect reliability and harshness is the amount of training on the topic of assessment or a particular assessment approach that a teacher has received (McMahon & Jones, [<reflink idref="bib30" id="ref24">30</reflink>]; Wiliam, [<reflink idref="bib39" id="ref25">39</reflink>]). Other factors that can impact reliability include training of teachers in assessment and the design of assessment tasks (McMahon & Jones, [<reflink idref="bib30" id="ref26">30</reflink>]). Relevant to the objectives of this study, Choi ([<reflink idref="bib14" id="ref27">14</reflink>]) maintains that there is a need to provide professional development whereas McMahon and Jones ([<reflink idref="bib30" id="ref28">30</reflink>]) have noted that when pairwise comparisons are used as a method of assessment, reliable classroom assessments can occur without the use of training or professional development (see also Jones & Alcock, [<reflink idref="bib24" id="ref29">24</reflink>]; Jones, Inglis, Gilmore, & Hodgen, [<reflink idref="bib26" id="ref30">26</reflink>]).</p> <hd id="AN0138541329-5">Pairwise comparisons and its applications in education</hd> <p>The method of pairwise comparisons is also referred to as comparative judgment because a model for its application was first formulated by Thurstone ([<reflink idref="bib37" id="ref31">37</reflink>]) and referred to as the Law of Comparative Judgment. In educational contexts, pairwise comparisons entail judges comparing pairs of performances and judging which performance is of a higher quality. The data are then analyzed to scale the performances, typically using the Bradley–Terry–Luce model (Bradley & Terry, [<reflink idref="bib9" id="ref32">9</reflink>]; Luce, [<reflink idref="bib28" id="ref33">28</reflink>]). Background to the use of the method of pairwise comparisons in education is provided by Bramley et al. ([<reflink idref="bib10" id="ref34">10</reflink>]), Heldsinger and Humphry ([<reflink idref="bib20" id="ref35">20</reflink>]), and Pollitt ([<reflink idref="bib31" id="ref36">31</reflink>]).</p> <p>There is a relatively small but growing body of work on the application of pairwise comparisons in educational contexts. Much of the research focuses on applications of pairwise comparisons to produce scales in a range of areas. Perhaps the most focus to date has been on judging and scaling students' writing performances (Heldsinger & Humphry, [<reflink idref="bib20" id="ref37">20</reflink>]; Steedle & Ferrara, [<reflink idref="bib35" id="ref38">35</reflink>]; van Daal, Lesterhuis, Coertjens, Donche, & De Maeyer, [<reflink idref="bib38" id="ref39">38</reflink>]) and students' problem solving in mathematics (Bisson, Gilmore, Inglis, & Jones, [<reflink idref="bib6" id="ref40">6</reflink>]; Jones & Inglis, [<reflink idref="bib25" id="ref41">25</reflink>]; Jones, Swan, & Pollitt, [<reflink idref="bib27" id="ref42">27</reflink>]). Other areas of focus have included peer assessment and feedback (Potter et al., [<reflink idref="bib32" id="ref43">32</reflink>]; Seery, Canty, & Phelan, [<reflink idref="bib33" id="ref44">33</reflink>]), the assessment of creative performances (Tarricone & Newhouse, [<reflink idref="bib36" id="ref45">36</reflink>]), and the assessment of oral narrative performances (Humphry, Heldsinger, & Dawkins, [<reflink idref="bib22" id="ref46">22</reflink>]). There has also been some research into the application of pairwise comparisons in setting grade boundaries (Benton & Elliot, [<reflink idref="bib5" id="ref47">5</reflink>]). To date, the internal consistency obtained using the method has typically ranged from high to very high, with most applications obtaining very high levels (Steedle & Ferrara, [<reflink idref="bib35" id="ref48">35</reflink>]).</p> <p>There is some emerging research into technical issues including adaptive approaches to comparative judgment (Pollitt, [<reflink idref="bib31" id="ref49">31</reflink>]), dependencies and related issues that arise in the main adaptive approach employed to date (Bramley & Vitello, [<reflink idref="bib11" id="ref50">11</reflink>]), and models and associated estimation procedures (Cattelan, [<reflink idref="bib13" id="ref51">13</reflink>]).</p> <p>Although the method of pairwise comparisons has typically produced very high levels of internal reliability, in their standard form pairwise comparisons are likely to be too time‐consuming and tedious to be viable as a general method for classroom assessment (Bramley et al., [<reflink idref="bib10" id="ref52">10</reflink>]; McMahon & Jones, [<reflink idref="bib30" id="ref53">30</reflink>]). Findings reported by Steedle and Ferrara ([<reflink idref="bib35" id="ref54">35</reflink>]) indicate the time to assess at least some performances may not be too much larger using pairwise comparisons than a rubric; however, it is difficult to identify from those findings a like‐with‐like comparison of time. Another practical consideration is that the method of pairwise comparison does not produce readily available diagnostic information.</p> <hd id="AN0138541329-6">Rationale for the present study</hd> <p>Aiming to capitalize on the reliability obtained by the method of pairwise comparison while increasing efficiency and the availability of diagnostic information, Heldsinger and Humphry ([<reflink idref="bib21" id="ref55">21</reflink>]) introduced an alternative two‐staged method of assessment. In stage 1 of the method, a relatively large number of performances are calibrated by asking teachers to compare performances in pairs, each time selecting the performance that is of a higher quality, and then analyzing their judgments using the Bradley–Terry–Luce (BTL) model. After calibration, a qualitative analysis of the calibrated performances is used to derive empirically based descriptions of the features of development evident in the performances, and a subset of performances is selected to act as exemplars. In stage 2 of the method, teachers assess student performances, which are separate from those in stage 1, by comparing these separate performances to the calibrated exemplars and with the support of the performance descriptors. The same two‐staged method is adopted in the present study.</p> <hd id="AN0138541329-7">Data Analysis</hd> <p>In stage 1 of this study, the data were analyzed using the BTL model, which is based on case 5 of Thurstone's ([<reflink idref="bib37" id="ref56">37</reflink>]) original model, which he referred to as the "law" of comparative judgment in relation to the original psychophysical context. Thurstone's model for pairwise comparison is in some respects a predecessor of the Rasch model and other IRT models (Andrich, [<reflink idref="bib2" id="ref57">2</reflink>]; Bock, [<reflink idref="bib8" id="ref58">8</reflink>]). In applying Thurstone's model for pairwise comparisons, the cumulative normal distribution was originally used. Substituting the numerically near‐equivalent logistic function for the cumulative normal gives the BTL model.</p> <p>The BTL model is identical to the conditional form of the Rasch model that is obtained when the item parameter has been eliminated. The BTL model is</p> <p> <ephtml> <math display="block" altimg="urn:x-wiley:00220655:media:jedm12223:jedm12223-math-0001" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>p</mi><mi>ij</mi></msub><mo>=</mo><mo form="prefix">Pr</mo><mrow><mo>{</mo><mi>i</mi><mo>></mo><mi>j</mi><mo>}</mo></mrow><mo>=</mo><mfrac><mrow><mo form="prefix">exp</mo><mspace width="0.16em" /><mo>(</mo><msub><mi>θ</mi><mi>i</mi></msub><mo>−</mo><msub><mi>θ</mi><mi>j</mi></msub><mo>)</mo></mrow><mrow><mn>1</mn><mo>+</mo><mo form="prefix">exp</mo><mspace width="0.16em" /><mo>(</mo><msub><mi>θ</mi><mi>i</mi></msub><mo>−</mo><msub><mi>θ</mi><mi>j</mi></msub><mo>)</mo></mrow></mfrac><mspace width="0.16em" /><mo>,</mo></mrow></math> </ephtml> </p> <p>where <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12223:jedm12223-math-0002" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>θ</mi><mi>i</mi></msub></math> </ephtml> and <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12223:jedm12223-math-0003" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>θ</mi><mi>j</mi></msub></math> </ephtml> are relative locations of persons <emph>i</emph> and <emph>j</emph> on the latent trait and <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12223:jedm12223-math-0004" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>p</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub></math> </ephtml> is the probability that performance <emph>i</emph> is judged better than performance <emph>j</emph>. To estimate model parameters from the data, custom software implementing maximum likelihood estimation of the person parameters was used. Fischer information and standard errors are obtained as a by‐product of the estimation procedure, which allows calculation of a person separation index defined as follows:</p> <p> <ephtml> <math display="block" altimg="urn:x-wiley:00220655:media:jedm12223:jedm12223-math-0005" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>S</mi><mo>=</mo><mfrac><mrow><mi> var </mi><mo>[</mo><mover accent="true"><mi>θ</mi><mo>̂</mo></mover><mo>]</mo><mo>−</mo><mi>MSE</mi></mrow><mrow><mi> var </mi><mo>[</mo><mover accent="true"><mi>θ</mi><mo>̂</mo></mover><mo>]</mo></mrow></mfrac><mo>,</mo></mrow></math> </ephtml> </p> <p>where <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12223:jedm12223-math-0006" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>MSE</mi><mo>=</mo><msub><mo>∑</mo><mi>n</mi></msub><msubsup><mi>σ</mi><mi>n</mi><mn>2</mn></msubsup></mrow></math> </ephtml> is the mean squared standard error of the estimates. The person separation index is directly analogous to Cronbach's <emph>α</emph> (Andrich, [<reflink idref="bib3" id="ref59">3</reflink>]) and indicates the level of internal consistency. This formula is the same as applied in Rasch software such as Winsteps and RUMM.</p> <hd id="AN0138541329-8">Method</hd> <p></p> <hd id="AN0138541329-9">Stage 1: Design and Development of the Persuasive Writing Assessment</hd> <p>The first stage of the study focused on the calibration of the exemplars to be used in the second stage.</p> <hd id="AN0138541329-10">Collection of persuasive performances</hd> <p>The first stage was conducted in five government and one private Western Australian primary schools in both metropolitan Perth (four) and regional Western Australian (two) areas. Due to the limited number of schools in this study, it was not feasible to employ a full stratified or other sampling design. Nevertheless, to obtain some heterogeneity, the schools were selected to reflect a mix of socioeconomic contexts as indicated by their values on the Index of Community Socio‐Educational Advantage (ICSEA; <emph>M</emph> = 1,070, <emph>SD</emph> = 75). This index (ICSEA) is used in the Australian National Assessment Program—Literacy and Numeracy (NAPLAN). The median for Australian schools is 1,000 and the standard deviation is 100 (Australian Curriculum Assessment and Reporting Authority, [<reflink idref="bib4" id="ref60">4</reflink>]).</p> <p>Participating teachers collected the persuasive performances in term 1 and samples were collected from 162 children (gender was not identified) in pre‐primary to Year 7. Children were aged from 5 to 12 years of age. The rationale for including this age range is it provides all primary school teachers with a set of exemplars relevant to all students in primary schools. NAPLAN Writing assessments are used for the same range of ages, indicating the viability of using a common writing prompt for all students in the range.</p> <p>The task was administered individually by classroom teachers. The administration was not strictly standardized; however, teachers were given administration instructions as a guide, which they could follow or adapt where appropriate. This included a guide for time allocation, persuasive topics, and instructions to read to the students. Teachers were provided with a range of persuasive topics to choose from such as the following: Social media is destroying friendships. To what extent do you agree with this statement? Write to convince the reader of your opinions. (<reflink idref="bib2" id="ref61">2</reflink>) Should we experiment on animals? Antibiotics, insulin, vaccines, heart bypass surgery were all tested on animals but is it ethical to use animals in experiments? What do you think? Write to convince the reader of your opinions. (<reflink idref="bib3" id="ref62">3</reflink>) Shops should be allowed to sell junk food. Doctors tell us that some food sold to us is bad for our health. Do you think shops should be stopped from selling these foods? Or do people have a right to choose what they want to eat? Write to convince the reader of your opinions. (<reflink idref="bib4" id="ref63">4</reflink>) Australia is a great place to live. To what extent do you agree with this statement? Write a letter explaining why Australia is or is not a great place to live. (<reflink idref="bib5" id="ref64">5</reflink>) Is it OK for the government to spy on people? Write to convince the reader of your opinions.</p> <hd id="AN0138541329-11">Calibration of the persuasive scale</hd> <p>Eighteen teachers from 12 metropolitan and regional schools, some of whom were involved in collecting samples, participated as judges in the study.  All were experienced classroom teachers.  Each received 30 minutes of training to make holistic judgments about students' persuasive writing skills, and use assessment and reporting software to make and record judgments. In making this judgment, they were asked to consider both authorial choices and conventions. Authorial choices include such aspects as subject matter, language choices, character and setting, and the reader–writer relationship. Conventions included aspects where the writer is expected to largely follow rules including spelling, punctuation, correct formation of sentences, and clarity of referencing. These criteria for assessment are aligned with the Australian Curriculum.</p> <p>Judges were given online access to specific pairs of persuasive performances to be compared. The pairs were generated randomly from the list of all pairs of performances. Judges worked individually and compared pairs of performances based on holistic judgments as to which performance displayed more advanced writing skill. A total of 3,228 pairwise comparisons were made by all judges combined as a basis for forming the scale. Pairs were generated so that each performance was compared with others on more than 30 occasions in order to ensure adequately low standard errors. The total number of pairs results from the target number per performance given the number of performances.</p> <hd id="AN0138541329-12">Pairwise comparison procedure</hd> <p>The locations of performances are inferred from the number of judgments in favor of each performance when compared with others using the BTL model. Performances are placed on the scale from weakest to strongest. If all performances were compared with each other, the strongest performance would be the one that compared better than others for the highest proportion of comparisons. However, given the application of scaling techniques there is not necessarily a strict monotonic correspondence between proportions and scale estimates.</p> <hd id="AN0138541329-13">Selection of exemplars and development of descriptors</hd> <p>The last part of stage 1 involved the development of an assessment tool that could be readily used by classroom teachers. This was comprised of three subcomponents.</p> <p>First, a descriptive qualitative analysis of all 162 persuasive performances on the scale was undertaken to examine the features of persuasive writing development as evidenced in the empirical data from student performances. Qualitative analysis examined both aspects of writing (authorial choices and conventions). This led to the development of empirically based performance descriptors.</p> <p>Second, from the original 162 performances, a subset of 17 performances was selected as exemplars. The aim was to select exemplars that captured reasonably typical features of writing development at given points on the scale. The descriptors were made available as an aid during assessment so that teachers could use both summary‐level descriptors and full exemplars.</p> <p>Lastly, a linear transformation was applied to the scale obtained from the analysis of pairwise data so that the exemplars had a range of approximately 160–480. This is simply so that the range is not interpreted by teachers as a percentage, and to avoid negative scale locations. The transformation results in a change of the arbitrary unit and origin of the interval scale obtained by application of the BTL model.</p> <hd id="AN0138541329-14">Stage 2: Assessment of Student Performances Against Calibrated Exemplars and Performance Desc...</hd> <p>Twenty‐five written persuasive performances were selected from the original 162 performances used in the pairwise comparisons, excluding any performances that had been selected as calibrated exemplars. The scale locations had a mean of.645 and standard deviation of 3.457. Performances were stratified into the lowest third, middle third, and highest third in terms of scale ranges and a selection was made of 8, 9, and 8, respectively, from these strata.</p> <p>These performances are referred to as the <emph>assessment sample</emph> performances. Eight judges, none of whom had participated in stage 1, compared performances in the assessment sample to the calibrated exemplars and the performance descriptors. The 17 exemplars and descriptors were displayed adjacent to a vertical scale in customized software for which the assessment display is shown in Figure . In this display, performances to be assessed appear on the right‐hand side and descriptors appear on the left‐hand side. Thumbnails of calibrated exemplars appear adjacent to the scale in the center. The descriptors also provide diagnostic information, as noted later.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/MEA/01sep19/jedm12223-fig-0001.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jedm12223-fig-0001.jpg" title="Stage 2 assessment display." /> </p> <p></p> <p>The judges were asked to make an on‐balance judgment, based on their analysis of the strengths and weaknesses of the sample, to determine which exemplar the sample was closest to or which two exemplars it fell between. The judges had the option to make one of four comparisons: the sample was exactly the same level (in terms of writing ability) as the exemplar, or it was slightly better than the exemplar, or it fell between both exemplars, or it was slightly weaker than the next exemplar. They scored each sample accordingly.</p> <p>The judges were provided with a guide to help make their judgments. This guide contains all the calibrated exemplars, the performance descriptors, and a close qualitative analysis of each exemplar. It was designed to help participants familiarize themselves with the exemplars and understand the particular features of each.</p> <hd id="AN0138541329-16">Results</hd> <p></p> <hd id="AN0138541329-17">Stage 1: Analysis of Pairwise Data</hd> <p>A total of 18 primary classroom teachers compared a total of 162 performances. A total of 3,228 comparisons were made. The Person Separation Index was.959, indicating a high level of internal consistency (Andrich, [<reflink idref="bib3" id="ref65">3</reflink>]; Heldsinger & Humphry, [<reflink idref="bib20" id="ref66">20</reflink>]).</p> <p>After applying the BTL scaling model, the original mean and <emph>SD</emph> for the scale locations were 0 and 4.091, respectively. A linear transformation was made such that the mean and <emph>SD</emph> were 355 and 85, respectively, to make the range more readily interpretable for classroom teachers by avoiding negative numbers and decimal places. A subset of 17 performances at equal intervals of 20 from 160 to 480 was selected as exemplar. Subsequently, 25 performances were selected for the assessment sample with a mean and standard deviation after transformation of 368.4 and 71.8, respectively.</p> <hd id="AN0138541329-18">Stage 2</hd> <p>Table  shows the location estimates of all 25 performances for Judges 1 to 8 (J1 to J8).</p> <p>Performance Estimates, Means, and Standard Deviations</p> <p> <ephtml> <table><thead><tr><th /><th align="center">J1</th><th align="center">J2</th><th align="center">J3</th><th align="center">J4</th><th align="center">J5</th><th align="center">J6</th><th align="center">J7</th><th align="center">J8</th><th align="center">Mean</th><th align="center"><italic>SD</italic></th></tr></thead><tbody><tr><td>1</td><td>420</td><td>450</td><td>300</td><td>365</td><td>350</td><td>315</td><td>330</td><td>395</td><td>365.6</td><td>52.5</td></tr><tr><td>2</td><td>420</td><td>430</td><td>430</td><td>390</td><td>440</td><td>350</td><td>470</td><td>390</td><td>415.0</td><td>37.0</td></tr><tr><td>3</td><td>465</td><td>500</td><td>450</td><td>430</td><td>440</td><td>400</td><td>450</td><td>460</td><td>449.4</td><td>28.8</td></tr><tr><td>4</td><td>440</td><td>490</td><td>430</td><td>455</td><td>480</td><td>450</td><td>480</td><td>440</td><td>458.1</td><td>22.4</td></tr><tr><td>5</td><td>445</td><td>440</td><td>420</td><td>430</td><td>280</td><td>325</td><td>420</td><td>400</td><td>395.0</td><td>59.9</td></tr><tr><td>6</td><td>430</td><td>460</td><td>410</td><td>420</td><td>410</td><td>425</td><td>430</td><td>430</td><td>426.9</td><td>15.8</td></tr><tr><td>7</td><td>430</td><td>500</td><td>400</td><td>430</td><td>420</td><td>430</td><td>440</td><td>410</td><td>432.5</td><td>30.1</td></tr><tr><td>8</td><td>435</td><td>440</td><td>320</td><td>390</td><td>320</td><td>420</td><td>420</td><td>405</td><td>393.8</td><td>48.2</td></tr><tr><td>9</td><td>390</td><td>430</td><td>320</td><td>380</td><td>180</td><td>385</td><td>430</td><td>400</td><td>364.4</td><td>82.1</td></tr><tr><td>10</td><td>350</td><td>300</td><td>385</td><td>345</td><td>320</td><td>285</td><td>295</td><td>360</td><td>330.0</td><td>35.5</td></tr><tr><td>11</td><td>365</td><td>370</td><td>290</td><td>310</td><td>260</td><td>250</td><td>300</td><td>350</td><td>311.9</td><td>46.0</td></tr><tr><td>12</td><td>330</td><td>300</td><td>250</td><td>325</td><td>260</td><td>310</td><td>320</td><td>390</td><td>310.6</td><td>43.6</td></tr><tr><td>13</td><td>340</td><td>330</td><td>205</td><td>285</td><td>220</td><td>295</td><td>270</td><td>260</td><td>275.6</td><td>47.7</td></tr><tr><td>14</td><td>380</td><td>320</td><td>320</td><td>315</td><td>250</td><td>315</td><td>290</td><td>380</td><td>321.3</td><td>43.2</td></tr><tr><td>15</td><td>335</td><td>320</td><td>270</td><td>295</td><td>260</td><td>280</td><td>290</td><td>340</td><td>298.8</td><td>29.9</td></tr><tr><td>16</td><td>395</td><td>400</td><td>290</td><td>430</td><td>280</td><td>320</td><td>350</td><td>430</td><td>361.9</td><td>60.5</td></tr><tr><td>17</td><td>355</td><td>300</td><td>300</td><td>315</td><td>280</td><td>335</td><td>280</td><td>360</td><td>315.6</td><td>31.4</td></tr><tr><td>18</td><td>190</td><td>180</td><td>170</td><td>180</td><td>180</td><td>170</td><td>210</td><td>180</td><td>182.5</td><td>12.8</td></tr><tr><td>19</td><td>180</td><td>240</td><td>180</td><td>200</td><td>180</td><td>170</td><td>200</td><td>160</td><td>188.8</td><td>24.7</td></tr><tr><td>20</td><td>330</td><td>270</td><td>280</td><td>465</td><td>500</td><td>495</td><td>490</td><td>410</td><td>405.0</td><td>98.2</td></tr><tr><td>21</td><td>325</td><td>260</td><td>255</td><td>255</td><td>240</td><td>185</td><td>275</td><td>200</td><td>249.4</td><td>43.5</td></tr><tr><td>22</td><td>230</td><td>280</td><td>265</td><td>210</td><td>220</td><td>245</td><td>310</td><td>330</td><td>261.3</td><td>43.2</td></tr><tr><td>23</td><td>335</td><td>450</td><td>290</td><td>260</td><td>180</td><td>230</td><td>260</td><td>285</td><td>286.3</td><td>80.2</td></tr><tr><td>24</td><td>310</td><td>250</td><td>240</td><td>275</td><td>220</td><td>230</td><td>270</td><td>340</td><td>266.9</td><td>41.1</td></tr><tr><td>25</td><td>340</td><td>270</td><td>270</td><td>335</td><td>420</td><td>320</td><td>310</td><td /><td>323.6</td><td>51.0</td></tr><tr><td><bold>Mean</bold></td><td>358.6</td><td>359.2</td><td>309.6</td><td>339.6</td><td>303.6</td><td>317.4</td><td>343.6</td><td>354.4</td><td /><td /></tr><tr><td><bold><italic>SD</italic></bold></td><td>75.2</td><td>94.2</td><td>79.3</td><td>82.2</td><td>100.9</td><td>88.0</td><td>86.9</td><td>82.0</td><td /><td /></tr><tr><td><bold>Mean of others</bold></td><td>332.5</td><td>332.4</td><td>339.5</td><td>335.2</td><td>340.3</td><td>338.4</td><td>334.6</td><td>333.1</td><td /><td /></tr></tbody></table> </ephtml> </p> <hd id="AN0138541329-19">Interrater reliability</hd> <p>Interrater reliability is here defined as the correlation between a pair of raters. The average interrater correlation between each of the 28 unique pairs of raters was.757, indicating a generally high level of interrater reliability, though with notable variability (see Table ). Despite a relatively small sample, all correlations are higher than expected by chance alone (<emph>p</emph> <.05 for all bivariate correlations between pairs of judges).</p> <p>Pairwise Intermarker Reliabilities and Average Intermarker Reliabilities by Judge</p> <p> <ephtml> <table><thead><tr><th>Judge</th><th align="center">J1</th><th align="center">J2</th><th align="center">J3</th><th align="center">J4</th><th align="center">J5</th><th align="center">J6</th><th align="center">J7</th><th align="center">J8</th><th align="center">Average</th></tr></thead><tbody><tr><td>J1</td><td /><td /><td /><td /><td /><td /><td /><td /><td>.872</td></tr><tr><td>J2</td><td>.8540003</td><td /><td /><td /><td /><td /><td /><td /><td>.749</td></tr><tr><td>J3</td><td>.8300003</td><td>.7850003</td><td /><td /><td /><td /><td /><td /><td>.841</td></tr><tr><td>J4</td><td>.8540003</td><td>.7070003</td><td>.7690003</td><td /><td /><td /><td /><td /><td>.939</td></tr><tr><td>J5</td><td>.5920003</td><td>.4450002</td><td>.6630003</td><td>.7750003</td><td /><td /><td /><td /><td>.735</td></tr><tr><td>J6</td><td>.7260003</td><td>.6170003</td><td>.6600003</td><td>.8980003</td><td>.7830003</td><td /><td /><td /><td>.875</td></tr><tr><td>J7</td><td>.7470003</td><td>.6940003</td><td>.7780003</td><td>.8980003</td><td>.7690003</td><td>.8980003</td><td /><td /><td>.907</td></tr><tr><td>J8</td><td>.8230003</td><td>.6930003</td><td>.7570003</td><td>.8710003</td><td>.6730003</td><td>.8300003</td><td>.7990003</td><td /><td>.869</td></tr></tbody></table> </ephtml> </p> <p>1 Mean pairwise intermarker reliability:.757.</p> <ulist> <item>2 <sups>*</sups>Correlation is significant at the.05 level (2‐tailed).</item> <item>3 <sups>**</sups>Correlation is significant at the.01 level (2‐tailed).</item> </ulist> <p>Table  shows the correlations between each rater's performance estimate and the average estimate of all other raters (rater‐average correlations). The raters' estimates were compared with the mean from all other judges by excluding the rater's own estimate from the mean in computing the judge–mean correlations. The mean correlation was.848 (range from.735 to.939) for persuasive writing, as shown in Table  in the column headed "Average," which again indicates a generally high level of interrater reliability.</p> <hd id="AN0138541329-20">Discussion</hd> <p>The aim of this study was to ascertain the reliability of classroom assessments using the two‐stage method described in the article. In stage 1, the method of pairwise comparisons was used to calibrate a large sample of 162 persuasive performances and the high reliability obtained is similar to findings from earlier studies using the paired comparison process cited earlier in the article. The findings from stage 2 indicate that high interrater reliability is achieved when teachers use calibrated exemplars and performance descriptors to assess performances. In practice, once stage 1 has been completed, classroom teachers only need to use stage 2 to assess their students. The assessments in stage 2 required no specific training, other than time needed for teachers to become familiar with the exemplars and accompanying descriptors in the custom software.</p> <p>This research extends tests undertaken in earlier research. Specifically, similar levels of interrater reliability were obtained in stage 2 of this study as were obtained by Heldsinger and Humphry ([<reflink idref="bib21" id="ref67">21</reflink>]) in relation to the assessment of early childhood recount writing, using a paper version of the two‐stage method of assessment. The present study involves older students who wrote longer texts in the context of persuasive writing, which is probably more demanding than recount writing.</p> <p>The present study also builds upon the research reported by McGrane, Humphry, and Heldsinger ([<reflink idref="bib29" id="ref68">29</reflink>]) in which 21 teachers all assessed a common set of 18 persuasive performances. In that study, the same two‐stage method was employed except that teachers assessed performances against a set of exemplars only, without descriptors, and they employed two separate assessment criteria (authorial choices and conventions) rather than one. The mean rater‐average correlation reported for the assessment of persuasive writing in this previous study was.907 based on the combined assessment on the two criteria. In the present study the mean rater‐average correlation was.848. However, in the previous research only highly experienced raters participated in matching, whereas in the present study participants were classroom teachers with typical levels of experience in assessing writing. Also, in the previous study by McGrane et al. ([<reflink idref="bib29" id="ref69">29</reflink>]), a much greater number of calibrated exemplars was used (36 vs. 17 in the present study). It was noted in this previous research that teachers did not always agree with the ordering of the larger number of exemplars, and they reported difficulties in making assessments (despite the high levels of interrater reliability). Also relevant here, the previous research did not use descriptors, which provide summary teaching points for next‐steps and therefore more readily accommodate formative assessment.</p> <p>A significant number of schools have used stage 2 of the method outlined here in classrooms, indicating the practical viability of the method. The reliability studies were conducted by a subset of the schools already using the method within their schools. Qualitative feedback has been obtained from 15 schools about various aspects of usage, indicating both benefits and challenges; however, it is beyond the scope of the present article to report this feedback in any further detail.</p> <p>The present study provides empirical results in response to recommendations for research focusing on the reliability of classroom assessments by Brookhart ([<reflink idref="bib12" id="ref70">12</reflink>]), Harlen ([<reflink idref="bib18" id="ref71">18</reflink>]), and Johnson ([<reflink idref="bib23" id="ref72">23</reflink>]). In this context the teachers were not assessing their own students' work; however, the approach readily translates to such contexts, provided there are not biases involved with assessing students known to the teachers.</p> <p>The two‐staged method was developed as a more time‐effective process than pairwise comparisons and one which allows for the provision of diagnostic information about students' work. With respect to diagnostic information, after assessments have been completed the descriptors shown in Figure  indicate what students can and cannot achieve in a given range of the scale (i.e., within their zone of proximal development) with respect to the criteria for the original pairwise comparisons, with those criteria including conventions and authorial choices. Teachers can also use descriptors from the next higher range as teaching points with respect to the same performance criteria.</p> <p>With respect to time‐effectiveness, it took a maximum of approximately seven minutes per assessment for teachers familiar with the exemplars and descriptors to assess a performance, with 25 assessments made in an allotted 3 hours. Exact times were not recorded but teachers reported finishing the tasks inside the allotted time. This amount of time is similar to the time taken for assessment of similar performances in large‐scale programs, for which the authors have overseen training, with rates of approximately 10 performances per hour by assessors using a rubric, excluding training time in advance of assessments. It is superior to the initial median time of 16 minutes reported in McGrane et al. ([<reflink idref="bib29" id="ref73">29</reflink>]) and comparable to the overall median time of approximately six minutes in that study. (This previous study found that time per rating reduced over time as each rater made 346 assessments, presumably due to practice and familiarity).</p> <p>Using the two‐stage method as detailed requires familiarity with exemplars but little or no training is involved, with teachers simply made aware of the exemplars, descriptors, and task required of them. The more performances teachers assess using the calibrated exemplars and descriptors, the more familiar they would be expected to become with the exemplars.</p> <p>The findings are limited to the context of primary schools. Further research is needed to ascertain the levels of reliability of the two‐staged method applied in contexts such as secondary schools and tertiary institutions, where performances are longer. The maximum length of the performances in this study was four pages. Also, a limitation of the study is that the sample size in stage 2 is relatively small. Further studies would be useful to ascertain whether similar levels of reliability are found in larger samples and other contexts. Nevertheless, combined with the related previous research referred to in the article, the findings indicate that the method is reliable for the assessment of writing in primary school.</p> <p>Variability in the levels of rater reliability was noted. It would also be useful to examine whether variations in levels of reliability are associated with background factors of teachers, such as teaching experience, academic background, specialist English skills, and so forth.</p> <hd id="AN0138541329-21">Conclusion</hd> <p>There is increasing use of pairwise comparisons in educational contexts. Stage 1 of the study involved conducting pairwise comparisons of 162 persuasive performances for which there was a high level of internal consistency. This is similar to previous findings related to the application of pairwise comparisons for essays. In practice, stage 1 is only required once and thereafter teachers can assess using stage 2 of the method.</p> <p>In stage 2, to assess a common set of 25 performances a group of teachers compared separate performances from the original sample against the set of calibrated exemplars and referred to the empirically derived performance descriptors. Generally, a high level of interrater reliability was obtained. The results of this study and previous research indicate that the two‐stage method of assessment enables teachers to make reliable assessments of writing in a manner that also enables formative assessment. The method requires little training to obtain good reliability and support materials, including instructional videos, can be made for this purpose.</p> <p>The study is consistent with the recommendation by Harlen ([<reflink idref="bib18" id="ref74">18</reflink>]) and others to explore the effectiveness of different approaches as a basis for improving the reliability of judgments by classroom teachers. The results of this research suggest that the particular method of assessment that is available to classroom teachers may affect the levels of reliability obtained. Both stages of the method require little or no training to obtain high levels of consistency and reliability.</p> <p>The two‐stage method is designed to capitalize on the high levels of internal consistency in pairwise comparisons as well as reduce the time taken to conduct the comparisons. In the second stage, implicit comparisons are made in the assessment, but once a scale has been constructed the average time to assess a performance is likely reduced. In addition, teachers are directly involved in the assessment and have access to formative information from the empirically derived performance descriptors and the ordered performance exemplars. Together, this information enables teachers to focus on specific skills that students need to develop to move further up the scale. The approach makes it possible to provide students with summaries of skills and exemplars of work somewhat better than their own so they know where to focus on making improvements.</p> <p>The two‐stage method provides a clear and explicit basis for moderation of teacher judgments because all teachers refer to the same set of exemplars during assessments. Given that one of the principal concerns about classroom assessments is whether they are sufficiently reliable, the reliability of the two‐stage method makes it a promising approach for achieving the widely desired objective in education of valuing professional expertise noted by Johnson ([<reflink idref="bib23" id="ref75">23</reflink>]) and others.</p> <ref id="AN0138541329-22"> <title> References </title> <blist> <bibl id="bib1" idref="ref1" type="bt">1</bibl> <bibtext> Allal, L. (2013). Teachers' professional judgement in assessment: A cognitive act and a socially situated practice. Assessment in Education: Principles, Policy and Practice, 20 (1), 20 – 34. https://doi.org/10.1080/0969594X.2012.736364</bibtext> </blist> <blist> <bibl id="bib2" idref="ref9" type="bt">2</bibl> <bibtext> Andrich, D. (1978). Relationships between the Thurstone and Rasch approaches to item scaling. Applied Psychological Measurement, 2, 449 – 460. https://doi.org/10.1177/014662167800200319</bibtext> </blist> <blist> <bibl id="bib3" idref="ref59" type="bt">3</bibl> <bibtext> Andrich, D. (1988). Rasch models for measurement. Beverly Hills, CA : Sage.</bibtext> </blist> <blist> <bibl id="bib4" idref="ref60" type="bt">4</bibl> <bibtext> Australian Curriculum Assessment and Reporting Authority. (2017). Guide to understanding 2013 Index of Community Socio‐Educational Advantage (ICSEA) values. Retrieved from https://acaraweb.blob.core.windows.net/resources/Guide_to_understanding_2013_ICSEA_values.pdf</bibtext> </blist> <blist> <bibl id="bib5" idref="ref47" type="bt">5</bibl> <bibtext> Benton, T., & Elliot, G. (2016). The reliability of setting grade boundaries using comparative judgement. Research Papers in Education, 31 (3), 352 – 376. https://doi.org/10.1080/02671522.2015.1027723</bibtext> </blist> <blist> <bibl id="bib6" idref="ref40" type="bt">6</bibl> <bibtext> Bisson, M. J., Gilmore, C., Inglis, M., & Jones, I. (2016). Measuring conceptual understanding using comparative judgement. International Journal of Research in Undergraduate Mathematics Education, 2 (2), 141 – 164. https://doi.org/10.1007/s40753-016-0024-3</bibtext> </blist> <blist> <bibl id="bib7" idref="ref2" type="bt">7</bibl> <bibtext> Black, P., & Wiliam, D. (2010). Inside the black box: Raising standards through classroom assessment. Phi Delta Kappan, 92 (1), 81 – 90. https://doi.org/10.1177/003172171009200119</bibtext> </blist> <blist> <bibl id="bib8" idref="ref58" type="bt">8</bibl> <bibtext> Bock, R. D. (1997). A brief history of item theory response. Educational Measurement: Issues and Practice, 16 (4), 21 – 33. https://doi.org/10.1111/j.1745-3992.1997.tb00605</bibtext> </blist> <blist> <bibl id="bib9" idref="ref32" type="bt">9</bibl> <bibtext> Bradley, R. A., & Terry, M. E. (1952). Rank analysis of incomplete block designs, I. The method of paired comparisons. Biometrika, 39, 324 – 345. https://doi.org/10.2307/2334029</bibtext> </blist> <blist> <bibtext> Bramley, T., Bell, J. F., & Pollitt, A. (1998). Assessing changes in standards over time using Thurstone's paired comparisons. Education Research and Perspectives, 25 (2), 1 – 24.</bibtext> </blist> <blist> <bibtext> Bramley, T., & Vitello, S. (2019). The effect of adaptivity on the reliability coefficient in adaptive comparative judgement. Assessment in Education: Principles, Policy and Practice, 26 (1), 43 – 58.</bibtext> </blist> <blist> <bibtext> Brookhart, S. M. (2013). The use of teacher judgement for summative assessment in the USA. Assessment in Education: Principles, Policy and Practice, 20, 69 – 90. https://doi.org/10.1080/0969594X.2012.703170</bibtext> </blist> <blist> <bibtext> Cattelan, M. (2012). Models for paired comparison data: A review with emphasis on dependent data. Statistical Science, 27, 412 – 433. https://doi.org/10.1214/12-STS396</bibtext> </blist> <blist> <bibtext> Choi, C. C. (1999). Public examinations in Hong Kong. Assessment in Education: Principles, Policy and Practice, 6, 405 – 417. https://doi.org/10.1080/09695949992829</bibtext> </blist> <blist> <bibtext> Connolly, S., Klenowski, V., & Wyatt‐Smith, C. (2012). Moderation and consistency of teacher judgement: Teachers' views. British Educational Research Journal, 38, 593 – 614. https://doi.org/10.1080/01411926.2011.569006</bibtext> </blist> <blist> <bibtext> Du Four, R. (2007). Once upon a time: A tale of excellence in assessment. In D. Reeves (Ed.), Ahead of the curve: The power of assessment to transform teaching and learning (pp. 253 – 267). Bloomington, IN : Solution Tree Press.</bibtext> </blist> <blist> <bibtext> Harlen, W. (2004). A systematic review of the evidence of reliability and validity of assessment by teachers used for summative purposes. London, UK : EPPI‐Centre, Social Science Research Unit, Institute of Education, University of London.</bibtext> </blist> <blist> <bibtext> Harlen, W. (2005). Trusting teachers' judgement: research evidence of the reliability and validity of teachers' assessment used for summative purposes. Research Papers in Education, 20 (3), 245 – 270. https://doi.org/10.1080/0267152050019374</bibtext> </blist> <blist> <bibtext> Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research, 77 (1), 81 – 112. https://doi.org/10.3102/003465430298487</bibtext> </blist> <blist> <bibtext> Heldsinger, S. A., & Humphry, S. M. (2010). Using the method of pairwise comparison to obtain reliable teacher assessments. Australian Educational Researcher, 37 (2), 1 – 19. https://doi.org/10.1007/BF03216919</bibtext> </blist> <blist> <bibtext> Heldsinger, S. A., & Humphry, S. M. (2013). Using calibrated exemplars in the teacher‐assessment of writing: An empirical study. Educational Research, 55 (3), 219 – 235. https://doi.org/10.1080/00131881.2013.825159</bibtext> </blist> <blist> <bibtext> Humphry, S., Heldsinger, S., & Dawkins, S. (2017). A two‐stage assessment method for assessing oral language in early childhood. Australian Journal of Education, 61 (2), 124 – 140. https://doi.org/10.1177/0004944117712777</bibtext> </blist> <blist> <bibtext> Johnson, S. (2013). On the reliability of high‐stakes teacher assessment. Research Papers in Education, 28 (1), 91 – 105. https://doi.org/10.1080/02671522.2012.754229</bibtext> </blist> <blist> <bibtext> Jones, I., & Alcock, L. (2014). Peer assessment without assessment criteria. Studies in Higher Education, 39, 1774 – 1787. https://doi.org/10.1080/03075079.2013.821974</bibtext> </blist> <blist> <bibtext> Jones, I., & Inglis, M. (2015). The problem of assessing problem solving: Can comparative judgement help? Educational Studies in Mathematics, 89, 337 – 355. https://doi.org/10.1007/s10649-015-9607-1</bibtext> </blist> <blist> <bibtext> Jones, I., Inglis, M., Gilmore, C., & Hodgen, J. (2013). Measuring conceptual understanding: The case of fractions. In A. M. Lindmeier and A. Heinze (Eds.), Proceedings of the 37th Conference of the International Group for the Psychology of Mathematics Education (Vol. 3, pp. 113 – 120). Kiel, Germany : PME.</bibtext> </blist> <blist> <bibtext> Jones, I., Swan, M., & Pollitt, A. (2015). Assessing mathematical problem solving using comparative judgement. International Journal of Science and Mathematics Education, 13 (1), 151 – 177. https://doi.org/10.1007/s10763-013-9497-6</bibtext> </blist> <blist> <bibtext> Luce, R. D. (1959). Individual choice behaviours: A theoretical analysis. New York, NY : John Wiley.</bibtext> </blist> <blist> <bibtext> McGrane, J. A., Humphry, S. M., & Heldsinger, S. A. (2018). Applying a Thurstonian, two‐stage method in the standardized assessment of writing. Applied Measurement in Education, 31 (4), 297 – 311. https://doi.org/10.1080/08957347.2018.1495216</bibtext> </blist> <blist> <bibtext> McMahon, S., & Jones, I. (2015). A comparative judgement approach to teacher assessment. Assessment in Education: Principles, Policy and Practice, 22, 368 – 389. https://doi.org/10.1080/0969594X.2014.978839.</bibtext> </blist> <blist> <bibtext> Pollitt, A. (2012). The method of adaptive comparative judgement. Assessment in Education: Principles, Policy and Practice, 19, 281 – 300. https://doi.org/10.1080/0969594X.2012.665354</bibtext> </blist> <blist> <bibtext> Potter, T., Englund, L., Charbonneau, J., MacLean, M. T., Newell, J., & Roll, I. (2017). ComPAIR: A new online tool using adaptive comparative judgement to support learning with peer feedback. Teaching and Learning Inquiry, 5 (2), 89 – 113.</bibtext> </blist> <blist> <bibtext> Seery, N., Canty, D., & Phelan, P. (2012). The validity and value of peer assessment using adaptive comparative judgement in design driven practical education. International Journal of Technology and Design in Education, 22 (2), 205 – 226. https://doi.org/10.1007/s10798-011-9194-0</bibtext> </blist> <blist> <bibtext> Standards and Testing Agency. (2015). Government response: Commission on Assessment Without Levels. Retrieved from https://<ulink href="http://www.gov.uk/government/publications/commission-on-assessment-without-levels-government-response">www.gov.uk/government/publications/commission-on-assessment-without-levels-government-response</ulink></bibtext> </blist> <blist> <bibtext> Steedle, J. T., & Ferrara, S. (2016). Evaluating comparative judgement as an approach to essay scoring. Applied Measurement in Education, 29 (3), 211 – 223. https://doi.org/10.1080/08957347.2016.1171769</bibtext> </blist> <blist> <bibtext> Tarricone, P., & Newhouse, C. P. (2016). Using comparative judgement and online technologies in the assessment and measurement of creative performance and capability. International Journal of Educational Technology in Higher Education, 13 (1), 16. https://doi.org/10.1186/s41239-016-0018-x</bibtext> </blist> <blist> <bibtext> Thurstone, L. L. (1927). A law of comparative judgment. Psychological Review, 34 (4), 273 – 286. https://doi.org/10.1037/h0070288</bibtext> </blist> <blist> <bibtext> van Daal, T., Lesterhuis, M., Coertjens, L., Donche, V., & De Maeyer, S. (2016). Validity of comparative judgement to assess academic writing: Examining implications of its holistic character and building on a shared consensus. Assessment in Education: Principles, Policy and Practice, 26, 1 – 16. https://doi.org/10.1080/0969594X.2016.1253542</bibtext> </blist> <blist> <bibtext> Wiliam, D. (2001). Reliability, validity, and all that jazz. Education 3–13, 29 (3), 17 – 21. https://doi.org/10.1080/03004270185200311</bibtext> </blist> </ref> <aug> <p>By Stephen Mark Humphry and Sandy Heldsinger</p> <p>Reported by Author; Author</p> <p></p> <p>STEPHEN MARK HUMPHRY is Associate Professor, Graduate School of Education, The University of Western Australia, M428 35 Stirling Hwy, Perth, WA 6009, Australia;. His primary research interests include educational assessment and psychometric methods.</p> <p>SANDY HELDSINGER is Senior Research Associate, Graduate School of Education, The University of Western Australia, M428 35 Stirling Hwy, Perth, WA 6009, Australia;. Her primary research interests include educational assessment.</p> </aug> <nolink nlid="nl1" bibid="bib16" firstref="ref3"></nolink> <nolink nlid="nl2" bibid="bib19" firstref="ref4"></nolink> <nolink nlid="nl3" bibid="bib12" firstref="ref5"></nolink> <nolink nlid="nl4" bibid="bib17" firstref="ref6"></nolink> <nolink nlid="nl5" bibid="bib23" firstref="ref7"></nolink> <nolink nlid="nl6" bibid="bib35" firstref="ref8"></nolink> <nolink nlid="nl7" bibid="bib10" firstref="ref10"></nolink> <nolink nlid="nl8" bibid="bib21" firstref="ref11"></nolink> <nolink nlid="nl9" bibid="bib34" firstref="ref13"></nolink> <nolink nlid="nl10" bibid="bib18" firstref="ref16"></nolink> <nolink nlid="nl11" bibid="bib15" firstref="ref22"></nolink> <nolink nlid="nl12" bibid="bib30" firstref="ref24"></nolink> <nolink nlid="nl13" bibid="bib39" firstref="ref25"></nolink> <nolink nlid="nl14" bibid="bib14" firstref="ref27"></nolink> <nolink nlid="nl15" bibid="bib24" firstref="ref29"></nolink> <nolink nlid="nl16" bibid="bib26" firstref="ref30"></nolink> <nolink nlid="nl17" bibid="bib37" firstref="ref31"></nolink> <nolink nlid="nl18" bibid="bib28" firstref="ref33"></nolink> <nolink nlid="nl19" bibid="bib20" firstref="ref35"></nolink> <nolink nlid="nl20" bibid="bib31" firstref="ref36"></nolink> <nolink nlid="nl21" bibid="bib38" firstref="ref39"></nolink> <nolink nlid="nl22" bibid="bib25" firstref="ref41"></nolink> <nolink nlid="nl23" bibid="bib27" firstref="ref42"></nolink> <nolink nlid="nl24" bibid="bib32" firstref="ref43"></nolink> <nolink nlid="nl25" bibid="bib33" firstref="ref44"></nolink> <nolink nlid="nl26" bibid="bib36" firstref="ref45"></nolink> <nolink nlid="nl27" bibid="bib22" firstref="ref46"></nolink> <nolink nlid="nl28" bibid="bib11" firstref="ref50"></nolink> <nolink nlid="nl29" bibid="bib13" firstref="ref51"></nolink> <nolink nlid="nl30" bibid="bib29" firstref="ref68"></nolink>
Header DbId: eric
DbLabel: ERIC
An: EJ1227713
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: A Two-Stage Method for Classroom Assessments of Essay Writing
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Humphry%2C+Stephen+Mark%22">Humphry, Stephen Mark</searchLink><br /><searchLink fieldCode="AR" term="%22Heldsinger%2C+Sandy%22">Heldsinger, Sandy</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Journal+of+Educational+Measurement%22"><i>Journal of Educational Measurement</i></searchLink>. Fall 2019 56(3):505-520.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Wiley-Blackwell. 350 Main Street, Malden, MA 02148. Tel: 800-835-6770; Tel: 781-388-8598; Fax: 781-388-8232; e-mail: cs-journals@wiley.com; Web site: http://www.wiley.com/WileyCDA
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 16
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2019
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Evaluative
– Name: Audience
  Label: Education Level
  Group: Audnce
  Data: <searchLink fieldCode="EL" term="%22Elementary+Education%22">Elementary Education</searchLink>
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Essays%22">Essays</searchLink><br /><searchLink fieldCode="DE" term="%22Elementary+School+Students%22">Elementary School Students</searchLink><br /><searchLink fieldCode="DE" term="%22Writing+%28Composition%29%22">Writing (Composition)</searchLink><br /><searchLink fieldCode="DE" term="%22Writing+Evaluation%22">Writing Evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Elementary+School+Teachers%22">Elementary School Teachers</searchLink><br /><searchLink fieldCode="DE" term="%22Persuasive+Discourse%22">Persuasive Discourse</searchLink><br /><searchLink fieldCode="DE" term="%22Achievement+Rating%22">Achievement Rating</searchLink><br /><searchLink fieldCode="DE" term="%22Interrater+Reliability%22">Interrater Reliability</searchLink><br /><searchLink fieldCode="DE" term="%22Expertise%22">Expertise</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1111/jedm.12223
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0022-0655
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: To capitalize on professional expertise in educational assessment, it is desirable to develop and test methods of rater-mediated assessment that enable classroom teachers to make reliable and informative judgments. Accordingly, this article investigates the reliability of a two-stage method used by classroom teachers to assess primary school students' persuasive writing. Stage 1 involves pairwise comparisons and stage 2 involves rating against calibrated exemplars from stage 1 plus performance descriptors. A high level of interrater reliability among teachers was obtained. This is consistent with previous evidence that the two-stage method is a viable classroom assessment method without extensive training and moderation. Implications for assessment practices in education are discussed with a focus on the widely expressed desire to value professional expertise.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2019
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1227713
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1227713
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1111/jedm.12223
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 16
        StartPage: 505
    Subjects:
      – SubjectFull: Essays
        Type: general
      – SubjectFull: Elementary School Students
        Type: general
      – SubjectFull: Writing (Composition)
        Type: general
      – SubjectFull: Writing Evaluation
        Type: general
      – SubjectFull: Elementary School Teachers
        Type: general
      – SubjectFull: Persuasive Discourse
        Type: general
      – SubjectFull: Achievement Rating
        Type: general
      – SubjectFull: Interrater Reliability
        Type: general
      – SubjectFull: Expertise
        Type: general
    Titles:
      – TitleFull: A Two-Stage Method for Classroom Assessments of Essay Writing
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Humphry, Stephen Mark
      – PersonEntity:
          Name:
            NameFull: Heldsinger, Sandy
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2019
          Identifiers:
            – Type: issn-print
              Value: 0022-0655
          Numbering:
            – Type: volume
              Value: 56
            – Type: issue
              Value: 3
          Titles:
            – TitleFull: Journal of Educational Measurement
              Type: main
ResultId 1