Implementing Confidence Assessment in Low-Stakes, Formative Mathematics Assessments

Saved in:
Bibliographic Details
Title: Implementing Confidence Assessment in Low-Stakes, Formative Mathematics Assessments
Language: English
Authors: Foster, Colin (ORCID 0000-0003-1648-7485)
Source: International Journal of Science and Mathematics Education. Oct 2022 20(7):1411-1429.
Availability: Springer. Available from: Springer Nature. One New York Plaza, Suite 4600, New York, NY 10004. Tel: 800-777-4643; Tel: 212-460-1500; Fax: 212-460-1700; e-mail: customerservice@springernature.com; Web site: https://link.springer.com/
Peer Reviewed: Y
Page Count: 19
Publication Date: 2022
Document Type: Journal Articles
Reports - Research
Education Level: Secondary Education
Descriptors: Mathematics Tests, Mathematics Education, Secondary School Students, Meta Analysis, Scores, Student Attitudes, Comparative Analysis, Effect Size, Formative Evaluation, Learning Processes, Test Items, Scoring, Self Efficacy, Educational Strategies, Bayesian Statistics
DOI: 10.1007/s10763-021-10207-9
ISSN: 1571-0068
1573-1774
Abstract: Confidence assessment (CA) involves students stating alongside each of their answers a confidence rating (e.g. 0 low to 10 high) to express how certain they are that their answer is correct. Each student's score is calculated as the sum of the confidence ratings on the items that they answered correctly, minus the sum of the confidence ratings on the items that they answered incorrectly; this scoring system is designed to incentivize students to give truthful confidence ratings. Previous research found that secondary-school mathematics students readily understood the negative-marking feature of a CA instrument used during one lesson, and that they were generally positive about the CA approach. This paper reports on a quasi-experimental trial of CA in four secondary-school mathematics lessons (N = 475 students) across time periods ranging from 3 weeks up to one academic year, compared to business-as-usual controls. A meta-analysis of the effect sizes across the four schools gave an aggregated Cohen's d of -0.02 [95% CI -0.22, 0.19] and an overall Bayes Factor B[subscript 01] of 8.48. This indicated substantial evidence for the null hypothesis that there was no difference between the attainment gains of the intervention group and the control group, relative to the alternative hypothesis that the gains were different. I conclude that incorporating confidence assessment into low-stakes classroom mathematics formative assessments does not appear to be detrimental to students' attainment, and I suggest reasons why a clear positive outcome was not obtained.
Abstractor: As Provided
Entry Date: 2022
Accession Number: EJ1348640
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwEF4iWQ5wsx86gXUut-fUpJAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDL_MkxzfQ9706rE7qQIBEICBm5b9EZCkKm7I8yknqTqga-Ny759W5H1M4eDbhOkySYGM883qNEGl2kcrSXtQo-zqh41TQyjoJMlV1Ckykyb7MTnTD5hpqRnhwcWSgsS0t5SPUISubdxnUrJljkwv3pqXvYM6GXTQ-DzgUdBq_FqSHiLcDQ3afmovzmsz0nle8jg_l_pyEDUM12rZf8M86qrF9mOdR0jSA-akHEDt
Text:
  Availability: 1
  Value: <anid>AN0159160694;[3d0g]01oct.22;2022Sep20.02:39;v2.2.500</anid> <title id="AN0159160694-1">Implementing Confidence Assessment in Low-Stakes, Formative Mathematics Assessments </title> <p>Confidence assessment (CA) involves students stating alongside each of their answers a confidence rating (e.g. 0 low to 10 high) to express how certain they are that their answer is correct. Each student's score is calculated as the sum of the confidence ratings on the items that they answered correctly, minus the sum of the confidence ratings on the items that they answered incorrectly; this scoring system is designed to incentivize students to give truthful confidence ratings. Previous research found that secondary-school mathematics students readily understood the negative-marking feature of a CA instrument used during one lesson, and that they were generally positive about the CA approach. This paper reports on a quasi-experimental trial of CA in four secondary-school mathematics lessons (N = 475 students) across time periods ranging from 3 weeks up to one academic year, compared to business-as-usual controls. A meta-analysis of the effect sizes across the four schools gave an aggregated Cohen's d of –0.02 [95% CI –0.22, 0.19] and an overall Bayes Factor B<sub>01</sub> of 8.48. This indicated substantial evidence for the null hypothesis that there was no difference between the attainment gains of the intervention group and the control group, relative to the alternative hypothesis that the gains were different. I conclude that incorporating confidence assessment into low-stakes classroom mathematics formative assessments does not appear to be detrimental to students' attainment, and I suggest reasons why a clear positive outcome was not obtained.</p> <p>Keywords: Confidence assessment; Formative assessment; Low-stakes assessments; Mathematics education; School mathematics</p> <hd id="AN0159160694-2">Introduction</hd> <p> <emph>Confidence assessment</emph> (CA) is a pedagogical practice involving a modification to the usual ways of conducting low-stakes (i.e. not-for-credit) formative assessments during school mathematics lessons. Students are asked to state a confidence rating (e.g. 0 low to 10 high) alongside each of their answers to express how certain they are that each answer is correct (Foster, [<reflink idref="bib14" id="ref1">14</reflink>]). Each student's score is then calculated as the sum of the confidence ratings for the items that they answered correctly, minus the sum of the confidence ratings for the items that they answered incorrectly.[<reflink idref="bib1" id="ref2">1</reflink>] The purpose of this scoring system is to incentivize students to give confidence ratings that are as truthful as possible. In the long term, students cannot systematically 'game' CA scores by consistently over- or under-stating their true confidence levels, so CA provides the possibility of accessing students' genuine beliefs about their degree of confidence at the level of individual items on an assessment. Students' <emph>calibration</emph> refers to the correlation between their confidence rating and their mean facility (the mean number of questions that they answered correctly) (Fischhoff et al., [<reflink idref="bib13" id="ref3">13</reflink>]), and it might be expected that students using CA over an extended period of time would gradually become better calibrated.</p> <p>Previous research testing a simple CA instrument over the course of one mathematics lesson on the topic of directed (positive and negative) numbers (Foster, [<reflink idref="bib14" id="ref4">14</reflink>]) found that secondary-school students were generally well calibrated in this topic, giving higher confidence scores on average for items that they answered correctly. Most students also readily understood the negative-marking aspect of CA, and were positive about the CA approach, alleviating concerns that students would find such an unfamiliar approach alien and unacceptable. Since then, CA has gained prominence amongst teachers (e.g. Foster et al., [<reflink idref="bib17" id="ref5">17</reflink>]), and, anecdotally, at teacher conferences and on social media, mathematics teachers tend to express considerable enthusiasm for the idea of CA, and it has been described enthusiastically on teacher-oriented podcasts (e.g. Barton, [<reflink idref="bib2" id="ref6">2</reflink>]; see also Baker, [<reflink idref="bib1" id="ref7">1</reflink>]).</p> <p>However, the extent to which CA 'works', in the sense of improving students' mathematics attainment in <emph>summative</emph> assessments when it is used over an extended period of time in <emph>formative</emph> assessments, is not known. Here, by 'summative' I mean an assessment which is used primarily as a retrospective evaluation of what has been learned up to that point, whereas a 'formative' assessment is integral to the teaching process and is "used to make changes to what would have happened in the absence of such information" (Wiliam, [<reflink idref="bib47" id="ref8">47</reflink>], p. 284; Wiliam, [<reflink idref="bib48" id="ref9">48</reflink>]).</p> <p>Clearly, students' mathematics attainment, as measured by general summative assessments, will depend on myriad factors, but literature from the use of CA in higher education (e.g. Gardner-Medwin, [<reflink idref="bib18" id="ref10">18</reflink>], [<reflink idref="bib19" id="ref11">19</reflink>], [<reflink idref="bib20" id="ref12">20</reflink>], [<reflink idref="bib21" id="ref13">21</reflink>]; Gardner-Medwin & Gahan, [<reflink idref="bib23" id="ref14">23</reflink>]; Gardner-Medwin & Curtin, [<reflink idref="bib22" id="ref15">22</reflink>]) suggests that it could have potential in school mathematics. Helping school students to consider how sure they are of the answers that they give might encourage them to self-check and to develop higher levels of self-awareness, which could enable them to target areas of weakness more effectively, which could increase their overall mathematics attainment. It could also be that being better calibrated is a desirable educational goal in its own right, since "knowing what you know and what you do not know", and therefore what you do or do not need to look up or seek support with, is an essential part of becoming a more educated person.</p> <p>One of the advantages of CA is that it is easy to implement, as it does not require redesigning assessment instruments (Barton, [<reflink idref="bib2" id="ref16">2</reflink>]; Foster, [<reflink idref="bib14" id="ref17">14</reflink>], Foster et al., [<reflink idref="bib17" id="ref18">17</reflink>]). Any classroom formative assessment method in which students write their answers (on paper, or even on mini-whiteboards [see McCrea, [<reflink idref="bib33" id="ref19">33</reflink>]]) can easily be modified by asking the students to write a confidence rating from 0 (low) to 10 (high) alongside each answer to indicate how sure they are that they are correct. This suggests that it could be possible to conduct a 'naturalistic' trial of CA in schools that minimally interferes with schools' existing assessment processes. If successful, such an intervention would be a low/no-cost 'easy win' for schools to implement (see Wiliam, [<reflink idref="bib49" id="ref20">49</reflink>]). Such an approach to trialling increases ecological validity (Robson & McCartan, [<reflink idref="bib39" id="ref21">39</reflink>]) and respects schools' autonomy and teachers' professionalism (Newmark, [<reflink idref="bib37" id="ref22">37</reflink>]), since the researcher does not attempt to take over and impose new systems and materials from the outside, without reference to the context of the students and teacher, and this may also increase compliance. Additionally, since it is easy for schools to agree to participate in such a trial, a closer-to-random sample of schools could likely be obtained, with fewer refusals than with standard educational trials. CA would seem to provide an ideal opportunity to test out such a 'naturalistic' trial.</p> <hd id="AN0159160694-3">Confidence Assessment in Mathematics</hd> <p></p> <hd id="AN0159160694-4">Confidence of Response</hd> <p>There are many similar and overlapping constructs in the literature relating to confidence at the fine grain size of individual items on an assessment (see e.g. Clarkson et al., [<reflink idref="bib9" id="ref23">9</reflink>]; Dirkzwager, [<reflink idref="bib11" id="ref24">11</reflink>]; Marsh et al., [<reflink idref="bib32" id="ref25">32</reflink>]; Stankov et al., [<reflink idref="bib43" id="ref26">43</reflink>]). For the purposes of CA, a pupil's "confidence of response" may be defined as "how certain they are that the answer that they have just given is correct" (Foster, [<reflink idref="bib14" id="ref27">14</reflink>], p. 274), and this may be represented on a scale from 0 (completely uncertain; i.e., just guessing) to 10 (absolutely certain). Since students' scores are calculated by <emph>summing</emph> these ratings (positively for correct answers and negatively for incorrect answers), it may be reasonable to treat this as a linear scale. This is plausible if we suppose that students are seeking to maximise their total score, leading to, for a 10-item formative assessment, a range of possible scores from –100 to 100. It may seem over-optimistic to expect students to discriminate their confidence on a 10-point scale (Preston & Colman, [<reflink idref="bib38" id="ref28">38</reflink>]); however, a 0–10 scale is likely to be considerably easier for students to use to calculate and interpret their scores in the classroom, since it is easy to calculate that for <emph>n</emph> questions the highest possible obtainable score will be 10<emph>n</emph>, and this may be considered to be what their score is 'out of'.</p> <hd id="AN0159160694-5">Uses of Confidence Assessment</hd> <p>CA has been used in higher education, particularly in medicine and related disciplines where it is critically important to discourage guessing in life-and-death matters (Gardner-Medwin, [<reflink idref="bib18" id="ref29">18</reflink>], [<reflink idref="bib19" id="ref30">19</reflink>], [<reflink idref="bib20" id="ref31">20</reflink>], [<reflink idref="bib21" id="ref32">21</reflink>]; Schoendorfer & Emmett, [<reflink idref="bib41" id="ref33">41</reflink>]). University students are often found to be poorly calibrated and to tend towards overconfidence (Ehrlinger et al., [<reflink idref="bib12" id="ref34">12</reflink>]). However, with repeated use of CA over time, calibration tends to improve (Gardner-Medwin & Curtin, [<reflink idref="bib22" id="ref35">22</reflink>]). It has been found that CA, often implemented in a multiple-choice context, can encourage self-checking, self-explanation and higher-level reasoning (Gardner-Medwin & Curtin, [<reflink idref="bib22" id="ref36">22</reflink>]; Sparck et al., [<reflink idref="bib42" id="ref37">42</reflink>]), and improve test validity by reducing gender bias (Hassmén & Hunt, [<reflink idref="bib24" id="ref38">24</reflink>]). Consequently, it has been suggested (Foster, [<reflink idref="bib14" id="ref39">14</reflink>], [<reflink idref="bib15" id="ref40">15</reflink>]; Foster et al., [<reflink idref="bib17" id="ref41">17</reflink>]) that CA might have considerable potential benefits for formative assessment in the school mathematics classroom. Although secondary-school students are younger and more diverse than university students, particularly those studying medicine and related areas, previous research (Foster, [<reflink idref="bib14" id="ref42">14</reflink>]; Foster et al., [<reflink idref="bib17" id="ref43">17</reflink>]) showed that they are capable of making valid judgments about their levels of confidence in a confidence assessment. So, it seems plausible that repeated use of CA over a period of time could have similar effects with secondary students to those reported for university students (Gardner-Medwin & Curtin, [<reflink idref="bib22" id="ref44">22</reflink>]).</p> <p>The potential benefits of incorporating confidence assessment into low-stakes formative assessments in mathematics include:</p> <p></p> <ulist> <item> <emph>Discouraging guessing</emph>, which adds noise to formative assessments and trivialises the learning of the subject (Foster, [<reflink idref="bib15" id="ref45">15</reflink>]).</item> <p></p> <item> <emph>Reducing over-confidence</emph>, which may inhibit students, through complacency, from learning and improving, and may lead them to embed errors and misconceptions without correcting them, thus hampering their future development in the subject.</item> <p></p> <item> <emph>Reducing under-confidence</emph>, which may prevent students from gaining as much satisfaction from the subject as they otherwise would, perhaps leading to lower motivation and a disinclination to pursue mathematics beyond the compulsory school years. Under-confidence may also trap students within repetitive cycles of practising content that is already secure, thus keeping them from accessing more demanding learning material.</item> </ulist> <p>It is important to stress that the pedagogical aim of using CA is not simply to raise all students' confidence, which would be highly undesirable. The aim is to encourage <emph>appropriate</emph> levels of self-confidence and self-awareness to help students engage in more effective future learning, since, "for secure development of procedural fluency it is important not only that a pupil can obtain the correct answer in a reasonable amount of time but that they have an accurate sense of their reliability with the procedure" (Foster, [<reflink idref="bib14" id="ref46">14</reflink>], p. 272). This means that, in some circumstances, the intention of CA would be to <emph>reduce</emph> students' confidence, at least temporarily, to help them gain a clearer picture of their weaknesses and difficulties, with a view to more targeted and effective subsequent learning. Many authors have drawn attention to <emph>illusory superiority effects</emph>, such as the <emph>Dunning–Kruger effect</emph> (see Ehrlinger et al., [<reflink idref="bib12" id="ref47">12</reflink>]), in which someone with low ability at a task overestimates their ability because they 'do not know what they do not know'. Common quotes to this effect within popular culture include "It ain't what you don't know that gets you into trouble. It's what you know for sure that just ain't so" (Mark Twain). CA has a potential role to play in recalibrating students to a more accurate assessment of what they do and do not currently know.</p> <p>In theory, we would expect the same CA intervention to improve the calibration of both under-confident and over-confident students. In this way, it should address the fact that students have different personalities: some may be more optimistic and others more pessimistic in their general outlook. Whilst this may affect their scoring in the short term, the intention is that over time all students would become better calibrated.</p> <hd id="AN0159160694-6">The Hypercorrection Effect</hd> <p>Another potential benefit of CA would be harnessing the <emph>hypercorrection effect</emph>, which is the observation that errors made with high confidence are more easily corrected than are errors made with low confidence (Butterfield & Metcalfe, [<reflink idref="bib4" id="ref48">4</reflink>]). This effect is surprising, since it might be expected that when a student is incorrect but very sure of their incorrect answer this might indicate a firmly-held belief, which would be difficult to dislodge. However, the effect has been reliably demonstrated in studies involving people of varied ages and in a variety of contexts (e.g. Metcalfe & Finn, [<reflink idref="bib34" id="ref49">34</reflink>]; van Loon et al., [<reflink idref="bib45" id="ref50">45</reflink>]), including in an authentic mathematics learning situation (Foster et al., [<reflink idref="bib17" id="ref51">17</reflink>]). The hypercorrection effect has been attributed to several possible mechanisms, including the memorable nature of a shock or surprise (Butterfield & Metcalfe, [<reflink idref="bib5" id="ref52">5</reflink>]). To see benefits from the hypercorrection effect in the mathematics classroom, it may be necessary that students <emph>reflect</emph> on the fact that they are confident, which CA provides an opportunity for them to do.</p> <hd id="AN0159160694-7">Research Question</hd> <p>Foster ([<reflink idref="bib14" id="ref53">14</reflink>]) found that secondary school students were overwhelmingly positive about the CA process, although, as this study was conducted over a single lesson, this may be at least partly a result of novelty effects (see Clark, [<reflink idref="bib8" id="ref54">8</reflink>]), and it was not possible to test for any long-term benefits over, say, time periods of up to a school year. Consequently, the research question for this study was: <emph>Does repeated use of CA in formative assessments in school mathematics lessons over an extended time period result in improvements in students' mathematics attainment in summative assessments?</emph></p> <p>The answer could be 'no' if the use of CA is distracting for students or the teacher, or takes valuable time away from teaching input, if it makes students become unduly risk averse (Kahneman & Tversky, [<reflink idref="bib29" id="ref55">29</reflink>]), or if students find the CA process discouraging, because their CA score is lower than the simple total of their correct answers. Less confident students may be repeatedly reminded of their low confidence by the CA method, and this could impair their performance on the questions. It may also be that it is unreasonable to expect CA to have a measurable impact on a distant outcome, such as overall mathematics attainment in a summative assessment. The CA process could also be contrary to equity if, as is plausible, high-attaining students are more likely to be better judges of their level of knowledge, and therefore more able to benefit (Ben-Simon et al., [<reflink idref="bib3" id="ref56">3</reflink>]). It may also be that teacher professional development is necessary in order to see students' attainment improve measurably from CA. Alternatively, the hope is that CA could be 'low-hanging fruit', which can offer schools an 'easy win' to benefit their students at little or no opportunity cost (<emph>cf</emph> Wiliam, [<reflink idref="bib49" id="ref57">49</reflink>]).</p> <hd id="AN0159160694-8">Method</hd> <p>A 'naturalistic' quasi-experimental trial (Christensen et al., [<reflink idref="bib7" id="ref58">7</reflink>]) was used, in which summative assessment data was obtained from schools whilst they incorporated CA into their regular in-class formative assessments. The intervention was minimally invasive, and aimed to preserve most features of schools' typical classroom and assessment practices. Rather than providing schools with researcher-designed CA tools and instruments, or requesting that schools develop their own, schools were instead asked to continue to use whatever assessment systems they had currently in place, both for formative assessment and for summative assessment. The only requested change was a simple modification to schools' formative assessment procedures, which did not require the re-design of any of the materials, in which students were asked to write a confidence rating next to each of their answers as they completed them during low-stakes, in-class tests. This small change to normal practice could be easily accommodated within lessons, constituting a minimal intrusion into schools' existing routines. This had the advantage of being easy for schools to commit to doing, as well as preserving as much ecological validity (Robson & McCartan, [<reflink idref="bib39" id="ref59">39</reflink>]) as possible. All teachers contacted agreed that the intervention would be easy to implement in their schools. For example, one head of department commented, "Students won't need any extra time to write in a confidence level and it'll be interesting for the teachers."</p> <hd id="AN0159160694-9">Intervention</hd> <p>Schools were asked to identify groups of "parallel classes" (as determined by them) of similar age and attainment, and to assign one of these to a business-as-usual control, where students would continue to experience the usual formative assessment practices currently operating in the school, without any modification. Schools were asked to modify the experience of the other, parallel, groups in only one way: for their regular formative, low-stakes in-class tests, schools were asked to continue to use their existing assessments but to ask students to also write a confidence rating (0 low to 10 high) beside each of their answers to indicate how sure they were that their given answer was correct (Foster, [<reflink idref="bib14" id="ref60">14</reflink>]). On completing their formative assessments, students were asked to calculate their scores by adding up the confidence ratings for the questions that they answered correctly and subtracting from this the total of the confidence ratings for the questions that they answered incorrectly. The scoring process, and the students' involvement in this, was an important aspect of the intervention, in order to incentivize truthful confidence ratings, by rewarding accurate confidence and penalising students for being over- or under-confident in their answers (see Foster, [<reflink idref="bib14" id="ref61">14</reflink>]). It also maintained the 'low-stakes' aspect of existing classroom culture in relation to in-class assessments, in which the teacher does not formally collect in the tests, mark them and record the results.</p> <p>The outcome measure was to compare students' marks in their normal termly or half-termly summative assessments, before and after a period of time in which CA was used in the students' low-stakes in-class formative assessments. (None of the schools' summative assessments contained any CA, and these were not modified in any way for either condition in this study.) Because schools used different summative assessments (often produced in-house) and administered these at different frequencies and times during the school year, data points were not synchronised between schools, or directly comparable between schools. Schools provided data on each student from the most-recent summative assessment point before embarking on CA, and this (converted to a percentage) was used as the pre-test score. They also provided data from the most recent summative assessment point before conclusion of the trial period, and this (also converted to a percentage) was used as the post-test score. (All schools were intending to continue using CA beyond the timescale of the trial, so CA was still in use at the conclusion of the study.)</p> <p>The intention was to test whether using CA in formative assessments over an extended period of time (half a term or so) might raise students' general mathematics attainment, as measured by their normal (non-CA) summative assessments, by comparing students' pre-test to post-test gain scores between students in the two conditions (confidence assessment and business-as-usual control). The CA formative assessments were the intervention, but not the data sources; the pre- and post- test summative scores were the data sources, and these were ordinary, non-CA summative assessments.</p> <hd id="AN0159160694-10">Participants</hd> <p>Emails were sent to schools across several mathematics teacher networks in the UK, and an invitation was posted on <emph>Twitter</emph> and widely retweeted by several contacts with large teacher followings:</p> <p>Now recruiting schools to trial 'confidence assessment', where students put a rating next to their answers to say how sure they are that their answer is right. Lots of potential benefits – see https://tinyurl.com/ydb2lhjx Please email me if you could help.</p> <p>Following this, 55 schools contacted me, and I replied to each, enclosing a 1-page explanation of what I was requesting (available in full in Appendix A). All of these teachers replied positively, saying that they wanted to try out the technique, but not all were able to take part fully in the trial. The most commonly-stated reasons for this were:</p> <p></p> <ulist> <item> no suitable parallel classes: often classes of similarly-attaining students were available, but they were not taught by teachers deemed to be comparable (e.g. one was a mathematics specialist and the other was not) or the school was too small to have parallel classes;</item> <p></p> <item> anticipated or actual lack of continuity in some classes during the period of the trial (e.g. extended periods of teacher absence or teachers leaving the school);</item> <p></p> <item> the teacher who was taking a lead on CA leaving the school or obtaining an unforeseen promotion to a more senior role;</item> <p></p> <item> a desire to implement CA with <emph>entire</emph> year groups, and not a subset of a year group.</item> </ulist> <p>The last-mentioned reason (<reflink idref="bib4" id="ref62">4</reflink>) was the most common. Although schools appreciated that simple controlled experiments (i.e. A/B tests) would be the only way to directly measure the effectiveness of CA, some had ethical concerns about denying students an intervention that they felt sure would be beneficial, and a small number felt that "experimenting on children" was morally wrong and did not wish to "use children as guinea pigs". In some cases, these schools rolled out CA to all of their mathematics students in a particular year group. In other schools, the reason for wanting to implement with full year groups was more practical, relating to efficiency, equity and simplicity of provision, and communication of assessment systems and outcomes to parents.</p> <p>In the end, complete data was obtained from four co-educational secondary schools in England, who participated fully in the trial, and this involved a total of 475 students over time periods ranging from 3 weeks to an entire school year (see Table 1 for details of the participating schools). A total of 11 students were not present at both the immediately-preceding summative assessment (the pre-test) and the summative assessment following the trial (the posttest), and data from these students were excluded from the analysis (see Table 1). All test marks were converted to percentages for analysis.</p> <p>Table 1 Schools participating in the study</p> <p> <ephtml> <table frame="hsides" rules="groups"><thead><tr><th><p>School</p></th><th><p>School description</p></th><th><p>Year group<sup>1</sup></p></th><th><p>Duration of trial<sup>2</sup></p></th><th><p><italic>N</italic></p><p>control</p></th><th><p><italic>N</italic></p><p>intervention</p></th></tr></thead><tbody><tr><td><p>A</p></td><td><p>comprehensive secondary school</p></td><td><p>7</p></td><td><p>3 weeks</p></td><td><p>111</p><p>(+ 5 excluded)</p></td><td><p>55</p></td></tr><tr><td><p>B</p></td><td><p>independent day school</p></td><td><p>7</p></td><td><p>1 term</p></td><td><p>57</p><p>(+ 1 excluded)</p></td><td><p>77</p><p>(+ 3 excluded)</p></td></tr><tr><td><p>C</p></td><td><p>comprehensive secondary school</p></td><td><p>8</p></td><td><p>2 terms</p></td><td><p>59</p></td><td><p>28</p><p>(+ 1 excluded)</p></td></tr><tr><td><p>D</p></td><td><p>comprehensive secondary school</p></td><td><p>7</p></td><td><p>3 terms</p></td><td><p>56</p></td><td><p>32</p><p>(+ 1 excluded)</p></td></tr></tbody></table> </ephtml> </p> <p>(Students were excluded because they did not complete either the pre-test or the post-test.) <sups>1</sups>Year 7 is aged 11–12; Year 8 is aged 12–13 <sups>2</sups>There are 3 terms in a school year and each term lasts about 13 weeks in total</p> <hd id="AN0159160694-11">Analysis</hd> <p>Some schools provided data on considerably more participants in the business-as-usual control classes than in the intervention classes. This was because schools C and D were trialling CA with only one class each, but each school had two other classes in the same year group who were not trialling the CA, and it was convenient for the schools to provide data for all of the cohort together. In School A, two classes trialled CA, and control data was provided for a large number of students from all of the other parallel classes. Consequently, the opportunity presented itself to create a dataset matched by pre-test scores, in order to compare more accurately the gain scores for students beginning at a similar attainment level at pre-test.</p> <p>Consequently, a matched sample (Table 2) for each of schools A, C, and D was produced, using the <emph>MatchIt</emph> package in <emph>R</emph> (Ho et al., [<reflink idref="bib25" id="ref63">25</reflink>]) and using the "nearest-neighbour" method. (<emph>R</emph> code for all of the analyses described in this paper is available in Appendix B.) Matching was <emph>not</emph> used for school B, because there were fewer control students than intervention students. Distribution plots for raw and matched samples (Appendix C) show that the matching process worked well for schools A and C and satisfactorily for D.</p> <p>Table 2 Original and matched datasets. Empty lines under 'matched dataset' indicate where matching was not used. Intervention datasets were not altered</p> <p> <ephtml> <table frame="hsides" rules="groups"><thead><tr><th /><th /><th colspan="3"><p>Original dataset</p></th><th colspan="3"><p>Matched dataset</p></th></tr><tr><th><p>School</p></th><th><p>Condition</p></th><th><p><italic>N</italic></p></th><th><p>pre-test</p><p>Mean</p></th><th><p>pre-test</p><p><italic>SD</italic></p></th><th><p><italic>N</italic></p></th><th><p>pre-test</p><p>Mean</p></th><th><p>pre-test</p><p><italic>SD</italic></p></th></tr></thead><tbody><tr><td><p>A</p></td><td><p>control</p></td><td><p>111</p></td><td><p>61.17</p></td><td><p>18.60</p></td><td><p>55</p></td><td><p>59.91</p></td><td><p>17.97</p></td></tr><tr><td /><td><p>intervention</p></td><td><p>55</p></td><td><p>59.49</p></td><td><p>18.04</p></td><td /><td /><td /></tr><tr><td><p>B</p></td><td><p>control</p></td><td><p>57</p></td><td><p>76.35</p></td><td><p>14.45</p></td><td /><td /><td /></tr><tr><td /><td><p>intervention</p></td><td><p>77</p></td><td><p>74.35</p></td><td><p>14.33</p></td><td /><td /><td /></tr><tr><td><p>C</p></td><td><p>control</p></td><td><p>59</p></td><td><p>50.94</p></td><td><p>16.01</p></td><td><p>28</p></td><td><p>55.78</p></td><td><p>14.88</p></td></tr><tr><td /><td><p>intervention</p></td><td><p>28</p></td><td><p>56.69</p></td><td><p>15.25</p></td><td /><td /><td /></tr><tr><td><p>D</p></td><td><p>control</p></td><td><p>56</p></td><td><p>62.61</p></td><td><p>9.81</p></td><td><p>32</p></td><td><p>68.75</p></td><td><p>6.41</p></td></tr><tr><td /><td><p>intervention</p></td><td><p>32</p></td><td><p>75.78</p></td><td><p>11.22</p></td><td /><td /><td /></tr></tbody></table> </ephtml> </p> <p>Ethical approval was obtained from the University of Leicester Ethics and Integrity Committee, and the full and matched datasets are freely available at https://doi.org/10.6084/m9.figshare.15027966.v1.</p> <hd id="AN0159160694-12">Results</hd> <p>Pre-test and post-test scores were all converted to percentages, and gain scores were calculated for each student as (post-test score – pre-test score) (Table 3). Analysis was conducted for each school using Welch's independent samples <emph>t</emph> tests on matched datasets for schools A, C, and D, and the original dataset for school B, where matching was not carried out. In schools A and B, students improved more in the intervention condition than in the control condition, but in schools C and D the reverse was the case, as shown by the sign of the <emph>t</emph> values and Cohen's <emph>d</emph> effect sizes in Table 4. Separate <emph>t</emph> tests for each school revealed that none of the effect sizes was significantly different from zero, providing no evidence for any effect of the CA intervention.</p> <p>Table 3 Pre-, post- and gain scores for each school</p> <p> <ephtml> <table frame="hsides" rules="groups"><thead><tr><th><p>School</p></th><th><p>Condition</p></th><th><p><italic>N</italic></p></th><th><p>pre-test</p><p>Mean</p></th><th><p>pre-test</p><p><italic>SD</italic></p></th><th><p>post-test</p><p>Mean</p></th><th><p>post-test</p><p><italic>SD</italic></p></th><th><p>Gain</p><p>Mean</p></th><th><p>Gain</p><p><italic>SD</italic></p></th></tr></thead><tbody><tr><td><p>A</p></td><td><p>control</p></td><td><p>55</p></td><td><p>59.91</p></td><td><p>17.97</p></td><td><p>45.06</p></td><td><p>19.07</p></td><td><p>–17.15</p></td><td><p>11.33</p></td></tr><tr><td /><td><p>intervention</p></td><td><p>55</p></td><td><p>59.48</p></td><td><p>18.04</p></td><td><p>43.15</p></td><td><p>19.72</p></td><td><p>–16.33</p></td><td><p>10.99</p></td></tr><tr><td><p>B</p></td><td><p>control</p></td><td><p>57</p></td><td><p>76.35</p></td><td><p>14.45</p></td><td><p>76.18</p></td><td><p>13.55</p></td><td><p>–0.18</p></td><td><p>14.08</p></td></tr><tr><td /><td><p>intervention</p></td><td><p>77</p></td><td><p>74.35</p></td><td><p>14.33</p></td><td><p>74.60</p></td><td><p>15.83</p></td><td><p>0.25</p></td><td><p>12.04</p></td></tr><tr><td><p>C</p></td><td><p>control</p></td><td><p>28</p></td><td><p>55.78</p></td><td><p>14.88</p></td><td><p>50.15</p></td><td><p>12.70</p></td><td><p>–1.44</p></td><td><p>10.13</p></td></tr><tr><td /><td><p>intervention</p></td><td><p>28</p></td><td><p>56.69</p></td><td><p>15.25</p></td><td><p>53.98</p></td><td><p>12.89</p></td><td><p>–2.71</p></td><td><p>9.36</p></td></tr><tr><td><p>D</p></td><td><p>control</p></td><td><p>32</p></td><td><p>68.75</p></td><td><p>6.41</p></td><td><p>81.60</p></td><td><p>8.99</p></td><td><p>14.97</p></td><td><p>8.35</p></td></tr><tr><td /><td><p>intervention</p></td><td><p>32</p></td><td><p>75.78</p></td><td><p>11.22</p></td><td><p>89.06</p></td><td><p>6.22</p></td><td><p>13.28</p></td><td><p>10.75</p></td></tr></tbody></table> </ephtml> </p> <p>Table 4 Analysis of gain scores for each school. Note: A matched sample was not used for school B. Welch independent-sample <emph>t</emph> tests were used</p> <p> <ephtml> <table frame="hsides" rules="groups"><thead><tr><th><p>School</p></th><th><p><italic>N</italic></p><p>control</p></th><th><p><italic>N</italic></p><p>intervention</p></th><th><p><italic>t</italic></p></th><th><p><italic>df</italic></p></th><th><p><italic>p</italic></p></th><th><p><italic>d</italic></p></th><th colspan="2"><p>95% CI for <italic>d</italic></p></th><th><p>Bayes Factor</p><p><italic>B</italic><sub>01</sub></p></th></tr></thead><tbody><tr><td><p>A</p></td><td><p>55</p></td><td><p>55</p></td><td><p>0.384</p></td><td><p>107.90</p></td><td><p>.701</p></td><td><p>0.073</p></td><td><p>–0.301</p></td><td><p>0.447</p></td><td><p>4.63</p></td></tr><tr><td><p>B</p></td><td><p>57</p></td><td><p>77</p></td><td><p>0.182</p></td><td><p>109.41</p></td><td><p>.856</p></td><td><p>0.032</p></td><td><p>–0.310</p></td><td><p>0.375</p></td><td><p>5.26</p></td></tr><tr><td><p>C</p></td><td><p>28</p></td><td><p>28</p></td><td><p>–0.485</p></td><td><p>53.665</p></td><td><p>.630</p></td><td><p>–0.130</p></td><td><p>–0.653</p></td><td><p>0.395</p></td><td><p>3.36</p></td></tr><tr><td><p>D</p></td><td><p>32</p></td><td><p>32</p></td><td><p>–0.702</p></td><td><p>58.426</p></td><td><p>.486</p></td><td><p>–0.175</p></td><td><p>–0.666</p></td><td><p>0.316</p></td><td><p>3.18</p></td></tr></tbody></table> </ephtml> </p> <p>Bayesian <emph>t</emph> tests were also conducted on the gain scores, comparing the fit of the data under the null hypothesis (that the gain scores are equal under the CA and control conditions) and the alternative hypothesis (that the gain scores are <emph>not</emph> equal under the two conditions). A Bayes factor <emph>B</emph> indicates the relative strength of evidence for two hypotheses (Dienes, [<reflink idref="bib10" id="ref64">10</reflink>]; Lambert, [<reflink idref="bib30" id="ref65">30</reflink>]; Rouder et al., [<reflink idref="bib40" id="ref66">40</reflink>]), and the interpretation is that the data are <emph>B</emph> times as likely under the null hypothesis as they are under the alternative hypothesis. (For a recent example of a similar use of Bayes factors, including a brief explanation of their meaning and interpretation, please see Foster, [<reflink idref="bib16" id="ref67">16</reflink>].)</p> <p>Bayes factors <emph>B</emph><subs><emph>01</emph></subs> in favour of the null hypothesis of equal gains under the two conditions, relative to the alternative hypothesis of unequal gains, were calculated using the default settings in <emph>JASP</emph> (JASP Team, [<reflink idref="bib26" id="ref68">26</reflink>]), with a Cauchy prior width of.707 (see Table 4). All of the estimated Bayes factors fell in the 3–10 range described by Jeffreys ([<reflink idref="bib27" id="ref69">27</reflink>]) as providing "substantial" evidence for the null hypothesis of no difference between the classes doing CA and the control classes. Note that this is not the same as the <emph>inconclusive</emph> result obtained from <emph>p</emph> values greater than.05 in the frequentist <emph>t</emph> tests, where we cannot conclude that either group outperformed the other. The Bayesian result provides <emph>positive evidence of no difference</emph>, not merely lack of evidence of a difference (see Dienes, [<reflink idref="bib10" id="ref70">10</reflink>]; Lambert, [<reflink idref="bib30" id="ref71">30</reflink>]).</p> <p>Standardised effect sizes from each school (Table 4) were combined, using both frequentist and Bayesian meta-analysis. The reason for treating each school as a separate study was that each was using different summative assessment measures, and pooling the raw data would not have been meaningful. Frequentist meta-analysis of the four effect sizes, using the <emph>metafor</emph> package in <emph>R</emph> (Viechtbauer, [<reflink idref="bib46" id="ref72">46</reflink>]) and the random-effects model (<emph>k</emph> = 4; τ<sups>2</sups> estimator: restricted maximum likelihood), gave an overall effect size not significantly different from zero (<emph>d</emph> = –0.02 [95% CI –0.22, 0.19], <emph>p</emph> =.870 with <emph>Q</emph>(<reflink idref="bib3" id="ref73">3</reflink>) = 0.878, <emph>p</emph> =.831) (see Fig. 1). Note that, in the conventional interpretation of frequentist hypothesis testing, it is not possible to conclude from this that the intervention had no effect, only that it was not possible to detect any effect.</p> <p>Graph: Fig. 1. Forest plot of Cohen's d effect sizes for each school. Bars indicate 95% confidence intervals</p> <p>Consequently, Bayesian meta-analysis was conducted using the <emph>BayesFactor</emph> package in <emph>R</emph> (Morey et al., [<reflink idref="bib36" id="ref74">36</reflink>])<emph>.</emph> This gave an overall Bayes Factor <emph>B</emph><subs><emph>01</emph></subs> in favour of the null hypothesis of 8.48, meaning that there was "substantial" (Jeffreys, [<reflink idref="bib27" id="ref75">27</reflink>]) evidence for the null hypothesis that the intervention group improved the same amount as the control group, relative to the alternative hypothesis that it did not. Note again that this is <emph>not</emph> the same as an inconclusive result, where we cannot say whether one group did significantly better than the other, such as was obtained from the frequentist meta-analysis. The Bayesian meta-analysis indicates positive evidence that the two groups did equally well.</p> <p>Schools were not explicitly asked to comment on their experiences of using CA with some of their classes. However, a few schools did comment when providing the data, and this was invariably positive. For example, one school noted, "The classes enjoyed doing the confidence scores on the Low Stake Quizzes and we found it useful to identify categories of student in the group and so how to help them: Confident and correct, Confident but incorrect, Not confident but correct & Not confident and incorrect." Another stated that "[the students] seemed intrigued by the scoring system so I'm hopeful it may have an impact on their self-reflection and thus on the follow up consolidation work they do ... They are very keen to not get any negatives so may be having the desired effect." No schools reported any problems implementing CA with their classes.</p> <hd id="AN0159160694-13">Discussion</hd> <p>The research question for this study was: <emph>Does repeated use of CA in formative assessments in school mathematics lessons over an extended time period result in improvements in students</emph>' <emph>mathematics attainment in summative assessments?</emph> To answer this, a 'naturalistic' quasi-experimental trial of CA involving four secondary schools (<emph>N</emph> = 475 students) was conducted over time periods ranging from 3 weeks up to one academic year, with students experiencing CA compared to students in business-as-usual control groups. A frequentist meta-analysis of the effect sizes across the four schools (Cohen's <emph>d</emph> = –0.02 [95% CI –0.22, 0.19]) is consistent with positive or negative effect sizes of small or moderately important size, and so, by itself, is inconclusive (see Lortie-Forgues & Inglis, [<reflink idref="bib31" id="ref76">31</reflink>], for discussion of this phenomenon). However, a Bayesian meta-analysis (Bayes Factor <emph>B</emph><subs><emph>01</emph></subs> of 8.48) revealed "substantial" (Jeffreys, [<reflink idref="bib27" id="ref77">27</reflink>]) evidence for the null hypothesis that there was no difference between the attainment gain of the intervention group and that of the control group, relative to the alternative hypothesis that the gains were different. This means that we can conclude that CA was not detrimental to students' attainment, but neither was it beneficial.</p> <p>There were reasons to think that CA might have turned out to have been detrimental to students' mathematics learning, through distracting them or their teachers from more effective aspects of the lesson and taking time away from teaching, but we did not find any evidence for this opportunity cost (see Wiliam, [<reflink idref="bib49" id="ref78">49</reflink>]). Similarly, CA could have had an adverse effect if it had led to students becoming damagingly risk averse (Kahneman & Tversky, [<reflink idref="bib29" id="ref79">29</reflink>]) and thereby leaving questions unanswered and supplying a zero confidence rating for them. Additionally, the requirement to think about confidence scores could have disturbed students' flow when answering the questions and distracted them from their learning. It might also have been anticipated that the students could have found their confidence scores discouraging, as these are likely to be lower than a simple total of questions answered correctly, and this could have led to disengagement from learning, but again this does not appear to have happened to any measurable degree.</p> <p>There were also good reasons for supposing that CA might have had a beneficial effect on students' mathematics attainment – indeed, this was the intention motivating this programme of research, and is frequently expressed enthusiastically by teachers (e.g. Baker, [<reflink idref="bib1" id="ref80">1</reflink>]; Barton, [<reflink idref="bib2" id="ref81">2</reflink>]). However, again, this does not appear to have happened. It may be that the attainment measure was too distant from the intervention, and that it was too optimistic to expect that CA used, from time to time, in formative assessments in mathematics lessons would lead to a detectable effect in termly/half-termly summative assessments, when myriad other factors are likely to be important to students' overall attainment. It may also be that without a strong imperative from the senior leadership in the school backing CA there was a lack of focus and the teachers' agendas were dominated by other promoted strategies and approaches that were given a higher status within the school. This is potentially one of the disadvantages of the 'naturalistic' nature of this trial, in which minimal disturbance is made to the school system. In this scenario, we are unlikely to benefit from Hawthorne-like effects (Robson & McCartan, [<reflink idref="bib39" id="ref82">39</reflink>]) and the general enthusiasm deriving from a concentrated focus on a high-profile intervention trial.</p> <p>Another factor may be that some of the schools may have been implementing CA infrequently and without sufficiently careful attention. The constraints within the naturalistic nature of this study meant that there was no opportunity for monitoring the fidelity of the implementation in a process evaluation, and it may be that heads of department were overestimating the compliance with CA within their schools when reporting back to me. It is also conceivable that there was 'leakage' from the intervention condition to the control condition, with teachers talking about CA in the staffroom, leading to teachers of control classes also trying it. This seems unlikely, however, given that heads of department were clear about the purpose of the trial, had ownership of it, and believed that their colleagues were keen to support this.</p> <p>The research question for this study refers to "repeated use of CA in formative assessments" and "over an extended time period", and neither of these terms can be defined precisely, due to the naturalist nature of this trial. This means that it is possible that more intensive use of CA, such as every lesson, rather than once a week, might be needed in order to see a measurable improvement in attainment in summative assessments. Alternatively, or additionally, it may be that the "extended time period" necessary is greater than the 3 weeks up to one academic year used in this study, although an intervention which shows no effect even after a year is probably of little practical value to schools.</p> <p>A clear overriding factor may be that teacher professional development is likely to be necessary in order for students to benefit significantly from CA. We know that, in general, for teaching strategies to be implemented effectively teachers need time to think through the pedagogical rationale and discuss approaches that they will use in the classroom (Joubert & Sutherland, [<reflink idref="bib28" id="ref83">28</reflink>]; Timperley et al., [<reflink idref="bib44" id="ref84">44</reflink>]; Yoon et al., [<reflink idref="bib50" id="ref85">50</reflink>]). In the present study, professional development was completely absent, and the only guidance to schools was a single sheet of paper outlining the approach (the entirety of this is presented in Appendix B). It is highly plausible that important features of CA might consequently have been 'lost in translation' or that teachers might not have sufficiently 'bought in' to the strategy.</p> <hd id="AN0159160694-14">Conclusion</hd> <p>A 'naturalistic' quasi-experimental trial of CA use within regular, low-stakes mathematics formative assessments across four secondary schools (<emph>N</emph> = 475 students) over time periods ranging from 3 weeks up to one academic year was conducted. A Bayesian meta-analysis of the effect sizes revealed substantial evidence of no effect on students' overall mathematics attainment, meaning that CA was not detrimental to students' attainment, but neither was it beneficial. Like a number of other high-profile interventions (see Lortie-Forgues & Inglis, [<reflink idref="bib31" id="ref86">31</reflink>]), CA, at least in this simple format, does not appear to be a quick, easy win for schools. To see benefits of CA, a closer-to-intervention measure may be needed and/or a more comprehensive implementation package, involving at least some professional development for teachers to set out the rationale for the process and generate some commitment. The sample of schools for this study was self-selecting, and consequently at least the lead teacher at each school was likely to be enthusiastic about the idea of CA. This might have been expected to have led to a stronger effect than if CA were implemented across 'typical' schools, but it may be that, without some professional development, enthusiasm by itself is insufficient to lead to effective implementation. Alternatively, it may simply be the case that use of CA does not in fact raise student attainment, and further studies would be needed to investigate these different possibilities.</p> <p>A novel, 'naturalistic' minimal-intervention approach to trialling was employed in this study, with mixed success. Some schools found it easy to find 'parallel' classes and to set up alternative conditions for the students, and using schools' existing data from summative assessments was unproblematic and very low-cost. All schools found conducting such a trial minimally invasive and fully compatible with their normal working processes. However, some schools were unwilling or unable to offer the intervention to a subset of the students, and in some cases this was because of uneasiness over the quasi-experimental methodology, as has been reported previously (Meyer et al., [<reflink idref="bib35" id="ref87">35</reflink>]).</p> <p>It is important that the educational research literature reports null findings, to combat the file-drawer problem of bias in the literature (Chambers, [<reflink idref="bib6" id="ref88">6</reflink>]), and this would seem to be particularly relevant for pedagogical approaches such as CA that currently command increasing attention in teacher-oriented literature. In future research, I plan to collect further data from schools who are persisting with CA and examine whether students' <emph>calibration</emph> (in addition to attainment) improves over extended periods of CA use. I also plan to interview students and teachers in schools who are persisting with CA to try to understand their perspectives on the approach. I also intend to develop associated professional development materials to support use of CA and to test the effectiveness of these. The ease with which CA can be adopted in practice, with minimal disturbance to existing classroom routines, makes it seem an attractive option to teachers, and it addresses the important issue of students' confidence. However, how to make it effective in practice remains an unsolved problem.</p> <hd id="AN0159160694-15">Note</hd> <p>The full dataset is available at https://doi.org/10.6084/m9.figshare.15027966.v1.</p> <hd id="AN0159160694-16">A. Appendices</hd> <p></p> <hd id="AN0159160694-17">A: Instructions Sent to Schools</hd> <p></p> <hd id="AN0159160694-18">Confidence Assessment in Mathematics: School-Based Trial</hd> <p> <emph>Confidence Assessment</emph> </p> <p>The idea is for students to give a confidence rating on a scale of 0 to 10 (0 is just guessing; 10 is absolutely certain) alongside their answers in their school mathematics assessments. Then, when they mark it, instead of simply counting the number of correct answers, they calculate their mark as the sum of the confidence ratings for the questions they got right <bold><emph>minus</emph></bold> the sum of the confidence ratings for the questions they got wrong. It shouldn't be possible to 'game' this system, as it rewards accurate assessment of confidence and penalises both over-confidence and under-confidence.</p> <p>There are potentially four benefits of building this into formative assessments:</p> <p></p> <ulist> <item> It improves students' calibration—being more confident about the ones they get right and less confident about the ones they get wrong—so that they 'know what they know and what they don't know' better, enabling better metacognition, more targeted revision, and more secure future learning;</item> <p></p> <item> It promotes self-checking—students often correct their answer when asked how sure they are;</item> <p></p> <item> It discourages guessing, which adds 'noise' to formative assessments;</item> <p></p> <item> It capitalises on the 'hypercorrection effect'—if a student states that they are very sure about something, and then they discover that they are wrong, they remember the correct answer better.</item> </ulist> <p>These benefits are plausible, and there is small-scale and anecdotal evidence for them, but we don't know if/how these things will work out in practice over time in real schools.</p> <p> <emph>What I would like you to do</emph> </p> <p>I would like you to trial confidence assessment with some classes across a year group. The only way to get convincing evidence whether confidence assessment 'works' is to do it with some classes and not with others and compare their progress in their school assessments. For example, if your Year 8 cohort were set in two bands, could you have the classes in one band continue as normal and those in the other band try confidence assessment for a term, or longer? Then see how their scores on half-termly assessments, say, compare between the two?</p> <p>The great thing about confidence assessment is that it should be very easy to implement. You wouldn't need to alter your school-based assessments at all, and students in the 'intervention' half of the year would just be asked to write down a confidence rating beside each answer and then work out their confidence score themselves. There would be no additional tests or marking, and I would be happy to crunch the numbers for the comparison of the students' half-termly/termly marks in the 'intervention' and 'control' groups if you sent me an anonymised spreadsheet. (I wouldn't need data on the confidence assessments themselves.)</p> <p>Obviously, you would need to decide what is feasible in terms of the number of classes and which year group, and the number of assessment points would be controlled by what you do in your school. I don't want to try to impose constraints on these things, because I don't want to increase anybody's workload and I want to see how confidence assessment might work in the 'natural' setting of real schools.</p> <p>I am extremely grateful for any help with this that you can give. If the results from this are promising, I intend to apply for funding to do a larger-scale, more robust trial.</p> <p>I have written a couple of articles about confidence assessment, if you want to read more—available free at https://tinyurl.com/ydb2lhjx and https://tinyurl.com/y8njoz5z – and I spoke about it on the Mr Barton Maths podcast: https://tinyurl.com/ybpbb68v.</p> <p>Please let me know if you can help, and get back to me with any questions.</p> <hd id="AN0159160694-19">B: R Code Used in the Analysis</hd> <p></p> <hd id="AN0159160694-20">Creating the Matched Data Sets</hd> <p>install.packages("MatchIt")</p> <p>library(MatchIt)</p> <p>Data <- read.csv("Fulldata.csv", header = TRUE, sep=",")</p> <p>Data <- subset(Data, School=="A")</p> <p>m.out <- matchit(treat ~ pre, data = Data, method = "nearest")</p> <p>summary(m.out)</p> <p>head(m.out)</p> <p>plot(m.out, type = "hist")</p> <p>m.data <- match.data(m.out)</p> <p>write.csv(m.data, file = "matched.csv")</p> <hd id="AN0159160694-21">Frequentist Meta-Analysis</hd> <p>install.packages("metafor")</p> <p>library(metafor)</p> <p>Data <- read.csv("Metadata.csv", header = TRUE, sep=",")</p> <p>Data$effectsize <- as.numeric(as.character(Data$Cohend))</p> <p>Data$var <- as.numeric(as.character(Data$Var))</p> <p>res <- rma(yi=effectsize, vi=var, data=Data, slab=paste(School, Ntotal, sep=", "), method="REML")</p> <p>res</p> <p>forest(res, xlab="Effect size (d)")</p> <p>mtext(bquote(paste("Summary: (Q = ",</p> <p>.(formatC(res$QE, digits=2, format="f")), ", df = ",.(res$k - res$p),</p> <p>", p = ",.(formatC(res$QEp, digits=2, format="f")), "; ", I^2, " = ",</p> <p>.(formatC(res$I2, digits=1, format="f")), "%)")))</p> <hd id="AN0159160694-22">Bayesian Meta-Analysis</hd> <p>install.packages("BayesFactor")</p> <p>library(BayesFactor)</p> <p>bf <- meta.ttestBF(Data$t, Data$Ncontrol, Data$Ntreat, rscale=.7071)</p> <p>bf[<reflink idref="bib1" id="ref89">1</reflink>]</p> <hd id="AN0159160694-23">C: Raw and Matched Distributions</hd> <p>Graph</p> <ref id="AN0159160694-24"> <title> References </title> <blist> <bibl id="bib1" idref="ref2" type="bt">1</bibl> <bibtext> Baker J. Using low-stakes quizzes with confidence assessment. Mathematics Teaching. 2019; 268: 11-14</bibtext> </blist> <blist> <bibl id="bib2" idref="ref6" type="bt">2</bibl> <bibtext> Barton, C. (2019). Conference takeaways: ResearchEd Blackpool 2019. [Audio podcast episode]. Accessed August 4, 2021 from <ulink href="http://www.mrbartonmaths.com/blog/conference-takeaways-researched-blackpool-2019/">http://www.mrbartonmaths.com/blog/conference-takeaways-researched-blackpool-2019/</ulink>.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref56" type="bt">3</bibl> <bibtext> Ben-Simon A, Budescu DV, Nevo B. A comparative study of measures of partial knowledge in multiple-choice tests. Applied Psychological Measurement. 1997; 21; 1: 65-88. 10.1177/0146621697211006</bibtext> </blist> <blist> <bibl id="bib4" idref="ref48" type="bt">4</bibl> <bibtext> Butterfield B, Metcalfe J. Errors committed with high confidence are hypercorrected. Journal of Experimental Psychology: Learning, Memory, and Cognition. 2001; 27: 1491-1494. 10.1037/0278-7393.27.6.1491</bibtext> </blist> <blist> <bibl id="bib5" idref="ref52" type="bt">5</bibl> <bibtext> Butterfield B, Metcalfe J. The correction of errors committed with high confidence. Metacognition and Learning. 2006; 1; 1: 69-84. 10.1007/s11409-006-6894-z</bibtext> </blist> <blist> <bibl id="bib6" idref="ref88" type="bt">6</bibl> <bibtext> Chambers C. The seven deadly sins of psychology: A manifesto for reforming the culture of scientific practice. 2017; Princeton University Press. 10.1515/9781400884940</bibtext> </blist> <blist> <bibl id="bib7" idref="ref58" type="bt">7</bibl> <bibtext> Christensen LB, Johnson RB, Turner LA. Research methods, design, and analysis. 2015; Pearson</bibtext> </blist> <blist> <bibl id="bib8" idref="ref54" type="bt">8</bibl> <bibtext> Clark RE. Reconsidering research on learning from media. Review of Educational Research. 1983; 53; 4: 445-459. 10.3102/00346543053004445</bibtext> </blist> <blist> <bibl id="bib9" idref="ref23" type="bt">9</bibl> <bibtext> Clarkson, L. M. C, Love, Q. U, & Ntow, F. D. (2017). How confidence relates to mathematics achievement: A new framework. In A. Chronaki (Ed.), Mathematics Education and Life at Times of Crisis, Proceedings of the Ninth International Mathematics Education and Society Conference (Vol. 2, pp. 441-451). University of Thessaly Press. Retrieved from https://muep.mau.se/bitstream/handle/2043/24043/MES9_Proceedings_low_Volume2.pdf?sequence=2&isAllowed=y#page=103</bibtext> </blist> <blist> <bibtext> Dienes Z. Using Bayes to get the most out of non-significant results. Frontiers in Psychology. 2014; 5: 781. 10.3389/fpsyg.2014.00781</bibtext> </blist> <blist> <bibtext> Dirkzwager A. Multiple evaluation: A new testing paradigm that exorcizes guessing. International Journal of Testing. 2003; 3; 4: 333-352. 10.1207/S15327574IJT0304_3</bibtext> </blist> <blist> <bibtext> Ehrlinger J, Johnson K, Banner M, Dunning D, Kruger J. Why the unskilled are unaware: Further explorations of (absent) self-insight among the incompetent. Organizational Behavior and Human Decision Processes. 2008; 105; 1: 98-121. 10.1016/j.obhdp.2007.05.002</bibtext> </blist> <blist> <bibtext> Fischhoff B, Slovic P, Lichtenstein S. Knowing with certainty: The appropriateness of extreme confidence. Journal of Experimental Psychology. 1977; 3; 4: 552-564. 10.1037/0096-1523.3.4.552</bibtext> </blist> <blist> <bibtext> Foster, C. (2016). Confidence and competence with mathematical procedures. Educational Studies in Mathematics, 91(2), 271–288. https://doi.org/10.1007/s10649-015-9660-9</bibtext> </blist> <blist> <bibtext> Foster, C. (2017). The guessing game. Teach Secondary, 6(8), 85. https://<ulink href="http://www.foster77.co.uk/Foster,%20Teach%20Secondary,%20The%20guessing%20game.pdf">www.foster77.co.uk/Foster,%20Teach%20Secondary,%20The%20guessing%20game.pdf</ulink></bibtext> </blist> <blist> <bibtext> Foster, C. (2018). Developing mathematical fluency: Comparing exercises and rich tasks. Educational Studies in Mathematics, 97(2), 121–141. https://doi.org/10.1007/s10649-017-9788-x</bibtext> </blist> <blist> <bibtext> Foster, C, Woodhead, S, Barton, C, & Clark-Wilson, A. (2021). School students' confidence when answering diagnostic questions online. Educational Studies in Mathematics. https://doi.org/10.1007/s10649-021-10084-7</bibtext> </blist> <blist> <bibtext> Gardner-Medwin AR. Confidence assessment in the teaching of basic science. Research in Learning Technology. 1995; 3; 1: 80-85. 10.3402/rlt.v3i1.9597</bibtext> </blist> <blist> <bibtext> Gardner-Medwin AR. Updating with confidence: Do your students know what they don't know?. Healthcare Informatics. 1998; 4: 45-46</bibtext> </blist> <blist> <bibtext> Gardner-Medwin ARBryan C, Clegg K. Confidence-based marking: Towards deeper learning and better exams. Innovative assessment in higher education. 2006; Routledge: 141-149</bibtext> </blist> <blist> <bibtext> Gardner-Medwin TBryan C, Clegg K. Certainty-based marking: Stimulating thinking and improving objective tests. Innovative assessment in higher education: A handbook for academic practitioners. 20192; Routledge: 141-150. 10.4324/9780429506857-13</bibtext> </blist> <blist> <bibtext> Gardner-Medwin, A. R, & Curtin, N. A. (2007). Certainty-based marking (CBM) for reflective learning and proper knowledge assessment. Paper presented at the REAP International Online Conference on Assessment Design for Learner Responsibility. Accessed August 4, 2021 from https://ewds.strath.ac.uk/REAP/reap07/Portals/2/CSL/t2%20-%20great%20designs%20for%20assessment/raising%20students%20meta-cognition/Certainty_based_marking_for_reflective_learning_and_knowledge_assessment.pdf.</bibtext> </blist> <blist> <bibtext> Gardner-Medwin AR, Gahan MChristie J. Formative and summative confidence-based assessment. Proceedings of the 7th international computer-aided assessment conference. 2003; Loughborough University: 147-155</bibtext> </blist> <blist> <bibtext> Hassmén P, Hunt DP. Human self-assessment in multiple-choice testing. Journal of Educational Measurement. 1994; 31; 2: 149-160. 10.1111/j.1745-3984.1994.tb00440.x</bibtext> </blist> <blist> <bibtext> Ho DE, Imai K, King G, Stuart EA. MatchIt: Nonparametric preprocessing for parametric causal inference. Journal of Statistical Software. 2011; 42; 8: 1-28. 10.18637/jss.v042.i08</bibtext> </blist> <blist> <bibtext> JASP Team (2020). JASP (Version 0.12.2) [Computer software]. Accessed August 4, 2021.</bibtext> </blist> <blist> <bibtext> Jeffreys H. The theory of probability. 19613; Oxford University Press</bibtext> </blist> <blist> <bibtext> Joubert, M, & Sutherland, R. (2009). A perspective on the literature: CPD for teachers of mathematics. NCETM. Accessed August 4, 2021 from https://<ulink href="http://www.ncetm.org.uk/media/1y2dv0zx/ncetm-recme-final-report.pdf">www.ncetm.org.uk/media/1y2dv0zx/ncetm-recme-final-report.pdf</ulink>.</bibtext> </blist> <blist> <bibtext> Kahneman D, Tversky A. Choices, values, and frames. American Psychologist. 1984; 39; 4: 341-350. 10.1037/0003-066X.39.4.341</bibtext> </blist> <blist> <bibtext> Lambert B. A student's guide to Bayesian statistics. 2018; SAGE Publications Ltd.</bibtext> </blist> <blist> <bibtext> Lortie-Forgues H, Inglis M. Rigorous large-scale educational RCTs are often uninformative: Should we be concerned?. Educational Researcher. 2019; 48; 3: 158-166. 10.3102/0013189X19832850</bibtext> </blist> <blist> <bibtext> Marsh HW, Pekrun R, Parker PD, Murayama K, Guo J, Dicke T, Arens AK. The murky distinction between self-concept and self-efficacy: Beware of lurking jingle-jangle fallacies. Journal of Educational Psychology. 2019; 111; 2: 331-353. 10.1037/edu0000281</bibtext> </blist> <blist> <bibtext> McCrea E. Making every maths lesson count: Six principles to support great maths teaching. 2019; Crown House Publishing Limited</bibtext> </blist> <blist> <bibtext> Metcalfe J, Finn B. Hypercorrection of high confidence errors in children. Learning and Instruction. 2012; 22; 4: 253-261. 10.1016/j.learninstruc.2011.10.004</bibtext> </blist> <blist> <bibtext> Meyer MN, Heck PR, Holtzman GS, Anderson SM, Cai W, Watts DJ, Chabris CF. Objecting to experiments that compare two unobjectionable policies or treatments. Proceedings of the National Academy of Sciences. 2019; 116; 22: 10723-10728. 10.1073/pnas.1820701116</bibtext> </blist> <blist> <bibtext> Morey, R. D, Rouder, J. N, Jamil, T, & Morey, M. R. D. (2015). Package 'BayesFactor'. Accessed August 4, 2021 from ftp://192.218.129.11/pub/CRAN/web/packages/BayesFactor/BayesFactor.pdf.</bibtext> </blist> <blist> <bibtext> Newmark B. Why Teach?. 2019; John Catt Educational Ltd.</bibtext> </blist> <blist> <bibtext> Preston CC, Colman AM. Optimal number of response categories in rating scales: Reliability, validity, discriminating power, and respondent preferences. Acta Psychologica. 2000; 104; 1: 1-15. 10.1016/S0001-6918(99)00050-5</bibtext> </blist> <blist> <bibtext> Robson C, McCartan K. Real world research. 20154; John Wiley & Sons</bibtext> </blist> <blist> <bibtext> Rouder JN, Speckman PL, Sun D, Morey RD, Iverson G. Bayesian t tests for accepting and rejecting the null hypothesis. Psychonomic Bulletin & Review. 2009; 16; 2: 225-237. 10.3758/PBR.16.2.225</bibtext> </blist> <blist> <bibtext> Schoendorfer, N, & Emmett, D. (2012). Use of certainty-based marking in a second-year medical student cohort: A pilot study. Advances in Medical Education and Practice, 3, 139-143. https://doi.org/10.2147/AMEP.S35972.</bibtext> </blist> <blist> <bibtext> Sparck EM, Bjork EL, Bjork RA. On the learning benefits of confidence-weighted testing. Cognitive Research: Principles and Implications. 2016; 1; 1: 3. 10.1186/s41235-016-0003-x</bibtext> </blist> <blist> <bibtext> Stankov L, Lee J, Luo W, Hogan DJ. Confidence: A better predictor of academic achievement than self-efficacy, self-concept and anxiety?. Learning and Individual Differences. 2012; 22; 6: 747-758. 10.1016/j.lindif.2012.05.013</bibtext> </blist> <blist> <bibtext> Timperley, H, Wilson, A, Barrar, H, & Fung, I. (2008). Teacher professional learning and development (Vol. 18). International Academy of Education. Accessed August 4, 2021 from https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.524.3566&rep=rep1&type=pdf.</bibtext> </blist> <blist> <bibtext> van Loon MH, Dunlosky J, van Gog T, van Merriënboer JJ, de Bruin AB. Refutations in science texts lead to hypercorrection of misconceptions held with high confidence. Contemporary Educational Psychology. 2015; 42: 39-48. 10.1016/j.cedpsych.2015.04.003</bibtext> </blist> <blist> <bibtext> Viechtbauer W. Conducting meta-analyses in R with the metafor package. Journal of Statistical Software. 2010; 36; 3: 1-48. 10.18637/jss.v036.i03</bibtext> </blist> <blist> <bibtext> Wiliam D. Formative assessment: Getting the focus right. Educational Assessment. 2006; 11; 3-4: 283-289. 10.1080/10627197.2006.9652993</bibtext> </blist> <blist> <bibtext> Wiliam D. Embedded formative assessment: Strategies for classroom assessment that drives student engagement and learning. 20172; Solution Tree Press</bibtext> </blist> <blist> <bibtext> Wiliam D. Creating the schools our children need. 2018; Learning Sciences International</bibtext> </blist> <blist> <bibtext> Yoon, K. S, Duncan, T, Lee, S. W. Y, Scarloss, B, & Shapley, K. L. (2007). Reviewing the evidence on how teacher professional development affects student achievement. Issues & answers (REL 2007-No. 033). Regional Educational Laboratory Southwest (NJ1). Accessed August 4, 2021 from https://files.eric.ed.gov/fulltext/ED498548.pdf.</bibtext> </blist> </ref> <ref id="AN0159160694-25"> <title> Footnotes </title> <blist> <bibtext> Other scoring rules and tariff matrices are possible; for simplicity, this was the one used in this research.</bibtext> </blist> </ref> <aug> <p>By Colin Foster</p> <p>Reported by Author</p> </aug> <nolink nlid="nl1" bibid="bib14" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib13" firstref="ref3"></nolink> <nolink nlid="nl3" bibid="bib17" firstref="ref5"></nolink> <nolink nlid="nl4" bibid="bib47" firstref="ref8"></nolink> <nolink nlid="nl5" bibid="bib48" firstref="ref9"></nolink> <nolink nlid="nl6" bibid="bib18" firstref="ref10"></nolink> <nolink nlid="nl7" bibid="bib19" firstref="ref11"></nolink> <nolink nlid="nl8" bibid="bib20" firstref="ref12"></nolink> <nolink nlid="nl9" bibid="bib21" firstref="ref13"></nolink> <nolink nlid="nl10" bibid="bib23" firstref="ref14"></nolink> <nolink nlid="nl11" bibid="bib22" firstref="ref15"></nolink> <nolink nlid="nl12" bibid="bib33" firstref="ref19"></nolink> <nolink nlid="nl13" bibid="bib49" firstref="ref20"></nolink> <nolink nlid="nl14" bibid="bib39" firstref="ref21"></nolink> <nolink nlid="nl15" bibid="bib37" firstref="ref22"></nolink> <nolink nlid="nl16" bibid="bib11" firstref="ref24"></nolink> <nolink nlid="nl17" bibid="bib32" firstref="ref25"></nolink> <nolink nlid="nl18" bibid="bib43" firstref="ref26"></nolink> <nolink nlid="nl19" bibid="bib38" firstref="ref28"></nolink> <nolink nlid="nl20" bibid="bib41" firstref="ref33"></nolink> <nolink nlid="nl21" bibid="bib12" firstref="ref34"></nolink> <nolink nlid="nl22" bibid="bib42" firstref="ref37"></nolink> <nolink nlid="nl23" bibid="bib24" firstref="ref38"></nolink> <nolink nlid="nl24" bibid="bib15" firstref="ref40"></nolink> <nolink nlid="nl25" bibid="bib34" firstref="ref49"></nolink> <nolink nlid="nl26" bibid="bib45" firstref="ref50"></nolink> <nolink nlid="nl27" bibid="bib29" firstref="ref55"></nolink> <nolink nlid="nl28" bibid="bib25" firstref="ref63"></nolink> <nolink nlid="nl29" bibid="bib10" firstref="ref64"></nolink> <nolink nlid="nl30" bibid="bib30" firstref="ref65"></nolink> <nolink nlid="nl31" bibid="bib40" firstref="ref66"></nolink> <nolink nlid="nl32" bibid="bib16" firstref="ref67"></nolink> <nolink nlid="nl33" bibid="bib26" firstref="ref68"></nolink> <nolink nlid="nl34" bibid="bib27" firstref="ref69"></nolink> <nolink nlid="nl35" bibid="bib46" firstref="ref72"></nolink> <nolink nlid="nl36" bibid="bib36" firstref="ref74"></nolink> <nolink nlid="nl37" bibid="bib31" firstref="ref76"></nolink> <nolink nlid="nl38" bibid="bib28" firstref="ref83"></nolink> <nolink nlid="nl39" bibid="bib44" firstref="ref84"></nolink> <nolink nlid="nl40" bibid="bib50" firstref="ref85"></nolink> <nolink nlid="nl41" bibid="bib35" firstref="ref87"></nolink>
Header DbId: eric
DbLabel: ERIC
An: EJ1348640
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Implementing Confidence Assessment in Low-Stakes, Formative Mathematics Assessments
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Foster%2C+Colin%22">Foster, Colin</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0003-1648-7485">0000-0003-1648-7485</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22International+Journal+of+Science+and+Mathematics+Education%22"><i>International Journal of Science and Mathematics Education</i></searchLink>. Oct 2022 20(7):1411-1429.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Springer. Available from: Springer Nature. One New York Plaza, Suite 4600, New York, NY 10004. Tel: 800-777-4643; Tel: 212-460-1500; Fax: 212-460-1700; e-mail: customerservice@springernature.com; Web site: https://link.springer.com/
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 19
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2022
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Audience
  Label: Education Level
  Group: Audnce
  Data: <searchLink fieldCode="EL" term="%22Secondary+Education%22">Secondary Education</searchLink>
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Mathematics+Tests%22">Mathematics Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Mathematics+Education%22">Mathematics Education</searchLink><br /><searchLink fieldCode="DE" term="%22Secondary+School+Students%22">Secondary School Students</searchLink><br /><searchLink fieldCode="DE" term="%22Meta+Analysis%22">Meta Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Scores%22">Scores</searchLink><br /><searchLink fieldCode="DE" term="%22Student+Attitudes%22">Student Attitudes</searchLink><br /><searchLink fieldCode="DE" term="%22Comparative+Analysis%22">Comparative Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Effect+Size%22">Effect Size</searchLink><br /><searchLink fieldCode="DE" term="%22Formative+Evaluation%22">Formative Evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Learning+Processes%22">Learning Processes</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Items%22">Test Items</searchLink><br /><searchLink fieldCode="DE" term="%22Scoring%22">Scoring</searchLink><br /><searchLink fieldCode="DE" term="%22Self+Efficacy%22">Self Efficacy</searchLink><br /><searchLink fieldCode="DE" term="%22Educational+Strategies%22">Educational Strategies</searchLink><br /><searchLink fieldCode="DE" term="%22Bayesian+Statistics%22">Bayesian Statistics</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1007/s10763-021-10207-9
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 1571-0068<br />1573-1774
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Confidence assessment (CA) involves students stating alongside each of their answers a confidence rating (e.g. 0 low to 10 high) to express how certain they are that their answer is correct. Each student's score is calculated as the sum of the confidence ratings on the items that they answered correctly, minus the sum of the confidence ratings on the items that they answered incorrectly; this scoring system is designed to incentivize students to give truthful confidence ratings. Previous research found that secondary-school mathematics students readily understood the negative-marking feature of a CA instrument used during one lesson, and that they were generally positive about the CA approach. This paper reports on a quasi-experimental trial of CA in four secondary-school mathematics lessons (N = 475 students) across time periods ranging from 3 weeks up to one academic year, compared to business-as-usual controls. A meta-analysis of the effect sizes across the four schools gave an aggregated Cohen's d of -0.02 [95% CI -0.22, 0.19] and an overall Bayes Factor B[subscript 01] of 8.48. This indicated substantial evidence for the null hypothesis that there was no difference between the attainment gains of the intervention group and the control group, relative to the alternative hypothesis that the gains were different. I conclude that incorporating confidence assessment into low-stakes classroom mathematics formative assessments does not appear to be detrimental to students' attainment, and I suggest reasons why a clear positive outcome was not obtained.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2022
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1348640
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1348640
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1007/s10763-021-10207-9
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 19
        StartPage: 1411
    Subjects:
      – SubjectFull: Mathematics Tests
        Type: general
      – SubjectFull: Mathematics Education
        Type: general
      – SubjectFull: Secondary School Students
        Type: general
      – SubjectFull: Meta Analysis
        Type: general
      – SubjectFull: Scores
        Type: general
      – SubjectFull: Student Attitudes
        Type: general
      – SubjectFull: Comparative Analysis
        Type: general
      – SubjectFull: Effect Size
        Type: general
      – SubjectFull: Formative Evaluation
        Type: general
      – SubjectFull: Learning Processes
        Type: general
      – SubjectFull: Test Items
        Type: general
      – SubjectFull: Scoring
        Type: general
      – SubjectFull: Self Efficacy
        Type: general
      – SubjectFull: Educational Strategies
        Type: general
      – SubjectFull: Bayesian Statistics
        Type: general
    Titles:
      – TitleFull: Implementing Confidence Assessment in Low-Stakes, Formative Mathematics Assessments
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Foster, Colin
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 10
              Type: published
              Y: 2022
          Identifiers:
            – Type: issn-print
              Value: 1571-0068
            – Type: issn-electronic
              Value: 1573-1774
          Numbering:
            – Type: volume
              Value: 20
            – Type: issue
              Value: 7
          Titles:
            – TitleFull: International Journal of Science and Mathematics Education
              Type: main
ResultId 1