Uncovering Key Statistical Concepts from an Investigation of Four Standardized Tests

Saved in:
Bibliographic Details
Title: Uncovering Key Statistical Concepts from an Investigation of Four Standardized Tests
Language: English
Authors: R.-M. Gibeau (ORCID 0000-0002-3649-4626), D. Cousineau
Source: Teaching Statistics: An International Journal for Teachers. 2025 47(3):219-234.
Availability: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 16
Publication Date: 2025
Document Type: Journal Articles
Reports - Research
Education Level: Higher Education
Postsecondary Education
Descriptors: Statistics Education, Student Evaluation, Standardized Tests, Mathematics Tests, Mathematical Concepts, Scores, Test Items, Accuracy, Teaching Methods, Learning Objectives, Evaluation Methods, Higher Education
DOI: 10.1111/test.12410
ISSN: 0141-982X
1467-9639
Abstract: To this date, few standardized tests measuring students' performance with regards to statistics exist. Only four tests have been proposed for college or university students. The goal of the present study is to investigate these tests. University professors or instructors experienced in teaching statistics were asked to list the concepts they think are being assessed by each item of the tests. A total of 708 responses were obtained from 42 participants. Using thematic analysis, 18 fundamental statistical concepts were identified. Unplanned analyses on participants' scores were also computed. The results suggest that there is no consensus on some of the items' correct answers. The study has practical implications for teaching statistics, from learning goals to assessment methods.
Abstractor: As Provided
Notes: https://osf.io/t6w7n
Entry Date: 2025
Accession Number: EJ1483737
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwFFVzd6Xfiy-j2keRdiLY-BAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDDVD5lupYgwCFFTKNgIBEICBm2KZriqGYXgsHcV0iFKOZOnFtmc0VWpBXi8J5lnmTu07OIoUbGyRHXY6w4Kv23bCxjwso1WHmjPOVeaTDYFjqIY_ubLuW7cT1IJzvyS6gmUPzSD7XUPH2zbB2K9JbG7vhzEtzQO_P43TAXFiv8cqNMaies2bn2_nLd38wrcdZ8huHtLvhB3Z_809CgP17Soc-y-YAXSGIkGd2Jvx
Text:
  Availability: 1
  Value: <anid>AN0188002603;d8y01sep.25;2025Sep18.05:53;v2.2.500</anid> <title id="AN0188002603-1">Uncovering key statistical concepts from an investigation of four standardized tests </title> <p>To this date, few standardized tests measuring students' performance with regards to statistics exist. Only four tests have been proposed for college or university students. The goal of the present study is to investigate these tests. University professors or instructors experienced in teaching statistics were asked to list the concepts they think are being assessed by each item of the tests. A total of 708 responses were obtained from 42 participants. Using thematic analysis, 18 fundamental statistical concepts were identified. Unplanned analyses on participants' scores were also computed. The results suggest that there is no consensus on some of the items' correct answers. The study has practical implications for teaching statistics, from learning goals to assessment methods.</p> <p>Keywords: assessment; standardized statistics tests; statistics education</p> <hd id="AN0188002603-2">INTRODUCTION</hd> <p>In this Big Data Era, statistical information is everywhere.[[<reflink idref="bib41" id="ref1">41</reflink>], [<reflink idref="bib57" id="ref2">57</reflink>]] This information is used to promote certain ideologies, justify decisions, encourage or discourage people to adopt a certain behavior, or simply to share events or conditions happening in the world.[[<reflink idref="bib51" id="ref3">51</reflink>]] As such, understanding the meaning and implications of this statistical information is a prerequisite for informed exchanges. However, it is commonly reported that statistics are among the most difficult subjects to teach and to learn.[[<reflink idref="bib14" id="ref4">14</reflink>], [<reflink idref="bib26" id="ref5">26</reflink>], [<reflink idref="bib40" id="ref6">40</reflink>]] Over the last few decades, many researchers in education and psychology have investigated why the teaching and learning of statistics are so challenging and how they can be improved.[[<reflink idref="bib5" id="ref7">5</reflink>], [<reflink idref="bib19" id="ref8">19</reflink>], [<reflink idref="bib24" id="ref9">24</reflink>], [<reflink idref="bib44" id="ref10">44</reflink>]]</p> <hd id="AN0188002603-3">Research avenues to improve statistics education</hd> <p>To circumvent these difficulties, some research examined the teaching approach and the teacher's characteristics, and how they influence students' learning.[[<reflink idref="bib29" id="ref11">29</reflink>], [<reflink idref="bib34" id="ref12">34</reflink>]] Their work suggests that the use of humor by instructors facilitates students' understanding, which in turn may improve their evaluation of the class and of the instructor.[[<reflink idref="bib3" id="ref13">3</reflink>], [<reflink idref="bib20" id="ref14">20</reflink>]] Others argued that manipulating real‐world data, doing demonstrations, and allowing students to manipulate data and probabilities help to develop statistical reasoning.[[<reflink idref="bib45" id="ref15">45</reflink>], [<reflink idref="bib54" id="ref16">54</reflink>], [<reflink idref="bib59" id="ref17">59</reflink>]]</p> <p>On the other hand, instead of alleviating the difficulties, other studies tried to identify the stumbling blocks which impede learning. Kahneman and Tversky[[<reflink idref="bib35" id="ref18">35</reflink>], [<reflink idref="bib55" id="ref19">55</reflink>]] and Konold[<reflink idref="bib37" id="ref20">37</reflink>] started this line of research by assuming that humans have heuristic modes of thinking (defined as simplistic ways of reasoning) and that these heuristics (reviewed in[<reflink idref="bib32" id="ref21">32</reflink>]) are keeping them from reasoning statistically, even after formal instruction. Kahneman and Tversky[<reflink idref="bib35" id="ref22">35</reflink>] examined people's ability to understand and manipulate probabilistic information. The authors concluded that heuristics are adopted when facing uncertainty. Furthermore, when researchers questioned students who correctly solved a problem, it was found that the correct reasoning did not underlie their correct response.[<reflink idref="bib38" id="ref23">38</reflink>]</p> <p>A third avenue of research, one that has been seldom taken, is to identify the core concepts underlying statistical understanding and which representations would best make them accessible to learners. Rather than alleviating learning difficulties or circumventing inadequate reasoning, this line of research takes the complementary approach of identifying what has to be conveyed to learners (i.e., the learning goals). Gigerenzer,[<reflink idref="bib28" id="ref24">28</reflink>] in particular, argued that probabilities (expressed with percentages) do not fit the learners' internal representations, which may be more aptly evoked with frequencies (expressed with counts). Consequently, identifying a list of the core concepts and how they can be distinguished from heuristic thinking would be a useful step in developing valid and discriminant assessment tools that go beyond mere memorization capacity and procedural skills. Indeed, it may be beneficial to invest efforts in developing assessments of conceptual understanding of statistics.</p> <hd id="AN0188002603-4">Procedural skills and conceptual understanding</hd> <p>The term <emph>procedural skills</emph> can be defined as "the ability to execute action sequences to solve problems"[<reflink idref="bib47" id="ref25">47</reflink>] (p. 346). In statistics, an example of procedural skills is the ability to calculate the mean given a set of values by applying a formula. To evaluate this type of learning, it is suggested to use typical problems for which students have step‐by‐step procedures to solve them. With technological advancements, specifically the existence of statistics software, it is argued that it is no longer necessary and even useful to teach students how to compute certain statistics by hand.[[<reflink idref="bib9" id="ref26">9</reflink>], [<reflink idref="bib12" id="ref27">12</reflink>], [<reflink idref="bib43" id="ref28">43</reflink>]] It is instead recommended to teach students how to use technology, how to choose the appropriate tool, how to interpret the outputs and how to communicate the results. The motivation to move from learning internalized procedural skills to learning technologically supported procedural skills, according to some authors, is to reduce cognitive effort to allow the learners to invest more cognitive capacity onto conceptual learning.</p> <p>A <emph>concept</emph> can be defined as how the understanding of something is internalized (typically a word or a symbol). It generally includes the characteristics of that thing as well as how it associates with other things.[<reflink idref="bib36" id="ref29">36</reflink>] As such, <emph>conceptual understanding</emph> can be defined as the understanding of the representation of similarities and differences (or characteristics) of elements and the association between them or how they influence each other.[<reflink idref="bib47" id="ref30">47</reflink>] For example, the concept <emph>mean</emph> is the representation of a response that is representative of a group of individual responses that are both similar and different from each other; a change in their individual values will change the value of the representative response and the visual representation of it. To evaluate this type of learning, it is suggested to use new and original problems for which students do not have a step‐by‐step procedure to solve them.[<reflink idref="bib47" id="ref31">47</reflink>] To complete these assessments, students must rely on conceptual knowledge to create an appropriate procedure that will solve the problem.</p> <hd id="AN0188002603-5">Evaluating students in statistics courses</hd> <p>According to Allen,[<reflink idref="bib2" id="ref32">2</reflink>] instructors too often assess students' memorization capacity based on how well they can remember formulas, rules, definitions or assumptions. However, focusing on students' conceptual understanding would influence their learning strategies and foster transferable knowledge.[[<reflink idref="bib11" id="ref33">11</reflink>], [<reflink idref="bib33" id="ref34">33</reflink>]] For this reason, some researchers argue that instructors do not assess students correctly.[[<reflink idref="bib12" id="ref35">12</reflink>], [<reflink idref="bib22" id="ref36">22</reflink>]] It can even be argued that the assessments used in introductory statistics courses are not valid for evaluating conceptual understanding. For example, anyone seeing the response choice "Correlation is not causation" will select it simply because it has been repeated frequently and memorized.</p> <p>On a related note, statistics anxiety questionnaires indicate that the anxiety of interpreting statistical results is a major component of the difficulty to understand statistics (an example item is "how anxious are you when interpreting the meaning of a table").[[<reflink idref="bib10" id="ref37">10</reflink>], [<reflink idref="bib27" id="ref38">27</reflink>], [<reflink idref="bib56" id="ref39">56</reflink>]] Assessing students in an appropriate way, avoiding rote learning and computation, could reduce anxiety which, according to some, is the most detrimental influencer of their performance.[[<reflink idref="bib5" id="ref40">5</reflink>], [<reflink idref="bib10" id="ref41">10</reflink>], [<reflink idref="bib14" id="ref42">14</reflink>], [<reflink idref="bib26" id="ref43">26</reflink>], [<reflink idref="bib31" id="ref44">31</reflink>], [<reflink idref="bib44" id="ref45">44</reflink>]]</p> <p>To assess students' understanding of statistics, it has been proposed to use specific items that would measure three basic learning goals: statistical literacy, statistical reasoning, and statistical thinking.[[<reflink idref="bib19" id="ref46">19</reflink>], [<reflink idref="bib30" id="ref47">30</reflink>]] The definition of statistical literacy varies in the literature. Aspects that often come up are the ability to read statistical information, critically evaluate it, form a reflection about such information and communicate this reflection.[<reflink idref="bib50" id="ref48">50</reflink>] To evaluate students' statistical literacy, instructors can use the terms <emph>identify</emph>, <emph>describe</emph>, <emph>rephrase</emph>, <emph>translate</emph>, <emph>interpret</emph>, and <emph>read</emph>.[<reflink idref="bib19" id="ref49">19</reflink>] Statistical reasoning is defined as the way people reason with statistical information, which involves interpreting results and graphs, and understanding the tools available.[<reflink idref="bib23" id="ref50">23</reflink>] To assess statistical reasoning, the terms <emph>why</emph>, <emph>how</emph> or <emph>explain the process</emph> can be used.[<reflink idref="bib19" id="ref51">19</reflink>] Finally, statistical thinking is defined as "what a statistician does"[<reflink idref="bib13" id="ref52">13</reflink>] (p. 4). In other words, it is the ability to explore data, to go beyond the teachings, being naturally curious about statistical problems and to see the process of doing statistics (analyses, interpretations, etc.) as a whole. The terms <emph>apply</emph>, <emph>criticize</emph>, <emph>evaluate</emph>, and <emph>generalize</emph> are thought to measure this learning goal.[<reflink idref="bib19" id="ref53">19</reflink>] It can be argued that the first and second learning goals are a way to develop students' conceptual understanding, while statistical thinking is a way to develop students' procedural skills.</p> <p>Other research, based on Garfield's[<reflink idref="bib24" id="ref54">24</reflink>] assessment framework, provided ways to evaluate students depending on the instructor's purpose for the assessment.[[<reflink idref="bib12" id="ref55">12</reflink>], [<reflink idref="bib26" id="ref56">26</reflink>]] Based on this work, instructors can choose between eight different ways to evaluate their students: homework, quizzes and exams, projects, critiques and writing assessments, lab reports, collaborative assessment tasks, minute papers, and student reflections. Additionally, instructors have three teaching strategies at their disposal according to Marson,[<reflink idref="bib39" id="ref57">39</reflink>] namely, immediate feedback, repetition, and original data (collected by learners). The author reported that immediate feedback helps students the most in their learning of statistics. This beneficial effect is presumably caused by students' awareness of their use of invalid heuristics. As interesting as these evaluation methods are, they do not concretely solve the problem of evaluating students' conceptual understanding instead of procedural skills.</p> <hd id="AN0188002603-6">Assessments in statistics education</hd> <p>As mentioned above, conceptual understanding includes statistical literacy and statistical reasoning. To our knowledge, there are four tests that claim to measure one of these two components, described in the Methods section in more detail: the <emph>Statistical Reasoning Assessment</emph> (SRA),[<reflink idref="bib25" id="ref58">25</reflink>] the <emph>Statistics Concept Inventory</emph> (SCI),[<reflink idref="bib2" id="ref59">2</reflink>] the <emph>Comprehensive Assessment of Outcomes for a First Course in Statistics</emph> (CAOS),[<reflink idref="bib18" id="ref60">18</reflink>] and the <emph>Statistical Reasoning in Biology Concept Inventory</emph> (SRBCI).[<reflink idref="bib17" id="ref61">17</reflink>] However, when examining the tests' items, the underlying concept each item is measuring is unspecified and often unclear; the developers do not clearly state what each item measures. They provide general topics or dimensions under which items were created without specifying the concepts measured. Some items, across all four tests, seem to measure students' rote learning and computational skills (procedural skills) rather than conceptual learning. Thus, they may not be valid instruments for measuring students' conceptual understanding of statistics.</p> <p>There are other instruments that have been developed to measure statistics learning.[[<reflink idref="bib49" id="ref62">49</reflink>], [<reflink idref="bib58" id="ref63">58</reflink>]] Some were developed for children[<reflink idref="bib58" id="ref64">58</reflink>] or include questions on research methods.[<reflink idref="bib49" id="ref65">49</reflink>] They often confuse performance with learning—because a student with a good performance does not necessarily understand the subject—and deviate very little from problems given in statistics classes. For exmaple Sabbag and colleagues[<reflink idref="bib49" id="ref66">49</reflink>] are focused on measuring statistical literacy and statistical reasoning and the relation between these two learning goals; it is therefore mute as to the concepts underpinning these goals.</p> <p>The aim of the present study is to investigate the following question: which concepts do statistics instructors think are being measured in standardized statistical instruments? An additional exploratory research question is: Do statistics instructors agree on these concepts and on the items' correct answer? In other words, the goal is to identify the underlying statistical concepts measured in these four tests through qualitative and quantitative data using statistics instructors since they are the ones teaching the subject and administering those tests. Their perspective is essential for understanding the tests mentioned above. Indeed, not being the developers of the tests and having pedagogical knowledge and experience in statistics teaching, make them the sample of choice. They understand how to construct evaluations based on their teaching and, inversely, what has been taught based on the evaluations. Their answers will clarify the content of those tests. As such, a sample of experienced statistics educators was invited to list which concepts are being tested in each item (qualitative) and optionally to answer the tests' items (quantitative). With their qualitative answers, the goal is to identify the building blocks of the statistical problems.</p> <hd id="AN0188002603-7">METHODS</hd> <p>Prior to data collection, this study underwent ethical review and was approved by the <emph>Ethics Research Board</emph> of the University of Ottawa.</p> <hd id="AN0188002603-8">Participants</hd> <p>In non‐math university programs, like psychology, it is the department's professors that are teaching the statistics courses and not statisticians or mathematicians. To increase the study's ecological validity, professors or instructors in psychology who have taught one or more statistics course(s) prior to the study were recruited by an invitation email sent through the mailing lists of the <emph>Canadian Society for Brain, Behavior and Cognitive Science</emph> (CSBBCS) and the <emph>Société Québécoise pour la Recherche en Psychologie</emph> (SQRP). In addition, all psychology professors at the authors' university and their past collaborators were contacted. Some reported in the comments that they have taught statistics classes many times; one also mentioned that he had just retired. In this invitation email, details about the aim of the study, what is required from participants, and the URL link to the online questionnaire were provided.</p> <p>The chosen population is better suited than any other population to participate in this study for two main reasons. Firstly, they are psychology professors, implying that they are frequent users of statistics for their own research projects. They likely appreciate the usefulness of such tools in research, the omnipresence of such information in daily life, and the necessity of understanding statistics and forming statistically qualified citizens. Secondly, they have taught one or more statistics course(s) to university students. This means that they have some pedagogical comprehension of the subject, they have experience with their field's curriculum and learning goals, and they are aware of concerns, fears and capabilities of students. These criteria ensure that participants are not statisticians and ensure that they are not influenced by the hindsight bias (the tendency of considering an information as being predictable and obvious after it was known[<reflink idref="bib1" id="ref67">1</reflink>]), in which case it would decrease the study's validity by focusing on experts instead of instructors.</p> <hd id="AN0188002603-9">Material</hd> <p>The four standardized tests listed above were assembled into one questionnaire in Qualtrics (a software used to create and distribute surveys). They are all multiple‐choice tests measuring statistical understanding. In total, there were 97 items. The tests are detailed here.</p> <hd id="AN0188002603-10">SRA</hd> <p>The <emph>Statistics Reasoning Assessment</emph> is composed of 20 multiple‐choice items (available in the validation article[<reflink idref="bib25" id="ref68">25</reflink>]). Most items are short scenarios with one or more correct answers. As an example, the second item is:</p> <p>The following message is printed on a bottle of prescription medication:</p> <p> <emph>WARNING: For applications to skin areas there is a 15% chance of developing a rash. If a rash develops, consult your physician</emph>.</p> <p>Which of the following is the best interpretation of this warning?</p> <p>The response choices are: "don't use the medication on your skin, there's a good chance of developing a rash," "for application to the skin, apply only 15% of the recommended dose," "if a rash develops, it will probably involve only 15% of the skin," "about 15 of 100 people who use this medication develop a rash," and "there is hardly a chance of getting a rash using this medication." The correct answer is the fourth response choice.</p> <p>The instrument was developed by Konold[<reflink idref="bib37" id="ref69">37</reflink>] for high school students and validated by Garfield.[<reflink idref="bib25" id="ref70">25</reflink>] It was developed in three steps: (i) experts reviewed the items for content validity, (ii) the items were administered to a group of students who answered open‐ended questions on the justification of their responses, which were used to form the SRA's response choices, and (iii) multiple pilot tests were run.</p> <p>The SRA aims to measure two dimensions, <emph>Correct Reasoning Skills</emph> and <emph>Misconceptions</emph>, which were divided into eight skills according to the authors. The first dimension includes: Correct interpretation of probabilities, understanding how to select the right average, correct computation of probabilities (further includes understanding probabilities as ratios and using combinatorial reasoning), understanding independence, understanding sampling variability, distinguishing between correlation and causation, correctly interpreting two‐way tables, and understanding the importance of large samples. The second dimension includes: Misconceptions involving averages (further includes the belief that averages are the most common number, the failure to consider outliers when computing the mean, group comparison using their averages, and confusion between mean and median), outcome orientation misconception, misconception about good samples needing a high percentage of the population, law of small numbers, representativeness misconception, correlation implies causation, equiprobability bias, and groups can only be compared if they have the same sample size. Notably, each skill is not measured by one or multiple questions, but by one or multiple response choice. For example, the correct interpretation of probabilities is measured by response choice <emph>d</emph> of item 2 and item 3.</p> <p>The internal consistency of the SRA has raised some issues and "yielded incomplete results"[<reflink idref="bib25" id="ref71">25</reflink>] (p. 31) but good test–retest reliability was observed for participants' correct responses score (<emph>r</emph> = 0.70) and incorrect responses score (<emph>r</emph> = 0.75). Criterion‐related validity was also examined by asking students enrolled in an introductory statistics course to complete the test. Students' scores on the SRA were then correlated with their grade on assignments. The author of the validation article mentioned that the correlations were low, without providing the coefficients.</p> <hd id="AN0188002603-11">SCI</hd> <p>The <emph>Statistics Concept Inventory</emph> has 25 multiple‐choice items. They are short scenarios with at least one correct answer. The SCI was developed as part of a doctoral dissertation[<reflink idref="bib2" id="ref72">2</reflink>] (the items are also available in the dissertation). Most participants were engineering students, some were psychology students, and a small number of participants were in other programs. As an example, the first item is:</p> <p>A certain diet plan claims that subjects lose an average of 20 pounds in 6 months on their plan. A dietician wishes to test this claim and recruits 15 people to participate in an experiment. Their weight is measured before and after the 6‐month period. Which is the appropriate test statistic to test the diet company's claim?</p> <p>The response choices are: "two‐sample <emph>Z</emph> test," "paired comparison <emph>t</emph> test," and "two‐sample <emph>t</emph> test." The correct answer is the second response choice.</p> <p>The items and the response choices were created based on textbooks, journal articles, and personal experiences. Also, the items are divided into four topics: inferential statistics, probability, descriptive statistics, and graphs. Through a process of focus groups, analysis of correct and incorrect answers, and expert opinions, the author revised the items of the questionnaire. The SCI was tested and revised six times from Fall 2002 to Spring 2006 with engineering and mathematics students. To keep the present text as short as possible, only the initial and final versions are discussed below.</p> <p>The author reported the alpha and omega reliability indices with the initial 38‐item SCI (Cronbach's <emph>α</emph> = 0.77 and McDonald's <emph>Ω</emph> = 0.89[[<reflink idref="bib4" id="ref73">4</reflink>], [<reflink idref="bib16" id="ref74">16</reflink>]]). An exploratory factor analysis suggested three solutions: a five‐factor model with an oblique rotation where the loadings vary between 0.30 and 0.66 with one cross‐loading; a five‐factor model with an orthogonal rotation where the loadings vary between 0.30 and 0.64 with four cross‐loadings; and a four‐factor model without rotation where the loadings vary between 0.30 and 0.59 with one cross‐loading. A total of 13 items were removed because previous research suggested a smaller number of factors compared to the first version of the SCI, which questioned face validity. After removing these 13 items, resulting in a 25‐item test, the author did a second confirmatory factor analysis which supported a one‐factor model (GFI = 0.92, PGFI = 0.84, RMSEA = 0.03). Next, the reliability of the 25‐item SCI was examined (<emph>α</emph> = 0.76). Finally, content validity was assessed through student interviews and faculty surveys. The author then claimed that this modified SCI would measure 14 important statistical concepts.</p> <hd id="AN0188002603-12">CAOS</hd> <p>The <emph>Comprehensive Assessment of Outcomes for a First Course in Statistics</emph> has 40 multiple‐choice items. They are short scenarios with at least one correct answer. This instrument was developed as part of the <emph>Assessment Resource Tools for Improving Statistical Thinking</emph> (ARTIST) project, which aimed at providing items and instruments for statistics education.[<reflink idref="bib18" id="ref75">18</reflink>] It was developed for post‐secondary students enrolled in an introductory statistics course. The items are available upon request to the first author of.[<reflink idref="bib18" id="ref76">18</reflink>] As an example, the seventh item is:</p> <p>A recent research study randomly divided participants into groups who were given different levels of Vitamin E to take daily. One group received only a placebo pill. The research study followed the participants for eight years to see how many developed a particular type of cancer during that period. Which of the following responses gives the best explanation as to the purpose of randomization in this study?</p> <p>The response choices are: "to increase the accuracy of the research results," "to ensure that all potential cancer patients had an equal chance of being selected for the study," "to reduce the amount of sampling error," "to produce treatment groups with similar characteristics," and "to prevent skewness in the results." The correct answer is the fourth response choice.</p> <p>The CAOS was developed over the course of 3 years. First, it was decided to focus the test on variability through 10 topics. Then, the authors assembled items from statistics instructors and created additional items for the topics that were not covered. The initial version had 34 items. This version was pilot tested with students enrolled in an introductory statistics course. The data allowed the authors to revise this version, which resulted in a 37‐item test. The CAOS was revised and tested two more times. The final version has 40 items divided into 10 topics: data collection and design, descriptive statistics, graphs, boxplots, normal distribution, bivariate data, probability, sampling variability, confidence intervals, and significance testing.</p> <p>Reliability of this final version was examined using Cronbach's alpha. The results suggest a good internal consistency with college students. Content validity was examined by asking 18 members of the <emph>Consortium for the Advancement of Undergraduate Statistics Education</emph> (CAUSE), who are both statisticians and statistics instructors, to act as experts and rate the items. These members revised the test, answered a few questions about its validity and unanimously agreed with the statements "CAOS measures important outcomes that are common to most first courses in statistics" and "CAOS measures outcomes for which I would be disappointed if they were not achieved by students who succeed in my statistics courses" (p. 31). Also, 94% of these experts agreed with "CAOS measures important outcomes that are common to most first courses in statistics"[<reflink idref="bib18" id="ref77">18</reflink>] (p. 31).</p> <hd id="AN0188002603-13">SRBCI</hd> <p>The <emph>Statistical Reasoning in Biology Concept Inventory</emph> is a 12‐item multiple‐choice measure. The items are divided into three blocks of four items each related to the same graphic and scenario. This instrument was developed for assessing undergraduate biology students' understanding of statistics.[<reflink idref="bib17" id="ref78">17</reflink>] The items are available on this website: https://q4b.biology.ubc.ca/concept-inventories/statistical-reasoning-in-biology/. As an example, the ninth item is:</p> <p>You tested whether temperature affected the growth of sunflowers over a 6‐month period after germination. You tested 24 plants in each of three treatments (12°C, 13°C, 14°C). You recorded the height (cm) of each plant 6 months after germination, and then calculated averages (means) and a measure of variation around the averages (95% confidence intervals).</p> <p>At 12°C, plants grew 114 cm ± 3 cm, at 13°C they grew 122 cm ± 4 cm, and at 14°C they grew 120 cm ± 8 cm. Which of the following is the most accurate interpretation of these results?</p> <p>The response choices are: "Temperature did affect growth rate because plants grew significantly taller at 13°C than at 12°C," "temperature did not affect growth rate because plants grew tallest in the middle temperature (13°C)", "temperature did affect growth rate because the average (mean) growth was different in each of the three temperatures," and "temperature did not affect growth rate because plants grown in the coolest (12°C) and hottest (14°C) temperatures did not grow significantly differently." The correct answer is the first choice.</p> <p>The items are divided into four "core conceptual grouping[s]"[<reflink idref="bib17" id="ref79">17</reflink>] (p. 4): repeatability of results, variation in data, hypotheses and predictions, and sample size. To create the items, the authors first examined the literature, looked at statistics exams and assignments and surveyed faculty members on important statistical concepts. The results indicate that there are four key concepts, which are the focus of the SRBCI: data variability, replication of results, hypotheses and predictions, and sample size. Then, the authors created three items for each key concept for a total of 12. The first and only version was tested with science students using interviews. Validity of this version was tested with expert revision. No clear indication of reliability or validity is provided. However, the authors ran a Rasch model analysis[[<reflink idref="bib7" id="ref80">7</reflink>], [<reflink idref="bib46" id="ref81">46</reflink>], [<reflink idref="bib48" id="ref82">48</reflink>]] where the results suggest that the data fits the one‐dimensional model well. Pairwise item correlations were calculated and were all within recommendations except for the correlations between item 1 and item 6, between item 3 and item 11, and between item 7 and item 8. Then, the infit and outfit mean‐square's measures of fit of each item were calculated. These varied between 0.6 and 1.4, which is within recommendations. Next, the Pearson product–moment correlations were computed between participants' score on the SRBCI and the Rasch model estimates. The authors reported almost linear correlations. Lastly, item difficulty was examined. It was concluded that the SRBCI is an appropriate instrument for evaluating students' statistical reasoning.</p> <hd id="AN0188002603-14">Procedure</hd> <p>The four tests were assembled into a single survey. The SRA's 20 items were divided into 18 blocks of item(s) as items 5 and 6 share the same scenario, and similarly for items 9 and 10. In the SCI, 25 blocks of one item each were formed since each item has its own scenario. In the CAOS, the 40 items were divided into 25 blocks of item(s) with the following items grouped together because they have the same scenario: 1 and 2, 3 to 5, 8 to 10, 11 to 13, 14 and 15, 23 and 24, 25 to 27, 28 to 31, and finally 34 and 35. Concerning the SRBCI, as said above, its 12 items are grouped into 3 blocks of 4 items each: items 1 to 4, 5 to 8, and 9 to 12. Consequently, all items were spread over 71 blocks.</p> <p>Because 71 is a high number of blocks, a random subset of 20 of these blocks were presented to each participant. All 71 blocks were divided randomly across participants. Further, the order of the selected blocks was random. Participants were uninformed to which items of which scale they answered. However, participants were informed that they did not have to answer the items. What really mattered in each block was an open‐ended question where they were asked, "what are the concept(s) or prerequisite(s) needed to answer this question?" Their answer to this question provided the qualitative data kept for the thematic analysis. A screen capture of a block is presented in Figure 1. After participants completed the questionnaire, their quantitative scores for each item were given to them along with the correct answer, based on the tests' developers. Participants received one point for each correctly answered item. Finally, they could provide comments on the last page.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/D8Y/01sep25/test12410-fig-0001.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="test12410-fig-0001.jpg" title="1 The first item of the SRA with the added question." /> </p> <p></p> <hd id="AN0188002603-16">Planned analyses</hd> <p></p> <hd id="AN0188002603-17">Qualitative</hd> <p>The text entries of each block of item(s) are the focus of this research. To analyze them, a thematic analysis divided into 6 phases was done, as described by Braun and Clarke.[<reflink idref="bib8" id="ref83">8</reflink>] The first and last phases include familiarization with the dataset and reporting the themes that were identified. The second phase involves generating codes for each block of item(s). Those codes are the statistical concepts listed by participants. Some participants wrote complete sentences explaining their reasoning, while others only provided the concepts they thought were necessary to respond to the item. As such, the codes helped to differentiate between words that are pertinent to the research objectives and words that are not. The third phase concerns the identification of potential themes in each block of item(s). The themes are codes or groups of codes that could be united. In other words, the themes can be identified through patterns of responses, that is, statistical concepts that are often explicitly or implicitly mentioned by participants. In the fourth phase, the list of potential themes is revised. It is required that themes do not only reflect codes but are also pertinent to the objectives. Themes must reveal something meaningful about the data; they must reflect the entire dataset in a coherent way. Finally, in the fifth phase, the themes are defined and named so that they include all the codes. Note that the analysis was realized by the first author and revised by the second author. Also, this study is exploratory because there is no a priori framework from which one can rely on.</p> <hd id="AN0188002603-18">Quantitative</hd> <p>Participants scores were calculated. Basic frequencies, sums, and means were used.</p> <hd id="AN0188002603-19">RESULTS</hd> <p></p> <hd id="AN0188002603-20">Data screening</hd> <p>Before data cleaning, 117 instructors clicked on the link. However, most of them did not finish the questionnaire. Specifically, some did not answer all open‐ended questions and others simply did not answer any question. Only 42 participants, or 35.9%, who responded to all the 20 blocks were kept for analysis. Neither demographic data nor their number of years of experience as a statistics instructor was collected. Those 42 participants provided qualitative responses to 710 blocks with 130 missing responses on the tests' items: 169 in the SRA (response rate of 88.5%), 259 in the SCI (response rate of 83.0%), 247 in the CAOS (response rate of 84.3%) and 35 in the SRBCI (response rate of 79.5%); remember that the participants were told that responding to the items was optional. Tables S1 to S4 (available on OSF; https://osf.io/t6w7n/) show the frequencies of qualitative responses per block of item(s).</p> <p>Three examples of responses are as follows. In response to the first item of the SRA: "In this question one would need to know how to identify a genuine outlier and further recognize that the mean is likely the most accurate answer. On a related note, we don't teach writing up statistics in stats class usually. This would be a great one to follow with asking them how they would write that up in a results section." For the items 34 and 35 of the CAOS, one response was: "understanding of random sampling and sampling distributions, the central limit theorem." Regarding the items 1 to 4 of the SRBCI, a participant wrote, "confidence intervals, sampling distributions, ANOVA, interpreting null effects."</p> <p>Participants took on average 3 h and 53 min to complete the questionnaire. Since Qualtrics saves participants' progression, some must have completed it in an "on and off" manner or over multiple sessions, resulting in four durations exceeding 20 h. The smallest amount of time taken is 21 min and the highest amount of time taken is 43 h and 57 min. The median completion time is 55 min.</p> <hd id="AN0188002603-21">Themes per tests</hd> <p></p> <hd id="AN0188002603-22">Identification of codes</hd> <p>In the SRA, a total of 60 codes were identified for an average of 3 codes per block. In the SCI, 69 codes were identified with approximately 3 per item. For the CAOS, there are 87 codes in total, likewise for an average of 3 codes per block. Finally, in the SRBCI, 23 codes were identified in total for an average of 8 codes per block. The tables of codes for each test (Tables S5–S8) are available on OSF at https://osf.io/t6w7n/.</p> <p>The codes of the SRA are somewhat similar to the 16 skills mentioned in the Methods section. Some other codes are very different from what is measured according to the authors. For example, item 24 of the CAOS is thought to measure understanding of statistical significance, statistical power, causality, type I and II errors. However, the authors mentioned that this item measure "understanding that an experimental design with random assignment supports causal inference"[<reflink idref="bib18" id="ref84">18</reflink>] (p. 44). Another example concerns items 5 to 8 of the SRBCI. Table S8 indicates that these items measure understanding of CIs, variance and variability, sample size, NHST, central limit theorem, and standard error. However, the authors mentioned that they measure understanding of "hypotheses and predictions" (items 5, 7, and 8), "repeatability of results" (item 6)[<reflink idref="bib17" id="ref85">17</reflink>] (p. 4). These differences suggest that when one examines the items alone, the statistical concepts measured by those tests are not clear. The codes in bold represent procedural skills. As can be seen, some tests measure more procedural skills than others.</p> <hd id="AN0188002603-23">Potential themes</hd> <p>The potential themes of each questionnaire across items are shown in Table 1. The recurrence was calculated at the questionnaire level, meaning that 44% of the SRA's responses on the open‐ended questions concern probability and trial independence. The SRA has eight themes, with three that have a recurrence of 14% and higher (i.e., <emph>Probability and Trial Independence</emph>, <emph>Sample and Sampling</emph>, and <emph>Central Tendency</emph>). The SCI has 12 themes, with 3 that have a recurrence above 10%. These are: <emph>Central Tendency</emph>, <emph>Standard deviation</emph> (SD)/<emph>Variance</emph>, and <emph>Distribution</emph> (<emph>normal</emph>, <emph>sampling</emph>, etc.). The CAOS (which is also the longest test with 40 items) has 13 different themes, with 4 that have a recurrence of 10% or more, namely, <emph>Null‐hypothesis Statistical Testing</emph> (NHST)/<emph>p‐value</emph>, <emph>Distribution</emph> (<emph>normal</emph>, <emph>sampling</emph>, etc.), <emph>Correlation/Causality</emph>, and <emph>Error</emph> (<emph>sampling, types, standard</emph>). Lastly, the SRBCI (the shortest test with 12 items) has only three themes (i.e., <emph>NHST</emph> with a recurrence of 33%, <emph>Sample Size</emph> with a recurrence of 28%, and <emph>Confidence intervals</emph> [<emph>CIs</emph>] with a recurrence of 28%).</p> <p>1 TABLE Potential themes of each questionnaire with their percentage of recurrence (percentages may add to more than 100% as some question items were identified with more than one theme).</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">SRA</th><th align="left">SCI</th><th align="left">CAOS</th><th align="left">SRBCI</th></tr></thead><tbody valign="top"><tr><td align="left">Probability and trial independence (44.1%)</td><td align="left">Central tendency (14.9%)</td><td align="left">NHST/p‐value (17.3%)</td><td align="left">NHST (33.3%)</td></tr><tr><td align="left">Sample and sampling (17.3%)</td><td align="left">SD/variance (14.9%)</td><td align="left">Distribution (normal, sampling, etc.; 17.0%)</td><td align="left">Sample size (28.2%)</td></tr><tr><td align="left">Central tendency (14.5%)</td><td align="left">Distribution (normal, sampling, etc.; 14.2%)</td><td align="left">Correlation/causality (12.2%)</td><td align="left">CIs (28.2%)</td></tr><tr><td align="left">Outlier (7.8%)</td><td align="left">Correlation (7.6%)</td><td align="left">Error (sampling, types, standard) (10.3%)</td></tr><tr><td align="left">Error and bias (7.3%)</td><td align="left">Probability (7.3%)</td><td align="left">Graphs/visualization (9.6%)</td></tr><tr><td align="left">Sample size (6.2%)</td><td align="left">Sample size (5.9%)</td><td align="left">Research methods/design (7.8%)</td></tr><tr><td align="left">Statistical tests (6.2%)</td><td align="left">CIs (5.2%)</td><td align="left">Variation/variability (7.4%)</td></tr><tr><td align="left">Math (5.6%)</td><td align="left">NHST/p‐value (5.2%)</td><td align="left">Central tendency (5.5%)</td></tr><tr><td align="left">Outliers (4.5%)</td><td align="left">Asymmetry (4.4%)</td></tr><tr><td align="left">Rank order (4.5%)</td><td align="left">Outliers (4.4%)</td></tr><tr><td align="left">Graphs (4.2%)</td><td align="left">Descriptive statistics (4.1%)</td></tr><tr><td align="left">Error/bias (3.5%)</td><td align="left">Probability (3.7%)</td></tr><tr><td align="left">Regression (3.7%)</td></tr></tbody></table> </ephtml> </p> <p>1 <emph>Note</emph>: The procedural skills are in bold.</p> <p>2 Abbreviations: CI, confidence interval; NHST, null hypothesis statistical tests; SD, standard deviation.</p> <p>Given the number of items, the SCI seems to be the one exploring the largest number of statistical concepts (12 themes for 25 items) whereas the SRBCI is the narrowest test (3 themes for 12 items). Two of the three themes in the SRBCI also happen to be high‐level themes which need ample prior training to be mastered or understood (specifically <emph>Confidence Intervals</emph> and <emph>Null Hypothesis Statistical Testing</emph>). Also, the test least focused on procedural skills and most focused on conceptual understanding seems to be the SRA.</p> <hd id="AN0188002603-24">Themes across tests</hd> <p>The authors identified 994 codes from the participants responses across all tests. At first, to construct this final list, only the tests' themes with a recurrence of 10% and higher were kept. Some of those themes were merged to form a unique theme. Ten different themes were identified. After discussion, it was realized that the themes with lower recurrence are not necessarily less important in statistics. Their low recurrences may have been caused by how the items were formulated and how the tests were developed. As such, after merging some themes together, the final list has 20 themes (including one theme not further delineated into multiple codes). These are shown in Table 2.</p> <p>2 TABLE Global themes across tests with percentage of recurrence.</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Themes</th><th align="left">Recurrence</th></tr></thead><tbody valign="top"><tr><td align="left">Probability and trial independence</td><td align="char" char="(">114 (11.6%)</td></tr><tr><td align="left">Central tendency</td><td align="char" char="(">98 (10.0%)</td></tr><tr><td align="left">Distributions</td><td align="char" char="(">97 (9.9%)</td></tr><tr><td align="left">NHST</td><td align="char" char="(">75 (7.6%)</td></tr><tr><td align="left">Sample and sampling</td><td align="char" char="(">70 (7.1%)</td></tr><tr><td align="left">Correlation and causality</td><td align="char" char="(">69 (7.0%)</td></tr><tr><td align="left">Statistical tests and computation</td><td align="char" char="(">66 (6.7%)</td></tr><tr><td align="left">Error and biases</td><td align="char" char="(">61 (6.2%)</td></tr><tr><td align="left">Confidence intervals</td><td align="char" char="(">48 (4.9%)</td></tr><tr><td align="left">Descriptive statistics</td><td align="char" char="(">48 (4.9%)</td></tr><tr><td align="left">Graphs and visualization</td><td align="char" char="(">44 (4.5%)</td></tr><tr><td align="left">Outliers</td><td align="char" char="(">42 (4.3%)</td></tr><tr><td align="left">Variability and dispersion</td><td align="char" char="(">35 (3.6%)</td></tr><tr><td align="left">Research methods/designs</td><td align="char" char="(">29 (3.0%)</td></tr><tr><td align="left">Randomness and randomization</td><td align="char" char="(">25 (2.5%)</td></tr><tr><td align="left">Asymmetry</td><td align="char" char="(">18 (1.8%)</td></tr><tr><td align="left">Representativity and generalizability</td><td align="char" char="(">12 (1.2%)</td></tr><tr><td align="left">Central limit theorem</td><td align="char" char="(">10 (1.0%)</td></tr><tr><td align="left">Statistical power</td><td align="char" char="(">8 (0.8%)</td></tr><tr><td align="left">Unclassified</td><td align="char" char="(">14 (1.4%)</td></tr><tr><td align="left">Total</td><td align="char" char="(">983 (100.0%)</td></tr></tbody></table> </ephtml> </p> <p>3 <emph>Note</emph>: The themes that were not further delineated into multiple codes include: (in)dependence/codependence, linear relation, prediction, law of large numbers, interpretation, calibration, outcome. The procedural skills are in bold.</p> <p>As can be seen, the tests have a few common themes: probability, central tendency, distributions, NHST, and samples. These themes are somewhat similar to the first dimension of the SRA and to the four dimensions of the SRBCI, but are quite different from the topics covered in the SCI—which are mostly computational. Regarding the CAOS, some themes are somewhat similar, and others are quite different.</p> <p>These 20 themes were then revised (see Table 3). Through this revision, some themes were paired, some stayed the same, and others were renamed to ensure that they truly captured the data and the codes. In total, 18 revised themes were defined. These themes represent the concepts that are measured in the four tests (SRA, SCI, CAOS and SRBCI).</p> <p>3 TABLE Revised global themes across tests and their definition.</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Themes</th><th align="left">Definitions</th></tr></thead><tbody valign="top"><tr><td align="left">Probability and randomness</td><td align="left">Some events are deterministic, others random. Random events are unpredictable in detail, but some events are less random than others, in the sense that they are more likely than others. Random events can occur with varying degrees of certainty (quantification level of the concept). The degree of certainty of occurrence of a random event (Bayesian definition); establishing the frequency of its occurrence in identical reproductions (frequentist definition).</td></tr><tr><td align="left">Central tendency</td><td align="left">The population may have a more typical observation (conceptual level). It is possible to estimate the most frequent, the most central, or the least erroneous magnitude of random events (quantification level). These three estimates are slightly different and correspond to the mode, median and mean respectively. Populations also have theoretical central tendencies; we can only estimate these from a sample (population parameter level).</td></tr><tr><td align="left">Distributions</td><td align="left">Conceiving the repartition of distinct events (sampling distribution), the repartition of levels of inaccuracy (error distribution). Understanding that these random draws result in regularities (sampling laws). Conceiving the bell curve (normal distribution) as an example of a distribution.</td></tr><tr><td align="left">NHST and p‐value</td><td align="left">NHST is a particular form of statistical inference (conceptual understanding). We can pose an a priori hypothesis about the result of a statistic, and state its "null" version. We choose a level of risk to adopt for the decision. There are several procedures (tests) for testing a null hypothesis, and these procedures are based on assumptions (quantification level). We can apply the procedure to obtain inferential statistics; establish the level of plausibility of the data when the null hypothesis is assumed (p‐value); make a decision based on this level of plausibility.</td></tr><tr><td align="left">Sample and sampling</td><td align="left">Randomly draw instances from the target population. For each draw, these instances will be different. There are many ways one can randomly draw instances from a population</td></tr><tr><td align="left">Shared variance and correlation</td><td align="left">Two characteristics of a subject can be associated; these associations can be positive or negative (conceptual understanding). These associations are expressed as co‐variation. These statistics are part of descriptive statistics but require bi‐variate data. The degree of association can be quantified by covariation, and by correlation (level of quantification). These degrees of association can be compared. Bivariate populations also have a covariation/correlation (level of a population parameter).</td></tr><tr><td align="left">Error and its margins</td><td align="left">Whenever a descriptive statistic is quantified, there ought to be some error in this quantification (conceptual understanding). This is due to chance and variability in the population examined, so the sample is more or less perfectly representative of the population. This sampling error leads to errors as soon as statistical results are produced from this sample. It may be in descriptive statistics (quantification level: estimation error, which can be quantified with the standard error and the precision interval) or in the hypothetical value given to a parameter (quantification level: inference error, which can be quantified with the confidence interval). It can also be described by a range of values adapted to a level of risk (quantification level: margin of error can be more or less narrow, depending on whether you want a margin of error that most probably contains the true value (high confidence) or fairly probably contains the true value (lower confidence). Any result is probably wrong to some extent, and so an estimate of the error should always accompany the result. These estimates can be quantified (standard error, confidence interval, or other technique; quantification level). These margins of error can be interpreted. Margins of error depend on the experimental design.</td></tr><tr><td align="left">Descriptive statistics</td><td align="left">A sample can be described. Central tendencies, variabilities and dispersions and asymmetries are possible descriptions, but there are others (e.g., quantiles and percentile ranks; conceptual understanding). These descriptive statistics are distinct from population parameters, but often serve to approximately quantify these parameters when they are unknown (with a range of values). These descriptive statistics can be quantified and compared (quantification level). Populations also have other aspects (level of a population parameter).</td></tr><tr><td align="left">Graphs and visualization</td><td align="left">Determining what is displayed on the graph; in particular, distinguishing between Venne diagrams, scatter graphs and summary graphs. Understanding the notion of dependent variable (the one to be read) and independent variable(s) (the one(s) that form(s) the data grouping(s)) or the notion of bivariate. Understanding whether what is displayed is raw data or descriptive statistics. When descriptive statistics are used, (i) determining whether margins of error accompany the results and (ii) being able to compare different statistics (conceptual understanding).</td></tr><tr><td align="left">Outliers and contamination</td><td align="left">Sometimes, the sample is contaminated by elements (participants) who are not actually members of the target population. These contaminants may be visible (outlying) or undetectable (inlying). An absence or a low level of contamination is preferable (conceptual understanding).</td></tr><tr><td align="left">Variability and dispersion</td><td align="left">Observations (participants) differ from one another, and the degree of difference may vary according to the population studied and the sample drawn from the population (conceptual understanding). This degree of difference is quantified by the typical deviation between observations. These typical deviations can be compared (quantification level). Populations also have theoretical dispersion trends; we can only estimate them from a sample (population parameter level).</td></tr><tr><td align="left">Research, research methods and research designs</td><td align="left">The conditions under which the data are obtained (e.g., whether subjects are measured once or several times; conceptual understanding). Data can be independent groups, repeated measures or paired. The specification is important to ensure that the data are valid (free of confounds), representative of the population and generalizable to that population (quantification level).</td></tr><tr><td align="left">Asymmetry</td><td align="left">Sampled scores can be asymmetrical, that is, more frequent on one side of the distribution than the other, more long‐range or more squashed. This asymmetry may be caused by a floor effect, a ceiling effect, contaminants and outliers, for example (conceptual understanding). This asymmetry can be quantified using the relevant formula (skewness or kurtosis); these figures can be compared (quantification level). Populations also have a theoretical asymmetry (level of a population parameter). A normal population has no asymmetry (its value is zero).</td></tr><tr><td align="left">Population</td><td align="left">Each study has one or more target groups of people, animals or objects. It is this (these) group(s) that is (are) typically recruited to participate. The results of a study are relevant to this (these) group(s) (conceptual understanding).</td></tr><tr><td align="left">Modeling</td><td align="left">The population and its relations are idealized by a modeling process. This process posits theoretical distributions, means, and/or correlations or covariations (conceptual understanding). Models built on data can be used to summarize the relations that exist between variables, and how they influence each other (quantification level).</td></tr><tr><td align="left">Data and observation</td><td align="left">Data are observations made on subjects participating in the study. Data can be univariate (composed of a single number quantifying a single characteristic of the subject, such as empathy level), bivariate (composed of two numbers quantifying two characteristics of the subject, such as height and weight) or multivariate (two or more numbers quantifying two or more characteristics of the subject). Each individual piece of data is unpredictable, as the subject is chosen at random (conceptual understanding).</td></tr><tr><td align="left">Decisions, risks and consequences</td><td align="left">An informed decision must be based on considerations of risk levels and the consequences of these risks. For example, a decision to undergo ovarian ablation following a single positive diagnosis of ovarian cancer must consider the risk of the diagnosis being a false positive, and the possible consequences of surgery. Understanding the notion of false positives, false negatives, as well as true positives and true negatives. These 4 possible outcomes are probabilistic and not certainties in most situations (conceptual understanding).</td></tr><tr><td align="left">Statistical inference</td><td align="left">Inference is the act of choosing an interpretation from observations. Statistical inference is based on observations sampled from a population, which are the result of a random phenomenon calculated through various statistical tests mediated by a model of the population. These tests are chosen according to the type of observations, the study design and the research question (conceptual understanding).</td></tr></tbody></table> </ephtml> </p> <p>4 <emph>Note</emph>: The procedural skills are in bold.</p> <hd id="AN0188002603-25">Lack of consensus around correct responses</hd> <p>As a complementary analysis, the participants' scores were examined. The participants were informed that the focus was on their assessments regarding the concepts, but many nonetheless answered the tests' items.</p> <p>Average scores are relatively low, ranging from 66% in the SRA to 83% in the SRBCI. Corresponding tables (Tables S9–S12) are available on OSF at https://osf.io/t6w7n/. These scores represent the average percentage of participants' score. Remember that the participants are all statistics instructors. It indicates that the tests are not as clear as one would assume. Indeed, there does not seem to be a consensus on what the correct answer for each item is among experts. This can be further examined in participants' text entries presented. Also, note that the best‐succeeded questionnaire (SRBCI) was built for biology students, whereas the recruited participants are researchers and instructors in the psychological sciences. Hence, the low scores in the other three tests cannot be attributed to a lack of competency from the participants.</p> <p>For some items, there does not seem to be an agreement between participants and the tests' developers on what the correct answer is. Some participants said that there was no single correct answer for an item. An example of such comment concerns items one to four of the SRBCI (note that all comments below come directly from the participants' responses):</p> <p>Given some of the wording there is no set answer to some of the questions, so it depends on how a student was educated. Not having the test in front of me I can't recall exactly the issues was one where there was an increasing effect, and the question is whether that increasing effect was significant (I think). A rational look would argue that taking all of the CIs into effect and the overall increasing effect would argue that you have something there. But some people, under a stricter hypothesis testing (and multiple testing) education that may even defy logic at times would focus on one of the increases being large enough that CIs don't overlap [...].</p> <p>Other participants argue that the correct answer is none of the response choices given. One participant said that item 4 of the SRA does not have the actual correct answer in its choices:</p> <p>This could be seen as a trick question, as the use of the word 'typical' hints that the mode is being sought. [...] Frankly, given that outlie[r] of 22, I'd be more inclined to choose the median as my measure of central tendency but that doesn't appear as an option.</p> <p>There also seems to be confusion in the participants' responses to some items. Participants do not seem sure which response choice they would choose. Some argue that the possible responses don't match the theory behind a concept. For example, one said for item 22 of the CAOS:</p> <p>I'm guessing maybe you wanted the first answer. None of the others are as good. But even though the design doesn't permit causal conclusions one can sometimes draw logical conclusions that are reasonable. One can conclude that earning more money does cause more recycling because the opposite is ridiculous. What one cannot do is draw the strong conclusion that the cause is direct and psychological. So, the alternative presented here with a conclusion is clearly wrong because it goes after a deeper direct cause that's completely unwarranted. This requires a surface understanding that correlation does not mean causation.</p> <p>This spontaneous discussion of correlation vs. causation shows that despite the validation and reliability testing done with CAOS, the wording of the items is not ideal.</p> <hd id="AN0188002603-26">DISCUSSION</hd> <p>The goal of the present study was to understand what is evaluated by the four standardized statistics understanding tests. A subsidiary goal was to examine participants' scores on those tests. In other words, the objective was to examine, from the perspective of experts, which statistical concepts are measured and whether experts agree with the tests' developers on the correct answers. The codes—representing the concepts participants listed—of each item were extracted, and then the themes—representing the concepts measured by the tests—were synthesized into a list of basic statistical concepts given in Table 2. Those concepts may be thought to be the building blocks of statistics, at least if these four tests are <emph>adequate</emph> and <emph>exhaustive</emph> reflections of statistics <emph>understanding</emph>.</p> <p>Three intermediate conclusions can be reached regarding these assumptions. (i) Some participants expressed concerns that the material used for assessments may not be adequate. They suggested that the items may use biased formulation or skewed options. Their low scores on the tests (in particular for the SRA with only two‐thirds of correct answers) lend support for this critique. Using four distinct tests hopefully alleviated this potential problem. (ii) The SRBCI encompassed only three statistical concepts, an incredibly small number given that a total of 19 (plus one miscellaneous) themes were identified. Whether such tests lack exhaustiveness or the themes given are "high‐level" concepts which hide multiple drawer concepts, so to speak, remains to be clarified. (iii) The results suggest that memorization of concepts plays a critical role in the tests. For example, many items rely on remembering what a sample is (its definition) but not necessarily the implications or consequences of <emph>varying</emph> the sample sizes. Likewise, many assumed that the short test's names were known and remembered (<emph>t</emph> test, <emph>z</emph> test, chi‐square test) which does not measure the understanding of the inner cogs of these procedures, only the ability to recognize their labels. As such, they may measure statistical literacy more than statistical understanding, or they may measure more procedural skills than conceptual understanding.</p> <hd id="AN0188002603-27">Methodological limitations</hd> <p>As pertinent as the results may be, it is necessary to consider the limitations of the present study. First, the sample size was smaller than expected. Most instructors who clicked on the link did not finish the questionnaire, and so were excluded from analysis. Small samples have some disadvantages. Specifically, the sample size may have affected the proportions and diversity of codes and themes that were identified. Indeed, with a larger sample, the list of fundamental statistical concepts may be different. The small sample size may have affected the randomization of blocks of item(s). When looking at the number of responses for each block, we see that some blocks of item(s) have more responses than others. For example, the block of items 3 to 5 of the CAOS has 14 responses, but its block of item 16 has only 8. This unequal distribution may have affected the codes' recurrence, their proportion, the themes identified, and participants' mean score. Also, the small number of responses per item (between 6 and 15) could have affected the qualitative and quantitative results. A larger sample would have minimized these issues.</p> <p>Another limitation concerns the target population. It is believed that the chosen sample was better suited to participate in the present study for the reasons mentioned in the Methods section. However, the less experienced participants in teaching statistics may have provided more incorrect answers to the tests' items. Also, the perspective of two other populations could have provided more balanced data. Indeed, students and administrators could have provided interesting responses, which could have influenced the results. It is possible that additional codes and themes could have been identified, or the recurrence of the themes could have been influenced, which would modify the current list of basic statistical concepts.</p> <p>A last limitation concerns the tests chosen as the study material. It was acknowledged that there are other tests measuring statistics understanding, and while these did not fit with the study's objectives, there is one that could have been added. This test is the LOCUS project. However, the test was developed based on the GAISE recommendations for PreK‐12 statistics education within the mathematics curriculum.[<reflink idref="bib21" id="ref86">21</reflink>] Considering that the objective of the present study was to investigate tests measuring statistics understanding of college or university students in psychology, the LOCUS was not used. Also, while the same statistical concepts may be taught across disciplines, it is necessary to keep in mind that these tests were developed for different populations; the SRBCI, for example, was developed for biology students.</p> <hd id="AN0188002603-28">Theoretical limitations</hd> <p>Naming the statistical concepts that must be learned will allow the development of assessment tools that are valid with respect to content, which in turn will allow proper evaluation of students' understanding. The approach taken herein makes the overarching assumptions that there are distinct concepts in statistics that can be isolated. However, this assumption could be wrong. Maybe statistical concepts are an intrinsically intertwined web of interrelated concepts. Take the confidence interval for example. To understand that this concept provides a range of plausible values for a descriptive statistic with a certain risk level, one needs to understand sampling, central tendency, variability, and probability. Yet, to figure out variability, a range of values is an apt representation which is precisely the desired understanding to reach from CIs. In other words, the concepts could be like Kekulé's snake biting its own tail. As such, there may not be one ideal teaching sequence for statistics, or one ideal curriculum ordered in an ideal way. As another example, some quantities are meta‐statistics: the standard error is the standard deviation of descriptive statistics, that is, a measure of spread of a second order. Thus, the concept of spread has multiple layers of existence. It may be that repetition is the key to teaching statistics. By being taught the concepts and their use multiple times, each additional iteration presenting more abstract applications, it may become possible to gain a deeper understanding of those concepts. In other words, maybe learning the same concepts multiple times is a necessity in statistics education to create a conceptual map that gets bigger, clearer, and more nuanced every time.</p> <p>Another assumption adopted herein is the view that statistical concepts embody static knowledge. In this view, a concept is a fixed declaration, like, for example, "2 + 2 = 4". Maybe this assumption is wrong as well. Statistical concepts could be operations that make elements move into new states. For example, sampling could be seen as the displacement of a few objects from a general pool to a smaller pool, called "the sample" (Stuart[<reflink idref="bib53" id="ref87">53</reflink>] made a similar argument for <emph>population</emph>). Thus, what would matter in statistics teaching is not the transmission of fixed concepts, but the transmission of transformations. Students may not be well equipped (in terms of cognitive abilities) to manipulate statistics components. Consequently, teaching how to mentally manipulate elements could facilitate the understanding of statistics.[<reflink idref="bib27" id="ref88">27</reflink>] To investigate this possibility, one may look at the contribution of mental manipulation abilities to statistics understanding (a study currently in preparation). In any case, a better questionnaire measuring conceptual understanding of statistics instead of memorization and procedural skill capacities is needed.</p> <hd id="AN0188002603-29">Implications</hd> <p>This list of basic statistical concepts (seen in Table 2) has implications for statistics education. As mentioned in the introduction, there are many issues in the education of statistics, and many studies argue that the curriculum is not well developed. Due to advancements in technology, many no longer recommend teaching (and assessing) students (on) how to compute a statistic by hand (e.g.,[<reflink idref="bib12" id="ref89">12</reflink>]). In other words, many no longer recommend teaching statistical procedural skills. Also, some research argues that there are missing or poorly described concepts in the curriculum, such as the concept of variability.[<reflink idref="bib42" id="ref90">42</reflink>] The present study is a starting point toward the improvement of the teaching of statistics through conceptual understanding. With an in‐depth examination of four tests on statistics understanding, it was possible to identify specific statistical concepts that are being measured in these tests. Those concepts can be considered as building blocks of statistics and organized with prerequisites. For example, to understand the concept of variability, one needs to understand the concept of population and sample/sampling first. As such, an instructor can first teach the concept of population, then of sample/sampling, and then of variability with different methods to measure it (e.g., CIs, standard deviation, etc.).</p> <p>Having the building blocks of statistics—what must be taught to students and so assessed—would solve many issues in statistics education. The list of basic concepts provides insights into what is or should be typically taught in introductory statistics courses. In essence, it provides professors and educators with a curriculum. In turn, knowing what is taught gives us insights into the learning objectives of those courses. By knowing what to teach, instructors will be better prepared and will be able to build better assessments and assignments that will support their students' learning, fostering deeper conceptual understanding. Furthermore, it may stimulate research into statistics education.</p> <p>Finally, the qualitative results of the present study indicate the need for better developed instruments that measure students' conceptual understanding rather than memorization capacity and procedural skills. Indeed, some themes identified did not match the concepts the tests were designed to measure. When looking at the quantitative data, the results suggest that experts did not agree on the correct answers. Some even argued what the correct answer should be and why. This could either mean that statistics instructors in social sciences are not as good at statistics and hold some misconceptions that they transfer to students, or that the tests are not as clear as first thought. Both explanations have big implications for the education of statistics. The first explanation brings to light the difficulty and non‐intuitive nature of statistics, which is not new. What is novel, however, is that these characteristics persist even after many years of instruction and experience. This may point to the need for further research on professors' statistical ability. For example, it would be interesting to do the same study with statistics professors coming from statistics departments and compare the results. Likewise, it would be interesting to investigate if there is a discrepancy between what statisticians say the learning goals should be and what is taught in classes, as a reviewer of the present text pointed out. Another possible study is to examine if it is an issue with the training of non‐statistician statistics instructors, like psychology professors. The second explanation supports the conclusion drawn from the qualitative results, which points to more research on assessments in statistics.</p> <hd id="AN0188002603-30">Conclusion</hd> <p>Through an investigation of four standardized tests on statistics understanding, it was possible to identify basic concepts that are typically taught to university students. Statistics instructors were asked to identify the concepts measured in each item. They also had the possibility of answering the items. The results suggested that there are 19 concepts measured in those four tests. Also, the instructors that answered the items had scores of 69% for the SRA, 80% for the SCI, 75% for the CAOS, and 83% for the SRBCI, for a grand mean of 76.8%. The weighted average is 76%. This large variance across tests calls for additional research on devising statistics assessments. In the end, this research may benefit the teaching and learning of statistics to foster understanding over memorization, form more efficient citizens, and future instructors to teach future students.</p> <hd id="AN0188002603-31">ACKNOWLEDGMENTS</hd> <p>This research began while the first author was completing a Practicum in Research taught by Jean‐François Bureau at the University of Ottawa during the 2021–2022 school year. She thanks him for the opportunity and the classmate who reviewed an earlier version anonymously.</p> <hd id="AN0188002603-32">FUNDING INFORMATION</hd> <p>This work was supported by the Social Sciences and Humanities Research Council under Grant number RGPIN‐2024‐03733.</p> <hd id="AN0188002603-33">CONFLICT OF INTEREST STATEMENT</hd> <p>The authors declare no conflicts of interest.</p> <hd id="AN0188002603-34">DATA AVAILABILITY STATEMENT</hd> <p>The data is available on OSF (https://osf.io/t6w7n/).</p> <ref id="AN0188002603-35"> <title> REFERENCES </title> <blist> <bibl id="bib1" idref="ref67" type="bt">1</bibl> <bibtext> R. Ackerman, D. M. Bernstein, and R. Kumar, Metacognitive hindsight bias, Mem. Cognit. 48 (2020), no. 5, 731 – 744. https://doi.org/10.3758/s13421-020-01012-w.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref32" type="bt">2</bibl> <bibtext> K. Allen. The statistics concept inventory: Development and analysis of a cognitive assessment instrument in statistics, Doctoral dissertation, University of Oklahoma. SSRN Journal. 2006 https://doi.org/10.2139/ssrn.2130143.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref13" type="bt">3</bibl> <bibtext> F. Bakar and V. Kumar, The use of humour in teaching and learning in higher education classrooms: Lecturers' perspectives, J. Engl. Acad. Purp. 40 (2019), 15 – 25. https://doi.org/10.1016/j.jeap.2019.04.006.</bibtext> </blist> <blist> <bibl id="bib4" idref="ref73" type="bt">4</bibl> <bibtext> S. Béland, D. Cousineau, and N. Loye, Utiliser le Coefficient Omega de McDonald à la Place de l'Alpha de Cronbach, McGill J. Educ. 52 (2017), no. 3, 791 – 804. https://doi.org/10.7202/1050915ar.</bibtext> </blist> <blist> <bibl id="bib5" idref="ref7" type="bt">5</bibl> <bibtext> J. Benson, Structural components of statistical test anxiety in adults: An exploratory model, J. Exp. Educ. 57 (1989), no. 3, 247 – 261. https://doi.org/10.1080/00220973.1989.10806509.</bibtext> </blist> <blist> <bibl id="bib6" type="bt">6</bibl> <bibtext> M. Birenbaum and S. Eylath, Who is afraid of statistics? Correlates of statistics anxiety among students of educational sciences, Educ. Res. 36 (1994), no. 1, 93 – 98. https://doi.org/10.1080/0013188940360110.</bibtext> </blist> <blist> <bibl id="bib7" idref="ref80" type="bt">7</bibl> <bibtext> W. J. Boone, Rasch analysis for instrument development: Why, when, and how? CBE Life Sci. Educ. 15 (2016), no. 4, 1 – 7. https://doi.org/10.1187/cbe.16-04-0148.</bibtext> </blist> <blist> <bibl id="bib8" idref="ref83" type="bt">8</bibl> <bibtext> V. Braun and V. Clarke, " Thematic analysis ," APA handbook of research methods in psychology, Vol. 2. Research designs: Quantitative, qualitative, neuropsychological, and biological, H. Cooper, P. M. Camic, D. L. Long, A. T. Panter, D. Rindskopf, and K. J. Sher (eds.), American Psychological Association, Washington, DC, 2012, pp. 57 – 71.</bibtext> </blist> <blist> <bibl id="bib9" idref="ref26" type="bt">9</bibl> <bibtext> P. Burckhardt, R. Nugent, and C. R. Genovese, Teaching statistical concepts and modern data analysis with a computing‐integrated learning environment, J. Stat. Data Sci. Educ. 29 (2021), no. sup1, S61 – S73. https://doi.org/10.1080/10691898.2020.1854637.</bibtext> </blist> <blist> <bibtext> M. Cantinotti, D. Lalande, M.‐A. Ferlatte, and D. Cousineau, Validation de la Version Francophone du Questionnaire d'Anxiété Statistique (SAS‐F‐24), Can. J. Behav. Sci. 49 (2017), no. 2, 133 – 142. https://doi.org/10.1037/cbs0000074.</bibtext> </blist> <blist> <bibtext> S. K. Carpenter, Testing enhances the transfer of learning, Curr. Dir. Psychol. Sci. 21 (2012), no. 5, 279 – 283. https://doi.org/10.1177/0963721412452728.</bibtext> </blist> <blist> <bibtext> R. Carver, M. Everson, J. Gabrosek, N. Horton, R. Lock, M. Mocko, A. Rossman, G. H. Roswell, P. Velleman, J. Witmer, and B. Wood, Guidelines for assessment and instruction in statistics education (GAISE). College Report 2016, American Statistical Association, Alexandria, VA, 2016.</bibtext> </blist> <blist> <bibtext> B. L. Chance, Components of statistical thinking and implications for instruction and assessment, J. Stat. Educ. 10 (2002), no. 3, 1 – 14. https://doi.org/10.1080/10691898.2002.11910677.</bibtext> </blist> <blist> <bibtext> K. M. T. Collins and A. J. Onwuegbuzie, I cannot read my statistics textbook: The relationship between reading ability and statistics anxiety, J. Negro Educ. 76 (2007), no. 2, 118 – 129. https://<ulink href="http://www.jstor.org/stable/40034551">www.jstor.org/stable/40034551</ulink>.</bibtext> </blist> <blist> <bibtext> D. Cousineau and B. Harding, Pourquoi les Statistiques sont‐elles Difficiles à Enseigner et à Comprendre? Quelques Réflexions, Rev. Psychoeduc. 46 (2017), no. 2, 397 – 419. https://doi.org/10.7202/1042257ar.</bibtext> </blist> <blist> <bibtext> L. J. Cronbach, Coefficient alpha and the internal structure of tests, Psychometrika 16 (1951), no. 3, 297 – 334. https://doi.org/10.1007/BF02310555.</bibtext> </blist> <blist> <bibtext> T. Deane, K. Nomme, E. Jeffery, C. Pollock, and G. Birol, Development of the statistical reasoning in biology concept inventory (SRBCI), CBE Life Sci. Educ. 15 (2016), 1 – 13. https://doi.org/10.1187/cbe.15-06-0131.</bibtext> </blist> <blist> <bibtext> R. delMas, J. Garfield, A. Ooms, and B. Chance, Assessing students' conceptual understanding after a first course in statistics, Stat. Educ. Res. J. 6 (2007), no. 2, 28 – 58. <ulink href="http://www.stat.auckland.ac.nz/serj">http://www.stat.auckland.ac.nz/serj</ulink>.</bibtext> </blist> <blist> <bibtext> R. delMas, Statistical literacy, reasoning, and thinking: A commentary, J. Stat. Educ. 10 (2002), no. 2, 1 – 11. https://doi.org/10.1080/10691898.2002.11910674.</bibtext> </blist> <blist> <bibtext> T. A. DeVaney, Anxiety and attitude of graduate students in on‐campus vs. online statistics courses, J. Stat. Educ. 18 (2010), no. 1, 1 – 15. https://doi.org/10.1080/10691898.2010.11889472.</bibtext> </blist> <blist> <bibtext> C. Franklin, G. Kader, D. Mewborn, J. Moreno, R. Peck, M. Perry, and R. Scheaffer, Guidelines for assessment and instruction in statistics education (GAISE) report: a pre‐K‐12 curriculum framework, American Statistical Association, Alexandria, VA, 2007.</bibtext> </blist> <blist> <bibtext> J. Garfield and B. Chance, Assessment in statistics education: Issues and challenges, Math. Think. Learn. 2 (2000), no. 1–2, 99 – 125. https://doi.org/10.1207/S15327833MTL0202_5.</bibtext> </blist> <blist> <bibtext> J. Garfield and I. Gal, " Teaching and assessing statistical reasoning ," Developing mathematical reasoning in grades K‐12, L. Stiff and F. R. Curcio (eds.), National Council of Teachers of Mathematics, Reston, VA, 1999, pp. 1 – 17.</bibtext> </blist> <blist> <bibtext> J. Garfield, Beyond testing and grading: Using assessment to improve student learning, J. Stat. Educ. 2 (1994), no. 1, 1 – 10. https://doi.org/10.1080/10691898.1994.11910462.</bibtext> </blist> <blist> <bibtext> J. Garfield, Assessing statistical reasoning, Stat. Educ. Res. J. 2 (2003), no. 1, 22 – 38. https://doi.org/10.52041/serj.v2i1.557.</bibtext> </blist> <blist> <bibtext> J. Garfield and D. Ben‐Zvi, Developing students' statistical reasoning: Connecting research and teaching practice, Springer, New York, 2008.</bibtext> </blist> <blist> <bibtext> R.‐M. Gibeau, E. A. Maloney, S. Béland, D. Lalande, M. Cantinotti, A. Williot, L. Chanquoy, J. Simon, M.‐A. Boislard‐Pépin, and D. Cousineau, The correlates of statistics anxiety: Relationships with spatial anxiety, mathematics anxiety and gender, J. Numer. Cogn. 9 (2023), no. 1, 16 – 43. https://doi.org/10.5964/jnc.8199.</bibtext> </blist> <blist> <bibtext> G. Gigerenzer, How to make cognitive illusions disappear: Beyond "heuristics and biases ", Eur. Rev. Soc. Psychol. 2 (1991), no. 1, 83 – 115. https://doi.org/10.1080/14792779143000033.</bibtext> </blist> <blist> <bibtext> H. Harth. An analysis of the teaching of introductory statistics at university in "context," Doctoral dissertation, Loughborough University. 2018.</bibtext> </blist> <blist> <bibtext> P. J. Holmes. The effect of motivation and attitude towards statistics on conceptual understanding of statistics, Doctoral dissertation, University of Georgia. 2014.</bibtext> </blist> <blist> <bibtext> R. Hubbard, Assessment and the process of learning statistics, J. Stat. Educ. 5 (1997), no. 1, 1 – 9. https://doi.org/10.1080/10691898.1997.11910522.</bibtext> </blist> <blist> <bibtext> T. Husereau, D. Cousineau, S. Zang, and R.‐M. Gibeau, A compendium of common heuristics, misconceptions, and biased reasoning used in statistical thinking, Quant. Methods Psychol. 20 (2024), no. 1, 57 – 75. https://doi.org/10.20982/tqmp.20.1.p057.</bibtext> </blist> <blist> <bibtext> J. L. Jensen, M. A. McDaniel, S. M. Woodard, and T. A. Kummer, Teaching to the test ... or testing to teach: Exams requiring higher order thinking skills encourage greater conceptual understanding, Educ. Psychol. Rev. 26 (2014), no. 2, 307 – 329. https://doi.org/10.1007/s10648-013-9248-9.</bibtext> </blist> <blist> <bibtext> X. Jia, W. Hu, F. Cai, H. Wang, J. Li, M. A. Runco, and Y. Chen, The influence of teaching methods on creative problem finding, Think. Skills Creat. 24 (2017), 86 – 94. https://doi.org/10.1016/j.tsc.2017.02.006.</bibtext> </blist> <blist> <bibtext> D. Kahneman and A. Tversky, Subjective probability: A judgment of representativeness, Cogn. Psychol. 3 (1972), no. 3, 430 – 454. https://doi.org/10.1016/0010-0285(72)90016-3.</bibtext> </blist> <blist> <bibtext> R. Konicek‐Moran and P. Keeley, Teaching for conceptual understanding in science, National Science Teachers Association, Richmond, VA, 2015.</bibtext> </blist> <blist> <bibtext> C. Konold, Informal conceptions of probability, Cogn. Instr. 6 (1989), no. 1, 59 – 98. https://doi.org/10.1207/s1532690xci0601_3.</bibtext> </blist> <blist> <bibtext> C. Konold, Issues in assessing conceptual understanding in probability and statistics, J. Stat. Educ. 3 (1995), no. 1, 1 – 9. https://doi.org/10.1080/10691898.1995.11910479.</bibtext> </blist> <blist> <bibtext> S. M. Marson, Three empirical strategies for teaching statistics, J. Teach. Soc. Work. 27 (2007), no. 3–4, 199 – 213. https://doi.org/10.1300/J067v27n03_13.</bibtext> </blist> <blist> <bibtext> M. M. McIntyre, Statistics is puzzling: Testing a novel approach to statistics learning, Scholarsh. Teach. Learn. Psychol. 9 (2020), no. 2, 1 – 9. https://doi.org/10.1037/stl0000204.</bibtext> </blist> <blist> <bibtext> P. Mcmahon, T. Zhang, and R. Dwight, Requirements for big data adoption for railway asset management, IEEE Access 8 (2020), 15543 – 15564. https://doi.org/10.1109/ACCESS.2020.2967436.</bibtext> </blist> <blist> <bibtext> M. M. Meletiou. Developing students' conceptions of variation: An untapped well in statistical reasoning, Doctoral dissertation, University of Texas. 2000.</bibtext> </blist> <blist> <bibtext> D. Nolan and D. Temple Lang, Computing in the statistics curricula, Am. Stat. 64 (2010), no. 2, 97 – 107. https://doi.org/10.1198/tast.2010.09132.</bibtext> </blist> <blist> <bibtext> A. J. Onwuegbuzie and M. A. Seaman, The effect of time constraints and statistics test anxiety on test performance in a statistics course, J. Exp. Educ. 63 (1995), no. 2, 115 – 124. https://doi.org/10.1080/00220973.1995.9943816.</bibtext> </blist> <blist> <bibtext> R. Park, Practical teaching strategies for hypothesis testing, Am. Stat. 73 (2019), no. 3, 282 – 287. https://doi.org/10.1080/00031305.2018.1424034.</bibtext> </blist> <blist> <bibtext> G. Rasch, Probabilistic models for some intelligence and attainment tests, MESA Press, San Diego, CA, 1993.</bibtext> </blist> <blist> <bibtext> B. Rittle‐Johnson, R. S. Siegler, and M. W. Alibali, Developing conceptual understanding and procedural skill in mathematics: An iterative process, J. Educ. Psychol. 93 (2001), no. 2, 346 – 362. https://doi.org/10.1037/0022-0663.93.2.346.</bibtext> </blist> <blist> <bibtext> A. Robitzsch, T. Kiefer, and M. Wu. TAM: Test analysis modules [R package version 4.2‐21]. 2024 https://CRAN.R-project.org/package=TAM.</bibtext> </blist> <blist> <bibtext> A. Sabbag, J. Garfield, and A. Zieffler, Assessing statistical literacy and statistical reasoning: The REALI instrument, Stat. Educ. Res. J. 17 (2018), no. 2, 141 – 160. https://doi.org/10.52041/serj.v17i2.163.</bibtext> </blist> <blist> <bibtext> S. Sharma, Definitions and models of statistical literacy: A literature review, Open Rev. Educ. Res. 4 (2017), no. 1, 118 – 133. https://doi.org/10.1080/23265507.2017.1354313.</bibtext> </blist> <blist> <bibtext> M. J. Shaughnessy, Statistics for all—The flip side of quantitative reasoning, National Council of Teachers of Mathematics, Reston, VA, 2010. https://<ulink href="http://www.nctm.org/News‐and‐Calendar/Messages‐from‐the‐President/Archive/J%5f‐Michael‐Shaughnessy/Statistics‐for‐All—the‐Flip‐Side‐of‐Quantitative‐Reasoning/">www.nctm.org/News‐and‐Calendar/Messages‐from‐the‐President/Archive/J%5f‐Michael‐Shaughnessy/Statistics‐for‐All—the‐Flip‐Side‐of‐Quantitative‐Reasoning/</ulink>.</bibtext> </blist> <blist> <bibtext> K. Slootmaeckers, B. Kerremans, and J. Adriaensen, Too afraid to learn: Attitudes towards statistics as a barrier to learning statistics and to acquiring quantitative skills, Politics 34 (2014), no. 2, 191 – 200. https://doi.org/10.1111/1467-9256.12042.</bibtext> </blist> <blist> <bibtext> M. Stuart, Changing the teaching of statistics, J. R. Stat. Soc. Ser. D Stat. 44 (1995), no. 1, 45 – 54. https://doi.org/10.2307/2348615.</bibtext> </blist> <blist> <bibtext> L. Theis and A. Savard, Linking probability to real‐world situations: How do teachers make use of the mathematical potential of simulation programs? [Contributed paper refereed], Eighth International Conference on Teaching Statistics (ICOTS8), Ljubljana, Slovenia, 2010.</bibtext> </blist> <blist> <bibtext> A. Tversky and D. Kahneman, Judgment under uncertainty: Heuristics and biases, Science 185 (1974), no. 4157, 1124 – 1131. https://<ulink href="http://www.jstor.org/stable/1738360">www.jstor.org/stable/1738360</ulink>.</bibtext> </blist> <blist> <bibtext> A. Vigil‐Colet, U. Lorenzo‐Seva, and L. Condon, Development and validation of the statistical anxiety scale, Psicothema 20 (2008), no. 1, 174 – 180. https://reunido.uniovi.es/index.php/PST/article/view/8638.</bibtext> </blist> <blist> <bibtext> Á. F. Villarejo‐Ramos, J.‐P. Cabrera‐Sánchez, J. Lara‐Rubio, and F. Liébana‐Cabanillas, Predicting big data adoption in companies with an explanatory and predictive model, Front. Psychol. 12 (2021), 1 – 12. https://doi.org/10.3389/fpsyg.2021.651398.</bibtext> </blist> <blist> <bibtext> J. Watson and R. Callingham, Statistical literacy: A complex hierarchical construct, Stat. Educ. Res. J. 2 (2003), no. 2, 3 – 46. https://doi.org/10.52041/serj.v2i2.553.</bibtext> </blist> <blist> <bibtext> M. C. Wittmann, J. T. Morgan, and R. E. Feeley, Laboratory‐tutorial activities for teaching probability, Phys. Rev. Spec. Top. Phys. Educ. Res. 2 (2006), 1 – 8. https://doi.org/10.1103/PhysRevSTPER.2.020104.</bibtext> </blist> </ref> <aug> <p>By R.‐M. Gibeau and D. Cousineau</p> <p>Reported by Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib41" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib57" firstref="ref2"></nolink> <nolink nlid="nl3" bibid="bib51" firstref="ref3"></nolink> <nolink nlid="nl4" bibid="bib14" firstref="ref4"></nolink> <nolink nlid="nl5" bibid="bib26" firstref="ref5"></nolink> <nolink nlid="nl6" bibid="bib40" firstref="ref6"></nolink> <nolink nlid="nl7" bibid="bib19" firstref="ref8"></nolink> <nolink nlid="nl8" bibid="bib24" firstref="ref9"></nolink> <nolink nlid="nl9" bibid="bib44" firstref="ref10"></nolink> <nolink nlid="nl10" bibid="bib29" firstref="ref11"></nolink> <nolink nlid="nl11" bibid="bib34" firstref="ref12"></nolink> <nolink nlid="nl12" bibid="bib20" firstref="ref14"></nolink> <nolink nlid="nl13" bibid="bib45" firstref="ref15"></nolink> <nolink nlid="nl14" bibid="bib54" firstref="ref16"></nolink> <nolink nlid="nl15" bibid="bib59" firstref="ref17"></nolink> <nolink nlid="nl16" bibid="bib35" firstref="ref18"></nolink> <nolink nlid="nl17" bibid="bib55" firstref="ref19"></nolink> <nolink nlid="nl18" bibid="bib37" firstref="ref20"></nolink> <nolink nlid="nl19" bibid="bib32" firstref="ref21"></nolink> <nolink nlid="nl20" bibid="bib38" firstref="ref23"></nolink> <nolink nlid="nl21" bibid="bib28" firstref="ref24"></nolink> <nolink nlid="nl22" bibid="bib47" firstref="ref25"></nolink> <nolink nlid="nl23" bibid="bib12" firstref="ref27"></nolink> <nolink nlid="nl24" bibid="bib43" firstref="ref28"></nolink> <nolink nlid="nl25" bibid="bib36" firstref="ref29"></nolink> <nolink nlid="nl26" bibid="bib11" firstref="ref33"></nolink> <nolink nlid="nl27" bibid="bib33" firstref="ref34"></nolink> <nolink nlid="nl28" bibid="bib22" firstref="ref36"></nolink> <nolink nlid="nl29" bibid="bib10" firstref="ref37"></nolink> <nolink nlid="nl30" bibid="bib27" firstref="ref38"></nolink> <nolink nlid="nl31" bibid="bib56" firstref="ref39"></nolink> <nolink nlid="nl32" bibid="bib31" firstref="ref44"></nolink> <nolink nlid="nl33" bibid="bib30" firstref="ref47"></nolink> <nolink nlid="nl34" bibid="bib50" firstref="ref48"></nolink> <nolink nlid="nl35" bibid="bib23" firstref="ref50"></nolink> <nolink nlid="nl36" bibid="bib13" firstref="ref52"></nolink> <nolink nlid="nl37" bibid="bib39" firstref="ref57"></nolink> <nolink nlid="nl38" bibid="bib25" firstref="ref58"></nolink> <nolink nlid="nl39" bibid="bib18" firstref="ref60"></nolink> <nolink nlid="nl40" bibid="bib17" firstref="ref61"></nolink> <nolink nlid="nl41" bibid="bib49" firstref="ref62"></nolink> <nolink nlid="nl42" bibid="bib58" firstref="ref63"></nolink> <nolink nlid="nl43" bibid="bib16" firstref="ref74"></nolink> <nolink nlid="nl44" bibid="bib46" firstref="ref81"></nolink> <nolink nlid="nl45" bibid="bib48" firstref="ref82"></nolink> <nolink nlid="nl46" bibid="bib21" firstref="ref86"></nolink> <nolink nlid="nl47" bibid="bib53" firstref="ref87"></nolink> <nolink nlid="nl48" bibid="bib42" firstref="ref90"></nolink>
Header DbId: eric
DbLabel: ERIC
An: EJ1483737
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Uncovering Key Statistical Concepts from an Investigation of Four Standardized Tests
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22R%2E-M%2E+Gibeau%22">R.-M. Gibeau</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-3649-4626">0000-0002-3649-4626</externalLink>)<br /><searchLink fieldCode="AR" term="%22D%2E+Cousineau%22">D. Cousineau</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Teaching+Statistics%3A+An+International+Journal+for+Teachers%22"><i>Teaching Statistics: An International Journal for Teachers</i></searchLink>. 2025 47(3):219-234.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 16
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2025
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Audience
  Label: Education Level
  Group: Audnce
  Data: <searchLink fieldCode="EL" term="%22Higher+Education%22">Higher Education</searchLink><br /><searchLink fieldCode="EL" term="%22Postsecondary+Education%22">Postsecondary Education</searchLink>
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Statistics+Education%22">Statistics Education</searchLink><br /><searchLink fieldCode="DE" term="%22Student+Evaluation%22">Student Evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Standardized+Tests%22">Standardized Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Mathematics+Tests%22">Mathematics Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Mathematical+Concepts%22">Mathematical Concepts</searchLink><br /><searchLink fieldCode="DE" term="%22Scores%22">Scores</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Items%22">Test Items</searchLink><br /><searchLink fieldCode="DE" term="%22Accuracy%22">Accuracy</searchLink><br /><searchLink fieldCode="DE" term="%22Teaching+Methods%22">Teaching Methods</searchLink><br /><searchLink fieldCode="DE" term="%22Learning+Objectives%22">Learning Objectives</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluation+Methods%22">Evaluation Methods</searchLink><br /><searchLink fieldCode="DE" term="%22Higher+Education%22">Higher Education</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1111/test.12410
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0141-982X<br />1467-9639
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: To this date, few standardized tests measuring students' performance with regards to statistics exist. Only four tests have been proposed for college or university students. The goal of the present study is to investigate these tests. University professors or instructors experienced in teaching statistics were asked to list the concepts they think are being assessed by each item of the tests. A total of 708 responses were obtained from 42 participants. Using thematic analysis, 18 fundamental statistical concepts were identified. Unplanned analyses on participants' scores were also computed. The results suggest that there is no consensus on some of the items' correct answers. The study has practical implications for teaching statistics, from learning goals to assessment methods.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: Note
  Label: Notes
  Group: Note
  Data: https://osf.io/t6w7n
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2025
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1483737
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1483737
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1111/test.12410
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 16
        StartPage: 219
    Subjects:
      – SubjectFull: Statistics Education
        Type: general
      – SubjectFull: Student Evaluation
        Type: general
      – SubjectFull: Standardized Tests
        Type: general
      – SubjectFull: Mathematics Tests
        Type: general
      – SubjectFull: Mathematical Concepts
        Type: general
      – SubjectFull: Scores
        Type: general
      – SubjectFull: Test Items
        Type: general
      – SubjectFull: Accuracy
        Type: general
      – SubjectFull: Teaching Methods
        Type: general
      – SubjectFull: Learning Objectives
        Type: general
      – SubjectFull: Evaluation Methods
        Type: general
      – SubjectFull: Higher Education
        Type: general
    Titles:
      – TitleFull: Uncovering Key Statistical Concepts from an Investigation of Four Standardized Tests
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: R.-M. Gibeau
      – PersonEntity:
          Name:
            NameFull: D. Cousineau
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 07
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 0141-982X
            – Type: issn-electronic
              Value: 1467-9639
          Numbering:
            – Type: volume
              Value: 47
            – Type: issue
              Value: 3
          Titles:
            – TitleFull: Teaching Statistics: An International Journal for Teachers
              Type: main
ResultId 1