Group-Difference Effect Sizes: Gauging the Practical Importance of Findings from Group-Experimental Research

Saved in:
Bibliographic Details
Title: Group-Difference Effect Sizes: Gauging the Practical Importance of Findings from Group-Experimental Research
Language: English
Authors: Cook, Bryan G. (ORCID 0000-0001-9294-0873), Cook, Lysandra, Therrien, William J. (ORCID 0000-0003-0594-5129)
Source: Learning Disabilities Research & Practice. May 2018 33(2):56-63.
Availability: Wiley-Blackwell. 350 Main Street, Malden, MA 02148. Tel: 800-835-6770; Tel: 781-388-8598; Fax: 781-388-8232; e-mail: cs-journals@wiley.com; Web site: http://www.wiley.com/WileyCDA
Peer Reviewed: Y
Page Count: 8
Publication Date: 2018
Document Type: Journal Articles
Reports - Descriptive
Descriptors: Effect Size, Learning Disabilities, Evaluation Methods, Groups, Research Methodology, Group Structure, Context Effect
DOI: 10.1111/ldrp.12167
ISSN: 0938-8982
Abstract: Effect sizes are powerful tools for evaluating the practical importance of study findings that should be considered in the context of study characteristics such as participants, dependent variables, and comparison condition. In this article, we discuss how group-difference effect sizes are used to gauge the practical importance of group experimental studies. We first define different types of group-difference effect sizes and discuss how they can provide valuable information for research consumers. Second, we present guidelines for interpreting group-difference effect sizes. Third, we discuss important contextual variables that should be taken into account when interpreting group-difference effect sizes reported in the literature. Last, we provide two examples of how group-difference effect sizes have been used in the learning disabilities research base.
Abstractor: As Provided
Entry Date: 2018
Accession Number: EJ1178972
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwHIBpeoDHFcqm4Lc5T8eRqfAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDHmZJX52x_vosc8swAIBEICBm5WUxC6D-hoktA451HhGBVhY6x1AdBE2cAPWmxI7xP4xpYrSw5zWvEHaBPlXbG8x0epiFG5y0WdUv5dVSt_H157i-aAQ8VGxHVdVaR4xHcDuKKhQMo35R2zNRhI2Iqbyvgav82n1sFM4RLepR5yl3PWjoWlCD14Xg4yIlPOt2T4XrB65Jt-Jp4mFbg1gtduOdDcVAUB0xNWBha2k
Text:
  Availability: 1
  Value: <anid>AN0129573248;7mj01may.18;2018May14.10:29;v2.2.500</anid> <title id="AN0129573248-1">Group‐Difference Effect Sizes: Gauging the Practical Importance of Findings from Group‐Experimental Research </title> <p>Abstract: Effect sizes are powerful tools for evaluating the practical importance of study findings that should be considered in the context of study characteristics such as participants, dependent variables, and comparison condition. In this article, we discuss how group‐difference effect sizes are used to gauge the practical importance of group experimental studies. We first define different types of group‐difference effect sizes and discuss how they can provide valuable information for research consumers. Second, we present guidelines for interpreting group‐difference effect sizes. Third, we discuss important contextual variables that should be taken into account when interpreting group‐difference effect sizes reported in the literature. Last, we provide two examples of how group‐difference effect sizes have been used in the learning disabilities research base.</p> <p>Ms. Perez, a resource teacher at Jefferson Elementary School, was pleased when she heard that her principal had scheduled an in‐service training on effective teaching practices based onJohn Hattie's ([<reflink idref="bib9" id="ref1">9</reflink>] ) Visible Learning. Jefferson Elementary School was a large school with a mix of new and experienced teachers who often had differing opinions on what instructional approaches were most effective and should be used. She had heard Hattie's work discussed before, and knew that he determined the effectiveness of instructional approaches based on the results of hundreds of meta‐analyses that included tens of thousands of research studies. She was hopeful that she and her colleagues could learn about the effectiveness of different instructional approaches from a definitive, research‐based source. At the workshop, the presenter talked a lot about effect sizes. For example, the presenter said that ability grouping had an effect size of 0.12, which seemed to be treated as a small effect. The presenter spent more time talking about approaches with larger effect sizes, like feedback (effect size = 0.75). As the presenter went on to discuss this and other practices with similar effect sizes, Ms. Perez couldn't help but wonder, what do effect sizes of 0.12 and 0.75 mean? Does that mean she should expect her students to improve 12 percent if she used ability grouping, or that 75 percent of students would be successful if she used feedback? Without understanding what effect sizes mean, Ms. Perez wasn't sure which if any of the interventions were worth using.</p> <p>In this article, we provide a practical overview of effect sizes (ESs) in group‐experimental research. As the term suggests, ESs represent the size of the effect or the strength of the relation between variables in a research study (see Table for a definition and example of key terms, such as effect size). Although p values derived from null hypothesis significance testing (e.g., p < .05) traditionally have been used to interpret findings of statistical analyses in research studies (Travers, Cook, & Cook, [<reflink idref="bib16" id="ref2">16</reflink>] ), ESs should also be considered because they indicate the practical importance of research findings in a way that p values cannot. Although researchers commonly reported p values without providing ESs in the past, reporting ESs is now recommended in the fields of psychology (Wilkinson & The Task Force on Statistical Inference, [<reflink idref="bib19" id="ref3">19</reflink>] ) and education (American Educational Research Association, [<reflink idref="bib1" id="ref4">1</reflink>] ), and is an indicator of high‐quality group‐experimental research in special education (Cook et al., [<reflink idref="bib7" id="ref5">7</reflink>] ; Gersten et al., [<reflink idref="bib8" id="ref6">8</reflink>] ). Our take‐home message in this article is that ESs enable research consumers to evaluate the practical importance of study findings when considered appropriately in the context of study characteristics such as participants, dependent variables, and comparison condition.</p> <p>Definitions and Examples of Key Terms</p> <p> <ephtml> <table border="1" cellpadding="3"><tr><th>Term</th><th>Definition</th><th>Example</th></tr><tr><td>Cohen's d</td><td>One of the group‐difference effect sizes, Cohen's d is the most popular effect size used in the special education literature.</td><td>The researchers reported a Cohen's d of 0.95, indicating that the intervention had a large effect on student outcomes.</td></tr><tr><td>Effect size</td><td>A measure in quantitative research indicating the size or magnitude of the effect of an intervention or strength of the relation between variables.</td><td>Many professional organizations recommend reporting effect sizes for quantitative analyses.</td></tr><tr><td>Group‐difference effect size</td><td>A type of effect size reflecting the difference between groups. Group‐difference effect sizes, such as Cohen's d and Hedges’ g, report the difference between groups in SD units.</td><td>To indicate the magnitude of effect in a group‐experimental study, researchers should report a group‐difference effect size.</td></tr><tr><td>Hedges’ g</td><td>One of the group‐difference effect sizes, Hedges’ g corrects for the slight overestimation of effects in Cohen's d.</td><td>The researchers reported a Hedges’ g of 0.93, indicating that the intervention had a large effect on student outcomes.</td></tr><tr><td>Meta‐analysis</td><td>A review of research that averages effect sizes across multiple studies.</td><td>In their meta‐analysis of 17 group‐experimental studies, the authors reported an overall effect size of 0.42 for the intervention.</td></tr><tr><td>Standard deviation</td><td>A measure of spread or variability. Conceptually, the SD is the typical difference of scores from the mean average. Group‐difference effect sizes are reported in SD units.</td><td>An effect size of 0.42 means that the intervention resulted in average improvement of 0.42 SDs.</td></tr><tr><td>Statistically significant</td><td>Results of a statistical analysis are typically considered statistically significant when the p value is <.05 (i.e., when the probability that study findings occurred when the null hypothesis is true is less than 5 percent). Statistical significance is influenced by sample size.</td><td>Findings from the experiment were statistically significant (p < .05) even though the effect size was small because of the large number of participants in the study.</td></tr></table> </ephtml> </p> <p>Effect size is a generic term, and a variety of ESs are available that express the size of effect or relation in different types of research studies. For example, ESs can be used to indicate the strength of association between variables in correlational research (e.g., r, β, odds ratios). Additionally, a variety of ESs have been developed to indicate the effect of an intervention in single‐case design research (e.g., Tau‐U). In this article, we focus on group‐difference ESs, which indicate the size or magnitude of differences between groups. Group‐difference ESs address the question, “How much of a difference is there between groups?” For the sake of clarity, we restrict our discussion and examples to group‐difference ESs used to indicate the size of the effect in group‐experimental research. Rather than wade into the technical details of calculating group‐difference ESs, our focus here is on an introductory, nontechnical discussion aimed at aiding the reader's understanding of the practical implications of group‐experimental research. Specifically, in this article we: (<reflink idref="bib1" id="ref7">1</reflink>) discuss what group‐difference ESs are, (<reflink idref="bib2" id="ref8">2</reflink>) consider why group‐difference ESs provide important information above and beyond p values, (<reflink idref="bib3" id="ref9">3</reflink>) present guidelines for interpreting group‐difference ESs, (<reflink idref="bib4" id="ref10">4</reflink>) note contextual factors to consider when interpreting group‐difference ESs, and (<reflink idref="bib5" id="ref11">5</reflink>) discuss two examples of how group‐difference ESs have been used in the LD research base.</p> <hd id="AN0129573248-2">WHAT ARE GROUP‐DIFFERENCE EFFECT SIZES?</hd> <p>Group‐difference ESs simply represent the difference between the average performance of groups, reported using a standardized metric based on the variability or spread of the outcome. In the following sections, we break down these two aspects of group‐difference ESs—(<reflink idref="bib1" id="ref12">1</reflink>) the difference between groups and (<reflink idref="bib2" id="ref13">2</reflink>) standardizing ESs into SD units—and then discuss (<reflink idref="bib3" id="ref14">3</reflink>) how group‐difference ESs are reported and (<reflink idref="bib4" id="ref15">4</reflink>) related group‐difference ESs such as Cohen's d and Hedges’ g.</p> <hd id="AN0129573248-3">The Difference between Groups</hd> <p>At the core of group‐difference ESs is a simple equation: the mean (M) of Group 1 (e.g., the experimental or treatment group) – M of Group 2 (e.g., a comparison group). In classic group experiments, researchers randomly assign half a group of study participants to a treatment group, who receive the experimental treatment (e.g., graphic organizers). The other half of the participants is a comparison group who receive no instruction (i.e., a control condition) or a comparison intervention. Although not always done in experimental research, researchers often administer a pretest to both groups (e.g., a social studies unit test). Then, both groups receive instruction (e.g., the treatment group with graphic organizers, a control group without graphic organizers). At the end of instruction for the unit, both groups take a posttest (e.g., the same unit test; see Cook & Cook, [<reflink idref="bib6" id="ref16">6</reflink>] , for a more detailed discussion of experimental design). If the performance of participants in the treatment group on the unit test improved by 20 points and the control group improved on average by 15 points, the difference between the groups is, obviously enough, 5.</p> <hd id="AN0129573248-4">A Standardized Metric</hd> <p>Although knowing that the treatment and control groups differed by five points on their improvement on the social studies unit test can be meaningful if we know that test well, it is not very informative if we are not familiar with the outcome measure being used, or if we want to compare or combine ESs across studies that use different outcome measures. Therefore, when calculating group‐difference ESs, we convert the raw difference between groups to a common scale. This is done by dividing the difference by a measure of how much the scores vary or are spread out (i.e., the SD), in essence adjusting the difference between groups by the variability in performance. The SD is a statistical term indicating how much participants’ performance differs (or deviates) from the average. Assume students in the control group improved by an average of 17.5 points (15 points in the control group, 20 in the treatment group) on the social studies unit test in our example. But not everyone in the two groups improved exactly 15 and 20 points, respectively. Some improved by 22 points, others by 9 points; and although a couple of students improved by more than 30 points, a couple of other students did not improve at all, and one even did worse on the posttest. If, in our example, the participants’ improvement scores differed from the typical amount of improvement by 7 points on average, the group‐difference ES would be 5 (difference between groups)/7 (SD), or 0.71. In other words, students in the graphic organizer group improved by 0.71 SDs more than those in the control group. So, rather than report the actual or raw difference between groups, group‐difference ESs indicate the difference between groups in SD units. Reporting group‐difference ESs using a common, standardized metric (e.g., difference in SD units) has a number of benefits, including facilitating comparison of findings from different studies and synthesis of ESs across studies in meta‐analyses.</p> <hd id="AN0129573248-5">How Group‐Difference Effect Sizes are Reported</hd> <p>Although they can be reported to any number of decimals, group‐difference ESs are commonly reported to two decimal points (e.g., 0.71). Negative ESs, though relatively rare, occur when the comparison group outperforms the treatment group. An ES of 0 indicates that the treatment and control groups’ average performance was exactly the same. Therefore, ESs near 0 (e.g., −0.05, 0.07) indicate that there was little difference in the performance of the treatment and comparison groups. In general, the larger the ES, the larger and more practically important the effect. Conceptually, there is no upper boundary on ESs, though in practice it is relatively rare to see group‐difference ESs larger than 2.00. Guidelines for interpreting group‐difference ESs are discussed in the subsequent Interpreting Group‐Difference Effect Sizes section.</p> <hd id="AN0129573248-6">Different Group‐Difference Effect Sizes</hd> <p>Cohen's d, named after Jacob Cohen, is the most widely used ES for representing the difference between groups in a group‐experimental study. In Cohen's d, the SD is estimated using a somewhat complicated formula that, in essence, averages the SDs of the comparison and treatment groups. Cohen's d has been shown to slightly overestimate effects, especially in studies involving a small number of participants (Lakens, [<reflink idref="bib10" id="ref17">10</reflink>] ). Hedges’ g, named after Larry Hedges, is another group‐difference ES that was developed to correct for the slight tendency of Cohen's d to overestimate effects. Conceptually, Hedges’ g is the same as Cohen's d, in that it expresses the difference between groups in units of SD. The only difference between the two measures is that in Hedges’ g the SD is calculated in a slightly different way. Glass's Δ, named after Gene Glass, is another group‐difference ES that readers might see in the research literature. Glass's delta is very similar to Cohen's d and Hedges’ g, except that it uses the SD from just the control group, rather than calculating an average SD across groups. Because of their strong similarities, all these group‐difference ESs are interpreted using the same guidelines, which we discuss later in this article.</p> <hd id="AN0129573248-7">UNIQUE INFORMATION PROVIDED BY EFFECT SIZES</hd> <p>In this section, we discuss two primary benefits of considering ESs to interpret the findings of group‐experimental research: (<reflink idref="bib1" id="ref18">1</reflink>) they indicate the practical importance of study findings, which statistical significance testing does not; and (<reflink idref="bib2" id="ref19">2</reflink>) using a common, standardized metric helps stakeholders understand, compare, and combine study effects.</p> <hd id="AN0129573248-8">Effect Sizes Indicate Practical Importance</hd> <p>Researchers have traditionally relied on null hypothesis significance testing and p values when evaluating the effects of group experiments. In essence, p values indicate the probability that the null hypothesis (e.g., that graphic organizers do not have an effect on test scores) is true given study outcomes, but this provides little insight as to whether the intervention had a practically meaningful effect on participants (Travers et al., [<reflink idref="bib16" id="ref20">16</reflink>] ). Because sample size plays an important role in determining p values and statistical significance, it is frequently the case that relatively small, unimportant differences between groups end up being statistically significant (i.e., p < .05) in studies with large numbers of participants. For example, imagine a study reporting that an intervention resulted in the treatment group improving from an average of 70 words read correctly per minute to 76 words, whereas a control group improved from an average of 70 to 74 words. Most educators would not consider that to be a practically important difference that would justify investing time and resources into the intervention. Yet if the sample size was large enough (e.g., both groups included many hundreds of students), the results could be statistically significant. Conversely, large and practically important differences between groups may not result in statistically significant findings when a study involves a small number of participants (e.g., <20 per group)—something that is common in special education research, given the relatively small number of students with a particular disability available to participate in many research studies.</p> <p>Additionally, statistical significance is often presented as a yes or no dichotomy; that is, findings are often considered as either statistically significant or not. As such, it is difficult to discern degrees of effectiveness. In contrast, ESs represent the precise difference between groups without considering sample size, and therefore allow us to compare the size of effects across generally effective interventions. In this way, rather than just addressing whether a practice is or is not effective, group‐difference ESs can be used to answer “how effective is the practice?” and “which intervention is most effective?” Once we understand ESs, we can readily gauge the strength of an intervention's effect, regardless of whether findings are statistically significant. Accordingly, it is important not to rely solely on p values and statistical significance when interpreting research findings, but also to consider ESs to determine whether, and the degree to which, research findings are practically important.</p> <hd id="AN0129573248-9">Using a Common Metric Helps Interpretation, Comparison, and Syntheses</hd> <p>When one is familiar with the outcome measure and its scale, and is concerned only with results from one particular study (i.e., is not interested in comparing outcomes with findings from other studies, or combining the findings with the results of other studies), the raw difference between group means is sufficient, and there is no benefit to using a group‐difference ES. However, a raw‐difference score is not meaningful if one does not know the scale of the measure (a difference of 5 points between groups would likely be practically important on a 25‐point test, but much less so on a 250‐point exam). Similarly, a raw difference is not meaningful if one is not familiar with the type of metric used (e.g., a difference of 5 normal curve equivalents is not meaningful to most educators). Because group‐difference ESs convert all group difference scores to a common scale, one can interpret ESs without knowing the outcome measure or the metric used. As described in the following section, an ES of, for example, 1.00 is considered large regardless of whether the assessment contained few items, contained many items, or used a measurement scale that one has never heard of.</p> <p>Converting group differences to a common scale also facilitates comparison and synthesis. For example, imagine a teacher is comparing the effectiveness of Interventions A, B, and C for improving mathematics performance for students with LD, and reads studies showing that Intervention A resulted in an improvement of 6 normal curve equivalents on a test that is not described, Intervention B caused an improvement of 7 points on a standardized test of achievement, and Intervention C resulted in an improvement of 9 percentile points on the state proficiency test. Based on these raw difference scores, the teacher cannot meaningfully compare findings and discern which is the most effective intervention, because of the different types of outcome measures and scales. However, if researchers reported Cohen's ds of 0.41, 0.82, and 0.18 for Interventions A, B, and C, respectively, it is clear which is the most effective intervention, because all effects are reported on the same, standardized scale.</p> <p>Using a common scale also allows researchers to combine findings across different studies through meta‐analyses. Because no research study is perfect, it is important to examine the effectiveness of practices by considering findings across entire bodies of research studies. Meta‐analyses combine the findings of multiple studies by reporting an average ES across studies (Banda & Therrien, [<reflink idref="bib2" id="ref21">2</reflink>] ). Because the average effect derived from a body of research is more trustworthy than the findings of a single study, meta‐analyses have become a popular approach for determining what works in many fields, including special education. Calculating an overall ES across studies necessitates converting effects from individual studies to a common metric, such as Cohen's d.</p> <hd id="AN0129573248-10">INTERPRETING GROUP‐DIFFERENCE EFFECT SIZES</hd> <p>When interpreting group‐difference ESs, we encourage readers to use established benchmarks, translate ESs into percentile rank improvements, and consider study context.</p> <hd id="AN0129573248-11">Benchmarks for Interpreting Group‐Difference Effect Sizes</hd> <p>We have suggested that ESs are well suited for examining the practical importance of group‐experimental research. This raises the question of how large an ES should be before it is considered important. Cohen ([<reflink idref="bib5" id="ref22">5</reflink>] ) provided some benchmarks that are commonly used, though he noted that they should be used with caution and that other factors—such as characteristics of study participants, type of outcome measure, and comparison condition—should be considered when interpreting group‐difference ESs. Cohen's benchmarks for group‐difference ESs are that an ES of 0.20 is considered a small effect, an ES of 0.50 is medium, and an ES of 0.80 is large. Rosenthal ([<reflink idref="bib14" id="ref23">14</reflink>] ) added a benchmark of 1.30 for very large effects. Additionally, in their reviews of instructional interventions and programs in education, the What Works Clearinghouse ([<reflink idref="bib18" id="ref24">18</reflink>] ) considers a Cohen's d of 0.25 or greater as “substantively important” (p. 14).</p> <hd id="AN0129573248-12">Translating Group‐Difference Effect Sizes to Percentile Rank Improvements</hd> <p>Group‐difference ESs can be interpreted in other ways (Coe, [<reflink idref="bib4" id="ref25">4</reflink>] ). We find Cohen's ([<reflink idref="bib5" id="ref26">5</reflink>] ) U<subs>3</subs>, which indicates the improvement in percentile ranking associated with the treatment, particularly informative. Assuming that scores are normally distributed, and that the treatment and comparison groups started with the same scores at pretest, Cohen's U<subs>3</subs> indicates at what percentile ranking the average participant in the treatment group would be in the comparison group. For example, when Cohen's d is 0 (i.e., the intervention had no effect), the score of the participant who is at the 50<sups>th</sups> percentile in the treatment group would also be at the 50<sups>th</sups> percentile in the comparison group. When d = 0.50 (Cohen's benchmark for a medium ES), however, the 50<sups>th</sups> percentile in the treatment group is the equivalent of the 73<sups>rd</sups> percentile in the control group; in other words, the intervention caused an increase of 23 percentiles for the average participant in the treatment group. Table provides Cohen's U<subs>3</subs> values and interpretive benchmarks for a variety of Cohen's d values. However, savvy research consumers should also consider study context when interpreting ESs.</p> <p>Cohen's U3 Values and Interpretive Benchmarks for Cohen's d Values</p> <p> <ephtml> <table border="1" cellpadding="3"><tr><th>Cohen's d Value</th><th>Cohen's U3 Value*</th><th>Interpretive Benchmark</th></tr><tr><td>0</td><td>50.00</td><td /></tr><tr><td>0.10</td><td>53.98</td><td /></tr><tr><td>0.20</td><td>57.93</td><td>Small</td></tr><tr><td>0.30</td><td>61.79</td><td /></tr><tr><td>0.40</td><td>65.54</td><td /></tr><tr><td>0.50</td><td>69.15</td><td>Medium</td></tr><tr><td>0.60</td><td>72.57</td><td /></tr><tr><td>0.70</td><td>75.80</td><td /></tr><tr><td>0.80</td><td>78.81</td><td>Large</td></tr><tr><td>0.90</td><td>81.59</td><td /></tr><tr><td>1.00</td><td>84.13</td><td /></tr><tr><td>1.30</td><td>90.32</td><td>Very large</td></tr><tr><td>2.00</td><td>97.72</td><td /></tr></table> </ephtml> </p> <p>1 *Cohen's U3 indicates the percentile ranking in the control group for the average (50th percentile) score in the treatment group.</p> <hd id="AN0129573248-13">Considering Study Context</hd> <p>Although ESs allow us to gauge the practical importance of treatment effects, they do not give us license to do so without thinking critically. Simply examining the magnitude of a treatment effect alone does not provide sufficient information to make an informed decision about the practical importance of an intervention, or about whether we should implement a practice with a particular group of students. Context is the key. We need to ask ourselves: What were the study characteristics that played a role in determining the ESs? Important characteristics to take into consideration include the study participants, the dependent variables, and the comparison condition (Therrien, Zaman, & Banda, [<reflink idref="bib15" id="ref27">15</reflink>] ).</p> <p>No matter what the intervention or the outcome variable is, the characteristics of study participants likely influence the magnitude of the ES. In other words, some students are likely to learn a task or skill more easily, and some are more likely to have a more difficult time. Student age is a prime example of a student characteristic that can affect ES magnitude. Take reading, for example; there is typically tremendous growth in reading fluency in the lower grades, but very little growth typically occurs in high school. In fact, students in kindergarten and 1<sups>st</sups> grade typically make a yearly ES gain of 1.52, whereas the average ES in reading for students in 11<sups>th</sups> and 12<sups>th</sups> grades is only 0.06 (Lipsey et al., [<reflink idref="bib11" id="ref28">11</reflink>] ). If we only interpreted these ESs using Cohen's guidelines, we could be making a critical mistake. An intervention in kindergarten that results in an ES of 0.80 (a large effect by Cohen's standards) would underperform the gains we would expect to see at this age, whereas an ES of 0.30 (a small effect by Cohen's standards) in 12<sups>th</sups> grade is much larger than the typical gain made at this age.</p> <p>Disability status is another important participant characteristic to take into consideration. Compared to typically achieving students, students with learning disabilities (LD), for example, often need more intensive interventions to make comparable improvements in achievement (Vaughn & Wanzek, [<reflink idref="bib17" id="ref29">17</reflink>] ). We need to consider this reality when examining ES magnitude. An intervention might have an ES increase of 0.80 for students without disabilities, but, without increased intensity, only an ES of 0.25 for students with LD. We therefore would be remiss if we did not consider disability status (e.g., LD vs. no LD) when evaluating the relative magnitude of an ES. When interpreting ES magnitude in a particular study, then, we need to consider for whom the practice was found to be effective, and juxtapose this information with the characteristics of the participants (Mathews, Hirsch, & Therrien, [<reflink idref="bib13" id="ref30">13</reflink>] ).</p> <p>Another important consideration when evaluating the relative importance of an ES is the dependent variable(s) or outcome measure(s). Researchers use a variety of different assessments to measure student performance and ascertain whether their interventions are effective, ranging from: (<reflink idref="bib1" id="ref31">1</reflink>) researcher‐generated measures that are closely aligned to the intervention to (<reflink idref="bib2" id="ref32">2</reflink>) norm‐referenced, distal assessment measures. An outcome measure is considered closely aligned when it measures exactly what the intervention is designed to change. In contrast, a measure is considered distal when it does not directly reflect the outcomes targeted by the intervention and measures student performance more broadly. An intervention is more likely to result in a large ES when the outcome measure is closely aligned than when a distal, overall measurement is used. For example, an ES of 0.30 on a proximal assessment of 10 spelling words that were directly taught in an intervention is not as meaningful as the same ES obtained on a distal measure of spelling achievement, such as the spelling subtest on a standardized achievement test that does not contain any of the specific words taught in the intervention. Lipsey and colleagues’ ([<reflink idref="bib11" id="ref33">11</reflink>] ) review of studies examining achievement outcomes for mainstream K‐12 students in the United States showed that researcher‐developed assessments, which are typically closely aligned with interventions, had an average ES of 0.39, whereas standardized, broad‐scope assessments had an average ES of only 0.08. Taking this information into consideration, we might conclude that an ES of 0.30 on a researcher‐developed spelling test lacks practical significance, whereas the same ES on a standardized achievement test represents an important increase in student performance.</p> <p>In group‐experimental research, education researchers compare the growth made by students in their intervention to students in a comparison condition. What students receive in the comparison condition also needs to be taken into account when evaluating the relative importance of an ES. We must ask ourselves: Compared to what? In some studies, students in the comparison condition receive no instruction (i.e., a control group), whereas in other studies students in the comparison group receive a competing intervention so the researchers can evaluate which intervention is more effective. Study participants in a true control group are not likely to improve as much on an outcome measure as would a comparison group that receives instruction on an established intervention. Accordingly, ESs for studies using “no‐treatment” control groups are likely to be larger than ESs for studies using comparison groups that receive a treatment. Thus, obtaining an ES of 0.30 may or may not be meaningful, depending on whether the comparison group received no intervention at all or received a different program that was well established as effective in previous research.</p> <p>In summary, larger ESs are required to indicate practically important effects when participants are younger, participants do not have disabilities, outcome measures are closely aligned with the intervention, and/or the study used a control group that did not receive any treatment. Conversely, smaller ESs may indicate practically important effects when participants are older, participants have disabilities, outcome measures are distal to the intervention, and/or the study used a comparison group that received treatment.</p> <hd id="AN0129573248-14">EXAMPLES FROM THE LITERATURE</hd> <p>ESs give educators valuable information that can be used to guide practice and policy when considered in the context of study characteristics. Below, we provide two examples from the literature that highlight both the usefulness of and the issues surrounding the interpretation of ESs. In conjunction with reading the summaries below, we encourage you to read the original studies.</p> <p>Bulgren, Marquis, Deshler, Lenz, and Schumaker ([<reflink idref="bib3" id="ref34">3</reflink>] ) examined the effectiveness of the question exploration routine (QER) on the performance of secondary students with and without disabilities (including 13 students with LD) on an English unit on Shakespeare's Romeo and Juliet. The authors reported a variety of ESs on researcher‐generated assessments (e.g., multiple choice, short essay tests). In general, the ESs reported were medium to large (ES = 0.73 to 1.23) based on Cohen's standards. What can we glean from these results? To answer this question, we need to examine the study context—namely participants, dependent measures, and comparison condition.</p> <p>Study participants included both students with and students without disabilities. In the narrative, the authors discuss results for these groups of students separately, but they do not report ESs for students with disabilities only (although these ESs could be calculated with the statistics provided in one of the tables). Although the intervention was associated with positive effects for students with disabilities on some outcome measures, performance gains for students with disabilities were not as dramatic as for students without disabilities. In fact, there were no differences at all between students with disabilities in the treatment and comparison conditions on some outcomes. Additionally, the dependent measures are researcher‐generated, so we should expect higher ESs than if the authors had used norm‐referenced assessments. Finally, students in the comparison condition received instruction on the same content as those in the intervention condition, which means that the ESs represent increases above and beyond what we might expect students to make via typical instruction. Therefore, when we take context into consideration, we conclude that the QER appears to result in meaningful improvements on proximal measures for secondary students overall—results that are especially impressive given that these are older students. Because the results for students with disabilities were not as dramatic, these students will likely need additional support and instruction to benefit at a commensurate level as their nondisabled peers. Additionally, effects on more distal measures, like standardized achievement tests, are likely to be smaller.</p> <p>Little and colleagues ([<reflink idref="bib12" id="ref35">12</reflink>] ) examined the effectiveness of a Tier 2 reading intervention (early reading intervention [ERI]) that they modified during instruction to meet the changing needs of participating kindergarten students who were identified as being at risk for reading difficulty. The authors reported 10 ESs for different curriculum‐based measures (CBMs) and norm‐referenced assessments that were given either at pre‐ and posttest, or at posttest only. The ESs reported ranged from no effect to a medium effect (ES = 0.00 to 0.59) based on Cohen's standards. As with the previous article, we need to examine the participants, dependent measures, and comparison condition to ascertain the relative importance of the ESs reported. Students involved in this study were at‐risk kindergartners who, at least at that point in time, did not have identified disabilities. The dependent variables implemented were not researcher‐generated, but were instead a variety of distal measures (e.g., CBMs, norm‐referenced assessments). These measures are harder to affect, meaning that lower ESs may still be of importance. Students in the comparison condition also received additional Tier 2 reading instruction, so the ESs represent gains students in the treatment condition made over and beyond their school's typical Tier 2 reading intervention. Knowing this is critically important because if the ESs just represented gains made from pre‐ to posttest, they likely would not have been impressive considering the large growth that kindergarten students without disabilities typically make in reading (Lipsey et al., [<reflink idref="bib11" id="ref36">11</reflink>] ).</p> <p>Based on the above, what might we conclude about the ESs reported by Little and colleagues ([<reflink idref="bib12" id="ref37">12</reflink>] )? Overall, it appears that the effects of the ERI Tier 2 intervention varied by dependent variable. In areas such as word blending (ES = 0.00), and alphabet (ES = 0.08) and letter sound (ES = 0.07) knowledge, there was little to no effect. However, it appears that the intervention resulted in at‐risk kindergartners making meaningful improvements in other critical areas, such as sound matching (ES = 0.59) and spelling (ES = 0.50), above and beyond what they would have achieved in typical Tier 2 reading instruction. However, there is an important caveat to this conclusion. None of the ESs reported were statistically significant (p values ranged from .06 to .49). The nonsignificant results may be due to the relatively small number of teachers (n = 21) and students (n = 90) involved in the study, which made it difficult to statistically detect difference in groups (Travers et al., [<reflink idref="bib16" id="ref38">16</reflink>] ). Therefore, it is important that other studies replicate these findings before implementing the program on a large scale.</p> <hd id="AN0129573248-15">CONCLUSION</hd> <p>ESs provide consumers with valuable information regarding the practical importance of an intervention's effect on student achievement. Because they are standardized, ESs from different studies can be evaluated by using the same general benchmarks, such as those provided by Cohen ([<reflink idref="bib5" id="ref39">5</reflink>] ). In addition, standardization allows us to directly compare ESs within and across studies, as well as synthesize ESs across studies in meta‐analyses. However, it is important to take into consideration study characteristics such as the participants, dependent measures, and comparison condition when evaluating the relative importance of individual ESs. Context is particularly important when using ESs to ascertain the potential efficacy of a particular practice for students with LD. Readers interested in learning more about group‐difference ESs should see the resources listed in Figure .</p> <ref id="AN0129573248-16"> <title>REFERENCES</title> <blist> <bibl id="bib1" idref="ref4" type="bt">1</bibl> <bibtext>American Educational Research Association. (2006). Standards for reporting on empirical social science research in AERA publications. Educational Researcher, 35(6), 33–40. https://doi.org/10.3102/0013189X035006033 </bibtext> </blist> <blist> <bibl id="bib2" idref="ref8" type="bt">2</bibl> <bibtext>Banda, D. R., & Therrien, W. J. (2008). A teacher's guide to meta‐analysis. Teaching Exceptional Children, 41(2), 66–71. https://doi.org/10.1177/004005990804100208 </bibtext> </blist> <blist> <bibl id="bib3" idref="ref9" type="bt">3</bibl> <bibtext>Bulgren, J. A., Marquis, J. G., Deshler, D. D., Lenz, B. K., & Schumaker, J. B. (2013). The use and effectiveness of a question exploration routine in secondary‐level English language arts classrooms. Learning Disabilities Research & Practice, 28, 156–169. https://doi.org/10.1111/ldrp.12018 </bibtext> </blist> <blist> <bibl id="bib4" idref="ref10" type="bt">4</bibl> <bibtext>Coe, R. (2002). It's the effect size, stupid. What effect size is and why it is important. Retrieved from <ulink href="http://www.leeds.ac.uk/educol/documents/00002182.htm">http://www.leeds.ac.uk/educol/documents/00002182.htm</ulink></bibtext> </blist> <blist> <bibl id="bib5" idref="ref11" type="bt">5</bibl> <bibtext>Cohen, J. (1988). Statistical power for the behavioral sciences (2nd ed.). Hillsdale, NJ: Erlbaum. </bibtext> </blist> <blist> <bibl id="bib6" idref="ref16" type="bt">6</bibl> <bibtext>Cook, B. G., & Cook, L. (2016). Research designs and special education research: Different designs address different questions. Learning Disabilities Research & Practice, 31, 190–198. https://doi.org/10.1111/ldrp.12110 </bibtext> </blist> <blist> <bibl id="bib7" idref="ref5" type="bt">7</bibl> <bibtext>Cook, B. G., Buysse, V., Klingner, J. K., Landrum, T. J., McWilliam, R. A., Tankersley, M., et al. (2015). CEC's standards for classifying the evidence base of practices in special education. Remedial and Special Education, 36, 220–234. https://doi.org/10.1177/074193251455727 </bibtext> </blist> <blist> <bibl id="bib8" idref="ref6" type="bt">8</bibl> <bibtext>Gersten, R., Fuchs, L. S., Compton, D., Coyne, M., Greenwood, C., & Innocenti, M. S. (2005). Quality indicators for group experimental and quasi‐experimental research in special education. Exceptional Children, 71, 149–164. https://doi.org/10.1177/001440290507100202 </bibtext> </blist> <blist> <bibl id="bib9" idref="ref1" type="bt">9</bibl> <bibtext>Hattie, J. (2008). Visible learning: A synthesis of over 800 meta‐analyses relating to achievement. London: Routledge. </bibtext> </blist> <blist> <bibl id="bib10" idref="ref17" type="bt">10</bibl> <bibtext>Lakens, D. (2013). Calculating and reporting effect sizes to facilitate cumulative science: A practical primer for t‐tests and ANOVAs. Frontiers in Psychology, 4, 863. https://doi.org/10.3389/fpsyg.2013.00863 </bibtext> </blist> <blist> <bibl id="bib11" idref="ref28" type="bt">11</bibl> <bibtext>Lipsey, M. W., Puzio, K., Yun, C., Hebert, M. A., Steinka‐Fry, K., Cole, M. W., et al. (2012). Translating the statistical representation of the effects of education interventions into more readily interpretable forms. Washington, DC: National Center for Special Education Research, Institute of Education Sciences, U.S. Department of Education. Retrieved from https://ies.ed.gov/ncser/pubs/20133000/pdf/20133000.pdf </bibtext> </blist> <blist> <bibl id="bib12" idref="ref35" type="bt">12</bibl> <bibtext>Little, M. E., Rawlinson, D. A., Simmons, D. C., Kim, M., Kwok, O. M., Hagan‐Burke, S., et al. (2012). A comparison of responsive interventions on kindergarteners’ early reading achievement. Learning Disabilities Research & Practice, 27, 189–202. https://doi.org/10.1111/j.1540-5826.2012.00366.x </bibtext> </blist> <blist> <bibl id="bib13" idref="ref30" type="bt">13</bibl> <bibtext>Mathews, H. M., Hirsch, S. E., & Therrien, W. J. (2017). Becoming critical consumers of research: Understanding replication. Intervention in School and Clinic. Advance online publication. https://doi.org/10.1177/1053451217736863 </bibtext> </blist> <blist> <bibl id="bib14" idref="ref23" type="bt">14</bibl> <bibtext>Rosenthal, J. A. (1996). Qualitative descriptors of strength of association and effect size. Journal of Social Service Research, 21(4), 37–59. https://doi.org/10.1300/J079v21n04_02 </bibtext> </blist> <blist> <bibl id="bib15" idref="ref27" type="bt">15</bibl> <bibtext>Therrien, W. J., Zaman, M., & Banda, D. R. (2011). How can meta‐analyses guide practice? A review of the learning disability research base. Remedial and Special Education, 32, 206–218. https://doi.org/10.1177/0741932510361266 </bibtext> </blist> <blist> <bibl id="bib16" idref="ref2" type="bt">16</bibl> <bibtext>Travers, J. C., Cook, B. G., & Cook, L. (2017). Null hypothesis significance testing and p‐values. Learning Disabilities Research & Practice, 32, 208–215. https://doi.org/10.1111/ldrp.12147 </bibtext> </blist> <blist> <bibl id="bib17" idref="ref29" type="bt">17</bibl> <bibtext>Vaughn, S., & Wanzek, J. (2014). Intensive interventions in reading for students with reading disabilities: Meaningful impacts. Learning Disabilities Research & Practice, 29, 46–53. https://doi.org/10.1111/ldrp.12031 </bibtext> </blist> <blist> <bibl id="bib18" idref="ref24" type="bt">18</bibl> <bibtext>What Works Clearinghouse. (2017). Procedures handbook, version 4.0. Retrieved from https://ies.ed.gov/ncee/wwc/Docs/referenceresources/wwc_procedures_handbook_v4.p </bibtext> </blist> <blist> <bibl id="bib19" idref="ref3" type="bt">19</bibl> <bibtext>Wilkinson & Task Force on Statistical Inference. (1999). Statistical methods in psychological journals: Guidelines and explanations. American Psychologist, 54, 594–604. https://doi.org/10.1037/0003-066X.54.8.594 </bibtext> </blist> </ref> <p>PHOTO (COLOR): Resources related to group‐difference effect sizes.</p> <aug> <p>By Bryan G. Cook; Lysandra Cook and William J. Therrien</p> <p></p> <p>Bryan G. Cook, Professor of Special Education at the University of Hawaii, earned his PhD from the University of California at Santa Barbara. He is Past President of CEC's Division for Research and coedits Behavioral Disorders (the research journal of CEC's Council for Children with Behavioral Disorders). His scholarly interests include evidence‐based practice, bridging the research‐to‐practice gap, open science, and meta‐research.</p> <p>Lysandra Cook, Associate Professor of Special Education at the University of Hawaii, earned her PhD from Kent State University. She is the coordinator for Project Laulima, a federal grant supporting the formation of a fully merged, cotaught teacher preparation program in elementary general and special education. Her scholarly interests include teacher preparation and evidence‐based practices.</p> <p>William J. Therrien is a Professor of Special Education at the University of Virginia. He earned his PhD from Penn State University. His main research interest is empirically investigating academic instructional interventions for students with learning disabilities.</p> </aug>
Header DbId: eric
DbLabel: ERIC
An: EJ1178972
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Group-Difference Effect Sizes: Gauging the Practical Importance of Findings from Group-Experimental Research
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Cook%2C+Bryan+G%2E%22">Cook, Bryan G.</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0001-9294-0873">0000-0001-9294-0873</externalLink>)<br /><searchLink fieldCode="AR" term="%22Cook%2C+Lysandra%22">Cook, Lysandra</searchLink><br /><searchLink fieldCode="AR" term="%22Therrien%2C+William+J%2E%22">Therrien, William J.</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0003-0594-5129">0000-0003-0594-5129</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Learning+Disabilities+Research+%26+Practice%22"><i>Learning Disabilities Research & Practice</i></searchLink>. May 2018 33(2):56-63.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Wiley-Blackwell. 350 Main Street, Malden, MA 02148. Tel: 800-835-6770; Tel: 781-388-8598; Fax: 781-388-8232; e-mail: cs-journals@wiley.com; Web site: http://www.wiley.com/WileyCDA
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 8
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2018
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Descriptive
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Effect+Size%22">Effect Size</searchLink><br /><searchLink fieldCode="DE" term="%22Learning+Disabilities%22">Learning Disabilities</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluation+Methods%22">Evaluation Methods</searchLink><br /><searchLink fieldCode="DE" term="%22Groups%22">Groups</searchLink><br /><searchLink fieldCode="DE" term="%22Research+Methodology%22">Research Methodology</searchLink><br /><searchLink fieldCode="DE" term="%22Group+Structure%22">Group Structure</searchLink><br /><searchLink fieldCode="DE" term="%22Context+Effect%22">Context Effect</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1111/ldrp.12167
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0938-8982
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Effect sizes are powerful tools for evaluating the practical importance of study findings that should be considered in the context of study characteristics such as participants, dependent variables, and comparison condition. In this article, we discuss how group-difference effect sizes are used to gauge the practical importance of group experimental studies. We first define different types of group-difference effect sizes and discuss how they can provide valuable information for research consumers. Second, we present guidelines for interpreting group-difference effect sizes. Third, we discuss important contextual variables that should be taken into account when interpreting group-difference effect sizes reported in the literature. Last, we provide two examples of how group-difference effect sizes have been used in the learning disabilities research base.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2018
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1178972
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1178972
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1111/ldrp.12167
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 8
        StartPage: 56
    Subjects:
      – SubjectFull: Effect Size
        Type: general
      – SubjectFull: Learning Disabilities
        Type: general
      – SubjectFull: Evaluation Methods
        Type: general
      – SubjectFull: Groups
        Type: general
      – SubjectFull: Research Methodology
        Type: general
      – SubjectFull: Group Structure
        Type: general
      – SubjectFull: Context Effect
        Type: general
    Titles:
      – TitleFull: Group-Difference Effect Sizes: Gauging the Practical Importance of Findings from Group-Experimental Research
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Cook, Bryan G.
      – PersonEntity:
          Name:
            NameFull: Cook, Lysandra
      – PersonEntity:
          Name:
            NameFull: Therrien, William J.
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 05
              Type: published
              Y: 2018
          Identifiers:
            – Type: issn-print
              Value: 0938-8982
          Numbering:
            – Type: volume
              Value: 33
            – Type: issue
              Value: 2
          Titles:
            – TitleFull: Learning Disabilities Research & Practice
              Type: main
ResultId 1