Making Sense of Single-Case Design Effect Sizes

Saved in:
Bibliographic Details
Title: Making Sense of Single-Case Design Effect Sizes
Language: English
Authors: Maggin, Daniel M. (ORCID 0000-0001-6018-4866), Cook, Bryan G. (ORCID 0000-0001-9294-0873), Cook, Lysandra
Source: Learning Disabilities Research & Practice. Aug 2019 34(3):124-132.
Availability: Wiley-Blackwell. 350 Main Street, Malden, MA 02148. Tel: 800-835-6770; Tel: 781-388-8598; Fax: 781-388-8232; e-mail: cs-journals@wiley.com; Web site: http://www.wiley.com/WileyCDA
Peer Reviewed: Y
Page Count: 9
Publication Date: 2019
Document Type: Journal Articles
Reports - Research
Descriptors: Effect Size, Research Design, Research Reports
DOI: 10.1111/ldrp.12204
ISSN: 0938-8982
Abstract: Single-case research methods provide a basis for demonstrating that an intervention produces a reliable change in a targeted outcome for individual cases. To supplement visual analysis of data in single-case studies, researchers frequently report statistics--often referred to as effect sizes--to summarize study findings. The recent proliferation of effect sizes used in single-case research can be confusing. In this article, after reviewing single-case research, we provide an overview of common types of effect sizes used in single-case research, including overlap metrics and within- and between-participant effect sizes, and conclude with examples of these effect sizes in the single-case literature. Our take-home message is that effect sizes are useful complements to visual analysis when interpreting results of single-case design research studies.
Abstractor: As Provided
Entry Date: 2019
Accession Number: EJ1223697
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwG1jKAaEF4cqA4zlZwp4k-7AAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDLNlnGpKdnZXALVmiAIBEICBmyHv9DACQAcp_tL-Nre8qjWhTtveOy7HgVl7-jtmDEx8IspmElOCQXAQWzSBdpl0CVw6IizQYfTSPByN-Hi0cgtVRQpsfRKDxEMER5dujMVX-JTRP37eFQYzJkXHw3AhkXRBOEKl8O9oebT1Q1rNAZohddZ3j6M473f4O8e8Z2jfr2rTnBePE8Vg4v2eXnKvqKeZYlgn8P0zVOFr
Text:
  Availability: 1
  Value: <anid>AN0137846215;7mj01aug.19;2019Aug03.05:48;v2.2.500</anid> <title id="AN0137846215-1">Making Sense of Single‐Case Design Effect Sizes </title> <p>Single‐case research methods provide a basis for demonstrating that an intervention produces a reliable change in a targeted outcome for individual cases. To supplement visual analysis of data in single‐case studies, researchers frequently report statistics—often referred to as effect sizes—to summarize study findings. The recent proliferation of effect sizes used in single‐case research can be confusing. In this article, after reviewing single‐case research, we provide an overview of common types of effect sizes used in single‐case research, including overlap metrics and within‐ and between‐participant effect sizes, and conclude with examples of these effect sizes in the single‐case literature. Our take‐home message is that effect sizes are useful complements to visual analysis when interpreting results of single‐case design research studies.</p> <p>Mr. Katz is a special education teacher at Fox Middle School. Charles, a student with a specific learning disability in mathematics who struggles with computational aspects of long division, is on his caseload. Mr. Katz searches through peer‐reviewed journal articles to identify an effective, research‐based intervention to use with Charles. Thinking back to a research class in his Master's program, Mr. Katz remembers that there are two types of research designs appropriate for evaluating whether interventions cause improved student outcomes: group experiments and single‐case design research. Mr. Katz feels he understands single‐case designs better than group experimental research, so he searches the single‐case literature for an intervention that has been demonstrated to work for students whose learning profiles and challenges are similar to Charles's. Mr. Katz notices that the authors of single‐case research studies often report effect sizes, something that he doesn't remember from his Master's class. For example, the authors of one article reported 78 percent points of nonoverlapping data, and another article reported a within‐participant effect size of 1.63. He really isn't sure what the term "effect size" means, or how to interpret the reported values, making it challenging for Mr. Katz to determine whether the interventions he is reading about are effective or not.</p> <p>The purpose of this article is to provide an overview of <emph>effect sizes</emph> (see Table  for definitions and examples of key terms, such as "effect size," that are used in this article) employed in single‐case design (SCD) studies, with the goal of assisting educators like Mr. Katz in interpreting, evaluating, and consuming SCD research. Effect sizes are useful complements to visual analysis when interpreting results of SCD studies. Before describing common effect sizes used in SCD research, we briefly review the purpose and underlying tenets of SCD research to provide a context for interpreting effect sizes. We conclude with brief examples of SCD effect sizes used in recent publications in special education.</p> <p>Definitions and Examples of Key Terms</p> <p> <ephtml> <table><thead><tr><th>Term</th><th align="center">Definition</th><th align="center">Example</th></tr></thead><tbody><tr><td>Baseline phase</td><td>Consecutive data points collected when the intervention is not implemented.</td><td>Data during the baseline phase indicated that Rosa's academic engagement was decreasing.</td></tr><tr><td>Causal inference</td><td>Conclusion that an intervention is responsible for changes in the outcome based on the results of a rigorous experiment.</td><td>Dr. Jerome drew a causal inference from the SCD study because of the functional relation between the intervention and student performance.</td></tr><tr><td>Effect size</td><td>A measure or index indicating the size or magnitude of the effect of an intervention.</td><td>Dr. Jerome reported a large effect for the study.</td></tr><tr><td>Experimental design</td><td>Research designs that examine whether an intervention causes changes in a targeted outcome.</td><td>SCD is a frequently used type of experimental design in special education research.</td></tr><tr><td>Functional relation</td><td>When a dependent variable changes repeatedly and predictably corresponding to the introduction and withdrawal of the independent variable.</td><td>Dr. Jerome's SCD study established a functional relation between the token economy intervention and Rosa's academic engagement.</td></tr><tr><td>Intervention phase</td><td>Consecutive data points collected when the intervention is implemented.</td><td>Data during the intervention phase indicated that Rosa's academic engagement increased.</td></tr><tr><td>Multiple‐baseline design</td><td>A type of SCD in which the intervention is introduced to participants at different points in time to increase confidence that it is the independent variable that led to changes in the outcome.</td><td>Dr. Jerome conducted a multiple‐baseline design to examine whether a functional relation existed between the token economy intervention and students' academic engagement.</td></tr><tr><td>Overlap metrics</td><td>Class of single‐case effect size metric based on the extent to which the data in phases overlap. Common overlap metrics include PND, IRD, NAP, PAND, and Tau‐U.</td><td>Dr. Jerome used PND, a type of overlap metric, to represent the effect of the intervention for study participants.</td></tr><tr><td>Repeated measures</td><td>The ongoing, continuous assessment of the target outcome across time and conditions.</td><td>Dr. Jerome took repeated measures of academic engagement in baseline and intervention phases.</td></tr><tr><td>Standard deviation</td><td>Index of the extent of variability or fluctuation across a set of data. Large standard deviations indicate a high degree of variability and low standard deviations indicate little variability.</td><td>On average, Rosa was academically engaged 55% of the time during baseline, with a <italic>SD</italic> of 8%.</td></tr><tr><td>Standardized mean difference</td><td>Specific type of effect size metric that indexes data in standard deviation units.</td><td>Dr. Jerome used within‐participant standardized mean difference effect sizes to represent the effect of the intervention for study participants.</td></tr><tr><td>Visual analysis</td><td>An approach to data analysis in SCD that involves visually examining graphed data of participant outcomes.</td><td>Visual analysis suggested a functional relation between the token economy and academic engagement.</td></tr><tr><td>Within‐study replication</td><td>Repeatedly examining intervention effects within an SCD study.</td><td>Within‐study replication of the positive effects of the token economy suggested a functional relation.</td></tr></tbody></table> </ephtml> </p> <p>1 <emph>Note</emph>. PND, percentage of nonoverlapping data; IRD, improvement rate difference; NAP, nonoverlap of all pairs; PAND, percent of all nonoverlapping data.</p> <hd id="AN0137846215-2">SINGLE‐CASE DESIGN RESEARCH</hd> <p>In a previous article in the <emph>Learning Disabilities Research & Practice</emph> research‐to‐practice series, Maggin, Cook, and Cook ([<reflink idref="bib10" id="ref1">10</reflink>]) described SCD research, emphasizing how it is useful for informing the selection of instructional interventions and practices. To review, SCDs are a set of <emph>experimental designs</emph> in which the individual (or sometimes a group of individuals) serves as their own comparison. SCD methods provide a basis for making <emph>causal inferences</emph> about the relation between an outcome and an instructional intervention by comparing how an individual performs when the intervention is and is not being implemented (Kilgus, Riley‐Tillman, & Kratochwill, [<reflink idref="bib8" id="ref2">8</reflink>]). In SCD research, causal inferences are supported through demonstrations of a <emph>functional relation</emph>, which occurs when data demonstrate that student performance changes in correspondence with (or as a function of) changes in an intervention (e.g., a student's on‐task behavior reliably improves when an intervention is implemented, but decreases when the intervention is withdrawn). What makes SCD research unique is that the responses of the individual—rather than group averages—are compared to estimate the effects of an intervention.</p> <p>To provide a basis for comparing student responses, SCD research requires the collection of repeated measures of the targeted outcome. Therefore, SCD researchers examine outcomes that they can measure repeatedly and frequently, often on a daily basis, such as the number of words read correctly per minute, the number of math problems solved correctly on a homework assignment, and the proportion of time a student is on task during math class. Collecting repeated measures of the target behavior allows the researcher to estimate the individual's performance over time, and to determine whether meaningful changes in the level, trend, and variability of performance correspond with manipulating (e.g., introducing or withdrawing) the intervention (Kazdin, [<reflink idref="bib7" id="ref3">7</reflink>]). Figure  presents data from Rodriguez and Anderson ([<reflink idref="bib13" id="ref4">13</reflink>]), in which the authors examined the effects of an interdependent contingency intervention, where all students in a group earn reinforcement if the collective behavior of the group meets a criterion for the number of problem behaviors demonstrated during a set period of time. The authors implemented a <emph>multiple‐baseline design</emph>, a common approach for evaluating the effect of an intervention on an outcome (see Maggin et al., [<reflink idref="bib10" id="ref5">10</reflink>]), across five small groups of students. The graphed data points represent the percentage of 10‐second intervals in which a problem behavior occurred in the small group. The researchers collected this data virtually every day over more than 30 days to estimate participant performance during <emph>baseline</emph> (when the group contingency was not implemented) and <emph>intervention</emph> (when the group contingency was implemented) phases.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/7MJ/01aug19/ldrp12204-fig-0001.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="ldrp12204-fig-0001.jpg" title="Data from Rodriguez and Anderson's ([13]) study. From "Integrating a social behavior intervention during small group academic instruction using a total group criterion intervention," by B. J. Rodriguez & C. M. Anderson, [13], Journal of Positive Behavior Interventions, 16, p. 240. Copyright 2013 by Hammill Institute on Disabilities. Reprinted with permission." /> </p> <p></p> <p>In addition to repeated measures, SCDs emphasize the role of <emph>within‐study replication</emph>. Replication refers to the process of repeating an experimental effect that demonstrates the intervention resulted in the predicted change in participant performance. Within‐study replication provides repeated demonstrations in the same study that the introduction and/or withdrawal of intervention results in predicted changes in the target outcome, thereby increasing confidence that the intervention—rather than some alternative explanation—is responsible for changes in the target outcome. Researchers use particular SCD designs to document within‐participant replication (see Maggin et al., [<reflink idref="bib10" id="ref6">10</reflink>]). Returning to Figure , Rodriguez and Anderson ([<reflink idref="bib13" id="ref7">13</reflink>]) arranged baseline and intervention phases to repeatedly demonstrate, or replicate, the experimental effect of the group contingency intervention across the five small groups at different points in time.</p> <p> <emph>Visual analysis</emph> of data patterns to determine whether, and the degree to which, an experimental effect is replicated within a study is the primary basis for analyzing SCD data and concluding whether a functional relation exists between the intervention and target outcomes. See Maggin et al. ([<reflink idref="bib10" id="ref8">10</reflink>]) for additional information on how to visually analyze SCD research by considering patterns related to level, trend, and variability of performance.</p> <p> <emph>Mr. Katz continues to search the SCD literature for an intervention to assist Charles with his long‐division struggles. He examines the graphs in multiple SCD studies to determine whether the designs include enough within‐participant replications and whether the data patterns support the interventions he is reading about as effective. In some cases, the designs and data patterns clearly indicated that the intervention was effective, whereas others seemed less clear. After reading several studies with inconclusive visual results, Mr. Katz begins reviewing the effect sizes reported in the articles more carefully. He notices a range of indices designed to capture the extent of overlap across phases, as well as metrics used to quantify the magnitude of effect between baseline and intervention phases. His curiosity piqued, Mr. Katz searches for resources to assist with interpreting each of these types of effect sizes</emph>.</p> <hd id="AN0137846215-4">SINGLE‐CASE RESEARCH EFFECT SIZES</hd> <p>Maggin et al. ([<reflink idref="bib10" id="ref9">10</reflink>]) emphasized that the traditional approach for determining whether data in an SCD study indicate a functional relation is visual analysis, or visually examining data level, trend, and variability within and between phases. Increasingly, however, researchers are including statistics—often referred to as effect sizes—to provide another way to describe and analyze the graphed results. The term <emph>effect size</emph> refers to a statistic that indexes a magnitude of effect on a common scale (Cook, Cook, & Therrien, [<reflink idref="bib3" id="ref10">3</reflink>]). In SCD research, effect size metrics represent the magnitude of difference between phases in which the intervention is and is not present. When describing effect sizes used for group‐experimental research, Cook et al. ([<reflink idref="bib3" id="ref11">3</reflink>]) noted that group‐comparison effect sizes, such as Cohen's <emph>d</emph> and Hedges' <emph>g</emph>, index the difference between two groups—typically those receiving the intervention and those serving as the control—in <emph>standard deviation</emph> units. Reporting effect sizes presented in a common metric such as standard deviation units assists research consumers in interpreting study findings, facilitates comparisons of findings across studies, and allows findings from different studies to be synthesized in meta‐analyses.</p> <p>The extension of effect sizes to SCD research is intuitively appealing because effect sizes have the same potential benefits for SCDs, including assisting research consumers in evaluating the findings both within a particular study and across multiple studies, and allowing for meta‐analysis of findings from SCD studies. Yet the characteristics and design principles of SCD research present statistical challenges for computing accurate effect sizes. For instance, SCD research is predicated on the collection of repeated measures of a target behavior for the same individual in the same setting over time. Group‐experimental effect size cannot be directly applied to SCD research because, by their very nature, the data in SCD research violate important statistical assumptions of these approaches. The relatively low number of participants and data points in SCD studies compounds the challenges to computing meaningful effect sizes. Computational challenges notwithstanding, developing and applying common, standardized effect sizes in SCD research papers affords many benefits and has become commonplace.</p> <p>Due to the statistical challenges presented by SCD research, and the different aspects of SCD research that can be considered when estimating the effects of the intervention, researchers have developed an array of approaches to measure effect sizes. Given the sheer volume of SCD effect size measures available, it is not possible to review each one that readers may encounter in the SCD literature. Therefore, we focus our attention on some of the most commonly used SCD effect sizes, including those based on the extent to which data between study phases overlap, and those designed to mirror effect sizes derived from group‐based experimental research. In the following sections, we provide an overview of these two general approaches, with an emphasis on conceptual principles and interpretation.</p> <hd id="AN0137846215-5">Overlap Metrics</hd> <p> <emph>Overlap metrics</emph> were among the initial approaches developed to quantify differences between baseline and intervention phases in SCD research (Scruggs & Mastropieri, [<reflink idref="bib14" id="ref12">14</reflink>]). As the name suggests, overlap (and nonoverlap) methods refer to the extent to which data in adjacent phases contain values that are within the same range. Scruggs, Mastropieri, and Casto ([<reflink idref="bib15" id="ref13">15</reflink>]) developed the initial—and perhaps most intuitive—overlap method, called the percentage of nonoverlapping data (PND). PND is computed by identifying the most extreme value in the baseline phase in the intended or therapeutic direction (e.g., the data point indicating the most words read correctly, or the lowest proportion of intervals in which problem behavior occurs) and comparing it to all data points in the subsequent intervention phase. If the best baseline data point equals or exceeds an intervention data point, it is counted as overlapping data. The number of nonoverlapping data points in the intervention phase, divided by total intervention phase points, is the PND. Using the data from Figure , the lowest value for problem behavior in an interval for Deborah's group in the baseline phase is 40, a value that is higher than all of the data points in the intervention phase. In other words, Deborah's group performed better in every intervention session than they did in their best‐performing baseline session. As such, the PND for Deborah's group is 100 percent, indicating that there is no overlap between baseline and intervention phases, suggesting that the intervention was highly effective for this group. In contrast, for Candice's group the lowest data point in baseline is 8, a value that is lower than all of the data points in intervention. Because all of the data points overlap with the most extreme baseline data point, the resulting PND is 0 percent, indicating that the intervention was ineffective for this group. PNDs for each of the groups in Figure  are provided in Table . Conventions suggest that a PND of 80 percent or more indicates an effective intervention, 60–80 percent indicates a moderate effect, and values smaller than 60 percent suggest no effect (Scruggs & Mastropieri, [<reflink idref="bib14" id="ref14">14</reflink>]).</p> <p>Overlap Effect Size Definitions, Conventions, and Application to Rodriguez and Anderson () Data</p> <p> <ephtml> <table><thead><tr><th /><th /><th>Small Group</th></tr><tr valign="bottom"><th>Effect Size</th><th align="center">Definition and Conventions for Interpretation</th><th align="center">Deborah</th><th align="center">Amy</th><th align="center">Barbara</th><th align="center">Natasha</th><th align="center">Candice</th></tr></thead><tbody><tr><td>PND</td><td>PND represents the proportion of overlapping data between the phases compared based on the most extreme baseline value. PND ranges from 0% to 100%, with values greater than 70% considered large, between 50% and 70% considered moderate, and below 50% considered small.</td><td align="center">100%</td><td align="center">90%</td><td align="center">94%</td><td align="center">64%</td><td align="center">0%</td></tr><tr><td>IRD</td><td>IRD represents the ratio of improved to nonimproved data points in the phases compared. IRD ranges from 0 to 1.0, with values greater than 0.70 considered large, between 0.50 and 0.70 considered moderate, and below 0.50 considered small.</td><td>1.0</td><td>0.85</td><td>0.93</td><td>0.78</td><td>0.62</td></tr><tr><td>NAP</td><td>NAP represents the ratio of improved to nonimproved data points, with all data points in each phase compared individually. NAP ranges from 0 to 1.0, with values greater than 0.90 considered large, between 0.60 and 0.90 considered moderate, and below 0.60 considered small.</td><td>1.0</td><td>0.98</td><td>0.99</td><td>0.96</td><td>0.81</td></tr><tr><td>PAND</td><td>PAND represents the proportion of data points removed from the intervention phase to the total number of data points in the intervention and baseline phase. PAND ranges from 0.50 to 1.0, with values greater 0.90 considered large, between 0.60 and 0.90 considered moderate, and below 0.60 considered small.</td><td>1.0</td><td>0.94</td><td>0.97</td><td>0.89</td><td>0.81</td></tr><tr><td>Tau‐U</td><td>Tau‐U represents the proportion of data that improved between baseline and intervention phases after controlling for trends in the baseline data. Tau‐U ranges from 0 to 1.0 (though values can exceed 1.0), with values greater than 0.90 considered large, between 0.60 and 0.90 considered moderate, and below 0.60 considered small.</td><td>0.99</td><td>1.09</td><td>1.00</td><td>0.83</td><td>0.43</td></tr></tbody></table> </ephtml> </p> <p>2 <emph>Note</emph>. PND, percentage of nonoverlapping data; IRD, improvement rate difference; NAP, nonoverlap of all pairs; PAND, percent of all nonoverlapping data.</p> <p>PND specifically, and overlap metrics in general, are appealing because they are relatively easy to understand and compute, in that they represent different ways to quantify overlapping data, a critical element when determining whether a functional relation exists. Despite these strengths, overlap statistics have received criticism because they account for only a single characteristic used in visual analysis, and there are circumstances where overlap does not fully capture the effects of an intervention. For instance, an intervention that produces small but stable change where none of the data points between phases overlap (e.g., a student increases from an average of eight math problems completed in baseline to an average of 11 problems in the intervention phase) results in the same PND as for a participant with much larger changes (e.g., from an average of eight math problems completed in baseline to an average of 30 problems in the intervention phase). PND is also sensitive to outliers, data points that deviate from the general pattern of performance. Returning to Figure , one might argue that the intervention appears to be more effective for Barbara's group than for Deborah's, because the baseline data tend to be more extreme, and because intervention data are lower and more stable for Barbara's group. Because there is one outlying data point in baseline and one in the intervention phase, however, the PND values suggest that the intervention was more effective for Deborah's group (PND = 100 percent) than for Barbara's group (PND = 94 percent). Thus, under certain conditions, PND may not provide an accurate index of intervention effect.</p> <hd id="AN0137846215-6">Overlap Extensions</hd> <p>The intuitive framework of data overlap, and its alignment with a critical element of visual analysis, led researchers to develop additional metrics that addressed many of PND's limitations. Table  provides definitions, conventions for interpretation, and values for the groups in Figure  for some of the most common overlap metrics that readers are likely to encounter when reading SCD studies: improvement rate difference (IRD), nonoverlap of all pairs (NAP), percent of all nonoverlapping data (PAND), and Tau‐U. Each of these effect sizes extends the original PND approach by taking into account the overlap of multiple data points in the baseline rather than just one. Because the overlap extensions compare multiple data points, the effect sizes are on different scales that generally range from 0 to 1, though there are special circumstances noted in Table . Rather than detail how each effect size is computed, we focus on the application of these overlap metrics to the data included in the example.</p> <p>Recall that (a) the purpose of the intervention in the Rodriguez and Anderson ([<reflink idref="bib13" id="ref15">13</reflink>]) study was to reduce the occurrence of disruptive behaviors, and therefore it is desirable to have data with low values, and (b) each of these effect sizes represents the extent to which the data in adjacent phases overlap. As such, it is easy to see why the data for Deborah's, Amy's, and Barbara's groups each demonstrate consistent and large effects across the overlap metrics. The data associated with Natasha's and Candice's small groups, however, show greater overlap between the baseline and intervention phases. Particularly instructive are the data for Candice's group, in which there was one baseline data point that was particularly low, resulting in a PND of 0, suggesting no effect for the intervention. Moreover, the data for Candice's group tend to move downward during baseline (i.e., performance was already changing in the desired direction before the intervention was implemented), and because Tau‐U accounts for trend, the Tau‐U indicates a smaller change in the outcome than either IRD, NAP, or PAND, all of which suggest moderate change. Finally, it is noteworthy that the Tau‐U effect size for Amy's group exceeds 1.0. This outcome occurs because Tau‐U accounts for baseline trend in the data and because—as can be seen—the data for this group was trending upward (indicating that behavior was getting worse) before implementation of the intervention. As such, Tau‐U "rewards" data patterns with baseline trends moving in the nontherapeutic direction with larger effects.</p> <hd id="AN0137846215-7">Single‐Case Design Effect Sizes Comparable to Group‐Comparison Effect Sizes</hd> <p>Overlap metrics have been the most prevalent effect size measures reported in SCD research (Maggin, O'Keeffe, & Johnson, [<reflink idref="bib12" id="ref16">12</reflink>]). In recent years, however, researchers have focused on developing new metrics that more closely align with the common effect sizes used in group‐based experimental research and are reported in standard deviation units (see Cook et al., [<reflink idref="bib3" id="ref17">3</reflink>]). The extension of effect sizes expressed in standard deviation units to SCD research is appealing because it strengthens the alignment between effect sizes in group and SCD research, potentially making interpretation of effects in the two types of research comparable. The extension of the standardized mean difference approach to SCD effects sizes has resulted in two approaches for estimating SCD effect sizes, each with its own interpretation. The first approach, within‐participant effect sizes, indexes an individual's change in performance between intervention and baseline conditions in relation to the variability of the individual's performance. The second approach, between‐participant effect sizes, indexes differences between baseline and intervention phases in relation to the variability in multiple individuals' performance (Shadish, Hedges, Horner, & Odom, [<reflink idref="bib16" id="ref18">16</reflink>]).</p> <hd id="AN0137846215-8">Within‐Participant Effect Sizes</hd> <p>Initially, extending standardized mean difference effect sizes from group experiments to SCD research was accomplished by subtracting the mean performance in intervention from mean performance in baseline, and dividing by the standard deviation of the individual's baseline data (Busk & Serlin, [<reflink idref="bib2" id="ref19">2</reflink>]). The result is an effect size in standard deviation units that is conceptually similar to Cohen's <emph>d</emph> and Hedges' <emph>g</emph> in group experimental research (see Cook et al., [<reflink idref="bib3" id="ref20">3</reflink>]).</p> <p>Returning to the five small groups in the Rodriguez and Anderson ([<reflink idref="bib13" id="ref21">13</reflink>]) study, the standardized mean difference effect size for Deborah's group is –3.25, indicating a reduction of 3.25 <emph>SD</emph>s in problem behavior between baseline and control; in other words, the change in behavior for Deborah's group was 3.25 <emph>SD</emph>s as large as the typical difference of baseline data points from the baseline average. Thus, the difference in performance after the intervention was implemented is much larger than the naturally occurring variability in group performance during baseline, indicating that the intervention had large effects. Results for the other groups included effect sizes of –2.45 for Amy's students, –3.02 for Barbara's, –1.60 for Natasha's, and –0.98 for Candice's.</p> <p>Because within‐participant effect sizes are expressed in standard deviation units for the individual participant (or case), it is important to recognize that these effect sizes are specific to each individual (or case) and cannot be meaningfully combined or compared with within‐participant effect sizes for other participants (or cases). For example, when considering Deb's, Amy's, and Barb's groups, although the largest difference between average group performance during baseline and intervention phases is for Amy's group, that group has the smallest within‐participant effect size among the three groups. That outcome is because the within‐participants effect is calculated in baseline standard deviation units for each individual case, and because there is more variability—or fluctuation—in baseline performance for Amy's group, resulting in a smaller effect size despite larger differences in baseline and intervention data.</p> <p>Because variability in baseline performance is different for each individual, within‐participant effect sizes are inherently individualized and should not be used to compare or combine the effects of an intervention across multiple individuals. Presently, conventions for interpreting within‐participant standardized mean difference effect sizes for SCD research suggest that values of 2.0 or greater indicate an effective intervention (Jenson, Clark, Kircher, & Kristjannson, [<reflink idref="bib6" id="ref22">6</reflink>]). It is worth noting that SCD effect sizes using the standardized mean difference approach tend to be quite a bit larger than those drawn from group‐based experimental research (Maggin, Lane, & Pustejovsky, [<reflink idref="bib11" id="ref23">11</reflink>]).</p> <hd id="AN0137846215-9">Between‐Participant Effect Size</hd> <p>By using data from multiple participants, the between‐participant SCD effect size estimates the effects of an intervention across participants (or cases) in a study. In this way, between‐participant SCD effect sizes are the most conceptually consistent with the effect sizes, such as Cohen's <emph>d</emph> and Hedges' <emph>g</emph>, commonly used in group‐experimental research. Returning to the Rodriguez and Anderson ([<reflink idref="bib13" id="ref24">13</reflink>]) study, using a formula developed by Hedges, Pustejovsky, and Shadish ([<reflink idref="bib5" id="ref25">5</reflink>]), the overall between‐group effect size for all small groups in Rodriguez and Anderson's study is –1.99, suggesting that, on average, the intervention was associated with a reduction in problem behavior across groups that was almost twice as large (i.e., two standard deviations) as the variability in naturally occurring problem behavior across student groups in the baseline. Experts have not reached consensus on conventions for interpreting SCD between‐groups effect sizes, though effects approaching 2.0 are not uncommon and support an intervention as being effective (e.g., Barton, Pustejovsky, Maggin, & Reichow, [<reflink idref="bib1" id="ref26">1</reflink>]).</p> <p>The nuanced, but important, distinction between the within‐ and between‐case standardized mean difference effect sizes in SCD research relates to the approach used to estimate the variability in baseline. Baseline variability—or the fluctuation of data points within the baseline phase—is the basis for standardizing the difference between data in the baseline and intervention phases. Thus, how standard deviation is determined is critical, and leads to different interpretations and uses of the metrics. Because between‐participant effect sizes are expressed in terms of normally occurring (baseline) variability across all participants in a study (rather than an individual case), they are not unique to an individual. Accordingly, they may be comparable in interpretation to effect sizes from group‐based experimental research and can be used to compare and combine effects across studies. In other words, research consumers can use between‐case effect sizes to evaluate the overall or generalized effects of the intervention, as well as indexing individual effects of an intervention. Research consumers will increasingly encounter between‐participant effect sizes in SCD research. It is important to note that although there are several methods for computing between‐participant effect sizes in SCD research, each estimates effect in standard deviation units across multiple cases and is comparable to common effect sizes in group experimental research, such as Cohen's <emph>d</emph> and Hedges' <emph>g</emph> (Shadish et al., [<reflink idref="bib16" id="ref27">16</reflink>]).</p> <hd id="AN0137846215-10">EXAMPLES FROM THE LITERATURE</hd> <p>In this section, we briefly review examples of recent SCD studies in special education. As we have discussed, there are many SCD effect sizes that use different approaches to estimate study effects, each with its own strengths and weaknesses. To avoid giving research consumers an overly narrow perspective on intervention effects, SCD researchers often report multiple effect sizes. For example, Gage, Grasley‐Boy, and MacSuga‐Gage ([<reflink idref="bib4" id="ref28">4</reflink>]) reported PND, Tau‐U, and a within‐participants standardized mean difference effect size in their multiple‐baseline study across three teachers examining the effect of professional development on teachers' use of behavior‐specific praise. PND ranged from 90 to 100 percent for the three teachers, Tau‐U ranged from.97 to 1.0, and within‐participants standardized mean difference effect sizes ranged from 2.7 to 10.8 (the teacher with an effect size of 10.8 almost never used behavior‐specific praise in the baseline phase, resulting in little baseline variability and a very large within‐participants effect size). All effect sizes corresponded with visual analysis of the graphed data and indicated an effective intervention with large effects across teachers.</p> <p>Losinski, Ennis, Sanders, and Wiseman ([<reflink idref="bib9" id="ref29">9</reflink>]) used a multiple‐baseline design across three schools to examine the effects of a self‐regulated strategy development intervention on the fraction calculations of students with or at risk for disabilities, four of whom were identified as having LD. The authors calculated both PND for individual students and a between‐participant effect size across all 16 participants. The authors interpreted PND ≥ 70 percent as indicating an effective intervention. PNDs ranged from 0 to 100 percent across the student participants, with all but two PNDs ≥ 70 percent—including three of four of the students with LD. PND for one student with LD was 50 percent, a result that the authors interpreted as indicating questionable effectiveness. The authors interpreted the between‐participant effect size across all participants of 1.53 as a large effect.</p> <hd id="AN0137846215-11">SUMMARY AND CONCLUSION</hd> <p>SCD methods provide a basis for demonstrating that a practice produces a reliable change in a targeted outcome for an individual or group. Traditional methods for evaluating SCD research rely on visual analysis of graphed data. Though visual analysis uses standardized procedures, it can present challenges for consistent interpretation and summarization, and does not lend itself to objective synthesis of multiple SCD studies (e.g., meta‐analysis). In response, researchers increasingly report statistics—often referred to as effect sizes—to summarize findings of SCD research and augment visual analysis. The large number of effect sizes used in SCD research can be confusing and can make interpretation challenging for research consumers. To address this issue, we provided an overview of common effect sizes for summarizing SCD research results. Interested readers can read more about SCD effects in the resources in Table .</p> <p>Recommended Resources with Information on Single‐Case Effect Sizes</p> <p> <ephtml> <table><thead><tr><th>Resource</th><th align="center">Brief Description</th></tr></thead><tbody><tr><td>Scruggs and Mastropieri (<xref ref-type="bibr" rid="bibr14">14</xref>)</td><td>The authors discuss effect sizes in SCD research generally, but especially percentage of nonoverlapping data.</td></tr><tr><td>Shadish et al. (<xref ref-type="bibr" rid="bibr16">16</xref>)</td><td>The authors provide a thorough discussion of both within‐ and between‐participant standardized effect sizes in SCD research.</td></tr><tr><td>Vannest and Ninci (<xref ref-type="bibr" rid="bibr17">17</xref>)</td><td>The authors describe, discuss, and provide interpretive guidelines for 5 overlap effect sizes in SCD research.</td></tr><tr><td>Zimmerman et al. (<xref ref-type="bibr" rid="bibr18">18</xref>)</td><td>The authors calculate and compare 6 SCD effect sizes for a study investigating the effectiveness of sensory‐based interventions for young children.</td></tr><tr><td><ext-link href="http://www.singlecaseresearch.org/" /></td><td>This site provides calculators for 5 different SCD effect sizes, including improvement rate difference, nonoverlap of all pairs, and Tau‐U.</td></tr><tr><td><ext-link href="https://jepusto.shinyapps.io/SCD-effect-sizes/" /></td><td>Users can enter baseline and intervention data into this online calculator and select a range of nonoverlap and parametric SCD effect sizes to be calculated.</td></tr></tbody></table> </ephtml> </p> <p>3 <emph>Note</emph>. SCD, single‐case design.</p> <p>Although overlap metrics and standardized mean difference approaches (both within‐ and between‐participant) provide useful summaries of the data, each has limitations and is calculated and interpreted differently. It is important, therefore, that consumers continue to rely on visually analyzing data patterns and research design to determine whether there is a functional relation (Maggin et al., [<reflink idref="bib10" id="ref30">10</reflink>]). Our take‐home message is that effect sizes are useful complements to visual analysis when interpreting results of SCD research studies.</p> <ref id="AN0137846215-12"> <title> Footnotes </title> <blist> <bibl id="bib1" idref="ref26" type="bt">1</bibl> <bibtext> Requests for reprints should be sent to Daniel M. Maggin, University of Illinois at Chicago. Electronic inquiries should be sent to dmaggin@uic.edu.</bibtext> </blist> </ref> <ref id="AN0137846215-13"> <title> REFERENCES </title> <blist> <bibtext> Barton, E. E., Pustejovsky, J. E., Maggin, D. M., & Reichow, B. (2017). Technology‐aided instruction and intervention for students with ASD: A meta‐analysis using novel methods of estimating effect sizes for single‐case research. Remedial and Special Education, 38, 371 – 386. https://doi.org/10.1177/0741932517729508</bibtext> </blist> <blist> <bibl id="bib2" idref="ref19" type="bt">2</bibl> <bibtext> Busk, P., & Serlin, R. (1992). Meta‐analysis for single‐participant research. In T. R. Kratochwill & J. R. Levin (Eds.), Single‐case research design and analysis: New directions for psychology and education (pp. 187 – 212). Mahwah, NJ : Erlbaum.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref10" type="bt">3</bibl> <bibtext> Cook, B. G., Cook, L., & Therrien, W. J. (2018). Group‐difference effect sizes: Gauging the practical importance of findings from group‐experimental research. Learning Disabilities Research & Practice, 33, 56 – 63. https://doi.org/10.1111/ldrp.12167</bibtext> </blist> <blist> <bibl id="bib4" idref="ref28" type="bt">4</bibl> <bibtext> Gage, N. A., Grasley‐Boy, N. M., & MacSuga‐Gage, A. S. (2018). Professional development to increase teacher behavior‐specific praise: A single‐case design replication. Psychology in the Schools, 55, 264 – 277. https://doi.org/10.1177/1098300717693568</bibtext> </blist> <blist> <bibl id="bib5" idref="ref25" type="bt">5</bibl> <bibtext> Hedges, L. V., Pustejovsky, J. E., & Shadish, W. R. (2012). A standardized mean difference effect size for single case designs. Research Synthesis Methods, 3, 224 – 239. https://doi.org/10.1002/jrsm.1052</bibtext> </blist> <blist> <bibl id="bib6" idref="ref22" type="bt">6</bibl> <bibtext> Jenson, W. R., Clark, E., Kircher, J. C., & Kristjansson, S. D. (2007). Statistical reform: Evidence‐based practice, meta‐analyses, and single subject designs. Psychology in the Schools, 44, 483 – 493. https://doi.org/10.1002/pits.20240</bibtext> </blist> <blist> <bibl id="bib7" idref="ref3" type="bt">7</bibl> <bibtext> Kazdin, A. E. (2011). Single‐case research designs: Methods for clinical and applied settings. New York : Oxford University Press.</bibtext> </blist> <blist> <bibl id="bib8" idref="ref2" type="bt">8</bibl> <bibtext> Kilgus, S. P., Riley‐Tillman, T. C., & Kratochwill, T. R. (2016). Establishing interventions via a theory‐driven single case design research cycle. School Psychology Review, 45, 477 – 498. https://doi.org/10.17105/SPR45-4.477-498</bibtext> </blist> <blist> <bibl id="bib9" idref="ref29" type="bt">9</bibl> <bibtext> Losinski, M., Ennis, R. P., Sanders, S., & Wiseman, N. (2019). An investigation of SRSD to teach fractions to students with disabilities. Exceptional Children, 85, 291 – 308. https://doi.org/10.1177/0014402918813980</bibtext> </blist> <blist> <bibtext> Maggin, D. M., Cook, B. G., & Cook, L. (2018). Using single‐case research designs to examine the effects of interventions in special education. Learning Disabilities Research & Practice, 33, 182 – 191. https://doi.org/10.1111/ldrp.12184</bibtext> </blist> <blist> <bibtext> Maggin, D. M., Lane, K. L., & Pustejovsky, J. E. (2017). Introduction to the special issue on single‐case systematic reviews and meta‐analyses. Remedial and Special Education, 38, 323 – 330. https://doi.org/10.1177/0741932517717043</bibtext> </blist> <blist> <bibtext> Maggin, D. M., O'Keeffe, B. V., & Johnson, A. H. (2011). A quantitative synthesis of methodology in the meta‐analysis of single‐subject research for students with disabilities: 1985–2009. Exceptionality, 19, 109 – 135. https://doi.org/10.1080/09362835.2011.565725</bibtext> </blist> <blist> <bibtext> Rodriguez, B. J., & Anderson, C. M. (2014). Integrating a social behavior intervention during small group academic instruction using a total group criterion intervention. Journal of Positive Behavior Interventions, 16, 234 – 245. https://doi.org/10.1177/1098300713492858</bibtext> </blist> <blist> <bibtext> Scruggs, T. E., & Mastropieri, M. A. (2013). PND at 25: Past, present, and future trends in summarizing single‐subject research. Remedial and Special Education, 34, 9 – 19. https://doi.org/10.1177/0741932512440730</bibtext> </blist> <blist> <bibtext> Scruggs, T. E., Mastropieri, M. A., & Casto, G. (1987). The quantitative synthesis of single‐subject research: Methodology and validation. Remedial and Special Education, 8 (2), 24 – 33. https://doi.org/10.1177/074193258700800206</bibtext> </blist> <blist> <bibtext> Shadish, W. R., Hedges, L. V., Horner, R. H., & Odom, S. L. (2015). The role of between‐case effect size in conducting, interpreting, and summarizing single‐case research (NCER 2015‐002). Washington, D.C.: National Center for Education Research, Institute of Education Sciences, U.S. Department of Education. https://files.eric.ed.gov/fulltext/ED562991.pdf</bibtext> </blist> <blist> <bibtext> Vannest, K. J., & Ninci, J. (2015). Evaluating intervention effects in single‐case research designs. Journal of Counseling & Development, 93 (4), 403 – 411.</bibtext> </blist> <blist> <bibtext> Zimmerman, K. N., Pustejovsky, J. E., Ledford, J. R., Barton, E. E., Severini, K. E., & Lloyd, B. P. (2018). Single‐case synthesis tools II: Comparing quantitative outcome measures. Research in Developmental Disabilities, 79, 65 – 76.</bibtext> </blist> </ref> <aug> <p>By Daniel M. Maggin; Bryan G. Cook and Lysandra Cook</p> <p>Reported by Author; Author; Author</p> <p></p> <p>Daniel M. Maggin is an Associate Professor at the University of Illinois at Chicago. His research focuses on the evaluation of evidence‐based practice for students with emotional and behavioral disorders. Currently, he co‐edits Behavioral Disorders.</p> <p>Bryan G. Cook is Professor of Special Education at the University of Virginia. He is Past President of CEC's Division for Research and co‐edits, with Dan Maggin,  Behavioral Disorders. His scholarly interests include evidence‐based practice, bridging the research‐to‐practice gap, open science, and meta‐research.</p> <p>Lysandra Cook, Associate Professor of Special Education at the University of Virginia, earned her PhD from Kent State University. Her scholarly interests include evidence‐based practices, teacher preparation, bridging the research‐to‐practice gap, and effective co‐teaching at the post‐secondary level.</p> </aug> <nolink nlid="nl1" bibid="bib10" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib13" firstref="ref4"></nolink> <nolink nlid="nl3" bibid="bib14" firstref="ref12"></nolink> <nolink nlid="nl4" bibid="bib15" firstref="ref13"></nolink> <nolink nlid="nl5" bibid="bib12" firstref="ref16"></nolink> <nolink nlid="nl6" bibid="bib16" firstref="ref18"></nolink> <nolink nlid="nl7" bibid="bib11" firstref="ref23"></nolink>
Header DbId: eric
DbLabel: ERIC
An: EJ1223697
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Making Sense of Single-Case Design Effect Sizes
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Maggin%2C+Daniel+M%2E%22">Maggin, Daniel M.</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-6018-4866">0000-0001-6018-4866</externalLink>)<br /><searchLink fieldCode="AR" term="%22Cook%2C+Bryan+G%2E%22">Cook, Bryan G.</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-9294-0873">0000-0001-9294-0873</externalLink>)<br /><searchLink fieldCode="AR" term="%22Cook%2C+Lysandra%22">Cook, Lysandra</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Learning+Disabilities+Research+%26+Practice%22"><i>Learning Disabilities Research & Practice</i></searchLink>. Aug 2019 34(3):124-132.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Wiley-Blackwell. 350 Main Street, Malden, MA 02148. Tel: 800-835-6770; Tel: 781-388-8598; Fax: 781-388-8232; e-mail: cs-journals@wiley.com; Web site: http://www.wiley.com/WileyCDA
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 9
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2019
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Effect+Size%22">Effect Size</searchLink><br /><searchLink fieldCode="DE" term="%22Research+Design%22">Research Design</searchLink><br /><searchLink fieldCode="DE" term="%22Research+Reports%22">Research Reports</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1111/ldrp.12204
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0938-8982
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Single-case research methods provide a basis for demonstrating that an intervention produces a reliable change in a targeted outcome for individual cases. To supplement visual analysis of data in single-case studies, researchers frequently report statistics--often referred to as effect sizes--to summarize study findings. The recent proliferation of effect sizes used in single-case research can be confusing. In this article, after reviewing single-case research, we provide an overview of common types of effect sizes used in single-case research, including overlap metrics and within- and between-participant effect sizes, and conclude with examples of these effect sizes in the single-case literature. Our take-home message is that effect sizes are useful complements to visual analysis when interpreting results of single-case design research studies.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2019
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1223697
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1223697
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1111/ldrp.12204
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 9
        StartPage: 124
    Subjects:
      – SubjectFull: Effect Size
        Type: general
      – SubjectFull: Research Design
        Type: general
      – SubjectFull: Research Reports
        Type: general
    Titles:
      – TitleFull: Making Sense of Single-Case Design Effect Sizes
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Maggin, Daniel M.
      – PersonEntity:
          Name:
            NameFull: Cook, Bryan G.
      – PersonEntity:
          Name:
            NameFull: Cook, Lysandra
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 08
              Type: published
              Y: 2019
          Identifiers:
            – Type: issn-print
              Value: 0938-8982
          Numbering:
            – Type: volume
              Value: 34
            – Type: issue
              Value: 3
          Titles:
            – TitleFull: Learning Disabilities Research & Practice
              Type: main
ResultId 1