An Individual Participant Data Meta-Analysis to Support Power Analyses for Randomized Intervention Studies in Preschool: Cognitive and Socio-Emotional Learning Outcomes
Saved in:
| Title: | An Individual Participant Data Meta-Analysis to Support Power Analyses for Randomized Intervention Studies in Preschool: Cognitive and Socio-Emotional Learning Outcomes |
|---|---|
| Language: | English |
| Authors: | Martin Brunner (ORCID |
| Source: | Educational Psychology Review. 2025 37. |
| Availability: | Springer. Available from: Springer Nature. One New York Plaza, Suite 4600, New York, NY 10004. Tel: 800-777-4643; Tel: 212-460-1500; Fax: 212-460-1700; e-mail: customerservice@springernature.com; Web site: https://link.springer.com/ |
| Peer Reviewed: | Y |
| Page Count: | 38 |
| Publication Date: | 2025 |
| Document Type: | Journal Articles Information Analyses |
| Education Level: | Early Childhood Education Preschool Education |
| Descriptors: | Preschools, Social Emotional Learning, Outcomes of Education, Cognitive Objectives, Child Care Centers, Foreign Countries, Preschool Children, Child Development, Journal Articles, Statistical Analysis, Early Childhood Education, Educational Research, Randomized Controlled Trials |
| Geographic Terms: | Germany |
| DOI: | 10.1007/s10648-024-09981-z |
| ISSN: | 1040-726X 1573-336X |
| Abstract: | There is a need for robust evidence about which educational interventions work in preschool to foster children's cognitive and socio-emotional learning (SEL) outcomes. Lab-based individually randomized experiments can develop and refine such interventions, and field-based randomized experiments (e.g., cluster randomized trials) evaluate their effectiveness in real-world daycare center settings. Applying reliable estimates of design parameters in the context of a priori power analyses is essential to ensure that the sample size of these studies is adequate to support strong statistical conclusions regarding the strength of the intervention effect. However, there is little knowledge on relevant design parameters with preschool children. We therefore utilized a systematic collection of individual participant data from four German probability samples (554 [less than or equal to] N [less than or equal to] 2928) with preschool children (aged two to six years) to estimate and meta-analyze design parameters. These parameters are relevant for planning single-level (e.g., in non-clustered lab-based settings), two-level (children nested in daycare centers), and three-level (children nested in groups, with groups nested in daycare centers) randomized intervention studies targeting cognitive and SEL outcomes assessed with three methods (standardized tests, parent ratings, and educator ratings). The design parameters depict between-group and -center differences as well as the proportion of variance in the outcomes explained by different covariate sets (socio-demographic characteristics, baseline measures, and their combination) at the child, group, and center level. In conclusion, this paper provides a rich source of design parameters, recommendations, and illustrations to support a priori power analyses for randomized intervention studies in early childhood education research. |
| Abstractor: | As Provided |
| Entry Date: | 2025 |
| Accession Number: | EJ1457332 |
| Database: | ERIC |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwFSC_j083nEeIED3VU0cBD_AAAA4jCB3wYJKoZIhvcNAQcGoIHRMIHOAgEAMIHIBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDMp9AocOomxea5o45gIBEICBmnnHIov1SQzXK5olkQYDSoftlx0f8gM5Tm-8jFLQgdZZrZDUdsTinmPo83v5A1Erls2SVqxGSw_VHY5MUsc50o4tNIjQmM68qGiEtuYm2I_1Re3gBAvtT87xEBTYIINMgvNBsab4pco0rwLtnKBB1_c1xBWAASWGtbSXxQYfemFISjJel1AYuShsDZhAcmXh3nQGruPecfIwfzs= Text: Availability: 1 Value: <anid>AN0182226205;epv01mar.25;2025Mar27.13:05;v2.2.500</anid> <title id="AN0182226205-1">An Individual Participant Data Meta-Analysis to Support Power Analyses for Randomized Intervention Studies in Preschool: Cognitive and Socio-Emotional Learning Outcomes </title> <p>There is a need for robust evidence about which educational interventions work in preschool to foster children's cognitive and socio-emotional learning (SEL) outcomes. Lab-based individually randomized experiments can develop and refine such interventions, and field-based randomized experiments (e.g., cluster randomized trials) evaluate their effectiveness in real-world daycare center settings. Applying reliable estimates of design parameters in the context of a priori power analyses is essential to ensure that the sample size of these studies is adequate to support strong statistical conclusions regarding the strength of the intervention effect. However, there is little knowledge on relevant design parameters with preschool children. We therefore utilized a systematic collection of individual participant data from four German probability samples (554 ≤ N ≤ 2928) with preschool children (aged two to six years) to estimate and meta-analyze design parameters. These parameters are relevant for planning single-level (e.g., in non-clustered lab-based settings), two-level (children nested in daycare centers), and three-level (children nested in groups, with groups nested in daycare centers) randomized intervention studies targeting cognitive and SEL outcomes assessed with three methods (standardized tests, parent ratings, and educator ratings). The design parameters depict between-group and -center differences as well as the proportion of variance in the outcomes explained by different covariate sets (socio-demographic characteristics, baseline measures, and their combination) at the child, group, and center level. In conclusion, this paper provides a rich source of design parameters, recommendations, and illustrations to support a priori power analyses for randomized intervention studies in early childhood education research.</p> <p>Supplementary Information The online version contains supplementary material available at https://doi.org/10.1007/s10648-024-09981-z.</p> <hd id="AN0182226205-2">Introduction</hd> <p>Early educational interventions may be particularly beneficial for fostering preschool children's cognitive, academic, and socio-emotional learning (SEL) outcomes (Barnett, [<reflink idref="bib4" id="ref1">4</reflink>]; OECD, [<reflink idref="bib68" id="ref2">68</reflink>]). However, not all of these interventions are equally successful or work equally well in all preschools (Barnett, [<reflink idref="bib4" id="ref3">4</reflink>]; Sabol et al., [<reflink idref="bib80" id="ref4">80</reflink>]). Thus, further research is needed to develop, refine, and evaluate educational interventions aiming at fostering preschool children's development. To this end, randomized experiments are indispensable because they allow for strong causal inferences about the impact of educational interventions (Slavin, [<reflink idref="bib88" id="ref5">88</reflink>]). Many randomized experiments have been conducted in the last two decades (Connolly et al., [<reflink idref="bib25" id="ref6">25</reflink>]) because major funding agencies (e.g., the UK Education Endowment Foundation and the US Institute of Education Sciences) have emphasized the importance of randomized field experiments to provide strong evidence for what fosters student learning (Hedges &amp; Schauer, [<reflink idref="bib45" id="ref7">45</reflink>]; Slavin, [<reflink idref="bib88" id="ref8">88</reflink>]). However, a review of large-scale intervention studies (Lortie-Forgues &amp; Inglis, [<reflink idref="bib57" id="ref9">57</reflink>]) showed that a large majority of these studies were "underpowered," meaning they were not sensitive enough to detect typical interventions effects.</p> <p>To tackle this problem, applying reliable estimates of design parameters in a priori power analyses is highly recommended to guarantee that the sample size of an intervention study is large enough to allow strong statistical conclusions regarding the strength of the intervention effect in the target population and to detect a meaningful intervention effect if it exists (Bloom et al., [<reflink idref="bib11" id="ref10">11</reflink>]; Hedges &amp; Hedberg, [<reflink idref="bib42" id="ref11">42</reflink>]; Hedges &amp; Rhoads, [<reflink idref="bib43" id="ref12">43</reflink>]; Raudenbush et al., [<reflink idref="bib75" id="ref13">75</reflink>]). To achieve this, researchers in early childhood education require design parameters that align with the specific design features of the planned intervention study. Depending on the research objective, researchers can select from various study designs that employ different strategies for (a) sampling preschool children and (b) randomizing these children to experimental groups (Hedges &amp; Rhoads, [<reflink idref="bib43" id="ref14">43</reflink>]). Lab-based randomized experiments conducted under well-controlled conditions by well-trained staff are a valuable methodological tool when the research goal is to develop and refine educational interventions, or test their efficacy. These studies typically assume simple random sampling of preschool children and apply individual random assignment of the sampled children to experimental conditions. In the remainder of this paper, we refer to this type of lab-based randomized experiment as a single-level design. However, evaluating the effectiveness of interventions on a larger scale—such as when implemented during regular preschool days by teachers or educators—requires randomized experiments in field settings, particularly in daycare centers. These randomized experiments utilize hierarchical sampling methods. For example, researchers may first draw a (random) sample of daycare centers, followed by a (random) sample of children within the selected centers. Randomization in these field settings can take various forms, leading to different multilevel designs. Multisite randomized trials (MSRTs) use blocked random assignment to allocate either individual children or entire groups to experimental conditions within daycare centers. Cluster randomized trials (CRTs), also known as group-randomized trials, randomly assign entire daycare centers to different experimental groups. Notably, if randomization is carried out <emph>within</emph> daycare centers, the validity of causal inferences may be threatened, as children in the control group may be (unintentionally) exposed to the intervention. Moreover, many social interventions operate at the group level, such as reforms affecting whole daycare centers (Hedges &amp; Rhoads, [<reflink idref="bib43" id="ref15">43</reflink>]). In such cases, CRTs may be the only reasonable design option.</p> <p>Crucially, although reliable knowledge of design parameters is essential for a priori power analyses during the planning stage of randomized intervention studies, this knowledge is largely lacking for the target population of preschool children in early childhood education. Except for parameters relevant to the very specific target population of socioeconomically disadvantaged preschool children in the United States (Jacob et al., [<reflink idref="bib47" id="ref16">47</reflink>]; Spybrook et al., [<reflink idref="bib93" id="ref17">93</reflink>]), there is little empirically-based information on design parameters that researchers can use to plan studies for broader target populations or in other countries. The primary goal of this paper is, therefore, to address this significant research gap. To this end, we provide a comprehensive, systematic collection of reliable design parameters for conducting power analyses for intervention studies in both lab settings (with simple random sampling and individual random assignment) and multilevel field settings (implementing CRTs or MSRTs).[<reflink idref="bib1" id="ref18">1</reflink>] These studies may target preschool children's cognitive or SEL outcomes, assessed by standardized tests, parents, and/or educators. To enhance their reliability and generalizability, we derived design parameters from the analysis and meta-analytic integration of individual participant data (IPD) from several large-scale studies involving probability samples of children attending German daycare centers—the preschool setting that most children in Germany and many other countries attend. Notably, our paper is accompanied by extensive online supplementary materials (OSM) in the Open Science Framework (OSF),[<reflink idref="bib2" id="ref19">2</reflink>] including details on the applied systematic search for IPD, measures, and methods (OSM A), the R code (R Core Team, [<reflink idref="bib74" id="ref20">74</reflink>]) for reproducibility and replicability, and detailed interactive tables of the design parameters that are required for sample size planning (OSM B). By embedding the present design parameters within the robust methodological literature on the statistical power of randomized experiments, we thoroughly discuss and illustrate how these parameters can be applied in power analyses for randomized intervention studies involving preschool children. In summary, this paper makes several significant contributions that can assist researchers in early childhood education to conduct effective and efficient studies on the impact of cognitive and socio-emotional development interventions for preschool children.</p> <hd id="AN0182226205-3">A Spotlight on the History of Randomized Field Trials in Education</hd> <p>Building on substantial methodological developments, randomized experiments in education have a long but complex history (Bloom, [<reflink idref="bib9" id="ref21">9</reflink>]; Bloom et al., [<reflink idref="bib10" id="ref22">10</reflink>]; Boruch, [<reflink idref="bib14" id="ref23">14</reflink>]; Cook, [<reflink idref="bib26" id="ref24">26</reflink>], [<reflink idref="bib27" id="ref25">27</reflink>]; Hedges &amp; Schauer, [<reflink idref="bib45" id="ref26">45</reflink>]; Mosteller &amp; Boruch, [<reflink idref="bib64" id="ref27">64</reflink>]). Particularly from the 1960s onward, randomized experiments were recognized as a key source of evidence on the impact of educational interventions in the United States. Consequently, from the 1960s to the 1980s, many randomized field trials were conducted to evaluate educational reforms and policies (Hedges &amp; Schauer, [<reflink idref="bib45" id="ref28">45</reflink>]). Some of these trials continue to have a lasting impact on U.S. educational policy today. For instance, the Perry Preschool Project involving low-income children demonstrated that high-quality early childhood education can have long-term beneficial effects on key life outcomes, such as higher earnings and reduced criminal activity, for the students (Schweinhart, [<reflink idref="bib86" id="ref29">86</reflink>]), as well as positive effects on the lives of their own children (García et al., [<reflink idref="bib36" id="ref30">36</reflink>]). However, during the 1980s and 1990s, there was a paradigm shift in educational research towards qualitative approaches. Furthermore, policy funding for randomized trials in the United States was significantly reduced. Consequently, the number of randomized field experiments (but not lab-based experiments) decreased sharply (Hedges &amp; Schauer, [<reflink idref="bib45" id="ref31">45</reflink>]). There were, however, two notable exceptions. The policy-relevant MSRT—Project STAR (Student–Teacher Achievement Ratio)—demonstrated a positive impact of small class sizes on students' achievement in elementary school and beyond (Nye et al., [<reflink idref="bib67" id="ref32">67</reflink>]). Furthermore, drawing on national probability samples, an MSRT examined whether participation in the Upward Bound program improved high school outcomes for low-income children (Myers &amp; Schirm, [<reflink idref="bib65" id="ref33">65</reflink>]). Importantly, with the onset of the new millennium, new laws took effect—the No Child Left Behind Act in 2001 and the Education Sciences Reform Act in 2002—reflecting a strong policy demand for robust scientific evidence on the impact of educational interventions and programs. Randomized experiments were regarded as one of the standard methods for generating this type of evidence. As a result, the newly established U.S. Institute of Education Sciences launched funding strategies that led to numerous randomized experiments (Hedges &amp; Schauer, [<reflink idref="bib45" id="ref34">45</reflink>]). An overview of these trials is provided by Spybrook and colleagues (Spybrook &amp; Raudenbush, [<reflink idref="bib91" id="ref35">91</reflink>]; Spybrook et al., [<reflink idref="bib92" id="ref36">92</reflink>]).</p> <p>The timing of when randomized experiments have been recognized as an important source of rigorous evidence on the impact of educational interventions—if at all—varies considerably across countries. For example, in the United Kingdom, the Education Endowment Foundation was established in 2011 to initiate major funding schemes for randomized experiments (Dawson et al., [<reflink idref="bib28" id="ref37">28</reflink>]). Around the same time, policy-initiated or policy-funded randomized trials were conducted in Denmark, Norway, and Sweden (Pontoppidan et al., [<reflink idref="bib71" id="ref38">71</reflink>]). Finally, in Germany, there have only been a few randomized trials (e.g., the CRT by Gaspard et al., [<reflink idref="bib37" id="ref39">37</reflink>]), likely because the policy demand for generating this type of robust experimental evidence is much lower than in other countries, particularly the United States (Standing Scientific Commission on Education Policy, [<reflink idref="bib96" id="ref40">96</reflink>]). In summary, the importance of randomized experiments for advancing knowledge on the impact of educational interventions appears to be widely recognized by researchers and policymakers, although some cross-national variation exists. This is evident in the significant overall increase in randomized studies conducted in recent years, as well as the varying number of randomized trials across countries (Connolly et al., [<reflink idref="bib25" id="ref41">25</reflink>]).</p> <hd id="AN0182226205-4">Which Design Parameters Do We Need in Power Analyses?</hd> <p>One key goal in planning a randomized intervention study is to ensure that it can effectively assess the strength of the intervention effect. In this regard, design sensitivity (or simply <emph>sensitivity</emph>) is an umbrella term that covers several interrelated statistical concepts (Hedges &amp; Hedberg, [<reflink idref="bib44" id="ref42">44</reflink>]): the precision (i.e., standard error) with which the intervention effect can be estimated, the statistical power to detect the intervention effect if it exists, and the minimum detectable effect size (<emph>MDES</emph>; Bloom, [<reflink idref="bib8" id="ref43">8</reflink>])<emph>.</emph> Power analyses are essential for evaluating and assuring the sensitivity of randomized studies, for example, to determine the sample size needed to achieve a certain <emph>MDES</emph> with confidence (e.g., with level of statistical significance α = 0.05 and power of 1 − β = 0.80). To this end, researchers need to set a meaningful value for the <emph>MDES,</emph> which, for a standardized test or questionnaire, is often given in terms of the standardized mean difference <emph>SMD</emph> to depict the effect of an educational intervention (Hedges &amp; Rhoads, [<reflink idref="bib43" id="ref44">43</reflink>]). For example, based on his review of over 1,000 randomized intervention studies on student achievement, Kraft ([<reflink idref="bib51" id="ref45">51</reflink>]) proposes considering an intervention effect of <emph>SMD</emph> &lt; 0.05 as "small", 0.05 ≤ <emph>SMD</emph> &lt; 0.20 as "medium", and <emph>SMD</emph> ≥ 0.20 as "large." Further, when evaluating the meaningfulness of the intervention effect it is also recommended to take into account the cost and scalability of the intervention (Kraft, [<reflink idref="bib51" id="ref46">51</reflink>]; Lipsey et al., [<reflink idref="bib56" id="ref47">56</reflink>]). Finally, it may be helpful to compare the expected effect to empirical benchmark values, such as normative expectations of academic growth, performance differences between socio-demographic groups, or performance differences between preschools with weak and average performance levels (Brunner et al., [<reflink idref="bib17" id="ref48">17</reflink>]; Dong et al., [<reflink idref="bib32" id="ref49">32</reflink>]; Lipsey et al., [<reflink idref="bib56" id="ref50">56</reflink>]).</p> <p>Power analyses of randomized experiments require several design parameters to determine the <emph>MDES</emph> or the required sample size to achieve a certain statistical power (e.g., Dong &amp; Maynard, [<reflink idref="bib31" id="ref51">31</reflink>]; Hedges &amp; Rhoads, [<reflink idref="bib43" id="ref52">43</reflink>]).[<reflink idref="bib3" id="ref53">3</reflink>] First, most field-based randomized studies, including CRTs and MSRTs in preschool settings, implement hierarchical sampling strategies to reflect the multilevel nature of these environments. Children can be nested in daycare centers (two-level design), or they can be nested in groups, with those groups nested in daycare centers (three-level design). Clustering implies that the target outcome measures of children belonging to the same group or daycare center tend to be more similar to each other than to children in other groups or centers. Intraclass correlations (<emph>ICC</emph>s; ρ) quantify the degree of similarity between children in the same group or daycare center, measuring how much the target outcome tend to cluster together. ICC values can range from zero to one. For example, when using vocabulary test scores as the outcome in a single-level lab-based randomized experiment, the scores are not clustered because children's test scores are independently sampled, with no higher-order units (e.g., daycare centers) involved by definition. Hence, ρ = 0. Moreover, in a two-level design where children are nested in daycare centers, a value of ρ = 0 indicates that there are no mean differences in test scores between daycare centers, and all variability in the test scores is observed within the centers. Conversely, a value of ρ = 1 indicates that all children in the same daycare center achieve the same test score, meaning that the total observed variability in the test scores is due to mean-level differences between daycare centers rather than within. Importantly, when estimating the effect of an intervention, the similarity between children in the same group and/or daycare center makes the clustered data from CRTs less efficient than the unclustered data from single-level lab-based randomized experiment with the same total sample size (Hedges &amp; Hedberg, [<reflink idref="bib42" id="ref54">42</reflink>]). In other words, all else being equal, CRTs with clustered data require a larger sample size than a single-level lab-based randomized experiment with unclustered data to achieve the same <emph>MDES</emph> or statistical power. Moreover, with MSRTs, the efficiency of estimating the intervention effect depends on the variance proportions attributable to (a) mean-level differences in the target outcome and (b) variations in the magnitude of the intervention effect across daycare centers. Specifically, due to the use of blocked random assignment in MSRTs, the variance attributable to mean-level differences between daycare centers in the target outcome is not considered when calculating the standard error of the intervention effect. As a result, when the magnitude of the intervention effect is the same (or very similar) across daycare centers, MSRTs can be more efficient than lab-based randomized experiments (Moerbeek &amp; Teerenstra, [<reflink idref="bib63" id="ref55">63</reflink>]). However, when the heterogeneity of the intervention effect across daycare centers (and the corresponding variance proportion) becomes sufficiently large, MSRTs are less efficient than lab-based randomized experiments in many practical applications (Hedges &amp; Rhoads, [<reflink idref="bib43" id="ref56">43</reflink>]; Moerbeek &amp; Teerenstra, [<reflink idref="bib63" id="ref57">63</reflink>]). In summary, clustering has important consequences for power analyses. Specifically, the preschool setting that most children attend in Germany (Autor:innengruppe Bildungsberichterstattung, [<reflink idref="bib3" id="ref58">3</reflink>]) and elsewhere (e.g., the United States; U.S. Department of Education, National Center for Education Statistics, [<reflink idref="bib101" id="ref59">101</reflink>]) is a group within a daycare center. Therefore, the most relevant design parameters for planning CRTs and MSRTs in preschool settings are <emph>ICC</emph>s, which represent the proportion of total variance in children's outcomes attributable to differences between (a) groups (ρ<subs><emph>Group</emph></subs>) within daycare centers and (b) daycare centers themselves (ρ<subs><emph>Center</emph></subs>).</p> <p>Second, covariates may substantially improve the sensitivity of randomized experiments in general, and single-level, lab-based studies with unclustered data (Porter &amp; Raudenbush, [<reflink idref="bib72" id="ref60">72</reflink>]) and CRTs and MSRTs (Dong &amp; Maynard, [<reflink idref="bib31" id="ref61">31</reflink>]; Hedges &amp; Hedberg, [<reflink idref="bib42" id="ref62">42</reflink>]; Hedges &amp; Rhoads, [<reflink idref="bib43" id="ref63">43</reflink>]; Raudenbush et al., [<reflink idref="bib75" id="ref64">75</reflink>]) in particular. Covariates remove noise in the variance of the outcome measure, improving the signal of the treatment effect (Raudenbush et al., [<reflink idref="bib75" id="ref65">75</reflink>], p. 18). Covariates are not required for randomized experiments, but when they explain a substantial proportion of variance in outcomes (<emph>R</emph><sups>2</sups>), they are a very efficient way to improve (i.e., decrease) the <emph>MDES</emph> and to reduce the required sample size to achieve a certain level of statistical power (Hedges &amp; Hedberg, [<reflink idref="bib42" id="ref66">42</reflink>]; Porter &amp; Raudenbush, [<reflink idref="bib72" id="ref67">72</reflink>]; Raudenbush et al., [<reflink idref="bib75" id="ref68">75</reflink>]). Values of <emph>R</emph><sups>2</sups> can range from zero to one. In single-level lab-based studies, information is needed on the proportion of total variance in the outcome (<emph>R</emph><sups>2</sups><subs><emph>Total</emph></subs>) that can be explained by covariates to guide researchers' decisions about inclusion. In CRTs and MSRTs, covariates may operate at various levels. To decide about the inclusion of covariates, researchers therefore need information on the proportion of variance in the outcome that can be explained by covariates at the individual child level (<emph>R</emph><sups>2</sups><subs><emph>Child</emph></subs>), the group level (<emph>R</emph><sups>2</sups><subs><emph>Group</emph></subs>), and the daycare center level (<emph>R</emph><sups>2</sups><subs><emph>Center</emph></subs>).</p> <hd id="AN0182226205-5">What Can Go Wrong in Power Analyses?</hd> <p>Design parameters, in terms of <emph>R</emph><sups>2</sups><subs><emph>Total</emph></subs> as well as <emph>ICC</emph>s and <emph>R</emph><sups>2</sups>s at various levels, are essential for power analyses of randomized studies. Two major problems can occur when using the "wrong" estimates for these design parameters. First, the actual degree of clustering may be larger and/or proportion of explained variance may be smaller than expected. This leads to reduced precision in estimating the intervention effect, diminished statistical power to detect that effect if it exists, and a <emph>MDES</emph> that is larger than the actual intervention effect. Consequently, the design sensitivity of the study is compromised, undermining the research team's statistical conclusions regarding the strength of the intervention effect in the target population. Second, the actual degree of clustering may be smaller and/or the proportion of explained variance may be larger than expected. Although this may result in very high design sensitivity with small standard errors, high statistical power, and an estimated <emph>MDES</emph> that is smaller than the actual intervention effect, it may also make the experiment inefficient from a cost perspective due to an unnecessarily large sample size. Hence, the research team and funding agencies invested more resources in the study than necessary. For these reasons, leading methodologists recommend that researchers base power analyses on reliable empirical estimates of design parameters (Bloom et al., [<reflink idref="bib11" id="ref69">11</reflink>]; Hedges &amp; Hedberg, [<reflink idref="bib42" id="ref70">42</reflink>]; Hedges &amp; Rhoads, [<reflink idref="bib43" id="ref71">43</reflink>]; Raudenbush et al., [<reflink idref="bib75" id="ref72">75</reflink>]), because research has shown that the value of design parameters depends strongly on the target outcome and target population (Brunner et al., [<reflink idref="bib15" id="ref73">15</reflink>]; Stallasch et al., [<reflink idref="bib94" id="ref74">94</reflink>], [<reflink idref="bib95" id="ref75">95</reflink>]).</p> <hd id="AN0182226205-6">The Empirical Body of Knowledge on Design Parameters for Preschool Children</hd> <p></p> <hd id="AN0182226205-7">Cognitive and Socio-Emotional Learning Outcomes</hd> <p>Educational interventions in preschool may have a broad, positive impact on children's cognitive and socio-emotional development (Barnett, [<reflink idref="bib4" id="ref76">4</reflink>]; OECD, [<reflink idref="bib68" id="ref77">68</reflink>]). We therefore aim to provide reliable design parameters for two broad outcome domains: cognitive and SEL outcomes. We use cognitive outcomes as an umbrella term to cover children's skills and knowledge in various subdomains, including (a) math and science skills (e.g., counting skills), (b) verbal skills (e.g., vocabulary, sentence comprehension), as well as (c) general cognitive skills and knowledge (e.g., working memory, reasoning skills, general knowledge). In addition, drawing on the widely-accepted Cattell-Horn-Carroll (CHC) taxonomy of cognitive abilities (Flanagan &amp; Dixon, [<reflink idref="bib34" id="ref78">34</reflink>]), we also consider children's psychomotor skills as a subdomain of cognitive outcomes. This subdomain covers children's more general psychomotor skills (e.g., throwing a ball, jumping with two feet) as well as psychomotor skills that are relevant in everyday situations, for example, the ability to close a zipper or to walk up stairs (Sparrow et al., [<reflink idref="bib90" id="ref79">90</reflink>]).</p> <p>SEL refers to the learning of knowledge, skills, and attitudes that enable individuals to regulate thoughts, emotions, and behavior; establish and manage interpersonal relationships; and achieve personal, academic, and collective goals (Durlak et al., [<reflink idref="bib33" id="ref80">33</reflink>]). According to the taxonomy by Schoon ([<reflink idref="bib85" id="ref81">85</reflink>]), important SEL outcomes reflect affective, cognitive, or behavioral manifestations of socio-emotional characteristics. These characteristics can be grouped into three broad subdomains: (a) self-orientation (e.g., self-control, emotion regulation, neuroticism, and conscientiousness), (b) other-orientation (e.g., empathy, pro-social behavior, externalizing problems, aggressive behavior, disruptive behavior, extraversion, and agreeableness), and (c) task-orientation (e.g., interest, persistence, and openness). Drawing on these definitions and frameworks, we review research on single- and multilevel design parameters for the target population of preschool children.</p> <hd id="AN0182226205-8">Previous Research on Design Parameters: An Overview</hd> <p>Previous research on design parameters has contributed substantial knowledge across various target populations, including students in elementary and secondary school at national (Dong et al., [<reflink idref="bib32" id="ref82">32</reflink>]; Hedberg, [<reflink idref="bib40" id="ref83">40</reflink>]; Hedges &amp; Hedberg, [<reflink idref="bib42" id="ref84">42</reflink>]; Jacob et al., [<reflink idref="bib47" id="ref85">47</reflink>]; Stallasch et al., [<reflink idref="bib94" id="ref86">94</reflink>], [<reflink idref="bib95" id="ref87">95</reflink>]; Westine et al., [<reflink idref="bib108" id="ref88">108</reflink>]) and international levels (Brunner et al., [<reflink idref="bib15" id="ref89">15</reflink>]; Kelcey et al., [<reflink idref="bib48" id="ref90">48</reflink>]; Zopluoglu, [<reflink idref="bib112" id="ref91">112</reflink>]), as well as teachers (Westine et al., [<reflink idref="bib109" id="ref92">109</reflink>]) and students in community colleges (Somers et al., [<reflink idref="bib89" id="ref93">89</reflink>]). In particular, research on design parameters with student populations in elementary school—the target population most closely related to preschool children, which are the focus of this paper—has shown that design sensitivity can be substantially improved by including two types of covariate sets: (a) socio-demographic (SD) characteristics (e.g., children's age, socioeconomic status [SES], gender, or migration background) and (b) baseline measures of the target outcome (Dong et al., [<reflink idref="bib32" id="ref94">32</reflink>]; Hedges &amp; Hedberg, [<reflink idref="bib42" id="ref95">42</reflink>]; Jacob et al., [<reflink idref="bib47" id="ref96">47</reflink>]; Kelcey et al., [<reflink idref="bib48" id="ref97">48</reflink>]; Stallasch et al., [<reflink idref="bib94" id="ref98">94</reflink>], [<reflink idref="bib95" id="ref99">95</reflink>]; Westine et al., [<reflink idref="bib108" id="ref100">108</reflink>]). This research has also highlighted significant variation in design parameters. Specifically, cross-national variation has been observed, which limits the ability to transfer these parameters from one country to another (Kelcey et al., [<reflink idref="bib48" id="ref101">48</reflink>]; Stallasch et al., [<reflink idref="bib94" id="ref102">94</reflink>]; Zopluoglu, [<reflink idref="bib112" id="ref103">112</reflink>]). Even within the same nation, design parameters for achievement or SEL outcomes vary across achievement (Kelcey et al., [<reflink idref="bib48" id="ref104">48</reflink>]; Stallasch et al., [<reflink idref="bib94" id="ref105">94</reflink>]) and SEL domains (Dong &amp; Maynard, [<reflink idref="bib31" id="ref106">31</reflink>]), particularly when using different types of assessments (e.g., parent and teacher reports; Hedberg, [<reflink idref="bib40" id="ref107">40</reflink>]). Consequently, design parameters cannot be easily generalized across domains or types of assessments. In summary, these findings emphasize the importance of developing and applying design parameters that align closely with the target population, outcome, and assessment of the planned study.</p> <hd id="AN0182226205-9">Design Parameters for Single-Level Randomized Experiments with Preschool Children</hd> <p>Design parameters for single-level lab-based studies (with individual random assignment) require knowledge about how much covariates enhance their sensitivity. Therefore, researchers need reliable estimates of <emph>R</emph><sups>2</sups><subs><emph>Total</emph></subs> for a target population and outcome, based on a specified set of covariates. In contrast to the body of knowledge for students in elementary or secondary education (Stallasch et al., [<reflink idref="bib95" id="ref108">95</reflink>]), there is no compilation of such single-level design parameters for SD characteristics and baseline measures as covariate sets for the target population of preschool children.</p> <p>Nevertheless, relevant results can be found in several large-scale studies and meta-analyses. First, previous research has shown that SD characteristics are correlated with (and therefore explain variance in) cognitive and SEL outcomes in preschool children. In particular, children's age, especially during preschool, is substantively related to their cognitive and SEL outcomes. When children grow older their cognitive skills generally improve in all domains (Tucker-Drob, [<reflink idref="bib100" id="ref109">100</reflink>]), and their socio-emotional characteristics become more differentiated (Caspi et al., [<reflink idref="bib21" id="ref110">21</reflink>]), demonstrating a multidimensional age-based developmental pattern (Bleidorn et al., [<reflink idref="bib7" id="ref111">7</reflink>]). Further, meta-analytic results by Letourneau et al. ([<reflink idref="bib53" id="ref112">53</reflink>]) indicate that higher values on family SES measures are associated with better cognitive and somewhat better SEL outcomes in preschool children (e.g., lower levels of externalizing and internalizing behavior problems). This pattern of results was also confirmed by an international large-scale assessment with representative samples from the United States, England, and Estonia (OECD, [<reflink idref="bib68" id="ref113">68</reflink>]). Moreover, this latter study also showed that girls in preschool have higher levels of verbal skills and SEL outcomes (e.g., better prosocial and less disruptive behavior). Further, this study also found that preschool children with migration background had lower levels of verbal and mathematical skills. Finally, these children were reported by their educators to demonstrate less prosocial, but also less disruptive behavior (OECD, [<reflink idref="bib68" id="ref114">68</reflink>]).</p> <p>Second, it is well established that preschool children's baseline measures substantially predict their future cognitive performance (e.g., when using prior knowledge as baseline measure; Simonsmeier et al., [<reflink idref="bib87" id="ref115">87</reflink>]) and socio-emotional characteristics (e.g., when using other-reports as a baseline measure of children's temperament or personality; see Table S7 in Bleidorn et al., [<reflink idref="bib7" id="ref116">7</reflink>]).</p> <hd id="AN0182226205-10">Design Parameters for Multilevel Randomized Experiments with Preschool Children</hd> <p>In stark contrast to target student populations in elementary or secondary school, little is known about multilevel design parameters for children attending preschool. The most comprehensive source on design parameters is the (largely unknown) data supplement that comes along with the Optimal Design power analysis software (Spybrook et al., [<reflink idref="bib93" id="ref117">93</reflink>]). These data were obtained from three large-scale longitudinal studies of children enrolled in Head Start centers in the US, with several waves of measurement between 1997 and 2006. In addition, Jacob et al. ([<reflink idref="bib47" id="ref118">47</reflink>]) provided a few multilevel design parameters for cognitive outcomes using data (collected between 2004 and 2009) with US samples of preschool children from low-income families living in Chicago. Design parameters for populations in other countries are even more difficult to obtain because relevant results are (a) scattered across individual studies, (b) usually available only for between-center differences (i.e., ρ<subs><emph>Center</emph></subs>) but not for other key design parameters (i.e., ρ<subs><emph>Group</emph></subs>, <emph>R</emph><sups>2</sups><subs><emph>Child</emph></subs><emph>, R</emph><sups>2</sups><subs><emph>Group</emph></subs> or <emph>R</emph><sups>2</sups><subs><emph>Center</emph></subs>), and (c) not yet systematically summarized. For example, between-center differences have been reported as auxiliary results in studies involving large-scale samples with preschool children in Germany (Leyendecker et al., [<reflink idref="bib54" id="ref119">54</reflink>]; Ulferts, [<reflink idref="bib102" id="ref120">102</reflink>])[<reflink idref="bib4" id="ref121">4</reflink>] and the United Kingdom (Sammons et al., [<reflink idref="bib81" id="ref122">81</reflink>], [<reflink idref="bib82" id="ref123">82</reflink>]).</p> <p>Integrating the results on multilevel design parameters from these selected studies reveals the following pattern (see also Figure A1 for a more detailed overview): First, in the United States, there were mostly small between-center differences in cognitive outcomes (e.g., when applying standardized tests: <emph>Mdn</emph> ρ<subs>Center</subs> = 0.03) and SEL outcomes (educator reports: <emph>Mdn</emph> ρ<subs>Center</subs> = 0.01; parent reports: <emph>Mdn</emph> ρ<subs>Center</subs> = 0.00). These between-center differences were (much) more pronounced in Germany and the United Kingdom for both cognitive (<emph>Mdn</emph> ρ<subs>Center</subs> = 0.13/0.17) and SEL outcomes (educator reports: <emph>Mdn</emph> ρ<subs>Center</subs> = 0.28/0.05). Second, the size of between-group differences in the United States depended on the combination of outcome domain and method of assessment. Typical values for cognitive outcomes were <emph>Mdn</emph> ρ<subs>Group</subs> = 0.05/0.08 when using standardized tests/parent reports; typical values for SEL outcomes were <emph>Mdn</emph> ρ<subs>Group</subs> = 0.19/0.00 when using teacher/parent reports. The studies from Germany and the United Kingdom used two-level designs (children nested in daycare centers) and thus cannot provide results for ρ<subs>Group</subs>. Third, for the United States, information was available for <emph>R</emph><sups>2</sups><subs><emph>Child</emph></subs> and <emph>R</emph><sups>2</sups><subs><emph>Center</emph></subs>, but not for <emph>R</emph><sups>2</sups><subs><emph>Group</emph></subs>. Further, most estimates were available for <emph>R</emph><sups>2</sups><subs><emph>Child</emph></subs> and using a covariate set comprising SD characteristics and a baseline measure. Applying these covariates, typical values for cognitive outcomes were <emph>Mdn</emph><emph>R</emph><sups>2</sups><subs><emph>Child</emph></subs> = 0.24/0.21 when using standardized tests/parent reports, and <emph>Mdn R</emph><sups>2</sups><subs><emph>Child</emph></subs> = 0.26/0.18 for SEL outcomes when using teacher/parent reports. Design parameters for <emph>R</emph><sups>2</sups><subs><emph>Center</emph></subs> were only available for cognitive outcomes measured with standardized tests: Median values of <emph>R</emph><sups>2</sups><subs>Center</subs> were in the range 0.20 ≤ <emph>R</emph><sups>2</sups><subs>Center</subs> ≤ 0.89. The studies from Germany and the United Kingdom did not report results on <emph>R</emph><sups>2</sups><emph>s</emph> at different levels.</p> <hd id="AN0182226205-11">The Present Study</hd> <p>There is a strong need for early childhood education research to provide robust evidence on which educational interventions and practices are effective in preschools. This requires researchers to conduct lab- and field-based randomized experiments to enable causal inferences about intervention effects with preschool children. Using reliable estimates of design parameters in a priori power analyses is essential to ensure that these studies can support strong statistical conclusions regarding the strength of the intervention effect in the target population (Hedges &amp; Hedberg, [<reflink idref="bib42" id="ref124">42</reflink>]; Hedges &amp; Rhoads, [<reflink idref="bib43" id="ref125">43</reflink>]). However, knowledge of these parameters has been largely lacking in early childhood education research. Moreover, our review of the limited available design parameters for preschool children, along with research on design parameters for elementary school children, highlights that these parameters can vary across target populations, outcomes, and types of assessments. As a result, design parameters specific to the target population, outcome, and type of assessment in early childhood education research are needed.</p> <p>Therefore, the overarching goals of this paper are to (a) significantly expand the knowledge base on design parameters relevant to preschool children, (b) synthesize these parameters to build a systematic body of knowledge, and (c) integrate this knowledge into the relevant methodological literature, providing a thorough discussion and illustration of how to apply it when conducting a priori power analyses during the planning stage of randomized experiments in both lab- and field-based settings. In doing so, our paper makes the following unique contributions to early childhood education research. First, we address a major gap in knowledge about design parameters for randomized experiments with preschool children. Specifically, there is currently no comprehensive compilation of design parameters (i.e., <emph>R</emph><sups>2</sups><subs>Total</subs>) for single-level (e.g., lab-based) randomized experiments targeting the population of preschool children. Furthermore, aside from the specific population of socioeconomically disadvantaged preschool children in the United States (Jacob et al., [<reflink idref="bib47" id="ref126">47</reflink>]; Spybrook et al., [<reflink idref="bib93" id="ref127">93</reflink>]), there is limited systematic knowledge about design parameters for two-level or three-level randomized field studies (e.g., CRTs and MSRTs) for broader populations involving preschool children in the United States or in other countries. To address these significant gaps in the literature, we provide a comprehensive collection of single- and multilevel design parameters for planning randomized experiments with preschool children. To this end, we conducted an IPD meta-analysis using several datasets based on probability samples of preschool children attending daycare centers. We focus our analyses on children attending daycare centers, as this is the preschool setting that most children (aged 3 years or older) attend in Germany (see Table C3-3web in Autor:innengruppe Bildungsberichterstattung, [<reflink idref="bib3" id="ref128">3</reflink>]) and many other countries (e.g., the United States; U.S. Department of Education, National Center for Education Statistics, [<reflink idref="bib101" id="ref129">101</reflink>]). A second contribution of our paper to early childhood education research is to offer design parameters for a large variety of cognitive and SEL outcomes because interventions in preschool may target different dimensions of children's development. A third contribution is our provision of design parameters for cognitive and socio-emotional characteristics, using both parent reports (representing the family context) and educator reports (representing the preschool context), acknowledging that these parameters may differ between sources (Hedberg, [<reflink idref="bib40" id="ref130">40</reflink>]) as children may exhibit different behaviors across contexts (Mischel &amp; Shoda, [<reflink idref="bib62" id="ref131">62</reflink>]). A fourth contribution is to offer design parameters for vital covariate sets, as covariates may substantially improve the sensitivity of randomized experiments (Hedges &amp; Hedberg, [<reflink idref="bib42" id="ref132">42</reflink>]; Porter &amp; Raudenbush, [<reflink idref="bib72" id="ref133">72</reflink>]; Raudenbush et al., [<reflink idref="bib75" id="ref134">75</reflink>]). In accordance with previous research on design parameters for the school context (Hedges &amp; Hedberg, [<reflink idref="bib42" id="ref135">42</reflink>]; Jacob et al., [<reflink idref="bib47" id="ref136">47</reflink>]; Stallasch et al., [<reflink idref="bib94" id="ref137">94</reflink>], [<reflink idref="bib95" id="ref138">95</reflink>]), we estimate single- and multilevel <emph>R</emph><sups>2</sups>'s for covariate sets including (a) sociodemographic characteristics, (b) baseline measures of the target outcome, and (c) sociodemographic characteristics and baseline measures combined. Given the strong developmental dynamics of cognitive and SEL outcomes in early childhood, it is not always possible to use identical baseline (IB) measures. We therefore also explored the predictive utility of proxy baseline (PB) measures. Specifically, we used children's vocabulary knowledge as a PB measure for cognitive outcomes because it is known to predict future learning in many cognitive domains (Peng &amp; Kievit, [<reflink idref="bib70" id="ref139">70</reflink>]). Further, children's problem behavior is theoretically, conceptually, and empirically linked to children's SEL outcomes (De Pauw &amp; Mervielde, [<reflink idref="bib29" id="ref140">29</reflink>]; Schoon, [<reflink idref="bib85" id="ref141">85</reflink>]; Tackett et al., [<reflink idref="bib98" id="ref142">98</reflink>]). We therefore used an important facet of children's (externalizing) problem behavior from the SEL subdomain "other-orientation" as a PB measure (see Schoon, [<reflink idref="bib85" id="ref143">85</reflink>]). Specifically, we chose children's disruptive behavior (e.g., a child interrupts or disturbs other children), as obtained from parent or educator reports, as a PB measure for SEL outcomes as assessed with same method. A fifth contribution of our paper is to provide standard errors for all design parameters, as any estimate of a design parameter is subject to sampling error (Jacob et al., [<reflink idref="bib47" id="ref144">47</reflink>]; Stallasch et al., [<reflink idref="bib94" id="ref145">94</reflink>]). These standard errors are essential for accounting for the statistical uncertainty of design parameters when used in an a priori power analysis. A sixth contribution of our paper to early childhood education research is to (a) build systematic knowledge on design parameters, (b) examine their generalizability, and (c) provide normative values. To achieve this, we meta-analytically integrated single-, two-, and three-level design parameters across domains, subdomains, applied measures, assessment methods, and age groups.</p> <hd id="AN0182226205-12">Method</hd> <p></p> <hd id="AN0182226205-13">Database Search for Large-Scale Studies</hd> <p>To build systematic knowledge on design parameters for preschool children, we applied a two-stage approach to the meta-analysis of IPD (Brunner et al., [<reflink idref="bib16" id="ref146">16</reflink>]). To this end, we first carried out a systematic search for IPD of German large-scale studies with preschool children (see Fig. 1). Specifically, we sought studies that met the following inclusion criteria: The studies should (a) include probability samples targeting the general population of preschool children in Germany or specific federal states to avoid sample selectivity bias and ensure coverage of the full range of target outcomes, thereby enhancing the generalizability of the results; (b) use daycare centers as the primary sampling unit to facilitate the estimation of multilevel design parameters; (c) be conducted in the year 2000 or later to ensure that the design parameters are current; and (d) apply an observational study design (i.e., not an experimental design) to align with the methodology used in most studies providing design parameters in the school context (e.g., Bloom et al., [<reflink idref="bib11" id="ref147">11</reflink>]; Brunner et al., [<reflink idref="bib15" id="ref148">15</reflink>]; Hedges &amp; Hedberg, [<reflink idref="bib42" id="ref149">42</reflink>]; see Stallasch et al., [<reflink idref="bib94" id="ref150">94</reflink>] and 2024 for comprehensive overviews), which also assures that design parameters are not affected by an educational intervention (see Jacob et al., [<reflink idref="bib47" id="ref151">47</reflink>]). To identify studies meeting these inclusion criteria, we searched (on January 29, 2024) two key German electronic data repositories (OSM A2). The search in the German Network of Educational Research Data repository[<reflink idref="bib5" id="ref152">5</reflink>] returned a total 15 studies (Tables A1 and A2). The search in the Research Data Centre (FDZ)[<reflink idref="bib6" id="ref153">6</reflink>] returned a total of six studies (Table A3). After removing eight duplicate studies (Table A4), we screened the description of 13 studies. We excluded 10 of these 13 studies because four studies did not cover the general target population, three studies were not based on samples with daycare centers as the primary sampling unit, one study implemented an experimental design, and two studies were collected before the year 2000 (Table A5). IPD from the remaining three studies were sought for retrieval and obtained: (a) the National Survey on Education, Care, and Development in Early Childhood (NUBBEK; Tietze et al., [<reflink idref="bib99" id="ref154">99</reflink>]), (b) the study on Educational Processes, Competence Development, and Selection Decisions in Preschool and School Age (BIKS; Weinert et al., [<reflink idref="bib105" id="ref155">105</reflink>]), and (c) the starting cohort of preschool children (starting cohort 2) of the National Educational Panel Study (NEPS; NEPS Network, [<reflink idref="bib66" id="ref156">66</reflink>]; Blossfeld &amp; Roßbach, [<reflink idref="bib12" id="ref157">12</reflink>]).</p> <p>Graph: Fig. 1 Flow Diagram to Identify the Large-Scale Studies That Were Used to Estimate Design Parameters for Preschool Children. Note. Adapted from the PRISMA 2020 (Page et al., [<reflink idref="bib69" id="ref158">69</reflink>]) and PRISMA IPD standards (Stewart et al., [<reflink idref="bib97" id="ref159">97</reflink>])</p> <hd id="AN0182226205-14">Studies and Samples</hd> <p>NUBBEK (carried out in eight federal states of Germany in 2010 and 2011) is a cross-sectional study with two samples representing two-year and four-year-old children (Leyendecker et al., [<reflink idref="bib54" id="ref160">54</reflink>]). BIKS (carried out in two federal states of Germany) is a longitudinal panel study with two cohorts (i.e., preschool children and students in primary education) that started in 2006. We used the data from the cohort of three-year old preschool children that were followed up to age six before they entered primary school. The time intervals between waves of measurement were about six months (Homuth et al., [<reflink idref="bib46" id="ref161">46</reflink>]). To estimate design parameters, we defined separate BIKS samples for each wave of measurement because the planned missing data design that was applied in this study resulted in a large number of missing values. Notably, for this reason we excluded data from the second wave of measurement because reliable imputation of missing data was not possible. Further, in each remaining wave of measurement we excluded data from those daycare centers for which no target outcome data were available. Finally, NEPS (carried out in all federal states of Germany) is an ongoing multi-cohort longitudinal panel study (Artelt &amp; Sixt, [<reflink idref="bib1" id="ref162">1</reflink>]; Blossfeld &amp; Roßbach, [<reflink idref="bib12" id="ref163">12</reflink>]). We used the data for the cohort of 4-year old children from the first two waves of measurement when they were in preschool. Data collection started in 2011; the time interval between waves was about nine months. All three studies followed a multistage sampling procedure where a random selection of daycare centers was drawn in the first stage of sampling. All children who attended the selected centers and who fulfilled the age-based inclusion criteria were invited to participate in a specific study. We used the IPD from the children who actually participated in these studies for our statistical analyses. As information on groups within daycare centers was missing for some children in the second wave of measurement in the NEPS, we defined separate samples for the first and second wave of measurement to estimate design parameters. In summary, the IPD of NUBBEK, BIKS, and NEPS were used to estimate design parameters for children aged 2, 3, 4, 5, and 6 years. Information on key characteristics for each sample can be found in Table 1.[<reflink idref="bib7" id="ref164">7</reflink>]</p> <p>Table 1 Description of the Samples that Were Applied to Estimate Design Parameters</p> <p> <ephtml> &lt;table frame="hsides" rules="groups"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left"&gt;&lt;p&gt;Statistic&lt;/p&gt;&lt;/th&gt;&lt;th align="left" colspan="5"&gt;&lt;p&gt;BIKS&lt;sup&gt;a&lt;/sup&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left" colspan="2"&gt;&lt;p&gt;NUBBEK&lt;/p&gt;&lt;/th&gt;&lt;th align="left" colspan="2"&gt;&lt;p&gt;NEPS&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th align="left" /&gt;&lt;th align="left"&gt;&lt;p&gt;Wave 1&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Wave 2&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Wave 3&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Wave 4&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Wave 5&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;2-year-olds&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;4-year-olds&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Wave 1&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Wave 2&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;italic&gt;Sample Size&lt;/italic&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;italic&gt;n&lt;/italic&gt;&lt;sub&gt;&lt;italic&gt;Children&lt;/italic&gt;&lt;/sub&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;518&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;468&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;460&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;457&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;425&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;564&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;714&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2928&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2727&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;italic&gt;n&lt;/italic&gt;&lt;sub&gt;&lt;italic&gt;Groups&lt;/italic&gt;&lt;/sub&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#8211;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#8211;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#8211;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#8211;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#8211;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#8211;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#8211;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;719&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;690&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;italic&gt;n&lt;/italic&gt;&lt;sub&gt;&lt;italic&gt;Centers&lt;/italic&gt;&lt;/sub&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;89&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;81&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;80&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;79&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;72&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;202&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;220&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;277&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;275&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;italic&gt;Mdn n&lt;/italic&gt;&lt;sub&gt;&lt;italic&gt;Children.Group&lt;/italic&gt;&lt;/sub&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#8211;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#8211;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#8211;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#8211;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#8211;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#8211;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#8211;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;4&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;3&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;italic&gt;Mdn n&lt;/italic&gt;&lt;sub&gt;&lt;italic&gt;Children.Center&lt;/italic&gt;&lt;/sub&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;6&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;6&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;6&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;6&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;6&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;3&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;3&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;10&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;9&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;italic&gt;Mdn n&lt;/italic&gt;&lt;sub&gt;&lt;italic&gt;Group.Center&lt;/italic&gt;&lt;/sub&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#8211;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#8211;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#8211;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#8211;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#8211;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#8211;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#8211;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;italic&gt;Age (in months)&lt;/italic&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;italic&gt;M&lt;/italic&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;42.2&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;54.0&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;60.0&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;66.0&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;72.0&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;33.1&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;53.9&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;57.8&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;66.7&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;italic&gt;SD&lt;/italic&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;4.1&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;4.1&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;4.1&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;4.1&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;4.0&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.0&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;3.7&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;3.9&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;3.8&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;% girls&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;48&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;48&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;47&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;47&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;47&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;49&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;51&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;49&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;50&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;% migration&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;20&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;18&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;19&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;19&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;16&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;19&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;30&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;31&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;30&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Years of education&lt;/p&gt;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;italic&gt;M&lt;/italic&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;15.0&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;15.0&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;15.0&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;15.0&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;15.1&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;15.6&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;14.9&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;14.5&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;14.5&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;italic&gt;SD&lt;/italic&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.5&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.5&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.5&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.5&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.5&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.8&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;3.0&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.4&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.4&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;% missing data&lt;/p&gt;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;italic&gt;Min&lt;/italic&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.0&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.0&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.0&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.0&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.0&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.0&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.0&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.6&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.1&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;italic&gt;25th Percentile&lt;/italic&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;4.1&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;6.4&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;1.3&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;14.4&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;5.4&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.4&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.1&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.5&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;1.5&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;italic&gt;Mdn&lt;/italic&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;5.0&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;6.4&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;7.6&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;14.7&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;17.5&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.5&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.9&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;10.3&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;10.0&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;italic&gt;75th Percentile&lt;/italic&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;10.2&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;11.7&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;29.3&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;18.3&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;18.5&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;4.6&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;3.2&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;15.4&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;14.8&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;italic&gt;Max&lt;/italic&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;19.7&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;39.1&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;29.3&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;45.3&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;39.8&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;18.9&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;3.4&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;23.4&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;28.7&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p> <emph>BIKS</emph> Educational Processes, Competence Development and Selection Decisions in Preschool and School Age. <emph>NUBBEK</emph> National Survey on Education, Care, and Development in Early Childhood. NEPS National Educational Panel Study (NEPS) – starting cohort 2 with kindergarten children. <emph>n</emph><subs><emph>Children</emph></subs> total number of children. <emph>n</emph><subs><emph>Groups</emph></subs> total number of groups within daycare centers. <emph>n</emph><subs><emph>center</emph></subs> total number of daycare centers. <emph>Mdn n</emph><subs><emph>Children.Group</emph></subs> median number of children per group within daycare centers. <emph>Mdn n</emph><subs><emph>Children.Center</emph></subs> median number of children per daycare center. <emph>Mdn n</emph><subs><emph>Group.Center</emph></subs> median number of groups per daycare center. % migration percentage of children with migration background. Years of education highest educational level of education in the family in terms of the completed years of education. % missing data percentage of missing data per variable <sups>a</sups>: Data of the original second wave of measurement of BIKS (when children were on average 48 months old) were not used in the present paper because of the large number of planned missing values. Wave 2 in this table refers to the third wave of measurement of BIKS when children were on average 54 months old</p> <hd id="AN0182226205-15">Measures</hd> <p>NUBBEK, BIKS, and NEPS utilized standardized tests to assess cognitive outcomes, covering the subdomains of early mathematics/science and verbal skills as well as general cognitive skills. Further, NUBBEK also provided parent and teacher report data on children's verbal skills and psycho-motor skills. Reliabilities for cognitive outcomes were 0.39 ≤ <emph>r</emph><subs>tt</subs> ≤ 0.90 (<emph>Mdn</emph> = 0.78) for standardized tests, 0.68 ≤ <emph>r</emph><subs>tt</subs> ≤ 0.88 (<emph>Mdn</emph> = 0.83) for parent reports, and 0.78 ≤ <emph>r</emph><subs>tt</subs> ≤ 0.95 (<emph>Mdn</emph> = 0.88) for educator reports. In addition, all studies provided rich data on children's SEL outcomes as assessed by parent and educator reports. SEL outcome measures targeted the subdomains self-orientation, other-orientation, and task-orientation. Reliabilities for SEL outcomes were 0.35 ≤ <emph>r</emph><subs>tt</subs> ≤ 0.90 (<emph>Mdn</emph> = 0.70) for parent reports and 0.55 ≤ <emph>r</emph><subs>tt</subs> ≤ 0.93 (<emph>Mdn</emph> = 0.81) for educator reports. Tables A6 to A10 in OSM A3 present further details on the applied measures.</p> <hd id="AN0182226205-16">Statistical Analyses</hd> <p>The two-stage approach of individual participant data (IPD) meta-analysis applied in this paper comprised two stages: In Stage 1, missing data were imputed, and design parameters were estimated separately for each sample and wave of measurement (when multiple waves were available). In Stage 2, these estimates were summarized meta-analytically.</p> <hd id="AN0182226205-17">Stage 1: Treatment of Missing Data</hd> <p>The percentage of missing values per variable varied from 0% to 45.3% (Table 1). To deal with missing data we used an adjusted cluster-means imputation approach for multilevel data (Grund et al., [<reflink idref="bib39" id="ref165">39</reflink>]) and generated 50 multiply imputed datasets for each sample (and wave of measurement) using the mice (van Buuren &amp; Groothuis-Oudshoorn, [<reflink idref="bib103" id="ref166">103</reflink>]) and miceadds (Robitzsch &amp; Grund, [<reflink idref="bib78" id="ref167">78</reflink>]) R packages. Notably, the applied multilevel imputation models were compatible with the models employed for estimating single- and multilevel design parameters. Design parameters were pooled across imputations using Rubin's ([<reflink idref="bib79" id="ref168">79</reflink>]) rules. We used the mitml package (Grund et al., [<reflink idref="bib38" id="ref169">38</reflink>]) to combine the estimates into a single set of results and to obtain standard errors that take into account within and between imputation variance. The methods for estimating the within-imputation variance (i.e., the standard errors for single- and multilevel design parameters) are detailed in OSM A4.</p> <hd id="AN0182226205-18">Stage 1: Estimation of Design Parameters</hd> <p></p> <hd id="AN0182226205-19">Single-Level Design Parameters</hd> <p>To estimate single-level design parameters (i.e., <emph>R</emph><sups>2</sups><subs><emph>Total</emph></subs>) for each sample (and wave of measurement), we used the R function lm (R Core Team, [<reflink idref="bib74" id="ref170">74</reflink>]) with ordinary least squares (OLS) to analyze each cognitive or SEL outcome using up to five sets of linear regression models. Model Set 1-SD comprised five SD characteristics, including children's age, gender, migration background, and two measures of parents' socioeconomic status (SES). The available SES measures varied somewhat across studies. We used the highest educational attainment within the family (i.e., the greatest number of years of schooling completed within a family) in all studies, the highest International Socio-Economic Index of Occupational Status within a family (HISEI; Ganzeboom &amp; Treiman, [<reflink idref="bib35" id="ref171">35</reflink>]) in BIKS and NEPS, and family income in NUBBEK. Model Set 2-IB comprised an identical baseline (IB) measure of the target cognitive or socio-emotional outcome. We selected baseline measures with the shortest possible time lag between waves of measurement and the same method of assessment (i.e., standardized test, educator report, or parent report). In Model Set 2-PB we used children's vocabulary knowledge as the PB measure for cognitive outcomes, and children's disruptive behavior (e.g., a child interrupts or disturbs other children) as obtained from parent or educator reports as the PB measure for SEL outcomes assessed with the same method. Notably, most results for Model Set 2-PB were obtained for the second wave of measurement of NEPS. Model Set 3-SD + IB and Model Set 3-SD + PB comprised SD characteristics and the selected IB or PB baseline measure of the target outcome. Using these procedures, we estimated a total of <emph>k</emph> = 237 single-level design parameters.</p> <hd id="AN0182226205-20">Two- and Three-Level Design Parameters</hd> <p>To ensure reliable estimation of multilevel design parameters, particularly the variance components of random effects, we followed the recommendation of McNeish and Stapleton ([<reflink idref="bib60" id="ref172">60</reflink>]) by using restricted maximum likelihood estimation (REML) as implemented in the R package lme4 (Bates et al., [<reflink idref="bib5" id="ref173">5</reflink>]). Specifically, we analyzed up to six sets of multilevel models for each sample (and wave of measurement). Model Set 0 comprised so-called "empty" multilevel models to estimate ICCs. We estimated five additional sets of multilevel models (Model Sets 1-SD, 2-IB, 2-PB, 3-SD + IB, and 3-SD + PB) for the covariates that we previously applied to estimate single-level design parameters. Given the available data and the study-specific sampling designs, we used the IPD of NEPS, BIKS, and NUBBEK to specify two-level models (children nested in daycare centers) for estimating the two-level design parameters (i.e., ρ<subs><emph>Center</emph></subs>, <emph>R</emph><sups>2</sups><subs><emph>Child</emph></subs><emph>,</emph> and <emph>R</emph><sups>2</sups><subs><emph>Center</emph></subs>) for each outcome and model set. In addition, we drew on the NEPS data to specify three-level models (children nested in groups, with groups nested in daycare centers) for estimating three-level design parameters (i.e., ρ<subs><emph>Group</emph></subs>, ρ<subs><emph>Center</emph></subs>, <emph>R</emph><sups>2</sups><subs><emph>Child</emph></subs><emph>, R</emph><sups>2</sups><subs><emph>Group</emph></subs> and <emph>R</emph><sups>2</sups><subs><emph>Center</emph></subs>) for each outcome and model set. All covariates were assessed at the child level. In addition to covariates at the child level, for each multilevel model set we entered daycare center means into the two-level models, and group and daycare center means into the three-level models. Further, we applied group-mean centering: covariates at the child level were centered around their respective daycare center/group means in the two-/three-level models, and group means were centered around their respective daycare center means in the three-level models (Raudenbush &amp; Bryk, [<reflink idref="bib76" id="ref174">76</reflink>]). Using these procedures, we estimated a total of <emph>k</emph> = 589/226 two-level/three-level design parameters.</p> <hd id="AN0182226205-21">Stage 2: Meta-Analytic Integration</hd> <p>To synthesize the sample-specific design parameters obtained in Stage 1, we provide numerous meta-analytical summaries (see Tables A11 to A15) using the R package metafor (Viechtbauer, [<reflink idref="bib104" id="ref175">104</reflink>]). Because random-effects models cannot be expected to reliably gauge the heterogeneity of a specific (true) design parameter when less than <emph>k</emph> = 10 observed design parameters are available (Langan et al., [<reflink idref="bib52" id="ref176">52</reflink>], p. 95), we applied (multivariate) fixed-effects models (Rice et al., [<reflink idref="bib77" id="ref177">77</reflink>]) when 2 ≤ <emph>k</emph> &lt; 10, and (multivariate) random-effects models (Hedges, [<reflink idref="bib41" id="ref178">41</reflink>]) when <emph>k</emph> ≥ 10 (see OSM A3 for further details).</p> <p>To assess the heterogeneity of design parameters, we computed the 95% prediction interval (95% PI), which provides a plausible range of values in which the true value of a specific design parameter of about 95% of all relevant populations will fall. We also estimated the standard deviation of the random effects (σ), <emph>I</emph><sups>2</sups>, and the Q statistic as additional heterogeneity measures (Borenstein et al., [<reflink idref="bib13" id="ref179">13</reflink>]). Notably, all design parameters have a theoretical range between zero and one. When the lower (or upper) bound value obtained for the 95% confidence interval (95% CI) for the meta-analytic average or the 95% PI was below zero (or above one), we truncated these values. Several design parameters were obtained for the same sample. To take the resulting within-sample dependencies into account we used the R package clubSandwich (Pustejovsky, [<reflink idref="bib73" id="ref180">73</reflink>]) to impute a working covariance matrix for the observed effect sizes (Hedges, [<reflink idref="bib41" id="ref181">41</reflink>]). We used <emph>r</emph> = 0.70 as a reasonable upper-bound estimate for the within-sample correlations among design parameters. Because we used an estimated working covariance matrix rather than an empirical one, we conducted sensitivity analyses (Hedges, [<reflink idref="bib41" id="ref182">41</reflink>]). These analyses corroborated that the meta-analytic statistics were fairly robust against the different values chosen for the correlation among design parameters (see OSM A5 for details).</p> <hd id="AN0182226205-22">Results and Discussion</hd> <p>Figure 2 presents the point estimates of design parameters that we estimated in Stage 1 as well as their meta-analytic summaries from Stage 2. Drawing on the empirical results, we discuss the use of the present design parameters for a priori power analyses to plan randomized intervention studies with preschool children, along with specific recommendations for early education researchers.</p> <p>Graph: Fig. 2 Single-, Two-, and Three-Level Design Parameters by Outcome Domain and Assessment Method. Note. The grey circles show the point estimates of design parameters. The black circles depict the meta-analytic averages; error bars in black/blue color depict their 95% CIs/95% PIs. Lower/upper bound values of these intervals outside the possible value range were truncated at 0/1, respectively. PIs were only computed when k ≥ 10 (see Method section). 1-SD = Model set 1 with five socio-demographic characteristics as covariates; 2-IB = Model set 2 with one identical baseline measure as covariate; 2-PB = Model set 2 with one proxy baseline measure as covariate; 3-SD + IB = Model Set 3 with five socio-demographic characteristics and one identical baseline measure as covariates; 3-SD + PB = Model Set 3 with five socio-demographic characteristics and one proxy baseline measure as covariates</p> <hd id="AN0182226205-23">Match Design Parameters to Key Characteristics of the Target Intervention</hd> <p>When carrying out an a priori power analysis for a randomized intervention study it is important to strive for an ideal match between the selected set of design parameters on the one side and key characteristics of the planned intervention on the other (Bloom et al., [<reflink idref="bib11" id="ref183">11</reflink>]; Brunner et al., [<reflink idref="bib15" id="ref184">15</reflink>]; Hedges &amp; Hedberg, [<reflink idref="bib42" id="ref185">42</reflink>]; Hedges &amp; Rhoads, [<reflink idref="bib43" id="ref186">43</reflink>]; Stallasch et al., [<reflink idref="bib94" id="ref187">94</reflink>], [<reflink idref="bib95" id="ref188">95</reflink>]). For one, this involves selecting design parameters that offer the best match to the assumptions of how the data of the planned study will be clustered. Specifically, single-level design parameters are required for lab-based randomized experiments where nonclustered data are assumed. Two-level parameters (children in daycare centers) and three-level parameters (children in groups, with groups in daycare centers) are needed for field-based randomized experiments.</p> <p>In addition, we often observed substantial heterogeneity of design parameters, which was reflected in the 95% PIs in the meta-analytic summaries (Fig. 2). For example, the 95% PI of ρ<subs><emph>Center</emph></subs> obtained for two-level designs was [0.00, 0.24] for cognitive outcomes assessed with standardized tests, and [0.01, 0.06]/[0.00, 0.18] for SEL outcomes assessed with parent/educator reports. Given their heterogeneity, design parameters should be matched to the target population and outcome of the intervention, ideally using a set of design parameter point estimates (e.g., for a two-level CRT: ρ<subs>Center</subs>, <emph>R</emph><sups>2</sups><subs>Child</subs>, and <emph>R</emph><sups>2</sups><subs>Center</subs>) that were obtained for the same or a very similar target measure and target age group.</p> <p>Yet, certain circumstances may limit this endeavor, such as the unavailability of suitable estimates for a specific data structure, measure, age group, or covariate set. Here, two strategies may be helpful. First, virtually all daycare centers in Germany use some sort of grouping for their children (e.g., by age, or by a specific team of educators). Hence, random assignment of daycare centers ideally requires three-level design parameters. When suitable parameters are not available (i.e., information on ρ<subs><emph>Group</emph></subs> and <emph>R</emph><sups>2</sups><subs>Group</subs> is missing), relevant two-level design parameters can be used because between-group differences in cognitive and SEL outcomes were often (very) small (Fig. 2), and the variance component attributable to these differences is reflected in the present two-level parameters (Zhu et al., [<reflink idref="bib111" id="ref189">111</reflink>]). Thus, planning a randomized experiment based on the available two-level information (though having a three-level data structure) will likely lead to very similar results (in terms of the <emph>MDES</emph> or required sample size), particularly when including covariates at the child level (Zhu et al., [<reflink idref="bib111" id="ref190">111</reflink>]).</p> <p>Second, the meta-analytic results obtained for broader outcome domains, subdomains, measures, or age groups which provide the best match to the application context can be used in power analysis when point estimates of design parameters are not (a) available, (b) useful (e.g., when the intervention targets verbal skills rather than a specific measure), or (c) reliable (i.e., when they have large standard errors). Doing so helps to determine sample sizes, power rates, or <emph>MDES</emph> values assuming typical values (i.e., meta-analytic averages) and plausible lower and upper bounds by drawing on the 95% PIs (if available). The 95% PIs may become especially important when researchers plan randomized experiments with target outcome measures or target populations of preschool children that differ substantially from those applied in the present study.</p> <hd id="AN0182226205-24">Beware of Between-Group and Between-Center Differences in the Target Outcome</hd> <p>Everything else being equal, clustering renders multilevel field-based studies less sensitive than single-level lab-based studies, where data are assumed not to be clustered. The most important design parameters for planning field-based randomized experiments in preschool settings therefore comprise information on between-group and between-daycare center differences in children's outcomes. When using two-level designs, the meta-analytic averages for between-center differences lay in the range ρ<subs><emph>Center</emph></subs> = 0.03 (parent report of SEL outcomes) and ρ<subs><emph>Center</emph></subs> = 0.24 (educator report of cognitive outcomes). When using three-level designs, the meta-analytic averages for between-center differences in three-level designs varied between ρ<subs><emph>Center</emph></subs> = 0.01 (parent report of SEL outcomes) and ρ<subs><emph>Center</emph></subs> = 0.11 (standardized tests cognitive outcomes).[<reflink idref="bib8" id="ref191">8</reflink>] Meta-analytic averages of between-group differences (within daycare centers) lay in the range ρ<subs><emph>Group</emph></subs> = 0.01 (parent report of SEL outcomes) and ρ<subs><emph>Group</emph></subs> = 0.09 (educator report of SEL outcomes).</p> <p>The impact of between-group and/or between-center differences on the sensitivity of randomized experiments with preschool children can be illustrated using the <emph>MDES</emph> (two-sided testing; α = 0.05; 1 − β = 0.80) as a measure of design sensitivity. For example, when using the meta-analytic average across all SEL outcomes assessed with educator reports, a two-level CRT (with balanced allocation of 50 daycare centers to experimental groups; 20 children per center) was considerably less sensitive (<emph>MDES</emph> = 0.30) than a corresponding single-level lab-based study with <emph>N</emph> = 1,000 children and individual random assignment (<emph>MDES</emph> = 0.18), although the total sample sizes of these studies were the same. Notably, Figure A6 in the OSM further illustrates the impact of between-group and/or between-center differences on the sensitivity of randomized experiments. To sum up, cognitive and SEL outcomes may demonstrate substantial between-group and between-center differences that imply that randomized experiments with preschool children may become noticeably less sensitive when using field-based intervention designs (e.g., two-level CRTs) rather than carrying out single-level lab-based randomized experiments.</p> <hd id="AN0182226205-25">Use Pre-Treatment Covariates to Improve the Sensitivity of Randomized Experiments</hd> <p>Adjusting for pre-treatment covariates to estimate the intervention effect in a randomized experiment does not change the size or nature of the estimated treatment effect. However, covariates can substantially improve the sensitivity of randomized experiments (Lin, [<reflink idref="bib55" id="ref192">55</reflink>]; Maxwell et al., [<reflink idref="bib58" id="ref193">58</reflink>]; Porter &amp; Raudenbush, [<reflink idref="bib72" id="ref194">72</reflink>]). For example, everything else being equal, larger values of <emph>R</emph><sups>2</sups><subs>Total</subs>, <emph>R</emph><sups>2</sups><subs>Child</subs>, <emph>R</emph><sups>2</sups><subs>Group</subs>, and <emph>R</emph><sups>2</sups><subs>Center</subs> lead to smaller values of the <emph>MDES.</emph> Importantly, when covariates are to be used it is essential to ensure that these covariates are measured before random assignment (i.e., pre-treatment covariates). Otherwise, covariates could be affected by the treatment and therefore act as "bad controls" when used to estimate an adjusted treatment effect (Cinelli et al., [<reflink idref="bib22" id="ref195">22</reflink>]).</p> <p>When used as pre-treatment covariates, all the covariate sets we investigated have the capacity to improve (at least somewhat) the sensitivity of randomized experiments with preschool children. The point estimates as well as the meta-analytic averages obtained for the various covariate sets showed that covariates typically explained some— and often a considerable proportion— of the variance in the outcome in total as well as at the child, group and center levels (see Fig. 2). Notably, the combination of SD and IB measures or SD and PB measures typically improved the explained amount of variance compared to using either set alone. Thus, when feasible, using Model Set 3-SD + IB or Model Set 3-SD + PB will generally lead to the most significant improvements in design sensitivity in most intervention scenarios compared to designs that do not incorporate covariates. For example, using model Set 3-SD + PB for cognitive outcomes measured with standardized tests improved the <emph>MDES</emph> from 0.18 to 0.16 in single-level designs, and from 0.33 to 0.26 in two-level and three-level CRTs, based on the sample specifications noted above. These improvements were achieved using the meta-analytic averages (shown in Fig. 2) of <emph>R</emph><sups>2</sups><subs>Total</subs> in single-level designs, <emph>R</emph><sups>2</sups><subs>Child</subs> and <emph>R</emph><sups>2</sups><subs>Center</subs> (along with ρ<subs>Center</subs>) in two-level CRTs, and <emph>R</emph><sups>2</sups><subs>Child</subs>, <emph>R</emph><sups>2</sups><subs>Group</subs> and <emph>R</emph><sups>2</sups><subs>Center</subs> (along with ρ<subs>Group</subs> and ρ<subs>Center</subs>) in three-level CRTs. Additional scenarios are illustrated in OSM A6.</p> <p>Importantly, the ability of covariates to improve the sensitivity of two- or three-level randomized experiments depends on (a) the level at which random assignment is implemented, (b) where the covariate is located, and (c) the degree of clustering in the data (Bulus &amp; Sahin, [<reflink idref="bib19" id="ref196">19</reflink>]; Konstantopoulos, [<reflink idref="bib50" id="ref197">50</reflink>]). Specifically, CRTs randomly assign entire daycare centers to experimental groups, whereas MSRTs offer more flexibility since random assignment can apply to individual children or groups within centers. Additionally, power analyses for MSRTs require assumptions about how much the intervention effect varies across daycare centers and how much of this variation can be explained by covariates (see also Discussion section; Dong &amp; Maynard, [<reflink idref="bib31" id="ref198">31</reflink>]; Konstantopoulos, [<reflink idref="bib49" id="ref199">49</reflink>]). Given these complexities and the limitations of the applied data, which did not allow for the estimation of these additional design parameters, we focus our discussion on CRTs that involve the random assignment of entire daycare centers.[<reflink idref="bib9" id="ref200">9</reflink>] When between-center differences in CRTs are ρ<subs>Center</subs> ≥ 0.10, large proportions of explained variance at the daycare center level (in terms of <emph>R</emph><sups>2</sups><subs>Center</subs>) result in significant improvements in design sensitivity. For such clustered data, covariates at the daycare center level outperform (group-mean centered) covariates at the child or group levels in many practical applications (Bulus &amp; Sahin, [<reflink idref="bib19" id="ref201">19</reflink>]; Konstantopoulos, [<reflink idref="bib50" id="ref202">50</reflink>]). Conversely, when between-center differences in CRTs are negligible to small (ρ<subs>Center</subs> &lt; 0.10), large proportions of explained variance at the daycare center level do not result in significant improvements in design sensitivity. This becomes evident in Figure A6 when looking at three-level designs with SEL outcomes assessed with parent reports. There is not much of a difference in <emph>MDES</emph> values between the models that do (all <emph>MDES</emph>s = 0.19) and do not use covariates (<emph>MDES</emph> = 0. 20) because the between-center differences were very small (ρ<subs><emph>Center</emph></subs> = 0.01). Moreover, as demonstrated by Konstantopoulos ([<reflink idref="bib50" id="ref203">50</reflink>]) and Bulus and Sahin ([<reflink idref="bib19" id="ref204">19</reflink>]), when between-center and between-group differences are very small (e.g., ρ<subs>Center</subs> ≤ 0.02 and ρ<subs>Group</subs> ≤ 0.02), (group-mean centered) covariates at the child level can often outperform those at the daycare center level in improving design sensitivity in many practical settings. In summary, adjusting for pre-treatment covariates in single-level lab-based experiments and multilevel CRTs and MSRTs typically improves the sensitivity of randomized experiments, at least to some extent (Bulus &amp; Sahin, [<reflink idref="bib19" id="ref205">19</reflink>]; Konstantopoulos, [<reflink idref="bib49" id="ref206">49</reflink>], [<reflink idref="bib50" id="ref207">50</reflink>]; Maxwell et al., [<reflink idref="bib58" id="ref208">58</reflink>]). To study the impact of covariates on design sensitivity, we recommend carrying out power analyses both assuming the application of different pre-treatment covariate sets and without covariates. This approach will also allow the detection of the rare situations when covariates do not explain sufficient variance in the outcome to outweigh the loss in degrees of freedom in the statistical tests, for example when using very small samples (Konstantopoulos, [<reflink idref="bib50" id="ref209">50</reflink>]; Maxwell et al., [<reflink idref="bib58" id="ref210">58</reflink>]).</p> <hd id="AN0182226205-26">Account for the Statistical Uncertainty of Design Parameters in Power Analyses</hd> <p>Taking advantage of large-scale probability samples implied that most standard errors obtained for the point estimates of the present design parameters were relatively small (see Figure A2).[<reflink idref="bib10" id="ref211">10</reflink>] How the statistical uncertainty of design parameters (as reflected by their standard errors and 95% CIs) is taken into account in a power analysis depends on the risk preferences of the intervention researchers and their funding agencies as well as the cost structure of the project (Jacob et al., [<reflink idref="bib47" id="ref212">47</reflink>]). For example, for high-profile interventions with strong interest in detecting the intervention effect (e.g., for policy implementation), researchers may want to apply conservative estimates of the <emph>MDES</emph> and required samples sizes. To this end, they can draw on the upper bound estimates of the 95% CI's for ρ<subs><emph>Group</emph></subs> and/or ρ<subs><emph>Center</emph></subs> in combination with the lower bound estimates of <emph>R</emph><sups>2</sups><subs>Total</subs><emph>, R</emph><sups>2</sups><subs>Child</subs>, <emph>R</emph><sups>2</sups><subs>Group</subs>, and <emph>R</emph><sups>2</sups><subs>Center.</subs> When standard errors are relatively large (e.g., larger than 0.05),[<reflink idref="bib11" id="ref213">11</reflink>] we recommend also looking for alternative sets of point estimates of design parameters that (a) match the target outcome and target population (e.g., other outcomes belonging to the same subdomain). In such cases, researchers can also use meta-analytic averages and their 95% CIs that were obtained for higher aggregate levels or (if standard errors for these averages are also large) alternative aggregate levels that could be estimated with higher statistical precision. For example, one could use the age-specific meta-analytic average for a subdomain rather than the age-specific point estimates obtained for a specific measure or apply the domain-specific meta-analytic average for a certain age-group rather than the meta-analytic average for a specific subdomain.</p> <hd id="AN0182226205-27">Application</hd> <p></p> <hd id="AN0182226205-28">How Should Appropriate Design Parameters be Selected?</hd> <p>To support power analyses for single-level and multilevel randomized experiments with preschool children, we created OSM B as a rich source with (a) point estimates (Tables B.CT, B.CP, B.CE, B.SP, and B.SE) and (b) meta-analytic summaries (Tables B.CT.M, B.CP.M, B.CE.M, B.SP.M, and B.SE.M) of design parameters. As discussed above, we recommend matching the design parameters to key characteristics of the target intervention. This involves (a) potential clustering of the data, (b) the target age group, (c) the target outcome domain, (d) the method to be used for assessing the outcome, and (e) the target measure. To find appropriate matches in OSM B, intervention researchers can first consult Tables A6 to A10 in OSM A that provide short descriptions of the measures for which point estimates of design parameters are available.</p> <p>When the target measures or other target characteristics of the planned study (e.g., target age group) do not match well to the characteristics of the studies that were used to estimate the present design parameters or when the available point estimates are associated with large standard errors, we recommend applying meta-analytic summaries of design parameters that match the target intervention. To do so, intervention researchers can consult Tables A11 to A15 that present overviews of available meta-analytic summaries.</p> <p>Finally, to select the design parameters or their meta-analytic summaries from the spreadsheets in OSM B, we recommend using the interactive filter functions of the spreadsheet software or adapting the R code provided in the OSF.</p> <hd id="AN0182226205-29">Application Scenarios</hd> <p>This section presents two scenarios to illustrate the use of the present design parameters for planning the sample size of randomized experiments with preschool children. For each scenario, we assumed a balanced design with children (Scenario 1) and daycare centers (Scenario 2) randomly assigned to the experimental groups in equal shares. Further, we set the desired power at 1 − β = 0.80 and used a two-tailed test (with significance level α = 0.05) to allow testing whether the intervention has unexpected negative effects on the outcome (Bland &amp; Altman, [<reflink idref="bib6" id="ref214">6</reflink>]). To compute the required sample sizes, we used the R package PowerUpR (Bulus et al., [<reflink idref="bib20" id="ref215">20</reflink>]) that is based on the power formulas provided in Dong and Maynard ([<reflink idref="bib31" id="ref216">31</reflink>]). Notably, other software tools that can be used for this purpose include the Excel worksheets by Dong and Maynard ([<reflink idref="bib31" id="ref217">31</reflink>]), the PowerUpR shiny app (Ataneka et al., [<reflink idref="bib2" id="ref218">2</reflink>]), and the Optimal Design software (Spybrook et al., [<reflink idref="bib93" id="ref219">93</reflink>]).[<reflink idref="bib12" id="ref220">12</reflink>] Tables A17 (Scenario 1) and A18 (Scenario 2) show the estimates of design parameters that were applied in the power analyses.</p> <hd id="AN0182226205-30">Scenario 1: How Many Children Are Required for a Lab-Based Randomized Experiment?</hd> <p>A research team developed an intervention to foster five- to six-year-old children's performance on standardized tests of verbal skills. To investigate its impact, the team plans a single-level lab-based randomized experiment to test whether the new method is generally effective under well-controlled conditions in the lab. The team expects that the intervention should have at least an effect that lies in the average range of effects relative to other randomized experiments on cognitive outcomes. Drawing on Kraft ([<reflink idref="bib51" id="ref221">51</reflink>]), the team therefore considers an intervention effect of <emph>SMD</emph> = 0.15 (i.e., a "medium" effect) as meaningful. The objective of the research team is to ensure that the intervention study can detect a treatment effect of <emph>MDES</emph> = 0.15 (with 1 − β = 0.80 and α = 0.05). Because the intervention targets verbal skills (and not a specific measure), the team selects meta-analytic averages of single-level design parameters for verbal skills (see Table B.CT.M in OSM B). Notably, the team uses meta-analytic estimates obtained by averaging across all age groups, as no meta-analytic summary for the target age group is available that includes all covariate sets. The research team begins by considering designs without covariates and learns that 1,398 children would be required to achieve <emph>MDES</emph> = 0.15 (Fig. 3a). Hence, 699 children should be randomly assigned to the intervention group, and 699 children to the control group. The team also tests the influence of different covariate sets (in terms of <emph>R</emph><sups>2</sups><subs><emph>Total</emph></subs>) on the required sample sizes. This results in a total sample size ranging from <emph>N</emph> = 916 when controlling for Set 3-SD + IB to <emph>N</emph> = 1,150 when controlling for Set 2-PB. The team also wants to take into account the statistical uncertainty associated with the applied design parameters and therefore determines conservative lower bound/optimistic upper bound estimates for <emph>N</emph> by drawing on the corresponding lower bounds/upper bounds of the 95% CIs of the meta-analytic averages of <emph>R</emph><sups>2</sups><subs>Total</subs>. When using lower bound values, the estimated total sample size lies between <emph>N</emph> = 1,034 (Set 3-SD + IB) and <emph>N</emph> = 1,200 (Set 1-SD). When using upper bound values, the research team would need between <emph>N</emph> = 798 (Set 3-SD + IB) and <emph>N</emph> = 1,118 (Set 2-PB) children. In sum, when it is not possible to use covariates, a total <emph>N</emph> of 1,398 preschool children is required to achieve <emph>MDES</emph> = 0.15. When covariates are an option, the required sample size depends on the risk preferences of the team. For example, opting for a conservative approach, the team should recruit a total <emph>N</emph> of 1,034 preschool children when using Set 3-SD + IB as covariates.</p> <p>Graph: Fig. 3 Results from Power Analyses to Estimate the Required Sample Size to Achieve MDES = 0.15 for (a) the Single-Level Randomized Experiment in Scenario 1 and (b) the CRT in Scenario 2 When Using Meta-Analytic Summaries of Design Parameters obtained for Different Sets of Covariates. Note. MDES = minimum detectable effect size with level of statistical significance α =.05 and power of 1 − β =.80. The value of the MDES = 0.15 was chosen based on Kraft ([<reflink idref="bib51" id="ref222">51</reflink>]), who suggested that an intervention effect size of a standardized mean difference of 0.15 (i.e., a medium" effect) is meaningful. N = total sample size. J = number of daycare centers. In Scenario 2 we assumed that 20 children are sampled in each daycare center, and, thus, N = J ⋅ 20. The points show the estimated sample sizes when drawing on the meta-analytic averages of design parameters. The error bars for Scenario 1 represents conservative/optimistic estimates when drawing on lower/upper bound values of the 95% CIs for the meta-analytic averages of R2Total. The error bars for Scenario 2 represents conservative/optimistic estimates when drawing on lower/upper bound values of the 95% CIs for the meta-analytic averages of R2Child and R2Center in combination with the upper/lower bound values of the 95% CI for the meta-analytic average of ρCenter. The values of the applied design parameters are shown in Tables A17 and A18 in OSM A. 1-SD = Model set 1 with five socio-demographic characteristics as covariates; 2-IB = Model set 2 with one identical baseline measure as covariate; 2-PB = Model set 2 with one proxy baseline measure as covariate; 3-SD + IB = Model Set 3 with five socio-demographic characteristics and one identical baseline measure as covariates; 3-SD + PB = Model Set 3 with five socio-demographic characteristics and one proxy baseline measure as covariates</p> <hd id="AN0182226205-31">Scenario 2: How Many Daycare Centers Are Required for a Two-Level CRT?</hd> <p>The research team was successful and found that their intervention improved children's verbal skills with <emph>SMD</emph> = 0.15. The team now plans to study whether the intervention also works when it is implemented by educators in the regular preschool context. To avoid unintentionally exposing children in the control group to the intervention, the research team wants to carry out a two-level CRT and to randomize whole daycare centers to the intervention or control group. The team intends to sample 20 children per daycare center. The researchers are now interested in <emph>J</emph>, the number of daycare centers necessary to detect an intervention effect that is identical to their lab-based study (<emph>SMD</emph> = 0.15). The team employs again meta-analytic averages of two-level design parameters for verbal skills (as obtained by averaging across all age groups) from Table B.CT.M from OSM B in their power analyses. The research team initially considers designs without covariates, uses the meta-analytic average of ρ<subs>Center</subs> = 0.15, and learns that the minimum number of daycare centers amounts to <emph>J</emph> = 278 (Fig. 3b). The team again tests the influence of different covariate sets on the required sample size using the meta-analytic averages obtained for ρ<subs>Center,</subs><emph>R</emph><sups>2</sups><subs>Child</subs> and <emph>R</emph><sups>2</sups><subs>Center</subs>. This leads to a total <emph>J</emph> in the range of <emph>J</emph> = 142 (Set 1-SD) to <emph>J</emph> = 202 (Set 2-PB). Given the high profile of the planned study, the team also wants to take into account the statistical uncertainty and therefore determines the upper bound estimates for <emph>J</emph> by using the upper bound of the 95% confidence interval for ρ<subs>Center</subs>, and the lower bound values for <emph>R</emph><sups>2</sups><subs>Child</subs> and <emph>R</emph><sups>2</sups><subs>Center.</subs> When using these values, the research team needs <emph>J</emph> = 238 when not using covariates, and between <emph>J</emph> = 214 and <emph>J</emph> = 294 when using Set 3-SD + IB or Set 2-IB.[<reflink idref="bib13" id="ref223">13</reflink>] To summarize, to achieve <emph>MDES</emph> = 0.15 when covariates are an option, the team should recruit a total of <emph>J</emph> = 142 daycare centers (total <emph>N</emph> = 142 ⋅ 20 = 2,840; Set 1-SD) when relying on the point estimates of the meta-analytic averages, or <emph>J</emph> = 214 daycare centers (<emph>N</emph> = 4,280; Set 3-SD + PB) when opting for a conservative approach.</p> <hd id="AN0182226205-32">General Discussion</hd> <p></p> <hd id="AN0182226205-33">A New Resource to Support Early Childhood Education Research</hd> <p>"Early learning remains one of the most neglected areas of educational research" (OECD, [<reflink idref="bib68" id="ref224">68</reflink>], p. 18). Particularly, there is a strong need for robust evidence about which educational interventions "work" or "work best" to promote preschool children's cognitive and socio-emotional development. Randomized experiments are key for drawing causal conclusions about the impact of such interventions. Hence, lab-based studies are required to develop and refine interventions, and field-based randomized experiments (e.g., CRTs or MSRTs) are beneficial for evaluating their effectiveness in real-world preschool settings. To ensure that these studies provide strong statistical conclusions regarding the size of the intervention effect in the target population, researchers should apply reliable estimates of design parameters when conducting a priori power analyses. However, there has been little research on relevant design parameters with preschool children. To address this significant research gap, we utilized a systematic collection of IPD from four large-scale German samples of preschool children to (a) estimate and (b) meta-analyze design parameters for single-level (non-clustered data), two-level (children in daycare centers), and three-level (children in groups, with groups nested in daycare centers) experimental designs. These parameters target cognitive and SEL outcomes, assessed using three methods: standardized tests, parent ratings, and educator ratings. The design parameters depict between-group and -center differences as well as the proportion of variance in the outcomes explained by five covariate sets including SD, IB, and PB measures, and the combination of SD and IB as well as SD and PB measures. Furthermore, we embedded the results on design parameters in the relevant methodological literature to discuss and illustrate their application in a priori power analyses. In summary, this paper offers a unique and rich resource to support researchers in early childhood education in carrying out a priori power analyses when planning lab-based and field-based randomized experiments.</p> <hd id="AN0182226205-34">Generalizability of the Present Design Parameters Across Countries?</hd> <p>The present design parameters address the target population of children attending daycare centers − the preschool setting that most children aged 3 years or older attend in Germany and many other countries. Comparing the multilevel design parameters from previous research and the present study shows that some country-specific estimates were similar, whereas others were (strikingly) different. For example, average between-center differences were ρ<subs>Center</subs> = 0.17/0.12/0.03 for cognitive outcomes (assessed by standardized tests) and ρ<subs>Center</subs> = 0.05/0.09/0.01 for SEL outcomes (assessed by educator reports) in the UK/Germany/US. One explanation for the large differences between the US and the other countries are the differences in the underlying samples. Design parameters for the United States targeted the population of socioeconomically disadvantaged preschool children (Jacob et al., [<reflink idref="bib47" id="ref225">47</reflink>]; see Appendix A.2 in Spybrook et al., [<reflink idref="bib93" id="ref226">93</reflink>]). By contrast, the studies from the United Kingdom and Germany were based on more heterogeneous samples. Importantly, even small differences in the value of design parameters may make a substantive difference in the required sample sizes. For example, consider a two-level CRT (with a balanced design, 20 children sampled per center; two-sided testing; α = 0.05; 1 − β = 0.80) to test the effectiveness of an intervention on cognitive outcomes (assessed with standardized tests) where the researchers expect a <emph>SMD</emph> = 0.15 and cannot use covariates. Whereas the UK average estimate of ρ<subs>Center</subs> = 0.17 would result in a required number of daycare centers of <emph>J</emph> = 298 (total N = 5,960), the German average estimate of ρ<subs>Center</subs> = 0.12 would result in a requirement of <emph>J</emph> = 232 daycare centers (total N 4,640), a difference of over 1,000 participants.</p> <p>This example illustrates that the present design parameters based on German samples should be applied very cautiously in countries other than Germany. To apply design parameters that align with their local context, researchers in the US or UK could draw on estimates from previous research (Jacob et al., [<reflink idref="bib47" id="ref227">47</reflink>]; Sammons et al., [<reflink idref="bib81" id="ref228">81</reflink>], [<reflink idref="bib82" id="ref229">82</reflink>]; Spybrook et al., [<reflink idref="bib93" id="ref230">93</reflink>]) that we used to generate Figure A1 and which are also available on our OSF. Furthermore, researchers can conduct a pilot study or apply the current analytic strategy using IPD from preschool children in their country (e.g., by utilizing our R code in the OSF as a template). If such alternatives are not feasible, it seems reasonable to use the present design parameters in a priori power analyses (e.g., employing conservative values). Considering their theoretical range from zero to one, the current design parameters define a plausible range relevant to the preschool context and certainly provide better estimates than nonspecific conventional benchmarks, such as categorizing <emph>R</emph><sups>2</sups><subs>Total</subs> = 0.02/0.13/0.26 as "small", "medium", or "large" (see Cohen, [<reflink idref="bib24" id="ref231">24</reflink>], who provided these values as well as a critical discussion).</p> <hd id="AN0182226205-35">Limitations and Outlook</hd> <p>Our study has several limitations that should be addressed in future research and nuances that should be considered when applying the present design parameters. First, the present paper provides design parameters that are highly relevant for power analyses of randomized intervention studies with preschool children. However, it was beyond the scope of our paper to elaborate on the statistical background of a priori power analyses of randomized experiments (e.g., Dong &amp; Maynard, [<reflink idref="bib31" id="ref232">31</reflink>]; Hedges &amp; Rhoads, [<reflink idref="bib43" id="ref233">43</reflink>]) or the process for estimating causal treatment effects once the empirical data have been collected (e.g., Ding, [<reflink idref="bib30" id="ref234">30</reflink>]; Lin, [<reflink idref="bib55" id="ref235">55</reflink>]).</p> <p>Second, we present design parameters necessary for planning randomized intervention studies with (a) simple random assignment of preschool children in single-level lab-based settings (assuming that the data are not clustered) and (b) random assignment in multilevel field settings, particularly in cluster randomized trials (CRTs) where entire daycare centers (including all sampled children or groups) are allocated either to the experimental or control group. Importantly, the present design parameters are also needed to plan MSRTs with blocked random assignment. In multi-site experiments the sample of children or groups is randomly assigned to experimental conditions within daycare centers. In addition to the design parameters (i.e., <emph>R</emph><sups>2</sups>s and ICCs) that we presented in this paper, power analyses of such experiments require information on the expected heterogeneity of the treatment effect across groups or daycare centers, as well as the extent to which covariates may explain this heterogeneity. General benchmarks for the potential magnitude of heterogeneity in treatment effects can be found in Weiss et al. ([<reflink idref="bib106" id="ref236">106</reflink>]). A vital task for future research is to provide systematic knowledge of these parameters for the preschool context, for example by using the analytic approach by Sabol et al. ([<reflink idref="bib80" id="ref237">80</reflink>]).</p> <p>Third, our applied methodology as well as the characteristics of the applied studies were associated with some limitations (see Stallasch et al., [<reflink idref="bib94" id="ref238">94</reflink>], [<reflink idref="bib95" id="ref239">95</reflink>]). Specifically, we analyzed and meta-analyzed design parameters utilizing IPD from several large-scale studies with probability samples, which mitigates potential bias due to sample selectivity and ensures coverage of the full range of target outcomes without variance restrictions. Overall, this approach supports reliable parameter estimation and the generalizability of results. However, we did not apply sampling weights when estimating design parameters because (a) information on weights was only available for NEPS and (b) the application of sampling weights is not possible with the lme4 package (Bates et al., [<reflink idref="bib5" id="ref240">5</reflink>]) that we used for the multilevel analyses. Therefore, our design parameters based on NEPS are representative only for the preschool children included in the present analyses and are likely somewhat less accurate than estimates derived from analyses using sampling weights (e.g., Wenger et al., [<reflink idref="bib107" id="ref241">107</reflink>]). Moreover, the NUBBEK and NEPS studies contained a relatively small number of children/groups per daycare center (see Table 1). However, robust evidence from simulation studies shows that accurate estimates of ρ<subs>Group</subs> and ρ<subs>Center</subs> can still be obtained in such data constellations, as the number of daycare centers exceeded 100 in the two-level models (McNeish, [<reflink idref="bib59" id="ref242">59</reflink>]) and 200 in the three-level models (Zhang et al., [<reflink idref="bib110" id="ref243">110</reflink>]). Nevertheless, larger sample sizes at all levels would have further enhanced the quality of the model parameters (McNeish &amp; Stapleton, [<reflink idref="bib60" id="ref244">60</reflink>]). Additionally, we drew on rather heterogeneous samples. Higher homogeneity may lead to smaller values for all design parameters under investigation due to range restrictions (Miciak et al., [<reflink idref="bib61" id="ref245">61</reflink>]). Thus, between-center and between-group differences— but also the amount of variance explained by covariates— may become smaller in more homogenous samples (Hedges &amp; Hedberg, [<reflink idref="bib42" id="ref246">42</reflink>]). Moreover, the time lag between baseline and outcome measures was between six and 12 months. Longer time lags are typically associated with (somewhat) smaller values of <emph>R</emph><sups>2</sups> at all levels of analyses (Bleidorn et al., [<reflink idref="bib7" id="ref247">7</reflink>]; Stallasch et al., [<reflink idref="bib95" id="ref248">95</reflink>]). Thus, relative to the present design parameters, somewhat larger values of <emph>R</emph><sups>2</sups> should be expected for shorter time intervals in the planned randomized experiment, and smaller values of <emph>R</emph><sups>2</sups> should be expected for longer time intervals. Finally, most measures in the social and behavioral sciences are affected by measurement error. The reliabilities (<emph>r</emph><subs>tt</subs>) for the measures of cognitive and SEL outcomes that we applied to estimate design parameters were fairly typical for applied (experimental) research, where <emph>r</emph><subs>tt</subs> ≥ 0.70 is desirable, but even smaller values may suffice for many research purposes (see Schmitt, [<reflink idref="bib83" id="ref249">83</reflink>] for a discussion). Notably, fallible measures usually lower <emph>R</emph><sups>2</sups> in total as well as at the child, group, or center levels (Cochran, [<reflink idref="bib23" id="ref250">23</reflink>]; Raudenbush &amp; Bryk, [<reflink idref="bib76" id="ref251">76</reflink>]). Thus, our estimates can be considered conservative estimates that may generalize well to empirical data of randomized studies. Of note, measurement error in pre-treatment covariates does not introduce bias in the estimated treatment effect but rather improves the precision with which it can be estimated compared to analyses that do not adjust for covariates (Maxwell et al., [<reflink idref="bib58" id="ref252">58</reflink>]).</p> <p>Fourth, we explored the explanatory power of two PB measures, namely vocabulary knowledge as a proxy baseline measure for cognitive outcomes and children's problem behavior (as assessed by parent or educator reports) as a proxy baseline measure for SEL outcomes. Our results showed that these measures may explain small to substantial proportions of variance in the outcome measures at all levels of analyses. Hence, PB measures may be a useful alternative to IB measures in the preschool context, where it may not be possible to apply identical pretest measures because of the strong developmental dynamics in cognitive and SEL outcomes among very young children. However, our design parameters for PB are confined to the applied or very similar measures. Thus, future research may profit from examining the explanatory power of a broader range of PB measures of cognitive or SEL outcomes.</p> <p>Fifth, we provided guidance and illustrative examples on how to account for the statistical uncertainty of design parameters in power analyses of randomized studies. Accordingly, we offered 95% CIs for point estimates and meta-analytic averages of design parameters, along with 95% PIs to depict the spread of the distribution of design parameters (in random-effects meta-analyses). The meta-analytic summaries—but not the point estimates—of design parameters were based on estimated working covariances which we computed using a reasonable upper-bound estimate of the correlation <emph>r</emph> among design parameters. Our sensitivity analyses showed that the standard errors of the meta-analytic averages and the standard deviations of the random effects (σ) increased with increasing values of <emph>r.</emph> However, the observed increase of these parameters was in most cases (very) small. The standard errors of the meta-analytic averages enter the estimation of 95% CIs and 95% PIs, and the standard deviations enter 95% PIs. Thus, most—but not all—95% CIs and 95% PIs can be considered to be fairly robust against the different values chosen for <emph>r</emph>. Notably, all meta-analytic results presented in this paper, along with the meta-analytic design parameters in OSM B, are based on using <emph>r</emph> = 0.70 as a reasonable upper bound for the within-sample correlations among design parameters. Thus, the 95% CIs and 95% PIs included in these tables represent conservative estimates of these intervals. Consequently, using the lower bound values for <emph>R</emph><sups>2</sups>s and/or upper bounds values of <emph>ICC</emph>'s from these intervals represents a conservative approach to account for the statistical uncertainty of these design parameters in a priori power analyses. When researchers wish to apply less conservative approaches, they can use the R code available in our OSF to extract meta-analytic design parameters based on smaller values for the within-sample correlation.</p> <p>Finally, we provided numerous meta-analytic summaries of design parameters to support the planning of randomized intervention studies with plausible meta-analytic averages and prediction intervals. However, it was beyond the scope of the present paper to investigate moderator variables that may explain the variability among design parameters. Further, such moderator analyses require a considerably larger number of studies than the three (i.e., BIKS, NUBBEK, and NEPS) we identified with our systematic search for the present paper. Nevertheless, an important next step for future research is to conduct meta-regression analyses to examine how moderator variables at the level of (a) effect sizes (e.g., specific characteristics of the applied measures, time lag between pre- and posttest, reliability of measures) or (b) studies (e.g., country, coverage of the target population, quality of the sampling process, observational or experimental study design, year of data collection) may explain the observed heterogeneity among design parameters.</p> <hd id="AN0182226205-36">Conclusion</hd> <p>This paper offers a unique and comprehensive resource for conducting a priori power analyses to plan sample sizes for both lab-based and field-based randomized experiments in early childhood education research. We hope that the design parameters, along with our recommendations for their application and the illustrative examples, will assist researchers in conducting randomized intervention studies that provide rigorous evidence to support preschool children's cognitive and socio-emotional development.</p> <hd id="AN0182226205-37">Authors' Contribution</hd> <p>All authors contributed to the study conception and design. Data preparation and analyses were performed by Martin Brunner and Sophie Stallasch. The first draft of the manuscript was written by Martin Brunner and all authors commented on subsequent versions of the manuscript. All authors read and approved the final manuscript.</p> <hd id="AN0182226205-38">Funding</hd> <p>Open Access funding enabled and organized by Projekt DEAL. This work was supported by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Grant 392108331.</p> <hd id="AN0182226205-39">Data Availability</hd> <p>This paper uses data from the National Educational Panel Study (NEPS; see Blossfeld &amp; Roßbach, 2019). The NEPS is carried out by the Leibniz Institute for Educational Trajectories (LIfBi, Germany) in cooperation with a nationwide network. The dataset for the study titled "Educational Processes, Competence Development and Selection Decisions in Preschool and School Age (BIKS-3–10, Weinert et al., 2024)" was made available by the Research Data Centre at the Institute for Educational Quality Improvement (FDZ at IQB). Permission from the dataset owners was granted to use these datasets for the research objectives of the present paper. The dataset for the National Survey on Education, Care, and Development in Early Childhood (NUBBEK) was made available by the Leibniz Institute for the Social Sciences in Mannheim (GESIS). The R code for reproducing all results as well as the data with the design parameters used in the present paper can be accessed via the Open Science Framework at https://osf.io/qz7fy.</p> <hd id="AN0182226205-40">Declarations</hd> <p></p> <hd id="AN0182226205-41">Ethics Approval</hd> <p>We used scientific use files of BIKS, NEPS, and NUBBEK. All ethical issues related to these data were handled by the scientific consortia who were responsible for collecting the data.</p> <hd id="AN0182226205-42">Consent</hd> <p>All authors agreed with the content and all authors gave explicit consent to submit.</p> <hd id="AN0182226205-43">Competing Interests</hd> <p>The authors have no relevant financial or non-financial interests to disclose.</p> <hd id="AN0182226205-44">Supplementary Information</hd> <p>Below is the link to the electronic supplementary material.</p> <p>Graph: Supplementary file1 (DOCX 1661 KB)</p> <hd id="AN0182226205-45">Publisher's Note</hd> <p>Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.</p> <ref id="AN0182226205-46"> <title> References </title> <blist> <bibl id="bib1" idref="ref18" type="bt">1</bibl> <bibtext> Artelt C, Sixt M. The National Educational Panel Study (NEPS)—Framework, design, and research potential. Zeitschrift Für Erziehungswissenschaft. 2023; 26; 2: 277-298. 10.1007/s11618-023-01156-w</bibtext> </blist> <blist> <bibl id="bib2" idref="ref19" type="bt">2</bibl> <bibtext> Ataneka, A, Kelcy, B, Dong, N, Bulus, M, &amp; Bai, F. (2023). PowerUp R Shiny App (v. 0.9) Manual. https://<ulink href="http://www.causalevaluation.org/uploads/7/3/3/6/73366257/r%5fshinnyapp%5fmanual%5f0.9.pdf">www.causalevaluation.org/uploads/7/3/3/6/73366257/r%5fshinnyapp%5fmanual%5f0.9.pdf</ulink></bibtext> </blist> <blist> <bibl id="bib3" idref="ref53" type="bt">3</bibl> <bibtext> Autor:innengruppe Bildungsberichterstattung. (2022). Bildung in Deutschland 2022 [Education in Germany 2022]. wbv Media. https://doi.org/10.3278/6001820hw</bibtext> </blist> <blist> <bibl id="bib4" idref="ref1" type="bt">4</bibl> <bibtext> Barnett WS. Effectiveness of early educational intervention. Science. 2011; 333; 6045: 975-978. 10.1126/science.1204534</bibtext> </blist> <blist> <bibl id="bib5" idref="ref152" type="bt">5</bibl> <bibtext> Bates D, Mächler M, Bolker B, Walker S. Fitting linear mixed-effects models using lme4. Journal of Statistical Software. 2015; 67: 1-48. 10.18637/jss.v067.i01gcrnkw</bibtext> </blist> <blist> <bibl id="bib6" idref="ref153" type="bt">6</bibl> <bibtext> Bland JM, Altman DG. Statistics Notes: One and two sided tests of significance. BMJ. 1994; 309; 6949: 248. 10.1136/bmj.309.6949.248</bibtext> </blist> <blist> <bibl id="bib7" idref="ref111" type="bt">7</bibl> <bibtext> Bleidorn W, Schwaba T, Zheng A, Hopwood CJ, Sosa SS, Roberts BW, Briley DA. Personality stability and change: A meta-analysis of longitudinal studies. Psychological Bulletin. 2022; 148; 7–8: 588-619. 10.1037/bul0000365</bibtext> </blist> <blist> <bibl id="bib8" idref="ref43" type="bt">8</bibl> <bibtext> Bloom HS. Minimum detectable effects: A simple way to report the statistical power of experimental designs. Evaluation Review. 1995; 19; 5: 547-556. 10.1177/0193841X9501900504</bibtext> </blist> <blist> <bibl id="bib9" idref="ref21" type="bt">9</bibl> <bibtext> Bloom, H. S. (2006). The core analytics of randomized experiments for social research. MDRC. <ulink href="http://www.mdrc.org/sites/default/files/full%5f533.pdf">http://www.mdrc.org/sites/default/files/full%5f533.pdf</ulink></bibtext> </blist> <blist> <bibtext> Bloom HS, Bos JM, Lee S-W. Using cluster random assignment to measure program impacts. Statistical implications for the evaluation of education programs. Evaluation Review. 1999; 23; 4: 445-469. 10.1177/0193841X9902300405</bibtext> </blist> <blist> <bibtext> Bloom HS, Richburg-Hayes L, Black AR. Using covariates to improve precision for studies that randomize schools to evaluate educational interventions. Educational Evaluation and Policy Analysis. 2007; 29; 1: 30-59. 10.3102/0162373707299550</bibtext> </blist> <blist> <bibtext> Blossfeld, H.-P, &amp; Roßbach, H.-G. (Eds.). (2019). Education as a lifelong process: The German National Educational Panel Study (NEPS) (2nd ed.). VS Verlag für Sozialwissenschaften. https://doi.org/10.1007/978-3-658-23162-0</bibtext> </blist> <blist> <bibtext> Borenstein M, Higgins JPT, Hedges LV, Rothstein HR. Basics of meta-analysis: I2 is not an absolute measure of heterogeneity. Research Synthesis Methods. 2017; 8; 1: 5-18. 10.1002/jrsm.1230</bibtext> </blist> <blist> <bibtext> Boruch R. Better evaluation for evidence-based policy: Place randomized trials in education, criminology, welfare, and health. The ANNALS of the American Academy of Political and Social Science. 2005; 599; 1: 6-18. 10.1177/0002716205275610</bibtext> </blist> <blist> <bibtext> Brunner M, Keller U, Wenger M, Fischbach A, Lüdtke O. Between-school variation in students' achievement, motivation, affect, and learning strategies: Results from 81 countries for planning group-randomized trials in education. Journal of Research on Educational Effectiveness. 2018; 11; 3: 452-478. 10.1080/19345747.2017.1375584gd4q25</bibtext> </blist> <blist> <bibtext> Brunner M, Keller L, Stallasch SE, Kretschmann J, Hasl A, Preckel F, Lüdtke O, Hedges LV. Meta-analyzing individual participant data from studies with complex survey designs: A tutorial on using the two-stage approach for data from educational large-scale assessments. Research Synthesis Methods. 2023; 14; 1: 5-35. 10.1002/jrsm.1584</bibtext> </blist> <blist> <bibtext> Brunner, M, Stallasch, S. E, &amp; Lüdtke, O. (2023b). Empirical benchmarks to interpret intervention effects on student achievement in elementary and secondary school: Meta-analytic results from Germany. Journal of Research on Educational Effectiveness, 17(1), 119–157. https://doi.org/10.1080/19345747.2023.2175753</bibtext> </blist> <blist> <bibtext> Bulus M. Minimum detectable effect size computations for cluster-level regression discontinuity studies: Specifications beyond the linear functional form. Journal of Research on Educational Effectiveness. 2022; 15; 1: 151-177. 10.1080/19345747.2021.1947425</bibtext> </blist> <blist> <bibtext> Bulus M, Sahin SG. Estimation and standardization of variance parameters for planning cluster-randomized trials: A short guide for researchers. Journal of Measurement and Evaluation in Education and Psychology. 2019; 10; 2: 2. 10.21031/epod.530642</bibtext> </blist> <blist> <bibtext> Bulus, M, Dong, N, Kelcey, B, &amp; Spybrook, J. (2021). PowerUpR: Power analysis tools for multilevel randomized experiments (Version 1.1.0) [Computer software]. https://cran.r-project.org/web/packages/PowerUpR/index.html</bibtext> </blist> <blist> <bibtext> Caspi A, Roberts BW, Shiner RL. Personality development: Stability and change. Annual Review of Psychology. 2005; 56: 453-484. 10.1146/annurev.psych.55.090902.141913cx7kjq</bibtext> </blist> <blist> <bibtext> Cinelli, C, Forney, A, &amp; Pearl, J. (2022). A crash course in good and bad controls. Sociological Methods &amp; Research, 53(3), 1071–1104. https://doi.org/10.1177/00491241221099552</bibtext> </blist> <blist> <bibtext> Cochran WG. Some effects of errors of measurement on multiple correlation. Journal of the American Statistical Association. 1970; 65; 329: 22-34. 10.1080/01621459.1970.10481059mmb8</bibtext> </blist> <blist> <bibtext> Cohen J. Statistical power analysis for the behavioral sciences. 1988; Lawrence Erlbaum</bibtext> </blist> <blist> <bibtext> Connolly P, Keenan C, Urbanska K. The trials of evidence-based practice in education: A systematic review of randomised controlled trials in education research 1980–2016. Educational Research. 2018; 60; 3: 276-291. 10.1080/00131881.2018.1493353gjrpjc</bibtext> </blist> <blist> <bibtext> Cook TD. Randomized experiments in educational policy research: A critical examination of the reasons the educational evaluation community has offered for not doing them. Educational Evaluation and Policy Analysis. 2002; 24; 3: 175-199. 10.3102/01623737024003175</bibtext> </blist> <blist> <bibtext> Cook TD. Emergent principles for the design, implementation, and analysis of cluster-based experiments in social science. The ANNALS of the American Academy of Political and Social Science. 2005. 10.1177/0002716205275738</bibtext> </blist> <blist> <bibtext> Dawson A, Yeomans E, Brown ER. Methodological challenges in education RCTs: Reflections from England's Education Endowment Foundation. Educational Research. 2018; 60; 3: 292-310. 10.1080/00131881.2018.1500079</bibtext> </blist> <blist> <bibtext> De Pauw SSW, Mervielde I. Temperament, personality and developmental psychopathology: A review based on the conceptual dimensions underlying childhood traits. Child Psychiatry &amp; Human Development. 2010; 41; 3: 313-329. 10.1007/s10578-009-0171-8fthv76</bibtext> </blist> <blist> <bibtext> Ding, P. (2023). A first course in causal inference (No. arXiv:2305.18793). arXiv. https://doi.org/10.48550/arXiv.2305.18793</bibtext> </blist> <blist> <bibtext> Dong N, Maynard R. PowerUp!: A tool for calculating minimum detectable effect sizes and minimum required sample sizes for experimental and quasi-experimental design studies. Journal of Research on Educational Effectiveness. 2013; 6; 1: 24-67. 10.1080/19345747.2012.673143gd4q27</bibtext> </blist> <blist> <bibtext> Dong N, Reinke WM, Herman KC, Bradshaw CP, Murray DW. Meaningful effect sizes, intraclass correlations, and proportions of variance explained by covariates for planning two- and three-level cluster randomized trials of social and behavioral outcomes. Evaluation Review. 2016; 40; 4: 334-377. 10.1177/0193841X16671283</bibtext> </blist> <blist> <bibtext> Durlak JA, Mahoney JL, Boyle AE. What we know, and what we need to find out about universal, school-based social and emotional learning programs for children and adolescents: A review of meta-analyses and directions for future research. Psychological Bulletin. 2022; 148: 765-782. 10.1037/bul0000383kx88</bibtext> </blist> <blist> <bibtext> Flanagan DP, Dixon SG. The Cattell-Horn-Carroll theory of cognitive abilities. 2014; John Wiley &amp; Sons, Ltd. mmcb. 10.1002/9781118660584.ese0431</bibtext> </blist> <blist> <bibtext> Ganzeboom HBG, Treiman DJ. Internationally comparable measures of occupational status for the 1988 International Standards Classification of Occupations. Social Science Research. 1996; 25: 201-239. 10.1006/ssre.1996.0010</bibtext> </blist> <blist> <bibtext> García JL, Heckman JJ, Ronda V. The lasting effects of early-childhood education on promoting the skills and social mobility of disadvantaged African Americans and their children. Journal of Political Economy. 2023; 131; 6: 1477-1506. 10.1086/722936</bibtext> </blist> <blist> <bibtext> Gaspard H, Dicke A-L, Flunger B, Brisson BM, Häfner I, Nagengast B, Trautwein U. Fostering adolescents' value beliefs for mathematics with a relevance intervention in the classroom. Developmental Psychology. 2015; 51; 9: 1226-1240. 10.1037/dev0000028</bibtext> </blist> <blist> <bibtext> Grund, S, Robitzsch, A, &amp; Lüdtke, O. (2021). Mitml: Tools for multiple imputation in multilevel modeling. R package version 0.4–3 [Computer software]. https://CRAN.R-project.org/package=lmeresampler</bibtext> </blist> <blist> <bibtext> Grund S, Lüdtke O, Robitzsch A. Handling missing data in cross-classified multilevel analyses: An evaluation of different multiple imputation approaches. Journal of Educational and Behavioral Statistics. 2023; 48: 454-489. 10.3102/10769986231151224kx9d</bibtext> </blist> <blist> <bibtext> Hedberg EC. Academic and behavioral design parameters for cluster randomized trials in kindergarten: An analysis of the Early Childhood Longitudinal Study 2011 kindergarten cohort (ECLS-K 2011). Evaluation Review. 2016; 40; 4: 279-313. 10.1177/0193841X16655657</bibtext> </blist> <blist> <bibtext> Hedges, L. V. (2019). Stochastically dependent effect sizes. In H. M. Cooper, L. V. Hedges, &amp; J. C. Valentine (Eds.), Handbook of research synthesis and meta-analysis (3rd Edition, pp. 245–280). Russell Sage Foundation. https://doi.org/10.7758/9781610448864.16</bibtext> </blist> <blist> <bibtext> Hedges LV, Hedberg EC. Intraclass correlation values for planning group-randomized trials in education. Educational Evaluation and Policy Analysis. 2007; 29; 1: 60-87. 10.3102/0162373707299706</bibtext> </blist> <blist> <bibtext> Hedges, L. V, &amp; Rhoads, C. (2010). Statistical power analysis in education research. NCSER 2010–3006. In National Center for Special Education Research. National Center for Special Education Research. <ulink href="http://files.eric.ed.gov/fulltext/ED509387.pdf">http://files.eric.ed.gov/fulltext/ED509387.pdf</ulink></bibtext> </blist> <blist> <bibtext> Hedges LV, Hedberg EC. Intraclass correlations and covariate outcome correlations for planning two- and three-level cluster-randomized experiments in education. Evaluation Review. 2013; 37; 6: 445-489. 10.1177/0193841X14529126</bibtext> </blist> <blist> <bibtext> Hedges LV, Schauer J. Randomised trials in education in the USA. Educational Research. 2018; 60; 3: 265-275. 10.1080/00131881.2018.1493350</bibtext> </blist> <blist> <bibtext> Homuth, C, Lehrl, S, Volodina, A, Weinert, S, &amp; Rossbach, H.-G. (2024). From Preschool to Vocational Training and Tertiary Education—Study Design of the BiKS-3–18 Study. In S. Weinert, H.-G. Rossbach, J. Von Maurice, H.-P. Blossfeld, &amp; C. Artelt (Eds.), Educational Processes, Decisions, and the Development of Competencies from Early Preschool Age to Adolescence (Vol. 16, pp. 21–53). Springer Fachmedien Wiesbaden. https://doi.org/10.1007/978-3-658-43414-4_2</bibtext> </blist> <blist> <bibtext> Jacob RT, Zhu P, Bloom HS. New empirical evidence for the design of group randomized trials in education. Journal of Research on Educational Effectiveness. 2010; 3; 2: 157-198. 10.1080/19345741003592428</bibtext> </blist> <blist> <bibtext> Kelcey B, Shen Z, Spybrook J. Intraclass correlation coefficients for designing cluster-randomized trials in Sub-Saharan Africa education. Evaluation Review. 2016; 40; 6: 500-525. 10.1177/0193841X16660246</bibtext> </blist> <blist> <bibtext> Konstantopoulos S. The power of the test for treatment effects in three-level block randomized designs. Journal of Research on Educational Effectiveness. 2008; 1; 4: 265-288. 10.1080/19345740802328216</bibtext> </blist> <blist> <bibtext> Konstantopoulos S. The impact of covariates on statistical power in cluster randomized designs: Which level matters more?. Multivariate Behavioral Research. 2012; 47; 3: 392-420. 10.1080/00273171.2012.673898</bibtext> </blist> <blist> <bibtext> Kraft MA. Interpreting effect sizes of education interventions. Educational Researcher. 2020; 49; 4: 241-253. 10.3102/0013189X20912798</bibtext> </blist> <blist> <bibtext> Langan D, Higgins JPT, Jackson D, Bowden J, Veroniki AA, Kontopantelis E, Viechtbauer W, Simmonds M. A comparison of heterogeneity variance estimators in simulated random-effects meta-analyses. Research Synthesis Methods. 2019; 10; 1: 83-98. 10.1002/jrsm.1316</bibtext> </blist> <blist> <bibtext> Letourneau NL, Duffett-Leger L, Levac L, Watson B, Young-Morris C. Socioeconomic status and child development: A meta-analysis. Journal of Emotional and Behavioral Disorders. 2013; 21; 3: 211-224. 10.1177/1063426611421007c9sdc3</bibtext> </blist> <blist> <bibtext> Leyendecker, B, Agache, A, &amp; Madsen, S. (2014). Nationale Untersuchung zur Bildung, Betreuung und Erziehung in der frühen Kindheit (NUBBEK) – Design, Methodenüberblick, Datenzugang und das Potenzial zu Mehrebenenanalysen [NUBBEK – a national German study on early childhood education and care: Design, methods overview, data access, and the potential for multilevel analyses]. ZfF – Zeitschrift für Familienforschung / Journal of Family Research, 26(2), 2. https://<ulink href="http://www.budrich-journals.de/index.php/zff/article/view/16528">www.budrich-journals.de/index.php/zff/article/view/16528</ulink></bibtext> </blist> <blist> <bibtext> Lin W. Agnostic notes on regression adjustments to experimental data: Reexamining Freedman's critique. The Annals of Applied Statistics. 2013; 7; 1: 295-318. 10.1214/12-AOAS583</bibtext> </blist> <blist> <bibtext> Lipsey, M. W, Puzio, K, Yun, C, Hebert, M. A, Steinka-Fry, K, Cole, M. W, Roberts, M, Anthony, K. S, &amp; Busick, M. D. (2012). Translating the statistical representation of the effects of education interventions into more readily interpretable forms. National Center for Special Education Research. <ulink href="http://eric.ed.gov/?id=ED537446">http://eric.ed.gov/?id=ED537446</ulink></bibtext> </blist> <blist> <bibtext> Lortie-Forgues H, Inglis M. Rigorous large-scale educational RCTs are often uninformative: Should we be concerned?. Educational Researcher. 2019; 48; 3: 158-166. 10.3102/0013189X19832850ggm84b</bibtext> </blist> <blist> <bibtext> Maxwell, S. E, Delaney, H. D, &amp; Kelley, K. (2018). Designing experiments and analyzing data: A model comparison perspective (Third edition). Routledge. https://doi.org/10.4324/9781315642956</bibtext> </blist> <blist> <bibtext> McNeish DM. Modeling sparsely clustered data: Design-based, model-based, and single-level methods. Psychological Methods. 2014; 19; 4: 552-563. 10.1037/met0000024</bibtext> </blist> <blist> <bibtext> McNeish DM, Stapleton LM. The effect of small sample size on two-level model estimates: A review and illustration. Educational Psychology Review. 2016; 28; 2: 295-314. 10.1007/s10648-014-9287-x</bibtext> </blist> <blist> <bibtext> Miciak J, Taylor WP, Stuebing KK, Fletcher JM, Vaughn S. Designing intervention studies: Selected populations, range restrictions, and statistical power. Journal of Research on Educational Effectiveness. 2016; 9; 4: 556-569. 10.1080/19345747.2015.1086916mmcc</bibtext> </blist> <blist> <bibtext> Mischel W, Shoda Y. A cognitive-affective system theory of personality: Reconceptualizing situations, dispositions, dynamics, and invariance in personality structure. Psychological Review. 1995; 102; 2: 246-268. 10.1037/0033-295X.102.2.246</bibtext> </blist> <blist> <bibtext> Moerbeek M, Teerenstra S. Power analysis of trials with multilevel data. 2016; CRC Press</bibtext> </blist> <blist> <bibtext> Mosteller, F, &amp; Boruch, R. (2002). Evidence matters: Randomized trials in education research. Brookings Institution Press. https://<ulink href="http://www.jstor.org/stable/10.7864/j.ctvc16n69">www.jstor.org/stable/10.7864/j.ctvc16n69</ulink></bibtext> </blist> <blist> <bibtext> Myers, D, &amp; Schirm, A. (1999). The Impacts of Upward Bound: Final Report for Phase I of the National Evaluation. https://eric.ed.gov/?id=ED432621</bibtext> </blist> <blist> <bibtext> NEPS Network. (2022). National Educational Panel Study, Scientific Use File of Starting Cohort Kindergarten (Version 10.0.0). LIfBi Leibniz Institute for Educational Trajectories. https://doi.org/10.5157/NEPS:SC2:10.0.0</bibtext> </blist> <blist> <bibtext> Nye B, Hedges LV, Konstantopoulos S. The long-term effects of small classes: A five-year follow-up of the Tennessee class size experiment. Educational Evaluation and Policy Analysis. 1999; 21; 2: 127-142. 10.3102/01623737021002127</bibtext> </blist> <blist> <bibtext> OECD. (2020). Early learning and child well-being: A study of five-year olds in England, Estonia, and the United States. OECD. https://doi.org/10.1787/3990407f-en</bibtext> </blist> <blist> <bibtext> Page, M. J, McKenzie, J. E, Bossuyt, P. M, Boutron, I, Hoffmann, T. C, Mulrow, C. D, Shamseer, L, Tetzlaff, J. M, Akl, E. A, Brennan, S. E, Chou, R, Glanville, J, Grimshaw, J. M, Hróbjartsson, A, Lalu, M. M, Li, T, Loder, E. W, Mayo-Wilson, E, McDonald, S, ..., Moher, D. (2021). The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ, 372, n71. gjkq9b.</bibtext> </blist> <blist> <bibtext> Peng P, Kievit RA. The development of academic achievement and cognitive abilities: A bidirectional perspective. Child Development Perspectives. 2020; 14; 1: 15-20. 10.1111/cdep.12352ggxvw3</bibtext> </blist> <blist> <bibtext> Pontoppidan M, Keilow M, Dietrichson J, Solheim OJ, Opheim V, Gustafson S, Andersen SC. Randomised controlled trials in Scandinavian educational research. Educational Research. 2018; 60; 3: 311-335. 10.1080/00131881.2018.1493351</bibtext> </blist> <blist> <bibtext> Porter AC, Raudenbush SW. Analysis of covariance: Its model and use in psychological research. Journal of Counseling Psychology. 1987; 34; 4: 383-392. 10.1037/0022-0167.34.4.383fn8rhp</bibtext> </blist> <blist> <bibtext> Pustejovsky, J. E. (2021). ClubSandwich: Cluster-robust (sandwich) variance estimators with small-sample corrections. R package version 0.5.3. [Computer software]. https://CRAN.R-project.org/package=clubSandwich</bibtext> </blist> <blist> <bibtext> R Core Team. (2024). R: A language and environment for statistical computing [Computer software]. R Foundation for Statistical Computing. https://<ulink href="http://www.R-project.org/">www.R-project.org/</ulink></bibtext> </blist> <blist> <bibtext> Raudenbush SW, Martinez A, Spybrook J. Strategies for improving precision in group-randomized experiments. Educational Evaluation and Policy Analysis. 2007; 29; 1: 5-29. 10.3102/0162373707299460</bibtext> </blist> <blist> <bibtext> Raudenbush SW, Bryk AS. Hierarchical linear models. 20022; Sage</bibtext> </blist> <blist> <bibtext> Rice K, Higgins JPT, Lumley T. A re-evaluation of fixed effect(s) meta-analysis. Journal of the Royal Statistical Society: Series A (Statistics in Society). 2018; 181; 1: 205-227. 10.1111/rssa.12275</bibtext> </blist> <blist> <bibtext> Robitzsch, A, &amp; Grund, S. (2023). miceadds: Some additional multiple imputation functions, especially for "mice" (Version 3.16-18) [Computer software].</bibtext> </blist> <blist> <bibtext> Rubin, D. B. (1987). Multiple imputation for nonresponse in surveys. J. Wiley &amp; Sons. https://doi.org/10.1002/9780470316696</bibtext> </blist> <blist> <bibtext> Sabol TJ, McCoy D, Gonzalez K, Miratrix L, Hedges L, Spybrook JK, Weiland C. Exploring treatment impact heterogeneity across sites: Challenges and opportunities for early childhood researchers. Early Childhood Research Quarterly. 2022; 58: 14-26. 10.1016/j.ecresq.2021.07.005mmcd</bibtext> </blist> <blist> <bibtext> Sammons, P, Sylva, K, Melhuish, E, Siraj-Blatchford, I, Taggart, B, &amp; Elliot, K. (2002). Measuring the impact of pre-school on children's cognitive progress over the preschool period: Technical Paper 8a. https://discovery.ucl.ac.uk/id/eprint/10005295/1/Sammons2003Effective(Tech.Paper8A).pdf</bibtext> </blist> <blist> <bibtext> Sammons, P, Sylva, K, Melhuish, E, Siraj-Blatchford, I, Taggart, B, &amp; Elliot, K. (2003). The Effective Provision of Pre-School Education (EPPE) Project: Measuring the Impact of Pre-School on Children's Social/Behavioural Development over the Pre-School Period. In Institute of Education, University of London/ Department for Education and Skills: London. [Report]. Institute of Education, University of London/ Department for Education and Skills. https://discovery.ucl.ac.uk/id/eprint/10005288/</bibtext> </blist> <blist> <bibtext> Schmitt N. Uses and abuses of coefficient alpha. Psychological Assessment. 1996; 8: 350-353. 10.1037/1040-3590.8.4.350cjqw3x</bibtext> </blist> <blist> <bibtext> Schochet PZ. Statistical power for regression discontinuity designs in education evaluations. Journal of Educational and Behavioral Statistics. 2009; 34; 2: 238-266. 10.3102/1076998609332748</bibtext> </blist> <blist> <bibtext> Schoon, I. (2021). Towards an integrative taxonomy of social-emotional competences. Frontiers in Psychology, 12, 515313. https://doi.org/10.3389/fpsyg.2021.515313</bibtext> </blist> <blist> <bibtext> Schweinhart, L. J. (Ed.). (2005). Lifetime effects: The High/Scope Perry preschool study through age 40. High/Scope Press.</bibtext> </blist> <blist> <bibtext> Simonsmeier BA, Flaig M, Deiglmayr A, Schalk L, Schneider M. Domain-specific prior knowledge and learning: A meta-analysis. Educational Psychologist. 2022; 57; 1: 31-54. 10.1080/00461520.2021.1939700jhm4</bibtext> </blist> <blist> <bibtext> Slavin RE. How evidence-based reform will transform research and practice in education. Educational Psychologist. 2020; 55; 1: 21-31. 10.1080/00461520.2019.1611432gf4hzh</bibtext> </blist> <blist> <bibtext> Somers, M.-A, Weiss, M. J, &amp; Hill, C. (2022). Design parameters for planning the sample size of individual-level randomized controlled trials in community colleges. Evaluation Review, 0193841X221121236. https://doi.org/10.1177/0193841X221121236</bibtext> </blist> <blist> <bibtext> Sparrow SS, Cicchetti DV, Balla DA. Vineland Adaptive Behavior Scales, Second Edition (Vineland-II): Survey forms manual. 2005; Pearson Assessments</bibtext> </blist> <blist> <bibtext> Spybrook J, Raudenbush SW. An examination of the precision and technical accuracy of the first wave of group-randomized trials funded by the Institute of Education Sciences. Educational Evaluation and Policy Analysis. 2009; 31; 3: 298-318. 10.3102/0162373709339524</bibtext> </blist> <blist> <bibtext> Spybrook J, Shi R, Kelcey B. Progress in the past decade: An examination of the precision of cluster randomized trials funded by the U.S. Institute of Education Sciences. International Journal of Research &amp; Method in Education. 2016; 39; 3: 255-267. 10.1080/1743727X.2016.1150454</bibtext> </blist> <blist> <bibtext> Spybrook, J, Bloom, H, Congdon, R, Hill, C, Martinez, A, &amp; Raudenbush, S. (2011). Optimal Design plus empirical evidence: Documentation for the "Optimal Design" software. <ulink href="http://hlmsoft.net/od/od301.zip">http://hlmsoft.net/od/od301.zip</ulink>. <ulink href="http://hlmsoft.net/od/od-manual-20111016-v300.pdf">http://hlmsoft.net/od/od-manual-20111016-v300.pdf</ulink></bibtext> </blist> <blist> <bibtext> Stallasch SE, Lüdtke O, Artelt C, Brunner M. Multilevel design parameters to plan cluster-randomized intervention studies on student achievement in elementary and secondary school. Journal of Research on Educational Effectiveness. 2021; 14; 1: 172-206. 10.1080/19345747.2020.1823539mmcf</bibtext> </blist> <blist> <bibtext> Stallasch SE, Lüdtke O, Artelt C, Hedges LV, Brunner M. Single- and multilevel perspectives on covariate selection in randomized intervention studies on student achievement. Educational Psychology Review. 2024; 36; 4: 112. 10.1007/s10648-024-09898-7</bibtext> </blist> <blist> <bibtext> Standing Scientific Commission on Education Policy. (2022). Impulspapier: Entwicklung von Leitlinien für das Monitoring und die Evaluation von Förderprogrammen im Bildungsbereich [Position paper on the development of guidelines for the monitoring and evaluation of support programs in the education sector]. https://<ulink href="http://www.swk-bildung.org/content/uploads/2024/02/SWK-2022-Impulspapier%5fMonitoring.pdf">www.swk-bildung.org/content/uploads/2024/02/SWK-2022-Impulspapier%5fMonitoring.pdf</ulink></bibtext> </blist> <blist> <bibtext> Stewart LA, Clarke M, Rovers M, Riley RD, Simmonds M, Stewart G, Tierney JF. Preferred reporting items for a systematic review and meta-analysis of individual participant data: The PRISMA-IPD statement. JAMA. 2015; 313; 16: 1657-1665. 10.1001/jama.2015.3656f7bhcv</bibtext> </blist> <blist> <bibtext> Tackett JL, Kushner SC, De Fruyt F, Mervielde I. Delineating personality traits in childhood and adolescence: Associations across measures, temperament, and behavioral problems. Assessment. 2013; 20; 6: 738-751. 10.1177/1073191113509686</bibtext> </blist> <blist> <bibtext> Tietze, W, Becker-Stoll, F, Bensel, J, Haug-Schnabel, G, Kalicki, B, Keller, H, &amp; Leyendecker, B. (2015). NUBBEK - Nationale Untersuchung zur Bildung, Betreuung und Erziehung in der frühen Kindheit [NUBBEK - National survey on education, care, and development in early childhood] (Version 3.0.0). GESIS Data Archive. https://doi.org/10.4232/1.12297</bibtext> </blist> <blist> <bibtext> Tucker-Drob EM. Cognitive aging and dementia: A life-span perspective. Annual Review of Developmental Psychology. 2019; 1; 1: 177-196. 10.1146/annurev-devpsych-121318-085204gh3gxg</bibtext> </blist> <blist> <bibtext> U.S. Department of Education, National Center for Education Statistics. (2021). Early childhood program participation: 2019 (NCES 2020–075REV), Table 1. National Center for Education Statistics. https://nces.ed.gov/fastfacts/display.asp?id=4</bibtext> </blist> <blist> <bibtext> Ulferts, H. (2017). Komponenten und Auswirkungen der Qualität mathematischer Bildung in frühkindlichen Bildungs- und Betreuungseinrichtungen [Components and impact of the quality of math education in early childhood education and care centers] [Freie Universität Berlin]. https://refubium.fu-berlin.de/handle/fub188/5432</bibtext> </blist> <blist> <bibtext> van Buuren, S, &amp; Groothuis-Oudshoorn, K. (2011). mice: Multivariate imputation by chained equations in R. Journal of Statistical Software, 45(3), 1–67. https://doi.org/10.18637/jss.v045.i03</bibtext> </blist> <blist> <bibtext> Viechtbauer, W. (2010). Conducting meta-analyses in R with the metafor package. Journal of Statistical Software, 36(3), 1–48. https://doi.org/10.18637/jss.v036.i03</bibtext> </blist> <blist> <bibtext> Weinert, S, Roßbach, H.-G, Faust, G, Blossfeld, H.-P, Artelt, C, &amp; Otto-Friedrich-Universität Bamberg. (2019). Educational Processes, Competence Development and Selection Decisions in Preschool and School Age (BiKS-3-10) (Version 6) [Dataset]. IQB - Institute for Educational Quality Improvement. https://doi.org/10.5159/IQB_BIKS_3_10_V6</bibtext> </blist> <blist> <bibtext> Weiss MJ, Bloom HS, Verbitsky-Savitz N, Gupta H, Vigil AE, Cullinan DN. How much do the effects of education and training programs vary across sites? Evidence from past multisite randomized trials. Journal of Research on Educational Effectiveness. 2017; 10; 4: 843-876. 10.1080/19345747.2017.1300719mmcg</bibtext> </blist> <blist> <bibtext> Wenger, M, Lüdtke, O, &amp; Brunner, M. (2018). Übereinstimmung, Variabilität und Reliabilität von Schülerurteilen zur Unterrichtsqualität auf Schulebene [Interrater agreement, variability, and reliability of student ratings of instructional quality at the school-level]. Zeitschrift für Erziehungswissenschaft, 21, 929–950. https://doi.org/10.1007/s11618-018-0813-3</bibtext> </blist> <blist> <bibtext> Westine CD, Spybrook J, Taylor JA. An empirical investigation of variance design parameters for planning cluster-randomized trials of science achievement. Evaluation Review. 2013; 37; 6: 490-519. 10.1177/0193841X14531584</bibtext> </blist> <blist> <bibtext> Westine CD, Unlu F, Taylor J, Spybrook J, Zhang Q, Anderson B. Design parameter values for impact evaluations of science and mathematics interventions involving teacher outcomes. Journal of Research on Educational Effectiveness. 2020; 13; 4: 816-839. 10.1080/19345747.2020.1821849</bibtext> </blist> <blist> <bibtext> Zhang H, Shen Z, Leite WL. The impacts of small teacher-level sample sizes in cluster-randomized trials. The Journal of Experimental Education. 2024; 0; 0: 1-20. 10.1080/00220973.2024.2376627</bibtext> </blist> <blist> <bibtext> Zhu P, Jacob R, Bloom H, Xu Z. Designing and analyzing studies that randomize schools to estimate intervention effects on student academic outcomes without classroom-level information. Educational Evaluation and Policy Analysis. 2012; 34; 1: 45-68. 10.3102/0162373711423786</bibtext> </blist> <blist> <bibtext> Zopluoglu C. Across-national comparison of intra-class correlation coefficient in educational achievement outcomes. Journal of Measurement and Evaluation in Education and Psychology. 2012; 3; 1: 242-278</bibtext> </blist> </ref> <ref id="AN0182226205-47"> <title> Footnotes </title> <blist> <bibtext> Notably, the present design parameters are also useful for a priori power analyses of multilevel quasi-experimental designs (Bulus, [18]; Dong &amp; Maynard, [31]; Schochet, [84]).</bibtext> </blist> <blist> <bibtext> https://osf.io/qz7fy. Tables and Figures presented in OSM A and B are indicated by corresponding letters (e.g., Table A1 in OSM A or Table B.CT.M in OSM B).</bibtext> </blist> <blist> <bibtext> In addition to the design parameters (i.e., <emph>R</emph><sups>2</sups>s and ICCs) that we present in this paper, power analyses of MSRTs require information on the expected heterogeneity of the treatment effect across daycare centers and the extent to which covariates may explain this heterogeneity (Dong &amp; Maynard, [31]; Hedges &amp; Rhoads, [43]). We further elaborate on this type of design parameters in the Discussion section.</bibtext> </blist> <blist> <bibtext> These studies drew on a subset of the data that we also used for estimating design parameters in the present paper (i.e., data from the BIKS and NUBBEK study). Of note, these previous studies presented only results for ρ<subs>Center</subs> for a limited set of outcome variables.</bibtext> </blist> <blist> <bibtext> https://<ulink href="http://www.forschungsdaten-bildung.de/en/studies/search">www.forschungsdaten-bildung.de/en/studies/search</ulink></bibtext> </blist> <blist> <bibtext> https://<ulink href="http://www.iqb.hu-berlin.de/fdz/studies/">www.iqb.hu-berlin.de/fdz/studies/</ulink></bibtext> </blist> <blist> <bibtext> Notably, we did not apply sampling weights when estimating design parameters because (a) information on weights was only available for NEPS and (b) the application of sampling weights is not possible with the lme4 package (Bates et al., [5]) that we used for the multilevel analyses. Therefore, our design parameters based on NEPS are representative only for the population of preschool children included in the present analyses.</bibtext> </blist> <blist> <bibtext> Of note, the variation in between-center differences between two- and three-level designs can be largely attributed to using only the subset of cognitive or SEL outcomes obtained for NEPS for estimating the three-level parameters.</bibtext> </blist> <blist> <bibtext> Bulus and Sahin ([19]) offer analytic solutions for evaluating the relative effectiveness of covariates at the child, group, or daycare center levels in enhancing the design sensitivity of two-level and three-level CRTs. For instance, their work identifies the conditions under which daycare center-level covariates are more effective in improving the design sensitivity of a two-level CRT compared to child-level covariates (and vice versa). Notably, Bulus ([18]) extends this analysis by providing analytic solutions to determine these conditions for quasi-experimental regression discontinuity designs.</bibtext> </blist> <blist> <bibtext> Nevertheless, some standard errors were relatively large. For example, most standard errors obtained for the point estimates and meta-analytic averages of <emph>R</emph><sups>2</sups><subs><emph>Group</emph></subs> and <emph>R</emph><sups>2</sups><subs><emph>Center</emph></subs> for SEL outcomes were quite large, particularly when using parent report (see Figures A2 and 1). Large standard errors of <emph>R</emph><sups>2</sups><subs><emph>Group</emph></subs><emph>,</emph> and <emph>R</emph><sups>2</sups><subs><emph>Center</emph></subs> (but not <emph>R</emph><sups>2</sups><subs><emph>Child</emph></subs>) were often observed for very small values of <subs><emph>Center</emph></subs> and <subs><emph>Group</emph></subs> (see Figure A3). Hence, there was not much variance for covariates to explain in the outcomes at the group or daycare center levels. In these cases, variance estimates at the group and daycare center levels became unstable (likely due to chance differences), resulting in large standard errors of the point estimates of <emph>R</emph><sups>2</sups><subs>Group</subs> and <emph>R</emph><sups>2</sups><subs>Center</subs> and their meta-analytic averages.</bibtext> </blist> <blist> <bibtext> It may be the case that no suitable alternative point estimates or meta-analytic averages are available for <emph>R</emph><sups>2</sups><subs><emph>Group</emph></subs> and <emph>R</emph><sups>2</sups><subs><emph>Center</emph></subs> that could be estimated with greater statistical precision. In this case, and when corresponding values for both ρ<subs><emph>Group</emph></subs> and ρ<subs><emph>Center</emph></subs> were negligible or at most very small (e.g., ρ<subs><emph>Group</emph></subs> ≤ .02 and ρ<subs><emph>Center</emph></subs> ≤ .02) we recommend drawing on the lower and upper bound estimates for <emph>R</emph><sups>2</sups><subs><emph>Group</emph></subs> and <emph>R</emph><sups>2</sups><subs><emph>Center</emph></subs> because doing so still leads to a plausible range of <emph>MDES</emph> values although these design parameters have large standard errors (see OSM A6 for a discussion).</bibtext> </blist> <blist> <bibtext> The Excel worksheets can be downloaded here: https://<ulink href="http://www.causalevaluation.org/power-analysis.html">www.causalevaluation.org/power-analysis.html</ulink>. The PowerUpR shiny app can be accessed here: https://powerupr.shinyapps.io/index. The Optimal Design software can be downloaded here: https://wtgrantfoundation.org/optimal-design-with-empirical-information-od.</bibtext> </blist> <blist> <bibtext> The standard errors for the meta-analytic average of <emph>R</emph><sups>2</sups><subs>Center</subs> obtained for Model Sets 1-SD, 2-IB, and 3-SD + IB were relatively large with <emph>SE</emph>s &gt; 0.05. However, the team did not find alternative meta-analytic summaries with SEs &lt; .05 for Model Sets 2-IB and 3-SD + IB and therefore considered the present values as the best guess.</bibtext> </blist> </ref> <aug> <p>Reported by Author; Author; Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib68" firstref="ref2"></nolink> <nolink nlid="nl2" bibid="bib80" firstref="ref4"></nolink> <nolink nlid="nl3" bibid="bib88" firstref="ref5"></nolink> <nolink nlid="nl4" bibid="bib25" firstref="ref6"></nolink> <nolink nlid="nl5" bibid="bib45" firstref="ref7"></nolink> <nolink nlid="nl6" bibid="bib57" firstref="ref9"></nolink> <nolink nlid="nl7" bibid="bib11" firstref="ref10"></nolink> <nolink nlid="nl8" bibid="bib42" firstref="ref11"></nolink> <nolink nlid="nl9" bibid="bib43" firstref="ref12"></nolink> <nolink nlid="nl10" bibid="bib75" firstref="ref13"></nolink> <nolink nlid="nl11" bibid="bib47" firstref="ref16"></nolink> <nolink nlid="nl12" bibid="bib93" firstref="ref17"></nolink> <nolink nlid="nl13" bibid="bib74" firstref="ref20"></nolink> <nolink nlid="nl14" bibid="bib10" firstref="ref22"></nolink> <nolink nlid="nl15" bibid="bib14" firstref="ref23"></nolink> <nolink nlid="nl16" bibid="bib26" firstref="ref24"></nolink> <nolink nlid="nl17" bibid="bib27" firstref="ref25"></nolink> <nolink nlid="nl18" bibid="bib64" firstref="ref27"></nolink> <nolink nlid="nl19" bibid="bib86" firstref="ref29"></nolink> <nolink nlid="nl20" bibid="bib36" firstref="ref30"></nolink> <nolink nlid="nl21" bibid="bib67" firstref="ref32"></nolink> <nolink nlid="nl22" bibid="bib65" firstref="ref33"></nolink> <nolink nlid="nl23" bibid="bib91" firstref="ref35"></nolink> <nolink nlid="nl24" bibid="bib92" firstref="ref36"></nolink> <nolink nlid="nl25" bibid="bib28" firstref="ref37"></nolink> <nolink nlid="nl26" bibid="bib71" firstref="ref38"></nolink> <nolink nlid="nl27" bibid="bib37" firstref="ref39"></nolink> <nolink nlid="nl28" bibid="bib96" firstref="ref40"></nolink> <nolink nlid="nl29" bibid="bib44" firstref="ref42"></nolink> <nolink nlid="nl30" bibid="bib51" firstref="ref45"></nolink> <nolink nlid="nl31" bibid="bib56" firstref="ref47"></nolink> <nolink nlid="nl32" bibid="bib17" firstref="ref48"></nolink> <nolink nlid="nl33" bibid="bib32" firstref="ref49"></nolink> <nolink nlid="nl34" bibid="bib31" firstref="ref51"></nolink> <nolink nlid="nl35" bibid="bib63" firstref="ref55"></nolink> <nolink nlid="nl36" bibid="bib101" firstref="ref59"></nolink> <nolink nlid="nl37" bibid="bib72" firstref="ref60"></nolink> <nolink nlid="nl38" bibid="bib15" firstref="ref73"></nolink> <nolink nlid="nl39" bibid="bib94" firstref="ref74"></nolink> <nolink nlid="nl40" bibid="bib95" firstref="ref75"></nolink> <nolink nlid="nl41" bibid="bib34" firstref="ref78"></nolink> <nolink nlid="nl42" bibid="bib90" firstref="ref79"></nolink> <nolink nlid="nl43" bibid="bib33" firstref="ref80"></nolink> <nolink nlid="nl44" bibid="bib85" firstref="ref81"></nolink> <nolink nlid="nl45" bibid="bib40" firstref="ref83"></nolink> <nolink nlid="nl46" bibid="bib108" firstref="ref88"></nolink> <nolink nlid="nl47" bibid="bib48" firstref="ref90"></nolink> <nolink nlid="nl48" bibid="bib112" firstref="ref91"></nolink> <nolink nlid="nl49" bibid="bib109" firstref="ref92"></nolink> <nolink nlid="nl50" bibid="bib89" firstref="ref93"></nolink> <nolink nlid="nl51" bibid="bib100" firstref="ref109"></nolink> <nolink nlid="nl52" bibid="bib21" firstref="ref110"></nolink> <nolink nlid="nl53" bibid="bib53" firstref="ref112"></nolink> <nolink nlid="nl54" bibid="bib87" firstref="ref115"></nolink> <nolink nlid="nl55" bibid="bib54" firstref="ref119"></nolink> <nolink nlid="nl56" bibid="bib102" firstref="ref120"></nolink> <nolink nlid="nl57" bibid="bib81" firstref="ref122"></nolink> <nolink nlid="nl58" bibid="bib82" firstref="ref123"></nolink> <nolink nlid="nl59" bibid="bib62" firstref="ref131"></nolink> <nolink nlid="nl60" bibid="bib70" firstref="ref139"></nolink> <nolink nlid="nl61" bibid="bib29" firstref="ref140"></nolink> <nolink nlid="nl62" bibid="bib98" firstref="ref142"></nolink> <nolink nlid="nl63" bibid="bib16" firstref="ref146"></nolink> <nolink nlid="nl64" bibid="bib99" firstref="ref154"></nolink> <nolink nlid="nl65" bibid="bib105" firstref="ref155"></nolink> <nolink nlid="nl66" bibid="bib66" firstref="ref156"></nolink> <nolink nlid="nl67" bibid="bib12" firstref="ref157"></nolink> <nolink nlid="nl68" bibid="bib69" firstref="ref158"></nolink> <nolink nlid="nl69" bibid="bib97" firstref="ref159"></nolink> <nolink nlid="nl70" bibid="bib46" firstref="ref161"></nolink> <nolink nlid="nl71" bibid="bib39" firstref="ref165"></nolink> <nolink nlid="nl72" bibid="bib103" firstref="ref166"></nolink> <nolink nlid="nl73" bibid="bib78" firstref="ref167"></nolink> <nolink nlid="nl74" bibid="bib79" firstref="ref168"></nolink> <nolink nlid="nl75" bibid="bib38" firstref="ref169"></nolink> <nolink nlid="nl76" bibid="bib35" firstref="ref171"></nolink> <nolink nlid="nl77" bibid="bib60" firstref="ref172"></nolink> <nolink nlid="nl78" bibid="bib76" firstref="ref174"></nolink> <nolink nlid="nl79" bibid="bib104" firstref="ref175"></nolink> <nolink nlid="nl80" bibid="bib52" firstref="ref176"></nolink> <nolink nlid="nl81" bibid="bib77" firstref="ref177"></nolink> <nolink nlid="nl82" bibid="bib41" firstref="ref178"></nolink> <nolink nlid="nl83" bibid="bib13" firstref="ref179"></nolink> <nolink nlid="nl84" bibid="bib73" firstref="ref180"></nolink> <nolink nlid="nl85" bibid="bib111" firstref="ref189"></nolink> <nolink nlid="nl86" bibid="bib55" firstref="ref192"></nolink> <nolink nlid="nl87" bibid="bib58" firstref="ref193"></nolink> <nolink nlid="nl88" bibid="bib22" firstref="ref195"></nolink> <nolink nlid="nl89" bibid="bib19" firstref="ref196"></nolink> <nolink nlid="nl90" bibid="bib50" firstref="ref197"></nolink> <nolink nlid="nl91" bibid="bib49" firstref="ref199"></nolink> <nolink nlid="nl92" bibid="bib20" firstref="ref215"></nolink> <nolink nlid="nl93" bibid="bib24" firstref="ref231"></nolink> <nolink nlid="nl94" bibid="bib30" firstref="ref234"></nolink> <nolink nlid="nl95" bibid="bib106" firstref="ref236"></nolink> <nolink nlid="nl96" bibid="bib107" firstref="ref241"></nolink> <nolink nlid="nl97" bibid="bib59" firstref="ref242"></nolink> <nolink nlid="nl98" bibid="bib110" firstref="ref243"></nolink> <nolink nlid="nl99" bibid="bib61" firstref="ref245"></nolink> <nolink nlid="nl100" bibid="bib83" firstref="ref249"></nolink> <nolink nlid="nl101" bibid="bib23" firstref="ref250"></nolink> |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1457332 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: An Individual Participant Data Meta-Analysis to Support Power Analyses for Randomized Intervention Studies in Preschool: Cognitive and Socio-Emotional Learning Outcomes – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Martin+Brunner%22">Martin Brunner</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0001-7182-5622">0000-0001-7182-5622</externalLink>)<br /><searchLink fieldCode="AR" term="%22Sophie+E%2E+Stallasch%22">Sophie E. Stallasch</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0002-4433-2600">0000-0002-4433-2600</externalLink>)<br /><searchLink fieldCode="AR" term="%22Cordula+Artelt%22">Cordula Artelt</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0001-7790-2502">0000-0001-7790-2502</externalLink>)<br /><searchLink fieldCode="AR" term="%22Oliver+Lüdtke%22">Oliver Lüdtke</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0001-9744-3059">0000-0001-9744-3059</externalLink>) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Educational+Psychology+Review%22"><i>Educational Psychology Review</i></searchLink>. 2025 37. – Name: Avail Label: Availability Group: Avail Data: Springer. Available from: Springer Nature. One New York Plaza, Suite 4600, New York, NY 10004. Tel: 800-777-4643; Tel: 212-460-1500; Fax: 212-460-1700; e-mail: customerservice@springernature.com; Web site: https://link.springer.com/ – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 38 – Name: DatePubCY Label: Publication Date Group: Date Data: 2025 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Information Analyses – Name: Audience Label: Education Level Group: Audnce Data: <searchLink fieldCode="EL" term="%22Early+Childhood+Education%22">Early Childhood Education</searchLink><br /><searchLink fieldCode="EL" term="%22Preschool+Education%22">Preschool Education</searchLink> – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Preschools%22">Preschools</searchLink><br /><searchLink fieldCode="DE" term="%22Social+Emotional+Learning%22">Social Emotional Learning</searchLink><br /><searchLink fieldCode="DE" term="%22Outcomes+of+Education%22">Outcomes of Education</searchLink><br /><searchLink fieldCode="DE" term="%22Cognitive+Objectives%22">Cognitive Objectives</searchLink><br /><searchLink fieldCode="DE" term="%22Child+Care+Centers%22">Child Care Centers</searchLink><br /><searchLink fieldCode="DE" term="%22Foreign+Countries%22">Foreign Countries</searchLink><br /><searchLink fieldCode="DE" term="%22Preschool+Children%22">Preschool Children</searchLink><br /><searchLink fieldCode="DE" term="%22Child+Development%22">Child Development</searchLink><br /><searchLink fieldCode="DE" term="%22Journal+Articles%22">Journal Articles</searchLink><br /><searchLink fieldCode="DE" term="%22Statistical+Analysis%22">Statistical Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Early+Childhood+Education%22">Early Childhood Education</searchLink><br /><searchLink fieldCode="DE" term="%22Educational+Research%22">Educational Research</searchLink><br /><searchLink fieldCode="DE" term="%22Randomized+Controlled+Trials%22">Randomized Controlled Trials</searchLink> – Name: Subject Label: Geographic Terms Group: Su Data: <searchLink fieldCode="DE" term="%22Germany%22">Germany</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1007/s10648-024-09981-z – Name: ISSN Label: ISSN Group: ISSN Data: 1040-726X<br />1573-336X – Name: Abstract Label: Abstract Group: Ab Data: There is a need for robust evidence about which educational interventions work in preschool to foster children's cognitive and socio-emotional learning (SEL) outcomes. Lab-based individually randomized experiments can develop and refine such interventions, and field-based randomized experiments (e.g., cluster randomized trials) evaluate their effectiveness in real-world daycare center settings. Applying reliable estimates of design parameters in the context of a priori power analyses is essential to ensure that the sample size of these studies is adequate to support strong statistical conclusions regarding the strength of the intervention effect. However, there is little knowledge on relevant design parameters with preschool children. We therefore utilized a systematic collection of individual participant data from four German probability samples (554 [less than or equal to] N [less than or equal to] 2928) with preschool children (aged two to six years) to estimate and meta-analyze design parameters. These parameters are relevant for planning single-level (e.g., in non-clustered lab-based settings), two-level (children nested in daycare centers), and three-level (children nested in groups, with groups nested in daycare centers) randomized intervention studies targeting cognitive and SEL outcomes assessed with three methods (standardized tests, parent ratings, and educator ratings). The design parameters depict between-group and -center differences as well as the proportion of variance in the outcomes explained by different covariate sets (socio-demographic characteristics, baseline measures, and their combination) at the child, group, and center level. In conclusion, this paper provides a rich source of design parameters, recommendations, and illustrations to support a priori power analyses for randomized intervention studies in early childhood education research. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2025 – Name: AN Label: Accession Number Group: ID Data: EJ1457332 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1457332 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1007/s10648-024-09981-z Languages: – Text: English PhysicalDescription: Pagination: PageCount: 38 Subjects: – SubjectFull: Preschools Type: general – SubjectFull: Social Emotional Learning Type: general – SubjectFull: Outcomes of Education Type: general – SubjectFull: Cognitive Objectives Type: general – SubjectFull: Child Care Centers Type: general – SubjectFull: Foreign Countries Type: general – SubjectFull: Preschool Children Type: general – SubjectFull: Child Development Type: general – SubjectFull: Journal Articles Type: general – SubjectFull: Statistical Analysis Type: general – SubjectFull: Early Childhood Education Type: general – SubjectFull: Educational Research Type: general – SubjectFull: Randomized Controlled Trials Type: general – SubjectFull: Germany Type: general Titles: – TitleFull: An Individual Participant Data Meta-Analysis to Support Power Analyses for Randomized Intervention Studies in Preschool: Cognitive and Socio-Emotional Learning Outcomes Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Martin Brunner – PersonEntity: Name: NameFull: Sophie E. Stallasch – PersonEntity: Name: NameFull: Cordula Artelt – PersonEntity: Name: NameFull: Oliver Lüdtke IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 03 Type: published Y: 2025 Identifiers: – Type: issn-print Value: 1040-726X – Type: issn-electronic Value: 1573-336X Numbering: – Type: volume Value: 37 Titles: – TitleFull: Educational Psychology Review Type: main |
| ResultId | 1 |