Meta-Regression Methods to Characterize Evidence Strength Using Meaningful-Effect Percentages Conditional on Study Characteristics

Saved in:
Bibliographic Details
Title: Meta-Regression Methods to Characterize Evidence Strength Using Meaningful-Effect Percentages Conditional on Study Characteristics
Language: English
Authors: Mathur, Maya B. (ORCID 0000-0001-6698-2607), VanderWeele, Tyler J. (ORCID 0000-0002-6112-0239)
Source: Research Synthesis Methods. Nov 2021 12(6):731-749.
Availability: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 19
Publication Date: 2021
Sponsoring Agency: National Institutes of Health (DHHS)
Contract Number: CA222147
P30CA124435
P30DK116074
UL1TR003142
Document Type: Journal Articles
Reports - Research
Descriptors: Regression (Statistics), Meta Analysis, Effect Size, Computation, Statistical Inference
DOI: 10.1002/jrsm.1504
ISSN: 1759-2879
Abstract: Meta-regression analyses usually focus on estimating and testing differences in average effect sizes between individual levels of each meta-regression covariate in turn. These metrics are useful but have limitations: they consider each covariate individually, rather than in combination, and they characterize only the mean of a potentially heterogeneous distribution of effects. We propose additional metrics that address both limitations. Given a chosen threshold representing a meaningfully strong effect size, these metrics address the questions: "For a given joint level of the covariates, what percentage of the population effects are meaningfully strong?" and "For any two joint levels of the covariates, what is the difference between these percentages of meaningfully strong effects?" We provide semiparametric methods for estimation and inference and assess their performance in a simulation study. We apply the proposed methods to meta-regression analyses on memory consolidation and on dietary behavior interventions, illustrating how the methods can provide more information than standard reporting alone. To facilitate implementing the methods in practice, we provide reporting guidelines and simple R code.
Abstractor: As Provided
Notes: https://osf.io/gs7fp
Entry Date: 2021
Accession Number: EJ1316105
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwHHd6B78LslMhGV0qSftKRCAAAA4jCB3wYJKoZIhvcNAQcGoIHRMIHOAgEAMIHIBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDMqF4upff1k_JsEh7QIBEICBmjnu_SHzebJnDpaHJxK04p4fHGpxTnG1UkPEAC7EZGMmZPeMMrQTclQkio-jBwJaeQLUyxiTNhXx6UC0fFiBgDNunWIBMlNitxzZosGe3syGNJbaLLhfAhIGhHOxMLZjWgx-EBPllJOAIyT55_C61_IqnkYlC0M3kuH8YcMnVPjzF_Y0ShcDP_M4wv7C4Wrkm9QUXs5D_eQ8K5g=
Text:
  Availability: 1
  Value: <anid>AN0153408964;[bdct]01nov.21;2021Nov08.04:51;v2.2.500</anid> <title id="AN0153408964-1">Meta‐regression methods to characterize evidence strength using meaningful‐effect percentages conditional on study characteristics </title> <p>Meta‐regression analyses usually focus on estimating and testing differences in average effect sizes between individual levels of each meta‐regression covariate in turn. These metrics are useful but have limitations: they consider each covariate individually, rather than in combination, and they characterize only the mean of a potentially heterogeneous distribution of effects. We propose additional metrics that address both limitations. Given a chosen threshold representing a meaningfully strong effect size, these metrics address the questions: "For a given joint level of the covariates, what percentage of the population effects are meaningfully strong?" and "For any two joint levels of the covariates, what is the difference between these percentages of meaningfully strong effects?" We provide semiparametric methods for estimation and inference and assess their performance in a simulation study. We apply the proposed methods to meta‐regression analyses on memory consolidation and on dietary behavior interventions, illustrating how the methods can provide more information than standard reporting alone. To facilitate implementing the methods in practice, we provide reporting guidelines and simple R code.</p> <p>Keywords: bootstrapping; effect sizes; heterogeneity; meta‐analysis; meta‐regression; semiparametric</p> <hd id="AN0153408964-2">Highlights</hd> <p></p> <hd id="AN0153408964-3">What is already known</hd> <p></p> <ulist> <item> Meta‐regression analyses usually focus on differences in average effects between levels of each covariate.</item> <p></p> <item> While useful, these metrics have limitations: they consider each covariate individually, rather than in combination, and they characterize only the mean of a potentially heterogeneous distribution of effects.</item> </ulist> <hd id="AN0153408964-4">What is new</hd> <p></p> <ulist> <item> We propose new metrics that address the questions: "For a given joint level of the meta‐regression covariates, what percentage of the population effects are meaningfully strong?" and "For any two joint levels of the covariates, what is the difference between these percentages of meaningfully strong effects?"</item> <p></p> <item> These metrics characterize the heterogeneous distribution of effects conditional on the covariates (not just the distribution's mean), and they enable direct comparison of joint levels of covariates.</item> <p></p> <item> The first and second metrics above can be applied to meta‐regressions with at least 10 and at least 20 studies, respectively. The metrics should be reported along with confidence intervals. Caution is warranted if the point estimates are both clustered and skewed. When using the second metric, the specified contrast should not use extreme covariate values.</item> </ulist> <hd id="AN0153408964-5">Potential impact for RSM readers outside the authors' field</hd> <p></p> <ulist> <item> These metrics could facilitate assessing and communicating how evidence strength differs for studies with different sets of characteristics.</item> </ulist> <hd id="AN0153408964-6">INTRODUCTION</hd> <p>Meta‐regression analyses usually focus on estimating and testing differences in average effect sizes between individual levels of each meta‐regression covariate in turn.1 These estimates are certainly useful, but do have limitations as standalone metrics. First, they consider each meta‐regression covariate individually and do not directly characterize differences in effect sizes associated with combinations of covariates that are of scientific interest. For example, if the covariates represent possible components of a behavior intervention, it may be useful to consider the strength of effects in studies with a particular combination of components rather than only estimating average effects of each component individually. This approach could be particularly useful given recent calls to conduct meta‐analyses that deliberately include studies representing a broad range of interventions, populations, and environments.2 Similarly, when meta‐regression is used to assess the association of studies' risk‐of‐bias characteristics with their effect sizes,1 it would often be useful to consider the strength of effects in studies with low risks of bias on all measures jointly, rather than individually (although one can never be certain that all relevant risks of bias have been assessed, nor that they have been rated with complete accuracy).</p> <p>Second, and more fundamentally, the usual estimates of average effect sizes characterize only the mean of a potentially heterogeneous distribution of effects. We therefore propose additional metrics that supplement standard reporting by directly addressing questions of fundamental interest in meta‐regression and characterizing the potentially heterogeneous distribution of effect sizes conditional on specified levels of the meta‐regression covariates. Specifically, in a manner we formalize below, the proposed metrics address two questions: (<reflink idref="bib1" id="ref1">1</reflink>) For a given joint level of the covariates, what percentage of the population effects are "meaningfully strong"? (<reflink idref="bib2" id="ref2">2</reflink>) For any two joint levels of the covariates, what is the difference between these percentages of meaningfully strong effects?</p> <p>These metrics extend methods we previously proposed for standard meta‐analysis.3–5 Specifically, we had previously recommended choosing a minimum threshold representing a meaningfully strong effect size ( <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0001" xmlns="http://www.w3.org/1998/Math/MathML"><mi>q</mi></math> </ephtml> ) and estimating the percentage or proportion of population effects above this threshold. Second, we and others3,6 have suggested estimating the percentage of effects below a second, possibly symmetric, threshold in the opposite direction from the estimated mean. We discussed a number of methods to choose these thresholds, which included considering the size of discrepancies between naturally occurring groups of interest, effect sizes produced by well‐evidenced interventions, cost‐effectiveness analyses, or minimum subjectively perceptible thresholds.3 We demonstrated that these percentage metrics could help convey evidence strength for meaningfully strong effects under effect heterogeneity in a manner that provides more information than meta‐analytic point estimates alone3 and also that they could sometimes help adjudicate apparent conflicts between meta‐analyses.7,8 In practice, the metrics have been successfully and informatively applied to meta‐analyses on a variety of topics.7,9–13</p> <p>Here, we provide extensions to meta‐regression that address the two questions above by characterizing, for a chosen level of the meta‐regression covariates, the percentage of population effects that are above or below the threshold <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0002" xmlns="http://www.w3.org/1998/Math/MathML"><mi>q</mi></math> </ephtml> . This metric helps characterize evidence strength for meaningfully strong effects in studies with a particular level of the covariates, and it could also be used as a hypothesis‐generating method to identify which joint levels of the covariates are associated with the largest estimated percentages of meaningfully strong effects. Naturally, this metric also allows one to characterize the complementary cumulative distribution function of the population effects (i.e., the percentage of effects stronger than any arbitrary threshold), which could be displayed graphically. The methods also allow one to estimate the difference in these percentages for any two joint levels of the covariates.</p> <p>We provide methods to estimate these metrics along with inference (Section 2.2). The methods involve first fitting a standard meta‐regression (e.g., using semiparametric methods that do not make assumptions on the distribution of population effects14–16), using the resulting estimates to appropriately "shrink" studies' point estimates toward the mean, and then estimating the proposed metrics using the empirical distribution of these shrunken estimates. We illustrate by re‐analyzing data from two previously published meta‐analyses,13,17 demonstrating that the proposed metrics can provide more information than standard reporting alone (Section 4). We assess the methods' performance in a simulation study that includes a variety of realistic and challenging scenarios (Section 5) and use the results to inform practical reporting guidelines (Section 2.3). The methods are straightforward to implement in practice, and we provide simple example R code to do so (Supplementary material or https://osf.io/gs7fp/).</p> <hd id="AN0153408964-7">METHODS</hd> <p></p> <hd id="AN0153408964-8">Existing methods for standard meta‐analysis</hd> <p>We first briefly review the previously proposed methods3–5 to estimate the percentage of meaningfully strong effects, termed <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0003" xmlns="http://www.w3.org/1998/Math/MathML"><mo>"</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub></math> </ephtml> ", in the context of standard meta‐analysis. Let <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0004" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>θ</mi><mi>i</mi></msub></math> </ephtml> , <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0005" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>θ</mi><mo>^</mo></mover><mi>i</mi></msub></math> </ephtml> , and <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0006" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>σ</mi><mo>^</mo></mover><mi>i</mi></msub></math> </ephtml> respectively denote the population effect size, point estimate, and estimated standard error of the <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0007" xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>i</mi><mi mathvariant="italic">th</mi></msup></math> </ephtml> study. Consider a standard meta‐analysis of <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0008" xmlns="http://www.w3.org/1998/Math/MathML"><mi>k</mi></math> </ephtml> independent studies, with <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0009" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true"><mi>μ</mi><mo>^</mo></mover></math> </ephtml> denoting the estimated mean and <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0010" xmlns="http://www.w3.org/1998/Math/MathML"><msup><mover accent="true"><mi>τ</mi><mo>^</mo></mover><mn>2</mn></msup></math> </ephtml> denoting the estimated heterogeneity (i.e., the estimated variance of the population effects). To estimate <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0011" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub></math> </ephtml> , existing methods begin by calculating a "calibrated" estimate5 for each meta‐analyzed study, defined as</p> <p>2.1 <ephtml> <math display="block" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0012" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>θ</mi><mo>˜</mo></mover><mi>i</mi></msub><mo>=</mo><mover accent="true"><mi>μ</mi><mo>^</mo></mover><mo>+</mo><msqrt><mfrac><msup><mover accent="true"><mi>τ</mi><mo>^</mo></mover><mn>2</mn></msup><mrow><msup><mover accent="true"><mi>τ</mi><mo>^</mo></mover><mn>2</mn></msup><mo>+</mo><msubsup><mover accent="true"><mi>σ</mi><mo>^</mo></mover><mi>i</mi><mn>2</mn></msubsup></mrow></mfrac></msqrt><mfenced open="(" close=")"><mrow><msub><mover accent="true"><mi>θ</mi><mo>^</mo></mover><mi>i</mi></msub><mo>−</mo><mover accent="true"><mi>μ</mi><mo>^</mo></mover></mrow></mfenced></math> </ephtml></p> <p>Intuitively, the calibrated estimate <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0013" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>θ</mi><mo>˜</mo></mover><mi>i</mi></msub></math> </ephtml> shrinks the point estimate <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0014" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>θ</mi><mo>^</mo></mover><mi>i</mi></msub></math> </ephtml> toward the estimated mean <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0015" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true"><mi>μ</mi><mo>^</mo></mover></math> </ephtml> with a degree of shrinkage that is inversely proportional to the study's precision: relatively imprecise estimates <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0016" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>θ</mi><mo>^</mo></mover><mi>i</mi></msub></math> </ephtml> (i.e., those with large <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0017" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mover accent="true"><mi>σ</mi><mo>^</mo></mover><mi>i</mi><mn>2</mn></msubsup></math> </ephtml> ) receive strong shrinkage toward <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0018" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true"><mi>μ</mi><mo>^</mo></mover></math> </ephtml> , while relatively precise estimates receive less shrinkage and remain closer to their original values. Thus, these calibrated estimates have been appropriately shrunk to correct the overdispersion such that their variance is equal to, the estimated variance of the population effects.5 The shrinkage factor <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0019" xmlns="http://www.w3.org/1998/Math/MathML"><msqrt><mrow><msup><mover accent="true"><mi>τ</mi><mo>^</mo></mover><mn>2</mn></msup><mo>/</mo><mfenced open="(" close=")"><mrow><msup><mover accent="true"><mi>τ</mi><mo>^</mo></mover><mn>2</mn></msup><mo>+</mo><msubsup><mover accent="true"><mi>σ</mi><mo>^</mo></mover><mi>i</mi><mn>2</mn></msubsup></mrow></mfenced></mrow></msqrt></math> </ephtml> minimizes the distance between the empirical cumulative distribution function of the calibrated estimates and of the population effects,18 which is the relevant loss function for estimating <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0020" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub></math> </ephtml> . Then, <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0021" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub></math> </ephtml> can be easily estimated4 as a sample proportion of the calibrated estimates that are stronger than <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0022" xmlns="http://www.w3.org/1998/Math/MathML"><mi>q</mi></math> </ephtml> :</p> <p>2.2 <ephtml> <math display="block" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0023" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mo>=</mo><munderover><mrow><mo>∑</mo></mrow><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mi>k</mi></munderover><mn mathvariant="double-struck">1</mn><mfenced open="{" close="}"><mrow><mover><mi>μ</mi><mo>^</mo></mover><mo>+</mo><msqrt><mfrac><msup><mover><mi>τ</mi><mo>^</mo></mover><mn>2</mn></msup><mrow><msup><mover><mi>τ</mi><mo>^</mo></mover><mn>2</mn></msup><mo>+</mo><msubsup><mover><mi>σ</mi><mo>^</mo></mover><mi>i</mi><mn>2</mn></msubsup></mrow></mfrac></msqrt><mfenced><mrow><msub><mover><mi>θ</mi><mo>^</mo></mover><mi>i</mi></msub><mo>−</mo><mover><mi>μ</mi><mo>^</mo></mover></mrow></mfenced><mo>></mo><mi>q</mi></mrow></mfenced></math> </ephtml></p> <p>Naturally, analogous methods can be used to estimate the proportion of effects below another threshold, for example to consider the percentage of effects that are in the direction opposite the overall estimated mean.3,6 Inference can be conducted using bias‐corrected and accelerated (BCa) bootstrapping.4,19,20 If the point estimates are potentially non‐independent because, for example, some articles in the meta‐analysis contribute multiple estimates from studies with similar study designs or populations, then one should resample clusters of estimates (e.g., articles) with replacement while leaving intact the estimates within each cluster (Davison and Hinkley21 Section 3.8). We will refer to this as the "cluster bootstrap".</p> <hd id="AN0153408964-9">Extension to meta‐regression</hd> <p>We now extend the above methods to a meta‐regression with the mean model <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0024" xmlns="http://www.w3.org/1998/Math/MathML"><mi>E</mi><mfenced open="[" close="]" separators="|"><mrow><mi>θ</mi><mspace width="0.5em" /></mrow><mrow><mspace width="0.5em" /><mi>Z</mi></mrow></mfenced><mo>=</mo><msub><mi>β</mi><mn>0</mn></msub><mo>+</mo><mi>Z</mi><msub><mi>β</mi><mn>1</mn></msub></math> </ephtml> , where <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0025" xmlns="http://www.w3.org/1998/Math/MathML"><mi>Z</mi></math> </ephtml> is a <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0026" xmlns="http://www.w3.org/1998/Math/MathML"><mi>k</mi><mo>×</mo><mi>p</mi></math> </ephtml> matrix of study‐level covariates of any type (binary, categorical, continuous, etc.) with realized levels in study <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0027" xmlns="http://www.w3.org/1998/Math/MathML"><mi>i</mi></math> </ephtml> of <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0028" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>z</mi><mi>i</mi></msub></math> </ephtml> and where <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0029" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>β</mi><mn>1</mn></msub></math> </ephtml> is a <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0030" xmlns="http://www.w3.org/1998/Math/MathML"><mi>p</mi></math> </ephtml> ‐vector of coefficients. We assume that the residual variance of the population effects, <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0031" xmlns="http://www.w3.org/1998/Math/MathML"><mi>Var</mi><mfenced open="(" close=")" separators="|"><mrow><mi>θ</mi><mspace width="0.5em" /></mrow><mrow><mspace width="0.5em" /><mi>Z</mi></mrow></mfenced></math> </ephtml> , is a constant <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0032" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mi>τ</mi><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup></math> </ephtml> . This is a standard estimand in meta‐regression and is often simply called <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0033" xmlns="http://www.w3.org/1998/Math/MathML"><mo>"</mo><msup><mi>τ</mi><mn>2</mn></msup></math> </ephtml> " in the literature and in software; here, we adopt the notation <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0034" xmlns="http://www.w3.org/1998/Math/MathML"><mo>"</mo><msubsup><mi>τ</mi><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup></math> </ephtml> " to clarify that this is the residual heterogeneity conditional on the meta‐regression covariates rather than the marginal <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0035" xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>τ</mi><mn>2</mn></msup></math> </ephtml> of a standard meta‐analysis. One would first estimate the parameters <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0036" xmlns="http://www.w3.org/1998/Math/MathML"><mfenced open="(" close=")" separators=",,"><msub><mi>β</mi><mn>0</mn></msub><msub><mi>β</mi><mn>1</mn></msub><msubsup><mi>τ</mi><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup></mfenced></math> </ephtml> via standard meta‐regression; in practice, we would recommend using semiparametric methods similar to generalized estimating equations that do not require assumptions on the distribution of population effects.14–16,22 Asymptotic and finite‐sample theory establishing that this approach provides consistent coefficient estimates under arbitrary distributions was provided elsewhere.14,22</p> <p>The proportion of population effects above <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0037" xmlns="http://www.w3.org/1998/Math/MathML"><mi>q</mi></math> </ephtml> , conditional on level <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0038" xmlns="http://www.w3.org/1998/Math/MathML"><mi>z</mi></math> </ephtml> of the covariates, is <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0039" xmlns="http://www.w3.org/1998/Math/MathML"><mi>P</mi><mfenced open="(" close=")" separators="|"><mrow><mi>θ</mi><mo>></mo><mi>q</mi><mspace width="0.5em" /></mrow><mrow><mspace width="0.5em" /><mi>Z</mi><mo>=</mo><mi>z</mi></mrow></mfenced></math> </ephtml> , termed <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0040" xmlns="http://www.w3.org/1998/Math/MathML"><mo>"</mo><msub><mi>P</mi><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> ." While it may seem intuitive to estimate <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0041" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>P</mi><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> by simply analyzing a subset of the studies and applying the existing methods for standard meta‐analysis,3,4 that approach would preclude consideration of continuous covariates and would be inefficient, especially when considering specific combinations of covariates that do not occur frequently in the observed studies. Instead, we propose estimating <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0042" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>P</mi><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> as follows. Define a point estimate for each study that has been "shifted" to the chosen covariate level <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0043" xmlns="http://www.w3.org/1998/Math/MathML"><mi>z</mi></math> </ephtml> as <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0044" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>θ</mi><mo>^</mo></mover><mi>i</mi></msub><mfenced open="(" close=")"><mrow><mi>Z</mi><mo>=</mo><mi>z</mi></mrow></mfenced><mo>=</mo><msub><mover accent="true"><mi>θ</mi><mo>^</mo></mover><mi>i</mi></msub><mo>−</mo><mfenced open="(" close=")"><mrow><msub><mi>z</mi><mi>i</mi></msub><mo>−</mo><mi>z</mi></mrow></mfenced><msub><mover accent="true"><mi>β</mi><mo>^</mo></mover><mn>1</mn></msub></math> </ephtml> , such that <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0045" xmlns="http://www.w3.org/1998/Math/MathML"><mi>E</mi><mfenced open="[" close="]"><mrow><msub><mover accent="true"><mi>θ</mi><mo>^</mo></mover><mi>i</mi></msub><mfenced open="(" close=")"><mrow><mi>Z</mi><mo>=</mo><mi>z</mi></mrow></mfenced></mrow></mfenced><mo>=</mo><msub><mi>β</mi><mn>0</mn></msub><mo>+</mo><mi>z</mi><msub><mi>β</mi><mn>1</mn></msub></math> </ephtml> . An analog to the calibrated estimate <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0046" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>θ</mi><mo>˜</mo></mover><mi>i</mi></msub></math> </ephtml> that has been shifted to <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0047" xmlns="http://www.w3.org/1998/Math/MathML"><mi>Z</mi><mo>=</mo><mi>z</mi></math> </ephtml> is then:</p> <p>2.3 <ephtml> <math display="block" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0048" xmlns="http://www.w3.org/1998/Math/MathML">θ˜iZ=z=E^θZ=z+τ^ɛ2τ^ɛ2+σ^i2θ^iZ=z−E^[θZ=z]=β^0+zβ^1+τ^ɛ2τ^ɛ2+σ^i2θ^i−β^0+ziβ^1</math> </ephtml></p> <p>In the second line, upon canceling the <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0049" xmlns="http://www.w3.org/1998/Math/MathML"><mi>z</mi><msub><mover accent="true"><mi>β</mi><mo>^</mo></mover><mn>1</mn></msub></math> </ephtml> terms, the term <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0050" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>θ</mi><mo>^</mo></mover><mi>i</mi></msub><mo>−</mo><mfenced open="[" close="]"><mrow><msub><mover accent="true"><mi>β</mi><mo>^</mo></mover><mn>0</mn></msub><mo>+</mo><msub><mi>z</mi><mi>i</mi></msub><msub><mover accent="true"><mi>β</mi><mo>^</mo></mover><mn>1</mn></msub></mrow></mfenced></math> </ephtml> is simply the (unshifted) residual of <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0051" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>θ</mi><mo>^</mo></mover><mi>i</mi></msub></math> </ephtml> with respect to its estimated expectation conditional on its realized <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0052" xmlns="http://www.w3.org/1998/Math/MathML"><mi>Z</mi><mo>=</mo><msub><mi>z</mi><mi>i</mi></msub></math> </ephtml> . Highly imprecise point estimates (i.e., those with a small <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0053" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mover accent="true"><mi>τ</mi><mo>^</mo></mover><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup><mo>/</mo><mfenced open="(" close=")"><mrow><msubsup><mover accent="true"><mi>τ</mi><mo>^</mo></mover><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup><mo>+</mo><msubsup><mover accent="true"><mi>σ</mi><mo>^</mo></mover><mi>i</mi><mn>2</mn></msubsup></mrow></mfenced></math> </ephtml> ) are strongly shrunk toward <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0054" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true"><mi>E</mi><mo>^</mo></mover><mfenced open="[" close="]" separators="|"><mrow><mi>θ</mi><mspace width="0.5em" /></mrow><mrow><mspace width="0.5em" /><mi>Z</mi><mo>=</mo><mi>z</mi></mrow></mfenced><mo>=</mo><msub><mover accent="true"><mi>β</mi><mo>^</mo></mover><mn>0</mn></msub><mo>+</mo><mi>z</mi><msub><mover accent="true"><mi>β</mi><mo>^</mo></mover><mn>1</mn></msub></math> </ephtml> , while more precise estimates remain close to the studies' own residuals, <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0055" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>θ</mi><mo>^</mo></mover><mi>i</mi></msub><mo>−</mo><mfenced open="[" close="]"><mrow><msub><mover accent="true"><mi>β</mi><mo>^</mo></mover><mn>0</mn></msub><mo>+</mo><msub><mi>z</mi><mi>i</mi></msub><msub><mover accent="true"><mi>β</mi><mo>^</mo></mover><mn>1</mn></msub></mrow></mfenced></math> </ephtml> .</p> <p>Analogously to the fact that standard calibrated estimates approximately match the first two moments of the marginal distribution of population effects,5 the shifted, calibrated estimates <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0056" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>θ</mi><mo>˜</mo></mover><mi>i</mi></msub><mfenced open="(" close=")"><mrow><mi>Z</mi><mo>=</mo><mi>z</mi></mrow></mfenced></math> </ephtml> approximately match the first two moments of the distribution of the population effects conditional on <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0057" xmlns="http://www.w3.org/1998/Math/MathML"><mi>Z</mi><mo>=</mo><mi>z</mi></math> </ephtml> . That is, <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0058" xmlns="http://www.w3.org/1998/Math/MathML"><mi>E</mi><mfenced open="[" close="]" separators="|"><mrow><msub><mover accent="true"><mi>θ</mi><mo>˜</mo></mover><mi>i</mi></msub><mfenced open="(" close=")"><mrow><mi>Z</mi><mo>=</mo><mi>z</mi></mrow></mfenced><mspace width="0.5em" /></mrow><mrow><mspace width="0.5em" /><mi>Z</mi><mo>=</mo><mi>z</mi></mrow></mfenced><mo>=</mo><msub><mi>β</mi><mn>0</mn></msub><mo>+</mo><mi>z</mi><msub><mi>β</mi><mn>1</mn></msub><mo>=</mo><mi>E</mi><mfenced open="[" close="]" separators="|"><mrow><mi>θ</mi><mspace width="0.5em" /></mrow><mrow><mspace width="0.5em" /><mi>Z</mi><mo>=</mo><mi>z</mi></mrow></mfenced></math> </ephtml> , and for large k:</p> <p> <ephtml> <math display="block" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0059" xmlns="http://www.w3.org/1998/Math/MathML">Varθ˜iZ=zZ=z≈τ^ɛ2τ^ɛ2+σ^i2Varθ^iZ=z≈τ^ɛ2τ^ɛ2+σ^i2τ^ɛ2+σ^i2=τ^ɛ2≈VarθZ=z</math> </ephtml> </p> <p>The proportion of meaningfully strong effects can then be estimated as:</p> <p>2.4 <ephtml> <math display="block" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0060" xmlns="http://www.w3.org/1998/Math/MathML">P^>qz=P^β^0+zβ^1+τ^ɛ2τ^ɛ2+σ^i2θ^i−β^0+ziβ^1>q=∑i=1k1β0^+τ^ɛ2τ^ɛ2+σ^i2θ^i−ziβ^1−β^0>q−zβ^1</math> </ephtml></p> <p>The estimated difference in these proportions comparing level <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0061" xmlns="http://www.w3.org/1998/Math/MathML"><mi>z</mi></math> </ephtml> to a chosen reference level <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0062" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>z</mi><mn>0</mn></msub></math> </ephtml> is simply <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0063" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> . Again, analogous methods can be used to estimate the proportion of effects below another threshold, or their difference.</p> <p>To apply Equation (2.4) in practice, one could simply plug in the meta‐regression estimates <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0064" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true"><msub><mi>β</mi><mn>0</mn></msub><mo>^</mo></mover></math> </ephtml> , <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0065" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true"><msub><mi>β</mi><mn>1</mn></msub><mo>^</mo></mover></math> </ephtml> , and <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0066" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mover accent="true"><mi>τ</mi><mo>^</mo></mover><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup></math> </ephtml> ; we call this the "one‐stage" method. When considering multiple choices of <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0067" xmlns="http://www.w3.org/1998/Math/MathML"><mi>q</mi></math> </ephtml> or <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0068" xmlns="http://www.w3.org/1998/Math/MathML"><mi>z</mi></math> </ephtml> , one can simply compute a single value of <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0069" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true"><msub><mi>β</mi><mn>0</mn></msub><mo>^</mo></mover><mo>+</mo><msqrt><mrow><msubsup><mover accent="true"><mi>τ</mi><mo>^</mo></mover><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup><mo>/</mo><mfenced open="(" close=")"><mrow><msubsup><mover accent="true"><mi>τ</mi><mo>^</mo></mover><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup><mo>+</mo><msubsup><mover accent="true"><mi>σ</mi><mo>^</mo></mover><mi>i</mi><mn>2</mn></msubsup></mrow></mfenced></mrow></msqrt><mfenced open="(" close=")"><mrow><msub><mover accent="true"><mi>θ</mi><mo>^</mo></mover><mi>i</mi></msub><mo>−</mo><msub><mi>z</mi><mi>i</mi></msub><msub><mover accent="true"><mi>β</mi><mo>^</mo></mover><mn>1</mn></msub><mo>−</mo><msub><mover accent="true"><mi>β</mi><mo>^</mo></mover><mn>0</mn></msub></mrow></mfenced></math> </ephtml> for each study and then compare these to various shifted thresholds, <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0070" xmlns="http://www.w3.org/1998/Math/MathML"><mi>q</mi><mo>−</mo><mi>z</mi><msub><mover accent="true"><mi>β</mi><mo>^</mo></mover><mn>1</mn></msub></math> </ephtml> . An essentially equivalent method, which we call the "two‐stage" method, can further illustrate the connection between these methods and the existing methods for standard meta‐analysis.4,5 That is, rather than applying Equation (2.4) directly using the meta‐regression estimates <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0071" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true"><msub><mi>β</mi><mn>0</mn></msub><mo>^</mo></mover></math> </ephtml> , <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0072" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true"><msub><mi>β</mi><mn>1</mn></msub><mo>^</mo></mover></math> </ephtml> , and <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0073" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mover accent="true"><mi>τ</mi><mo>^</mo></mover><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup></math> </ephtml> , as in the one‐stage method, one could instead use only the meta‐regression estimate <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0074" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>β</mi><mo>^</mo></mover><mn>1</mn></msub></math> </ephtml> to shift the point estimates themselves to <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0075" xmlns="http://www.w3.org/1998/Math/MathML"><mi>Z</mi><mo>=</mo><mn>0</mn></math> </ephtml> , that is, <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0076" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>θ</mi><mo>^</mo></mover><mi>i</mi></msub><mfenced open="(" close=")"><mrow><mi>Z</mi><mo>=</mo><mn>0</mn></mrow></mfenced><mo>=</mo><msub><mover accent="true"><mi>θ</mi><mo>^</mo></mover><mi>i</mi></msub><mo>−</mo><msub><mi>z</mi><mi>i</mi></msub><msub><mover accent="true"><mi>β</mi><mo>^</mo></mover><mn>1</mn></msub></math> </ephtml> . Because <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0077" xmlns="http://www.w3.org/1998/Math/MathML"><mi>E</mi><mfenced open="[" close="]" separators="|"><mrow><msub><mover accent="true"><mi>θ</mi><mo>^</mo></mover><mi>i</mi></msub><mfenced open="(" close=")"><mrow><mi>Z</mi><mo>=</mo><mn>0</mn></mrow></mfenced><mspace width="0.5em" /></mrow><mrow><mspace width="0.5em" /><mi>Z</mi><mo>=</mo><mn>0</mn></mrow></mfenced><mo>=</mo><msub><mi>β</mi><mn>0</mn></msub></math> </ephtml> and <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0078" xmlns="http://www.w3.org/1998/Math/MathML"><mi>Var</mi><mfenced open="(" close=")" separators="|"><mrow><msub><mover accent="true"><mi>θ</mi><mo>^</mo></mover><mi>i</mi></msub><mfenced open="(" close=")"><mrow><mi>Z</mi><mo>=</mo><mn>0</mn></mrow></mfenced><mspace width="0.5em" /></mrow><mrow><mspace width="0.5em" /><mi>Z</mi><mo>=</mo><mn>0</mn></mrow></mfenced><mo>=</mo><msubsup><mi>τ</mi><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup><mo>+</mo><msubsup><mi>σ</mi><mi>i</mi><mn>2</mn></msubsup></math> </ephtml> , one could then apply Equation (2.4) by simply fitting a standard intercept‐only meta‐analysis (without covariates) to the <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0079" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>θ</mi><mo>^</mo></mover><mi>i</mi></msub><mfenced open="(" close=")"><mrow><mi>Z</mi><mo>=</mo><mn>0</mn></mrow></mfenced></math> </ephtml> and then using the pooled point estimate and estimated heterogeneity from this meta‐analysis, <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0080" xmlns="http://www.w3.org/1998/Math/MathML"><msup><mover accent="true"><mi>τ</mi><mo>^</mo></mover><mn>2</mn></msup></math> </ephtml> , to compute standard calibrated estimates5 (e.g., using the R package MetaUtility::calib_ests). These estimates could then be compared to the threshold that has also been shifted to <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0081" xmlns="http://www.w3.org/1998/Math/MathML"><mi>Z</mi><mo>=</mo><mn>0</mn></math> </ephtml> , that is, <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0082" xmlns="http://www.w3.org/1998/Math/MathML"><mi>q</mi><mo>−</mo><mi>z</mi><msub><mover accent="true"><mi>β</mi><mo>^</mo></mover><mn>1</mn></msub></math> </ephtml> . Although these methods are not exactly numerically equivalent,[<reflink idref="bib1" id="ref3">1</reflink>] simulation results (Section 5) indicated that they performed almost identically in practice. We use the one‐stage method for the applied example and code example below.</p> <p>As in standard meta‐analysis,4 inference for <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0083" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> or <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0084" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> can proceed via bootstrapping. Specifically, one can resample rows of the original sample, <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0085" xmlns="http://www.w3.org/1998/Math/MathML"><mfenced open="(" close=")" separators=",,"><msub><mover accent="true"><mi>θ</mi><mo>^</mo></mover><mi>i</mi></msub><msub><mover accent="true"><mi>σ</mi><mo>^</mo></mover><mi>i</mi></msub><msub><mi>z</mi><mi>i</mi></msub></mfenced></math> </ephtml> , fit a meta‐regression model to the resampled datasets to obtain new estimates <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0086" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>β</mi><mo>^</mo></mover><mn>0</mn></msub></math> </ephtml> , <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0087" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>β</mi><mo>^</mo></mover><mn>1</mn></msub></math> </ephtml> , and <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0088" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mover accent="true"><mi>τ</mi><mo>^</mo></mover><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup></math> </ephtml> , and finally estimate <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0089" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> via Equation (2.4). Note that meta‐regression estimation of <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0090" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true"><msub><mi>β</mi><mn>0</mn></msub><mo>^</mo></mover></math> </ephtml> , <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0091" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true"><msub><mi>β</mi><mn>1</mn></msub><mo>^</mo></mover></math> </ephtml> , and <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0092" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mover accent="true"><mi>τ</mi><mo>^</mo></mover><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup></math> </ephtml> must be included in the bootstrapping process to adequately capture the propagation of their sampling errors to the estimates of interest. If the point estimates are clustered, the cluster bootstrap should be used, as described in Section 2.1. A bias‐corrected and accelerated (BCa) confidence interval19,20 can then be constructed from the bootstrapped values of <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0093" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> or <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0094" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> . Informed by simulation results (Section 5), we provide the following four suggested guidelines for reporting <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0095" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> or <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0096" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> to help ensure that the metrics will provide accurate and interpretable results.</p> <hd id="AN0153408964-10">Guidelines for applying and reporting these metrics</hd> <p>We developed the guidelines below such that, based on an extensive simulation study including both realistic and extreme scenarios (Section 5), the metrics' performances conformed to the following thresholds: the bias was within <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0097" xmlns="http://www.w3.org/1998/Math/MathML"><mo>±</mo><mn>5</mn></math> </ephtml> percentage points in at least 90% of simulation scenarios, the coverage was no lower than 90% in at least 90% of simulation scenarios, and no more than 2% of simulation scenarios had coverage less than 85%. We required these criteria to hold regardless of the number of meta‐analyzed studies, subject to the fourth guideline below. Determining criteria for adequate estimator performance is inherently somewhat arbitrary; we consider these criteria to represent adequate performance based on the performance of existing standard estimators in meta‐regression (Section 5). Meta‐analysts who wish to apply the metrics according to more or less stringent criteria for the estimators' performance can browse comprehensive results for all simulation scenarios in a publicly available, documented dataset (https://osf.io/gs7fp/). The guidelines are:</p> <p></p> <ulist> <item> Include confidence intervals when reporting <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0098" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> or <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0099" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> . The confidence intervals help convey that while these metrics are overall unbiased, they may have considerable sampling variability in some settings and then could potentially be far from the truth for a given single analysis. For example, if <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0100" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> is large (e.g., 85%), but its confidence interval also includes small values (e.g., [15%, 100%]), or vice versa, investigators should comment on this when interpreting the metrics (as we demonstrate in the applied examples).</item> <p></p> <item> If point estimates are clustered (for example, within articles), investigate whether the population effects are also skewed by examining a density or cumulative distribution plot of the calibrated estimates (e.g., Section 4). The estimates <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0101" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> and <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0102" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> and/or their confidence intervals may not perform well when the population effects are both clustered and skewed. In such cases, consider eliminating clustering by averaging estimates and variances within clusters (Sutton et al. 1 Section 15.3) and using these average estimates to estimate <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0103" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> and <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0104" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> . [<reflink idref="bib2" id="ref4">2</reflink>]</item> <p></p> <item> When choosing a contrast to examine via <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0105" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> , avoid specifying extreme quantiles or rare values of the covariates (as described in Section 5.1.3). Choosing extremes can compromise the performance of <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0106" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> , though did not appear to compromise <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0107" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> or <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0108" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> themselves.</item> <p></p> <item> Apply the metric <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0109" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> only to meta‐regressions with at least 10 studies, and apply <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0110" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> only to meta‐regressions at least 20 studies. The metrics can perform poorly when there are fewer studies than this. In meta‐regressions of 10–20 studies, the metrics show adequate statistical performance as defined above but may have substantial sampling variability, such that the absolute error for any given sample may be large and the confidence intervals may accordingly be highly imprecise.</item> </ulist> <hd id="AN0153408964-11">ADDITIONAL CONSIDERATIONS</hd> <p></p> <hd id="AN0153408964-12">Types of meta‐regression covariates</hd> <p>At least two kinds of meta‐regression covariates may be of interest. First, some covariates may be scientifically interesting in their own right because they are hypothesized to be associated with a study's true population effect size; for example, the baseline clinical characteristics of a study's subjects might be associated with a treatment's effectiveness. Second, some covariates may not be of inherent scientific interest, but rather may be associated with a study's point estimate because they relate to the bias with which its true population effect is estimated; for example, observational studies might have typically larger or smaller estimates than randomized trials due to confounding. The proposed methods, like meta‐regression more broadly, aim only to characterize effect sizes conditional on study characteristics; they cannot isolate the causal effect that "changing" a study's characteristics would have on its population effect, the bias in its estimate, or both. As such, covariates falling into either or both categories can be handled identically in analysis. When considering specific biases in causal estimation, such as unmeasured confounding, the proposed methods could be combined with sensitivity analysis methods that do focus on causal estimation.4,24</p> <hd id="AN0153408964-13">Choices of effect‐size measures</hd> <p>There is a large literature on choosing effect‐size measures with which to conduct meta‐analyses; we summarize here only a few points that are not specific to the methods we have proposed here. First, analyses of binary outcomes can be conducted on either a multiplicative scale (e.g., risk ratios or odds ratios) or an additive scale (e.g., risk differences). For modeling purposes, the choice of an additive or multiplicative scale can be informed by scientific considerations regarding hypothesized mechanisms of the exposure or intervention, by statistical goodness of fit, and by parsimony.25 Multiplicative measures are also sometimes used in contexts when converting studies' estimates to a common, additive scale would invoke potentially unrealistic distributional assumptions, as was the case in the second applied example.13 Additive measures are often more relevant to assessing interventions' effects on public health, for example when estimating the "number needed to treat" based on a risk difference or when identifying which subgroup to treat based on an additive interaction measure.26 When assessing public health effectiveness in this sense, then regardless of the scale on which analyses are conducted, the effect measures should be converted to public health‐relevant measures before one considers whether effects are meaningfully strong.</p> <p>Second, analyses of continuous outcomes are often conducted on the standardized mean difference scale. This scale has limitations: for example, if two interventions produce the same absolute change in the same outcome measure, but are studied in different populations in which the variability on the outcome differs substantially, the interventions would produce different standardized mean differences.27,28 Some meta‐analysts argue against the use of <emph>SMD</emph>s27,28 or feel that the scale should never be used in any context (a point raised, for example, during the peer review of this paper). Whether and how <emph>SMD</emph>s should be used is a current debate in the field of meta‐analysis. Our views are as follows. When meta‐analyzing studies that measure the outcome on the same scale (e.g., blood pressure in terms of mmHg), it may often be preferable to use raw mean differences.28 However, in many scientific fields, studies do not measure outcomes on exactly the same scale, as in both applied examples provided here; in such cases, using standardized mean differences may enable some approximate comparison and synthesis of effect sizes across studies. Additionally, when outcomes use arbitrary or unitless raw measures (e.g., points on a Likert scale), expressing effect sizes using standardized mean differences may provide some sense of effect sizes relative to variability in that sample, similar to measures of genetic heritability. For some outcomes, such as income or grades in high school, absolute changes may in fact be less substantively meaningful than effects relative to variability in the population, for example expressed by <emph>SMD</emph>s with appropriately chosen denominators. Alternative metrics characterize effect sizes relative to a specified minimally important difference, rather than to sample variability,29 and we look forward to other metrics that might be proposed. Similar considerations and caveats apply when considering standardized versus absolute contrasts in continuous exposures.30</p> <p>If the meta‐analyst chooses to apply the metrics we propose with effect sizes on the <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0114" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">SMD</mi></math> </ephtml> scale, a few caveats should be kept in mind. Selecting a single threshold representing a meaningfully strong effect size across studies makes most sense when either the outcome has similar variability across studies or when, as described above, effects relative to the population are themselves of substantive interest. Second, population effects that exceed the chosen threshold may be those that arise in populations with very limited variability on the outcome measure rather than those in which the absolute effect size is very large.</p> <hd id="AN0153408964-14">APPLIED EXAMPLES</hd> <p>All data and R code required to reproduce the analyses and plots for both applied examples is publicly available and documented (https://osf.io/gs7fp/).</p> <hd id="AN0153408964-15">Sleeping targeted memory recall</hd> <p>In sleeping targeted memory recall (TMR), a specific sensory cue, such as an odor, is first paired with training stimuli during learning, and then the same sensory cue is presented again while the learner is sleeping. This is thought to aid natural processes of memory reactivation and consolidation during sleep. Hu et al.17 conducted a meta‐analysis investigating the effects of sleeping TMR on memory consolidation, the process by which recent learned experiences are crystallized into long‐term memory. They meta‐analyzed studies that measured memory improvements after a period of sleep during which sleeping TMR was either used or not used. They used subset analyses and meta‐regression to investigate various candidate covariates representing specific sleeping TMR methods and types of memory outcomes.</p> <p>For our re‐analysis, we focused on two candidate covariates: (<reflink idref="bib1" id="ref5">1</reflink>) the sleep stage during which the sensory cue was presented (dichotomized as slow‐wave sleep versus any other stage); and (<reflink idref="bib2" id="ref6">2</reflink>) the duration in hours that subjects were allowed to sleep between learning and testing. We considered effect sizes larger than <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0115" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">SMD</mi><mo>=</mo><mn>0.20</mn></math> </ephtml> to be meaningfully strong. We informed this choice of threshold by conventional criteria for a "small" effect size31 and by comparison to well‐evidenced interventions,3 namely conscious mnemonic methods that are known to improve memory consolidation. For example, mnemonics such as rehearsal and the method of loci produce average effect sizes of approximately <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0116" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">SMD</mi><mo>=</mo><mn>0.31</mn></math> </ephtml> compared to no training,32 and distributed practice produces effects of approximately <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0117" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">SMD</mi><mo>=</mo><mn>0.46</mn></math> </ephtml> compared to massed practice.33 Sleeping TMR is an unconscious and probably subtler method than these conscious mnemonics, so we selected an effect‐size threshold somewhat smaller than the <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0118" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">SMD</mi><mo>=</mo><mn>0.31</mn></math> </ephtml> to 0.46 seen for the latter. Of course, our choice of threshold is arbitrary. In practice, it is often reasonable to consider multiple thresholds or to present a plot of the estimated complementary cumulative distribution function of population effects conditional on <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0119" xmlns="http://www.w3.org/1998/Math/MathML"><mi>z</mi></math> </ephtml> (i.e., <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0120" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> as a function of the threshold <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0121" xmlns="http://www.w3.org/1998/Math/MathML"><mi>q</mi></math> </ephtml> ), as we illustrate below. A pointwise confidence interval could be constructed by bootstrapping selectively for the thresholds at which the value of <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0122" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> changes.</p> <p>We first robustly meta‐analyzed <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0123" xmlns="http://www.w3.org/1998/Math/MathML"><mi>k</mi><mo>=</mo><mn>208</mn></math> </ephtml> point estimates from 87 experiments,[<reflink idref="bib3" id="ref7">3</reflink>] using a working exchangeable correlation structure to model clustering of estimates within experiments,14,15 to estimate an overall average effect size on the standardized mean difference ( <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0124" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">SMD</mi></math> </ephtml> ) scale of 0.29 (95% CI: [0.19, 0.35]; <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0125" xmlns="http://www.w3.org/1998/Math/MathML"><mi>p</mi><mo><</mo><mn>0.0001</mn></math> </ephtml> ). We used existing methods for standard meta‐analysis4 with cluster‐bootstrapping to estimate that, overall, 49% (95% CI: [39%, 58%]) of effects were meaningfully strong by our chosen criterion.[<reflink idref="bib4" id="ref8">4</reflink>] To investigate effect‐measure modification using our proposed methods, we conducted a robust meta‐regression14,15 using the mean model <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0126" xmlns="http://www.w3.org/1998/Math/MathML"><mi>E</mi><mfenced open="[" close="]" separators="|"><mrow><mi>θ</mi><mspace width="0.5em" /></mrow><mrow><mspace width="0.5em" /><mi>Z</mi></mrow></mfenced><mo>=</mo><msub><mi>β</mi><mn>0</mn></msub><mo>+</mo><msub><mi>β</mi><mrow><mn>1</mn><mi>s</mi></mrow></msub><msub><mi>Z</mi><mi>s</mi></msub><mo>+</mo><msub><mi>β</mi><mrow><mn>1</mn><mi>d</mi></mrow></msub><msub><mi>Z</mi><mi>d</mi></msub></math> </ephtml> , where <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0127" xmlns="http://www.w3.org/1998/Math/MathML"><mi>Z</mi><mo>=</mo><mfenced open="(" close=")" separators=","><msub><mi>Z</mi><mi>s</mi></msub><msub><mi>Z</mi><mi>d</mi></msub></mfenced></math> </ephtml> , <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0128" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Z</mi><mi>s</mi></msub></math> </ephtml> indicated that the cue was presented during slow‐wave sleep versus any other sleep stage, and <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0129" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Z</mi><mi>d</mi></msub></math> </ephtml> was the duration of sleep in hours. We estimated the percentage of meaningfully strong effects when the cue was presented during slow‐wave sleep and subjects were allowed to sleep for 8 h (i.e., <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0130" xmlns="http://www.w3.org/1998/Math/MathML"><mi>Z</mi><mo>=</mo><mfenced open="(" close=")"><mrow><mn>1</mn><mo>,</mo><mn>8</mn></mrow></mfenced></math> </ephtml> ). We also estimated the percentage of meaningfully strong effects when the cue was presented during any other sleep stage and subjects were allowed to sleep for only 2 h (i.e., <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0131" xmlns="http://www.w3.org/1998/Math/MathML"><mi>Z</mi><mo>=</mo><mfenced open="(" close=")"><mrow><mn>0</mn><mo>,</mo><mn>2</mn></mrow></mfenced></math> </ephtml> ), and we estimated the difference between these two percentages.</p> <p>From the meta‐regression, the estimated intercept was <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0132" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true"><msub><mi>β</mi><mn>0</mn></msub><mo>^</mo></mover><mo>=</mo><mn>0.10</mn></math> </ephtml> (95% CI: [−0.16, 0.37]; <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0133" xmlns="http://www.w3.org/1998/Math/MathML"><mi>p</mi><mo>=</mo><mn>0.42</mn></math> </ephtml> ), the estimated effect of cue presentation during slow‐wave sleep was <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0134" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>β</mi><mo>^</mo></mover><mrow><mn>1</mn><mi>s</mi></mrow></msub><mo>=</mo><mn>0.13</mn></math> </ephtml> (95% CI: [−0.11, 0.36]; <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0135" xmlns="http://www.w3.org/1998/Math/MathML"><mi>p</mi><mo>=</mo><mn>0.27</mn></math> </ephtml> ), and the estimated effect of an additional hour of sleep was <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0136" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>β</mi><mo>^</mo></mover><mrow><mn>1</mn><mi>d</mi></mrow></msub><mo>=</mo><mn>0.003</mn></math> </ephtml> (95% CI: [−0.02, 0.03]; <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0137" xmlns="http://www.w3.org/1998/Math/MathML"><mi>p</mi><mo>=</mo><mn>0.80</mn></math> </ephtml> ). The estimated heterogeneity was <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0138" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mover accent="true"><mi>τ</mi><mo>^</mo></mover><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup><mo>=</mo><mn>0.07</mn></math> </ephtml> . We used these estimated regression coefficients and Equation (2.4) to calculate shifted, calibrated estimates (Figure 1) and to estimate that, for cue presentation during slow‐wave sleep and with 8 h of sleep, 53% of effects were meaningfully strong (95% CI: [34%, 72%]). Figure 2 shows the estimated complementary cumulative distribution function for such studies. In contrast, for cue presentation during any other sleep stage and with only 2 h of sleep, we estimated that 38% of effects were meaningfully strong (95% CI: [16%, 73%]). The estimated difference in the sleeping TMR effect, comparing these two joint levels of the covariates, was thus 15 percentage points (95% CI: [−24, 51]).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/BDCT/01nov21/jrsm1504-fig-0001.jpg?ephost1=dGJyMNHX8kSepq84v%2bvlOLCmsE6epq5Srqa4SK6WxWXS" alt="jrsm1504-fig-0001.jpg" title="1 For the applied example on memory consolidation, a smoothed density estimate for standard calibrated estimates5 that do not condition on covariates (black curve) and for calibrated estimates that have been shifted to covariate level Z=0 (orange curve), as in Equation 2.4. Solid red line: the shifted threshold q−zβ^1 for q=0.20 and for covariate level z=1,8. Dashed red line: the shifted threshold q−z0β^1 for reference covariate level z0=0,2 [Colour figure can be viewed at wileyonlinelibrary.com]" /> </p> <p></p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/BDCT/01nov21/jrsm1504-fig-0002.jpg?ephost1=dGJyMNHX8kSepq84v%2bvlOLCmsE6epq5Srqa4SK6WxWXS" alt="jrsm1504-fig-0002.jpg" title="2 For the applied example on memory consolidation, the estimated complementary cumulative distribution function of population effects in studies with z=1,8. The black dashed line represents the null. Shaded bands are 95% cluster‐bootstrapped pointwise confidence intervals" /> </p> <p></p> <p>Considering these results holistically, the point estimates from the meta‐regression suggested fairly small, "statistically nonsignificant" average increases in effect sizes associated with presentation during slow‐wave sleep ( <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0146" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>β</mi><mo>^</mo></mover><mrow><mn>1</mn><mi>s</mi></mrow></msub><mo>=</mo><mn>0.13</mn></math> </ephtml> (95% CI: [−0.11, 0.36]; <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0147" xmlns="http://www.w3.org/1998/Math/MathML"><mi>p</mi><mo>=</mo><mn>0.27</mn></math> </ephtml> ) and with an additional hour of sleep ( <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0148" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>β</mi><mo>^</mo></mover><mrow><mn>1</mn><mi>d</mi></mrow></msub><mo>=</mo><mn>0.003</mn></math> </ephtml> (95% CI: [−0.02, 0.03]; <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0149" xmlns="http://www.w3.org/1998/Math/MathML"><mi>p</mi><mo>=</mo><mn>0.80</mn></math> </ephtml> ). However, using our proposed metrics to compare joint levels of these two covariates simultaneously and to better characterize heterogeneity, rather than only average effect sizes, suggests that a majority of population effects were meaningfully strong when the cue was presented during slow‐wave sleep and with 8 h of sleep, with the confidence interval bounded above 34%. The proposed metrics also suggested that this percentage of meaningfully strong effects may have been somewhat larger than when the cue was presented during any other sleep stage and with only 2 h of sleep, though the confidence interval was wide.</p> <hd id="AN0153408964-18">Behavior interventions to reduce meat consumption</hd> <p>Mathur et al.13,34 conducted a meta‐analysis to assess the effectiveness of educational behavior interventions that attempt to reduce meat consumption by appealing to animal welfare. They meta‐analyzed 100 studies (from 34 articles) of such interventions, all of which measured behavioral or self‐reported outcomes related to meat consumption, purchase, or related intentions and had a control condition. Mathur et al.13 concluded that the interventions appeared to consistently reduce meat consumption, purchase, or related intentions at least in the short term with meaningfully large effects (meta‐analytic average risk ratio [ <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0150" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">RR</mi></math> </ephtml> ] = 1.22; 95% CI: [1.13, 1.33] with 71% of population effects estimated to be stronger than <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0151" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">RR</mi><mo>=</mo><mn>1.1</mn></math> </ephtml> ; 95% CI: [59%, 80%]) and additionally used meta‐regression to assess whether various characteristics regarding interventions' contents were associated with their effect sizes. As discussed in Section 3.2, these effect sizes should, when possible, also be considered on the risk difference scale before assessing substantive meaningfulness. Although not all studies reported sufficient information to calculate risk differences directly, most studies for which a risk ratio was extracted used a median split on the outcome in the control group, so the pooled <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0152" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">RR</mi><mo>=</mo><mn>1.22</mn></math> </ephtml> corresponds approximately to a risk difference <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0153" xmlns="http://www.w3.org/1998/Math/MathML"><mfenced open="[" close="]"><mi mathvariant="italic">RD</mi></mfenced><mo> </mo><mtext>of</mtext><mo> </mo><mn>0.11</mn></math> </ephtml> , and the threshold <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0154" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">RR</mi><mo>=</mo><mn>1.1</mn></math> </ephtml> corresponds approximately to <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0155" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">RD</mi><mo>=</mo><mn>0.05</mn></math> </ephtml> . Mathur et al. noted important methodological limitations in this field, including the predominant use of outcomes based on self‐reported food consumption or intentions, the potential for social desirability bias, and the potential for confounding in studies that were not randomized or that had differential dropout between intervention arms.</p> <p>For our re‐analysis, we estimated the percentage of meaningfully strong effects in interventions that contained graphic visual or verbal depictions of factory farms, a suspected effect‐measure modifier. This component has been controversial, as it could usefully invoke cognitive dissonance35 and harness deep‐seated connections between physical and moral disgust,36,37 or alternatively might backfire for some individuals.35 We excluded two studies for which the meta‐analysts had been unable to determine whether the intervention contained graphic content, leaving 98 analyzed studies, of which 61 (62%) had graphic interventions. We considered effect sizes larger than <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0156" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">RD</mi><mo>=</mo><mn>0.05</mn></math> </ephtml> (corresponding approximately to <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0157" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">RR</mi><mo>=</mo><mn>1.1</mn></math> </ephtml> in this meta‐analysis) to be meaningfully large based on the effect sizes of similar behavior interventions.[<reflink idref="bib5" id="ref9">5</reflink>]</p> <p>We first estimated the percentage of meaningfully strong effects among studies of graphic interventions while averaging over the distribution of all other study characteristics; that is, we fit a meta‐regression with an intercept term and a covariate indicating the use of graphic content. We again used cluster bootstrapping when estimating inference for the percentage metrics. In this first meta‐regression, we estimated that effect sizes in studies of graphic interventions were comparable to those in studies of non‐graphic interventions (effect‐measure modification <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0158" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">RR</mi><mo>=</mo><mn>0.96</mn></math> </ephtml> ; 95% CI: [0.79, 1.17]) and that 70% (95% CI: [49%, 86%]) of effects in studies of graphic interventions were stronger than <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0159" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">RR</mi><mo>=</mo><mn>1.1</mn></math> </ephtml> . Figure 3 shows the estimated complementary cumulative distribution function for studies of graphic interventions.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/BDCT/01nov21/jrsm1504-fig-0003.jpg?ephost1=dGJyMNHX8kSepq84v%2bvlOLCmsE6epq5Srqa4SK6WxWXS" alt="jrsm1504-fig-0003.jpg" title="3 For the applied example on meat consumption, the estimated complementary cumulative distribution function of population effects in studies whose interventions contained graphic content, regardless of risks of bias. The black dashed line represents the null. Shaded bands are 95% cluster‐bootstrapped pointwise confidence intervals" /> </p> <p></p> <p>Second, we estimated the percentage again upon more stringently considering studies that not only used graphic interventions, but that were also of relatively high methodological quality based on four risk‐of‐bias covariates. We speculated that these covariates might be associated with the bias in studies' estimates, though they could also be associated with measured or unmeasured effect‐measure modifiers. Specifically, we conditioned on studies' having received "low" risk‐of‐bias ratings13 (i.e., indicating higher methodological quality) with respect to the exchangeability of their intervention and control groups, their susceptibility to social desirability bias, and the external generalizability of their recruited subjects. We also conditioned on studies' use of direct behavioral measures of meat consumption (e.g., based on meal purchases at a university dining hall) or self‐reports (e.g., via food frequency questionnaires) rather than mere intentions. Although a number of studies (13–53%) fulfilled each of these risk‐of‐bias criteria individually, no single study of a graphic intervention fulfilled all four simultaneously, which would have precluded conducting a subset analysis. However, the proposed meta‐regression methods allowed us to estimate that, in hypothetical high‐quality studies of graphic interventions, the percentage of meaningfully strong effects would in fact increase to 97% (95% CI: [18%, 100%]). (See the Discussion for important considerations about assuming additive risks of bias.) This finding corroborates the meta‐analysts' observation that higher‐quality studies in fact tended to have somewhat larger effect sizes. However, it is also important to note that the confidence interval is quite wide given the near‐ceiling point estimate of 97% and is much wider than the confidence interval we obtained when conditioning only on graphic content (i.e., [49%, 86%]). This decrease in precision upon conditioning on risks of bias is itself informative: it quantitatively corroborates the meta‐analysts' impression that to more precisely and confidently characterize these interventions' effects will require that future studies prioritize methodological rigor over, perhaps, the introduction of new interventions.</p> <hd id="AN0153408964-20">SIMULATION STUDY</hd> <p></p> <hd id="AN0153408964-21">Simulation methods</hd> <p>We conducted an extensive simulation study to assess the performance of the proposed point estimation and inference methods for <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0160" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> and the difference <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0161" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> , including in scenarios with extreme values of the estimands, skewed population effects, and clustering.</p> <hd id="AN0153408964-22">Data generation</hd> <p>We considered a meta‐regression on two covariates, <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0162" xmlns="http://www.w3.org/1998/Math/MathML"><mi>Z</mi><mo>=</mo><mfenced open="(" close=")" separators=","><msub><mi>Z</mi><mi>c</mi></msub><msub><mi>Z</mi><mi>b</mi></msub></mfenced></math> </ephtml> , in which <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0163" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Z</mi><mi>c</mi></msub></math> </ephtml> was a standard normal variable and <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0164" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Z</mi><mi>b</mi></msub></math> </ephtml> was a binary variable with prevalence 50% that we generated independently of <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0165" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Z</mi><mi>c</mi></msub></math> </ephtml> . To generate data, we first assigned <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0166" xmlns="http://www.w3.org/1998/Math/MathML"><mi>k</mi></math> </ephtml> total studies ( <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0167" xmlns="http://www.w3.org/1998/Math/MathML"><mi>k</mi><mo>∈</mo><mfenced open="{" close="}"><mn>10,20,50,100,150</mn></mfenced></math> </ephtml> ) to <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0168" xmlns="http://www.w3.org/1998/Math/MathML"><mi>M</mi></math> </ephtml> clusters, where <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0169" xmlns="http://www.w3.org/1998/Math/MathML"><mi>M</mi></math> </ephtml> was either equal to <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0170" xmlns="http://www.w3.org/1998/Math/MathML"><mi>k</mi></math> </ephtml> (for no clustering) or equal to <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0171" xmlns="http://www.w3.org/1998/Math/MathML"><mi>k</mi><mo>/</mo><mn>2</mn></math> </ephtml> , such that each cluster contained two estimates.[<reflink idref="bib6" id="ref10">6</reflink>] We varied the residual heterogeneity[<reflink idref="bib7" id="ref11">7</reflink>] <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0172" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mi>τ</mi><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup><mspace width="0.25em" /><mo>∈</mo><mspace width="0.25em" /><mfenced open="{" close="}"><mn>0.0025,0.01,0.04,0.25,0.64</mn></mfenced></math> </ephtml> . For scenarios with clustering, we set the between‐cluster variance, <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0173" xmlns="http://www.w3.org/1998/Math/MathML"><mi>Var</mi><mfenced open="(" close=")"><mi>ζ</mi></mfenced></math> </ephtml> , to <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0174" xmlns="http://www.w3.org/1998/Math/MathML"><mn>0.75</mn><mo>×</mo><msubsup><mi>τ</mi><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup></math> </ephtml> , such that 75% of the residual heterogeneity was due to between‐cluster heterogeneity and 25% was due to within‐cluster heterogeneity. Within each cluster, studies' random intercepts were either normally or exponentially distributed, and the distribution of total sample sizes within each study was either <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0175" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>N</mi><mi>i</mi></msub><mo>∼</mo><mtext>Unif</mtext><mfenced><mn>50,150</mn></mfenced></math> </ephtml> or <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0176" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>N</mi><mi>i</mi></msub><mo>∼</mo><mtext>Unif</mtext><mfenced open="(" close=")"><mn>800,900</mn></mfenced></math> </ephtml> .</p> <p>We generated the population effect for the <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0177" xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>i</mi><mi mathvariant="italic">th</mi></msup></math> </ephtml> study in the <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0178" xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>m</mi><mi mathvariant="italic">th</mi></msup></math> </ephtml> cluster from the mean model:</p> <p> <ephtml> <math display="block" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0179" xmlns="http://www.w3.org/1998/Math/MathML">θmi=β0+β1cZc+β1bZb+ζm+γmi ζm~N0Varζ(cluster‐level random effects)γmi~N0τɛ2−Varζ or γmi~Expτɛ2−Varζ−1/2(study‐level random effects)</math> </ephtml> </p> <p>We fixed the intercept to <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0182" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>β</mi><mn>0</mn></msub><mo>=</mo><mn>0</mn></math> </ephtml> and covariate effect strengths to <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0183" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>β</mi><mrow><mn>1</mn><mi>c</mi></mrow></msub><mo>=</mo><mn>0.5</mn></math> </ephtml> and <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0184" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>β</mi><mrow><mn>1</mn><mi>b</mi></mrow></msub><mo>=</mo><mn>1</mn></math> </ephtml> . We held constant these effect sizes for all simulation scenarios but manipulated the parameters of interest, namely <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0185" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>P</mi><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> and the difference <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0186" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>P</mi><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mi>P</mi><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> , across a broad range by choosing the threshold <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0187" xmlns="http://www.w3.org/1998/Math/MathML"><mi>q</mi></math> </ephtml> appropriately. We considered three choices of covariate levels of interest, <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0188" xmlns="http://www.w3.org/1998/Math/MathML"><mfenced open="(" close=")" separators=","><msub><mi>Z</mi><mi>c</mi></msub><msub><mi>Z</mi><mi>b</mi></msub></mfenced></math> </ephtml> , which are described in Table 1. Within each cluster (suppressing the "m" subscript notation) and for each of <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0189" xmlns="http://www.w3.org/1998/Math/MathML"><mi>k</mi></math> </ephtml> meta‐analyzed studies, we generated a population effect, <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0190" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>θ</mi><mi>i</mi></msub></math> </ephtml> , on the raw mean difference scale from a normal distribution or a shifted exponential distribution. We chose each distribution's parameters to provide the desired intercept <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0191" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>β</mi><mn>0</mn></msub></math> </ephtml> and heterogeneity <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0192" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mi>τ</mi><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup></math> </ephtml> .</p> <p>1 TABLECovariate contrasts used in the simulation study. z0 and z list the value of the continuous covariate followed by the value of the binary covariate.</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Acronym for covariate contrast</th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0195" xmlns="http://www.w3.org/1998/Math/MathML"><msub xmlns=""><mi>z</mi><mn>0</mn></msub></math></p></th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0196" xmlns="http://www.w3.org/1998/Math/MathML"><mi xmlns="">z</mi></math></p></th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0197" xmlns="http://www.w3.org/1998/Math/MathML"><mi xmlns="">E</mi><mfenced open="[" close="]" separators="|" xmlns=""><mrow><mi>θ</mi><mspace width="0.25em" /></mrow><mrow><mspace width="0.25em" /><mi>Z</mi><mo>=</mo><msub><mi>z</mi><mn>0</mn></msub></mrow></mfenced></math></p></th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0198" xmlns="http://www.w3.org/1998/Math/MathML"><mi xmlns="">E</mi><mfenced open="[" close="]" separators="|" xmlns=""><mrow><mi>θ</mi><mspace width="0.25em" /></mrow><mrow><mspace width="0.25em" /><mi>Z</mi><mo>=</mo><mi>z</mi></mrow></mfenced><mspace width="0.25em" xmlns="" /></math></p></th><th align="left">Description</th></tr></thead><tbody valign="top"><tr><td>B ("binary")</td><td>(0, 0)</td><td>(0, 1)</td><td>0</td><td>1</td><td>Compares levels of the binary covariate while holding the continuous covariate constant at its mean.</td></tr><tr><td>BC ("binary and continuous")</td><td>(−0.5, 0)</td><td>(0.5, 1)</td><td>−0.25</td><td>1.25</td><td>Compares levels of the binary covariate along with a contrast in the continuous covariate from its 31st percentile to its 69th percentile.</td></tr><tr><td>BC‐rare ("binary and continuous involving a rare covariate value")</td><td>(2, 0)</td><td>(0.5, 1)</td><td>1</td><td>1.25</td><td>Compares levels of the binary covariate along with a contrast in the continuous covariate from its 98th percentile to its 69th percentile.</td></tr></tbody></table> </ephtml> </p> <p>We then simulated subject‐level data for a control group with mean 0 and for a treatment group with mean <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0199" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>θ</mi><mi>i</mi></msub></math> </ephtml> ; each group was of size <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0200" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>N</mi><mi>i</mi></msub><mo>/</mo><mn>2</mn></math> </ephtml> with a standard deviation of <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0201" xmlns="http://www.w3.org/1998/Math/MathML"><mn>1</mn></math> </ephtml> . Thus, the within‐study standard error of the estimated mean difference, <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0202" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>θ</mi><mo>^</mo></mover><mi>i</mi></msub></math> </ephtml> , was approximately <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0203" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>σ</mi><mo>^</mo></mover><mi>i</mi></msub><mo>=</mo><msqrt><mrow><mn>4</mn><mo>/</mo><msub><mi>N</mi><mi>i</mi></msub></mrow></msqrt></math> </ephtml> . For the meta‐regression, the proportion of the total residual variance attributable to residual effect heterogeneity,43,44 <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0204" xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>I</mi><mn>2</mn></msup></math> </ephtml> , was approximately <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0205" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mi>τ</mi><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup><mo>/</mo><mfenced open="(" close=")"><mrow><msubsup><mi>τ</mi><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup><mo>+</mo><mn>4</mn><mo>/</mo><mi>E</mi><mfenced open="[" close="]"><mi>N</mi></mfenced></mrow></mfenced></math> </ephtml> (Table 2). We chose values of <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0206" xmlns="http://www.w3.org/1998/Math/MathML"><mi>q</mi></math> </ephtml> to vary the first parameter of interest, <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0207" xmlns="http://www.w3.org/1998/Math/MathML"><mi>P</mi><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> , in {0.05, 0.10, 0.20, 0.50}. We then calculated the parameter <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0208" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>P</mi><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> and difference <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0209" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>P</mi><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mi>P</mi><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> based on <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0210" xmlns="http://www.w3.org/1998/Math/MathML"><mi>q</mi></math> </ephtml> , <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0211" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>β</mi><mrow><mn>1</mn><mi>c</mi></mrow></msub></math> </ephtml> , <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0212" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>β</mi><mrow><mn>1</mn><mi>b</mi></mrow></msub></math> </ephtml> , <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0213" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mi>τ</mi><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup></math> </ephtml> , and the appropriate distributional parameters. Table 3 summarizes the full‐factorial design of the simulation study. There were 2400 unique sets of parameters, each analyzed with all of the estimation methods described below.</p> <p>2 TABLEApproximate values of relative residual heterogeneity (I2) for each combination of simulation parameters pertaining to the mean within‐study sample size (EN) and residual heterogeneity (τɛ2)</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left" /><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0217" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup xmlns=""><mi>τ</mi><mi>ɛ</mi><mn>2</mn></msubsup><mo xmlns="">=</mo><mn xmlns="">0.0025</mn></math></p></th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0218" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup xmlns=""><mi>τ</mi><mi>ɛ</mi><mn>2</mn></msubsup><mo xmlns="">=</mo><mn xmlns="">0.01</mn></math></p></th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0219" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup xmlns=""><mi>τ</mi><mi>ɛ</mi><mn>2</mn></msubsup><mo xmlns="">=</mo><mn xmlns="">0.04</mn></math></p></th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0220" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup xmlns=""><mi>τ</mi><mi>ɛ</mi><mn>2</mn></msubsup><mo xmlns="">=</mo><mn xmlns="">0.25</mn></math></p></th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0221" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup xmlns=""><mi>τ</mi><mi>ɛ</mi><mn>2</mn></msubsup><mo xmlns="">=</mo><mn xmlns="">0.64</mn></math></p></th></tr></thead><tbody valign="top"><tr><td><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0222" xmlns="http://www.w3.org/1998/Math/MathML"><mi xmlns="">E</mi><mi xmlns="">N</mi><mo xmlns="">=</mo>100</math></p></td><td>0.06</td><td>0.20</td><td>0.50</td><td>0.86</td><td>0.94</td></tr><tr><td><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0223" xmlns="http://www.w3.org/1998/Math/MathML"><mi xmlns="">E</mi><mi xmlns="">N</mi><mo xmlns="">=</mo>850</math></p></td><td>0.25</td><td>0.68</td><td>0.89</td><td>0.98</td><td>0.99</td></tr></tbody></table> </ephtml> </p> <p>3 TABLEPossible values of data‐generation simulation parameters, manipulated in a full‐factorial design</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Covariate contrast</th><th align="left"><italic>k</italic></th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0224" xmlns="http://www.w3.org/1998/Math/MathML"><mi xmlns="">E</mi><mfenced open="[" close="]" xmlns=""><mi>N</mi></mfenced></math></p></th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0225" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup xmlns=""><mi>τ</mi><mi>ɛ</mi><mn>2</mn></msubsup><mspace width="0.25em" xmlns="" /></math></p></th><th align="left">Clustered</th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0226" xmlns="http://www.w3.org/1998/Math/MathML"><mi xmlns="">γ</mi></math></p> distribution</th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0227" xmlns="http://www.w3.org/1998/Math/MathML"><mi xmlns="">P</mi><mfenced open="(" close=")" xmlns=""><mi>z</mi></mfenced><mspace width="0.25em" xmlns="" /></math></p></th></tr></thead><tbody valign="top"><tr><td>B</td><td>10</td><td>100</td><td>0.0025</td><td>No: M = k</td><td>Normal</td><td>0.05</td></tr><tr><td>BC</td><td>20</td><td>850</td><td>0.01</td><td><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0228" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic" xmlns="">Var</mi><mi xmlns="">ζ</mi><mo xmlns="">=</mo>0.75<mo xmlns="">×</mo><mi xmlns="">τ</mi><mi xmlns="">ɛ</mi>2</math></p> Yes: M = k/2</td><td>Exponential</td><td>0.10</td></tr><tr><td>BC‐rare</td><td>50</td><td /><td>0.04</td><td /><td /><td>0.20</td></tr><tr><td /><td>100</td><td /><td>0.25</td><td /><td /><td>0.50</td></tr><tr><td /><td>150</td><td /><td>0.64</td><td /><td /><td /></tr></tbody></table> </ephtml> </p> <hd id="AN0153408964-23">Estimation procedures</hd> <p>We assessed both the one‐stage and the two‐stage methods described in Section 2.2. For both methods, for each scenario and iteration, we first estimated the meta‐regression parameters using a meta‐regression model with robust variance estimation as recommended in Section 2.2; for scenarios with clustering, we used a hierarchical working model.14 In this approach, weighted least squares is used to estimate the meta‐regression coefficients. The heterogeneity <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0229" xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>τ</mi><mn>2</mn></msup></math> </ephtml> is estimated using a method‐of‐moments estimator detailed elsewhere.14 We fit this model using the robumeta package in R.15</p> <p>Then, for the one‐stage method, we use the meta‐regression estimates to directly calculate shifted calibrated estimates as in Equation (2.3). For the two‐stage method, as described in Section 2.2, we shifted the point estimates themselves and then fit a standard intercept‐only meta‐analysis (without covariates) to the shifted estimates, thus obtaining calibrated estimates in the usual manner for a standard meta‐analysis.5 This second‐stage meta‐analysis used the Dersimonian‐Laird heterogeneity estimator,23 the usual choice for calculating calibrated estimates.5 We fit the second‐stage meta‐analysis using the R package metafor45 and obtained the calibrated estimates using MetaUtility.46</p> <p>To obtain inference, we bootstrapped the meta‐regression estimation process as well as, for the two‐stage method, the standard meta‐analysis estimation process. In scenarios with clustering, we used the cluster bootstrap and constructed bias‐corrected and accelerated confidence intervals19,20 as described in Section 2.1. We implemented bootstrapping using custom‐written code and the R package boot.47</p> <hd id="AN0153408964-24">High‐level simulation structure</hd> <p>We ran simulations representing all 2400 possible combinations of the varying data generation parameters (Table 3). For computational convenience, we generated separate datasets for the one‐stage and two‐stage methods, resulting in a total of 4800 "scenarios." We ran 500 simulation iterates per scenario[<reflink idref="bib8" id="ref12">8</reflink>] and used 1000 bootstrap iterates for all inference.</p> <hd id="AN0153408964-25">Exploration of bootstrap bias corrections</hd> <p>Second, we explored whether the bootstrap estimates could be used to correct any bias in the estimation of <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0230" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> and <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0231" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> ; we calculated the bias‐corrected versions of these estimates by subtracting from each original estimate the bootstrapped estimate of its bias (i.e., the difference between the mean of the bootstrapped sampling distribution and the estimate from the original sample itself; Davison & Hinkley,21 Section 2.1.2). Because the sampling distributions of <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0232" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> and <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0233" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> can be highly skewed when their estimands are close to 0 or 1, we speculated that applying a variance‐stabilizing transformation to the estimates might make the estimators more approximately pivotal and therefore improve the fidelity of the bootstrapped sampling distribution.48 Accordingly, we calculated bias‐corrected versions of the estimates with and without first taking the logit of the proportion estimates after truncating the proportions to <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0234" xmlns="http://www.w3.org/1998/Math/MathML"><mfenced open="[" close="]"><mn>0.001,0.999</mn></mfenced></math> </ephtml> .</p> <hd id="AN0153408964-26">Metrics of estimators' performance</hd> <p>For each scenario, we assessed the point estimators' performance and variability in terms of their mean bias and mean absolute error, defined as follows for a generic parameter <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0235" xmlns="http://www.w3.org/1998/Math/MathML"><mi>ω</mi></math> </ephtml> :</p> <p> <ephtml> <math display="block" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0236" xmlns="http://www.w3.org/1998/Math/MathML">Bias=1500∑r=1500ω^r−ωAbsolute error=1500∑r=1500∣ω^r−ω∣</math> </ephtml> </p> <p>where <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0237" xmlns="http://www.w3.org/1998/Math/MathML"><mi>r</mi></math> </ephtml> indexes simulation iterates. Second, to compare of our proposed estimators' performance to that of standard estimators from meta‐regression (i.e., coefficient estimates and heterogeneity estimates), we assessed relative bias, defined for a generic nonnegative parameter <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0238" xmlns="http://www.w3.org/1998/Math/MathML"><mi>ω</mi></math> </ephtml> as:</p> <p> <ephtml> <math display="block" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0239" xmlns="http://www.w3.org/1998/Math/MathML"><mtext>Relative bias</mtext><mo>=</mo><mfrac><mn>1</mn><mn>500</mn></mfrac><munderover><mo>∑</mo><mrow><mi>r</mi><mo>=</mo><mn>1</mn></mrow><mn>500</mn></munderover><mfrac><mrow><msub><mover><mi>ω</mi><mo>^</mo></mover><mi>r</mi></msub><mo>−</mo><mi>ω</mi></mrow><mi>ω</mi></mfrac></math> </ephtml> </p> <p>For each scenario, we assessed inference in terms of the mean coverage and mean width of 95% confidence intervals. When summarizing results across scenarios, we report medians because the metrics were often skewed across scenarios. Regarding inference, we also report the proportion of scenarios for which coverage was less than 85%. To help characterize variability in results across scenarios, we report <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0240" xmlns="http://www.w3.org/1998/Math/MathML"><msup><mn>10</mn><mi mathvariant="italic">th</mi></msup></math> </ephtml> and/or <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0241" xmlns="http://www.w3.org/1998/Math/MathML"><msup><mn>90</mn><mi mathvariant="italic">th</mi></msup></math> </ephtml> percentiles of the performance metrics across scenarios.[<reflink idref="bib9" id="ref13">9</reflink>] Throughout, we collapse over the results of the one‐stage and two‐stage method because they performed comparably.</p> <hd id="AN0153408964-27">RESULTS</hd> <p>Comprehensive results for all simulation scenarios, including additional performance metrics, are publicly available as a dataset (https://osf.io/gs7fp/). A small percentage of the 4800 total scenarios (2.5%) proved computationally infeasible to run because their data generation parameters consistently produced extreme datasets for which standard meta‐regression estimation failed. We analyzed the remaining 4679 scenarios, representing 2388 of the 2400 possible combinations of data‐generation parameters.</p> <hd id="AN0153408964-28">High‐level summary of all simulation results</hd> <p>In general, performance was better for <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0242" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> than for <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0243" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> . Performance was typically better in larger meta‐analyses (i.e., larger <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0244" xmlns="http://www.w3.org/1998/Math/MathML"><mi>k</mi></math> </ephtml> ), those with larger numbers of studies (i.e., larger <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0245" xmlns="http://www.w3.org/1998/Math/MathML"><mi>E</mi><mfenced open="[" close="]"><mi>N</mi></mfenced></math> </ephtml> ), and those with normal, independent population effects and was typically worse in meta‐analyses with other characteristics. Performance for <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0246" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> declined when the covariate contrast involved an extreme quantile of one of the covariates. (In the Supplementary material, we more specifically describe results of regressing these performance metrics on main effects of the aforementioned six meta‐analysis characteristics.)</p> <p>In the case of <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0247" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub></math> </ephtml> , coverage was conservative ( <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0248" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo><mn>95</mn><mo>%</mo></math> </ephtml> ) for small values of <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0249" xmlns="http://www.w3.org/1998/Math/MathML"><mi>k</mi></math> </ephtml> and declined somewhat as <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0250" xmlns="http://www.w3.org/1998/Math/MathML"><mi>k</mi></math> </ephtml> increased. Coverage was slightly below nominal for <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0251" xmlns="http://www.w3.org/1998/Math/MathML"><mi>k</mi><mo>=</mo><mn>150</mn></math> </ephtml> (93%; 10<sups><emph>th</emph></sups> percentile: 88%); additional diagnostics regarding the bootstrap samples preliminarily suggested that this might have reflected infidelity of the bootstrap sampling distribution to the actual sampling distribution (Section 5.2.2). Below, we speculate based on statistical theory on possible mechanisms for this finding, though precisely testing these proposed mechanisms was beyond the scope of the present simulation study. This could be investigated in future work.</p> <p>Figures 4 and 5 are violin plots (i.e., mirrored density plots) showing, for simulation scenarios fulfilling the guidelines given in Section 2.3, the distribution of each performance metric stratified by <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0253" xmlns="http://www.w3.org/1998/Math/MathML"><mi>k</mi></math> </ephtml> . The Supplementary material contains similar violin plots showing the distribution of the performance metrics across all scenarios, as well as stratified by additional characteristics of the meta‐analysis ( <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0254" xmlns="http://www.w3.org/1998/Math/MathML"><mi>k</mi></math> </ephtml> , <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0255" xmlns="http://www.w3.org/1998/Math/MathML"><mi>E</mi><mfenced open="[" close="]"><mi>N</mi></mfenced></math> </ephtml> , the presence of clustering, the population effect distribution, the covariate contrast, and the true <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0256" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>P</mi><mrow><mo>></mo><mi>q</mi></mrow></msub></math> </ephtml> ). The plots are provided to illustrate the extent of variability in performance metrics that might be expected across datasets (although some of the observed variability represents Monte‐Carlo error).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/BDCT/01nov21/jrsm1504-fig-0004.jpg?ephost1=dGJyMNHX8kSepq84v%2bvlOLCmsE6epq5Srqa4SK6WxWXS" alt="jrsm1504-fig-0004.jpg" title="4 For P^>qz, violin plots showing the number of meta‐analyzed studies (k) versus (a) bias, (b) absolute error, (c) 95% confidence interval coverage, and (d) 95% confidence interval width. White boxplots display the median, 25th percentile, and 75th percentile. Horizontal dashed reference lines represent perfect performance [Colour figure can be viewed at wileyonlinelibrary.com]" /> </p> <p></p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/BDCT/01nov21/jrsm1504-fig-0005.jpg?ephost1=dGJyMNHX8kSepq84v%2bvlOLCmsE6epq5Srqa4SK6WxWXS" alt="jrsm1504-fig-0005.jpg" title="5 For P^>qz−P^>qz0, violin plots showing the number of meta‐analyzed studies (k) versus (a) bias, (b) absolute error, (c) 95% confidence interval coverage, and (d) 95% confidence interval width. White boxplots display the median, 25th percentile, and 75th percentile. Horizontal dashed reference lines represent perfect performance [Colour figure can be viewed at wileyonlinelibrary.com]" /> </p> <p></p> <hd id="AN0153408964-31">Results for P^>qz</hd> <p> <bold>Point estimation</bold>. Table 4 shows results for <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0266" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> . Across all scenarios, <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0267" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> was approximately unbiased (bias = 0.00; 10<sups><emph>th</emph></sups> percentile: −0.02, 90<sups><emph>th</emph></sups> percentile: 0.08). However, because the estimator had substantial variability in some scenarios, it did have a non‐negligible absolute error of 0.07 (90<sups><emph>th</emph></sups> percentile: 0.23). The magnitude of the bias of <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0271" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> did not seem to differ systematically by characteristics of the simulation scenarios that would be observable in practice (i.e., not unobservable parameters).</p> <p>4 TABLEMain simulation results for P^>qz in all scenarios, in scenarios with normally distributed population effects, and in scenarios fulfilling the reporting guidelines. Values are medians across scenarios with 10 th and 90 th percentiles given in parentheses. No. scenarios: the number of scenarios analyzed. Coverage: coverage of 95% confidence intervals. Cov. <85%: proportion of scenarios with coverage <85%. CI width: width of 95% confidence intervals. Width >0.90: proportion of scenarios with very large average CI width (>0.90). These results are discussed in Section 5.2.2</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Scenarios</th><th align="left">No. scenarios</th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0279" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true" xmlns=""><msub><mi>P</mi><mi>z</mi></msub><mo>^</mo></mover></math></p> bias</th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0280" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true" xmlns=""><msub><mi>P</mi><mi>z</mi></msub><mo>^</mo></mover></math></p> abs. error</th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0281" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true" xmlns=""><msub><mi>P</mi><mi>z</mi></msub><mo>^</mo></mover></math></p> rel. bias</th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0282" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true" xmlns=""><mi>E</mi><mo>^</mo></mover><mfenced open="[" close="]" separators="|" xmlns=""><mrow><mi>θ</mi><mspace width="0.25em" /></mrow><mrow><mspace width="0.25em" /><mi>Z</mi><mo>=</mo><mi>z</mi></mrow></mfenced></math></p> rel. bias</th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0283" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup xmlns=""><mover accent="true"><mi>τ</mi><mo>^</mo></mover><mi>ɛ</mi><mn>2</mn></msubsup></math></p> rel. bias</th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0284" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true" xmlns=""><msub><mi>P</mi><mi>z</mi></msub><mo>^</mo></mover></math></p> coverage</th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0285" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true" xmlns=""><msub><mi>P</mi><mi>z</mi></msub><mo>^</mo></mover></math></p> cov. < 85%</th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0286" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true" xmlns=""><msub><mi>P</mi><mi>z</mi></msub><mo>^</mo></mover></math></p> CI width</th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0287" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true" xmlns=""><msub><mi>P</mi><mi>z</mi></msub><mo>^</mo></mover></math></p> width > 0.90</th></tr></thead><tbody valign="top"><tr><td>All</td><td>4679</td><td>0 (−0.02, 0.08)</td><td>0.07 (0.03, 0.23)</td><td>−0.01 (−0.18, 0.29)</td><td>0.05 (0.01, 0.17)</td><td>0.35 (0.14, 1.41)</td><td>0.95 (0.89, 1)</td><td>0.05</td><td>0.37 (0.12, 0.89)</td><td>0.10</td></tr><tr><td>Normal effects</td><td>2350</td><td>0 (−0.02, 0.02)</td><td>0.07 (0.03, 0.22)</td><td>−0.02 (−0.11, 0.20)</td><td>0.05 (0.01, 0.17)</td><td>0.32 (0.13, 1.41)</td><td>0.96 (0.91, 1)</td><td>0</td><td>0.37 (0.13, 0.90)</td><td>0.10</td></tr><tr><td>Not clustered exponential effects (i.e., scenarios fulfilling guidelines)</td><td>3522</td><td>0 (−0.02, 0.05)</td><td>0.07 (0.03, 0.22)</td><td>−0.02 (−0.13, 0.20)</td><td>0.05 (0.01, 0.17)</td><td>0.35 (0.14, 1.39)</td><td>0.96 (0.91, 1)</td><td>0.01</td><td>0.37 (0.12, 0.91)</td><td>0.11</td></tr></tbody></table> </ephtml> </p> <p> <bold>Exploration of bootstrap bias corrections</bold>. Bias‐correcting <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0288" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> via the bootstrap estimates did not improve and sometimes exacerbated bias, even when first taking the logit; we speculate this reflects the estimator's non‐pivotality,48 the non‐existence of an Edgeworth expansion for sample quantiles,21 and potentially a failure to reach the appropriate asymptotics regarding the bootstrap sampling distribution in these datasets of 10 to 150 observations. For these reasons, the slightly below‐nominal coverage seen at <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0289" xmlns="http://www.w3.org/1998/Math/MathML"><mi>k</mi><mo>=</mo><mn>150</mn></math> </ephtml> might reflect what is essentially bias in the bootstrap sampling distribution. Future work could consider using a double bootstrap to correct such bias.21</p> <p> <bold>Coverage</bold>. Across all scenarios, coverage was nominal (95%) on average, though was less than 85% in 5% of scenarios. The distribution of population effects and the presence of clustering appeared to affect coverage, with normal effects and unclustered effects producing the best coverage. Among scenarios with normally distributed population effects, none (0%) had coverage less than 85% (Table 4). When considering the scenarios fulfilling the reporting guidelines in Section 2.3 (i.e., scenarios whose population effects were not <emph>both</emph> exponentially distributed and clustered), 1% of scenarios had coverage less than 85%. In these recommended scenarios, the average confidence interval width was 0.37, and the confidence interval was highly imprecise (width <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0290" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo><mn>0.90</mn></math> </ephtml> ) in 11% of scenarios. Scenarios with <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0291" xmlns="http://www.w3.org/1998/Math/MathML"><mi>k</mi><mo>≤</mo><mn>20</mn></math> </ephtml> often had average confidence interval widths greater than 0.60, particularly when the meta‐analyzed studies had average sample sizes of 100 rather than 850. Specifically, with <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0292" xmlns="http://www.w3.org/1998/Math/MathML"><mi>k</mi><mo>≤</mo><mn>20</mn></math> </ephtml> , the average width was 0.74 ( <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0293" xmlns="http://www.w3.org/1998/Math/MathML"><msup><mn>90</mn><mi mathvariant="italic">th</mi></msup></math> </ephtml> percentile: 0.95), and the confidence intervals were highly imprecise in 24% of scenarios.</p> <p> <bold>Diagnostics regarding bootstrapped inference</bold>. The standard deviation of the bootstrap sampling distribution typically underestimated the empirical standard error of <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0294" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> , suggesting that when poor coverage occurred, it likely reflected, at least in part, infidelity of the bootstrap sampling distribution to the true sampling distribution. We also investigated whether characteristics of the bootstrap sampling distribution could be used as diagnostics for the confidence interval to perform poorly. However, characteristics such as the bootstrap distribution's skewness and the percentage of iterates for which the meta‐regression or <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0295" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> were not estimable were not associated with the performance of the confidence interval.</p> <hd id="AN0153408964-32">Results for P^>qz−P^>qz0</hd> <p> <bold>Point estimation</bold>. Table 5 shows results for <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0297" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> . Across all scenarios, <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0298" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> was approximately unbiased (bias = 0.00; 10<sups><emph>th</emph></sups> percentile: −0.02; 90<sups><emph>th</emph></sups> percentile: 0.08). Like <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0301" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub></math> </ephtml> , this estimator also had considerable sampling variability in some scenarios, resulting in an absolute error of 0.08 (90<sups><emph>th</emph></sups> percentile: 0.24). The performance of the point estimate <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0303" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> did appear somewhat related to observable characteristics of the simulation scenarios: as was the case for <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0304" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> , considering only scenarios with normal population effects somewhat improved the absolute error to 0.07 (vs. 0.08 in all scenarios). Additionally, avoiding very rare covariate values when defining the contrast of interest seemed to improve point estimation: excluding scenarios in which the <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0305" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>z</mi><mn>0</mn></msub></math> </ephtml> level involved the 98<sups><emph>th</emph></sups> quantile of the continuous covariate (termed "BC‐rare scenarios"; Table 1) improved the absolute error to 0.07. Including only scenarios fulfilling the reporting guidelines given in Section 2.3 (i.e., scenarios with <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0307" xmlns="http://www.w3.org/1998/Math/MathML"><mi>k</mi><mo>≥</mo><mn>20</mn></math> </ephtml> that did not use the BC‐rare contrast and did not have clustered exponential effects) slightly improved the absolute error to 0.06 (90<sups><emph>th</emph></sups> percentile: 0.17). As with <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0309" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> , bias‐correcting <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0310" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> via the bootstrap estimates was not effective.</p> <p>5 TABLEMain simulation results for P^>qz−P^>qz0 in all scenarios, in scenarios with normally distributed population effects, when excluding scenarios whose covariate contrast used an extreme quantile of the continuous covariate ("not BC‐rare"), and in scenarios fulfilling the reporting guidelines. Values are medians across scenarios with 10 th and 90 th percentiles given in parentheses. No. scenarios: the number of scenarios analyzed. Coverage: coverage of 95% confidence intervals. Cov. <85%: proportion of scenarios with coverage <85%. CI width: width of 95% confidence intervals. Width >0.90: proportion of scenarios with very large average CI width (>0.90). These results are discussed in Section 5.2.3</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Scenarios</th><th align="left">No. scenarios</th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0318" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true" xmlns=""><msub><mi>P</mi><mi>z</mi></msub><mo>^</mo></mover><mo xmlns="">−</mo><mover accent="true" xmlns=""><msub><mi>P</mi><msub><mi>z</mi><mn>0</mn></msub></msub><mo>^</mo></mover></math></p> bias</th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0319" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true" xmlns=""><msub><mi>P</mi><mi>z</mi></msub><mo>^</mo></mover><mo xmlns="">−</mo><mover accent="true" xmlns=""><msub><mi>P</mi><msub><mi>z</mi><mn>0</mn></msub></msub><mo>^</mo></mover><mspace width="0.25em" xmlns="" /></math></p> abs. error</th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0320" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true" xmlns=""><msub><mi>P</mi><mi>z</mi></msub><mo>^</mo></mover><mo xmlns="">−</mo><mover accent="true" xmlns=""><msub><mi>P</mi><msub><mi>z</mi><mn>0</mn></msub></msub><mo>^</mo></mover></math></p> rel. bias</th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0321" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true" xmlns=""><mi>E</mi><mo>^</mo></mover><mfenced open="[" close="]" separators="|" xmlns=""><mrow><mi>θ</mi><mspace width="0.25em" /></mrow><mrow><mspace width="0.25em" /><mi>Z</mi><mo>=</mo><mi>z</mi></mrow></mfenced></math></p> rel. bias</th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0322" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup xmlns=""><mover accent="true"><mi>τ</mi><mo>^</mo></mover><mi>ɛ</mi><mn>2</mn></msubsup></math></p> rel. bias</th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0323" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true" xmlns=""><msub><mi>P</mi><mi>z</mi></msub><mo>^</mo></mover><mo xmlns="">−</mo><mover accent="true" xmlns=""><msub><mi>P</mi><msub><mi>z</mi><mn>0</mn></msub></msub><mo>^</mo></mover></math></p> coverage</th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0324" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true" xmlns=""><msub><mi>P</mi><mi>z</mi></msub><mo>^</mo></mover><mo xmlns="">−</mo><mover accent="true" xmlns=""><msub><mi>P</mi><msub><mi>z</mi><mn>0</mn></msub></msub><mo>^</mo></mover></math></p> cov. < 85%</th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0325" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true" xmlns=""><msub><mi>P</mi><mi>z</mi></msub><mo>^</mo></mover><mo xmlns="">−</mo><mover accent="true" xmlns=""><msub><mi>P</mi><msub><mi>z</mi><mn>0</mn></msub></msub><mo>^</mo></mover></math></p> CI width</th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0287" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true" xmlns=""><msub><mi>P</mi><mi>z</mi></msub><mo>^</mo></mover></math></p> width > 0.90</th></tr></thead><tbody valign="top"><tr><td>All</td><td>4679</td><td>0 (−0.02, 0.08)</td><td>0.08 (0.03, 0.24)</td><td>−0.02 (−0.26, 0.33)</td><td>0.05 (0.01, 0.17)</td><td>0.35 (0.14, 1.41)</td><td>0.92 (0.67, 0.99)</td><td>0.24</td><td>0.39 (0.12, 0.90)</td><td>0.10</td></tr><tr><td>Normal effects</td><td>2350</td><td>0 (−0.03, 0.01)</td><td>0.07 (0.03, 0.23)</td><td>−0.03 (−0.18, 0.16)</td><td>0.05 (0.01, 0.17)</td><td>0.32 (0.13, 1.41)</td><td>0.92 (0.72, 0.99)</td><td>0.18</td><td>0.39 (0.13, 0.90)</td><td>0.10</td></tr><tr><td>Not BC‐rare</td><td>3098</td><td>0 (−0.02, 0.09)</td><td>0.07 (0.03, 0.23)</td><td>−0.01 (−0.17, 0.33)</td><td>0.05 (0.01, 0.17)</td><td>0.35 (0.14, 1.39)</td><td>0.93 (0.84, 0.99)</td><td>0.11</td><td>0.37 (0.12, 0.89)</td><td>0.09</td></tr><tr><td>Not BC‐rare nor clustered exponential effects; k ≥ 20 (i.e., scenarios fulfilling guidelines)</td><td>1908</td><td>0 (−0.02, 0.03)</td><td>0.06 (0.02, 0.17)</td><td>−0.02 (−0.14, 0.15)</td><td>0.04 (0.01, 0.13)</td><td>0.29 (0.13, 1.25)</td><td>0.94 (0.9, 0.99)</td><td>0.02</td><td>0.29 (0.11, 0.81)</td><td>0.05</td></tr></tbody></table> </ephtml> </p> <p> <bold>Coverage</bold>. Across all scenarios, coverage was somewhat below nominal (92%) on average and was less than 85% in 24% of scenarios. When restricting attention to scenarios fulfilling the guidelines, coverage was close to nominal (94%; <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0326" xmlns="http://www.w3.org/1998/Math/MathML"><msup><mn>10</mn><mi mathvariant="italic">th</mi></msup></math> </ephtml> percentile: 90%), and 2% of scenarios had coverage less than 85% (Table 5). As described above, this set of restrictions also produced the most improvement in the absolute error. In these scenarios, the average confidence interval width was 0.29, and the confidence interval was highly imprecise (width <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0327" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo><mn>0.90</mn></math> </ephtml> ) in 5% of scenarios. As with <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0328" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> , and characteristics of the bootstrap sampling distribution did not predict confidence interval coverage.</p> <hd id="AN0153408964-33">Exploratory investigations of other point estimation methods</hd> <p>As noted above, applying a bootstrap bias correction directly to <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0329" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> or <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0330" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> was not effective. Given the bias seen in <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0331" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mover accent="true"><mi>τ</mi><mo>^</mo></mover><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup></math> </ephtml> (Tables 4 and 5), we also experimented with using bootstrapping to bias‐correct the meta‐regression estimates <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0332" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mover accent="true"><mi>τ</mi><mo>^</mo></mover><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup></math> </ephtml> as well as <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0333" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>β</mi><mo>^</mo></mover><mn>0</mn></msub></math> </ephtml> , <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0334" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>β</mi><mo>^</mo></mover><mrow><mn>1</mn><mi>c</mi></mrow></msub></math> </ephtml> , and <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0335" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>β</mi><mo>^</mo></mover><mrow><mn>1</mn><mi>b</mi></mrow></msub></math> </ephtml><emph>before</emph> using them to calculate the calibrated estimates via Equation (2.4). (The bias in the latter three coefficient estimates was, however, typically small and was an order of magnitude smaller than that seen in <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0336" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mover accent="true"><mi>τ</mi><mo>^</mo></mover><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup></math> </ephtml> .) That is, we bootstrapped the meta‐regression model (again with 1000 iterates) and bias‐corrected each of these meta‐regression estimates by subtracting the bootstrapped estimate of its bias. We then calculated our proposed estimators using these bias‐corrected estimates. As a benchmark representing the best possible bias reduction in <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0337" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> and <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0338" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> that could be achieved if, hypothetically, all bias in the meta‐regression estimates were eliminated, we also calculated calibrated estimates using the actual parameters of the meta‐regression rather than estimates.</p> <p>We applied these approaches in the 32 "worst" scenarios from the main simulation, defined as the unique scenarios for which the relative bias of <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0339" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> , the relative bias of <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0340" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> , the coverage of the confidence interval for <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0341" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> , or the coverage of the confidence interval for <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0342" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> were among the 10 worst for that performance metric across all scenarios analyzed with the one‐stage method and in BC‐rare scenarios. We focused on BC‐rare scenarios because, as noted above, these scenarios seemed to particularly affect estimation of <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0343" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> ; we focused on the one‐stage method because it performed comparably to the two‐stage method. To avoid reversion to the mean49 that could arise from comparing results in the worst scenarios to those in the main simulations, we analyzed these scenarios via the usual one‐stage method again in the newly generated datasets (a replication of the main simulation results).</p> <p>Like the bootstrapped bias corrections applied directly to our proposed estimators in the main simulations, bias corrections on the meta‐regression estimates did not improve and sometimes exacerbated bias in the proposed estimators (Table 6). However, using the meta‐regression parameters rather than estimates improved bias by at least twofold. This suggests that for scenarios in which <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0344" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> and <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0345" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> performed poorly, a substantial portion, though not all, of their bias was attributable to bias in the standard meta‐regression estimates that propagated to <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0346" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> and <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0347" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> via the calibrated estimates. We return to this point in the Discussion. Because these methods did not improve bias, we did not assess their impact on inference.</p> <p>6 TABLESecondary simulation results for P^>q and P^>qz−P^>qz0 in the 32 worst BC‐rare scenarios when applying the basic one‐stage method used in the main simulations, when bias‐correcting the meta‐regression estimates, and when using the true meta‐regression parameters rather than estimates. Abs. bias: absolute error. Rel. bias: relative absolute error. These results are discussed in Section 5.3</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Method</th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0350" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true" xmlns=""><msub><mi>P</mi><mi>z</mi></msub><mo>^</mo></mover></math></p> bias</th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0351" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true" xmlns=""><msub><mi>P</mi><mi>z</mi></msub><mo>^</mo></mover></math></p> abs. error</th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0352" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true" xmlns=""><msub><mi>P</mi><mi>z</mi></msub><mo>^</mo></mover><mo xmlns="">−</mo><mover accent="true" xmlns=""><msub><mi>P</mi><msub><mi>z</mi><mn>0</mn></msub></msub><mo>^</mo></mover></math></p> bias</th><th align="left"><p><math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0353" xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true" xmlns=""><msub><mi>P</mi><mi>z</mi></msub><mo>^</mo></mover><mo xmlns="">−</mo><mover accent="true" xmlns=""><msub><mi>P</mi><msub><mi>z</mi><mn>0</mn></msub></msub><mo>^</mo></mover></math></p> abs. error</th></tr></thead><tbody valign="top"><tr><td>One‐stage</td><td>0.07</td><td>0.12</td><td>0.03</td><td>0.11</td></tr><tr><td>One‐stage with bias‐corrected meta‐regression estimates</td><td>0.08</td><td>0.13</td><td>0.04</td><td>0.11</td></tr><tr><td>One‐stage with meta‐regression parameters (benchmark)</td><td>0.01</td><td>0.06</td><td>0.00</td><td>0.04</td></tr></tbody></table> </ephtml> </p> <hd id="AN0153408964-34">DISCUSSION</hd> <p>We have proposed straightforward, easily interpreted statistical metrics for meta‐regression that characterize the percentage of meaningfully strong population effects for a specified level of the covariates and that further enable comparison of these percentages between two different levels of the covariates. As such, we believe the proposed metrics could usefully supplement standard reporting, which focuses only on average differences between levels of each covariate in turn. We provided simple R code to estimate the proposed metrics (Supplementary material or https://osf.io/gs7fp/).</p> <p>These methods have limitations, including those inherent to meta‐regression.50 For example, even when the meta‐analyzed studies are randomized, the covariates are not randomized across studies, so their effects on the outcome may be confounded.50,51 That is, it might not be the covariate itself that causally affects the study's primary exposure effect size, but rather another variable associated with it. The difference <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0354" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> therefore represents an average difference between studies with covariates <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0355" xmlns="http://www.w3.org/1998/Math/MathML"><mi>z</mi></math> </ephtml> versus those with <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0356" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>z</mi><mn>0</mn></msub></math> </ephtml> , <emph>not</emph> the causal effect of "changing" a study's covariates from <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0357" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>z</mi><mn>0</mn></msub></math> </ephtml> to <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0358" xmlns="http://www.w3.org/1998/Math/MathML"><mi>z</mi></math> </ephtml> . Similar considerations apply if using the proposed metrics to consider effect sizes in a hypothetical target population.52 Meta‐regression and our proposed metrics allow consideration of combinations of covariates that do not occur in the data, but one should not extrapolate unreasonably beyond the observed data. Our proposed methods could validly be used to estimate which levels of the covariates have the largest <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0359" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> ; but if doing so, <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0360" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> for the apparently "best" levels may then be biased upward due to statistical reversion to the mean,49 as is the case more generally with point estimation after conditioning on the size of the estimate. In relatively small meta‐regressions, the BCa bootstrap (particularly for <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0361" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> ) may fail to converge or may yield wide confidence intervals spanning most or all of the possible range <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0362" xmlns="http://www.w3.org/1998/Math/MathML"><mfenced open="[" close="]"><mrow><mn>0</mn><mo>,</mo><mn>1</mn></mrow></mfenced></math> </ephtml> , which should instill appropriate circumspection about what can be learned from a relatively small meta‐regression. Finally, as in meta‐analysis more broadly, limitations in the statistics reported in the meta‐analyzed papers often hampers extracting effect sizes on a scale that is comparable across studies and is substantively meaningful (Section 3.2); this problem would be mitigated if authors of original research were to publicly release deidentified datasets at the individual participant level.</p> <p>The simulation results suggested that <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0363" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> and <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0364" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> were approximately unbiased on average. However, it is important to note that due to the estimators' variability, they did show non‐negligible absolute errors of 0.07 and 0.06. When we specifically investigated the scenarios in which the estimators showed the most bias, almost all of the bias in <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0365" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced></math> </ephtml> and <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0366" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><mi>z</mi></mfenced><mo>−</mo><msub><mover accent="true"><mi>P</mi><mo>^</mo></mover><mrow><mo>></mo><mi>q</mi></mrow></msub><mfenced open="(" close=")"><msub><mi>z</mi><mn>0</mn></msub></mfenced></math> </ephtml> appeared to propagate to the estimators from the standard meta‐regression estimate <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0367" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mover accent="true"><mi>τ</mi><mo>^</mo></mover><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup></math> </ephtml> , which itself showed non‐negligible bias (e.g., relative bias = 0.35) and perhaps to a lesser extent from the meta‐regression coefficient estimates (Table 6). We used a moment‐type estimator that accommodates clustering;14,15 future work could investigate whether other heterogeneity estimators, such as the restricted maximum likelihood estimator, would perform better in this context.53 Heterogeneity estimation is indeed a longstanding challenge in meta‐analysis,16,53 and developing methods to allow robust and more efficient heterogeneity estimation is an active research area. As methods continue to improve, we expect that their use will also naturally propagate to the performance of the proposed estimators. When using current heterogeneity estimation methods,14 we provided practical guidelines informed by the simulation results regarding performance across numerous scenarios with varying characteristics.</p> <p>The simulation study itself was limited in scope for computational reasons: the 4800 scenarios we considered certainly do not represent the entire range of population effect distributions, sample sizes, and other characteristics that could potentially affect the estimators' performance. It would be useful to conduct more extensive simulation studies, for example assessing many more distributions of population effects, types of covariates (including clustered or correlated covariates), and choices of contrast.</p> <p>As we have noted, our proposed methods are semiparametric in that, when the meta‐regression is itself fit using semiparametric methods,14 the mean model must be correctly specified. If the mean model is misspecified, the meta‐regression estimates themselves may be biased,14 and this bias may also propagate to our proposed estimators. In particular, meta‐regression specifications usually include only main effects of the covariates; estimating interactions among the covariates with reasonable precision would typically be infeasible without quite large numbers of studies.26 By omitting potential interactions from the mean model, for example, any covariates representing risks of bias are assumed to operate additively on the effect size scale on which the meta‐analysis is conducted. Similarly to the applied example regarding dietary interventions, we might meta‐regress studies' log‐<emph>RR</emph> estimates on main effects of external generalizability and of susceptibility to social desirability bias, without including an interaction between the two covariates. This model assumes that the average biases produced by a lack of generalizability and by the susceptibility to social desirability bias are additive on the log‐<emph>RR</emph> scale and multiplicative on the <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0370" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">RR</mi></math> </ephtml> scale. However, in principle, the biases might in fact interact: subjects who were recruited because they were already highly motivated to change their behavior (representing poor generalizability) might be more likely to lie about their behavior to reduce cognitive dissonance or to conform to perceived pressure from the experimenters (representing a particularly pronounced social desirability bias effect), whereas a general sample of subjects who were not already motivated to change their behavior may report their behavior with less regard to perceived social desirability. Such effects would not be captured by a model with only main effects. However, with these caveats in mind, we hope that these methods will help investigators better understand how results from meta‐analyses might vary across study characteristics.</p> <hd id="AN0153408964-35">ACKNOWLEDGEMENTS</hd> <p>This research was supported by (<reflink idref="bib1" id="ref14">1</reflink>) the Pershing Square Fund for Research on the Foundations of Human Behavior; (<reflink idref="bib2" id="ref15">2</reflink>) NIH grant R01 CA222147; (<reflink idref="bib3" id="ref16">3</reflink>) the NIH‐funded Biostatistics, Epidemiology and Research Design (BERD) Shared Resource of Stanford University's Clinical and Translational Education and Research (UL1TR003142); (<reflink idref="bib4" id="ref17">4</reflink>) the Biostatistics Shared Resource (BSR) of the NIH‐funded Stanford Cancer Institute (P30CA124435); and (<reflink idref="bib5" id="ref18">5</reflink>) the Quantitative Sciences Unit through the Stanford Diabetes Research Center (P30DK116074).</p> <hd id="AN0153408964-36">CONFLICT OF INTEREST</hd> <p>The authors declare that they have no conflicts of interest.</p> <hd id="AN0153408964-37">DATA AVAILABILITY STATEMENT</hd> <p>All code and data required to reproduce the simulation study and applied examples are publicly available and documented, along with a simple code example (https://osf.io/gs7fp/).</p> <p>GRAPH: Supplementary information S1</p> <ref id="AN0153408964-38"> <title> Footnotes </title> <blist> <bibl id="bib1" idref="ref1" type="bt">1</bibl> <bibtext> In particular, in the two‐stage approach, <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0371" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mover accent="true"><mi>τ</mi><mo stretchy="true">^</mo></mover><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup></math> </ephtml> is estimated using a moment estimator that uses plug‐in estimates of the within‐cluster variances,[[14]] whereas in the second stage of the two‐stage method, it is estimated using the classical Dersimonian‐Laird moment estimator[[5], [23]] applied to the shifted point estimates. A simple simulated example available online (https://osf.io/gs7fp/) illustrates the methods' non‐equivalence in a small meta‐analysis with relatively low heterogeneity, in which the second stage of the two‐stage method estimates <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0372" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mover accent="true"><mi>τ</mi><mo stretchy="true">^</mo></mover><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup><mo>=</mo><mn>0</mn></math> </ephtml> , leading to calibrated estimates that are all exactly equal to the Dersimonian‐Laird pooled point estimate, whereas the one‐stage method estimates <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0373" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mover accent="true"><mi>τ</mi><mo stretchy="true">^</mo></mover><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup></math> </ephtml> slightly greater than 0, leading to calibrated estimates that differ slightly from one another.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref2" type="bt">2</bibl> <bibtext> This approach would be similar to scenarios in the simulation study with non‐clustered exponential effects, which are included in Table 4, row 3 and Table 5, row 4.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref7" type="bt">3</bibl> <bibtext> As in the original meta‐analysis,[17] we excluded four outlying point estimates from an original sample size of 212 estimates, leaving <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0374" xmlns="http://www.w3.org/1998/Math/MathML"><mi>k</mi><mo>=</mo><mn>208</mn></math> </ephtml> estimates used in analysis.</bibtext> </blist> <blist> <bibl id="bib4" idref="ref8" type="bt">4</bibl> <bibtext> It may seem counterintuitive that this percentage was less than 50% even though the estimated average effect size of <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0375" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">SMD</mi><mo>=</mo><mn>0.29</mn></math> </ephtml> exceeded <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0376" xmlns="http://www.w3.org/1998/Math/MathML"><mi>q</mi><mo>=</mo><mn>0.20</mn></math> </ephtml> . This reflects the effect sizes' skewness (Figure 1).</bibtext> </blist> <blist> <bibl id="bib5" idref="ref9" type="bt">5</bibl> <bibtext> As in the original meta‐analysis, we considered effect sizes larger than <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0377" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">RR</mi><mo>=</mo><mn>1.1</mn></math> </ephtml> to be meaningfully large based on the effect sizes of other health behavior interventions: for example, general nutritional "nudges" produce average effect sizes[38] of approximately <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0378" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">RR</mi><mo>=</mo><mn>1.15</mn></math> </ephtml> , and graphic warnings on cigarette boxes increase short‐term intentions to quit by approximately <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0379" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">RR</mi><mo>=</mo><mn>1.14</mn></math> </ephtml> upon conversion from the odds ratio scale.[[39], [41]]</bibtext> </blist> <blist> <bibl id="bib6" idref="ref10" type="bt">6</bibl> <bibtext> For realism, we informed this clustering structure by a corpus of 63 large meta‐analyses that were systematically sampled from journals representing a variety of disciplines.[42] In these meta‐analyses, each paper contributed a median of 1.5 studies per cluster (mean 2.7).</bibtext> </blist> <blist> <bibl id="bib7" idref="ref11" type="bt">7</bibl> <bibtext> For comparison, among the 28 meta‐analyses in the aforementioned corpus[42] with estimates on the standardized mean difference scale and for which <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0380" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mover accent="true"><mi>τ</mi><mo stretchy="true">^</mo></mover><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup></math> </ephtml> was statistically estimable,[[14]] 13 meta‐analyses (46%) had <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0381" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mover accent="true"><mi>τ</mi><mo stretchy="true">^</mo></mover><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup></math> </ephtml> >0 and hence would be candidates to apply our methods. Of these 13, the <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0382" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mover accent="true"><mi>τ</mi><mo stretchy="true">^</mo></mover><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup></math> </ephtml> estimates had a mean and median of 1.05 and 0.44 respectively, and 1 meta‐analysis had <ephtml> <math display="inline" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1504:jrsm1504-math-0383" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mover accent="true"><mi>τ</mi><mo stretchy="true">^</mo></mover><mi mathvariant="normal">ɛ</mi><mn>2</mn></msubsup><mo><</mo><mn>0.01</mn></math> </ephtml> . These estimates were from standard meta‐analysis rather than meta‐regression, so are best viewed as benchmarks for the amount of residual heterogeneity that might be expected if any meta‐regression covariates explain little of the heterogeneity.</bibtext> </blist> <blist> <bibl id="bib8" idref="ref12" type="bt">8</bibl> <bibtext> We used a relatively small number of iterates per scenario to enable computationally feasible assessment of the 4800 scenarios, which required 79 days of parallelized computational time on a high‐performance cluster. To provide a sense of the resulting Monte Carlo error, we would expect, for example, that 5% of simulation iterates would have estimated coverage percentages of less than 93% or more than 97% even if all confidence intervals in fact had exactly nominal coverage.</bibtext> </blist> <blist> <bibl id="bib9" idref="ref13" type="bt">9</bibl> <bibtext> We did not focus on minima and maxima across scenarios because, compared to 10<sups><emph>th</emph></sups> and 90<sups><emph>th</emph></sups> percentiles, minima and maxima across thousands of scenarios are highly unstable metrics, especially when there is some Monte Carlo error. Designing guidelines based on the very worst scenario, rather than based on less extreme percentiles as we did, would essentially overfit the characteristics of that single scenario, would be excessively dependent on the specific scenarios we chose to study, and would underestimate the metrics' actual worst‐case performance because of reversion to the mean.[49]</bibtext> </blist> <blist> <bibtext> Funding information National Institutes of Health, Grant/Award Numbers: CA222147, P30CA124435, P30DK116074, UL1TR003142; Pershing Square Foundation, Grant/Award Number: N/A</bibtext> </blist> </ref> <ref id="AN0153408964-39"> <title> REFERENCES </title> <blist> <bibtext> Sutton AJ, Abrams KR, Jones DR, Sheldon TA, Song F. Methods for Meta‐Analysis in Medical Research. Vol 348. New Jersey: Wiley Chichester ; 2000.</bibtext> </blist> <blist> <bibtext> Hedges LV, Saul JA, Cyr C, et al. Childhood obesity evidence base project: a rationale for taxonomic versus conventional meta‐analysis. Childhood Obes. 2020 ; 16 (S2): S2 ‐ S1.</bibtext> </blist> <blist> <bibtext> Mathur MB, Vander Weele TJ. New metrics for meta‐analyses of heterogeneous effects. Stat Med. 2019 ; 38 (8): 1336 ‐ 1342.</bibtext> </blist> <blist> <bibtext> Mathur MB, VanderWeele TJ. Robust metrics and sensitivity analyses for meta‐analyses of heterogeneous effects. Epidemiology. 2020 ; 31 (3): 356 ‐ 358.</bibtext> </blist> <blist> <bibtext> Wang C‐C, Lee W‐C. A simple method to estimate prediction intervals and predictive distributions: summarizing meta‐analyses beyond means and confidence intervals. Res Synth Methods. 2019 ; 10 (2): 255 ‐ 266.</bibtext> </blist> <blist> <bibtext> Ludema C, Cole SR, Poole C, Chu H, Eron JJ. Meta‐analysis of randomized trials on the association of prophylactic acyclovir and hiv‐1 viral load in individuals coinfected with herpes simplex virus‐2. AIDS (London, England). 2011 ; 25 (10): 1265.</bibtext> </blist> <blist> <bibtext> Mathur MB, VanderWeele TJ. Finding common ground in meta‐analysis wars on violent video games. Persp Psychol Sci. 2019 ; 14 (4): 705 ‐ 708.</bibtext> </blist> <blist> <bibtext> Lewis M, Mathur MB, VanderWeele TJ, Frank MC. The puzzling relationship between multi‐lab replications and meta‐analyses of the published literature. Preprint. https://psyarxiv.com/pbrdk/.</bibtext> </blist> <blist> <bibtext> Baumeister SE, Leitzmann MF, Linseisen J, Schlesinger S. Physical activity and the risk of liver cancer: a systematic review and meta‐analysis of prospective studies and a bias analysis. JNCI J Natl Cancer Inst. 2019 ; 111 (11): 1142 ‐ 1151.</bibtext> </blist> <blist> <bibtext> Noetel M, Griffith S, Delaney O, Sanders T, Parker P, del Pozo Cruz B, Lonsdale C. Are you better on YouTube? A systematic review of the effects of video on learning in higher education. Preprint. 2020. https://psyarxiv.com/kynez/</bibtext> </blist> <blist> <bibtext> Frederick DE, VanderWeele TJ. Longitudinal meta‐analysis of job crafting shows positive association with work engagement. Cogent Psychol. 2020 ; 7 (1): 1746733.</bibtext> </blist> <blist> <bibtext> VanderWeele TJ, Mathur MB, Chen Y. Media portrayals and public health i mplications for suicide and other behaviors. JAMA Psychiatry. 2019 ; 76 (9): 891 ‐ 892.</bibtext> </blist> <blist> <bibtext> Mathur MB, Peacock J, Reichling DB, et al. Interventions to reduce meat consumption by appealing to animal welfare: meta‐analysis and evidence‐based recommendations. Appetite. 2021 ;164: 105277.</bibtext> </blist> <blist> <bibtext> Hedges LV, Tipton E, Johnson MC. Robust variance estimation in meta‐regression with dependent effect size estimates. Res Synth Methods. 2010 ; 1 (1): 39 ‐ 65.</bibtext> </blist> <blist> <bibtext> Fisher Z, Tipton E. Robumeta: an R‐package for robust variance estimation in meta‐analysis. arXiv Preprint arXiv:1503.02220, 2015.</bibtext> </blist> <blist> <bibtext> Pustejovsky JE, Tipton E. Meta‐analysis with robust variance estimation: expanding the range of working models. Prev Sci. 2020.</bibtext> </blist> <blist> <bibtext> Hu X, Cheng LY, Chiu MH, Paller KA. Promoting memory consolidation during sleep: a meta‐analysis of targeted memory reactivation. Psychol Bull. 2020 ; 146 (3): 218.</bibtext> </blist> <blist> <bibtext> Louis TA. Estimating a population of parameter values using Bayes and empirical Bayes methods. J Am Stat Assoc. 1984 ; 79 (386): 393 ‐ 398.</bibtext> </blist> <blist> <bibtext> Efron B. Better bootstrap confidence intervals. J Am Stat Assoc. 1987 ; 82 (397): 171 ‐ 185.</bibtext> </blist> <blist> <bibtext> Carpenter J, Bithell J. Bootstrap confidence intervals: when, which, what? A practical guide for medical statisticians. Stat Med. 2000 ; 19 (9): 1141 ‐ 1164.</bibtext> </blist> <blist> <bibtext> Davison AC, Hinkley DV. Bootstrap Methods and their Application. Cambridge: Cambridge University Press ; 1997.</bibtext> </blist> <blist> <bibtext> Tipton E. Small sample adjustments for robust variance estimation with meta‐regression. Psychol Methods. 2015 ; 20 (3): 375.</bibtext> </blist> <blist> <bibtext> DerSimonian R, Laird N. Meta‐analysis in clinical trials. Control Clin Trials. 1986 ; 7 (3): 177 ‐ 188.</bibtext> </blist> <blist> <bibtext> Mathur MB, VanderWeele TJ. Sensitivity analysis for unmeasured confounding in meta‐analyses. J Am Stat Assoc. 2020 ; 115 (529): 163 ‐ 172.</bibtext> </blist> <blist> <bibtext> Spiegelman D, VanderWeele TJ. Evaluating public health interventions: 6. Modeling ratios or differences? Let the data tell us. Am J Public Health. 2017 ; 107 (7): 1087 ‐ 1091.</bibtext> </blist> <blist> <bibtext> VanderWeele T. Explanation in Causal Inference: Methods for Mediation and Interaction. Oxford: Oxford University Press ; 2015.</bibtext> </blist> <blist> <bibtext> Greenland S, Schlesselman JJ, Criqui MH. The fallacy of employing standardized regression coefficients and correlations as measures of effect. Am J Epidemiol. 1986 ; 123 (2): 203 ‐ 208.</bibtext> </blist> <blist> <bibtext> Cummings P. Arguments for and against standardized mean differences (effect sizes). Arch Pediatr Adolesc Med. 2011 ; 165 (7): 592 ‐ 596.</bibtext> </blist> <blist> <bibtext> Shrier I, Christensen R, Juhl C, Beyene J. Meta‐analysis on continuous outcomes in minimal important difference units: an application with appropriate variance calculations. J Clin Epidemiol. 2016 ; 80 : 57 ‐ 67.</bibtext> </blist> <blist> <bibtext> Maya B Mathur and Tyler J VanderWeele. A simple, interpretable conversion from Pearson's correlation to Cohen's d for continuous exposures. Epidemiology, 31 (2): e16 – e18, 2020.</bibtext> </blist> <blist> <bibtext> Cohen J. Statistical Power Analysis for the Behavioral Sciences. Cambridge, MA: Academic Press ; 2013.</bibtext> </blist> <blist> <bibtext> Gross AL, Parisi JM, Spira AP, et al. Memory training interventions for older adults: a meta‐analysis. Aging Mental Health. 2012 ; 16 (6): 722 ‐ 734.</bibtext> </blist> <blist> <bibtext> Donovan JJ, Radosevich DJ. A meta‐analytic review of the distribution of practice effect: now you see it, now you don't. J Appl Psychol. 1999 ; 84 (5): 795.</bibtext> </blist> <blist> <bibtext> Mathur MB, Robinson TN, Reichling DB, et al. Reducing meat consumption by appealing to animal welfare: protocol for a meta‐analysis and theoretical review. Syst Rev. 2020 ; 9 (1): 1 ‐ 8.</bibtext> </blist> <blist> <bibtext> Rothgerber H. Meat‐related cognitive dissonance: a conceptual framework for understanding how meat eaters reduce negative arousal from eating animals. Appetite. 2020 ; 146 : 104511. https://doi.org/10.1016/j.appet.2019.104511</bibtext> </blist> <blist> <bibtext> Chapman HA, Anderson AK. Things rank and gross in nature: a review and synthesis of moral disgust. Psychol Bull. 2013 ; 139 (2): 300. https://doi.org/10.1037/a0030964</bibtext> </blist> <blist> <bibtext> Feinberg M, Kovacheff C, Teper R, Inbar Y. Understanding the process of moralization: how eating meat becomes a moral issue. J Pers Soc Psychol. 2019 ; 117 : 50 ‐ 72. https://doi.org/10.1037/pspa0000149</bibtext> </blist> <blist> <bibtext> Arno A, Thomas S. The efficacy of nudge theory strategies in influencing adult dietary behaviour: a systematic review and meta‐analysis. BMC Public Health. 2016 ; 16 (1): 676. https://doi.org/10.1186/s12889-016-3272-x</bibtext> </blist> <blist> <bibtext> Brewer NT, Hall MG, Noar SM, et al. Effect of pictorial cigarette pack warnings on changes in smoking behavior: a randomized clinical trial. JAMA Int Med. 2016 ; 176 (7): 905 ‐ 912. https://doi.org/10.1001/jamainternmed.2016.2621</bibtext> </blist> <blist> <bibtext> VanderWeele TJ. On a square‐root transformation of the odds ratio for a common outcome. Epidemiology. 2017 ; 28 (6): e58 ‐ e60. https://doi.org/10.7326/0003-4819-154-10-201105170-00008</bibtext> </blist> <blist> <bibtext> VanderWeele TJ. Optimal approximate conversions of odds ratios and hazard ratios to risk ratios. Biometrics. 2019 ; 76 : 746 ‐ 752.</bibtext> </blist> <blist> <bibtext> Mathur MB, VanderWeele TJ. Estimating publication bias in meta‐analyses of peer‐reviewed studies: A meta‐meta‐analysis across disciplines and journal tiers. Preprint. 2020. https://osf.io/p3xyd/.</bibtext> </blist> <blist> <bibtext> Higgins PT, Thompson SG, Deeks JJ, Altman DG. Measuring inconsistency in meta‐analyses. BMJ. 2003 ; 327 (7414): 557 ‐ 560.</bibtext> </blist> <blist> <bibtext> Valentine JC, Pigott TD, Rothstein HR. How many studies do you need? A primer on statistical power for meta‐analysis. J Educ Behav Stat. 2010 ; 35 (2): 215 ‐ 247.</bibtext> </blist> <blist> <bibtext> Viechtbauer W. Conducting meta‐analyses in R with the metafor package. J Stat Softw. 2010 ; 36 (3): 1 ‐ 48.</bibtext> </blist> <blist> <bibtext> Mathur MB, Wang R, VanderWeele TJ. MetaUtility: utility functions for conducting and interpreting meta‐analyses. R Package Version 2.1.0, 2019.</bibtext> </blist> <blist> <bibtext> Canty A, Ripley B. Boot: Bootstrap Functions. R package version 1.3‐27, 2021.</bibtext> </blist> <blist> <bibtext> Hall P, Wilson SR. Two guidelines for bootstrap hypothesis testing. Biometrics. 1991 ;47(2): 757 ‐ 762.</bibtext> </blist> <blist> <bibtext> Samuels ML. Statistical reversion toward the mean: more universal than regression toward the mean. Am Stat. 1991 ; 45 (4): 344 ‐ 346.</bibtext> </blist> <blist> <bibtext> Thompson SG, Higgins JPT. How should meta‐regression analyses be undertaken and interpreted? Stat Med. 2002 ; 21 (11): 1559 ‐ 1573.</bibtext> </blist> <blist> <bibtext> VanderWeele TJ, Knol MJ. Interpretation of subgroup analyses in randomized trials: heterogeneity versus secondary interventions. Ann Internal Med. 2011 ; 154 (10): 680 ‐ 683.</bibtext> </blist> <blist> <bibtext> Steele RJ, Schnitzer ME, Shrier I. Importance of homogeneous effect modification for causal interpretation of meta‐analyses. Epidemiology. 2020 ; 31 (3): 353 ‐ 355.</bibtext> </blist> <blist> <bibtext> Veroniki AA, Jackson D, Viechtbauer W, et al. Methods to estimate the between‐study variance and its uncertainty in meta‐analysis. Res Synt Methods. 2016 ; 7 (1): 55 ‐ 79.</bibtext> </blist> </ref> <aug> <p>By Maya B. Mathur and Tyler J. VanderWeele</p> <p>Reported by Author; Author</p> </aug>
Header DbId: eric
DbLabel: ERIC
An: EJ1316105
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Meta-Regression Methods to Characterize Evidence Strength Using Meaningful-Effect Percentages Conditional on Study Characteristics
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Mathur%2C+Maya+B%2E%22">Mathur, Maya B.</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-6698-2607">0000-0001-6698-2607</externalLink>)<br /><searchLink fieldCode="AR" term="%22VanderWeele%2C+Tyler+J%2E%22">VanderWeele, Tyler J.</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-6112-0239">0000-0002-6112-0239</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Research+Synthesis+Methods%22"><i>Research Synthesis Methods</i></searchLink>. Nov 2021 12(6):731-749.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 19
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2021
– Name: SourceSuprt
  Label: Sponsoring Agency
  Group: SrcSuprt
  Data: National Institutes of Health (DHHS)
– Name: NumberContract
  Label: Contract Number
  Group: NumCntrct
  Data: CA222147<br />P30CA124435<br />P30DK116074<br />UL1TR003142
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Regression+%28Statistics%29%22">Regression (Statistics)</searchLink><br /><searchLink fieldCode="DE" term="%22Meta+Analysis%22">Meta Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Effect+Size%22">Effect Size</searchLink><br /><searchLink fieldCode="DE" term="%22Computation%22">Computation</searchLink><br /><searchLink fieldCode="DE" term="%22Statistical+Inference%22">Statistical Inference</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1002/jrsm.1504
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 1759-2879
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Meta-regression analyses usually focus on estimating and testing differences in average effect sizes between individual levels of each meta-regression covariate in turn. These metrics are useful but have limitations: they consider each covariate individually, rather than in combination, and they characterize only the mean of a potentially heterogeneous distribution of effects. We propose additional metrics that address both limitations. Given a chosen threshold representing a meaningfully strong effect size, these metrics address the questions: "For a given joint level of the covariates, what percentage of the population effects are meaningfully strong?" and "For any two joint levels of the covariates, what is the difference between these percentages of meaningfully strong effects?" We provide semiparametric methods for estimation and inference and assess their performance in a simulation study. We apply the proposed methods to meta-regression analyses on memory consolidation and on dietary behavior interventions, illustrating how the methods can provide more information than standard reporting alone. To facilitate implementing the methods in practice, we provide reporting guidelines and simple R code.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: Note
  Label: Notes
  Group: Note
  Data: https://osf.io/gs7fp
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2021
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1316105
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1316105
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1002/jrsm.1504
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 19
        StartPage: 731
    Subjects:
      – SubjectFull: Regression (Statistics)
        Type: general
      – SubjectFull: Meta Analysis
        Type: general
      – SubjectFull: Effect Size
        Type: general
      – SubjectFull: Computation
        Type: general
      – SubjectFull: Statistical Inference
        Type: general
    Titles:
      – TitleFull: Meta-Regression Methods to Characterize Evidence Strength Using Meaningful-Effect Percentages Conditional on Study Characteristics
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Mathur, Maya B.
      – PersonEntity:
          Name:
            NameFull: VanderWeele, Tyler J.
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 11
              Type: published
              Y: 2021
          Identifiers:
            – Type: issn-print
              Value: 1759-2879
          Numbering:
            – Type: volume
              Value: 12
            – Type: issue
              Value: 6
          Titles:
            – TitleFull: Research Synthesis Methods
              Type: main
ResultId 1