A Comparison of Three Popular Methods for Handling Missing Data: Complete-Case Analysis, Inverse Probability Weighting, and Multiple Imputation

Saved in:
Bibliographic Details
Title: A Comparison of Three Popular Methods for Handling Missing Data: Complete-Case Analysis, Inverse Probability Weighting, and Multiple Imputation
Language: English
Authors: Roderick J. Little (ORCID 0000-0001-9878-6977), James R. Carpenter, Katherine J. Lee
Source: Sociological Methods & Research. 2024 53(3):1105-1135.
Availability: SAGE Publications. 2455 Teller Road, Thousand Oaks, CA 91320. Tel: 800-818-7243; Tel: 805-499-9774; Fax: 800-583-2665; e-mail: journals@sagepub.com; Web site: https://sagepub.com
Peer Reviewed: Y
Page Count: 31
Publication Date: 2024
Document Type: Journal Articles
Reports - Research
Descriptors: Foreign Countries, Probability, Robustness (Statistics), Responses, Statistical Inference, Statistical Distributions, Evaluation Methods, Comparative Testing
Geographic Terms: United Kingdom (England), United Kingdom (Scotland), United Kingdom (Wales)
DOI: 10.1177/00491241221113873
ISSN: 0049-1241
1552-8294
Abstract: Missing data are a pervasive problem in data analysis. Three common methods for addressing the problem are (a) complete-case analysis, where only units that are complete on the variables in an analysis are included; (b) weighting, where the complete cases are weighted by the inverse of an estimate of the probability of being complete; and (c) multiple imputation (MI), where missing values of the variables in the analysis are imputed as draws from their predictive distribution under an implicit or explicit statistical model, the imputation process is repeated to create multiple filled-in data sets, and analysis is carried out using simple MI combining rules. This article provides a non-technical discussion of the strengths and weakness of these approaches, and when each of the methods might be adopted over the others. The methods are illustrated on data from the Youth Cohort (Time) Series (YCS) for England, Wales and Scotland, 1984-2002.
Abstractor: As Provided
Entry Date: 2024
Accession Number: EJ1434927
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwFIwtd8u1hQWnwR49YXefpVAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDDKzlzxBCVtoxag1ygIBEICBm7u34kr0VRXYfR8tZMPTcS9JBHiZKZgt1O-u-oCflvI3U1ACtmkmqgAUZgZCuNu8UKeQg1uzr6MKyFPbQHGpkUbMjfYLh8fha5woCgXlEf10ejGT_9XeNoxNFGN-D55Yuvwg_PZOyEiHsVxPVKv3CUaiPOlwz2x8NT7VPz4pm278aFmmvY7Qae20j2MQfGU6wn4y7N6kUGadaTwa
Text:
  Availability: 1
  Value: <anid>AN0178879686;som01aug.24;2024Aug09.05:11;v2.2.500</anid> <title id="AN0178879686-1">A Comparison of Three Popular Methods for Handling Missing Data: Complete-Case Analysis, Inverse Probability Weighting, and Multiple Imputation </title> <p>Missing data are a pervasive problem in data analysis. Three common methods for addressing the problem are (a) complete-case analysis, where only units that are complete on the variables in an analysis are included; (b) weighting, where the complete cases are weighted by the inverse of an estimate of the probability of being complete; and (c) multiple imputation (MI), where missing values of the variables in the analysis are imputed as draws from their predictive distribution under an implicit or explicit statistical model, the imputation process is repeated to create multiple filled-in data sets, and analysis is carried out using simple MI combining rules. This article provides a non-technical discussion of the strengths and weakness of these approaches, and when each of the methods might be adopted over the others. The methods are illustrated on data from the Youth Cohort (Time) Series (YCS) for England, Wales and Scotland, 1984–2002.</p> <p>Keywords: incomplete data; imputation; missing data; weighting</p> <hd id="AN0178879686-2">Preliminaries</hd> <p>Missing data are a pervasive problem in statistical analysis. The topic has an extensive literature – textbooks on the topic include [<reflink idref="bib16" id="ref1">16</reflink>], [<reflink idref="bib38" id="ref2">38</reflink>], [<reflink idref="bib22" id="ref3">22</reflink>], [<reflink idref="bib8" id="ref4">8</reflink>], and [<reflink idref="bib32" id="ref5">32</reflink>]. We consider and compare three common approaches to the analysis of data with missing values, namely complete-case analysis (henceforth CC), inverse probability weighting (henceforth IPW), and multiple imputation (henceforth MI). In CC — or complete-record analysis (e.g. [<reflink idref="bib8" id="ref6">8</reflink>], chapter 1) to avoid confusion with the terminology of cases and controls in medical studies — only units that are complete on the variables in an analysis are included; in IPW, the complete cases are weighted by the inverse of an estimate of the probability of being complete; and in MI, missing values of the variables in the analysis are imputed as draws from their predictive distribution under an implicit or explicit statistical model; the imputation process is repeated to create multiple filled-in data sets, and analysis is carried out using simple MI combining rules ([<reflink idref="bib28" id="ref7">28</reflink>]).</p> <p>All methods for handling missing data make unverifiable assumptions; perhaps the closest to an assumption-free method is that in [<reflink idref="bib10" id="ref8">10</reflink>], which presents bounds on parameter inferences based on best and worst-case values of the missing variables. This method is (as they acknowledge) very conservative, and is essentially limited to missing variables that have known finite support.</p> <p>Our focus here is principally on inference for regression coefficients and sample means. We restrict attention to models that assume the missingness mechanism is missing at random (MAR), as discussed in the next section, although it is important to recognize that missing not at random (MNAR) MI methods are possible as well. CC, IPW and MI are all quite general, in that (given sufficient information) they can be used to handle missing data in any statistical analysis an analyst might wish to perform on the data without missing values.</p> <p>Other valid approaches exist for handling missing data (by "valid" we mean that when its assumptions hold, a method yields consistent estimates of target parameters, confidence intervals with close to nominal coverage, and tests close to the stated size). Specifically, likelihoods can be defined for nonrectangular data sets with missing values, and hence methods based on these likelihoods can be implemented. In particular, maximum likelihood (ML) estimates can be computed, with standard errors based on the information matrix or sample re-use methods like the bootstrap; or a prior distribution can be added to the specification and inferences based on the Bayesian posterior distribution. Indeed, ML methods for missing data are quite widely used in the social sciences – often implicitly, as incorporated into structural equation modeling software like Mplus ([<reflink idref="bib19" id="ref9">19</reflink>]). ML is asymptotically equivalent to MI under the same model for the data, so it shares some of the properties of MI discussed here. However, MI is more flexible than ML in some settings, because it allows variables not included in the final analysis model to be included in the imputation model and readily extends to settings where data may be MNAR. Augmented inverse-probability weighted estimating equations ([<reflink idref="bib24" id="ref10">24</reflink>]; [<reflink idref="bib25" id="ref11">25</reflink>]) employ estimating equations that include model predictions of missing values and weighted residual terms, which provide some protection against model misspecification.</p> <p>We focus on the three methods described above because they are used extremely widely. In particular, CC is the default method in much statistical software, is intuitive and is simple to implement. IPW is the standard approach to handling unit nonresponse in surveys, and is also relatively simple to carry out. MI methods are more varied and complex, but increasingly common because of extensive availability in computer software packages. For example, MICE and other R packages ([<reflink idref="bib39" id="ref12">39</reflink>], [<reflink idref="bib36" id="ref13">36</reflink>]), IVEware ([<reflink idref="bib23" id="ref14">23</reflink>]), PROC MI in [<reflink idref="bib31" id="ref15">31</reflink>], and Stata (see https://<ulink href="http://www.stata.com/features/multiple-imputation/">www.stata.com/features/multiple-imputation/</ulink>).</p> <p>Additional modeling assumptions are unavoidable when analyzing data with missing values, so the most important step in dealing with missing data is to limit the extent of missing values, by careful design and data collection (e.g. [<reflink idref="bib20" id="ref16">20</reflink>]). Because some data are likely to be missing despite these efforts, it is important to try to collect covariates that are predictive of the missing values, so that an adequate adjustment can be made. In addition, the processes that lead to missing values should be assessed during the collection of data if possible (e.g. [<reflink idref="bib15" id="ref17">15</reflink>]), because this information plays a role in the choice of missing data adjustment method, as discussed further below.</p> <p>A basic assumption in all missing-data methods is that missingness of a particular value hides a true underlying value that is meaningful for analysis. Deciding whether a value is meaningful is not always as simple as it seems. For example, consider a longitudinal analysis of measures of quality of life; for subjects who leave the study because they move to a different location, it makes sense to consider quality of life as missing, whereas for subjects who die during the course of the study, it is not reasonable to consider quality of life after time of death as missing. Rather it is preferable to restrict the analysis of quality of life to individuals while they are alive. More complex missing data problems arise when individuals leave a study for unknown reasons, which may include relocation or death. Another example is nonresponse to opinion polls, where the target population consists of individuals who will vote – nonresponse for people who do not vote is arguably not missing data, since an imputed value is not meaningful for estimating the proportion of votes cast for each candidate.</p> <p>Despite the fact that CC, IPW and MI are common in practice, we believe that the principles underlying the choice between these methods are not as well understood as they might be. Therefore, this article provides a relatively nontechnical discussion of the strengths and weakness of CC, IPW and MI, and guidelines for when each of the methods might be favored over the others. For those who believe that the material is well known, here are four preliminary facts that may surprise some readers:</p> <p></p> <ulist> <item> We use the term <emph>auxiliary</emph> variables to mean fully-observed variables used for imputation or weighting but not included in the substantive model of interest. IPW based on auxiliary variables is widely viewed as reducing bias in CC estimates. However, in many realistic survey settings where the auxiliary variables are strongly related to the propensity to respond and weakly related to the survey variable of interest, IPW actually leads to worse inferences than CC (see the subsection "Inference for the Mean of an Incomplete Variable Y" for details).</item> <p></p> <item> In the statistics literature, CC is widely criticized and seen as inferior to methods like MI that use all available data. However, for some regression problems CC is optimal, and MI is actually less, not more, efficient (see the subsection "Missing data in Regression" and [<reflink idref="bib11" id="ref18">11</reflink>]).</item> <p></p> <item> CC is often described as biased unless the data are missing completely at random, as defined in the next section, and IPW is widely viewed as for reducing the bias of CC analysis. However, for some problems, CC analysis is actually less biased than IPW.</item> <p></p> <item> In some settings, a hybrid combination of CC and MI is less biased than CC, IPW or MI.</item> </ulist> <p>We expand upon points 2−4 in the subsection "Missing data in Regression". We illustrate the methods by analyzing data from a UK youth cohort study, and conclude by summarizing our recommendations concerning the methods.</p> <hd id="AN0178879686-3">Pattern and Mechanism of Missing Data</hd> <p>The <emph>pattern</emph> and <emph>mechanism</emph> of missing data are important features of the problem that play an important role in choosing between CC, IPW and MI. The <emph>pattern</emph> refers to which values in the data set are observed and which are missing. Specifically, let <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>Y</mi><mo>=</mo><mo stretchy="false">(</mo><msub><mi>y</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub><mo stretchy="false">)</mo></math> </ephtml> denote an <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><mi>n</mi><mo>×</mo><mi>p</mi><mo stretchy="false">)</mo></math> </ephtml> rectangular dataset without missing values, with <emph>i</emph>th row <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>y</mi><mi>i</mi></msub><mo>=</mo><mo stretchy="false">(</mo><msub><mi>y</mi><mrow><mi>i</mi><mn>1</mn></mrow></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>y</mi><mrow><mi>i</mi><mi>p</mi></mrow></msub><mo stretchy="false">)</mo></math> </ephtml> where <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>y</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub></math> </ephtml> is the value of variable <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mi>j</mi></msub></math> </ephtml> for subject <emph>i</emph>. With missing values, the pattern of missing data is defined by the <emph>response indicator matrix</emph><ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>R</mi><mo>=</mo><mo stretchy="false">(</mo><msub><mi>r</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub><mo stretchy="false">)</mo></math> </ephtml> , such that <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>r</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub><mo>=</mo><mn>1</mn></math> </ephtml> if <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>y</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub></math> </ephtml> is observed and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>r</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub><mo>=</mo><mn>0</mn></math> </ephtml> if <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>y</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub></math> </ephtml> is missing; equivalently, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mn>1</mn><mo>−</mo><msub><mi>r</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub></math> </ephtml> is the <emph>missing-data indicator</emph> for <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>y</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub></math> </ephtml> .</p> <p>Some methods for handling missing data apply to any pattern of missing data, whereas other methods assume a special pattern. A simple special pattern is <emph>univariate</emph> nonresponse, where missingness is confined to a single variable. Another example is <emph>monotone</emph> missing data, where the variables can be ordered so that <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mrow><mspace width=".1em" /><mi>j</mi><mo>+</mo><mn>1</mn></mrow></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>Y</mi><mi>p</mi></msub></math> </ephtml> are missing for all subjects where <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mi>j</mi></msub></math> </ephtml> is missing, for all <emph>j</emph> = 1,..., <emph>p</emph>-1. Looking at the matrix <emph>R,</emph> the result is a "staircase" pattern, where all variables and units to the left of the broken line forming the generally irregular staircase are observed, and all to the right are missing. This pattern arises in longitudinal data subject to attrition, where once a person drops out, no more data are observed for that person.</p> <p>The missingness <emph>mechanism</emph> addresses the reasons why values are missing, and whether these reasons relate to values in the data set. For example, subjects involved in a longitudinal intervention may be more likely to drop out of a study because they feel a treatment was ineffective, which might be related to a poor value of an outcome measure. [<reflink idref="bib27" id="ref19">27</reflink>] treated <emph>R</emph> as a random matrix, and characterized the missingness mechanism by the conditional distribution of <emph>R</emph> given <emph>Y</emph>, say <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo form="prefix" movablelimits="true">Pr</mo><mo stretchy="false">(</mo><mi>R</mi><mo fence="false" stretchy="false">|</mo><mi>Y</mi><mo>,</mo><mi>ϕ</mi><mo stretchy="false">)</mo></math> </ephtml> , where <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>ϕ</mi></math> </ephtml> denotes unknown parameters. When missingness does not depend on the values of the data <emph>Y</emph>, missing or observed, that is, <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><mo form="prefix" movablelimits="true">Pr</mo><mo stretchy="false">(</mo><mi>R</mi><mo fence="false" stretchy="false">|</mo><mi>Y</mi><mo>,</mo><mi>ϕ</mi><mo stretchy="false">)</mo><mo>=</mo><mo form="prefix" movablelimits="true">Pr</mo><mo stretchy="false">(</mo><mi>R</mi><mo fence="false" stretchy="false">|</mo><mi>ϕ</mi><mo stretchy="false">)</mo><mrow><mspace width="0.25em" /><mi mathvariant="normal">for</mi><mspace width=".1em" /><mi mathvariant="normal">all</mi><mspace width=".1em" /></mrow><mi>Y</mi><mo>,</mo><mi>ϕ</mi><mo>,</mo></math> </ephtml></p> <p>Graph</p> <p>the missingness is called missing completely at random (MCAR). An MCAR mechanism is plausible in some planned missing-data designs, but is a strong and often unrealistic assumption, especially when missing data do not occur by design, because missingness often does depend on values of variables.</p> <p>Using a slightly informal notation, let <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mrow><mo stretchy="false">(</mo><mn>1</mn><mo stretchy="false">)</mo></mrow></msub></math> </ephtml> denote the observed components of <emph>Y</emph> and let <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mrow><mo stretchy="false">(</mo><mn>0</mn><mo stretchy="false">)</mo></mrow></msub></math> </ephtml> denote the missing components of <emph>Y</emph>. A less restrictive assumption is that missingness depends only on values <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mrow><mo stretchy="false">(</mo><mn>1</mn><mo stretchy="false">)</mo></mrow></msub></math> </ephtml> that are observed, and given these not on values <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mrow><mo stretchy="false">(</mo><mn>0</mn><mo stretchy="false">)</mo></mrow></msub></math> </ephtml> that are missing. That is: <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><mo form="prefix" movablelimits="true">Pr</mo><mo stretchy="false">(</mo><mi>R</mi><mo fence="false" stretchy="false">|</mo><msub><mi>Y</mi><mrow><mo stretchy="false">(</mo><mn>1</mn><mo stretchy="false">)</mo></mrow></msub><mo>,</mo><msub><mi>Y</mi><mrow><mo stretchy="false">(</mo><mn>0</mn><mo stretchy="false">)</mo></mrow></msub><mo>,</mo><mi>ϕ</mi><mo stretchy="false">)</mo><mo>=</mo><mo form="prefix" movablelimits="true">Pr</mo><mo stretchy="false">(</mo><mi>R</mi><mo fence="false" stretchy="false">|</mo><msub><mi>Y</mi><mrow><mo stretchy="false">(</mo><mn>1</mn><mo stretchy="false">)</mo></mrow></msub><mo>,</mo><mi>ϕ</mi><mo stretchy="false">)</mo><mrow><mspace width="0.25em" /><mi mathvariant="normal">for</mi><mspace width=".1em" /><mi mathvariant="normal">all</mi><mspace width=".1em" /></mrow><msub><mi>Y</mi><mrow><mo stretchy="false">(</mo><mn>0</mn><mo stretchy="false">)</mo></mrow></msub><mo>,</mo><mi>ϕ</mi><mo>.</mo></math> </ephtml></p> <p>Graph</p> <p>The missing data are then called missing at random (MAR) at the observed values of <emph>R</emph> and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mrow><mo stretchy="false">(</mo><mn>1</mn><mo stretchy="false">)</mo></mrow></msub></math> </ephtml> . If (<reflink idref="bib1" id="ref20">1</reflink>) does not hold, the data are missing not at random (MNAR). Concerning our three compared methods, CC is always valid under MCAR, and in particular circumstances to be described, it may be valid under weaker assumptions about the missingness mechanism. The implementations of IPW and MI in widely-available software are valid under MAR (and hence under MCAR), and we restrict attention to these versions here; it is possible to develop versions of IPW and MI for MNAR mechanisms, but these lie outside the scope of this article.</p> <hd id="AN0178879686-4">Methods</hd> <p></p> <hd id="AN0178879686-5">Complete-Case (CC) Analysis</hd> <p>CC for a set of variables simply discards units where any of these variables are missing. It has the advantage of simplicity, and it is the default analysis in most statistical software packages. It has two main drawbacks. Firstly, the complete cases are not a random subsample of the original sample unless the data are MCAR. This is usually an unrealistic assumption, because cases with missing values often differ from complete cases in terms of the variables of interest. If the complete cases are not a random subsample, CC will give biased answers for simple summary measures (such as mean, sd) and may yield biased answers for regression models, although not in all situations, as discussed below. Secondly, CC discards information in the incomplete cases, which has typically cost non-trivial resources to collect, and which will often contain information for reducing bias or increasing the efficiency of CC estimates. A key question is thus how much information is contained in the incomplete cases – if nearly all the information is contained in the complete cases, CC might be a reasonable approach. Unfortunately, the answer to this question is not straightforward, because it depends on the fraction of complete cases, the distribution of variables (observed and missing) in the incomplete cases, the missingness mechanism, and the nature of the specific analysis of interest. Some specific examples are provided below.</p> <hd id="AN0178879686-6">Inverse Probability Weighting (IPW)</hd> <p>A modification of CC, commonly used to handle unit nonresponse in surveys, is inverse probability weighting (IPW), which weights complete units by the inverse of an estimate of the probability of response (see e.g. [<reflink idref="bib34" id="ref21">34</reflink>]). In particular, when estimating a population mean, the sample mean is replaced by the weighted mean. IPW can also be applied to estimators other than means, such as regression coefficients, or more generally, estimators for generalized estimating equations (weighted GEE).</p> <p>A simple approach for creating weights is to form adjustment cells (also called subclasses) based on background variables measured for both respondents and nonrespondents; for unit nonresponse adjustment, these are often based on geographical areas or groupings of similar areas based on aggregate socioeconomic data. All nonrespondents are assigned zero weight and the nonresponse weight for all respondents in an adjustment cell is then the inverse of the estimated response rate in that cell. For more details see [<reflink idref="bib16" id="ref22">16</reflink>], Example 3.6).</p> <p>With more extensive background information, a generalization of adjustment cell weighting is <emph>response propensity</emph> stratification, where (a) the indicator for unit nonresponse is regressed on the background variables, using the combined data for respondents and nonrespondents, using a method such as logistic regression appropriate for a binary outcome; (b) a predicted response probability is computed for each respondent based on the regression in (a); and (c) adjustment cells are formed based on a categorized version of the predicted response probability. The creation of adjustment cells can be useful to reduce extreme weights, which otherwise can inflate the variance of weighted estimates. Theory ([<reflink idref="bib26" id="ref23">26</reflink>]) suggests that this is an effective method for removing nonresponse bias attributable to the background variables when unit nonresponse is MAR. Adjustment cell weighting is a special case of this method when the adjustment cell variables are indicators of the cells. In both methods, weights are rescaled so that they sum to the number of respondents.</p> <p>Although IPW can be useful for reducing nonresponse bias, it does have serious limitations. First, information in the incomplete cases is only used to determine the weights (i.e. the weight model uses variables that are fully observed on both respondents and non-respondents), and partially observed cases are still discarded in the weighted analysis. This fact means that the method is generally inefficient when, as will often be the case, there is substantial information in these partially observed cases. Therefore, weighted estimates can have unacceptably high variance, especially when extreme values of a variable are given large weights. For modifications of the inverse probability weights to increase efficiency, see for example [<reflink idref="bib4" id="ref24">4</reflink>].</p> <p>Variance estimation for weighted estimates will ideally take into account uncertainty in the estimated weights, otherwise standard errors will be overestimated so inferences will be conservative. Approaches include Taylor series expansion ([<reflink idref="bib25" id="ref25">25</reflink>]) or computing bootstrap standard errors, with weights recalculated for each bootstrap sample ([<reflink idref="bib16" id="ref26">16</reflink>], Chapter 5).</p> <hd id="AN0178879686-7">Multiple Imputation (MI)</hd> <p>Methods that impute or fill in the missing values have the advantage that, unlike CC or IPW, observed values in the incomplete cases are retained to make full use of them in the analysis. In fact, the goal of MI is to preserve the information in the observed values for inference, not to get the best predictions of the missing values.</p> <p>Because we can never recover the actual missing value, a single imputation for each missing value cannot reflect the imputation uncertainty, and as a result standard errors of estimates based on analysis of a single filled-in data tend to be underestimated. Large-sample results show that for simple situations with 30% of the information missing, single imputation under the correct model results in nominal 90% confidence intervals having actual coverages below 80% ([<reflink idref="bib30" id="ref27">30</reflink>]). The inaccuracy of nominal levels is even more extreme in multiparameter testing problems ([<reflink idref="bib28" id="ref28">28</reflink>], Chapter 4).</p> <p>Multiple (as opposed to single) imputation fixes this problem ([<reflink idref="bib28" id="ref29">28</reflink>], [<reflink idref="bib29" id="ref30">29</reflink>], [<reflink idref="bib33" id="ref31">33</reflink>]). The basic steps of MI are to (a) estimate a predictive distribution for the missing values <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mrow><mo stretchy="false">(</mo><mn>0</mn><mo stretchy="false">)</mo></mrow></msub></math> </ephtml> given the observed values <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mrow><mo stretchy="false">(</mo><mn>1</mn><mo stretchy="false">)</mo></mrow></msub></math> </ephtml> in the data set— approaches to this are described below; (b) fill in, or impute, the missing values with draws from this predictive distribution (note that the imputations are draws, that is random selections from the predictive distribution, not means of the predictive distribution); (c) repeat step (b) <emph>M</emph> > 1 times (where, say, <emph>M</emph> = 10 or 20) to create <emph>M</emph> datasets, each containing different sets of draws of the missing values. For any particular analysis of the filled-in data, we then apply the standard complete-data analysis to each of the <emph>M</emph> datasets, yielding <emph>M</emph> estimates, say <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msup><mrow><mover><mi>θ</mi><mo stretchy="false">^</mo></mover></mrow><mrow><mo stretchy="false">(</mo><mn>1</mn><mo stretchy="false">)</mo></mrow></msup><mo>,</mo><mo>...</mo><mo>,</mo><msup><mrow><mover><mi>θ</mi><mo stretchy="false">^</mo></mover></mrow><mrow><mo stretchy="false">(</mo><mi>M</mi><mo stretchy="false">)</mo></mrow></msup><mo stretchy="false">)</mo></math> </ephtml> of parameters <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>θ</mi></math> </ephtml> ; (d) combine the parameter estimates to create an overall estimate of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>θ</mi></math> </ephtml> — a method for doing this is called a MI combining rule. In particular for scalar estimands, the MI estimate is the average <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mrow><mover><mi>θ</mi><mo stretchy="false">^</mo></mover></mrow><mrow><mrow><mi mathvariant="normal">MI</mi></mrow></mrow></msub><mo>=</mo><msubsup><mo movablelimits="false">∑</mo><mrow><mi>m</mi><mo>=</mo><mn>1</mn></mrow><mi>M</mi></msubsup><mrow><msup><mrow><mrow><mover><mi>θ</mi><mo stretchy="false">^</mo></mover></mrow></mrow><mrow><mo stretchy="false">(</mo><mi>m</mi><mo stretchy="false">)</mo></mrow></msup></mrow><mo>/</mo><mi>M</mi></math> </ephtml> of the estimates from the <emph>M</emph> datasets, and the sampling variance of the estimate is estimated as <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mover><mi>V</mi><mo stretchy="false">^</mo></mover></mrow><mo>=</mo><mrow><mover><mi>W</mi><mo stretchy="false">^</mo></mover></mrow><mo>+</mo><mo stretchy="false">(</mo><mn>1</mn><mo>+</mo><mn>1</mn><mo>/</mo><mi>M</mi><mo stretchy="false">)</mo><mrow><mover><mi>B</mi><mo stretchy="false">^</mo></mover></mrow></math> </ephtml> , where <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mover><mi>W</mi><mo stretchy="false">^</mo></mover></mrow></math> </ephtml> is the average of the estimated sampling variances from the <emph>M</emph> datasets, and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mover><mi>B</mi><mo stretchy="false">^</mo></mover></mrow></math> </ephtml> is the sample variance of the estimates across the <emph>M</emph> datasets; the factor 1 + 1/<emph>M</emph> is a small-<emph>M</emph> correction. The quantity <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><mn>1</mn><mo>+</mo><mn>1</mn><mo>/</mo><mi>M</mi><mo stretchy="false">)</mo><mrow><mover><mi>B</mi><mo stretchy="false">^</mo></mover></mrow></math> </ephtml> is crucial, because it estimates the increase in the variance from imputation uncertainty, which is omitted (i.e., set to zero) by single imputation methods. Other combining rules provide refinements of this basic method, and include combining rules for test statistics and p-values. See, for example, [<reflink idref="bib16" id="ref32">16</reflink>], Section 10.2).</p> <p>The imputation of draws from the predictive distribution creates the variability in the estimates over the MI data sets, allowing the appropriate assessment of imputation uncertainty. Imputing draws is inefficient, but the fact that <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mrow><mover><mi>θ</mi><mo stretchy="false">^</mo></mover></mrow><mrow><mrow><mi mathvariant="normal">MI</mi></mrow></mrow></msub></math> </ephtml> is averaged over datasets reduces this inefficiency, roughly by a factor of <emph>M.</emph> In fact, MI under a well-specified model is essentially fully efficient from a statistical perspective, providing <emph>M</emph> is sufficiently large. The appropriate choice of <emph>M</emph> depends on the fraction of missing information, which is estimated for each parameter <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>θ</mi></math> </ephtml> by <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><mn>1</mn><mo>+</mo><mn>1</mn><mo>/</mo><mi>M</mi><mo stretchy="false">)</mo><mrow><mover><mi>B</mi><mo stretchy="false">^</mo></mover></mrow><mo>/</mo><mrow><mover><mi>V</mi><mo stretchy="false">^</mo></mover></mrow></math> </ephtml> . Larger fractions of missing information require larger values of <emph>M</emph> to yield good estimates of the imputation uncertainty.</p> <p>Once the data are imputed, the remaining steps of MI are not much more difficult than doing a single imputation. The additional computing from repeating an analysis <emph>M</emph> times is not a major burden and MI combining rules for standard errors are standard in MI programs. Modern MI programs yield imputed data sets that lead to proper inferences, in the sense that they appropriately incorporate uncertainty in the parameter estimates in the imputation models; they also can be applied to a general missing data pattern. The imputation models generally assume the missing data are MAR, although MNAR mechanisms can also be incorporated (e.g. [<reflink idref="bib37" id="ref33">37</reflink>], [<reflink idref="bib9" id="ref34">9</reflink>]). Most of the work is in generating good predictive distributions for the missing values.</p> <p>There are three primary approaches to creating the predictive distributions for multiple imputation of the missing data. (<reflink idref="bib1" id="ref35">1</reflink>) <emph>Joint modeling</emph>, where predictive distributions are derived from an explicit parametric joint model <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>f</mi><mo stretchy="false">(</mo><mi>Y</mi><mo fence="false" stretchy="false">|</mo><mi>θ</mi><mo stretchy="false">)</mo></math> </ephtml> for the variables in the data set, indexed by parameters <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>θ</mi></math> </ephtml> . Examples of models include the multivariate normal model for continuous variables, loglinear models for categorical variables, and the general location model for mixture of continuous and categorical variables (see [<reflink idref="bib16" id="ref36">16</reflink>], Chapters 11–14). (<reflink idref="bib2" id="ref37">2</reflink>) <emph>Sequential regression imputation</emph> ([<reflink idref="bib23" id="ref38">23</reflink>]), also called <emph>chained-equation</emph> imputation ([<reflink idref="bib42" id="ref39">42</reflink>], [<reflink idref="bib39" id="ref40">39</reflink>]), or <emph>full conditional specification</emph>, where a model is specified for the conditional distribution <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>f</mi><mi>j</mi></msub><mo stretchy="false">(</mo><msub><mi>Y</mi><mi>j</mi></msub><mo fence="false" stretchy="false">|</mo><msub><mi>Y</mi><mrow><mo stretchy="false">(</mo><mi>j</mi><mo stretchy="false">)</mo></mrow></msub><mo stretchy="false">)</mo></math> </ephtml> of each variable <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mi>j</mi></msub></math> </ephtml> with missing values, given the other variables <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mrow><mo stretchy="false">(</mo><mi>j</mi><mo stretchy="false">)</mo></mrow></msub><mo>=</mo><mo stretchy="false">(</mo><msub><mi>Y</mi><mn>1</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>Y</mi><mrow><mspace width=".1em" /><mi>j</mi><mo>−</mo><mn>1</mn></mrow></msub><mo>,</mo><msub><mi>Y</mi><mrow><mspace width=".1em" /><mi>j</mi><mo>+</mo><mn>1</mn></mrow></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>Y</mi><mi>p</mi></msub><mo stretchy="false">)</mo></math> </ephtml> , for <emph>j</emph> = 1,...,<emph>p</emph>. These methods are iterative, and impute the missing values of each variable as draws from their conditional distribution, given the observed or most recently imputed values of the other variables. (<reflink idref="bib3" id="ref41">3</reflink>) <emph>Hot deck</emph> imputation, which matches each incomplete case (which we call the recipient) to a complete case (which we call the donor) based on some closeness metric. The values of missing variables for the recipient are then imputed with the corresponding values of those variables for the donor. A variety of metrics are used, but a common and principled choice is the distance between the predicted means from a regression of the missing variables on the observed variables (predictive mean matching, see [<reflink idref="bib13" id="ref42">13</reflink>]). Hot deck methods were originally defined for single imputation – for a review of these methods, see [<reflink idref="bib1" id="ref43">1</reflink>]; they can be extended to MI by defining a set of close donors for each recipient, and randomly picking a donor for each MI data set ([<reflink idref="bib13" id="ref44">13</reflink>]).</p> <p>Joint modeling is well-motivated theoretically – the underlying theory is Bayesian, which creates imputations that take into account uncertainty in model parameters. The approach is well suited to a monotone pattern, where the joint distribution of <emph>Y</emph> can be factored as <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>f</mi><mo stretchy="false">(</mo><msub><mi>Y</mi><mn>1</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>Y</mi><mi>p</mi></msub><mo stretchy="false">)</mo><mo>=</mo><msub><mi>f</mi><mn>1</mn></msub><mo stretchy="false">(</mo><msub><mi>Y</mi><mn>1</mn></msub><mo stretchy="false">)</mo><msub><mi>f</mi><mn>2</mn></msub><mo stretchy="false">(</mo><msub><mi>Y</mi><mn>2</mn></msub><mo fence="false" stretchy="false">|</mo><msub><mi>Y</mi><mn>1</mn></msub><mo stretchy="false">)</mo><mo>...</mo><mi>f</mi><mo stretchy="false">(</mo><msub><mi>Y</mi><mi>p</mi></msub><mo fence="false" stretchy="false">|</mo><msub><mi>Y</mi><mn>1</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>Y</mi><mrow><mspace width=".1em" /><mi>p</mi><mo>−</mo><mn>1</mn></mrow></msub><mo stretchy="false">)</mo></math> </ephtml> , with variables arranged from most observed ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>1</mn></msub><mo stretchy="false">)</mo></math> </ephtml> to least observed ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mi>p</mi></msub><mo stretchy="false">)</mo></math> </ephtml> . The distributions in this product can then be modeled using regressions appropriate for the outcome variable type – for example, normal linear regression for continuous outcomes, logistic regression for binary outcomes, and so on. These regressions are quite flexible, in that they can include polynomial terms and interactions as covariates. Imputations are created sequentially, first filling in missing values of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>1</mn></msub></math> </ephtml> as draws from <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>f</mi><mn>1</mn></msub><mo stretchy="false">(</mo><msub><mi>Y</mi><mn>1</mn></msub><mo stretchy="false">)</mo></math> </ephtml> , then filling in missing values of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>2</mn></msub></math> </ephtml> as draws from <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>f</mi><mn>2</mn></msub><mo stretchy="false">(</mo><msub><mi>Y</mi><mn>2</mn></msub><mo fence="false" stretchy="false">|</mo><msub><mi>Y</mi><mn>1</mn></msub><mo stretchy="false">)</mo></math> </ephtml> , conditioning on observed and previously imputed values of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>1</mn></msub></math> </ephtml> , and so on. The data analyst creating the MIs needs to provide appropriate specifications of these regressions, using subject-matter knowledge and regression diagnostic tools applied to the set of cases that are observed on the relevant set of variables.</p> <p>For non-monotone patterns, imputation algorithms are iterative and involve an application of Markov Chain Monte Carlo methods. This means that methods are needed to monitor convergence of the chain, and the methods can be computationally intensive if the data matrix is large. A challenge for the joint modeling approach is the limited availability of models for joint distribution of <emph>Y</emph>. For example, the popular multivariate normal model implies that the normal regression models for the imputations that are linear and additive in the covariates, with a constant residual variance. This limitation can be eased by strategies such as transformation of the variables, or more generally by using a latent normal model for binary and unordered categorical data, a flexible approach which has recently been shown to perform well ([<reflink idref="bib21" id="ref45">21</reflink>]).</p> <p>The chained equation approach sidesteps this limitation of joint modeling for non-monotone patterns by not requiring that the set of conditional distributions <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo fence="false" stretchy="false">{</mo><msub><mi>f</mi><mi>j</mi></msub><mo stretchy="false">(</mo><msub><mi>Y</mi><mi>j</mi></msub><mo fence="false" stretchy="false">|</mo><msub><mi>Y</mi><mrow><mo stretchy="false">(</mo><mi>j</mi><mo stretchy="false">)</mo></mrow></msub><mo stretchy="false">)</mo><mo fence="false" stretchy="false">}</mo></math> </ephtml> for each <emph>j</emph> corresponds to a coherent joint distribution for <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>Y</mi><mn>1</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>Y</mi><mi>p</mi></msub><mo stretchy="false">)</mo></math> </ephtml> . This allows much more flexibility in the choice of imputation model for each variable, at the expense of some theoretical coherence. In practice, simulations studies suggest that the approach does well, provide careful attention is given to specifying the set of imputation models so they are mutually consistent.</p> <p>Finally, hot deck approaches avoid the need to formally specify imputation models, and are potentially less vulnerable to model misspecification (although they still rely on the MAR assumption). These methods tend to perform well with large data sets, where potential donors that are close matches to recipients are plentiful. They are less useful (and results may have relatively high variance) in smaller datasets, where good matches are less plentiful. In such setting, the joint modeling or sequential regression approaches tend to be superior.</p> <hd id="AN0178879686-8">Methods Compared on Some Common Applications</hd> <p></p> <hd id="AN0178879686-9">Inference for the Mean of an Incomplete Variable Y</hd> <p>For inference about a mean (or other location parameter like the median), CC is vulnerable to bias unless the complete cases can be viewed as akin to a random sample of the original data, as when the missing data are MCAR. The bias of CC depends on the fraction of incomplete cases and the extent to which complete and incomplete cases differ on the variable of interest. Specifically, suppose a variable <emph>Y</emph> has missing values, and partition the population into strata consisting of respondents and nonrespondents to <emph>Y</emph>. Let <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>μ</mi><mrow><mi>C</mi><mi>C</mi></mrow></msub></math> </ephtml> and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>μ</mi><mrow><mi>I</mi><mi>C</mi></mrow></msub></math> </ephtml> denote the population means of <emph>Y</emph> in these strata, that is of the complete and incomplete cases, respectively. The overall mean is <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>μ</mi><mo>=</mo><msub><mi>π</mi><mrow><mi>C</mi><mi>C</mi></mrow></msub><msub><mi>μ</mi><mrow><mi>C</mi><mi>C</mi></mrow></msub><mo>+</mo><mo stretchy="false">(</mo><mn>1</mn><mo>−</mo><msub><mi>π</mi><mrow><mi>C</mi><mi>C</mi></mrow></msub><mo stretchy="false">)</mo><msub><mi>μ</mi><mrow><mi>I</mi><mi>C</mi></mrow></msub></math> </ephtml> , where <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>π</mi><mrow><mi>C</mi><mi>C</mi></mrow></msub></math> </ephtml> is the expected fraction of complete cases. Assuming a form of CC that yields an unbiased estimate of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>μ</mi><mrow><mi>C</mi><mi>C</mi></mrow></msub></math> </ephtml> , the bias of CC is: <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>μ</mi><mrow><mi>C</mi><mi>C</mi></mrow></msub><mo>−</mo><mi>μ</mi><mo>=</mo><mo stretchy="false">(</mo><mn>1</mn><mo>−</mo><msub><mi>π</mi><mrow><mi>C</mi><mi>C</mi></mrow></msub><mo stretchy="false">)</mo><mo stretchy="false">(</mo><msub><mi>μ</mi><mrow><mi>C</mi><mi>C</mi></mrow></msub><mo>−</mo><msub><mi>μ</mi><mrow><mi>I</mi><mi>C</mi></mrow></msub><mo stretchy="false">)</mo><mo>,</mo></math> </ephtml></p> <p>Graph</p> <p>the expected fraction of incomplete cases multiplied by the difference in the means for complete and incomplete cases. If the mechanism is MCAR then <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>μ</mi><mrow><mi>C</mi><mi>C</mi></mrow></msub><mo>=</mo><msub><mi>μ</mi><mrow><mi>I</mi><mi>C</mi></mrow></msub></math> </ephtml> and the bias is zero. For a given value of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>μ</mi><mrow><mi>C</mi><mi>C</mi></mrow></msub><mo>−</mo><msub><mi>μ</mi><mrow><mi>I</mi><mi>C</mi></mrow></msub><mo stretchy="false">)</mo></math> </ephtml> , the bias clearly increases with the expected fraction <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><mn>1</mn><mo>−</mo><msub><mi>π</mi><mrow><mi>C</mi><mi>C</mi></mrow></msub><mo stretchy="false">)</mo></math> </ephtml> of the incomplete cases; this is one reason why the fraction of incomplete cases is considered a useful indicator for the potential seriousness of the problem of missing data. However, the bias also depends on <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>μ</mi><mrow><mi>C</mi><mi>C</mi></mrow></msub><mo>−</mo><msub><mi>μ</mi><mrow><mi>I</mi><mi>C</mi></mrow></msub><mo stretchy="false">)</mo></math> </ephtml> , a quantity that we typically know little about.</p> <p>Suppose now we have a set of fully-observed auxiliary variables <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>X</mi><mo>=</mo><mo stretchy="false">(</mo><msub><mi>X</mi><mn>1</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>X</mi><mi>p</mi></msub><mo stretchy="false">)</mo></math> </ephtml> as well as the data on <emph>Y</emph>. The data are then <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>y</mi><mi>i</mi></msub><mo>,</mo><msub><mi>x</mi><mrow><mi>i</mi><mn>1</mn></mrow></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>x</mi><mrow><mi>i</mi><mi>p</mi></mrow></msub><mo stretchy="false">)</mo><mo>,</mo><mi>i</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><mi>r</mi></math> </ephtml> and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>x</mi><mrow><mi>i</mi><mn>1</mn></mrow></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>x</mi><mrow><mi>i</mi><mi>p</mi></mrow></msub><mo stretchy="false">)</mo></math> </ephtml> for <emph>i</emph> = <emph>r</emph><emph>+</emph> 1, ..., <emph>n</emph> where here <emph>r</emph> is the number of complete cases. CC inference for the mean of <emph>Y</emph> discards the units with <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>X</mi><mn>1</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>X</mi><mi>p</mi></msub></math> </ephtml> observed and <emph>Y</emph> missing. If the data are not MCAR, the distribution of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>X</mi><mn>1</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>X</mi><mi>p</mi></msub></math> </ephtml> differs for the complete and incomplete cases, and comparisons of these distributions, such as t tests comparing the means, provide tests of MCAR.</p> <p>Both IPW and MI exploit the information about <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>X</mi><mn>1</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>X</mi><mi>p</mi></msub></math> </ephtml> to potentially reduce the bias of CC. Specifically, in IPW the complete units are weighted by the inverse of the estimated probability that <emph>Y</emph> is observed, computed based on the response rate within adjustment cells or a regression of <emph>R</emph> on <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>X</mi><mn>1</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>X</mi><mi>p</mi></msub></math> </ephtml> . In MI, the missing values of <emph>Y</emph> are multiply-imputed as draws from the predictive distribution of <emph>Y</emph> given <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>X</mi><mn>1</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>X</mi><mi>p</mi></msub></math> </ephtml> .</p> <p>Table 1, an extension of a table in [<reflink idref="bib17" id="ref46">17</reflink>], compares the bias and variance of estimates of the mean of <emph>Y</emph> from MI and IPW relative to CC. It summarizes theoretical properties of the methods; for example, for an <emph>X</emph> variable to reduce bias it needs to be related to both <emph>Y</emph> and <emph>R</emph>, and to reduce variance it needs to be related to <emph>Y</emph>. Eight cells are displayed, based on strength of association between the auxiliary variables <emph>X</emph> and nonresponse <emph>R</emph> and outcome <emph>Y</emph>. For characterizing the association between <emph>X</emph> and <emph>Y</emph>, it is helpful to split <emph>X</emph> into the propensity, that is the best predictor of <emph>R</emph> in the regression of <emph>R</emph> on <emph>X</emph>, and components of <emph>X</emph> orthogonal to the propensity, say <emph>Z</emph>. Two types of association between <emph>X</emph> and <emph>Y</emph> are distinguished, the strength of association between the propensity to respond and <emph>Y</emph>, and the strength of association between <emph>Z</emph> and <emph>Y.</emph> With a single <emph>X</emph>, the propensity is a function of <emph>X</emph> and <emph>Z</emph> is a null set. Within each cell, the absolute bias and variance of IPW and MI estimates relative to CC are tabulated. See the footnote to the table for details.</p> <p>Graph</p> <p>Table 1. Bias and Variance of MI and IPW Relative to CC for Estimating a Mean, by Strength of Association of the Auxiliary Variables <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>X</mi><mo>=</mo><mo stretchy="false">(</mo><msub><mi>X</mi><mn>1</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>X</mi><mi>p</mi></msub><mo stretchy="false">)</mo></math> </ephtml> with Response (R) and Outcome (Y).</p> <p> <ephtml> <table><colgroup><col align="left" /><col align="center" /><col align="center" /><col align="center" /><col align="center" /></colgroup><thead><tr><th align="left" rowspan="2">Association of <italic>X</italic> with Response <italic>R</italic> (ie. strength of propensity to respond<italic>)</italic></th><th align="left" colspan="4">Association of outcome <italic>Y</italic> with (i) propensity to respond and (ii) Z (as defined in the text)</th></tr><tr><th align="left">Propensity: LowZ: Low</th><th align="left">Propensity: LowZ: High</th><th align="left">Propensity: HighZ: Low</th><th align="left">Propensity: High Z: High</th></tr></thead><tbody><tr><td><bold>Low</bold></td><td align="center">Cell LLLIPW MIBias: --- ---Var: --- ---</td><td align="center">Cell LLHIPW MIBias --- ---Var --- <p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false" xmlns="">↓</mo></math></p></td><td align="center">Cell LHLIPW MIBias: --- ---Var: <p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false" xmlns="">↓</mo></math></p><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false" xmlns="">↓</mo></math></p></td><td align="center">Cell LHHIPW MIBias: --- ---Var: <p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false" xmlns="">↓</mo></math></p><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false" xmlns="">↓↓</mo></math></p></td></tr><tr><td><bold>High</bold></td><td align="center">Cell HLLIPW MIBias: --- ---Var: <p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false" xmlns="">↑</mo></math></p> ---</td><td align="center">Cell HLHIPW MIBias: --- ---Var: <p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false" xmlns="">↑</mo></math></p><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false" xmlns="">↓</mo></math></p></td><td align="center">Cell HHLIPW MIBias: <p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false" xmlns="">↓</mo></math></p><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false" xmlns="">↓</mo></math></p>Var: <p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false" xmlns="">↓</mo></math></p><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false" xmlns="">↓</mo></math></p></td><td align="center">Cell HHHIPW MIBias: <p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false" xmlns="">↓</mo></math></p><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false" xmlns="">↓</mo></math></p>Var: <p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false" xmlns="">↓</mo></math></p><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false" xmlns="">↓↓</mo></math></p></td></tr></tbody></table> </ephtml> </p> <p>1 Notes:.</p> <ulist> <item>2 "---" for Bias (or Var) within a cell indicates that the estimate for the method has similar bias (or variance) to the estimate for CC.</item> <item>3 " <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">↓</mo></math> </ephtml> " for Bias (or Var) within a cell indicates that the estimate for the method has less absolute bias (or variance) than the estimate for CC.</item> <item>4 " <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">↑</mo></math> </ephtml> " for Bias (or Var) within a cell indicates that the estimate for the method has greater absolute bias (or variance) than the estimate for CC.</item> <item>5 In summary, " <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">↓</mo></math> </ephtml> " indicates that a method is better than CC, " <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">↑</mo></math> </ephtml> " indicates that a method is worse than CC, and "---" indicates that a method is similar to CC.</item> </ulist> <p>When <emph>X</emph> is weakly associated with both <emph>R</emph> and <emph>Y</emph> (Cell LLL), CC, IPW, and MI are similar, and CC may be preferred on grounds of simplicity. For either IPW or MI to reduce absolute bias of CC, the propensity to respond needs to be related to both <emph>R</emph> and <emph>Y</emph>, as in the cells HHL and HHH in Table 1. When the propensity is strongly associated with <emph>R</emph> but weakly associated with <emph>Y</emph> (Cells HLL and HLH), IPW actually makes things worse than CC in terms of variance, because the variability of the sample weights increases the sampling variance of the weighted mean without a compensating reduction in absolute bias. When the propensity is strongly related to <emph>Y</emph> (cells LHL, HHL, LHH and HHH), IPW can have lower variance than CC. MI does not have the increased variance of IPW in Cell HLL, and otherwise is more efficient than CC and IPW when there are auxiliary variables other than the propensity that predict <emph>Y</emph>, namely cells LLH, HLH, LHH and HHH). See [<reflink idref="bib3" id="ref47">3</reflink>] for an application where MI is more efficient than CC, and [<reflink idref="bib6" id="ref48">6</reflink>] for more on the utility of auxiliary variables for enhancing the precision of MI inferences.</p> <p>Note that MI is seen to be superior to IPW in Table 1, a result of the fact that both methods can reduce bias, but MI can reduce both bias and variance ([<reflink idref="bib13" id="ref49">13</reflink>]). However, this property relies on the assumption that the imputation model is well specified, for example, nonlinear terms and interactions among the <emph>X</emph>'s are included as predictors if they are needed. If the imputation model is misspecified but the model for <emph>R</emph> on <emph>X</emph> for IPW is well specified, then IPW may be superior to MI. Also, IPW based on a single regression of <emph>R</emph> on <emph>X</emph> can be applied to a set of variables <emph>Y</emph> with the same pattern of missing values, whereas MI requires a different imputation model for each <emph>Y</emph> - variable in the set.</p> <p>In summary, absolute bias is reduced by IPW and MI when there are auxiliary variables that are predictive of both response and <emph>Y</emph>. Sampling variance is reduced by IPW when the response propensity is predictive of <emph>Y</emph>, and MI is generally more efficient than IPW, particularly when auxiliary variables, <emph>Z</emph>, orthogonal to the propensity to respond are predictive of <emph>Y</emph>. These comments generally apply for subgroup means, with <emph>X</emph> interpreted as the set of auxiliary variables other than the variable used to form the subgroups.</p> <p>What if <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>X</mi><mn>1</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>X</mi><mi>p</mi></msub></math> </ephtml> also have missing values? Because MI can be applied to a general pattern of missing values, it can still be used to recover the information about the mean of <emph>Y</emph> in the observed auxiliary variables for units where <emph>Y</emph> is missing, although additional modeling is needed to develop imputation models for the missing values of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>X</mi><mn>1</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>X</mi><mi>p</mi></msub></math> </ephtml> . The weights in IPW can only condition on the subset of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>X</mi><mn>1</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>X</mi><mi>p</mi></msub></math> </ephtml> that are completely observed, and therefore may be less effective in reducing absolute bias or variance than MI. This is one reason why imputation is generally favored over weighting for item nonresponse, which often gives an unstructured "swiss-cheese" appearance to the matrix <emph>R</emph> of response indicators.</p> <hd id="AN0178879686-10">Missing Data in Regression</hd> <p>The available information in the incomplete cases is different if interest concerns the coefficients of the regression of <emph>Y</emph> on <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>X</mi><mo>=</mo><mo stretchy="false">(</mo><msub><mi>X</mi><mn>1</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>X</mi><mi>p</mi></msub><mo stretchy="false">)</mo></math> </ephtml> rather than the mean of <emph>Y</emph>. With <emph>X</emph> fully observed and missing values confined to <emph>Y</emph>, and assuming MAR, the likelihood has the form <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><mi>L</mi><mo stretchy="false">(</mo><mi>θ</mi><mo>,</mo><mi>ϕ</mi><mo fence="false" stretchy="false">|</mo><mrow><mi mathvariant="normal">data</mi><mo stretchy="false">)</mo></mrow><mo>=</mo><munderover><mo movablelimits="false">∏</mo><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mi>r</mi></munderover><mrow><mspace width=".1em" /><msub><mi>f</mi><mrow><mi>Y</mi><mo fence="false" stretchy="false">|</mo><mi>X</mi></mrow></msub><mo stretchy="false">(</mo><msub><mi>y</mi><mi>i</mi></msub><mo fence="false" stretchy="false">|</mo><msub><mi>x</mi><mrow><mi>i</mi><mn>1</mn></mrow></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>x</mi><mrow><mi>i</mi><mi>p</mi></mrow></msub><mo>,</mo><mi>θ</mi><mo stretchy="false">)</mo></mrow><mo>×</mo><munderover><mo movablelimits="false">∏</mo><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mi>n</mi></munderover><mrow><mspace width=".1em" /><msub><mi>f</mi><mi>X</mi></msub><mo stretchy="false">(</mo><msub><mi>x</mi><mrow><mi>i</mi><mn>1</mn></mrow></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>x</mi><mrow><mi>i</mi><mi>p</mi></mrow></msub><mo>,</mo><mi>ϕ</mi><mo stretchy="false">)</mo></mrow></math> </ephtml></p> <p>Graph</p> <p>where <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>f</mi><mrow><mi>Y</mi><mo fence="false" stretchy="false">|</mo><mi>X</mi></mrow></msub><mo stretchy="false">(</mo><msub><mi>y</mi><mi>i</mi></msub><mo fence="false" stretchy="false">|</mo><msub><mi>x</mi><mrow><mi>i</mi><mn>1</mn></mrow></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>x</mi><mrow><mi>i</mi><mi>p</mi></mrow></msub><mo>,</mo><mi>θ</mi><mo stretchy="false">)</mo></math> </ephtml> is the density of the conditional distribution of <emph>Y</emph> given <emph>X</emph> and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>f</mi><mi>X</mi></msub><mo stretchy="false">(</mo><msub><mi>x</mi><mrow><mi>i</mi><mn>1</mn></mrow></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>x</mi><mrow><mi>i</mi><mi>p</mi></mrow></msub><mo>,</mo><mi>ϕ</mi><mo stretchy="false">)</mo></math> </ephtml> is the density of the marginal distribution of <emph>X</emph>. It follows that provided <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>θ</mi></math> </ephtml> and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>ϕ</mi></math> </ephtml> are distinct parameters, the complete cases carry all the information for the parameters <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>θ</mi></math> </ephtml> of the regression of <emph>Y</emph> on <emph>X</emph>, so CC is in fact fully efficient. CC is also unbiased under MAR, not requiring the stronger MCAR assumption. Under MAR, MI is not necessary in this situation. IPW is sometimes advocated over CC because it is potentially more robust than CC when the regression of <emph>Y</emph> on <emph>X</emph> is misspecified but the model for the propensity to respond is correctly specified. From a robustness perspective, comparing the results from CC and IPW is sensible, if only as a specification check for the regression of <emph>Y</emph> on <emph>X</emph>.</p> <p>The incomplete cases have more information when there are missing values in the covariates <emph>X</emph> rather than missing values in the outcome <emph>Y</emph>. Suppose, for simplicity, that values of one of the covariates, say <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>X</mi><mn>1</mn></msub></math> </ephtml> , are missing, and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>X</mi><mn>2</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>X</mi><mi>p</mi></msub><mo>,</mo><mi>Y</mi></math> </ephtml> are fully observed. The incomplete cases with <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>X</mi><mn>1</mn></msub></math> </ephtml> missing then have considerable information for the intercept and coefficients of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>X</mi><mn>2</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>X</mi><mi>p</mi></msub></math> </ephtml> , but very limited information for the coefficient of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>X</mi><mn>1</mn></msub></math> </ephtml> ([<reflink idref="bib14" id="ref50">14</reflink>]). The incomplete cases are thus of limited value if the primary interest is in the coefficient of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>X</mi><mn>1</mn></msub></math> </ephtml> , but are of considerably more value if the primary interest is in other coefficients; in particular, if <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>X</mi><mn>1</mn></msub></math> </ephtml> is weakly associated with <emph>Y</emph>, the incomplete cases have about as much information as the complete cases for these other regression coefficients. For MI of covariates, it is important to include the outcome variable <emph>Y</emph> in the imputation model so as not to bias the estimated regression coefficients from the fitted data ([<reflink idref="bib14" id="ref51">14</reflink>]).</p> <p>CC also has the (perhaps unexpected) property of being unbiased for regression coefficients when the probability that a case is complete depends on the covariates but — given these — not the outcome, under a well-specified model ([<reflink idref="bib11" id="ref52">11</reflink>], [<reflink idref="bib16" id="ref53">16</reflink>], Example 3.3). For the simple missingness pattern discussed above, we consider three cases:</p> <p></p> <ulist> <item> If missingness of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>X</mi><mn>1</mn></msub></math> </ephtml> depends on <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>X</mi><mn>2</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>X</mi><mi>p</mi></msub><mo stretchy="false">)</mo></math> </ephtml> but not <emph>Y</emph> or <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>X</mi><mn>1</mn></msub></math> </ephtml> , then data are MAR, and CC, IPW and MI are all consistent;</item> <p></p> <item> If missingness of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>X</mi><mn>1</mn></msub></math> </ephtml> depends on <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>X</mi><mn>2</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>X</mi><mi>p</mi></msub><mo stretchy="false">)</mo></math> </ephtml> and <emph>Y</emph> but not on <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>X</mi><mn>1</mn></msub></math> </ephtml> , then data are again MAR; IPW and MI are consistent but CC is generally biased;</item> <p></p> <item> If missingness of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>X</mi><mn>1</mn></msub></math> </ephtml> depends on <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>X</mi><mn>1</mn></msub><mo>,</mo><msub><mi>X</mi><mn>2</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>X</mi><mi>p</mi></msub><mo stretchy="false">)</mo></math> </ephtml> but not <emph>Y</emph>, then data are MNAR; MI (based on an MAR model) and IPW are generally biased, but CC is consistent.</item> </ulist> <p>Thus CC is biased when missingness depends on <emph>Y</emph> (given <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>X</mi><mn>1</mn></msub><mo>,</mo><msub><mi>X</mi><mn>2</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>X</mi><mi>p</mi></msub><mo stretchy="false">)</mo></math> </ephtml> ), whereas IPW and MI methods are biased when missingness depends on <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>X</mi><mn>1</mn></msub></math> </ephtml> (given <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>X</mi><mn>2</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>X</mi><mi>p</mi></msub><mo stretchy="false">)</mo></math> </ephtml> and <emph>Y</emph>). MI is, however, more efficient than CC and IPW when it comes to the standard error of the regression coefficients, so it may be preferred from a mean square error perspective even if a moderate MNAR mechanism is suspected.</p> <p>If that <emph>p</emph> is large, and missing data are scattered over the covariates <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>X</mi><mn>1</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>X</mi><mi>p</mi></msub></math> </ephtml> in a haphazard way, such that the fraction of complete cases is relatively small. Then there could be substantial payoff in terms of increased precision in using MI to include the incomplete cases in the analysis. As before, IPW has more limited potential because weighting the complete cases does not exploit the information available in incomplete data.</p> <p>A hybrid of CC and MI called <emph>subset MI</emph> (SMI, [<reflink idref="bib18" id="ref54">18</reflink>]) has potential value in situations where something is known about the missingness mechanism. Partition the covariates into two sets, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>X</mi><mo>=</mo><mo stretchy="false">(</mo><mi>U</mi><mo>,</mo><mi>V</mi><mo stretchy="false">)</mo></math> </ephtml> , and suppose it is suspected that the probability that <emph>U</emph> is complete depends on the covariates <emph>U</emph> and <emph>V</emph> but not <emph>Y</emph>, and missing values of <emph>Y</emph> and <emph>V</emph> are MAR in the set of units where <emph>U</emph> is fully observed<emph>.</emph> Then cases that have missing values for any of the variables in <emph>U</emph> are discarded, and in the remaining cases, MI is applied to fill in any missing values in <emph>V</emph>. The resulting data set contains complete cases in <emph>U</emph> and some cases with imputed values in <emph>V</emph>. [<reflink idref="bib18" id="ref55">18</reflink>] show that this method is unbiased, whereas as CC, IPW and MI (applied to the full dataset) are all subject to bias. Thus, for example, if <emph>income</emph> is an incomplete covariate, and missingness of <emph>income</emph> is suspected to depend on the underlying <emph>income</emph> value, then one might discard units with <emph>income</emph> missing and apply MI to the resulting dataset. Another application of this idea arises when the outcome variable <emph>Y</emph> has missing values: apply MI to the whole data set, but then drop units with <emph>Y</emph> missing when estimating the regression of <emph>Y</emph> on <emph>X</emph>, because after MI these units carry no additional information for the regression ([<reflink idref="bib40" id="ref56">40</reflink>]). Here, dropping the incomplete cases avoids simulation error from multiply-imputing values of <emph>Y</emph>.</p> <hd id="AN0178879686-11">Bias and Precision of Complete-Case Inferences for an Odds Ratio</hd> <p>Suppose the data consist of two binary variables <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>1</mn></msub></math> </ephtml> and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>2</mn></msub></math> </ephtml> , and the complete cases have both <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>1</mn></msub></math> </ephtml> and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>2</mn></msub></math> </ephtml> observed, and the incomplete cases have either <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>1</mn></msub></math> </ephtml> or <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>2</mn></msub></math> </ephtml> missing, The parameter of interest is the odds ratio in the 2 × 2 table of counts classified by <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>1</mn></msub></math> </ephtml> and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>2</mn></msub></math> </ephtml> , The CC estimate is then the sample odds ratio based on the complete cases. This analysis is not subject to selection bias if the probability of response depends on <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>1</mn></msub></math> </ephtml> alone, or <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>2</mn></msub></math> </ephtml> alone, or more generally if the logarithm of the probability of response is an additive function of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>1</mn></msub></math> </ephtml> and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>2</mn></msub></math> </ephtml> . This result underpins the validity of case-control studies for estimating odds ratios from observational studies. In terms of precision, supplemental margins on <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>1</mn></msub></math> </ephtml> and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>2</mn></msub></math> </ephtml> provide very little information for the odds ratio, but can reduce bias and increase precision for estimating the marginal distributions of these variables, which may be of substantive interest ([<reflink idref="bib16" id="ref57">16</reflink>], Example 3.4) For further discussion see [<reflink idref="bib2" id="ref58">2</reflink>].</p> <p>Similar comments apply to the coefficient of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>1</mn></msub></math> </ephtml> in the logistic regression of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>2</mn></msub></math> </ephtml> on <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>1</mn></msub></math> </ephtml> and <emph>X</emph>, when the incomplete cases have either <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>1</mn></msub></math> </ephtml> or <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>2</mn></msub></math> </ephtml> missing and <emph>X</emph> are additional fully observed covariates. The exponent of the coefficient of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>1</mn></msub></math> </ephtml> in that case represents a conditional odds ratio, given <emph>X</emph>.</p> <hd id="AN0178879686-12">Missing Data in Repeated Measures</hd> <p>Longitudinal data are often subject to missing data because individuals drop out before the study ends. This form of missingness is often not MCAR, because the distribution of the study outcomes is different for those who do and do not drop out. A common analysis approach is CC, an approach whose validity again depends on the analysis of interest, and how much information the incomplete cases carry for that analysis.</p> <p>For example, consider a simple design with fully observed baseline measure on a study variable <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>0</mn></msub></math> </ephtml> and a single follow-up measure of the study variable, say <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>1</mn></msub></math> </ephtml> , which is missing for individuals who drop out. How much information is available in the dropouts with <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>0</mn></msub></math> </ephtml> observed but <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>1</mn></msub></math> </ephtml> missing? If <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>0</mn></msub></math> </ephtml> is highly predictive of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>1</mn></msub></math> </ephtml> but weakly predictive of the change variable <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>1</mn></msub><mo>−</mo><msub><mi>Y</mi><mn>0</mn></msub></math> </ephtml> , the incomplete cases are informative for the mean of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>1</mn></msub></math> </ephtml> but relatively uninformative for the mean of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>1</mn></msub><mo>−</mo><msub><mi>Y</mi><mn>0</mn></msub></math> </ephtml> , and (assuming MAR), as we have already seen, have no information at all for the regression of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>1</mn></msub></math> </ephtml> on <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>0</mn></msub></math> </ephtml> . So CC is justified for the latter two analyses but less justified for first analysis.</p> <p>As a more complex example, suppose the study involves fully-observed baseline measures <emph>X</emph> and <emph>K</emph> repeated measures <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>Y</mi><mo>=</mo><mo stretchy="false">(</mo><msub><mi>Y</mi><mn>1</mn></msub><mo>,</mo><msub><mi>Y</mi><mn>2</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>Y</mi><mi>K</mi></msub><mo stretchy="false">)</mo></math> </ephtml> on a study variable, and individuals dropping out between times <emph>k</emph> and <emph>k</emph> + 1 have <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mn>1</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>Y</mi><mi>k</mi></msub></math> </ephtml> observed and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mrow><mi>k</mi><mo>+</mo><mn>1</mn></mrow></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>Y</mi><mi>K</mi></msub></math> </ephtml> missing. The incomplete cases often have substantial information for the regression of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>Y</mi><mi>K</mi></msub></math> </ephtml> on <emph>X</emph>, and a repeated-measures model should be used to model the longitudinal distribution of <emph>Y</emph> given <emph>X</emph>. This is particularly so if the intermediate values of <emph>Y</emph> measured prior to drop-out are predictive of the missing values of <emph>Y</emph> after dropout. Specifically, use a repeated measures model (fully efficient if correctly specified) for all the observed data, with carefully chosen covariance structure and a parameterization of the mean model chosen to answer the scientific question; for examples, see [<reflink idref="bib5" id="ref59">5</reflink>] chapter 3. Because ML for the repeated-measures model is fully efficient, MI or IPW is not needed.</p> <hd id="AN0178879686-13">Other Analyses</hd> <p>We have focused attention on analysis of means and regression parameters here, because these analyses are common in social science data sets. CC, IPW and MI can all be applied to other types of analysis, such as loglinear models for contingency tables, time series modeling, or analyses that involve latent variables like factor analysis or latent structure analysis. To keep this paper a manageable length, we preclude a detailed discussion of these other kinds of analyses, but offer some general comments for completeness.</p> <p></p> <ulist> <item> CC analysis simply applies the analysis of interest to the complete cases, and is often a default option in computer packages, provided missing data codes created to represent missing values are recognized by the package – beware of having missing-data codes like −9999 treated as if they are real values!</item> <p></p> <item> IPW can be applied in any computer software provided the software allows for weighting the cases.</item> <p></p> <item> MI can be applied to fill in the missing values, and then the analysis method of interest applied to each of the filled-in data sets. Parameter estimates from the analysis of each data set, and associated standard errors, can be combined using standard MI combining rules. This approach should work well provided the predictive distributions for filling in each of the incomplete variables are reasonable – a MI approach like chained equations ([<reflink idref="bib40" id="ref60">40</reflink>]; [<reflink idref="bib24" id="ref61">24</reflink>]) is recommended, as these methods handle a general pattern of missing data, and allow for flexible modeling of the conditional distribution of each of the incomplete variables given the other variables in the data set.</item> <p></p> <item> An alternative is to apply ML for the analysis of interest. Examples are provided in [<reflink idref="bib16" id="ref62">16</reflink>]. This may be more efficient than chained equation MI because it reflects features of the analysis of interest in the joint distribution of the variables. On the other hand, as noted above, chained equation MI is more flexible, and can included auxiliary variables not included in the final analysis.</item> <p></p> <item> As discussed in above when comparing the methods on some common applications, the relative gain of MI over CC or IPW depends on how much information is contained for the analysis of interest in the incomplete cases, and this can vary quite dramatically depending on the context. For model-based analysis, information is measured by the observed or expected information matrix, which is the second derivative of the loglikelihood with respect to the parameters. The relative size of this measure of information for a particular parameter in complete and incomplete cases then determines the loss of efficiency of CC, and can be calculated for particular models and parameters, under assumptions about the missing data mechanism. For more details, See [<reflink idref="bib16" id="ref63">16</reflink>], Section 8.4.3).</item> </ulist> <hd id="AN0178879686-14">Illustrative Example</hd> <p></p> <hd id="AN0178879686-15">Introduction</hd> <p>We illustrate some of the points from previous sections with an analysis of data from the Youth Cohort (Time) Series (YCS) for England, Wales and Scotland, 1984 − 2002 ([<reflink idref="bib35" id="ref64">35</reflink>]). The raw data are freely available from the U.K. data archive, http://data-archive.ac.uk, study number SN 5765. The data come from two UK representative government-funded cohort studies set up to examine the effects of social, economic and policy change on young people's experiences of education and transitions to the labor market. For our analyses we use a subset of the data from school children attending all school types in England and Wales from five YCS cohorts, who reached the end of Year 11 (ie age 16 + years) in years 1990, 1993, 1995, 1997 and 1999). All our analyses use Stata 15.1.</p> <p>We compare estimates from CC, IPW and MI of the distribution of parental occupation, and the regression of Year 11 educational achievement (in the General Certificate of Secondary Education qualifications), on the covariates cohort, boy, ethnicity and a three-level classification of parental occupation derived from information provided by the school children. A description of these variables is given in Table 2.</p> <p>Graph</p> <p>Table 2. Description of Youth Cohort Series Variables.</p> <p> <ephtml> <table><colgroup><col align="left" /><col align="left" /></colgroup><thead><tr><th align="left">Variable name</th><th align="left">Description</th></tr></thead><tbody><tr><td>Educational Achievement Score</td><td>GCSE points score. Each pupil sits up to 15 GCSE exams. The results for each are converted into a score from 7 (highest grade) to 0 (fail). These are summed across a pupil's exams and capped at 84 (equivalent to 12 GCSEs at the top grade).</td></tr><tr><td>Cohort</td><td>year of data collection: 1990, 93, 95, 97, 99</td></tr><tr><td>Boy</td><td>indicator variable for boys</td></tr><tr><td>Occupation</td><td>parental occupation, categorized as managerial, intermediate or working</td></tr><tr><td>Ethnicity</td><td>Categorized as Bangladeshi, Black, Indian, other Asian, Other, Pakistani or White</td></tr></tbody></table> </ephtml> </p> <p>The dataset contains information on 76,891 children. Data on the covariates boy and cohort are complete; however, the other variables have some missing values, as shown in Table 3. We see that the principal missing data pattern is missing parental occupation (which was derived from a series of questions asked to the pupils).</p> <p>Graph</p> <p>Table 3. Principal Missing Data Patterns in the YCS.</p> <p> <ephtml> <table><colgroup><col align="left" /><col align="left" /><col align="left" /><col align="left" /><col align="left" /><col align="left" /></colgroup><thead><tr><th align="left">Pattern</th><th align="left">GCSE score</th><th align="left">Occupation</th><th align="left">Ethnicity</th><th align="left">n</th><th align="left">% of total</th></tr></thead><tbody><tr><td>1</td><td>√</td><td>√</td><td>√</td><td>66965</td><td>87%</td></tr><tr><td>2</td><td>√</td><td>?</td><td>√</td><td>7523</td><td>10%</td></tr><tr><td>3</td><td>?</td><td>√</td><td>√</td><td>760</td><td><1%</td></tr><tr><td>4</td><td>√</td><td>?</td><td>?</td><td>651</td><td><1%</td></tr><tr><td>5</td><td align="center" colspan="3">Other patterns</td><td>892</td><td><1%</td></tr></tbody></table> </ephtml> </p> <hd id="AN0178879686-16">Estimation of Weights for IPW Analyses</hd> <p>The weights for IPW analyses are the inverse of estimates of the probability of being complete, computed by a logistic regression with outcome 1 if a unit is complete and 0 otherwise. If we include all the 76, 891 units in this regression, we are restricted to variables that are fully observed, namely boy and cohort. To include other predictors in this regression that are relatively complete, we fit the logistic regression using the 74,488 units that have boy, cohort, occupation, ethnicity and GCSE score observed. Because of the size of the data set, all the coefficients in this model are highly significant and are kept in the final model. Table 4 shows the resulting distribution of inverse probability weights.</p> <p>Graph</p> <p>Table 4. Distribution of Inverse Probability Weights from Logistic Model for the Probability That a Unit is Complete.</p> <p> <ephtml> <table><colgroup><col align="left" /><col align="left" /><col align="left" /><col align="left" /><col align="left" /><col align="left" /><col align="left" /><col align="left" /><col align="left" /><col align="left" /></colgroup><thead><tr><th align="left" /><th align="left" colspan="9">Percentile of weight distribution</th></tr></thead><tbody><tr><td /><td>2.5</td><td>25</td><td>50</td><td>75</td><td>97.5</td><td>99</td><td>99.5</td><td>99.9</td><td>100</td></tr><tr><td /><td>1.02</td><td>1.05</td><td>1.07</td><td>1.13</td><td>1.39</td><td>1.72</td><td>2.06</td><td>3.10</td><td>6.19</td></tr></tbody></table> </ephtml> </p> <hd id="AN0178879686-17">Creation of Multiply-Imputed Data Sets for MI Analyses</hd> <p>MI is more computationally demanding, but still quite straightforward. We apply MI by chained equations, using the software program Stata. As there are three variables with missing data, we have three chained regression imputation equations which the imputation algorithm cycles through:</p> <p> <emph>linear regression of GCSE score on: ethnicity, parental occupation, sex, cohort</emph> </p> <p> <emph>multinomial regression of ethnicity on: GCSE score, parental occupation, sex, cohort</emph> </p> <p> <emph>multinomial regression of parental occupation on: GCSE score, ethnicity, sex, cohort</emph> </p> <p>Each of these properly imputes the missing data in the dependent variable, and then takes these imputed values through to the next model. We complete 10 cycles before imputing each data set and a further 10 cycles between each of our 20 imputations.</p> <p>A complication in this imputation is that the ethnicity variable has a number of relatively sparse categories, leading to quasi-complete separation. Unless this is corrected for, this can cause coefficients in the multinomial regression model to become large in magnitude, and corresponding SEs to be large too, leading to poor imputations. A relatively simple fix, which we use here, is to temporarily augment the data set with small number of observations at each point when this occurs ([<reflink idref="bib41" id="ref65">41</reflink>]).</p> <p>We now illustrate and compare the results of two analyses using CC, IPW and MI.</p> <hd id="AN0178879686-18">Estimated Marginal Distribution of Ethnicity</hd> <p>Because Ethnicity is the variable with the highest proportion of missing values, we compare estimates of the distribution of this variable from the different methods in Table 5. This analysis is similar to an analysis of means as discussed in above, because the proportion of cases in a particular category can be viewed as the mean of a binary variable indicating belonging to that category.</p> <p>Graph</p> <p>Table 5. Distribution of Parental Occupation, Estimated from CC, IPW and MI for (i) Whole Data Set (top) and (ii) Bangladeshi Ethnic Group (Bottom).</p> <p> <ephtml> <table><colgroup><col align="left" /><col align="left" /><col align="left" /><col align="left" /></colgroup><thead><tr><th align="left" /><th align="left" colspan="3">Parental Occupation (estimate, 95% CI)</th></tr><tr><th align="left">Estimated using:</th><th align="left">managerial and professional</th><th align="left">intermediate</th><th align="left">working</th></tr></thead><tbody><tr><td>CC (n = 66,965)</td><td>43.3% (42.9%−43.6%)</td><td>33.5% (33.2%−33.9%)</td><td>23.2% (22.8%−23.5%)</td></tr><tr><td>IPW (n = 66,965)</td><td>42.2% (41.9%−42.6%)</td><td>33.8% (33.4%−34.1%)</td><td>24.0% (23.7%−24.3%)</td></tr><tr><td>MI (n = 76,791)</td><td>41.8% (41.4%−42.1%)</td><td>33.8% (33.5%−34.2%)</td><td>24.4% (24.1%−24.7%)</td></tr><tr><td>Bangladeshi ethnicity only:</td><td /><td /><td /></tr><tr><td>CC (n = 246)</td><td>14.2% (10.4%−19.2%)</td><td>41.5% (35.4%−47.8%)</td><td>44.3% (38.2%−50.6%)</td></tr><tr><td>IPW (n = 246)</td><td>12.5% (8.8%−17.4%)</td><td>41.0% (34.5%−47.8%)</td><td>46.5% (39.8%−53.4%)</td></tr><tr><td>MI (n = 593–605)</td><td>11.0% (7.7%−13.6%)</td><td>40.4% (34.7%−46.1%)</td><td>48.9% (43.2%−54.6%)</td></tr></tbody></table> </ephtml> </p> <p>6 Note: As Bangladeshi ethnicity also has missing values (which are imputed) the number of cases of Bangladeshi ethnicity varies across the 20 imputations from 593–605 cases.</p> <p>Because of the relatively small proportion of values with parental occupation missing, overall the marginal distribution of the variable is similar for CC, IPW and ML analysis, both in terms of point estimates and standard errors. However, the proportion of missing parental occupation values is much higher in some ethnic groups, and highest in the Bangladeshi group. For the Bangladeshi group, in particular, point estimates for CC appear biased relative to those which assume MAR (IPW and MI) and we see that MI has notably narrower confidence intervals. As preliminary analysis suggests we are in cells HHL or HHH of Table 1, this is in line with what we expect.</p> <hd id="AN0178879686-19">Estimated Regression of GCSE Score on Covariates</hd> <p>The regression analysis results CC, IPW and MI are summarized in Table 6. For simplicity, we treat the weights as fixed rather than accounting for their sampling error, a strategy that tends to yield conservative standard errors (which is of no concern in a data set of this size). We focus here on the coefficients of ethnicity, because the coefficients of the other variables are similar for the three methods.</p> <p>Graph</p> <p>Table 6. Estimated Effects of Ethnicity on GCSE Score (Estimate, SE), Adjusted for Cohort, sex and Parental Occupation; (i) Left, CC Analysis (ii) Centre, IPW; (iii) Right, MI.</p> <p> <ephtml> <table><colgroup><col align="left" /><col align="center" /><col align="center" /><col align="center" /></colgroup><thead><tr><th align="left">Ethnic group (reference group: white)</th><th align="left">CC analysis</th><th align="left">IPW analysis</th><th align="left">MI analysis (20 imputations)</th></tr></thead><tbody><tr><td>Black (1.8%)</td><td>−5.77 (0.540)</td><td>−7.62 (0.593)</td><td>−7.57 (0.488)</td></tr><tr><td>Indian (2.9%)</td><td>4.27 (0.392)</td><td>3.71 (0.412)</td><td> 3.79 (0.380)</td></tr><tr><td>Pakistani (2.0%)</td><td>−2.10 (0.557)</td><td> −5.25 (0.657)</td><td>−4.05 (0.452)</td></tr><tr><td>Bangladeshi (0.9%)</td><td>0.61 (1.007)</td><td> −4.37 (1.262)</td><td>−3.91 (0.710)</td></tr><tr><td>Other Asian (1.3%)</td><td>6.20 (0.586)</td><td>5.07 (0.653)</td><td> 5.32 (0.548)</td></tr><tr><td>Other (1.2%)</td><td>−0.20 (0.630)</td><td> −1.72 (0.748)</td><td>−1.26 (0.590)</td></tr></tbody></table> </ephtml> </p> <p>The estimates of the ethnicity coefficients from CC differ markedly from the estimates from IPW and MI (which are similar). In particular, we see that the estimated coefficient for Bangladeshi ethnicity, which is not statistically significant in the CC analysis, is now statistically significant and similar to the coefficient for the Pakistani group; both coefficients are substantially more negative. Also, the coefficient for Black is slightly more negative, and that for Other Asian more positive. The key reason why the IPW/MI results are relatively unbiased compared to CC is that the probability of a CC (equivalent, for most individuals, to the probability that <emph>occupation</emph> is observed) is strongly influenced by the outcome (GCSE score) and <emph>ethnicity</emph>. In this case, theory (see above) suggests that ethnicity coefficients are likely to be biased. For CC to be less biased than IPW/MI, a strong MNAR mechanism would need to be operating. However, [<reflink idref="bib8" id="ref66">8</reflink>]:240) analyzed a different version of these data and showed that inferences are robust to departures from MAR.</p> <p>While coefficient estimates from IPW and MI are similar, Table 6 shows standard errors are considerably reduced with MI (especially in the Bangladeshi group) compared to the CC analysis. This is typical, and consistent with the discussion above: here MI is mostly bringing back into the analysis individuals with missing <emph>occupation</emph>, but observed data on other variables. Therefore, we would expect coefficients for ethnicity categories, to have greatest reduction in their standard error. This is further emphasized here for the Bangladeshi category, because this group is one of the smallest, yet has the one of the highest proportions of individuals with missing <emph>occupation.</emph></p> <p>By contrast SEs for IPW are larger than for the CC analysis. This is again typical; an intuitive explanation is that IPW only reweights cases with no missing data. Cases with one or more missing values are therefore discarded by IPW, whereas all the information is included in MI. Indeed, provide the imputation model is appropriately specified, theory and experience suggest it makes best use of the available data.</p> <p>Finally, note that both IPW and MI assume that the data are MAR. As discussed earlier, this is typically the natural assumption for a primary analysis. However, as it is untestable from the data at hand, it is often useful to perform sensitivity analysis. This can also be readily carried out using MI; see [<reflink idref="bib8" id="ref67">8</reflink>]:240) for an analysis of different version of these data, which shows that inferences are robust to the MAR assumption.</p> <hd id="AN0178879686-20">Conclusions</hd> <p>We have presented a non-technical discussion of three widely used approaches for handling missing data, namely CC, IPW and MI. In applications, we always begin by tabulating and graphing the data, and exploring the associations using complete case analyses. As we move to the definitive analysis, Table 7 summarizes how we choose between the approaches in our work.</p> <p>Graph</p> <p>Table 7. Summary of Recommendations.</p> <p> <ephtml> <table><colgroup><col align="left" /><col align="left" /><col align="left" /></colgroup><thead><tr><th align="left">Analysis method</th><th align="left">When to use</th><th align="left">When to avoid</th></tr></thead><tbody><tr><td>CC</td><td><list list-type="Bullet"><list-item><p>Unbiased (though not fully efficient) for estimating regression coefficients when the probability that a case is complete depends on the covariates but, given these, not the outcome.</p></list-item><list-item><p>Can be used to estimate an OR if the probability of a complete case depends the exposure or the outcome (but not both)</p></list-item></list></td><td><list list-type="Bullet"><list-item><p>when estimating the mean of an incomplete outcome if data are not MCAR (because likely to be biased)</p></list-item><list-item><p>when there are auxiliary variables that can be used to recover missing data (because MI or IPW will be more efficient)</p></list-item></list></td></tr><tr><td>IPW</td><td><list list-type="Bullet"><list-item><p>More efficient than a CC analysis if there are useful auxiliary variables</p></list-item><list-item><p>Valid for estimating a regression coefficient if the missingness mechanism is MAR</p></list-item></list></td><td><list list-type="Bullet"><list-item><p>When MI is also valid, IPW is generally less efficient because IPW (i) only reweights complete cases and (ii) cannot use incomplete auxiliary variables</p></list-item></list></td></tr><tr><td>MI</td><td><list list-type="Bullet"><list-item><p>More efficient than a CC analysis if there are useful auxiliary variables</p></list-item><list-item><p>Valid for estimating a regression coefficient if the missingness mechanism is MAR</p></list-item><list-item><p>When IPW and MI are valid, and correctly implemented, MI is typically more efficient.</p></list-item></list></td><td><list list-type="Bullet"><list-item><p>When we are not confident the imputation model is (i) consistent with the scientific model and (ii) well specified (because results at increased risk of bias)</p></list-item></list></td></tr></tbody></table> </ephtml> </p> <p>In particular, when data are plausibly MAR, MI and IPW can improve efficiency and reduce bias over a CC analysis. The relative gain of MI over CC or IPW depends on how much information is contained in the incomplete cases for the scientific analysis. Further, IPW and MI can both exploit information in auxiliary variables (which are not included in the scientific model) to (i) increase the plausibility of the MAR assumption, and hence reduce bias and (ii) further improve efficiency – especially with MI, when the auxiliary variables are good predictors of the variables with missing values in the scientific model. A further advantage of MI is that it can be used when these auxiliary variables themselves have missing values.</p> <p>However, the advantages of MI are contingent on the imputation model being well specified, in terms of assumed relationships between the missing and observed variables. In scientific models where there are, for example, non-linear effects, interactions, hierarchical (multilevel) structure and time-to-event outcomes, considerable thought needs to go into the specification of the imputation model. Analysts should also check that the distribution of imputed data is plausible in the scientific context (e.g. graphically). [<reflink idref="bib7" id="ref68">7</reflink>] discuss a number of examples in detail, and provide further references. One robust MI approach is Penalized Spline of Propensity Prediction, which imputes missing variables based on a model that includes a penalized spline of the estimated response propensity and other predictive covariates ([<reflink idref="bib43" id="ref69">43</reflink>], [<reflink idref="bib44" id="ref70">44</reflink>]).</p> <p>When MI is not indicated, and we are choosing between CC and IPW, we reiterate the following. First, for inferences about means, use IPW when auxiliary variables are available that are strongly related to both response and the variable with missing values. Second, for inferences about regression with all (or most) missing values in the outcome alone, CC is valid if the regression model is correctly specified. However, it is prudent to compare results from IPW and CC as a specification check. If the estimated regression coefficients from the two analyses are very different, (and a careful check of the IPW model does not highlight any concerns) the specification of the regression model needs to be checked for errors (for example, assumptions about linearity of absence of interactions may be invalid.) Third, for inferences about regression with missing values in the covariates, IPW is preferred if the missingness mechanism is MAR, CC is preferred if the missingness mechanism plausibly depends on the covariates but not (or only weakly on) the outcome.</p> <p>In our work, we typically compare the CC analysis with either MI or IPW (or both) and seek to understand and explain in our reporting why they differ, because such explanations typically give additional insights and hence improve confidence in the scientific findings.</p> <p>In practice the mechanism behind the missing data will not be known, and requires making an assumption of the most plausible mechanism. It is therefore important to conduct a sensitivity analysis to alternative plausible assumptions regarding the missingness mechanism. One approach to clarifying the assumptions regarding the missingness mechanism in the primary and sensitivity analysis is to use causal diagrams ([<reflink idref="bib12" id="ref71">12</reflink>]).</p> <p>Finally, if the data are suspected to be MNAR, then it is important to consider an MNAR model, at least as a sensitivity analysis. A discussion of MNAR models is beyond the scope of this manuscript, but MI provides a practical vehicle, as described by [<reflink idref="bib7" id="ref72">7</reflink>] and references therein.</p> <hd id="AN0178879686-21">Acknowledgments</hd> <p>James Carpenter is supported by UK Medical Research Council Grant MC_UU_00004/07. The manuscript is submitted on behalf of the STRengthening Analytical Thinking for Observational Studies (STRATOS) initiative (<ulink href="http://stratos-initiative.org">http://stratos-initiative.org</ulink>), which aims to provide accessible and accurate guidance documents for relevant topics in the design and analysis of observational studies. The authors thank the the editor, associate editor, three referees and two reviewers on the STRATOS publication panel for their helpful comments on the manuscript.</p> <ref id="AN0178879686-22"> <title> References </title> <blist> <bibl id="bib1" idref="ref20" type="bt">1</bibl> <bibtext> Andridge Rebecca H., Little Roderick J. 2010. " A Review of Hot Deck Imputation for Survey Nonresponse." International Statistical Review. 78(1):40‐64.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref37" type="bt">2</bibl> <bibtext> Bartlett Jonathan W., Harel Ofer, Carpenter James R. 2015. " Asymptotically Unbiased Estimation of Exposure Odds Ratios in Complete Records Logistic Regression." American Journal of Epidemiology. 182(8):730‐6.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref41" type="bt">3</bibl> <bibtext> Belin Tom R. 2009. " Missing Data: What a Little can do, and What Researchers can do in Response." American Journal of Ophthalmology. 148(6):820‐2.</bibtext> </blist> <blist> <bibl id="bib4" idref="ref24" type="bt">4</bibl> <bibtext> Cao Weihua, Tsiatis Anastasios A., Davidian Marie. 2009. " Improving Efficiency and Robustness of the Doubly Robust Estimator for a Population Mean with Incomplete Data." Biometrika. 96:723‐34.</bibtext> </blist> <blist> <bibl id="bib5" idref="ref59" type="bt">5</bibl> <bibtext> Carpenter James R., Kenward Michael G. 2008. " Missing Data in Clinical Trials – a Practical Guide." National Health Service Co-ordinating Centre for Research Methodology, url = https://researchonline.lshtm.ac.uk/id/eprint/4018500/.</bibtext> </blist> <blist> <bibl id="bib6" idref="ref48" type="bt">6</bibl> <bibtext> Collins Linda M., Schafer Joseph L., Kam Chi-Ming. 2001. " A Comparison of Inclusive and Restrictive Strategies in Modern Missing Data Procedures." Psychological Methods. 6(4):330‐51.</bibtext> </blist> <blist> <bibl id="bib7" idref="ref68" type="bt">7</bibl> <bibtext> Carpenter James R., Smuk Melanie. 2021. " Missing Data: A Statistical Framework for Practice." Biometrical Journal. 63:915‐47. https://doi.org/10.1002/bimj.202000196</bibtext> </blist> <blist> <bibl id="bib8" idref="ref4" type="bt">8</bibl> <bibtext> Carpenter James R., Kenward Michael G. 2013. Multiple Imputation and Its Application. New York: Wiley.</bibtext> </blist> <blist> <bibl id="bib9" idref="ref34" type="bt">9</bibl> <bibtext> Giusti Caterina, Little Roderick J. 2011. " A Sensitivity Analysis of Nonignorable Nonresponse to Income in a Survey with a Rotating Panel Design." Journal of Official Statistics. 27(2):211‐29.</bibtext> </blist> <blist> <bibtext> Horowitz Joel L., Manski Charles F. 1998. " Censoring of Outcomes and Regressors Due to Survey Nonresponse: Identification and Estimation Using Weights and Imputations." Journal of Econometrics. 84:37‐58.</bibtext> </blist> <blist> <bibtext> Hughes Rachael A., Heron Jon, Sterne Jonathan A.C., Tilling Kate. 2019. " Accounting for Missing Data in Statistical Analyses: Multiple Imputation is not Always the Answer." International Journal of Epidemiology. 48(4):1294‐304.</bibtext> </blist> <blist> <bibtext> Lee Katherine J., Tilling Kate M., Cornish Rosie P., Little Roderick J., Bell Melanie M., Goetghebeur Els, Hogan Joseph W., Carpenter James R. 2021. " Framework for the Treatment and Reporting of Missing Data in Observational Studies: The Treatment And Reporting of Missing Data in Observational Studies Framework." Journal of Clinical Epidemiology, 134: 79‐88.</bibtext> </blist> <blist> <bibtext> Little Roderick J. 1986. " Survey Nonresponse Adjustments." International Statistical Review, 54, 139‐157</bibtext> </blist> <blist> <bibtext> Little Roderick J. 1988. " Missing Data in Large Surveys (with Discussion)." Journal of Business and Economic Statistics. 6:287‐301.</bibtext> </blist> <blist> <bibtext> Little Roderick J. 1992. " Regression with Missing X's: A Review." Journal of the American Statistical Association. 87:1227‐37.</bibtext> </blist> <blist> <bibtext> Little Roderick J. 1995. " Modeling the Drop-Out Mechanism in Longitudinal Studies." Journal of the American Statistical Association. 90:1112‐21.</bibtext> </blist> <blist> <bibtext> Little Roderick J., Rubin Donald B. 2019. Statistical Analysis with Missing Data,</bibtext> </blist> <blist> <bibtext> Little Roderick J., Vartivarian Sonya. 2005. " Does Weighting for Nonresponse Increase the Variance of Survey Means? " Survey Methodology. 31:161‐8.</bibtext> </blist> <blist> <bibtext> Little Roderick J., Zhang Nanhua. 2011. " Subsample Ignorable Likelihood for Regression Analysis with Missing Data." Journal of the Royal Statistical Society: Series C: Applied Statistics. 60(4):591‐605.</bibtext> </blist> <blist> <bibtext> Muthen Linda K., Muthen Bengt. 2017. Mplus Version 8 User's Guide. Muthen and Muthen.</bibtext> </blist> <blist> <bibtext> National Research Council. 2010. The Prevention and Treatment of Missing Data in Clinical Trials. United States National Research Council.</bibtext> </blist> <blist> <bibtext> Quartagno Matteo, Carpenter James R. 2019. " Multiple Imputation for Discrete Data: Evaluation of the Joint Latent Normal Model." Biometrical Journal. 61(4):1003‐19.</bibtext> </blist> <blist> <bibtext> Raghunathan Trivellore E.2015. Missing Data Analysis in Practice. New York: Chapman and Hall / CRC.</bibtext> </blist> <blist> <bibtext> Raghunathan Trivellore E., Lepkowski James, Van Hoewyk John, Solenberger Peter W. 2001. " A Multivariate Technique for Multiply Imputing Missing Values Using a Sequence of Regression Models." Survey Methodology. 27(1):85‐95. For associated IVEWARE software see. <ulink href="http://www.isr.umich.edu/src/smp/ive/">http://www.isr.umich.edu/src/smp/ive/</ulink></bibtext> </blist> <blist> <bibtext> Robins James M., Rotnitzky Andrea. 1995. " Semiparametric Efficiency in Multivariate Regression Models with Missing Data." Journal of the American Statistical Association. 90:122‐9.</bibtext> </blist> <blist> <bibtext> Robins James M., Rotnitzky Andrea, Zhao Lue Ping. 1995. " Analysis of Semiparametric Regression Models for Repeated Outcomes in the Presence of Missing Data." Journal of the American Statistical Association. 90:106‐21.</bibtext> </blist> <blist> <bibtext> Rosenbaum Paul R., Rubin Donald B. 1983. " The Central Role of the Propensity Score in Observational Studies for Causal Effects." Biometrika. 70:41‐55.</bibtext> </blist> <blist> <bibtext> Rubin Donald B. 1976. " Inference and Missing Data (with Discussion)." Biometrika. 63:581‐92.</bibtext> </blist> <blist> <bibtext> Rubin Donald B.1987. Multiple Imputation for Nonresponse in Surveys. New York: Wiley.</bibtext> </blist> <blist> <bibtext> Rubin Donald B. 1996. " Multiple Imputation After 18 + Years (with Discussion)." Journal of the American Statistical Association. 91:473‐89.</bibtext> </blist> <blist> <bibtext> Rubin Donald B., Schenker Nathaniel. 1986. " Multiple Imputation for Interval Estimation from Simple Random Samples with Ignorable Nonresponse." Journal of the American Statistical Association. 81:366‐74.</bibtext> </blist> <blist> <bibtext> SAS. 2015. The MI Procedure. SAS/STAT 14.1 User's Guide, SAS Institute Inc., Cary, NC, USA.</bibtext> </blist> <blist> <bibtext> Schafer Joseph L.1997. Analysis of Incomplete Multivariate Data. New York: CRC Press.</bibtext> </blist> <blist> <bibtext> Schafer Joseph L. 1998. " Multiple Imputation: A Primer." Statistical Methods in Medical Research. 8:3‐15.</bibtext> </blist> <blist> <bibtext> Seaman Shaun R., White Ian R. 2011. " Review of Inverse Probability Weighting for Dealing with Missing Data." Statistical Methods in Medical Research. 22:278‐95.</bibtext> </blist> <blist> <bibtext> Shapira Marina, Iannelli Cristina, Croxford Linda, 2007. Youth Cohort Time Series for England, Wales and Scotland, 1984–2002. [Data Collection]. Scottish Centre for Social Research, University of Edinburgh, Centre for Educational Sociology, National Centre for Social Research, [original data producer(s)]. Scottish Centre for Social Research. SN: 5765, https://doi.org/10.5255/UKDA-SN-5765-1</bibtext> </blist> <blist> <bibtext> Su Yu-Sung, Gelman Andrew, Hill Jennifer, Yajima Masanao. 2011. " Multiple Imputation with Diagnostics (mi) in R: Opening Windows into the Black Box." Journal of Statistical Software. 45(2):1‐31.</bibtext> </blist> <blist> <bibtext> Tompsett Daniel M., Leacy Finbarr, Moreno-Betancur Margaritat, Heron Jon, White Ian R. 2018. " On the use of the not-at-Random Fully Conditional Specification (NARFCS) Procedure in Practice." Statistics in Medicine. 37(15):2338‐53.</bibtext> </blist> <blist> <bibtext> van Buuren Stef. 2018. Flexible Imputation of Missing Data.2nd Edition. FL: Boca-Raton: CRC/Chapman and Hall.</bibtext> </blist> <blist> <bibtext> van Buuren Stef, Groothuis-Oudshoorn Karen. 2011. " Multivariate Imputation by Chained Equations in R." Journal of Statistical Software. 45(4):1‐67. For associated software see <ulink href="http://www.multiple-imputation.com">http://www.multiple-imputation.com</ulink>.</bibtext> </blist> <blist> <bibtext> Von Hippel Paul T. (2007). " Regression with Missing Ys: An Improved Strategy for Analyzing Multiply Imputed Data." Sociological Methodology. 37(1):83‐117.</bibtext> </blist> <blist> <bibtext> White Ian R., Daniel Rhian, Royston Patrick. 2010. " Avoiding Bias Due to Perfect Prediction in Multiple Imputation of Incomplete Categorical Variables." Computational Statistics and Data Analysis. 54(10):2267‐75.</bibtext> </blist> <blist> <bibtext> White Ian R., Royston Patrick, Wood Angela M. 2011. " Multiple Imputation Using Chained Equations: Issues and Guidance for Practice." Statistics in Medicine. 30(4):377‐99.</bibtext> </blist> <blist> <bibtext> Zhang Guangyu, Little Roderick J. 2009. " Extensions of the Penalized Spline of Propensity Prediction Method of Imputation." Biometrics. 65(3):911‐8.</bibtext> </blist> <blist> <bibtext> Zhang Guangyu, Little Roderick J. 2011. " A Comparative Study of Doubly-Robust Estimators of the Mean with Missing Data." Journal of Statistical Computation and Simulation. 81(12):2039‐58.</bibtext> </blist> </ref> <ref id="AN0178879686-23"> <title> Footnotes </title> <blist> <bibtext> The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.</bibtext> </blist> <blist> <bibtext> The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the UK Medical Research Council, grant number MC_UU_00004/07.</bibtext> </blist> <blist> <bibtext> Roderick J. Little https://orcid.org/0000-0001-9878-6977</bibtext> </blist> </ref> <aug> <p>By Roderick J. Little; James R. Carpenter and Katherine J. Lee</p> <p>Reported by Author; Author; Author</p> <p></p> <p>Roderick J. Little is Richard D. Remington Distinguished University Professor of Biostatistics at the University of Michigan, where he also holds appointments in the Institute for Social Research and the Department of Statistics. His research focuses on methods for the analysis of data with missing values and model-based survey inference, and the application of statistics to diverse scientific areas, including medicine, demography, economics, psychiatry, aging and the environment.</p> <p>James R. Carpenter is Professor of Medical Statistics at the London School of Hygiene & Tropical Medicine, and Methodology Programme Leader at the MRC Clinical Trials unit at UCL, London UK. His research interests include methodology for multilevel modelling and missing data, with applications to observational data and clinical trials. He co-authored Multiple Imputation and its Application (Wiley, 2013) with Mike Kenward.</p> <p>Katherine J. Lee is Professor of Biostatistics at the Murdoch Children's Research Institute, Melbourne, and the University of Melbourne, Australia. Her research interests are in the method of multiple imputation for missing data and adaptive clinical trial designs.</p> </aug> <nolink nlid="nl1" bibid="bib16" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib38" firstref="ref2"></nolink> <nolink nlid="nl3" bibid="bib22" firstref="ref3"></nolink> <nolink nlid="nl4" bibid="bib32" firstref="ref5"></nolink> <nolink nlid="nl5" bibid="bib28" firstref="ref7"></nolink> <nolink nlid="nl6" bibid="bib10" firstref="ref8"></nolink> <nolink nlid="nl7" bibid="bib19" firstref="ref9"></nolink> <nolink nlid="nl8" bibid="bib24" firstref="ref10"></nolink> <nolink nlid="nl9" bibid="bib25" firstref="ref11"></nolink> <nolink nlid="nl10" bibid="bib39" firstref="ref12"></nolink> <nolink nlid="nl11" bibid="bib36" firstref="ref13"></nolink> <nolink nlid="nl12" bibid="bib23" firstref="ref14"></nolink> <nolink nlid="nl13" bibid="bib31" firstref="ref15"></nolink> <nolink nlid="nl14" bibid="bib20" firstref="ref16"></nolink> <nolink nlid="nl15" bibid="bib15" firstref="ref17"></nolink> <nolink nlid="nl16" bibid="bib11" firstref="ref18"></nolink> <nolink nlid="nl17" bibid="bib27" firstref="ref19"></nolink> <nolink nlid="nl18" bibid="bib34" firstref="ref21"></nolink> <nolink nlid="nl19" bibid="bib26" firstref="ref23"></nolink> <nolink nlid="nl20" bibid="bib30" firstref="ref27"></nolink> <nolink nlid="nl21" bibid="bib29" firstref="ref30"></nolink> <nolink nlid="nl22" bibid="bib33" firstref="ref31"></nolink> <nolink nlid="nl23" bibid="bib37" firstref="ref33"></nolink> <nolink nlid="nl24" bibid="bib42" firstref="ref39"></nolink> <nolink nlid="nl25" bibid="bib13" firstref="ref42"></nolink> <nolink nlid="nl26" bibid="bib21" firstref="ref45"></nolink> <nolink nlid="nl27" bibid="bib17" firstref="ref46"></nolink> <nolink nlid="nl28" bibid="bib14" firstref="ref50"></nolink> <nolink nlid="nl29" bibid="bib18" firstref="ref54"></nolink> <nolink nlid="nl30" bibid="bib40" firstref="ref56"></nolink> <nolink nlid="nl31" bibid="bib35" firstref="ref64"></nolink> <nolink nlid="nl32" bibid="bib41" firstref="ref65"></nolink> <nolink nlid="nl33" bibid="bib43" firstref="ref69"></nolink> <nolink nlid="nl34" bibid="bib44" firstref="ref70"></nolink> <nolink nlid="nl35" bibid="bib12" firstref="ref71"></nolink>
Header DbId: eric
DbLabel: ERIC
An: EJ1434927
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: A Comparison of Three Popular Methods for Handling Missing Data: Complete-Case Analysis, Inverse Probability Weighting, and Multiple Imputation
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Roderick+J%2E+Little%22">Roderick J. Little</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-9878-6977">0000-0001-9878-6977</externalLink>)<br /><searchLink fieldCode="AR" term="%22James+R%2E+Carpenter%22">James R. Carpenter</searchLink><br /><searchLink fieldCode="AR" term="%22Katherine+J%2E+Lee%22">Katherine J. Lee</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Sociological+Methods+%26+Research%22"><i>Sociological Methods & Research</i></searchLink>. 2024 53(3):1105-1135.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: SAGE Publications. 2455 Teller Road, Thousand Oaks, CA 91320. Tel: 800-818-7243; Tel: 805-499-9774; Fax: 800-583-2665; e-mail: journals@sagepub.com; Web site: https://sagepub.com
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 31
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2024
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Foreign+Countries%22">Foreign Countries</searchLink><br /><searchLink fieldCode="DE" term="%22Probability%22">Probability</searchLink><br /><searchLink fieldCode="DE" term="%22Robustness+%28Statistics%29%22">Robustness (Statistics)</searchLink><br /><searchLink fieldCode="DE" term="%22Responses%22">Responses</searchLink><br /><searchLink fieldCode="DE" term="%22Statistical+Inference%22">Statistical Inference</searchLink><br /><searchLink fieldCode="DE" term="%22Statistical+Distributions%22">Statistical Distributions</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluation+Methods%22">Evaluation Methods</searchLink><br /><searchLink fieldCode="DE" term="%22Comparative+Testing%22">Comparative Testing</searchLink>
– Name: Subject
  Label: Geographic Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22United+Kingdom+%28England%29%22">United Kingdom (England)</searchLink><br /><searchLink fieldCode="DE" term="%22United+Kingdom+%28Scotland%29%22">United Kingdom (Scotland)</searchLink><br /><searchLink fieldCode="DE" term="%22United+Kingdom+%28Wales%29%22">United Kingdom (Wales)</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1177/00491241221113873
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0049-1241<br />1552-8294
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Missing data are a pervasive problem in data analysis. Three common methods for addressing the problem are (a) complete-case analysis, where only units that are complete on the variables in an analysis are included; (b) weighting, where the complete cases are weighted by the inverse of an estimate of the probability of being complete; and (c) multiple imputation (MI), where missing values of the variables in the analysis are imputed as draws from their predictive distribution under an implicit or explicit statistical model, the imputation process is repeated to create multiple filled-in data sets, and analysis is carried out using simple MI combining rules. This article provides a non-technical discussion of the strengths and weakness of these approaches, and when each of the methods might be adopted over the others. The methods are illustrated on data from the Youth Cohort (Time) Series (YCS) for England, Wales and Scotland, 1984-2002.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2024
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1434927
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1434927
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1177/00491241221113873
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 31
        StartPage: 1105
    Subjects:
      – SubjectFull: Foreign Countries
        Type: general
      – SubjectFull: Probability
        Type: general
      – SubjectFull: Robustness (Statistics)
        Type: general
      – SubjectFull: Responses
        Type: general
      – SubjectFull: Statistical Inference
        Type: general
      – SubjectFull: Statistical Distributions
        Type: general
      – SubjectFull: Evaluation Methods
        Type: general
      – SubjectFull: Comparative Testing
        Type: general
      – SubjectFull: United Kingdom (England)
        Type: general
      – SubjectFull: United Kingdom (Scotland)
        Type: general
      – SubjectFull: United Kingdom (Wales)
        Type: general
    Titles:
      – TitleFull: A Comparison of Three Popular Methods for Handling Missing Data: Complete-Case Analysis, Inverse Probability Weighting, and Multiple Imputation
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Roderick J. Little
      – PersonEntity:
          Name:
            NameFull: James R. Carpenter
      – PersonEntity:
          Name:
            NameFull: Katherine J. Lee
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 08
              Type: published
              Y: 2024
          Identifiers:
            – Type: issn-print
              Value: 0049-1241
            – Type: issn-electronic
              Value: 1552-8294
          Numbering:
            – Type: volume
              Value: 53
            – Type: issue
              Value: 3
          Titles:
            – TitleFull: Sociological Methods & Research
              Type: main
ResultId 1