Multiple Imputation to Estimate Hierarchical Models from Data Missing at Random: Latent Covariates, Random Coefficients, and Statistical Interactions

Saved in:
Bibliographic Details
Title: Multiple Imputation to Estimate Hierarchical Models from Data Missing at Random: Latent Covariates, Random Coefficients, and Statistical Interactions
Language: English
Authors: Yongyun Shin (ORCID 0000-0003-0654-2620), Stephen W. Raudenbush
Source: Grantee Submission. 2025.
Peer Reviewed: Y
Page Count: 52
Publication Date: 2025
Sponsoring Agency: Institute of Education Sciences (ED)
Contract Number: R305D210022
Document Type: Reports - Research
Descriptors: Hierarchical Linear Modeling, Maximum Likelihood Statistics, Sampling, Error of Measurement, Algorithms, Equations (Mathematics), Computation, Bayesian Statistics
DOI: 10.3102/10769986251385210
Abstract: Consider the conventional multilevel model Y=C[gamma]+Zu+e where [gamma] represents fixed effects and (u,e) are multivariate normal random effects. The continuous outcomes Y and covariates C are fully observed with a subset Z of C. The parameters are [theta]=([gamma],var(u),var(e)). Dempster, Rubin and Tsutakawa (1981) framed the estimation as a missing data problem, where (Y,u) are the complete data and the random effects u are conceived as missing data. Viewed in this way, the Expectation-Maximization (EM) algorithm has proven to be a natural and popular approach to estimation. However, when C is partially observed or subject to measurement error, it is natural to formulate a multilevel model for C that includes random effects, [nu]. In this article, we extend this thinking to allow estimation of the joint distribution of data Y=(Y,C)=(U[subscript o],U[subscript m]) and random effects b=(u,v) from observed data Y[subscript o]=(Y[subscript o],C[subscript o]) and to generate multiple imputations of missing data (Y[subscript m],b) based on the estimated distribution under the assumption that the data Y are missing at random. This approach contributes to the literature on multiple imputation in three ways: (a) it allows random effects [nu] to be conceived as latent covariates, thus addressing measurement errors of C; (b) it allows non-linearities, including random coefficients, interaction effects, and other polynomial effects involving partially observed covariates; (c) it imputes (Y[subscript m],b) using two-step importance sampling. In these cases, the joint distribution of Y is not analytically tractable even if the analytic multilevel model of interest to the analyst follows a multivariate normal distribution. We prove that our method of maximizing the likelihood and imputing missing data ensures compatibility of the non-normal joint distribution with the analytic normal theory multilevel model via provisionally known random effects. We present and evaluate a sufficient condition under which the produced imputations are compatible with the analytic model. [This paper will be published in the "Journal of Educational and Behavioral Statistics."]
Abstractor: As Provided
IES Funded: Yes
Entry Date: 2025
Accession Number: ED677391
Database: ERIC
FullText Text:
  Availability: 0
Header DbId: eric
DbLabel: ERIC
An: ED677391
AccessLevel: 3
PubType: Report
PubTypeId: report
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Multiple Imputation to Estimate Hierarchical Models from Data Missing at Random: Latent Covariates, Random Coefficients, and Statistical Interactions
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Yongyun+Shin%22">Yongyun Shin</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-0654-2620">0000-0003-0654-2620</externalLink>)<br /><searchLink fieldCode="AR" term="%22Stephen+W%2E+Raudenbush%22">Stephen W. Raudenbush</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Grantee+Submission%22"><i>Grantee Submission</i></searchLink>. 2025.
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 52
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2025
– Name: SourceSuprt
  Label: Sponsoring Agency
  Group: SrcSuprt
  Data: Institute of Education Sciences (ED)
– Name: NumberContract
  Label: Contract Number
  Group: NumCntrct
  Data: R305D210022
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Hierarchical+Linear+Modeling%22">Hierarchical Linear Modeling</searchLink><br /><searchLink fieldCode="DE" term="%22Maximum+Likelihood+Statistics%22">Maximum Likelihood Statistics</searchLink><br /><searchLink fieldCode="DE" term="%22Sampling%22">Sampling</searchLink><br /><searchLink fieldCode="DE" term="%22Error+of+Measurement%22">Error of Measurement</searchLink><br /><searchLink fieldCode="DE" term="%22Algorithms%22">Algorithms</searchLink><br /><searchLink fieldCode="DE" term="%22Equations+%28Mathematics%29%22">Equations (Mathematics)</searchLink><br /><searchLink fieldCode="DE" term="%22Computation%22">Computation</searchLink><br /><searchLink fieldCode="DE" term="%22Bayesian+Statistics%22">Bayesian Statistics</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.3102/10769986251385210
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Consider the conventional multilevel model Y=C[gamma]+Zu+e where [gamma] represents fixed effects and (u,e) are multivariate normal random effects. The continuous outcomes Y and covariates C are fully observed with a subset Z of C. The parameters are [theta]=([gamma],var(u),var(e)). Dempster, Rubin and Tsutakawa (1981) framed the estimation as a missing data problem, where (Y,u) are the complete data and the random effects u are conceived as missing data. Viewed in this way, the Expectation-Maximization (EM) algorithm has proven to be a natural and popular approach to estimation. However, when C is partially observed or subject to measurement error, it is natural to formulate a multilevel model for C that includes random effects, [nu]. In this article, we extend this thinking to allow estimation of the joint distribution of data Y=(Y,C)=(U[subscript o],U[subscript m]) and random effects b=(u,v) from observed data Y[subscript o]=(Y[subscript o],C[subscript o]) and to generate multiple imputations of missing data (Y[subscript m],b) based on the estimated distribution under the assumption that the data Y are missing at random. This approach contributes to the literature on multiple imputation in three ways: (a) it allows random effects [nu] to be conceived as latent covariates, thus addressing measurement errors of C; (b) it allows non-linearities, including random coefficients, interaction effects, and other polynomial effects involving partially observed covariates; (c) it imputes (Y[subscript m],b) using two-step importance sampling. In these cases, the joint distribution of Y is not analytically tractable even if the analytic multilevel model of interest to the analyst follows a multivariate normal distribution. We prove that our method of maximizing the likelihood and imputing missing data ensures compatibility of the non-normal joint distribution with the analytic normal theory multilevel model via provisionally known random effects. We present and evaluate a sufficient condition under which the produced imputations are compatible with the analytic model. [This paper will be published in the "Journal of Educational and Behavioral Statistics."]
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: CodeSource
  Label: IES Funded
  Group: SrcInfo
  Data: Yes
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2025
– Name: AN
  Label: Accession Number
  Group: ID
  Data: ED677391
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=ED677391
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.3102/10769986251385210
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 52
    Subjects:
      – SubjectFull: Hierarchical Linear Modeling
        Type: general
      – SubjectFull: Maximum Likelihood Statistics
        Type: general
      – SubjectFull: Sampling
        Type: general
      – SubjectFull: Error of Measurement
        Type: general
      – SubjectFull: Algorithms
        Type: general
      – SubjectFull: Equations (Mathematics)
        Type: general
      – SubjectFull: Computation
        Type: general
      – SubjectFull: Bayesian Statistics
        Type: general
    Titles:
      – TitleFull: Multiple Imputation to Estimate Hierarchical Models from Data Missing at Random: Latent Covariates, Random Coefficients, and Statistical Interactions
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Yongyun Shin
      – PersonEntity:
          Name:
            NameFull: Stephen W. Raudenbush
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 10
              M: 12
              Type: published
              Y: 2025
          Titles:
            – TitleFull: Grantee Submission
              Type: main
ResultId 1