Externally-derived algorithms to impute missing data may be biased even under MCAR
Abstract
Administrative claims data often lack key information (e.g. laboratory results) needed for pharmacoepidemiologic analysis. Linkage to other data sources (e.g. electronic health records), can address this limitation but introduces data that may be missing not at random (MNAR). Alternatively, claims-based prediction models developed in external data can be used to ‘impute’ missing variables. We sought to evaluate the assumptions of these approaches with a Monte Carlo simulation.
A continuous variable, Y, was generated using two binary variables W (observed) and U (unobserved) in both primary (M) and external (N) datasets under the same data generating process. A linear model for Y using only W was fit in the external data and applied to the primary data to obtain fused predictions via G-computation. We induced 30% missingness for Y in the primary data: (1) randomly (i.e., MCAR); (2) dependent on W (i.e., MAR); and (3) dependent on U and W (i.e., MNAR). To examine the effect of population differences we varied the distributions of W and U across the primary and external datasets. We compared the complete case analysis means and fused means to the true mean across mechanisms. Bias, empirical standard error and root mean squared error were assessed.
When the primary and external data sets are fully exchangeable, the fused mean was unbiased under MCAR, MAR and MNAR. This is preserved when the data sources differ only on observed W. Where data sources differ by unobserved predictors, fused means were biased under all three missingness mechanisms.
External data fusion can mitigate bias due to MNAR, under alternative assumptions of (conditional) exchangeability of primary and external data on observed predictors. Externally-derived algorithms can introduce bias if these assumptions are not met or with model misspecification. These transportability assumptions are implicitly made when using externally-derived algorithms to impute missing data in epidemiology studies.

