Contents
What is multiple imputation used for?
Multiple imputation is a general approach to the problem of missing data that is available in several commonly used statistical packages. It aims to allow for the uncertainty about the missing data by creating several different plausible imputed data sets and appropriately combining results obtained from each of them.
How multiple imputation makes a difference?
Multiple imputation is substantially more efficient than listwise deletion because it (1) util- izes rather than discards data in incomplete observations and (2) allows analysts to incorporate extra information into the imputation model by including variables that are not in the analysis (auxiliary variables).
Should you impute outcome variables?
Outcome variables must not be imputed. Predictor variables must not be imputed. Multiple imputation must not be used because you will end up with several different outcomes of your statistical analysis.
How is multiple imputation used in prognostic variable selection?
Multiple imputation (MI) accounts for imputation uncertainty that allows for adequate statistical testing. We developed and tested a methodology combining MI with bootstrapping techniques for studying prognostic variable selection.
How is multiple imputation used in statistical inference?
Multiple imputation (MI) accounts for the uncertainty caused by the missing data, and when properly done, MI provides correct statistical inferences [ 6 ]. MI replaces each missing values by two or more imputations. The spread between the imputed values reflects the uncertainty about the missing data.
How is the spread between the imputations explained?
MI replaces each missing values by two or more imputations. The spread between the imputed values reflects the uncertainty about the missing data. MI proceeds by applying the complete-data analysis to each imputed data set, followed by pooling the results into a final estimate.
How does multiple imputation affect the inclusion frequency?
We found that the effect of imputation variation on the inclusion frequency was larger than the effect of sampling variation. When MI and bootstrapping were combined at the range of 0% (full model) to 90% of variable selection, bootstrap corrected c-index values of 0.70 to 0.71 and slope values of 0.64 to 0.86 were found.