2013Unpublished venueRequires access

Hot Deck Propensity Score Imputation For Missing Values

Benjamin Mayer

Open publisher page 5 citations

Abstract

The adequate handling of missing data in medical research still constitutes a major problem of statistical data analysis. Although different imputation strategies and methods have been developed in recent years, none of them can be used unhesitatingly in arbitrary missing data situations. A commonly observed drawback of the established imputation methods is their sensitivity against the number of missing values in the analysis set, the underlying missing data mechanism and scale level of the incomplete variable. The extent of this sensitivity is indeed different for individual methods. A novel approach for the imputation of missing data was proposed here which especially does not depend on the scale level of the variable that is affected by missing values. This approach in a multiple imputation setting was based on the calculation of propensity scores and their usage in creating adequate imputation values, resting upon a Hot Deck principle. In the course of a real-data based simulation study, the proposed method was compared to a standard multiple imputation approach and a complete case analysis strategy. The results showed that there was a dependency of imputation’s goodness from the assumed missing data proportion, too. However, the developed imputation method generated consistent estimations of the width of confidence intervals and could be applied for different kinds of incomplete variables straightforwardly, which is not possible with other imputation approaches that easily in general.

About this research paper

What this paper is about

The adequate handling of missing data in medical research still constitutes a major problem of statistical data analysis. Although different imputation strategies and methods have been developed in recent years, none of them can be used unhesitatingly in arbitrary missing data situations. A commonly observed drawback of the established imputation methods is their sensitivity against the number of missing values in the analysis set, the underlying missing data mechanism and scale level of the incomplete variable. The extent of this sensitivity is indeed different for individual methods. A novel approach for the imputation of missing data was proposed here which especially does not depend on the scale level of the variable that is affected by missing values. This approach in a multiple imputation setting was based on the calculation of propensity scores and their usage in creating adequate imputation values, resting upon a Hot Deck principle. In the course of a real-data based simulation study, the proposed method was compared to a standard multiple imputation approach and a complete case analysis strategy. The results showed that there was a dependency of imputation’s goodness from the assumed missing data proportion, too. However, the developed imputation method generated consistent estimations of the width of confidence intervals and could be applied for different kinds of incomplete variables straightforwardly, which is not possible with other imputation approaches that easily in general.

Why it matters

OpenAlex reports 5 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

The adequate handling of missing data in medical research still constitutes a major problem of statistical data analysis. Although different imputation strategies and methods have been developed in recent years, none of them can be used unhesitatingly in arbitrary missing data situations. A commonly observed drawback of the established imputation methods is their sensitivity against the number of missing values in the analysis set, the underlying missing data mechanism and scale level of the incomplete variable. The extent of this sensitivity is indeed different for individual methods. A novel approach for the imputation of missing data was proposed here which especially does not depend on the scale level of the variable that is affected by missing values. This approach in a multiple imputation setting was based on the calculation of propensity scores and their usage in creating adequate imputation values, resting upon a Hot Deck principle. In the course of a real-data based simulation study, the proposed method was compared to a standard multiple imputation approach and a complete case analysis strategy. The results showed that there was a dependency of imputation’s goodness from the assumed missing data proportion, too. However, the developed imputation method generated consistent estimations of the width of confidence intervals and could be applied for different kinds of incomplete variables straightforwardly, which is not possible with other imputation approaches that easily in general.

Key concepts: Imputation (statistics), Missing data, Statistics, Computer science, Data mining, Mathematics

Related papers

Back to paper searchBrowse research topicsOriginal source
Hot Deck Propensity Score Imputation For Missing Values — Research Paper | ScholarLens