Focused Information Criteria and Model Averaging for the Cox Hazard Regression Model
Nils Lid Hjort, Gerda Claeskens
Abstract
Nils Lid Hjort, Gerda Claeskens
Abstract
This article is concerned with variable selection methods for the Cox proportional hazards regression model. Including excessive covariates causes extra variability and inflated confidence intervals for regression parameters; thus regimes for discarding the less informative ones are needed. Our framework has p covariates designated as “protected,” while variables from a further set of q covariates are examined for possible inclusion or exclusion. We develop a focused information criterion (FIC) that for given interest parameter finds the best subset of covariates. Thus the FIC might find that the best model for predicting median survival time is different than the best model for estimating survival probabilities, and the best overall model for analyzing men's survival might not be the same as the best overall model for analyzing women's survival. Methodology is also developed for model averaging, wherein the final estimate of a quantity is a weighted average of estimates computed for a range of submodels. Our methods are illustrated in simulations and for a survival study of Danish skin cancer patients.
OpenAlex reports 124 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
This article is concerned with variable selection methods for the Cox proportional hazards regression model. Including excessive covariates causes extra variability and inflated confidence intervals for regression parameters; thus regimes for discarding the less informative ones are needed. Our framework has p covariates designated as “protected,” while variables from a further set of q covariates are examined for possible inclusion or exclusion. We develop a focused information criterion (FIC) that for given interest parameter finds the best subset of covariates. Thus the FIC might find that the best model for predicting median survival time is different than the best model for estimating survival probabilities, and the best overall model for analyzing men's survival might not be the same as the best overall model for analyzing women's survival. Methodology is also developed for model averaging, wherein the final estimate of a quantity is a weighted average of estimates computed for a range of submodels. Our methods are illustrated in simulations and for a survival study of Danish skin cancer patients.
Key concepts: Covariate, Proportional hazards model, Statistics, Regression analysis, Econometrics, Mathematics, Model selection, Confidence interval