2007•Unpublished venueRequires access

Variable Screening for Multinomial Logistic Regression on Very Large Data Sets as Applied to Direct Response Modeling

John A. Cherrie

Open publisher page 3 citations

Abstract

Beginning with version 8.2 SAS supports multinomial logistic regression as part of PROC LOGISTIC. As a direct response modeling provider, we have found that multinomial logistic regression models often provide us with our best solutions. However, using multinomial logistic regression presents some challenges. We are often faced with very large sample sizes and a large number of candidate predictor variables. The sheer sizes of our data sets make variable selection and scoring significant issues. The functionality that SAS has built into PROC LOGISTIC to handle variable selection and scoring is not feasible for the size of data sets that we use. I will present how we have addressed the issue of variable screening in our modeling environment. I will provide a comparison of gains charts that show that the multinomial logistic model is preferable to other solutions. I will discuss and share bits of code on how we utilize univariate screening, resampling methods, and clustering for variable reduction. Discussing the issues regarding creating efficient production scoring modules utilizing multinomial logistic regression is beyond the scope this presentation. Upon request I will supply code for a simple macro that creates a text file version of a coefficient file that can be incorporated into an efficient production scoring environment.

About this research paper

What this paper is about

Beginning with version 8.2 SAS supports multinomial logistic regression as part of PROC LOGISTIC. As a direct response modeling provider, we have found that multinomial logistic regression models often provide us with our best solutions. However, using multinomial logistic regression presents some challenges. We are often faced with very large sample sizes and a large number of candidate predictor variables. The sheer sizes of our data sets make variable selection and scoring significant issues. The functionality that SAS has built into PROC LOGISTIC to handle variable selection and scoring is not feasible for the size of data sets that we use. I will present how we have addressed the issue of variable screening in our modeling environment. I will provide a comparison of gains charts that show that the multinomial logistic model is preferable to other solutions. I will discuss and share bits of code on how we utilize univariate screening, resampling methods, and clustering for variable reduction. Discussing the issues regarding creating efficient production scoring modules utilizing multinomial logistic regression is beyond the scope this presentation. Upon request I will supply code for a simple macro that creates a text file version of a coefficient file that can be incorporated into an efficient production scoring environment.

Why it matters

OpenAlex reports 3 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Beginning with version 8.2 SAS supports multinomial logistic regression as part of PROC LOGISTIC. As a direct response modeling provider, we have found that multinomial logistic regression models often provide us with our best solutions. However, using multinomial logistic regression presents some challenges. We are often faced with very large sample sizes and a large number of candidate predictor variables. The sheer sizes of our data sets make variable selection and scoring significant issues. The functionality that SAS has built into PROC LOGISTIC to handle variable selection and scoring is not feasible for the size of data sets that we use. I will present how we have addressed the issue of variable screening in our modeling environment. I will provide a comparison of gains charts that show that the multinomial logistic model is preferable to other solutions. I will discuss and share bits of code on how we utilize univariate screening, resampling methods, and clustering for variable reduction. Discussing the issues regarding creating efficient production scoring modules utilizing multinomial logistic regression is beyond the scope this presentation. Upon request I will supply code for a simple macro that creates a text file version of a coefficient file that can be incorporated into an efficient production scoring environment.

Key concepts: Multinomial logistic regression, Logistic regression, Computer science, Multinomial distribution, Feature selection, Variable (mathematics), Statistics, Data mining

Related papers

Back to paper searchBrowse research topicsOriginal source
Variable Screening for Multinomial Logistic Regression on Very Large Data Sets as Applied to Direct Response Modeling — Research Paper | ScholarLens