Distributed Multinomial Regression
Matt A. Taddy
Abstract
Matt A. Taddy
Abstract
faculty.chicagobooth.edu/matt.taddy This article introduces a model-based approach to distributed computing for multinomial logis-tic regression. We treat counts for each response category as independent Poisson regressions via plug-in estimates for fixed effects shared across categories. The work is driven by the high-dimensional-response multinomial models that arise in analysis of a large number of random counts. Our archetypal applications are in text analysis, where documents are tokenized and the token counts are modeled as arising from a multinomial dependent upon document attributes. We estimate such models for a publicly available dataset of reviews from Yelp, with text re-gressed onto a large set of explanatory variables (user, business, and rating information). The fitted models serve as a basis for exploring the connection between words and variables of inter-est (e.g., star rating), for reducing dimension into supervised factor scores, and for prediction. ar
OpenAlex reports 77 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
faculty.chicagobooth.edu/matt.taddy This article introduces a model-based approach to distributed computing for multinomial logis-tic regression. We treat counts for each response category as independent Poisson regressions via plug-in estimates for fixed effects shared across categories. The work is driven by the high-dimensional-response multinomial models that arise in analysis of a large number of random counts. Our archetypal applications are in text analysis, where documents are tokenized and the token counts are modeled as arising from a multinomial dependent upon document attributes. We estimate such models for a publicly available dataset of reviews from Yelp, with text re-gressed onto a large set of explanatory variables (user, business, and rating information). The fitted models serve as a basis for exploring the connection between words and variables of inter-est (e.g., star rating), for reducing dimension into supervised factor scores, and for prediction. ar
Key concepts: Multinomial logistic regression, Multinomial distribution, Computer science, Softmax function, Dimension (graph theory), Set (abstract data type), Poisson regression, Data set