2016•Unpublished venueRequires access

Distributed Multinomial Regression

Matt A. Taddy

Open publisher page 77 citations

Abstract

faculty.chicagobooth.edu/matt.taddy This article introduces a model-based approach to distributed computing for multinomial logis-tic regression. We treat counts for each response category as independent Poisson regressions via plug-in estimates for fixed effects shared across categories. The work is driven by the high-dimensional-response multinomial models that arise in analysis of a large number of random counts. Our archetypal applications are in text analysis, where documents are tokenized and the token counts are modeled as arising from a multinomial dependent upon document attributes. We estimate such models for a publicly available dataset of reviews from Yelp, with text re-gressed onto a large set of explanatory variables (user, business, and rating information). The fitted models serve as a basis for exploring the connection between words and variables of inter-est (e.g., star rating), for reducing dimension into supervised factor scores, and for prediction. ar

About this research paper

What this paper is about

faculty.chicagobooth.edu/matt.taddy This article introduces a model-based approach to distributed computing for multinomial logis-tic regression. We treat counts for each response category as independent Poisson regressions via plug-in estimates for fixed effects shared across categories. The work is driven by the high-dimensional-response multinomial models that arise in analysis of a large number of random counts. Our archetypal applications are in text analysis, where documents are tokenized and the token counts are modeled as arising from a multinomial dependent upon document attributes. We estimate such models for a publicly available dataset of reviews from Yelp, with text re-gressed onto a large set of explanatory variables (user, business, and rating information). The fitted models serve as a basis for exploring the connection between words and variables of inter-est (e.g., star rating), for reducing dimension into supervised factor scores, and for prediction. ar

Why it matters

OpenAlex reports 77 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

faculty.chicagobooth.edu/matt.taddy This article introduces a model-based approach to distributed computing for multinomial logis-tic regression. We treat counts for each response category as independent Poisson regressions via plug-in estimates for fixed effects shared across categories. The work is driven by the high-dimensional-response multinomial models that arise in analysis of a large number of random counts. Our archetypal applications are in text analysis, where documents are tokenized and the token counts are modeled as arising from a multinomial dependent upon document attributes. We estimate such models for a publicly available dataset of reviews from Yelp, with text re-gressed onto a large set of explanatory variables (user, business, and rating information). The fitted models serve as a basis for exploring the connection between words and variables of inter-est (e.g., star rating), for reducing dimension into supervised factor scores, and for prediction. ar

Key concepts: Multinomial logistic regression, Multinomial distribution, Computer science, Softmax function, Dimension (graph theory), Set (abstract data type), Poisson regression, Data set

Related papers

Back to paper searchBrowse research topicsOriginal source
Distributed Multinomial Regression — Research Paper | ScholarLens