A Practical Approach in Modelling Count Data
W. Y. Wan Fairos, Wan Fairos Wan Yaacob, Mohamad Alias Lazim, Yap Bee Wah
Abstract
W. Y. Wan Fairos, Wan Fairos Wan Yaacob, Mohamad Alias Lazim, Yap Bee Wah
Abstract
Attempts to model count data have varied from the use of least square regression techniques to methods involving exponential distribution families including Poisson and Negative Binomial (NB) models. Given the nature of discrete, non-negative integer value of count data, the Poisson distribution has been verified to be the best distribution to describe count data. Though the Poisson regression model works well for count data, it still suffers one potential problem. This relates to the assumption of the equality of variance and mean in the use of Poisson model. When the assumption is violated, overdispersion occurs and the standard errors estimated will be biased which will then lead to incorrect test statistics. Much of the early work done to correct for over dispersion was to use the Poisson Quasi Likelihood or alternatively the Negative Binomial model. In real life application, count data often exhibits overdispersion and excess zeros. While negative binomial model can deal with overdispersion, the Zero-inflated Poisson and Negative Binomial Regression model can be used to handle excess zeros in their own way. Thus, this paper provides a road map of the practical approach for modeling count data and an illustration using doctor visits data.
OpenAlex reports 20 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Attempts to model count data have varied from the use of least square regression techniques to methods involving exponential distribution families including Poisson and Negative Binomial (NB) models. Given the nature of discrete, non-negative integer value of count data, the Poisson distribution has been verified to be the best distribution to describe count data. Though the Poisson regression model works well for count data, it still suffers one potential problem. This relates to the assumption of the equality of variance and mean in the use of Poisson model. When the assumption is violated, overdispersion occurs and the standard errors estimated will be biased which will then lead to incorrect test statistics. Much of the early work done to correct for over dispersion was to use the Poisson Quasi Likelihood or alternatively the Negative Binomial model. In real life application, count data often exhibits overdispersion and excess zeros. While negative binomial model can deal with overdispersion, the Zero-inflated Poisson and Negative Binomial Regression model can be used to handle excess zeros in their own way. Thus, this paper provides a road map of the practical approach for modeling count data and an illustration using doctor visits data.
Key concepts: Overdispersion, Count data, Quasi-likelihood, Negative binomial distribution, Poisson distribution, Poisson regression, Statistics, Mathematics