Models for Injury Count Data in the U.S. National Health Interview Survey
Jin Peng, Tianmeng Lyu, Junxin Shi, Haikady N. Nagaraja, Huiyun Xiang
Journal of Scientific Research and Reports · pp. 2286–2302 · Published 17 Jul 2014
10.9734/JSRR/2014/9490Abstract
Aims: To examine the best count data model for injury data in the National Health Interview Survey (NHIS). To compare the best count data model with traditional logistic regression model in analyzing injury data in NHIS. Data Source: 2006-2010 medically consulted non-occupational injury data from National Health Interview Survey (NHIS). Methodology: Six count data models (Poisson, negative binomial (NB), zero-inflated Poisson (ZIP), zero-inflated NB (ZINB), hurdle Poisson (HP), and hurdle NB (HNB)) were compared using Likelihood Ratio (LR) test and Vuong test. Injury count was used as the dependent variable in count data models. Independent variables included age, gender, marital status, race, education, poverty status, disability status and medical insurance coverage status. Dichotomized injury count was used as the dependent variable in logistic regression model. The same independent variables used in count data models were included in logistic regression model. The model fit of logistic regression was examined by Hosmer and Lemeshow goodness of fit test. Results: Among 248,850 participants aged 18-64, 98.37% have no medically consulted non-occupational injuries, 1.55% have 1 medically consulted non-occupational injury, 0.07% have 2 or more medically consulted non-occupational injuries. Zero-inflated negative binomial (ZINB) model offered the best fit. Logistic regression model provided a good fit but resulted in different estimates from ZINB model. Conclusion: Zero-inflated negative binomial (ZINB) model demonstrated the potential to be the best model for injury count data with excess zeros. Given the infrequent occurrence of multiple injuries in our data, the logistic regression model is appropriate for assessing injury burden and identifying injury risks. However, for more frequently-occurring injuries (e.g. sports injuries), logistic regression may undercount the total number of injuries and result in biased estimates. The evaluation procedure and model selection criteria presented in this paper provide a useful approach to modeling injury count data with excess zeros.
Cited by 2
Afet SÖZEN ÖZDEN, Elvan HAYAT · Pamukkale University Journal of Social Sciences Institute · 2022
Hiroko Mori, Joshua Wu, Motomu Ibaraki · International Journal of Environmental Research and Public Health · 2018
Related research
Article metrics
Real usage data collected on this platform.
0
Page views
0
PDF downloads
0
Outbound clicks
2
Citations
Views by country
Approximate, from request IP at view time — not citizenship or institution. Countries with fewer than 5 views are grouped as "Other".
No views recorded yet.
Traffic sources
Referring site, by host.
No traffic recorded yet.
Views and downloads exclude known bots/crawlers. Citations combines this platform's own DOI-resolved index with each external source's own reported total — see Cited by above for individually listed citing works. Last refreshed 0 seconds ago.