Skip to content
Research Article Open access

Novel Data Mining Techniques for Incomplete Clinical Data in Diabetes Management

Herbert F. Jelinek, Andrew Yatsko, Andrew Stranieri, Sitalakshmi Venkatraman

Current Journal of Applied Science and Technology · pp. 4591–4606 · Published 16 Sep 2014

10.9734/BJAST/2014/11744

Abstract

An important part of health care involves upkeep and interpretation of medical databases containing patient records for clinical decision making, diagnosis and follow-up treatment. Missing clinical entries make it difficult to apply data mining algorithms for clinical decision support. This study demonstrates that higher predictive accuracy is possible using conventional data mining algorithms if missing values are dealt with appropriately. We propose a novel algorithm using a convolution of sub-problems to stage a super problem, where classes are defined by Cartesian Product of class values of the underlying problems, and Incomplete Information Dismissal and Data Completion techniques are applied for reducing features and imputing missing values. Predictive accuracies using Decision Branch, Nearest Neighborhood and Naïve Bayesian classifiers were compared to predict diabetes, cardiovascular disease and hypertension. Data is derived from Diabetes Screening Complications Research Initiative (DiScRi) conducted at a regional Australian university involving more than 2400 patient records with more than one hundred clinical risk factors (attributes). The results show substantial improvements in the accuracy achieved with each classifier for an effective diagnosis of diabetes, cardiovascular disease and hypertension as compared to those achieved without substituting missing values. The gain in improvement is 7% for diabetes, 21% for cardiovascular disease and 24% for hypertension, and our integrated novel approach has resulted in more than 90% accuracy for the diagnosis of any of the three conditions. This work advances data mining research towards achieving an integrated and holistic management of diabetes.

Data mining missing value imputation diabetes management classifiers diagnosis accuracy.

Cited by 6

Data analytics identify glycated haemoglobin co-markers for type 2 diabetes mellitus diagnosis

Herbert F. Jelinek, Andrew Stranieri, Andrew Yatsko · Computers in Biology and Medicine · 2016

Detecting depression severity using weighted random forest and oxidative stress biomarkers

Mariam Bader, Moustafa Abdelwanis, Maher Maalouf · Scientific Reports · 2024

Research and Citation Analysis of Data Mining Technology Based on Bayes Algorithm

Mingyang Liu, Ming Qu, Bin Zhao · Mobile Networks and Applications · 2016

Identifying the Relation between Fasting Blood Glucose and Glycosylated Haemoglobin Levels in Greek Diabetic Patients

M Stamouli, A Pouliakis, A Mourtzikou · Annals of Cytology and Pathology · 2016

Diagnostic with incomplete nominal/discrete data

Herbert F. Jelinek, Andrew Yatsko, Andrew Stranieri · Artificial Intelligence Research · 2015

Article metrics

Real usage data collected on this platform.

0

Page views

0

PDF downloads

0

Outbound clicks

6

Citations

Views by country

Approximate, from request IP at view time — not citizenship or institution. Countries with fewer than 5 views are grouped as "Other".

No views recorded yet.

Traffic sources

Referring site, by host.

No traffic recorded yet.

Views and downloads exclude known bots/crawlers. Citations combines this platform's own DOI-resolved index with each external source's own reported total — see Cited by above for individually listed citing works. Last refreshed 0 seconds ago.