|
Communications on Applied Electronics
Foundation of Computer Science (FCS), NY, USA
|
| Volume 8 - Issue 2 |
| Published: September 2026 |
| Authors: Oluwaponmile H. Adefuwa, Olatunde D. Akinrolabu, Ayodeji O. Ayeni |
10.5120/cae1e52adc932f2
|
Oluwaponmile H. Adefuwa, Olatunde D. Akinrolabu, Ayodeji O. Ayeni . An Inclusive Diabetes Detection System using Machine Learning and Explainable Artificial Intelligence. Communications on Applied Electronics. 8, 2 (September 2026), 9-21. DOI=10.5120/cae1e52adc932f2
@article{ 10.5120/cae1e52adc932f2,
author = { Oluwaponmile H. Adefuwa,Olatunde D. Akinrolabu,Ayodeji O. Ayeni },
title = { An Inclusive Diabetes Detection System using Machine Learning and Explainable Artificial Intelligence },
journal = { Communications on Applied Electronics },
year = { 2026 },
volume = { 8 },
number = { 2 },
pages = { 9-21 },
doi = { 10.5120/cae1e52adc932f2 },
publisher = { Foundation of Computer Science (FCS), NY, USA }
}
%0 Journal Article
%D 2026
%A Oluwaponmile H. Adefuwa
%A Olatunde D. Akinrolabu
%A Ayodeji O. Ayeni
%T An Inclusive Diabetes Detection System using Machine Learning and Explainable Artificial Intelligence%T
%J Communications on Applied Electronics
%V 8
%N 2
%P 9-21
%R 10.5120/cae1e52adc932f2
%I Foundation of Computer Science (FCS), NY, USA
Diabetes mellitus is a chronic metabolic disorder affecting over 537 million adults globally, with prevalence expected to reach 783 million by 2045. Even with advances in diagnostic technology, equitable and early detection remains a challenge, particularly across demographically diverse populations. This study presents the design and development of an inclusive diabetes detection system that combines machine-learning classification, Local Interpretable Model-agnostic Explanations (LIME)-based explainability, and interactive Streamlit web deployment. A publicly available synthetic dataset of 100,000 patient records was preprocessed through encoding, Min-Max normalisation, and SMOTE-based class balancing. Four machine learning models (XGBoost, Random Forest, Naive Bayes, and Long Short-Term Memory (LSTM)) were trained and evaluated using accuracy, precision, recall, F1-score, and AUC-ROC. Hyperparameter tuning was conducted using Grid Search with 5-fold cross-validation. XGBoost achieved the best performance with 91.99% accuracy, precision of 1.0000, recall of 0.8665, F1-score of 0.9285, and AUC-ROC of 0.9438, clearly exceeding the research target of 70% accuracy. Demographic subgroup analysis across gender, ethnicity, age group, and income level confirmed equitable performance, with F1-score variation below 0.004 across most demographic dimensions. LIME generated clinically meaningful individual-level explanations, consistently identifying HbA1c and fasting glucose as the dominant predictive features. A broader evaluation was done, the trained XGBoost model was further tested on the Pima Indians Diabetes Dataset, a widely used benchmark with only 8 features, achieving 72.53% accuracy and AUC-ROC of 0.7969 using only 6 of 38 model features, demonstrating generalisation capability beyond the original training distribution. The complete system was deployed as a browser-based Streamlit application enabling real-time prediction and transparent explanation for clinical and non-clinical users.