Early detection of diabetes mellitus, a chronic condition affecting a large share of the world\'s population, plays a major role in keeping long-term complications and healthcare costs down. Here, eight commonly used supervised classifiers — namely Logistic Regression, Decision Tree, Random Forest, SVM, KNN, Naïve Bayes, AdaBoost, and Gradient Boosting — are put through a head-to-head comparison for the diabetes-prediction task. Training and evaluation were carried out on 768 patient records from the Kaggle-hosted Pima Indians Diabetes Dataset, spanning eight clinical attributes (Pregnancies, Glucose, Blood Pressure, Skin Thickness, Insulin, BMI, Diabetes Pedigree Function, Age). Accuracy, precision, recall, F1 score, AUC-ROC, and 10-fold cross-validation together formed the basis for judging each model. Among the eight, Gradient Boosting came out on top with 75.32% accuracy and an AUC-ROC of 0.8257, narrowly ahead of AdaBoost and Random Forest, both at 74.68% accuracy and 0.8253 AUC-ROC. When Random Forest was used to rank feature importance, Glucose, BMI, and Age surfaced as the strongest predictors of diabetes onset. Overall, boosting-based ensembles appear well suited to combining strong accuracy with clinical interpretability, pointing toward practical, low-cost, real-time screening tools.
Introduction
The text presents a comparative study of eight machine-learning algorithms for predicting diabetes using the Pima Indians Diabetes Dataset. The main motivation is that early diabetes detection can help prevent complications such as cardiovascular disease, kidney damage, retinal problems, and nerve damage, particularly in areas where conventional laboratory-based diagnosis may be difficult to access.
Objective
The study has two main goals:
Compare the performance of eight supervised machine-learning classifiers using the same preprocessing and validation procedure.
Identify the most important clinical features for diabetes prediction while considering the balance between model accuracy and interpretability.
Dataset and Methodology
The study uses 768 patient records containing eight clinical attributes:
Pregnancies
Glucose
Blood Pressure
Skin Thickness
Insulin
BMI
Diabetes Pedigree Function
Age
Zero values in physiologically impossible fields such as glucose, blood pressure, insulin, and BMI were treated as missing values. The data was standardized using z-score normalization, divided into training and test sets using an 80/20 stratified split, and evaluated using 10-fold cross-validation.
The eight models compared were:
Logistic Regression
Decision Tree
Random Forest
Support Vector Machine (RBF)
K-Nearest Neighbors
Gaussian Naïve Bayes
AdaBoost
Gradient Boosting
Performance was measured using Accuracy, Precision, Recall, F1-Score, and AUC-ROC.
Results
Gradient Boosting performed best overall on the held-out test set, achieving:
Accuracy: 75.32%
Precision: 66.00%
Recall: 61.11%
F1-score: 63.46%
AUC-ROC: 0.8257
AdaBoost and Random Forest followed with 74.68% accuracy.
An important finding was that different models offered different advantages:
Decision Tree achieved the highest precision (71.43%) but had low recall (46.30%), meaning it missed many diabetic cases.
Naïve Bayes achieved the highest recall (62.96%), making it potentially useful for screening where identifying as many at-risk patients as possible is important.
Logistic Regression had the highest 10-fold cross-validation accuracy (78.02%), despite achieving only 70.78% accuracy on the test set, suggesting relatively good generalization.
Random Forest provided useful feature-importance information while maintaining competitive predictive performance.
Feature Importance
Random Forest analysis identified the most influential clinical variables as:
Glucose — 29.46%
BMI — 16.86%
Age — 11.53%
Diabetes Pedigree Function — 11.11%
Insulin — 9.76%
Blood Pressure — 7.49%
Pregnancies — 6.99%
Skin Thickness — 6.80%
Glucose was by far the most important predictor, followed by BMI and age. These findings are consistent with established clinical understanding of diabetes risk.
Conclusion
Across this comparison of eight machine learning algorithms on the Pima Indians Diabetes Dataset, run through one consistent evaluation pipeline, Gradient Boosting emerged ahead of the pack on both accuracy and AUC-ROC, with AdaBoost and Random Forest not far off; Glucose, BMI, and Age stood out as the most clinically relevant predictors per the Random Forest feature-importance analysis. Together, these findings make a reasonable case for boosting-based ensembles as a foundation for interpretable, accurate diabetes screening — the kind that could realistically fit into low-cost clinical decision-support systems or mobile health apps.
A few directions could take this further: larger and more demographically varied datasets, SMOTE-based class balancing [10] to handle imbalance, hybrid CNN-LSTM architectures [13,14,15] for capturing more complex, non-linear relationships among clinical variables, and explainability tools like SHAP or LIME for instance-level interpretability along the lines of Chao and Liu\'s work [9]. Beyond that, building this framework into an actual mobile or web-based screening tool would be a natural next step toward real healthcare impact.
References
[1] World Health Organization. (2024). Diabetes Fact Sheet. WHO, Geneva.
[2] Pima Indians Diabetes Dataset. Kaggle / UCI Machine Learning Repository, National Institute of Diabetes and Digestive and Kidney Diseases.
[3] Sisodia, D., & Sisodia, D. S. (2018). Prediction of diabetes using classification algorithms. Procedia Computer Science, 132, 1578–1585.
[4] Khanam, J. J., & Foo, S. Y. (2021). A comparison of machine learning algorithms for diabetes prediction. ICT Express, 7(4), 432–439.
[5] Mujumdar, A., & Vaidehi, V. (2019). Diabetes prediction using machine learning algorithms. Procedia Computer Science, 165, 292–299.
[6] Kopitar, L., Kocbek, P., Cilar, L., Sheikh, A., & Stiglic, G. (2020). Early detection of type 2 diabetes mellitus using machine learning-based prediction models. Scientific Reports, 10, 11981.
[7] Rastogi, R., & Bansal, M. (2023). Diabetes prediction model using data mining techniques. Measurement: Sensors, 25, 100605.
[8] Application of Machine Learning Models for Early Detection and Accurate Classification of Type 2 Diabetes. (2023). Diagnostics, 13(14), 2383.
[9] Chao, Z., & Liu, J. (2026). Machine Learning-Based Diabetes Prediction Model Construction and Analysis: A Case Study of the Kaggle Diabetes Dataset. In Optimization and Data Science in Industrial Engineering (ODSIE 2025), Communications in Computer and Information Science, vol. 2854. Springer, Cham.
[10] Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic Minority Over-sampling Technique. Journal of Artificial Intelligence Research, 16, 321–357.
[11] Akula, R., Nguyen, N., & Garibay, I. (2019). Supervised Machine Learning based Ensemble Model for Accurate Prediction of Type 2 Diabetes. IEEE SoutheastCon 2019, 1–9.
[12] Patro, S., Sharma, V., & Vishwakarma, S. (2023). A comparative study of machine learning models for diabetes prediction. BMC Bioinformatics, 24, 372.
[13] Aslan, M. F., & Sabanci, K. (2023). A novel proposal for deep learning-based diabetes prediction: Converting clinical data to image data. Diagnostics, 13(4), 796.
[14] Fan, Y. (2025). Diabetes diagnosis using a hybrid CNN-LSTM-MLP ensemble. Scientific Reports, 15, 26765.
[15] Enhanced diabetes prediction using pre-trained CNNs, LSTM, and conditional GAN on transformed numerical data. (2026). Scientific Reports.
[16] Farnoosh, R., et al. (2025). DiabetesXpertNet: An innovative attention-based CNN for accurate type 2 diabetes prediction. PLOS ONE.
[17] Noh, M. J., & Kim, Y. S. (2025). Diabetes Prediction Through Linkage of Causal Discovery and Inference Model with Machine Learning Models. Biomedicines, 13(1), 124.
[18] Breiman, L. (2001). Random Forests. Machine Learning, 45(1), 5–32.
[19] Cortes, C., & Vapnik, V. (1995). Support-vector networks. Machine Learning, 20(3), 273–297.
[20] Freund, Y., & Schapire, R. E. (1997). A decision-theoretic generalization of on-line learning and an application to boosting. Journal of Computer and System Sciences, 55(1), 119–139.
[21] Pedregosa, F., et al. (2011). Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research, 12, 2825–2830.