Interpretable Machine Learning Models for Predicting Hypertension Risk

Authors

  • Joseph Bamikole Olojido Department of Computer Science, Rufus Giwa Polytecnic Owo, Ondo State, Nigeria. Author
  • Fayowole Ayanfe Aworetan Department of Computer Science, Rufus Giwa Polytecnic Owo, Ondo State, Nigeria Author
  • Oluwatoyin James Ojajuni Department of Computer Science, Rufus Giwa Polytecnic Owo, Ondo State, Nigeria Author

Keywords:

Hypertension, Machine Learning; Explainable AI (XAI), Risk Stratification, XGBoost, SHAP

Abstract

Hypertension ranks among the most significant contributors to cardiovascular disease, stroke, and premature mortality worldwide. Because the condition typically produces no noticeable symptoms until serious complications have already developed, identifying at-risk individuals early is critical. Traditional statistical approaches often struggle to capture the complex, nonlinear interactions among demographic, clinical, and behavioural variables, and although many machine learning methods outperform these classical models, their opaque, "black-box" nature limits clinician trust and slows adoption in practice. This study develops an interpretable machine learning framework for predicting hypertension risk using Framingham Heart Study records comprising 4,240 participants and 13 attributes. Following data cleaning, imputation, standardisation, and feature selection, four ensemble algorithms — Decision Tree, Random Forest, Gradient Boosting, and XGBoost — were trained and optimised through GridSearchCV on an 80:20 stratified split. XGBoost produced the strongest overall results, achieving a ROC AUC of 0.9582, accuracy of 0.8939, precision of 0.8078, recall of 0.8631, and an F1-score of 0.8346. Explanations generated through SHAP, LIME, and decision-tree visualisation consistently identified systolic and diastolic blood pressure as the leading predictors, a pattern broadly consistent with established clinical diagnostic criteria. Grouping predicted risk scores into Low, Medium, and High categories allowed the framework to correctly flag 86.31% of individuals in the High Risk group. Taken together, the results show that strong predictive performance and model transparency can be achieved simultaneously, providing a practical and reproducible tool for early hypertension risk assessment in clinical practice

References

Agrawal, T. (2020). Hyperparameter Optimization Using Scikit-Learn, 31-51. https://doi.org/10.1007/978-1-4842-6579-6_2.

Choi, S., Oh, M., Lee, D., Jee, S., & Jeon, J. (2025). Invasive and non-invasive variables prediction models for cardiovascular disease-specific mortality between machine learning vs. traditional statistics. Scientific Reports, 15. https://doi.org/10.1038/s41598-025-18853-7.

Donmez, T., & Kutlu, M. (2025). Explainable quantum-enhanced machine learning for hypertension prediction. The European Physical Journal Special Topics. https://doi.org/10.1140/epjs/s11734-025-01629-5.

Fang, M., Chen, Y., Xue, R., Wang, H., Chakraborty, N., Su, T., & Dai, Y. (2021). A hybrid machine learning approach for hypertension risk prediction. Neural Computing and Applications, 35, 14487-14497. https://doi.org/10.1007/s00521-021-06060-0.

Kanagachidambaresan, G., & Vinoothna, M. (2021). Visualizations. Programming with TensorFlow. https://doi.org/10.1007/978-3-030-57077-4_3.

Kaya, B. (2025). Hypertension measurement methods and differential diagnosis. Scientific Reports in Medicine. https://doi.org/10.37609/srinmed.48.

Kim, S., & Yu, J. (2023). Stratified importance sampling for a Bernoulli mixture model of portfolio credit risk. Annals of Operations Research, 322, 819-849. https://doi.org/10.1007/s10479-023-05174-z.

Partho, P., Bhowmik, P., & Nasir, M. (2025). An Interpretable Ensemble Framework Towards Efficient Hypertension Risk Assessment. 2025 2nd International Conference on Next-Generation Computing, IoT and Machine Learning (NCIM), 1-6. https://doi.org/10.1109/ncim65934.2025.11159863.

Ponce-Bobadilla, A., Schmitt, V., Maier, C., Mensing, S., & Stodtmann, S. (2024). Practical guide to SHAP analysis: Explaining supervised machine learning model predictions in drug development. Clinical and Translational Science, 17. https://doi.org/10.1111/cts.70056.

Rehman, S., Rehman, E., Mumtaz, A., & Zhang, J. (2022). Cardiovascular Disease Mortality and Potential Risk Factor in China: A Multi-Dimensional Assessment by a Grey Relational Approach. International Journal of Public Health, 67. https://doi.org/10.3389/ijph.2022.1604599.

Rolon-Mérette, D., Ross, M., Rolon-Merette, T., & Church, K. (2020). Introduction to Anaconda and Python: Installation and setup. The Quantitative Methods for Psychology. https://doi.org/10.20982/tqmp.16.5.s003.

Santhoshini, S., & Ravimaran, S. (2025). DATAPREPMATE: A PYTHON TOOL FOR AUTOMATED DATA CLEANING USING AI TECHNIQUES. International Journal of Engineering Applied Sciences and Technology. https://doi.org/10.33564/ijeast.2025.v10i03.005.

Vrigazova, B. (2021). The Proportion for Splitting Data into Training and Test Set for the Bootstrap in Classification Problems. Business Systems Research Journal, 12, 228 - 242. https://doi.org/10.2478/bsrj-2021-0015.

Vu, T. (2022). Jupyter Notebook. . https://doi.org/10.55277/researchhub.611vklh0.

Waburi, E., Muriithi, D., & Sundays, E. (2025). Integrating Explainable Machine Learning Models for Early Detection of Hypertension: A Transparent Approach to AI-Driven Healthcare. American Journal of Artificial Intelligence. https://doi.org/10.11648/j.ajai.20250902.17.

Downloads

Published

2026-04-30