Authors :
Aminu Usman Jibril; A. Senthil Kumar
Volume/Issue :
Volume 11 - 2026, Issue 7 - July
Google Scholar :
https://tinyurl.com/34uzsr7e
Scribd :
https://tinyurl.com/ytrezhaf
DOI :
https://doi.org/10.38124/ijisrt/26jul1418
Note : A published paper may take 4-5
working days from the publication date to appear in PlumX Metrics, Semantic Scholar, and
ResearchGate.
Abstract :
Missing values and significant class imbalance are common characteristics of healthcare datasets which
significantly impair predictive model performance and reduce their dependability in clinical decision-making. Creation of
reliable and broadly applicable healthcare prediction system depends on addressing this issues As to improve overall data
quality, this study suggests an integrated data preparation system that integrates cluster-aware oversampling methods with
Generative Adversarial Imputation Networks (GAIN). By using adversarial training to understand intricate underlying
data distributions GAIN model are used to estimate missing values while maintaining significant statistical correlations
between variables. Simultaneously, hybrid SMOTE-ENN method are used to remove ambiguous and noisy data and
efficiently handle class imbalance. Real-world diabetic readmission dataset are used to assess suggested methodology, and
show notable gains in data completeness distribution preservation, and prediction performance. Significant improvements
in accuracy, recall, and F1-score are revealed by experimental data, suggesting improved capacity to detect high-risk
individuals. As compared to traditional methods incorporation of sophisticated preprocessing technique enhances model
resilience and generalisation. This results highlight significance of integrating class balancing technique and intelligent
imputation into single framework. Overall, study emphasises how important sophisticated preprocessing are to enhancing
clinical applicability, robustness and dependability of predictive healthcare analytics system.
Keywords :
GAN; GAIN; Data Imputation; SMOTE-ENN; Class Imbalance; Predictive Healthcare Analytics; Diabetes Readmission.
References :
- Goodfellow, I., Pouget‑Abadie, J., Mirza, M., Xu, B., Warde‑Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2014). Generative adversarial nets. Advances in Neural Information Processing System, 27, 2672–2680. https://papers.nips.cc/paper/5423-generative-adversarial-nets
- Yoon, J., Jordon, J., & van der Schaar, M. (2018). GAIN: Missing data imputation using generative adversarial nets. Proceedings of 35th International Conference on Machine Learning 5689–5698. http://proceedings.mlr.press/v80/yoon18a.html
- Goodfellow, I., et. al., (2020). Generative adversarial networks. Communications of ACM, 63(11), 139–144. https://doi.org/10.1145/3422622
- Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 16, 321–357. https://doi.org/10.1613/jair.953
- Batista, G. E., Prati, R. C., & Monard, M. C. (2004). study of behavior of several methods for balancing machine learning training data. ACM SIGKDD Explorations 6(1), 20–29. https://doi.org/10.1145/1007730.1007735
- He, H., & Garcia, E. A. (2009). Learning from imbalanced data. IEEE Transactions on Knowledge and Data Engineering 21(9), 1263–1284. https://doi.org/10.1109/TKDE.2008.239
- King G., & Zeng L. (2001). Logistic regression in rare events data. Political Analysis 9(2), 137–163. https://doi.org/10.1093/oxfordjournals.pan.a004868
- Chen, T., & Guestrin, C. (2016). XGBoost: scalable tree boosting system. Proceedings of 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining 785–794. https://doi.org/10.1145/2939672.2939785
- Steinhubl, S. R., Muse, E. D., & Topol, E. J. (2015). emerging field of mobile health. Science Translational Medicine, 7(283), 283rv3. https://doi.org/10.1126/scitranslmed.aaa3487
- Esteva, A., Robicquet A., Ramsundar, B., Kuleshov, V., DePristo, M., Chou, K., Cui, C., Corrado, G., Thrun, S., & Dean, J. (2019). guide to deep learning in healthcare. Nature Medicine, 25, 24–29. https://doi.org/10.1038/s41591-018-0316-z
- Rajkomar, A., Oren, E., Chen, K., Dai, A. M., Hajaj, N., Hardt M., Liu, P. J., Liu, X., Marcus J., Sun, M., Sundberg P., Yee, H., Zhang K., Zhang Y., Flores G., Ledsam, J. R., et. al., (2018). Scalable and accurate deep learning with electronic health records. npj Digital Medicine, 1(18). https://doi.org/10.1038/s41746-018-0029-1
- Bates J., Saria, S., Ohno-Machado, L., Shah, A., & Escobar, G. (2014). Big data in health care. Health Affairs 33(7), 1123–1131. https://doi.org/10.1377/hlthaff.2014.0147
- Creswell, A., White, T., Dumoulin, V., Arulkumaran, K., Sengupta, B., & Bharath, A. A. (2018). Generative adversarial networks: An overview. IEEE Signal Processing Magazine, 35(1), 53–65. https://doi.org/10.1109/MSP.2017.2765202
- Breiman, L. (2001). Random forests. Machine Learning 45, 5–32. https://doi.org/10.1023/A:1010933404324
- Friedman, J. H. (2001). Greedy function approximation: gradient boosting machine. Annals of Statistics 29(5), 1189–1232. https://doi.org/10.1214/aos/1013203451
- LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521, 436–444. https://doi.org/10.1038/nature14539
- Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient estimation of word representations in vector space. ICLR Workshop. https://arxiv.org/abs/1301.3781
- Kingma, D. P., & Ba, J. (2015). Adam: method for stochastic optimization. ICLR. https://arxiv.org/abs/1412.6980
- Van Buuren, S., & Groothuis-Oudshoorn, K. (2011). MICE: Multivariate imputation by chained equations. Journal of Statistical Software, 45(3), 1–67. https://www.jstatsoft.org/article/view/v045i03
- Little, R. J., & Rubin, D. B. (2019). Statistical analysis with missing data (3rd ed.). Wiley. https://www.wiley.com/en-us/Statistical+Analysis+with+Missing+Data%2C+3rd+Edition-p-9781119407605
- Dong G., & Peng H. (2013). Principled missing data methods for researchers. Springer. https://doi.org/10.1007/978-1-4614-6841-3
- Schafer, J. L. (1997). Analysis of incomplete multivariate data. Chapman & Hall. https://doi.org/10.1007/978-1-4757-3542-0
- Fawcett T. (2006). An introduction to ROC analysis. Pattern Recognition Letters 27(8), 861–874. https://doi.org/10.1016/j.patrec.2005.10.010
- Chicco, D., & Jurman, G. (2020). advantages of Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. Bioinformatics 36(20), 517–520. https://doi.org/10.1093/bioinformatics/btaa106
- Sculley, D., Holt G., Golovin, D., Davydov, E., Phillips T., Ebner, D., Chaudhary, V., Young M., Crespo, J. F., & Dennison, D. (2015). Hidden technical debt in machine learning system. Advances in Neural Information Processing System. https://papers.nips.cc/paper/5656-hidden-technical-debt-in-machine-learning-system
- Strack, B., DeShazo, J., Gennings C., Olmo, J. L., Ventura, S., Cios K., & Re, C. (2014). Impact of HbA1c measurement on hospital readmission rates. BioMed Research International, 2014, Article ID 781670. https://doi.org/10.1155/2014/781670
- He, H., Bai, Y., Garcia, E., & Li, S. (2008). ADASYN: Adaptive synthetic sampling approach for imbalanced learning. IEEE International Joint Conference on Neural Networks 1322–1328. https://doi.org/10.1109/IJCNN.2008.4633969
- Krawczyk, J. (2016). Learning from imbalanced data: open challenges and future directions. Progress in Artificial Intelligence, 5, 221–232. https://doi.org/10.1007/s13748-016-0094-0
- Iqbal, Z., Rafique, A., Qaisar, S., et. al., (2025). Advancements and challenges in development of generative adversarial network (GANs) for deep learning. SN Computer Science. https://doi.org/10.1007/s44354-025-00007-w
- Ou, H., Yao, Y., & He, Y. (2024). Missing Data Imputation Method Combining Random Forest and Generative Adversarial Imputation Network. Sensors 24(4), 1112. https://doi.org/10.3390/s24041112
Missing values and significant class imbalance are common characteristics of healthcare datasets which
significantly impair predictive model performance and reduce their dependability in clinical decision-making. Creation of
reliable and broadly applicable healthcare prediction system depends on addressing this issues As to improve overall data
quality, this study suggests an integrated data preparation system that integrates cluster-aware oversampling methods with
Generative Adversarial Imputation Networks (GAIN). By using adversarial training to understand intricate underlying
data distributions GAIN model are used to estimate missing values while maintaining significant statistical correlations
between variables. Simultaneously, hybrid SMOTE-ENN method are used to remove ambiguous and noisy data and
efficiently handle class imbalance. Real-world diabetic readmission dataset are used to assess suggested methodology, and
show notable gains in data completeness distribution preservation, and prediction performance. Significant improvements
in accuracy, recall, and F1-score are revealed by experimental data, suggesting improved capacity to detect high-risk
individuals. As compared to traditional methods incorporation of sophisticated preprocessing technique enhances model
resilience and generalisation. This results highlight significance of integrating class balancing technique and intelligent
imputation into single framework. Overall, study emphasises how important sophisticated preprocessing are to enhancing
clinical applicability, robustness and dependability of predictive healthcare analytics system.
Keywords :
GAN; GAIN; Data Imputation; SMOTE-ENN; Class Imbalance; Predictive Healthcare Analytics; Diabetes Readmission.