⚠ Official Notice: www.ijisrt.com is the official website of the International Journal of Innovative Science and Research Technology (IJISRT) Journal for research paper submission and publication. Please beware of fake or duplicate websites using the IJISRT name.



Fairness Audit and Debiasing of Machine Learning Models for Sepsis Mortality Prediction Across Demographic Subgroups: A Multi-Metric Study with Accuracy–Fairness Tradeoff Analysis


Authors : Francis Mawutor Amuyao; Isaac Tosin Adisa

Volume/Issue : Volume 11 - 2026, Issue 7 - July


Google Scholar : https://tinyurl.com/2ksef8uk

Scribd : https://tinyurl.com/6rzwbbpw

DOI : https://doi.org/10.38124/ijisrt/26jul1334

Note : A published paper may take 4-5 working days from the publication date to appear in PlumX Metrics, Semantic Scholar, and ResearchGate.


Abstract : Machine learning (ML) models for sepsis mortality prediction are increasingly deployed in intensive care units, yet performance across demographic subgroups remains poorly evaluated, and algorithmic disparities can directly exacerbate health inequities in critical care settings. We trained three ML models—logistic regression (LR), XGBoost, and a multilayer perceptron (MLP)—on a 10,000-patient cohort calibrated to published MIMIC-IV sepsis statistics, and performed a four-metric fairness audit (equalized odds difference [EOD], demographic parity difference, predictive parity gap, and subgroup calibration error) across race/ethnicity, sex, and insurance type. Per-group threshold optimisation was applied for debiasing, and an accuracy–fairness Pareto tradeoff was quantified.

Keywords : Algorithmic Fairness; Sepsis Prediction; Machine Learning; MIMIC-IV; Equalized Odds; SHAP; Health Disparities; Debiasing; ICU.

References :

  1. K.E. Rudd, S.C. Johnson, K.M. Agesa, et al., “Global, regional, and national sepsis incidence and mortality, 1990–2017: analysis for the Global Burden of Disease Study,” Lancet, vol. 395, no. 10219, pp. 200–211, 2020.
  2. C. Fleischmann-Struzek, L. Mellhammar, N. Rose, et al., “Incidence and mortality of hospital- and ICU-treated sepsis: results from an updated and expanded systematic review and meta-analysis,” Intensive Care Medicine, vol. 46, no. 8, pp. 1552–1562, 2020.
  3. A. Kumar, D. Roberts, K.E. Wood, et al., “Duration of hypotension before initiation of effective antimicrobial therapy is the critical determinant of survival in human septic shock,” Critical Care Medicine, vol. 34, no. 6, pp. 1589–1596, 2006.
  4. Y. Su, C. Guo, S. Zhou, C. Li, N. Ding, “Early predicting 30-day mortality in sepsis in MIMIC-IV by an artificial neural networks model,” Internal and Emergency Medicine, vol. 20, no. 3, pp. 909–918, 2024.
  5. C. Meng, L. Trinh, N. Xu, J. Enouen, Y. Liu, “Interpretability and fairness evaluation of deep learning models on MIMIC-IV dataset,” Scientific Reports, vol. 12, no. 1, p. 7166, 2022.
  6. Z. Obermeyer, B. Powers, C. Vogeli, S. Mullainathan, “Dissecting racial bias in an algorithm used to manage the health of populations,” Science, vol. 366, no. 6464, pp. 447–453, 2019.
  7. F.B. Mayr, V.B. Talisa, V. Balakumar, et al., “Proportion and cost of unplanned 30-day readmissions after sepsis compared with other medical conditions,” JAMA, vol. 317, no. 5, pp. 530–531, 2017.
  8. A.M. Esper, M. Moss, G.S. Martin, “The effect of diabetes mellitus on organ dysfunction with sepsis: an epidemiological study,” Critical Care, vol. 13, no. 1, p. R18, 2009.
  9. A.E.W. Johnson, L. Bulgarelli, L. Shen, et al., “MIMIC-IV, a freely accessible electronic health record dataset,” Scientific Data, vol. 10, no. 1, p. 1, 2023.
  10. M. Singer, C.S. Deutschman, C.W. Seymour, et al., “The Third International Consensus Definitions for Sepsis and Septic Shock (Sepsis-3),” JAMA, vol. 315, no. 8, pp. 801–810, 2016.
  11. GUIDE Expert Panel, “Guidance for unbiased predictive information for healthcare decision-making and equity (GUIDE): considerations when race may be a prognostic factor,” npj Digital Medicine, 2024.
  12. A. Chouldechova, “Fair prediction with disparate impact: a study of bias in recidivism prediction instruments,” Big Data, vol. 5, no. 2, pp. 153–163, 2017.
  13. M. Hardt, E. Price, N. Srebro, “Equality of opportunity in supervised learning,” Advances in Neural Information Processing Systems, vol. 29, 2016.
  14. J. Wawira Gichoya, L.G. McCoy, L.A. Celi, M. Ghassemi, “Equity in essence: a call for operationalising fairness in machine learning for healthcare,” BMJ Health & Care Informatics, vol. 28, no. 1, p. e100289, 2021.
  15. Y. Chen, Y. Liu, L. Yao, “Explainable AI for fair sepsis mortality predictive model,” in Proc. 22nd Int. Conf. on Artificial Intelligence in Medicine (AIME 2024), 2024.
  16. J. Huang, G. Galal, M. Etemadi, M. Vaidyanathan, “Evaluation and mitigation of racial bias in clinical machine learning models: scoping review,” JMIR Medical Informatics, vol. 10, no. 5, p. e36388, 2022.
  17. N. Mehrabi, F. Morstatter, N. Saxena, K. Lerman, A. Galstyan, “A survey on bias and fairness in machine learning,” ACM Computing Surveys, vol. 54, no. 6, pp. 1–35, 2021.
  18. X. Liu, D.L. Armaignac, C. Becker, et al., “Monitoring fairness in machine learning models that predict patient mortality in the ICU,” arXiv preprint, 2024.
  19. S. Verma, J. Rubin, “Fairness definitions explained,” in Proc. Int. Workshop on Software Fairness (FairWare), 2018, pp. 1–7.
  20. Scoping Review Group on Algorithmic Bias, “Sociodemographic bias in clinical machine learning models: a scoping review,” Journal of Clinical Epidemiology, 2024.

Machine learning (ML) models for sepsis mortality prediction are increasingly deployed in intensive care units, yet performance across demographic subgroups remains poorly evaluated, and algorithmic disparities can directly exacerbate health inequities in critical care settings. We trained three ML models—logistic regression (LR), XGBoost, and a multilayer perceptron (MLP)—on a 10,000-patient cohort calibrated to published MIMIC-IV sepsis statistics, and performed a four-metric fairness audit (equalized odds difference [EOD], demographic parity difference, predictive parity gap, and subgroup calibration error) across race/ethnicity, sex, and insurance type. Per-group threshold optimisation was applied for debiasing, and an accuracy–fairness Pareto tradeoff was quantified.

Keywords : Algorithmic Fairness; Sepsis Prediction; Machine Learning; MIMIC-IV; Equalized Odds; SHAP; Health Disparities; Debiasing; ICU.

Paper Submission Last Date
30 - September - 2026

SUBMIT YOUR PAPER CALL FOR PAPERS
Video Explanation for Published paper

Never miss an update from Papermashup

Get notified about the latest tutorials and downloads.

Subscribe by Email

Get alerts directly into your inbox after each post and stay updated.
Subscribe
OR

Subscribe by RSS

Add our RSS to your feedreader to get regular updates from us.
Subscribe