ANALYSIS OF MORTALITY RISK FACTORS AND ENSEMBLE PREDICTION MODELING FOR DKD PATIENTS IN ICU.

Saved in:
Bibliographic Details
Title: ANALYSIS OF MORTALITY RISK FACTORS AND ENSEMBLE PREDICTION MODELING FOR DKD PATIENTS IN ICU.
Authors: LIN, YANTING1,2 (AUTHOR) 37220222203689@stu.xmu.edu.cn, DENG, YIWEN2 (AUTHOR) 37220222203576@stu.xmu.edu.cn, LIU, HAOLIN2 (AUTHOR) 37220222203695@stu.xmu.edu.cn, WANG, YONGCHAO3 (AUTHOR) 382294032@qq.com, ZHANG, SUZHEN1 (AUTHOR) 18959227290@163.com
Source: Journal of Mechanics in Medicine & Biology. Apr2026, Vol. 26 Issue 3, p1-24. 24p.
Subjects: Mortality risk factors, Diabetic nephropathies, Feature extraction, Intensive care units, Ensemble learning, Critical care medicine, Machine learning, Severity of illness index
Abstract: This study, based on the Medical Information Mart for Intensive Care IV (MIMIC-IV) database, aimed to investigate the risk factors associated with in-hospital mortality among patients with diabetic kidney disease (DKD) admitted to the intensive care unit (ICU), and to develop a high-precision predictive model. A total of 434 DKD patient records were meticulously selected through strict inclusion criteria. A filter-based feature engineering strategy was employed. Initially, multi-source heterogeneous data tables were integrated using Structured Query Language (SQL) queries in combination with Python for preprocessing. Missing values in numerical variables were imputed using Bayesian regression or median substitution, and all continuous variables were standardized. Subsequently, key prognostic features were identified through univariate statistical testing (t-test or χ 2 test) followed by multivariate logistic regression analysis. The significant predictors of mortality included demographic characteristics (e.g., age), laboratory parameters [e.g., the Minimum White Blood Cell count (WBC_min), the Maximum Blood Urea Nitrogen (BUN_max), the Mean Sodium (Na_mean)], the Sequential Organ Failure Assessment (SOFA) score, clinical interventions (e.g., mechanical ventilation, vasopressor use), and comorbidities [e.g., hypertension, chronic obstructive pulmonary disease (COPD)]. In terms of model construction strategy, this study first evaluated the performance of eight individual machine learning models, including Decision Tree (DT), Random Forest (RF), K-Nearest Neighbors (KNN), Support Vector Machine (SVM), Adaptive Boosting (AdaBoost), eXtreme Gradient Boosting (XGBoost), Naïve Bayes (NB), and Multi-Layer Perceptron (MLP). An innovative pairwise ensemble strategy was applied, resulting in 28 unique model combinations. Several combinations demonstrated significantly improved performance compared with single models. Notably, the SVM   +   RF combination achieved the highest performance with an Area Under the Curve (AUC) of 0.865, representing a 0.9 percentage point improvement over the best-performing single model, RF (AUC = 0. 8 5 6). Inspired by these findings, a comprehensive ensemble framework incorporating all eight models was subsequently developed. The final ensemble prediction framework adopted a dynamic weighting strategy based on validation-set AUCs, with hyperparameter optimization conducted using Optuna via Tree-structured Parzen Estimator (TPE) sampling over 100 iterations. The resulting model weights were as follows: SVM (0.86), XGBoost (0.58), RF (0.69), KNN (0.52), NB(0.40), MLP (0.40), AdaBoost (0.01), and DT (0.002). After weighted probability fusion, the ensemble model achieved an AUC of 0.8676 on the test set, representing a 0.26 percentage point improvement over the best-performing pairwise combination (SVM   +   RF). Feature importance analysis identified the average SOFA score (sofa_score_avg), the mean sodium level (Sodium_mean), the minimum creatinine level (Creatinine_min), and the mean potassium level (Potassium_mean) as the most critical features in predicting in-hospital mortality risk for DKD patients. Additionally, higher SOFA scores, lower WBC_min, and BUN_max were associated with an increased risk of mortality, consistent with the results from multivariable logistic regression analysis. This study effectively utilized filter-based feature engineering to identify key prognostic indicators and, inspired by pairwise model-combination results, developed a dynamic weighted ensemble framework that significantly enhanced the identification of high-risk patients. [ABSTRACT FROM AUTHOR]
Copyright of Journal of Mechanics in Medicine & Biology is the property of World Scientific Publishing Company and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Description
Abstract:This study, based on the Medical Information Mart for Intensive Care IV (MIMIC-IV) database, aimed to investigate the risk factors associated with in-hospital mortality among patients with diabetic kidney disease (DKD) admitted to the intensive care unit (ICU), and to develop a high-precision predictive model. A total of 434 DKD patient records were meticulously selected through strict inclusion criteria. A filter-based feature engineering strategy was employed. Initially, multi-source heterogeneous data tables were integrated using Structured Query Language (SQL) queries in combination with Python for preprocessing. Missing values in numerical variables were imputed using Bayesian regression or median substitution, and all continuous variables were standardized. Subsequently, key prognostic features were identified through univariate statistical testing (t-test or χ 2 test) followed by multivariate logistic regression analysis. The significant predictors of mortality included demographic characteristics (e.g., age), laboratory parameters [e.g., the Minimum White Blood Cell count (WBC_min), the Maximum Blood Urea Nitrogen (BUN_max), the Mean Sodium (Na_mean)], the Sequential Organ Failure Assessment (SOFA) score, clinical interventions (e.g., mechanical ventilation, vasopressor use), and comorbidities [e.g., hypertension, chronic obstructive pulmonary disease (COPD)]. In terms of model construction strategy, this study first evaluated the performance of eight individual machine learning models, including Decision Tree (DT), Random Forest (RF), K-Nearest Neighbors (KNN), Support Vector Machine (SVM), Adaptive Boosting (AdaBoost), eXtreme Gradient Boosting (XGBoost), Naïve Bayes (NB), and Multi-Layer Perceptron (MLP). An innovative pairwise ensemble strategy was applied, resulting in 28 unique model combinations. Several combinations demonstrated significantly improved performance compared with single models. Notably, the SVM   +   RF combination achieved the highest performance with an Area Under the Curve (AUC) of 0.865, representing a 0.9 percentage point improvement over the best-performing single model, RF (AUC = 0. 8 5 6). Inspired by these findings, a comprehensive ensemble framework incorporating all eight models was subsequently developed. The final ensemble prediction framework adopted a dynamic weighting strategy based on validation-set AUCs, with hyperparameter optimization conducted using Optuna via Tree-structured Parzen Estimator (TPE) sampling over 100 iterations. The resulting model weights were as follows: SVM (0.86), XGBoost (0.58), RF (0.69), KNN (0.52), NB(0.40), MLP (0.40), AdaBoost (0.01), and DT (0.002). After weighted probability fusion, the ensemble model achieved an AUC of 0.8676 on the test set, representing a 0.26 percentage point improvement over the best-performing pairwise combination (SVM   +   RF). Feature importance analysis identified the average SOFA score (sofa_score_avg), the mean sodium level (Sodium_mean), the minimum creatinine level (Creatinine_min), and the mean potassium level (Potassium_mean) as the most critical features in predicting in-hospital mortality risk for DKD patients. Additionally, higher SOFA scores, lower WBC_min, and BUN_max were associated with an increased risk of mortality, consistent with the results from multivariable logistic regression analysis. This study effectively utilized filter-based feature engineering to identify key prognostic indicators and, inspired by pairwise model-combination results, developed a dynamic weighted ensemble framework that significantly enhanced the identification of high-risk patients. [ABSTRACT FROM AUTHOR]
ISSN:02195194
DOI:10.1142/S0219519426400348