A preprocessing-enhanced stacking classifier for generalized cardiovascular disease detection across diverse datasets
Scientific Reports, cilt.16, sa.1, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 16 Sayı: 1
- Basım Tarihi: 2026
- Doi Numarası: 10.1038/s41598-026-41042-z
- Dergi Adı: Scientific Reports
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, BIOSIS, Chemical Abstracts Core, EMBASE, MEDLINE, Directory of Open Access Journals, Zoological Record, Academic Search Ultimate (EBSCO), Natural Science Collection (ProQuest), Biological Science Database (ProQuest), Biomedical Reference Collection: Corporate Edition (EBSCO), Health Research Premium Collection (ProQuest)
- Anahtar Kelimeler: Clustering, CVDs, Ensemble learning, ML, preprocessing, Stacking
- Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
- İstanbul Gelişim Üniversitesi Adresli: Evet
Özet
Cardiovascular diseases (CVDs) remain a major global health challenge, requiring early-detection models that are both accurate and generalizable across diverse data settings. This study introduces a preprocessing-enhanced stacking ensemble for binary CVD prediction, explicitly designed to improve robustness under heterogeneous feature distributions. The preprocessing pipeline incorporates feature transformation, derived attribute construction, encoding, and K-modes clustering, all applied after strict train–test separation to preserve evaluation validity. The proposed stacking architecture integrates three complementary tree-based base learners i.e., Random Forest, Decision Tree, and Extra Trees with Logistic Regression as a meta-learner to aggregate out-of-fold predictions. The framework was evaluated on three heterogeneous datasets. The ensemble achieved accuracies of 93.26%, 72%, and 99% on Datasets I, II, and III, respectively, with performance stability confirmed using 95% confidence intervals across five random seeds. Statistical significance analysis using McNemar’s test demonstrated that the proposed model significantly outperformed several strong baselines (p < 0.05), including Random Forest, Logistic Regression, and XGBoost on Dataset I; CNN on Dataset II; and Decision Tree and Logistic Regression on Dataset III. These results indicate that the proposed framework maintains consistent performance across varying data modalities, noise levels, and feature structures.