A preprocessing-enhanced stacking classifier for generalized cardiovascular disease detection across diverse datasets


Ashraf A., Masih A., Saddiqa A., Mahmood J., Ali A., Abdulnabi M. S. H., ...Daha Fazla

Scientific Reports, cilt.16, sa.1, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 16 Sayı: 1
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1038/s41598-026-41042-z
  • Dergi Adı: Scientific Reports
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, BIOSIS, Chemical Abstracts Core, EMBASE, MEDLINE, Directory of Open Access Journals, Zoological Record, Academic Search Ultimate (EBSCO), Natural Science Collection (ProQuest), Biological Science Database (ProQuest), Biomedical Reference Collection: Corporate Edition (EBSCO), Health Research Premium Collection (ProQuest)
  • Anahtar Kelimeler: Clustering, CVDs, Ensemble learning, ML, preprocessing, Stacking
  • Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
  • İstanbul Gelişim Üniversitesi Adresli: Evet

Özet

Cardiovascular diseases (CVDs) remain a major global health challenge, requiring early-detection models that are both accurate and generalizable across diverse data settings. This study introduces a preprocessing-enhanced stacking ensemble for binary CVD prediction, explicitly designed to improve robustness under heterogeneous feature distributions. The preprocessing pipeline incorporates feature transformation, derived attribute construction, encoding, and K-modes clustering, all applied after strict train–test separation to preserve evaluation validity. The proposed stacking architecture integrates three complementary tree-based base learners i.e., Random Forest, Decision Tree, and Extra Trees with Logistic Regression as a meta-learner to aggregate out-of-fold predictions. The framework was evaluated on three heterogeneous datasets. The ensemble achieved accuracies of 93.26%, 72%, and 99% on Datasets I, II, and III, respectively, with performance stability confirmed using 95% confidence intervals across five random seeds. Statistical significance analysis using McNemar’s test demonstrated that the proposed model significantly outperformed several strong baselines (p < 0.05), including Random Forest, Logistic Regression, and XGBoost on Dataset I; CNN on Dataset II; and Decision Tree and Logistic Regression on Dataset III. These results indicate that the proposed framework maintains consistent performance across varying data modalities, noise levels, and feature structures.