Federated Deep Learning for Privacy-Preserving Cardiovascular Disease Risk Prediction
2026-07-09 • Machine Learning
Machine LearningHuman-Computer Interaction
AI summaryⓘ
The authors developed a method to predict heart disease risk by training models on data from two different groups without sharing sensitive patient details. They used a technique called federated learning, which lets different hospitals work together on a model without combining their data in one place. Their approach improved prediction accuracy compared to models trained separately on each group. This shows that privacy-friendly collaboration can help make better health predictions even when data comes from very different sources.
Cardiovascular diseaseRisk prediction modelFederated learningPrivacy preservationDeep survival modelsPopulation-based cohortsC-statisticClinical outcomes
Authors
Hyunho Mo, Djura Smits, Mahlet A. Birhanu, Maarten J. G. Leening, Daniel Bos, Pim van der Harst, Esther E. Bron
Abstract
Cardiovascular disease risk prediction models often rely on data from a single institution or centrally pooled datasets. Extending these models across institutions could be limited by privacy regulations and constraints on sharing patient-level data. Federated learning enables collaborative model development without transferring sensitive patient data, but its application in healthcare remains challenging because datasets often differ in size, population characteristics, and outcome definitions. In this study, we present a federated deep learning approach for privacy-preserving cardiovascular disease risk prediction that integrates two population-based cohorts with different characteristics: Lifelines, including 148,230 participants meeting the study inclusion criteria with self-reported outcomes, and the Rotterdam Study, including a smaller cohort of 10,155 participants with digitally linked clinical outcomes. Model performance was primarily evaluated on the Rotterdam Study because of its complete follow-up. Deep survival models trained using federated learning achieved higher predictive performance than models trained locally without federation. For the Rotterdam Study, the C-statistic increased from 0.728 (95% CI: 0.717-0.739) to 0.739 (95% CI: 0.728-0.749). For Lifelines, the C-statistic increased from 0.783 (95% CI: 0.775-0.791) to 0.787 (95% CI: 0.780-0.792). These findings suggest that federated deep learning across heterogeneous cohorts can improve cardiovascular disease risk prediction while preserving the privacy of individual-level patient data.