Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering
2026-08-06 • Artificial Intelligence
Artificial IntelligenceMachine Learning
AI summaryⓘ
The authors created a system called Nimblemind Multi-Agent System (nMAS) to help automatically organize important heart failure information from electronic health records (EHRs). This system produces structured features linked to evidence and clinical guidelines, making it easier to analyze patient data. When tested on fake patient records, nMAS improved the accuracy of identifying two types of heart failure. The authors showed that this automated approach works well but note that it needs testing on data from other hospitals.
Electronic Health RecordFeature EngineeringHeart FailureHFrEFHFpEFMulti-Agent SystemLarge Language ModelsAUROCClinical PhenotypingEvidence Traceability
Authors
Soorya Ram Shimgekar, Michelle Hu, Dorisa Shehi, Daniel Kang, Roy Ka-Wei Lee, Koustuv Saha, Christian Poellabauer, Christopher Lee, Sajeev Singh, Piyum Zonooz, Navin Kumar, Zeeshan Ahmed, Priyadarshini Kachroo
Abstract
Electronic health record (EHR) feature engineering is a major bottleneck in clinical research and AI, accounting for 39-45% of data scientists' workload. This is especially pronounced in heart failure, which affects an estimated 6.7 million U.S. adults and requires integrating fragmented EHR data with disease-specific, guideline-based clinical reasoning. Existing rule-based and large language model (LLM)-based approaches offer only partial automation with limited maintainability and evidence traceability. We developed the Nimblemind Multi-Agent System (nMAS), an evidence-linked, rubric-grounded pipeline for automated heart-failure feature engineering, and evaluated it on 500 dummy patient records from nine EHR source tables. nMAS generated 132 structured and 70 rubric-scored aggregated features, verified for structural integrity, rubric compliance, and provenance, and audited by a restricted LLM. Adding the aggregated features improved held-out AUROC from 0.895 to 0.963 for HFrEF and 0.870 to 0.910 for HFpEF phenotyping, and an independent LLM-based rubric assessment of evidence support and methodological soundness scored the features at 81.5% of maximum points. These results demonstrate the feasibility of automated, auditable feature engineering for complex cardiovascular EHR data, though evaluation was limited to a single-institution cohort and external validation is needed.