ArXiv

Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering

Authors
Soorya Ram Shimgekar, Michelle Hu, Dorisa Shehi...
Categories
cs.AI, cs.LG
arXiv
https://arxiv.org/abs/2608.06366v1
PDF
https://arxiv.org/pdf/2608.06366v1

Brief

The paper introduces nMAS, an evidence-linked, rubric-grounded multi-agent pipeline to automate heart-failure feature engineering from fragmented EHR data. Applied to 500 dummy patient records across nine tables, nMAS generated 202 derived features (132 structured + 70 rubric-scored), improved HFrEF AUROC to 0.963 and HFpEF to 0.910, and achieved 81.5% on an LLM rubric audit. Full text was not available; results are promising but limited by single-institution evaluation and need external validation.

Why it matters

Nimblemind Multi-Agent System (nMAS) produced 132 structured and 70 rubric-scored aggregated heart-failure features from 500 dummy patient records spanning nine EHR source tables; outputs were verified for structural integrity, rubric compliance, provenance, and audited by a restricted LLM.

Key details

  • Adding nMAS aggregated features raised held-out AUROC for HFrEF phenotyping from 0.895 to 0.963 and for HFpEF from 0.870 to 0.910; an independent LLM-based rubric assessment scored the features at 81.5% of maximum points.
  • Motivation and scope: EHR feature engineering accounts for 39–45% of data scientists' workload and heart failure affects ~6.7 million U.S. adults; evaluation was limited to a single-institution cohort (500 dummy records) and requires external validation.
Source evidence

Abstract

Electronic health record (EHR) feature engineering is a major bottleneck in clinical research and AI, accounting for 39-45% of data scientists' workload. This is especially pronounced in heart failure, which affects an estimated 6.7 million U.S. adults and requires integrating fragmented EHR data with disease-specific, guideline-based clinical reasoning. Existing rule-based and large language model (LLM)-based approaches offer only partial automation with limited maintainability and evidence traceability. We developed the Nimblemind Multi-Agent System (nMAS), an evidence-linked, rubric-grounded pipeline for automated heart-failure feature engineering, and evaluated it on 500 dummy patient records from nine EHR source tables. nMAS generated 132 structured and 70 rubric-scored aggregated features, verified for structural integrity, rubric compliance, and provenance, and audited by a restricted LLM. Adding the aggregated features improved held-out AUROC from 0.895 to 0.963 for HFrEF and 0.870 to 0.910 for HFpEF phenotyping, and an independent LLM-based rubric assessment of evidence support and methodological soundness scored the features at 81.5% of maximum points. These results demonstrate the feasibility of automated, auditable feature engineering for complex cardiovascular EHR data, though evaluation was limited to a single-institution cohort and external validation is needed.