ArXiv

Doubly Robust Functional Representation Learning for Longitudinal Causal Inference with Irregular Histories

Authors
Mengfei Ran, Yifeng Shen, Ruijie Guan
Categories
stat.ML, cs.LG, stat.AP, stat.CO, stat.ME
arXiv
https://arxiv.org/abs/2607.28567v1
PDF
https://arxiv.org/pdf/2607.28567v1

Brief

Doubly Robust Functional Representation Learning (DR-FRL) addresses causal estimation from irregular, informative longitudinal fragments by learning estimand-targeted states with functional and temporal encoders and nuisance heads, then validating them with EIF-targeted diagnostics. The authors show representation error contributes as a second-order remainder so doubly robust, asymptotically linear inference holds under explicit conditions. Simulations and a VitalDB ICU audit demonstrate practical gains and limits.

Why it matters

Mengfei Ran, Yifeng Shen, and Ruijie Guan (arXiv 2607.28567v1, 2026-07-30) propose DR-FRL, a cross-fitted workflow that maps irregular functional histories (point clouds, sensor streams, lab fragments) into estimand-targeted states via functional and temporal encoders, with nuisance heads for outcome, treatment, and censoring and EIF-targeted validation/calibration/overlap/tail diagnostics.

Key details

  • They prove that if the learned state preserves EIF-relevant nuisance information, representation error enters the same second-order product remainder as ordinary nuisance error, so the mean estimator is asymptotically linear under explicit rate, overlap, calibration, and stability conditions; Catoni aggregation is treated separately as a bounded-influence point estimator.
  • Empirically, simulations show DR-FRL gains when functional confounding is high-dimensional, measurement is informative, support is weak, or pseudo-outcomes are heavy-tailed; a VitalDB ICU audit using irregular laboratory point clouds found scalar laboratory summaries already carry much of the endpoint-relevant information for an ICU-disposition outcome (a useful negative finding).
Source evidence

Abstract

Longitudinal causal studies often record histories as irregular functional fragments: laboratory values, physiologic signals, sensor streams, and image-derived summaries measured at unequal and informative times. Standard doubly robust estimators usually require scalar summaries, whereas sequence learners optimize prediction losses that need not stabilize the efficient influence function. We propose Doubly Robust Functional Representation Learning (DR-FRL), a cross-fitted workflow that turns irregular histories into estimand-targeted states for observed-history regimes. Functional and temporal encoders map point clouds and prior histories into states; nuisance heads estimate outcome, treatment, and censoring functions; and EIF-targeted validation, calibration, overlap, tail, and ablation diagnostics assess whether the state supports the estimating equation. If the selected state preserves the nuisance information needed by the EIF, representation error enters the same second-order product remainder as ordinary nuisance error, and the mean estimator is asymptotically linear under explicit rate, overlap, calibration, and stability conditions. Catoni aggregation is treated separately as a bounded-influence point estimator, not a replacement for Wald inference. Simulations show gains when functional confounding is high-dimensional, measurement is informative, support is weak, or pseudo-outcomes are heavy-tailed. A VitalDB audit shows that DR-FRL can use irregular laboratory point clouds and deliver a useful negative finding: for this ICU-disposition endpoint, scalar laboratory summaries already carry much endpoint-relevant information.