Doubly Robust Functional Representation Learning for Longitudinal Causal Inference with Irregular Histories

2026-07-30Machine Learning

Machine Learning
AI summary

The authors address how to analyze complex medical histories that are recorded irregularly over time, like lab results taken at uneven intervals. They created a method called DR-FRL that turns these irregular data points into useful summaries for studying cause-and-effect relationships. Their approach carefully checks these summaries to ensure they work well with statistical methods that estimate treatment effects. When tested with simulations and real ICU data, their method performed better in complicated settings but showed existing simple summaries were already informative for predicting patient outcomes. This work helps improve analyses when data are messy or incomplete over time.

doubly robust estimationcausal inferencefunctional data analysisefficient influence functioncross-fittinglongitudinal datanuisance parametersCatoni aggregationasymptotic linearityirregular time series
Authors
Mengfei Ran, Yifeng Shen, Ruijie Guan
Abstract
Longitudinal causal studies often record histories as irregular functional fragments: laboratory values, physiologic signals, sensor streams, and image-derived summaries measured at unequal and informative times. Standard doubly robust estimators usually require scalar summaries, whereas sequence learners optimize prediction losses that need not stabilize the efficient influence function. We propose Doubly Robust Functional Representation Learning (DR-FRL), a cross-fitted workflow that turns irregular histories into estimand-targeted states for observed-history regimes. Functional and temporal encoders map point clouds and prior histories into states; nuisance heads estimate outcome, treatment, and censoring functions; and EIF-targeted validation, calibration, overlap, tail, and ablation diagnostics assess whether the state supports the estimating equation. If the selected state preserves the nuisance information needed by the EIF, representation error enters the same second-order product remainder as ordinary nuisance error, and the mean estimator is asymptotically linear under explicit rate, overlap, calibration, and stability conditions. Catoni aggregation is treated separately as a bounded-influence point estimator, not a replacement for Wald inference. Simulations show gains when functional confounding is high-dimensional, measurement is informative, support is weak, or pseudo-outcomes are heavy-tailed. A VitalDB audit shows that DR-FRL can use irregular laboratory point clouds and deliver a useful negative finding: for this ICU-disposition endpoint, scalar laboratory summaries already carry much endpoint-relevant information.