ArXiv

A Mathematical Framework for Topological Causal Data Analysis

Authors
Hugo Gobato Souto, Ioannis Diamantis
Categories
stat.ME, math.ST, stat.ML
arXiv
https://arxiv.org/abs/2607.28161v1
PDF
https://arxiv.org/pdf/2607.28161v1

Brief

Topological Causal Data Analysis (TCDA) addresses causal inference for structured outcomes (images, point clouds, networks) where Y^1−Y^0 may be undefined. The authors separate modeling choices (observation space, causal class, topology, query), distinguish outcome- vs distribution-level contrasts, derive identification and doubly-robust estimators for Banach-valued summaries, prove g-formula-based identification with stability-transfer bounds, and formalize target-specific topological ignorability. Summary based on the paper abstract; full text was not reviewed.

Why it matters

Introduces Topological Causal Data Analysis (TCDA), a framework that separates observation space, causal-model class, topological representation, and causal query, and distinguishes outcome-level TCDA (transforms individual potential outcomes) from distribution-level TCDA (transforms interventional outcome laws) (Souto & Diamantis, 2026-07-30, arXiv:2607.28161v1).

Key details

  • Provides formal results: outcome-level identification and doubly robust representations for Banach-space-valued summaries; distribution-level identification via the standard causal g-formula plus stability-transfer bounds and plug-in consistency; and defines target-specific 'topological ignorability' to identify covariate-standardized coarse effects without full interventional laws.
  • Limits observational topology for causal discovery: observational topological summaries can aid diagnosis on restricted model classes but cannot by themselves identify causal structure.
Source evidence

Abstract

Many modern outcomes, including images, point clouds, networks, and spatial fields, are structured objects for which (Y^1-Y^0) may be undefined or scientifically inadequate. We introduce \emph{Topological Causal Data Analysis} (TCDA), a framework separating the observation space, causal-model class, topological representation, and causal query. Topology does not define interventions; it supplies stable, shape-sensitive summaries after causal assumptions have been specified. We distinguish outcome-level TCDA, which transforms individual potential outcomes, from distribution-level TCDA, which transforms interventional outcome laws, and characterize when outcome and distribution level contrasts agree. Building on recent outcome-level theory, we formulate identification and doubly robust representations for Banach-space-valued summaries. At the distribution level, we identify targets through the standard causal (g)-formula and derive stability-transfer bounds and plug-in consistency. We also place target-specific topological ignorability within the framework, clarifying when a covariate-standardized coarse effect can be identified without identifying the full interventional laws. Finally, we delimit the role of observational topology in causal discovery: it can assist diagnosis on restricted model classes but cannot by itself identify causal structure.