ArXiv

ESAR: Event-Based Synthetic Aperture Reconstruction

Authors
Harbir Antil, Daniel Blauvelt, David Sayre
Categories
eess.IV, cs.CV, eess.SP
arXiv
https://arxiv.org/abs/2607.15073v1
PDF
https://arxiv.org/pdf/2607.15073v1

Brief

ESAR (Event-Based Synthetic Aperture Reconstruction) frames monocular event-camera imaging as recovering a static log–radiance field θ via APθ = b + η, where P maps the scene into motion-dependent views and A is temporal differencing. Exploiting near-nadir motion (approximate shifts) and regularized inversion, experiments on simulated and Falcon Neuro data (Antil et al., 2026-07-16) recover coherent large-scale structure and suppress fine texture. Summary based on the abstract (full text not reviewed).

Why it matters

ESAR formulates monocular event-camera reconstruction as a synthetic-aperture inverse problem that recovers a static ground-domain log–radiance field θ ∈ ℝ^{N_g}, replacing the latent pixel-time volume v ∈ ℝ^{N_pN_t} with the geometric relation v = Pθ and the linearized measurement model APθ = b + η (A = temporal differencing, b = signed binned event counts).

Key details

  • Under near-nadir motion successive projections are approximately shifted views so the composite operator AP is ill-conditioned (spatial averaging combined with temporal differencing); the authors (Antil, Blauvelt, Sayre; arXiv 2026-07-16) use regularized inversion and show on simulated data and real Falcon Neuro near-nadir event recordings that the θ-based method recovers coherent large-scale spatial structure while suppressing fine-scale texture relative to dynamic latent-image and learned event-reconstruction baselines.
Source evidence

Abstract

Event cameras report asynchronous polarity events when changes in log--radiance exceed a fixed contrast threshold, producing signed temporal contrast measurements rather than conventional image frames. We formulate monocular event-based imaging as a synthetic-aperture inverse problem for a static ground-domain log--radiance field $θ\in \mathbb{R}^{Ng}$. Instead of reconstructing a latent pixel-time volume $v \in \mathbb{R}^{NpN_t}$, we impose the geometric relation $v=Pθ$, where $P$ maps the fixed scene into motion-dependent latent views. Aggregating events over finite time intervals gives the linearized model [ APθ= b+η, ] where $A$ is a temporal differencing operator, $b$ contains signed binned event counts, and $η$ represents measurement and modeling errors. This decomposition exposes a synthetic-aperture structure: under near-nadir motion, successive projections are approximately shifted views of a common scene, while the composite operator $AP$ remains ill-conditioned because it combines spatial averaging with temporal differencing. We therefore use regularized inversion to recover $θ$. Numerical experiments on simulated data and real near-nadir Falcon Neuro event data show that the proposed $θ$-based formulation recovers coherent large-scale spatial structure, relative to dynamic latent-image and learned event-reconstruction baselines, while suppressing fine-scale texture.