🤖 AI Summary
This work proposes a novel paradigm termed “temporal backward imaging,” which aims to infer recent human–environment interaction events from fading multimodal traces—specifically, physical residues observable across thermal, ultraviolet, and visible spectra. To enable this research direction, we introduce TRACE-HEI, the first synchronized trimodal video dataset capturing such residual signals, and develop a vision–language-guided diffusion generative model that reconstructs plausible past scene frames conditioned on structured textual descriptions. Experimental results demonstrate the feasibility of inferring recent activities from residual traces under multimodal complementary constraints. This study formalizes the task for the first time, establishes a benchmark dataset and methodological framework, and expands the frontier of scene understanding beyond instantaneous observation.
📝 Abstract
We introduce time-reversed imaging, a new paradigm that infers what just happened in a scene from fading multimodal traces. Instead of extrapolating or interpolating video frames, our goal is to infer past human-environment interactions from residual physical imprints observable in thermal, ultraviolet, and visible spectra. To study this problem, we present TRACE-HEI, the first proof-of-concept dataset for time-reversed imaging, containing synchronized tri-modal video sequences of actions such as sitting, touching, moving objects, and liquid spills, captured across diverse materials and recorded up to three minutes after contact. To establish the benchmark, we propose a multimodal inference approach that extracts structured textual descriptions of detected traces and uses them to constrain a vision-language-guided diffusion model for reconstructing plausible past frames. Experiments show that inferring recent events from fading traces is challenging but feasible when complementary modalities reduce solution ambiguity. This work defines the first computational and experimental foundation for time-reversed imaging, bridging vision, physics, and generative reasoning, and opening new directions for scene understanding beyond instantaneous observation.