Diffusion-Based Action Recognition Generalizes to Untrained Domains

📅 2025-09-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address insufficient cross-species, cross-view, and cross-scenario generalization in action recognition, this paper proposes a robust action recognition framework synergizing visual diffusion models (VDMs) and Transformers. The method leverages conditional generation at early diffusion timesteps to extract semantically rich features and introduces multi-stage conditional constraints to enhance cross-domain semantic consistency. Concurrently, a Transformer architecture efficiently aggregates temporal features. Evaluated on three major cross-domain benchmarks—cross-species, cross-view, and cross-scenario—the approach achieves state-of-the-art performance, significantly improving generalization to unseen domains. By jointly exploiting the representational power of diffusion-based semantic grounding and Transformer-based temporal modeling, the framework advances toward human-level robustness in action understanding.

Technology Category

Computer Vision: Diffusion Models for VisionIntelligent Robots: Multimodal Perception & Sensor FusionHumans and AI: Human-Aware Planning and Behavior Prediction

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Vertical and domain-specific searchGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Humans can recognize the same actions despite large context and viewpoint variations, such as differences between species (walking in spiders vs. horses), viewpoints (egocentric vs. third-person), and contexts (real life vs movies). Current deep learning models struggle with such generalization. We propose using features generated by a Vision Diffusion Model (VDM), aggregated via a transformer, to achieve human-like action recognition across these challenging conditions. We find that generalization is enhanced by the use of a model conditioned on earlier timesteps of the diffusion process to highlight semantic information over pixel level details in the extracted features. We experimentally explore the generalization properties of our approach in classifying actions across animal species, across different viewing angles, and different recording contexts. Our model sets a new state-of-the-art across all three generalization benchmarks, bringing machine action recognition closer to human-like robustness. Project page: $href{https://www.vision.caltech.edu/actiondiff/}{ exttt{vision.caltech.edu/actiondiff}}$ Code: $href{https://github.com/frankyaoxiao/ActionDiff}{ exttt{github.com/frankyaoxiao/ActionDiff}}$
Problem

Research questions and friction points this paper is trying to address.

Generalizing action recognition across species, viewpoints, and contexts
Overcoming deep learning limitations in cross-domain generalization
Enhancing semantic feature extraction using diffusion model timesteps
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision Diffusion Model generates semantic features
Transformer aggregates diffusion-based features
Early timestep conditioning enhances generalization capability