temporal feature aggregation

Designs and implements methods that aggregate, align, and fuse per-frame visual features over time to produce temporally consistent, richer feature representations for video sequences. This includes modeling temporal relationships and fusing fine-grained appearance cues across frames to increase feature completeness and robustness for small, fast-moving, or partially occluded objects.

temporalfeatureaggregation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing video multimodal fusion methods naively adapt static image fusion techniques, neglecting temporal dependencies and resulting in inter-frame inconsistency. To address this, we propose the first temporal modeling and vision–semantics co-learning framework specifically designed for video fusion. Our method introduces a vision–semantics interaction module and a temporal coordination module, implemented via dual distillation branches—DINOv2 for semantic representation and VGG19 for low-level visual features. We further incorporate a temporal enhancement mechanism, a temporal consistency loss, and a dedicated evaluation metric suite. Extensive experiments demonstrate that our approach significantly improves weak-information recovery and dynamic coherence, achieving state-of-the-art performance across multiple public video datasets. It attains superior results in all three critical dimensions: visual fidelity, semantic accuracy, and temporal continuity. The source code is publicly available.

Addressing video degradation within the fusion pipelineEnsuring temporal consistency in multi-modal video fusionIntegrating visual-semantic collaboration for enhanced representation

Not All Frame Features Are Equal: Video-to-4D Generation via Decoupling Dynamic-Static Features

Feb 12, 2025
LY
Liying Yang
🏛️ Macau University of Science and Technology | The University of Queensland | Institute of Automation, Chinese Academy of Sciences

In video-to-dynamic-4D-scene generation, static regions dominate the optimization process, leading to blurred dynamic details, texture distortion, and overfitting. To address this, we propose DS4D—the first framework to explicitly decouple dynamic and static features along the temporal dimension. Methodologically, DS4D introduces a Dynamic-Static Feature Decoupling (DSFD) module to isolate motion representations; a spatio-temporal similarity fusion (TSSF) module that adaptively aggregates multi-view dynamic information for improved motion modeling; and a Gaussian-based differentiable 4D reconstruction pipeline. Evaluated on real-world scene datasets, DS4D significantly enhances texture fidelity in dynamic regions and inter-frame motion consistency, achieving state-of-the-art performance on video-to-4D generation.

Decoupling dynamic and static featuresEnhancing dynamic region representationImproving 4D generation from video

DATA: Domain-And-Time Alignment for High-Quality Feature Fusion in Collaborative Perception

Jul 24, 2025
CT
Chengchang Tian
🏛️ Southeast University | Washington State University

In collaborative perception, hardware heterogeneity induces feature-domain shift, while communication latency causes temporal misalignment—jointly degrading feature quality and accumulating cross-node errors. To address these challenges at the feature-level fusion stage, we propose a systematic alignment framework: (1) a consistency-preserving domain alignment module mitigates inter-device feature distribution discrepancies; (2) a progressive temporal alignment module corrects dynamic timing offsets via multi-scale motion modeling and two-stage compensation; and (3) an observability-constrained discriminator and instance-aware hierarchical aggregation strategy enhance semantic consistency. Evaluated on three benchmark datasets, our method achieves state-of-the-art performance and demonstrates significantly improved robustness under high communication latency and pose estimation errors.

Address domain gaps from hardware diversity and deployment conditionsEnhance semantic feature quality for collaborative perception fusionMitigate temporal misalignment caused by transmission delays

A Unified Solution to Video Fusion: From Multi-Frame Learning to Benchmarking

May 26, 2025
ZZ
Zixiang Zhao
🏛️ ETH Zürich | Xi’an Jiaotong University | Shanghai Jiao Tong University | Nanjing University

Existing video fusion methods predominantly process frames independently, neglecting temporal correlations and consequently suffering from flickering and temporal inconsistency. To address this, we propose Unified Video Fusion (UniVF), the first framework enabling joint spatiotemporal optimization. Our contributions are threefold: (1) We introduce VF-Bench—the first comprehensive Video Fusion Benchmark covering four distinct modalities: infrared–visible, medical, remote sensing, and low-light fusion; (2) We design a flow-guided deformable feature alignment module to enable multi-frame collaborative learning; and (3) We propose a synthetic-data-driven protocol for video-pair construction and joint spatial–temporal quality assessment. Extensive experiments on VF-Bench demonstrate that UniVF consistently outperforms state-of-the-art methods, achieving significant improvements in both temporal consistency and spatial fidelity of fused videos.

Addressing temporal inconsistency in video fusion methodsIntroducing a benchmark for evaluating video fusion tasksProposing a unified framework for coherent video fusion

This work addresses the challenges of temporally inconsistent predictions—such as flickering—in human-centric dense video tasks under motion, occlusion, and illumination changes, compounded by the scarcity of multi-task paired video supervision. To this end, we propose a scalable, photorealistic synthetic human video generation method that, for the first time, provides both frame-level and sequence-level pixel-wise annotations, including depth, surface normals, and masks. Leveraging this data, we develop a unified Vision Transformer (ViT)-based dense prediction architecture that integrates CSE human geometric priors with a lightweight channel reweighting module. Our approach employs a two-stage training strategy—static pretraining followed by dynamic sequence fine-tuning—to jointly optimize spatial and temporal consistency. The method achieves state-of-the-art performance on THuman2.1 and Hi4D benchmarks and demonstrates strong generalization to in-the-wild real-world videos.

flickeringhuman-centric dense predictionpaired supervision

Latest Papers

What's happening recently
View more

This work addresses the challenges of ghosting and drift in infrared and visible video fusion, which arise from temporal misalignment, geometric rigidity, and error accumulation in diffusion models. The authors reformulate the fusion task as a history-conditioned motion generation problem and propose a spectral filtering framework that implicitly models motion dynamics to circumvent explicit alignment. Key innovations include stable historical guidance, a soft temporal anchoring mechanism, and a decoupled structure-motion adaptive strategy, complemented by a two-stage training scheme and latent space optimization. The method achieves state-of-the-art performance in both fusion quality and temporal consistency, effectively suppressing artifacts and drift.

driftingerror accumulationghosting artifacts

This study addresses the limitations of pretrained video models in instruction following, temporal coherence, and physical safety constraints by presenting the first systematic survey of post-training and alignment strategies for video generation. We propose a unified taxonomy encompassing both implicit and explicit alignment paradigms, comprehensively reviewing four major technical approaches: supervised fine-tuning, self-training distillation, preference-based reward mechanisms, and inference-time interventions. By consolidating existing dataset benchmarks and evaluation practices, this work identifies core open challenges, particularly scalable reward design and temporal consistency. Ultimately, it provides a systematic foundation for enhancing the controllability and reliability of video generation systems.

AlignmentHuman IntentPost-Training

This work proposes treating temporal flow rate as a learnable visual concept to enable perception and controllable generation of video playback speed. Through a self-supervised approach leveraging multimodal cues and temporal structure inherent in videos, the model accurately detects speed variations and estimates actual playback rates. Building upon this framework, the authors construct the largest wild slow-motion video dataset to date and demonstrate two key applications: synthesizing videos at user-specified playback speeds and performing temporal super-resolution on low-frame-rate videos to recover fine-grained dynamic details. This study is the first to model time flow rate as a manipulable perceptual dimension, significantly advancing capabilities in video temporal understanding and generation.

slow-motiontemporal controltemporal reasoning

This work addresses the ambiguity regarding whether existing single-stage video object detectors genuinely leverage temporal context, as standard evaluation metrics often fail to reveal their actual reliance on temporal information. To this end, we propose TemporalLens, a diagnostic framework that quantifies a model’s temporal dependency through controlled perturbations—including temporal shuffling, structured occlusion, and redundancy injection. Furthermore, we design YOLO-3D based on YOLOv8, explicitly preserving the temporal dimension within the backbone to enhance genuine temporal reasoning. Experiments demonstrate that TemporalLens effectively distinguishes between stacked 2D models and true temporal architectures, while YOLO-3D achieves an average mAP@50 improvement of 3.7 percentage points with 32-frame inputs, underscoring the critical role of temporal depth in performance gains.

model diagnosticssingle-stage detectorstemporal context

This work addresses the challenge of simultaneously preserving spatiotemporal consistency and high-frequency details in infrared and visible video fusion. To this end, the authors propose a frequency-aware fusion framework that decomposes features into high- and low-frequency components for separate modeling. The high-frequency branch captures motion cues and fine details through sparse cross-modal spatiotemporal interactions, while the low-frequency branch enhances robustness to dynamic artifacts such as flickering and jitter via a temporal perturbation strategy. Additionally, an offset-aware temporal consistency constraint is introduced to stabilize inter-frame representations. By jointly integrating frequency decomposition, sparse cross-modal interaction, and temporal perturbation—a combination not previously explored—this method achieves state-of-the-art performance on multiple public benchmarks, significantly improving temporal stability without compromising high-frequency detail preservation.

high-frequency detailsinfrared and visible video fusionspatial detail preservation

Hot Scholars

GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
HD

Henghui Ding

Fudan University
Computer VisionMachine LearningSegmentationAIGC
JG

Juergen Gall

University of Bonn, Lamarr - Institute for Machine Learning and Artificial Intelligence
Action recognitionVideo understandingAnticipationForecasting
NN

Nassir Navab

Professor of Computer Science, Technische Universität München
XL

Xiaodan Liang

Professor of Computer Science, Sun Yat-sen University, MBZUAI, CMU, NUS
Computer visionEmbodied AIMachine learning