Score
Designs and implements algorithms or modules that use estimated motion (e.g., optical flow or trajectories) from temporal image sequences to align, augment, or refine per-frame feature representations; the work produces motion-aware feature warping, attention/weighting, or fusion schemes that improve temporal consistency, object localization, and discriminative power of features for downstream tasks.
In collaborative perception, hardware heterogeneity induces feature-domain shift, while communication latency causes temporal misalignment—jointly degrading feature quality and accumulating cross-node errors. To address these challenges at the feature-level fusion stage, we propose a systematic alignment framework: (1) a consistency-preserving domain alignment module mitigates inter-device feature distribution discrepancies; (2) a progressive temporal alignment module corrects dynamic timing offsets via multi-scale motion modeling and two-stage compensation; and (3) an observability-constrained discriminator and instance-aware hierarchical aggregation strategy enhance semantic consistency. Evaluated on three benchmark datasets, our method achieves state-of-the-art performance and demonstrates significantly improved robustness under high communication latency and pose estimation errors.
This paper proposes a training-free video semantic editing framework that precisely injects semantic features of a user-provided reference object into designated spatiotemporal regions of a video, while rigorously preserving the original motion trajectories and visual context. Methodologically, it introduces a novel motion-aware concept alignment mechanism, integrating structured noise trajectory manipulation, momentum-based semantic correction, gamma-residual noise stabilization, and latent-space object detection and tracking; additionally, diagonal denoising scheduling and class-agnostic segmentation are incorporated to enhance controllability. Key contributions include: (1) the first CASS (Consistency-Aware Semantic Similarity) metric specifically designed for video semantic editing evaluation; (2) superior performance over state-of-the-art baselines on a newly constructed benchmark dataset, with significant improvements in spatial consistency, motion coherence, and CASS scores; and (3) zero-shot, high-fidelity, temporally consistent video editing without any model training.
To address the limited robustness of cross-video stream feature matching under noise, frame misalignment, and cross-modal (e.g., infrared–visible) conditions, this paper proposes a purely temporal, keypoint-free matching method. Instead of relying on spatial keypoint detection, our approach models motion signatures of pixel blocks across consecutive frames by jointly encoding optical flow and block-level temporal correlations, and employs Dynamic Time Warping (DTW) to enhance cross-video motion alignment. The method inherently achieves scale, rotation, and translation invariance, requires no training, and supports cross-modal matching. Experiments demonstrate that it significantly outperforms state-of-the-art methods across diverse challenging scenarios—achieving substantial gains in matching accuracy while reducing computational overhead by over 60%, thereby enabling real-time deployment.
Existing custom video generation methods struggle to simultaneously preserve subject appearance fidelity and ensure temporal motion consistency, primarily due to the lack of object-level subject-motion disentangled modeling. This paper proposes a subject-motion representation disentanglement and alignment framework for text-to-video generation. We introduce the first object-level representation alignment mechanism, design a sparse spatiotemporal LoRA injection strategy to minimize fine-tuning interference, and develop a collaborative self-supervised subject encoder and optical-flow-based motion encoder. The method integrates self-supervised representation learning, optical-flow-driven motion modeling, efficient LoRA-based fine-tuning, and spatiotemporally sparse adapters. Evaluated on multiple benchmarks, it achieves significant improvements in subject similarity (+12.6%) and motion consistency (+9.8%), enabling fine-grained, disentangled controllable generation with both high visual fidelity and temporal stability.
Existing video motion editing methods are largely confined to simple transformations (e.g., translation, scaling) and struggle to accurately transfer complex semantic motions—such as full-body gestures, facial expressions, object dynamics, or camera motion—using only a single reference video. This work proposes a semantic-level video motion transfer framework that enables precise motion transfer from one reference video to arbitrary target images—without requiring spatial alignment. Our approach leverages a pre-trained image-video diffusion model and introduces three key innovations: (1) motion-textual inversion, a novel technique that decouples appearance and motion representations via joint embedding of motion and textual semantics; (2) frame-wise dilated motion embeddings for high-temporal-resolution motion encoding; and (3) implicit motion encoding coupled with cross-attention modulation to enable zero-shot cross-domain transfer. Extensive experiments demonstrate significant improvements over state-of-the-art methods in motion fidelity, generalizability, and temporal consistency.
This work investigates the role of long-term motion cues in visual perception tasks and their advantages over static image representations. The authors construct a low-dimensional, efficient motion representation based on point trajectory estimation and systematically evaluate its performance across multiple tasks, including action recognition, object understanding, material classification, and spatial reasoning. The study demonstrates that this motion-based representation captures rich semantic information and exhibits significantly stronger generalization than conventional image features under low-data and zero-shot settings. Furthermore, when fused with standard video representations, it yields additional accuracy gains. Notably, the proposed approach achieves a superior trade-off between computational efficiency (measured in GFLOPs) and performance, highlighting its potential for efficient visual understanding.
该研究解决了视频中局部运动表示问题,通过全视频处理和基于空间掩码的区域查询方法,生成保留全局上下文的局部动态嵌入。
为解决视频生成中的运动不一致问题,提出MotionSpec方法,通过频谱轨迹一致性(STC)和局部流一致性(LFC)增强视频的运动连贯性和真实性。
本文提出COMET框架,通过显式时间表示、外观-运动融合及方向感知优化,解决了视频多模态大语言模型中细粒度运动-时间理解不足的问题。
研究通过分析V-JEPA 2和VideoMAE-v2模型,探讨了视频基础模型中时空表示的编码内容、出现位置及几何组织方式,并使用轻量级探针来发现三种时间属性。