motion-guided feature enhancement

Designs and implements algorithms or modules that use estimated motion (e.g., optical flow or trajectories) from temporal image sequences to align, augment, or refine per-frame feature representations; the work produces motion-aware feature warping, attention/weighting, or fusion schemes that improve temporal consistency, object localization, and discriminative power of features for downstream tasks.

motion-guidedfeatureenhancement

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

DATA: Domain-And-Time Alignment for High-Quality Feature Fusion in Collaborative Perception

Jul 24, 2025
CT
Chengchang Tian
🏛️ Southeast University | Washington State University

In collaborative perception, hardware heterogeneity induces feature-domain shift, while communication latency causes temporal misalignment—jointly degrading feature quality and accumulating cross-node errors. To address these challenges at the feature-level fusion stage, we propose a systematic alignment framework: (1) a consistency-preserving domain alignment module mitigates inter-device feature distribution discrepancies; (2) a progressive temporal alignment module corrects dynamic timing offsets via multi-scale motion modeling and two-stage compensation; and (3) an observability-constrained discriminator and instance-aware hierarchical aggregation strategy enhance semantic consistency. Evaluated on three benchmark datasets, our method achieves state-of-the-art performance and demonstrates significantly improved robustness under high communication latency and pose estimation errors.

Address domain gaps from hardware diversity and deployment conditionsEnhance semantic feature quality for collaborative perception fusionMitigate temporal misalignment caused by transmission delays

This paper proposes a training-free video semantic editing framework that precisely injects semantic features of a user-provided reference object into designated spatiotemporal regions of a video, while rigorously preserving the original motion trajectories and visual context. Methodologically, it introduces a novel motion-aware concept alignment mechanism, integrating structured noise trajectory manipulation, momentum-based semantic correction, gamma-residual noise stabilization, and latent-space object detection and tracking; additionally, diagonal denoising scheduling and class-agnostic segmentation are incorporated to enhance controllability. Key contributions include: (1) the first CASS (Consistency-Aware Semantic Similarity) metric specifically designed for video semantic editing evaluation; (2) superior performance over state-of-the-art baselines on a newly constructed benchmark dataset, with significant improvements in spatial consistency, motion coherence, and CASS scores; and (3) zero-shot, high-fidelity, temporally consistent video editing without any model training.

Bridging image-domain semantic mixing to video editingEnsuring temporal coherence in video frame transitionsInjecting reference image features while preserving motion

Flow Intelligence: Robust Feature Matching via Temporal Signature Correlation

Apr 16, 2025
JW
Jie Wang
🏛️ Tsinghua University | University of Electronic Science and Technology of China | New York University

To address the limited robustness of cross-video stream feature matching under noise, frame misalignment, and cross-modal (e.g., infrared–visible) conditions, this paper proposes a purely temporal, keypoint-free matching method. Instead of relying on spatial keypoint detection, our approach models motion signatures of pixel blocks across consecutive frames by jointly encoding optical flow and block-level temporal correlations, and employs Dynamic Time Warping (DTW) to enhance cross-video motion alignment. The method inherently achieves scale, rotation, and translation invariance, requires no training, and supports cross-modal matching. Experiments demonstrate that it significantly outperforms state-of-the-art methods across diverse challenging scenarios—achieving substantial gains in matching accuracy while reducing computational overhead by over 60%, thereby enabling real-time deployment.

Enabling cross-modal matching without extensive training dataOvercoming limitations of spatial feature-based methodsRobust feature matching across noisy video streams

SMRABooth: Subject and Motion Representation Alignment for Customized Video Generation

Dec 13, 2025
XX
Xuancheng Xu
🏛️ Nanjing University of Posts and Telecommunications | Peng Cheng Laboratory

Existing custom video generation methods struggle to simultaneously preserve subject appearance fidelity and ensure temporal motion consistency, primarily due to the lack of object-level subject-motion disentangled modeling. This paper proposes a subject-motion representation disentanglement and alignment framework for text-to-video generation. We introduce the first object-level representation alignment mechanism, design a sparse spatiotemporal LoRA injection strategy to minimize fine-tuning interference, and develop a collaborative self-supervised subject encoder and optical-flow-based motion encoder. The method integrates self-supervised representation learning, optical-flow-driven motion modeling, efficient LoRA-based fine-tuning, and spatiotemporally sparse adapters. Evaluated on multiple benchmarks, it achieves significant improvements in subject similarity (+12.6%) and motion consistency (+9.8%), enabling fine-grained, disentangled controllable generation with both high visual fidelity and temporal stability.

Ensuring subject appearance similarity in customized video generationMaintaining temporally consistent motion patterns from reference videosReducing interference between subject and motion representations alignment

Reenact Anything: Semantic Video Motion Transfer Using Motion-Textual Inversion

Aug 01, 2024
MK
Manuel Kansy
🏛️ ETH Zürich | DisneyResearch|Studios

Existing video motion editing methods are largely confined to simple transformations (e.g., translation, scaling) and struggle to accurately transfer complex semantic motions—such as full-body gestures, facial expressions, object dynamics, or camera motion—using only a single reference video. This work proposes a semantic-level video motion transfer framework that enables precise motion transfer from one reference video to arbitrary target images—without requiring spatial alignment. Our approach leverages a pre-trained image-video diffusion model and introduces three key innovations: (1) motion-textual inversion, a novel technique that decouples appearance and motion representations via joint embedding of motion and textual semantics; (2) frame-wise dilated motion embeddings for high-temporal-resolution motion encoding; and (3) implicit motion encoding coupled with cross-attention modulation to enable zero-shot cross-domain transfer. Extensive experiments demonstrate significant improvements over state-of-the-art methods in motion fidelity, generalizability, and temporal consistency.

Disentangling appearance and motion in video generationEnhancing motion granularity using motion-textual inversionTransferring motion from reference video to target image

Latest Papers

What's happening recently
View more

This work investigates the role of long-term motion cues in visual perception tasks and their advantages over static image representations. The authors construct a low-dimensional, efficient motion representation based on point trajectory estimation and systematically evaluate its performance across multiple tasks, including action recognition, object understanding, material classification, and spatial reasoning. The study demonstrates that this motion-based representation captures rich semantic information and exhibits significantly stronger generalization than conventional image features under low-data and zero-shot settings. Furthermore, when fused with standard video representations, it yields additional accuracy gains. Notably, the proposed approach achieves a superior trade-off between computational efficiency (measured in GFLOPs) and performance, highlighting its potential for efficient visual understanding.

long-term motionmotion representationtemporal information