Score
Designs and evaluates algorithms, retrieval systems, and loss functions that match and retrieve persistent identities across images, video frames, or camera views using natural-language descriptions and multimodal cues (text-based re-identification, language-guided re-id, and spatio-temporal re-identification). Builds training objectives and components—identity-consistency/contrastive/identity-preservation losses, appearance-coherence and temporal-identity losses, location-aware and location-informed matching, and robust pairwise scoring and spatio-temporal priors—to fuse visual, textual, temporal, and location information and stabilize identity across views and time.
Current single-modality person re-identification faces performance bottlenecks under challenging conditions such as low illumination and occlusion. This work systematically reviews the evolution of multimodal person re-identification, encompassing cross-modal tasks including visible-infrared, text-to-image, sketch-based, and non-line-of-sight scenarios. It further introduces, for the first time, a Transformer-based baseline framework for visible-infrared re-identification, which achieves effective matching through cross-modal alignment and modality-invariant feature learning. The proposed comprehensive survey framework and baseline model not only substantially enhance cross-modal retrieval performance but also establish a unified reference benchmark and offer forward-looking directions for future research in the field.
This work addresses the limited generalization of person re-identification models to unseen domains caused by cross-camera viewpoint variations. To this end, we propose the first multimodal joint learning framework tailored for image-based person re-identification. Our approach integrates multi-camera image data with single-camera image–text pairs and jointly optimizes three objectives: person re-identification, image–text matching, and text-guided image reconstruction. This synergistic training strategy effectively enriches the semantic representation of single-camera data and mitigates domain shift. Extensive experiments on multiple cross-domain person re-identification benchmarks demonstrate that our method significantly outperforms existing state-of-the-art approaches, achieving superior generalization performance.
This work addresses the limitations of existing video identity replacement methods, which are largely confined to single-person scenarios, rely on structural controls such as pose or masks, and lack paired data for multi-person settings, thereby hindering flexible multimodal editing. The authors propose Vorch-IR, a unified framework built upon the LTX2 model, which jointly conditions on driving videos, indexed reference images, and textual instructions to enable identity swapping for single subjects, pairs, and backgrounds within a single model—without requiring pose or layout alignment. Key innovations include text-guided reference character selection, an automated paired-data construction pipeline, and a non-autoregressive temporal-overlap inference mechanism capable of generating minute-long videos. Experiments demonstrate that the proposed method significantly outperforms existing approaches in identity fidelity, motion preservation, and temporal coherence.
This work addresses the challenges of identity confusion, unnatural motion, and “copy-paste” artifacts in existing personalized text-to-video generation methods when handling multi-identity scenarios. To overcome these limitations, the authors propose a video diffusion Transformer-based framework for multi-identity customization, incorporating multimodal identity alignment (visual and semantic), a spatially guided identity localization module, and a semantic-aware controller to enhance identity disentanglement and motion naturalness. The approach further introduces joint multi-face encoding, bounding box constraints, and mask regularization loss, alongside the construction of the first large-scale multi-identity video dataset. Experimental results demonstrate that the proposed method significantly outperforms current state-of-the-art approaches in both identity consistency and motion naturalness, enabling high-fidelity generation of videos featuring multiple coordinated characters.
In text-to-video generation, achieving simultaneous human identity consistency, spatial layout coherence, and temporal motion smoothness remains challenging; existing end-to-end approaches suffer from inherent spatiotemporal optimization trade-offs. To address this, we propose a spatiotemporally decoupled two-stage generation framework: first, decomposing the text prompt into spatial (image generation) and temporal (video generation) semantic components; then, introducing a semantic prompt optimization mechanism alongside spatial- and temporal-separate feature modeling to jointly enhance identity fidelity and motion naturalness. Our method achieves second place in the ACM Multimedia Challenge 2025, attaining state-of-the-art performance in human identity consistency, text-video alignment, and overall visual quality.
Addressing the challenges of identity preservation, reliance on fine-tuning, and data scarcity in text-to-video generation, this paper proposes a training-free triple-enhancement framework. First, GPT-4o–driven face-aware prompt enhancement bridges the semantic gap between textual descriptions and visual content. Second, a prompt-aware reference image optimization mechanism improves input consistency. Third, a unified gradient-guided strategy jointly optimizes identity fidelity and spatiotemporal coherence during diffusion model sampling—enabling inference-time refinement without architectural modification. The method requires no model training or fine-tuning. Extensive evaluation on a thousand-video benchmark demonstrates significant improvements in character identity consistency and video quality, outperforming state-of-the-art approaches in both automated metrics and human assessment. It ranked first in the ACM Multimedia 2025 Challenge, validating its strong generalizability and practical applicability.
This work addresses the challenge of distinguishing appearance-similar individuals in video-based person re-identification by proposing an input-aware, scalable mixture-of-experts architecture. The method employs an input-aware routing mechanism to dynamically select relevant experts and integrates a spatiotemporal feature selection strategy to enhance fine-grained discriminative modeling across both spatial and temporal dimensions. The designed module supports flexible expansion with new experts, thereby improving model generalization. Evaluated on two large-scale video person re-identification benchmarks, the proposed approach significantly outperforms existing state-of-the-art methods, achieving leading performance.
This work addresses the challenge that universal multimodal embeddings (UME) often lack effective visual identity discriminability, which hinders performance in tasks such as instance retrieval, re-identification, and identity-consistent content generation. To tackle this issue, we formally define and systematically study the problem of visual identity representation within UME for the first time, proposing a unified VisID modeling framework that jointly optimizes general multimodal and identity-specific representations through an identity-aware sampling mechanism. We further introduce MVEB, the first large-scale multimodal visual identity benchmark, encompassing both real-world and synthetic data, to facilitate training and evaluation. Extensive experiments demonstrate that our approach substantially enhances identity discriminability in UME while preserving strong general multimodal capabilities, confirming its effectiveness and generalization across diverse settings.
This work addresses the challenges of fragmented trajectories and poor identity continuity in thermal multi-object tracking, which stem from weak appearance cues and frequent detection interruptions. Building upon the YOLOv8 and SORT baseline, the authors propose a lightweight post-processing module that integrates online short-gap remapping with offline trajectory relinking driven by spatio-temporal, motion, and boundary cues. This approach significantly enhances identity preservation without relying on complex re-identification models or online association mechanisms. Experimental results on the PBVS thermal MOT benchmark demonstrate that the method improves IDF1 from 82.25 to 84.93 while maintaining stable MOTA, underscoring the critical role of scene-level spatio-temporal consistency in sustaining identity continuity in thermal video tracking.
This work addresses the performance limitations in person re-identification caused by conflicting optimization objectives between image and text modalities, which hinder effective shared representation learning. To resolve this, the authors propose a decoupled two-stage training strategy: first pretraining a single visual encoder on an image-to-image (I2I) task, then introducing textual supervision to optimize the text-to-image (T2I) alignment. By explicitly accounting for the fundamental differences between I2I and T2I tasks, and integrating cross-modal alignment loss with domain-mixed training, the method not only substantially improves T2I generalization but also reciprocally enhances I2I performance, achieving mutual gains across both tasks.