language-based re-identification

Designs and evaluates algorithms, retrieval systems, and loss functions that match and retrieve persistent identities across images, video frames, or camera views using natural-language descriptions and multimodal cues (text-based re-identification, language-guided re-id, and spatio-temporal re-identification). Builds training objectives and components—identity-consistency/contrastive/identity-preservation losses, appearance-coherence and temporal-identity losses, location-aware and location-informed matching, and robust pairwise scoring and spatio-temporal priors—to fuse visual, textual, temporal, and location information and stabilize identity across views and time.

language-basedre-identification

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.28
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$199K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limited generalization of person re-identification models to unseen domains caused by cross-camera viewpoint variations. To this end, we propose the first multimodal joint learning framework tailored for image-based person re-identification. Our approach integrates multi-camera image data with single-camera image–text pairs and jointly optimizes three objectives: person re-identification, image–text matching, and text-guided image reconstruction. This synergistic training strategy effectively enriches the semantic representation of single-camera data and mitigates domain shift. Extensive experiments on multiple cross-domain person re-identification benchmarks demonstrate that our method significantly outperforms existing state-of-the-art approaches, achieving superior generalization performance.

cross-domaindomain generalizationimage-based Re-ID

This work addresses the limitations of existing video identity replacement methods, which are largely confined to single-person scenarios, rely on structural controls such as pose or masks, and lack paired data for multi-person settings, thereby hindering flexible multimodal editing. The authors propose Vorch-IR, a unified framework built upon the LTX2 model, which jointly conditions on driving videos, indexed reference images, and textual instructions to enable identity swapping for single subjects, pairs, and backgrounds within a single model—without requiring pose or layout alignment. Key innovations include text-guided reference character selection, an automated paired-data construction pipeline, and a non-autoregressive temporal-overlap inference mechanism capable of generating minute-long videos. Experiments demonstrate that the proposed method significantly outperforms existing approaches in identity fidelity, motion preservation, and temporal coherence.

multi-person replacementmultimodal editingpaired training data

This work addresses the challenges of identity confusion, unnatural motion, and “copy-paste” artifacts in existing personalized text-to-video generation methods when handling multi-identity scenarios. To overcome these limitations, the authors propose a video diffusion Transformer-based framework for multi-identity customization, incorporating multimodal identity alignment (visual and semantic), a spatially guided identity localization module, and a semantic-aware controller to enhance identity disentanglement and motion naturalness. The approach further introduces joint multi-face encoding, bounding box constraints, and mask regularization loss, alongside the construction of the first large-scale multi-identity video dataset. Experimental results demonstrate that the proposed method significantly outperforms current state-of-the-art approaches in both identity consistency and motion naturalness, enabling high-fidelity generation of videos featuring multiple coordinated characters.

facial motion naturalnessidentity confusionidentity separation

Identity-Preserving Text-to-Video Generation Guided by Simple yet Effective Spatial-Temporal Decoupled Representations

Jul 07, 2025
YW
Yuji Wang
🏛️ Tencent YouTu Lab | Shanghai Jiao Tong University

In text-to-video generation, achieving simultaneous human identity consistency, spatial layout coherence, and temporal motion smoothness remains challenging; existing end-to-end approaches suffer from inherent spatiotemporal optimization trade-offs. To address this, we propose a spatiotemporally decoupled two-stage generation framework: first, decomposing the text prompt into spatial (image generation) and temporal (video generation) semantic components; then, introducing a semantic prompt optimization mechanism alongside spatial- and temporal-separate feature modeling to jointly enhance identity fidelity and motion naturalness. Our method achieves second place in the ACM Multimedia Challenge 2025, attaining state-of-the-art performance in human identity consistency, text-video alignment, and overall visual quality.

Balancing spatial coherence and temporal smoothness in videosDecoupling spatial and temporal features for consistent video outputSpatial-temporal trade-off in identity-preserving video generation

Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement

Sep 01, 2025
JG
Jiayi Gao
🏛️ Wangxuan Institute of Computer Technology | National Institute of Health Data Science

Addressing the challenges of identity preservation, reliance on fine-tuning, and data scarcity in text-to-video generation, this paper proposes a training-free triple-enhancement framework. First, GPT-4o–driven face-aware prompt enhancement bridges the semantic gap between textual descriptions and visual content. Second, a prompt-aware reference image optimization mechanism improves input consistency. Third, a unified gradient-guided strategy jointly optimizes identity fidelity and spatiotemporal coherence during diffusion model sampling—enabling inference-time refinement without architectural modification. The method requires no model training or fine-tuning. Extensive evaluation on a thousand-video benchmark demonstrates significant improvements in character identity consistency and video quality, outperforming state-of-the-art approaches in both automated metrics and human assessment. It ranked first in the ACM Multimedia 2025 Challenge, validating its strong generalizability and practical applicability.

Bridging semantic gap between video description and reference imageEnhancing identity preservation without costly fine-tuningImproving video quality while maintaining subject fidelity

Latest Papers

What's happening recently
View more

This work addresses the challenge of distinguishing appearance-similar individuals in video-based person re-identification by proposing an input-aware, scalable mixture-of-experts architecture. The method employs an input-aware routing mechanism to dynamically select relevant experts and integrates a spatiotemporal feature selection strategy to enhance fine-grained discriminative modeling across both spatial and temporal dimensions. The designed module supports flexible expansion with new experts, thereby improving model generalization. Evaluated on two large-scale video person re-identification benchmarks, the proposed approach significantly outperforms existing state-of-the-art methods, achieving leading performance.

appearance-similar identitiesfine-grained featuresspatial-temporal information

This work addresses the challenge that universal multimodal embeddings (UME) often lack effective visual identity discriminability, which hinders performance in tasks such as instance retrieval, re-identification, and identity-consistent content generation. To tackle this issue, we formally define and systematically study the problem of visual identity representation within UME for the first time, proposing a unified VisID modeling framework that jointly optimizes general multimodal and identity-specific representations through an identity-aware sampling mechanism. We further introduce MVEB, the first large-scale multimodal visual identity benchmark, encompassing both real-world and synthetic data, to facilitate training and evaluation. Extensive experiments demonstrate that our approach substantially enhances identity discriminability in UME while preserving strong general multimodal capabilities, confirming its effectiveness and generalization across diverse settings.

identity preservationinstance retrievalre-identification

This work addresses the challenges of fragmented trajectories and poor identity continuity in thermal multi-object tracking, which stem from weak appearance cues and frequent detection interruptions. Building upon the YOLOv8 and SORT baseline, the authors propose a lightweight post-processing module that integrates online short-gap remapping with offline trajectory relinking driven by spatio-temporal, motion, and boundary cues. This approach significantly enhances identity preservation without relying on complex re-identification models or online association mechanisms. Experimental results on the PBVS thermal MOT benchmark demonstrate that the method improves IDF1 from 82.25 to 84.93 while maintaining stable MOTA, underscoring the critical role of scene-level spatio-temporal consistency in sustaining identity continuity in thermal video tracking.

appearance cuesdetection interruptionsidentity continuity

This work addresses the performance limitations in person re-identification caused by conflicting optimization objectives between image and text modalities, which hinder effective shared representation learning. To resolve this, the authors propose a decoupled two-stage training strategy: first pretraining a single visual encoder on an image-to-image (I2I) task, then introducing textual supervision to optimize the text-to-image (T2I) alignment. By explicitly accounting for the fundamental differences between I2I and T2I tasks, and integrating cross-modal alignment loss with domain-mixed training, the method not only substantially improves T2I generalization but also reciprocally enhances I2I performance, achieving mutual gains across both tasks.

cross-modal retrievalmodality discrepancyoptimization conflict

Hot Scholars

DC

Daniel Cohen-Or

Professor of Computer Science, Tel Aviv University
GraphicsImagingGeometric Modeling
BK

Bernhard Kainz

FAU Erlangen-Nürnberg, Imperial College London
human-in-the-loop computingmachine learningmedical image analysis
XH

Xiaobin Hu

Tencent Youtu Lab;Technische Universität München (TUM)
Deep learningComputer visionVLMAgents
YL

Yebin Liu

Professor, Tsinghua University
Computer GraphicsComputational Photography3D VisionDigital Humans