HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement

📅 2026-07-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of simultaneously preserving subject fidelity and ensuring interaction plausibility in human-object interaction video generation, a task further complicated by the lack of explicit modeling of internal reference cues such as multi-view consistency. To this end, the authors propose HOMIE, a unified framework that synergistically integrates multimodal large language models (MLLMs) with video generative models to support both cross-subject and single-subject personalized generation. The key innovation lies in incorporating global multimodal guidance into the self-attention mechanism and designing a modality-reference embedding that effectively aligns MLLM-derived semantic features with VAE image tokens, thereby circumventing controllability loss typically incurred by conventional text encoders. Experimental results demonstrate that HOMIE achieves state-of-the-art performance across diverse human-object interaction video generation tasks, significantly enhancing both subject consistency and interaction realism.
📝 Abstract
Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on inter-subject personalization still struggle to strike a balance between high subject fidelity and accurate interaction patterns between humans and diverse objects, especially when objects represent abstract concepts such as logos. Second, while intra-subject references (e.g., OCR maps, multi-view inputs) are expected to enhance subject fidelity, most existing works lack mechanisms to understand such latent correspondence. To address both challenges, we propose HOMIE, an HOCVP framework that tackles both inter- and intra-subject input settings in a unified manner. Compared to previous approaches, HOMIE proposes a better MLLM integration strategy to extract knowledge of reference-level relationships without compromising the controllability of text encoders or incurring costly re-alignment. Specifically, we introduce global multimodal guidance within self-attention to better align MLLM-derived semantic features with VAE tokens. Furthermore, we propose modality-reference embedding to differentiate tokens from MLLM features and VAE tokens and associate intra-subject reference image tokens. Extensive experiments validate that our method achieves state-of-the-art performance across various HOCVP tasks. Project Page: https://yiyangcai.github.io/homie-page.github.io/
Problem

Research questions and friction points this paper is trying to address.

Human-object centric video personalization
subject fidelity
interaction patterns
intra-subject reference
multimodal correspondence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Human-object centric video personalization
Multimodal Large Language Model (MLLM)
Modality-reference embedding
Global multimodal guidance
Subject-driven video generation