Score
Designs, implements, and evaluates mechanisms that condition generative models on subject-specific representations so outputs reflect, persist, and control distinct subjects; this includes embedding- and token-based interfaces, multimodal and vocoder conditioning, memory tokens, and positional-shift techniques (e.g., RoPE-shift). Builds pipelines and diagnostics to inject or retrieve subject tokens, balance training across varying subject counts, and measure or mitigate identity leakage and copy‑paste artifacts.
To address weak data scalability and poor subject generalization in multi-subject image generation, this paper proposes the UNO framework. Methodologically, it first introduces an in-context multi-subject pairing data synthesis approach based on diffusion Transformers to enable zero-shot subject extension. Second, it designs the UNO model, integrating a progressive cross-modal alignment mechanism with universal rotary position encoding to support controllable single- and multi-subject joint generation. Departing from conventional fine-tuning paradigms, UNO enhances data efficiency via iterative multi-image conditional generation and in-context learning. Experiments demonstrate that UNO significantly outperforms existing methods in subject consistency, text fidelity, and layout controllability. Notably, it exhibits strong generalization capability in few-shot and even zero-shot multi-subject scenarios, establishing new state-of-the-art performance in controllable multi-subject image generation.
This work addresses the challenge of jointly achieving fine-grained control over identity and semantic attributes (e.g., pose, style, illumination) in multi-subject text-to-image generation. We propose a reference-image-guided, token-level text-flow modulation method for DiT-based architectures. Specifically, a lightweight image-to-offset mapping network generates reference-driven modulation offsets for each text token, enabling disentangled modeling and independent controllability of identity and semantic attributes. Compared to existing approaches, our method significantly alleviates attribute entanglement and editing artifacts, thereby improving generation fidelity, cross-subject consistency, and editability. Extensive experiments demonstrate superior personalized control and synthesis quality—particularly in complex multi-subject scenarios—while maintaining computational efficiency and architectural compatibility with diffusion transformer backbones.
Existing subject-driven image generation methods often struggle to simultaneously preserve identity and follow textual instructions due to the separate encoding of text and reference images, frequently resulting in copy-paste artifacts. To address this limitation, this work proposes a diffusion-based generation framework that leverages a multimodal large language model (MLLM) to jointly encode textual prompts and reference images. The approach integrates identity conditions extracted via a VAE and dynamically fuses semantic and fine-grained details during the denoising process. A novel dual-layer aggregation (DLA) module is introduced to effectively combine multi-level MLLM features, complemented by a multi-stage denoising strategy that balances semantic fidelity with identity preservation. Experimental results demonstrate that the proposed method substantially mitigates artifact generation and achieves clear superiority over current state-of-the-art approaches in human preference evaluations.
This work addresses the challenge of efficiently generating high-quality training data required for supervised learning in text-to-image generation models. We propose the Guided Adversarial Prompts (GAP) framework—a closed-loop data generation system integrating three core mechanisms: (1) adversarial prompt optimization guided by supervised model loss, (2) target distribution alignment via feature matching or discriminator-based guidance, and (3) online feedback adaptation. GAP is the first method to synergistically couple adversarial generation with explicit distributional constraints, shifting data synthesis from open-loop, static prompting to closed-loop, adaptive refinement. Empirical evaluation across diverse settings—including multi-task learning, heterogeneous model architectures, and distribution shifts (e.g., spurious correlations, unseen domains)—demonstrates substantial improvements in downstream model generalization. Data utilization efficiency increases by up to 3.2× compared to baseline approaches.
Existing text-to-image diffusion models struggle to simultaneously achieve precise spatial localization and fine-grained, continuous control over specific object attributes. To address this, we propose a model-agnostic disentangled control method that requires no architectural modification. We first discover transferable token-level semantic directions within CLIP text embeddings; leveraging this insight, we establish an optimization-free direction identification framework coupled with a model-driven direction modeling mechanism. By applying controlled perturbations in the prompt embedding space, our approach enables parallel, continuous intensity adjustment of multiple attributes for a single subject. This unified framework resolves the inherent trade-off between global semantic control and local attribute localization, significantly improving both accuracy and flexibility of attribute manipulation while preserving subject identity. Code and an interactive demo are publicly available.
This study addresses whether generative AI systems, by virtue of memorizing training data, produce outputs that constitute legally actionable “copies” of copyrighted works. Integrating insights from the memory mechanisms and probabilistic generation behaviors of large language models, the paper offers the first systematic interdisciplinary analysis arguing that copyright law should adopt a functional standard to determine whether an AI model contains a “copy.” The research demonstrates that current legal frameworks typically recognize copying only when specific protected works can be readily extracted from the model, thereby exposing significant limitations in the applicability of existing doctrines in the AI era. Building on this finding, the work proposes targeted legal reforms to better align copyright enforcement with the technical realities of modern generative systems.
This study addresses the difficulty of attributing model behaviors to their origins due to the lack of generative provenance in synthetic speech data. We propose a compact provenance contract and auditing protocol, formally establishing provenance as a necessary but insufficient condition for behavioral attribution. Methodologically, we construct synthetic research objects that bind source specifications to content within a Japanese nursing care scenario, implementing audits through immutable manifests, disjoint versioning of scenario seeds, and multimodal asset linkage. Experimentally, we audit 1.55 hours of speech, revealing impediments to precise upstream attribution and establishing a candidate causal graph framework. This work provides a novel paradigm for enhancing the traceability and causal analysis of synthetic data.
本文提出MSR方法,通过为每个参考图像分配独立的潜在标记组,并使用时隙感知条件方案来解决视频生成中多图像条件下的外观保持问题。
This study addresses the limitation of existing multi-subject image generation methods, which merely verify subject presence without ensuring that attributes, actions, and relationships are correctly bound to their reference subjects. To overcome this, the authors propose Visual Jev, a reference-binding-based visual reward mechanism. Leveraging the Qwen3.5-4B verifier and the MICo-150K dataset, binary supervision signals are constructed offline through fixed question-answering pairs. The probability of affirmative responses from the language model is then utilized as a reinforcement learning reward within the GRPO framework, effectively translating visual judgments into training signals. Evaluated on an 897-task subset, the approach improves the GPT-5.4 composite score from 41.78 to 52.50. While the work provides a comprehensive implementation and evaluation framework, the statistical significance of these results warrants further verification.
This work proposes a general-purpose physics-based character controller capable of executing tasks with natural, realistic, and diverse motions. The approach discretizes motion data using Finite Scalar Quantization (FSQ) and integrates a GPT-style autoregressive Transformer with end-to-end reinforcement learning to build a transferable generative motion controller amenable to downstream task fine-tuning. A key innovation lies in the joint optimization of the discrete action vocabulary and the control policy, replacing conventional pipeline-based training procedures. Experimental results demonstrate that the method achieves a 99.98% motion reproduction success rate on large-scale motion datasets and exhibits robust behaviors such as perturbation response and fall recovery, proving effective across a range of downstream control tasks.