Score
Techniques to represent, encode, and inject diverse control modalities (e.g., keyboard/mouse, high‑level instructions, spatial profiles) into generative models so outputs can be steered predictably. This includes designing conditioning formats for diffusion or world generators to support mid‑stream reprompting, distance profiles, and geometry‑aware controls.
Robot manipulation faces three fundamental bottlenecks: data scarcity, difficulty in long-horizon task planning, and weak multimodal reasoning. To address these, this paper proposes the first three-tier generative AI framework for robotic manipulation: (1) a foundational layer for synthetic data and reward generation; (2) a middleware layer for joint language, vision, and state modeling; and (3) a policy layer for grasp and trajectory generation. We systematically benchmark GANs, VAEs, diffusion models, normalizing flows, and autoregressive models across manipulation tasks, delineating their respective applicability boundaries. Drawing on over 100 state-of-the-art works, we rigorously analyze performance limits in data augmentation, cross-modal instruction grounding, and embodied policy learning. Furthermore, we introduce AwesomeGAIManipulation—an open-source resource repository integrating curated papers, benchmarks, and code—to accelerate community progress in generative AI–driven robotics.
The field of motion generation lacks a systematic survey grounded in generative methodology. To address this gap, we propose the first deep taxonomy centered on generative strategies, synthesizing state-of-the-art works from top-tier conferences (CVPR, ICCV, SIGGRAPH, CoRL) since 2023. Our framework uniformly analyzes four dominant paradigms—GANs, autoencoders, autoregressive models, and diffusion models—across three dimensions: architectural design, conditional modeling mechanisms, and evaluation protocols for motion sequence synthesis. We consolidate widely adopted datasets and metrics, identifying key challenges including motion coherence, physical plausibility, and cross-domain generalization. By establishing a comprehensive, comparable analytical benchmark with well-defined dimensions, our survey significantly enhances methodological comparability and facilitates precise problem diagnosis. This work serves as a foundational reference for researchers advancing generative motion modeling.
The rapid advancement of generative AI—including GANs, VAEs, and diffusion models—has led to an overwhelming and fragmented literature, necessitating a systematic synthesis. This survey proposes a unified technical taxonomy that integrates the evolutionary trajectories, architectural variants, and hybridization strategies of these three dominant paradigms, clarifying shared optimization principles for generation quality, diversity, and controllability. It introduces, for the first time, a multi-dimensional classification framework spanning model architecture, training mechanisms, and application domains. Furthermore, incorporating ethical considerations and societal impact, the survey identifies three key frontiers: scalability, trustworthy generation, and human-AI collaboration. By unifying conceptual foundations and highlighting emerging challenges, this work delivers a structured, forward-looking technical roadmap for researchers and practitioners in generative AI.
This work addresses the challenge of efficiently generating high-quality training data required for supervised learning in text-to-image generation models. We propose the Guided Adversarial Prompts (GAP) framework—a closed-loop data generation system integrating three core mechanisms: (1) adversarial prompt optimization guided by supervised model loss, (2) target distribution alignment via feature matching or discriminator-based guidance, and (3) online feedback adaptation. GAP is the first method to synergistically couple adversarial generation with explicit distributional constraints, shifting data synthesis from open-loop, static prompting to closed-loop, adaptive refinement. Empirical evaluation across diverse settings—including multi-task learning, heterogeneous model architectures, and distribution shifts (e.g., spurious correlations, unseen domains)—demonstrates substantial improvements in downstream model generalization. Data utilization efficiency increases by up to 3.2× compared to baseline approaches.
This work addresses the challenge of spatial controllability in image generation models, systematically investigating the unified modeling capability of Transformer-based diffusion, flow, and autoregressive architectures under fine-grained spatial conditions—such as edge maps and pose keypoints. We propose control-token pre-filling as an efficient, general-purpose baseline; identify classifier-free guidance scaling and softmax truncation as critical for improving control consistency; and re-validate adapter-based fine-tuning for mitigating task forgetting. Experiments on ImageNet demonstrate that our approach significantly enhances controllability consistency while preserving high-fidelity generation under data-limited regimes. By decoupling the effects of architecture, training methodology, and guidance strategies, this study establishes a reproducible benchmark framework and provides practical design principles for controllable image synthesis.
Scientific simulations often lack topological controllability in generative modeling. Method: This paper proposes a vector-field topology-guided conditional diffusion model. It is the first to embed topology signals—such as critical point locations and Poincaré indices—encoded via coordinate-based neural networks (SIRENs) into the diffusion denoising process, combined with gradient-guided sampling for explicit, precise control over 2D vector field topology. Contributions/Results: (1) Generated fields strictly satisfy user-specified critical point types and positions—achieving 100% constraint adherence; (2) Topological consistency is rigorously guaranteed while preserving fidelity to the underlying data distribution; (3) Enables topology-aware alignment across ensembles, significantly enhancing scientific exploration efficiency in fluid dynamics and related domains.
CAD modeling remains highly manual, lacking multimodal interaction and automation support. Method: This paper introduces the first end-to-end image-to-parametric-CAD-command-sequence framework for editable and manufacturable 3D shape generation. It innovatively integrates CLIP-style contrastive representation learning, latent diffusion priors, and an autoregressive Transformer architecture to enable image-driven CAD command sequence generation with geometric constraint-aware decoding. Contributions/Results: (1) Generates topologically valid, parameter-tunable, and manufacturing-ready CAD models from a single input image; (2) Enables cross-modal CAD retrieval, improving image-to-model accuracy by 32.7% on large-scale CAD databases; (3) Outperforms all state-of-the-art methods on both unconditional and image-conditioned CAD generation benchmarks. This work advances AI-driven design-to-manufacturing closed-loop automation.
Existing video generation models rely heavily on text prompts, which lack precise spatiotemporal control over dynamic motion and complex action composition. To address this, we propose Motion Prompting—a novel conditioning framework that leverages variable-granularity motion trajectories (sparse/dense, object-level/global/temporal) to enable fine-grained control over camera/object motion, image interaction, motion transfer, and editing. Methodologically, we introduce the first trajectory encoder coupled with a spatiotemporal attention fusion architecture, complemented by motion-guided latent-space optimization and a semantic-driven motion prompt expansion mechanism that automatically maps high-level semantics into detailed motion signals. Quantitative evaluations and human studies across multiple tasks demonstrate significant improvements over state-of-the-art baselines. Generated videos exhibit enhanced physical plausibility and emergent behaviors, establishing a new paradigm for interactive video generation in embodied world modeling.
This work addresses the challenge of achieving precise and controllable image generation with diffusion models in the absence of large-scale annotated data. The authors propose a self-conditioning mechanism leveraging pretrained self-supervised representations, which identifies semantic directions in the representation space to guide the diffusion process without requiring labeled conditions. This approach not only enhances unconditional generation quality but also constructs a smooth and disentangled controllable generation space. Experimental results demonstrate that the proposed method achieves superior performance in image generation and editing tasks, excelling in controllability, smoothness, and disentanglement compared to existing alternatives.
Existing 3D generative models struggle with intuitive and precise geometric control: text prompts are often ambiguous, while image-based editing is cumbersome and inefficient. This paper introduces the first training-free, test-time spatial control framework that supports diverse spatial inputs—from simple primitives to complex meshes—and directly injects them into pretrained 3D generative models for explicit geometric guidance during synthesis. Our method leverages differentiable rendering and feature alignment to integrate geometric priors via a spatial conditioning injection mechanism, enabling plug-and-play integration and adjustable trade-offs between geometric fidelity and visual realism. Experiments demonstrate substantial improvements in geometric accuracy over fine-tuning- or optimization-based baselines. User studies confirm superior intuitiveness and editing efficiency. The framework enables real-time interactive editing—from quadric surfaces to textured 3D assets—without model retraining.
Existing diffusion models struggle to simultaneously achieve high fidelity and compositional consistency under multi-modal conditional control (e.g., text, reference images, pose, and spatial layout). This paper introduces CanvasDiff, a multi-task diffusion generation framework built upon a unified canvas representation. It encodes heterogeneous control signals into a single structured canvas image and incorporates a vision-spatial joint reasoning module alongside a multi-task canvas training strategy to enable end-to-end cross-modal joint modeling. CanvasDiff significantly improves identity preservation, pose accuracy, and layout controllability under complex conditions. It outperforms state-of-the-art methods on challenging tasks including multi-person synthesis, fine-grained pose control, and semantic layout-constrained generation. To foster reproducibility and further research, the code and pretrained models are publicly released.
This work addresses the reliability challenges of multimodal generative models when required to adhere to structured, domain-specific, or safety-critical knowledge. It introduces, for the first time, a four-layer knowledge injection framework grounded in a structural view of the generation process, decomposing it into input/output boundaries, transition functions, intermediate states, and model parameters—corresponding respectively to the surface, trajectory, latent space, and parameter layers. The authors establish principled guidelines for multi-layer combinatorial design and implement knowledge injection methods for the first three layers using diffusion models and multimodal knowledge graphs. Experimental results demonstrate that coordinated injection across these three layers reduces knowledge-violating outputs by 70.97%, confirming both the framework’s efficacy and the complementary roles of its constituent layers.
Existing diffusion-based methods for interactive 3D world generation suffer from excessive parameter counts, high inference step requirements, and unbounded historical context growth—leading to poor real-time performance and limited fine-grained textual control. This paper introduces the first end-to-end, explorable 3D world generation framework supporting single-image or text input and keyboard-driven real-time navigation. Our approach addresses these limitations through three core innovations: (1) a long-video modeling architecture integrating unified context compression with linear attention fusion; (2) a streaming inference mechanism leveraging bidirectional attention distillation and enhanced text embedding guidance; and (3) an event-level, text-guided paradigm for dynamic world evolution. Experiments demonstrate substantial reductions in model parameters and sampling steps, enabling millisecond-scale interactive response while preserving high visual fidelity and ensuring full-text controllability throughout generation.