control signal encoding

Techniques to represent, encode, and inject diverse control modalities (e.g., keyboard/mouse, high‑level instructions, spatial profiles) into generative models so outputs can be steered predictably. This includes designing conditioning formats for diffusion or world generators to support mid‑stream reprompting, distance profiles, and geometry‑aware controls.

controlsignalencoding

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Motion Generation: A Survey of Generative Approaches and Benchmarks

Jul 07, 2025
AK
Aliasghar Khani
🏛️ Autodesk Research

The field of motion generation lacks a systematic survey grounded in generative methodology. To address this gap, we propose the first deep taxonomy centered on generative strategies, synthesizing state-of-the-art works from top-tier conferences (CVPR, ICCV, SIGGRAPH, CoRL) since 2023. Our framework uniformly analyzes four dominant paradigms—GANs, autoencoders, autoregressive models, and diffusion models—across three dimensions: architectural design, conditional modeling mechanisms, and evaluation protocols for motion sequence synthesis. We consolidate widely adopted datasets and metrics, identifying key challenges including motion coherence, physical plausibility, and cross-domain generalization. By establishing a comprehensive, comparable analytical benchmark with well-defined dimensions, our survey significantly enhances methodological comparability and facilitates precise problem diagnosis. This work serves as a foundational reference for researchers advancing generative motion modeling.

Compare advantages and limitations of diverse modeling paradigmsProvide structured review of recent motion generation advancementsSurvey generative approaches for realistic motion sequence synthesis

Generative AI in Depth: A Survey of Recent Advances, Model Variants, and Real-World Applications

Oct 23, 2025
SY
Shamim Yazdani
🏛️ Florida International University | Manipal Institute of Technology | University of Southern California | Indian Statistical Institute | University of Wollongong

The rapid advancement of generative AI—including GANs, VAEs, and diffusion models—has led to an overwhelming and fragmented literature, necessitating a systematic synthesis. This survey proposes a unified technical taxonomy that integrates the evolutionary trajectories, architectural variants, and hybridization strategies of these three dominant paradigms, clarifying shared optimization principles for generation quality, diversity, and controllability. It introduces, for the first time, a multi-dimensional classification framework spanning model architecture, training mechanisms, and application domains. Furthermore, incorporating ethical considerations and societal impact, the survey identifies three key frontiers: scalability, trustworthy generation, and human-AI collaboration. By unifying conceptual foundations and highlighting emerging challenges, this work delivers a structured, forward-looking technical roadmap for researchers and practitioners in generative AI.

Addressing difficulties in tracking rapid progress of generative modelsExamining ethical concerns and societal impacts of synthetic mediaSurveying recent advances in generative AI models and variants

Must-Read Papers

Most classic and influential ideas
View more

Controlled Training Data Generation with Diffusion Models

Mar 22, 2024
TY
Teresa Yeo
🏛️ Swiss Federal Institute of Technology Lausanne | MIT

This work addresses the challenge of efficiently generating high-quality training data required for supervised learning in text-to-image generation models. We propose the Guided Adversarial Prompts (GAP) framework—a closed-loop data generation system integrating three core mechanisms: (1) adversarial prompt optimization guided by supervised model loss, (2) target distribution alignment via feature matching or discriminator-based guidance, and (3) online feedback adaptation. GAP is the first method to synergistically couple adversarial generation with explicit distributional constraints, shifting data synthesis from open-loop, static prompting to closed-loop, adaptive refinement. Empirical evaluation across diverse settings—including multi-task learning, heterogeneous model architectures, and distribution shifts (e.g., spurious correlations, unseen domains)—demonstrates substantial improvements in downstream model generalization. Data utilization efficiency increases by up to 3.2× compared to baseline approaches.

Automate closed-loop feedback for adversarial prompt generationControl text-to-image models for supervised training dataGuide generation to match target data distributions

A Practical Investigation of Spatially-Controlled Image Generation with Transformers

Jul 21, 2025
GX
Guoxuan Xia
🏛️ Huawei Noah’s Ark Lab

This work addresses the challenge of spatial controllability in image generation models, systematically investigating the unified modeling capability of Transformer-based diffusion, flow, and autoregressive architectures under fine-grained spatial conditions—such as edge maps and pose keypoints. We propose control-token pre-filling as an efficient, general-purpose baseline; identify classifier-free guidance scaling and softmax truncation as critical for improving control consistency; and re-validate adapter-based fine-tuning for mitigating task forgetting. Experiments on ImageNet demonstrate that our approach significantly enhances controllability consistency while preserving high-fidelity generation under data-limited regimes. By decoupling the effects of architecture, training methodology, and guidance strategies, this study establishes a reproducible benchmark framework and provides practical design principles for controllable image synthesis.

Addressing knowledge gaps in transformer-based control methodsComparing performance across different generation paradigmsEnabling spatially-controlled image generation with transformers

Scientific simulations often lack topological controllability in generative modeling. Method: This paper proposes a vector-field topology-guided conditional diffusion model. It is the first to embed topology signals—such as critical point locations and Poincaré indices—encoded via coordinate-based neural networks (SIRENs) into the diffusion denoising process, combined with gradient-guided sampling for explicit, precise control over 2D vector field topology. Contributions/Results: (1) Generated fields strictly satisfy user-specified critical point types and positions—achieving 100% constraint adherence; (2) Topological consistency is rigorously guaranteed while preserving fidelity to the underlying data distribution; (3) Enables topology-aware alignment across ensembles, significantly enhancing scientific exploration efficiency in fluid dynamics and related domains.

Control generative models to produce fields with specified topological featuresEnable visual analysis by ensuring outputs adhere to user-defined topologiesReduce computational cost of numerical simulations using generative modeling

CAD modeling remains highly manual, lacking multimodal interaction and automation support. Method: This paper introduces the first end-to-end image-to-parametric-CAD-command-sequence framework for editable and manufacturable 3D shape generation. It innovatively integrates CLIP-style contrastive representation learning, latent diffusion priors, and an autoregressive Transformer architecture to enable image-driven CAD command sequence generation with geometric constraint-aware decoding. Contributions/Results: (1) Generates topologically valid, parameter-tunable, and manufacturing-ready CAD models from a single input image; (2) Enables cross-modal CAD retrieval, improving image-to-model accuracy by 32.7% on large-scale CAD databases; (3) Outperforms all state-of-the-art methods on both unconditional and image-conditioned CAD generation benchmarks. This work advances AI-driven design-to-manufacturing closed-loop automation.

Enabling image-based retrieval of CAD modelsEnhancing modifiability and manufacturability of CAD designsGenerating editable 3D CAD models from images

Motion Prompting: Controlling Video Generation with Motion Trajectories

Dec 03, 2024
DG
Daniel Geng
🏛️ University of Michigan | Google DeepMind | Brown University

Existing video generation models rely heavily on text prompts, which lack precise spatiotemporal control over dynamic motion and complex action composition. To address this, we propose Motion Prompting—a novel conditioning framework that leverages variable-granularity motion trajectories (sparse/dense, object-level/global/temporal) to enable fine-grained control over camera/object motion, image interaction, motion transfer, and editing. Methodologically, we introduce the first trajectory encoder coupled with a spatiotemporal attention fusion architecture, complemented by motion-guided latent-space optimization and a semantic-driven motion prompt expansion mechanism that automatically maps high-level semantics into detailed motion signals. Quantitative evaluations and human studies across multiple tasks demonstrate significant improvements over state-of-the-art baselines. Generated videos exhibit enhanced physical plausibility and emergent behaviors, establishing a new paradigm for interactive video generation in embodied world modeling.

Control video generation using motion trajectories instead of text promptsEncode flexible motion representations for object-specific or global scene motionTranslate high-level user requests into detailed motion prompts for diverse applications

Latest Papers

What's happening recently
View more

This work addresses the challenge of achieving precise and controllable image generation with diffusion models in the absence of large-scale annotated data. The authors propose a self-conditioning mechanism leveraging pretrained self-supervised representations, which identifies semantic directions in the representation space to guide the diffusion process without requiring labeled conditions. This approach not only enhances unconditional generation quality but also constructs a smooth and disentangled controllable generation space. Experimental results demonstrate that the proposed method achieves superior performance in image generation and editing tasks, excelling in controllability, smoothness, and disentanglement compared to existing alternatives.

controllable image generationdiffusion modelsgeneration control

SpaceControl: Introducing Test-Time Spatial Control to 3D Generative Modeling

Dec 04, 2025
EF
Elisabetta Fedele
🏛️ ETH Zurich | Stanford University | Technion

Existing 3D generative models struggle with intuitive and precise geometric control: text prompts are often ambiguous, while image-based editing is cumbersome and inefficient. This paper introduces the first training-free, test-time spatial control framework that supports diverse spatial inputs—from simple primitives to complex meshes—and directly injects them into pretrained 3D generative models for explicit geometric guidance during synthesis. Our method leverages differentiable rendering and feature alignment to integrate geometric priors via a spatial conditioning injection mechanism, enabling plug-and-play integration and adjustable trade-offs between geometric fidelity and visual realism. Experiments demonstrate substantial improvements in geometric accuracy over fine-tuning- or optimization-based baselines. User studies confirm superior intuitiveness and editing efficiency. The framework enables real-time interactive editing—from quadric surfaces to textured 3D assets—without model retraining.

Balances geometric fidelity with visual realism in outputsProvides explicit spatial control for 3D generationUses geometric inputs like primitives or meshes for precision

Canvas-to-Image: Compositional Image Generation with Multimodal Controls

Nov 26, 2025
YD
Yusuf Dalva
🏛️ Snap Inc. | UC Merced | Virginia Tech

Existing diffusion models struggle to simultaneously achieve high fidelity and compositional consistency under multi-modal conditional control (e.g., text, reference images, pose, and spatial layout). This paper introduces CanvasDiff, a multi-task diffusion generation framework built upon a unified canvas representation. It encodes heterogeneous control signals into a single structured canvas image and incorporates a vision-spatial joint reasoning module alongside a multi-task canvas training strategy to enable end-to-end cross-modal joint modeling. CanvasDiff significantly improves identity preservation, pose accuracy, and layout controllability under complex conditions. It outperforms state-of-the-art methods on challenging tasks including multi-person synthesis, fine-grained pose control, and semantic layout-constrained generation. To foster reproducibility and further research, the code and pretrained models are publicly released.

Enables high-fidelity compositional image generation with multimodal controlsGeneralizes to multi-control scenarios through joint multi-task trainingUnifies heterogeneous controls into a single canvas for integrated visual-spatial reasoning

This work addresses the reliability challenges of multimodal generative models when required to adhere to structured, domain-specific, or safety-critical knowledge. It introduces, for the first time, a four-layer knowledge injection framework grounded in a structural view of the generation process, decomposing it into input/output boundaries, transition functions, intermediate states, and model parameters—corresponding respectively to the surface, trajectory, latent space, and parameter layers. The authors establish principled guidelines for multi-layer combinatorial design and implement knowledge injection methods for the first three layers using diffusion models and multimodal knowledge graphs. Experimental results demonstrate that coordinated injection across these three layers reduces knowledge-violating outputs by 70.97%, confirming both the framework’s efficacy and the complementary roles of its constituent layers.

intervention layersiterative generationknowledge infusion

Yume-1.5: A Text-Controlled Interactive World Generation Model

Dec 26, 2025
XM
Xiaofeng Mao
🏛️ Shanghai AI Laboratory | Fudan University | Shanghai Innovation Institute

Existing diffusion-based methods for interactive 3D world generation suffer from excessive parameter counts, high inference step requirements, and unbounded historical context growth—leading to poor real-time performance and limited fine-grained textual control. This paper introduces the first end-to-end, explorable 3D world generation framework supporting single-image or text input and keyboard-driven real-time navigation. Our approach addresses these limitations through three core innovations: (1) a long-video modeling architecture integrating unified context compression with linear attention fusion; (2) a streaming inference mechanism leveraging bidirectional attention distillation and enhanced text embedding guidance; and (3) an event-level, text-guided paradigm for dynamic world evolution. Experiments demonstrate substantial reductions in model parameters and sampling steps, enabling millisecond-scale interactive response while preserving high visual fidelity and ensuring full-text controllability throughout generation.

Enables text-controlled world event generationGenerates interactive worlds from text promptsReduces model size and inference steps for real-time performance

Hot Scholars

JZ

Jialong Zuo

Zhejiang University
Speech SynthesisVoice Conversion
RH

Rongjie Huang

FAIR, Zhejiang University
Multimedia ComputingSpeechNatural Language Processing
ZZ

Zhou Zhao

Zhejiang University
Machine LearningData MiningMultimedia Computing
MF

Minghui Fang

Zhejiang University
SpeechMulti-Modal LearningInformation Retrieval