coarse-to-fine generation

Designs and builds generative pipelines that produce sequences or trajectories by first synthesizing coarse, semantically structured stages (decodable latent states) and then progressively refining those stages into detailed outputs; the systems support inspection and local edits between stages and can be used to simulate or generate agent‑centric, trajectory‑based behaviors.

coarse-to-finegeneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.4
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing generative models typically optimize only the final output, lacking semantic controllability over intermediate generation steps. This work proposes Trajectory Forcing, a novel framework that explicitly models the generation path as a semantically editable trajectory progressing from global layout to fine-grained details. The approach constructs a coarse-to-fine teacher hierarchy based on DINOv2 clustering and trains a hierarchical, step-conditioned flow-matching model to enable structured and editable intermediate state generation. Experiments demonstrate that the framework maintains high-fidelity image synthesis while supporting cross-scale local editing. Furthermore, it introduces trajectory-aware metrics for consistency and controllability, moving beyond conventional endpoint-only evaluations such as FID and thereby addressing their inherent limitations.

controllable generationgenerative pathintermediate dynamics

Existing text-to-simulation methods for generating high-fidelity, executable rare traffic scenarios often suffer from semantic inaccuracies, insufficient iterative refinement, and limited robustness. This work proposes a multi-stage closed-loop framework that leverages large language models (LLMs) to progressively translate natural language descriptions into executable Scenic scripts. The generated scenarios are automatically validated in a simulation environment, which provides structured diagnostic feedback to guide iterative LLM refinement. This approach significantly enhances the semantic accuracy, executability, and diversity of the generated scenarios. Experimental results demonstrate superior performance over existing baselines in rare traffic scenario generation, highlighting improved reliability and automation capabilities.

executable scenario generationnatural language to simulationrare scenario synthesis

AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation

Aug 01, 2024
MH
Mengkang Hu
🏛️ The University of Hong Kong | Microsoft Corporation

To address the weak planning capability and poor generalization of LLM-based agents, this paper proposes AgentGen—a novel framework for agent instruction tuning. First, it constructs a domain-inspired heuristic environment model that automatically generates structurally diverse simulated environments from an inspiration corpus. Second, it introduces a Bidirectional Evolution (Bi-Evol) algorithm that jointly optimizes task difficulty progression—forward-growing task complexity and backward validation—to produce smooth, pedagogically grounded task sequences. Third, it performs environment-task co-driven instruction fine-tuning. Experiments show that AgentGen-finetuned Llama-3.1-8B outperforms GPT-3.5, while Llama-3.1-70B achieves state-of-the-art performance across multi-domain planning benchmarks, with significant gains in stepwise reasoning and cross-environment generalization. The core contributions are: (i) a principled mechanism for generating diverse, semantically rich environments, and (ii) a bidirectional task evolution paradigm that bridges curriculum learning and robust generalization.

Automating environment and task generationEnhancing LLM planning abilitiesImproving task difficulty diversity

AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials

Dec 12, 2024
YX
Yiheng Xu
🏛️ University of Hong Kong | Salesforce Research

High-quality multi-step GUI interaction trajectories for training GUI agents are scarce and prohibitively expensive to annotate manually. Method: This paper proposes a web-tutorial-based automated trajectory synthesis framework: (1) crawling open-source online tutorials and parsing them into structured, multi-step task specifications; (2) orchestrating a vision-language model (VLM) agent to execute tasks and generate trajectories in real GUI environments; and (3) employing a VLM-based evaluator for end-to-end automatic trajectory validation. We introduce “guided replay”—the first paradigm enabling fully automated conversion of unstructured textual tutorials into executable, verifiable GUI trajectories without human annotation. Contribution/Results: Experiments demonstrate that synthesized trajectories significantly improve agent performance in GUI element localization and multi-step planning, outperforming prior methods across multiple benchmarks. Moreover, the per-trajectory data cost is reduced by over an order of magnitude, enabling scalable, low-cost GUI agent training.

Enhances GUI agent performance with multimodal data.Generates web agent trajectories using web tutorials.Reduces data collection costs without human annotation.

Latest Papers

What's happening recently
View more

This work proposes a unified framework for understanding and developing generative artificial intelligence models capable of producing multimodal content, including images, text, video, and molecular structures. Addressing the current fragmentation in generative modeling, the study integrates core methodologies—such as variational autoencoders, generative adversarial networks, diffusion models, and large language models—into a cohesive theoretical system grounded in mathematical principles, architectural design, and mechanisms for controllable generation. This framework not only advances a systematic understanding of multimodal generative processes but also provides robust theoretical foundations and practical pathways for generating high-quality, controllable digital content, with direct implications for applications in scientific discovery and beyond.

ArchitecturesArtificial IntelligenceFoundational Principles

This study systematically investigates the security challenges arising as generative AI transitions from content generation to performing real-world actions, introducing a novel tripartite threat taxonomy encompassing content-level, model-level, and agent-level risks. Through integrated threat modeling, evaluation of technical countermeasures—including detection, watermarking, alignment techniques, and agent-specific safeguards—and analysis of governance structures, the work reveals a pervasive gap between the rapid expansion of attack surfaces and the current state of defensive capabilities. Most existing technical solutions remain contingent on nascent institutional coordination mechanisms that have yet to mature. The research underscores the necessity for parallel evolution of technical and governance approaches and highlights the critical importance of cross-layer collaborative defense strategies to effectively mitigate emerging threats.

Agentic actionAttack surfaceGenerative AI

Existing trajectory synthesis methods predominantly focus on write-intensive, multi-turn tasks, overlooking the read-intensive challenges posed by high evidential loads in single-decision scenarios. This work proposes WRIT, a novel framework that decouples trajectory complexity into two orthogonal axes: the number of write decisions and the evidential burden per decision. WRIT systematically generates write–read-intensive training trajectories through task generation, diverse user behavior modeling, and executable environment simulation, establishing an efficient synthetic pipeline. Remarkably, a 4B-parameter model trained on only 2K WRIT-generated trajectories outperforms GPT-5.1 no-think on the τ²-bench while substantially reducing inference token consumption, thereby demonstrating the efficacy of evidence-driven decision modeling.

evidence-grounded decision makingmulti-turn agentsread-intensive

This work addresses the limitations of existing GUI interaction datasets, which suffer from insufficient coverage and short temporal horizons, making it difficult to capture rare state transitions and complex multi-step operations. To overcome this, the paper introduces the SEE framework, which constructs an explicit UI state transition graph by integrating vision-language models with UI element awareness. Guided by this graph structure, the framework employs path planning and controlled sampling to efficiently synthesize diverse, long-horizon interaction trajectories. This approach avoids unproductive loops and enables reproducible, interpretable data generation. Evaluated across multiple real-world applications, SEE generates high-quality trajectories averaging 14.8 steps in length, significantly improving agent task success rates and generalization to unseen interfaces.

data synthesisGUI agentinteraction trajectories

This work proposes a method for automatically mining readable, structured skill libraries from user GUI interaction trajectories to enhance agent policy performance. The approach comprises a three-stage pipeline: trajectory segmentation, unsupervised clustering to generate candidate skills, and skill-aware policy training, integrating trajectory representation learning, offline reward modeling, and the GRPO algorithm. It presents the first systematic validation of the feasibility of unsupervised extraction of interpretable skills from real-world interactions. Experimental results on the InteraSkill Workflows benchmark show that five out of eight clusters achieve purity above 0.95, yet the resulting policy improvement remains marginal—skill-step accuracy in IW tasks increases only from 18.5% to 20.5%—highlighting a disconnect between skill interpretability and effective policy transfer.

computer-using agentsGUI automationinteraction trajectory

Hot Scholars

PH

Pu Hua

IIIS, Tsinghua University
Robot Learning
TM

Ting Ma

Harbin Institute of Technology (Shenzhen)
Computational neuroscienceneuroimagebrain-computer-interfacemedical image analysis
YQ

Yiming Qin

EPFL, LTS4
generative modelsgraph machine learningdrug discovery
HX

Huazhe Xu

Tsinghua University
Embodied AIReinforcement LearningComputer VisionDeep Learning