vision-language skill sequencing

Designs and builds systems that select and order skills from a predefined library by using vision-language models to condition which skill to run next on visual observations and language-conditioned goals. Implements algorithms that detect task state and skill boundaries and produce long-horizon sequences or plans of skills based on head-camera or other visual input and contextual instructions.

vision-languageskillsequencing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.3
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the vulnerability of Vision-Language-Action (VLA) models to visual shortcuts, which cause them to disregard instructions and struggle to generalize to unseen skill compositions. To overcome this limitation, this work proposes the CRAFT framework, which leverages counterfactual data augmentation to eliminate spurious visual correlations. Furthermore, it introduces a novel skill-representation-based cross-execution supervision transfer mechanism that facilitates effective skill decoupling and reliable reuse. Extensive evaluations across multiple simulation benchmarks and real-world robotic platforms demonstrate that the proposed approach significantly improves execution success rates for undemonstrated skill compositions while preserving high performance on demonstrated tasks. Consequently, this method effectively resolves the combinatorial generalization challenge inherent in VLA models.

Compositional GeneralizationCounterfactual PairsSkill Alignment

This study addresses the data scarcity and generalization bottlenecks encountered by small models in robotic manipulation by proposing an LLM-SLM orchestration framework based on task decomposition and skill composition. The core methodology leverages large language models to drive data synthesis for training small language models, designs a skill-aware context-free grammar (CFG) to constrain the action space, and introduces a progressive orchestration strategy to ensure reliable task planning and natural language instruction parsing. Experimental evaluations conducted on both unmanned aerial vehicles and ground vehicles demonstrate that the proposed approach significantly outperforms existing baselines, substantially enhancing the model's zero-shot generalization capability to unseen tasks.

Dataset ConstructionGeneralizationRobot Operation

Interactive Task Planning with Language Models

Oct 16, 2023
BL
Boyi Li
🏛️ University of California, Berkeley

Existing interactive robotic task planners suffer from poor generalization, reliance on predefined modules, and labor-intensive prompt engineering. Method: This paper proposes a long-horizon task planning framework for real-world scenarios, introducing a novel functional architecture that jointly integrates high-level language-based planning with low-level skill invocation, grounded in vision-language multimodal scene understanding. Contribution/Results: The framework enables dynamic goal interpretation, online replanning, and zero-shot cross-task transfer—requiring only lightweight task instructions without domain-specific fine-tuning or manual prompt design. Evaluated on a real-world bubble tea preparation task, it successfully generates and executes unseen goals, accommodates mid-task user requests for modification, performs precise online replanning, and generalizes to diverse household service scenarios.

Adapting to new tasks without complex engineeringInteractive robot task planning generalizationLanguage models for open-ended planning

Bootstrapping Object-level Planning with Large Language Models

Sep 18, 2024
DP
D. Paulius
🏛️ Brown University | University of Innsbruck

Direct generation of PDDL goals or task sequences by large language models (LLMs) often yields semantically abstract, non-executable outputs, hindering integration with robot task and motion planning (TAMP). Method: We propose an LLM knowledge distillation framework that extracts object-level state-change knowledge via prompt engineering, constructs a function-oriented object network (FOON), and automatically compiles it into semantically aligned, executable PDDL subgoals. Contribution/Results: This FOON-PDDL joint representation establishes the first structured synergy between LLM-derived high-level semantics and classical planners’ action-object constraints. Evaluated on simulated pick-and-place tasks, our approach improves subgoal success rate by 37%, significantly enhancing planning feasibility and cross-task generalization.

Extracts knowledge from LLMs for object-level planning.Generates PDDL subgoals from functional object-oriented networks.Improves task and motion planning in pick-and-place tasks.

Latest Papers

What's happening recently
View more

This study addresses the challenge faced by vision-language models (VLMs) in organizing discrete events from video streams into persistent skill libraries by proposing a streaming embodied skill discovery paradigm. Methodologically, it introduces the Counterfactual Library State Rebalancing (CLaRe) algorithm to optimize skill grouping, which, combined with supervised fine-tuning, enables automatic skill localization, transition, and reuse decision-making. To systematically evaluate these capabilities, the authors construct the Video2Skill benchmark. Experimental results reveal existing skill-scaling bottlenecks in current VLMs and demonstrate that accurately recognizing insufficient available skills constitutes a core challenge for general-purpose planning. Ultimately, this work provides a novel pathway toward constructing reusable embodied skill libraries.

Embodied Skill DiscoveryManipulationSkill Reusability

This work addresses the challenge that existing vision-language-action models struggle to accurately execute natural language instructions involving spatiotemporal and logical constraints, while also lacking interpretability. The authors propose a hierarchical framework that, for the first time, deeply integrates Signal Temporal Logic (STL) between language understanding and robotic execution. The approach decomposes high-level instructions into subtasks and generates verifiable, optimizable, and correctable STL specifications, which dynamically schedule low-level policies. By combining vision-language models, STL, model predictive control, and learned policies, the method enables an end-to-end mapping from natural language instructions to formal specifications, supporting online monitoring and replanning. Experiments in real-world tabletop environments demonstrate significant improvements in accuracy, reliability, and interpretability of language-guided robotic tasks.

interpretabilitynatural language instructionsprecise specification

This study addresses the challenges of large control-level action search spaces, error accumulation, and reliance on manual annotations for symbolic skills in long-horizon planning. To this end, it proposes a flow matching-based latent world model coupled with a hierarchical planning framework. By designing four skill abstraction mechanisms that map noise seeds to symbolic labels, the method automatically generates skill actions without domain knowledge, enabling single-execution state transitions that effectively integrate representation learning with symbolic reasoning. Evaluated on simulated block rearrangement tasks, the proposed unsupervised approach maintains high success rates across multi-skill scenarios, validating the effectiveness of the introduced abstraction mechanisms.

error accumulationlong-horizon planningsearch space

This work addresses two prevailing paradigms in robot learning—“weight internalization” and “code-based skill self-generation”—by clarifying their distinctions and evolutionary trajectories, and tackling open challenges such as ambiguous skill definitions, ill-defined self-improvement mechanisms, and cross-platform adaptability. The authors propose a unified taxonomic framework that articulates five distinct interpretations of “skills,” operationalizes self-improvement for the first time, and systematically evaluates six technical families: vision-language-action models, zero-shot program synthesis, closed-loop self-repair, persistent skill memory, reinforcement learning–based skill discovery, and large language model–driven skill repositories with evolutionary search. A review of 77 representative systems reveals that only a few—including ASPIRE, ENPIRE, and RoboClaw—achieve open-ended self-improvement loops, while highlighting critical gaps in current commercial skill markets regarding adaptability and safety validation.

code-as-policyrobot learningself-improvement

Hot Scholars

KK

Kris Kitani

Carnegie Mellon University, Meta FAIR
Computer VisionAIMachine Learning
HD

Haodong Duan

Shanghai AI Lab | CUHK | PKU
Computer VisionVideo UnderstandingMultimodal LearningGenerative AI
ZP

Zihao Pan

Meituan-M17 LongCat Team; Sun Yat-sen University
Generative ModelingMultimodal Large Language ModelsDiffusion Models