Institution profile

Knowin AI

Industry research
Research library9linked papers
Opportunities0open roles
Selected work

Representative Papers

RoboAware: Learning to Coordinate Embodied Skills from Counterfactual Outcomes

Oct 08, 2026

This study addresses the challenge faced by embodied agents in anticipating which policy family is optimal under current physical states when composing modular skills. To this end, it proposes a hierarchical MDP framework based on the P5 architecture to unify skill semantics, and introduces State-Conditioned Counterfactual Branching (SCB) to compensate for missing observations during training data generation. Furthermore, an Execution-Aware Learning (EAL) mechanism is designed to integrate Monte Carlo Tree Search with Q-learning for value function distillation, enabling dynamic coordination between frozen policies and code generation. Experimental results demonstrate that the proposed approach achieves an overall success rate of 77.0% across 100 tasks and attains state-of-the-art performance of 90.0% on both RoboSuite and RoboTwin benchmarks, significantly outperforming existing baselines.

0 citationsRead paper

Are Frontier VLM Agents Ready to Be Robot Generalists? An Empirical Study with the Embodied Agent Arena

Sep 30, 2026

This study addresses the significant gap between the localized perception capabilities of frontier vision-language models (VLMs) and their ability to execute complete tasks as general-purpose robots. To bridge this divide, we construct Embodied Agent Arena, a benchmark comprising thousands of cases, and introduce the GeoProbe evaluation framework. By employing hybrid validation through Blender rendering and real-world scenes, our approach decouples perception accuracy from task completion via minimal testing, enabling a systematic assessment of seven prominent VLMs in embodied scenarios. This work explicitly delineates the boundary between perceptual precision and functional deployment, revealing fundamental shortcomings of current models in coordinated, goal-directed actions. Ultimately, these findings identify critical challenges that must be addressed to advance toward general-purpose robotic agents.

0 citationsRead paper

Seg3DParts: Segmentation-Grounded Controllable Part-Level 3D Generation

Sep 29, 2026

This study addresses the challenges of occlusion-induced structural ambiguity and insufficient inter-part coherence in single-image part-level 3D generation. To overcome these limitations, this work proposes a controllable 3D mesh generation framework that leverages segmentation as an explicit anchor. Methodologically, it pioneers the use of segmentation as an explicit localization signal for generation and introduces global context exchange alongside structured cross-part attention mechanisms. Combined with a newly constructed large-scale dataset, PartObjectNet, the proposed approach achieves high-quality part assembly without requiring post-processing alignment. Experimental results demonstrate that this method significantly outperforms existing approaches in geometric quality, part coherence, and controllable decomposition.

0 citationsRead paper

Robo-Harness K1: Harnessing Robot-Use Agents via Perception Augmentation

Sep 24, 2026

This study addresses the heavy reliance of Vision-Language-Action (VLA) models on extensive demonstrations and the high cost of pure-RGB control by proposing the Robo-Harness K1 framework. Its core innovation lies in encapsulating 3D perception as a tool interface, enabling agents to select universal actions by querying geometric evidence such as depth and anchor points, thereby achieving robotic manipulation without modifying or fine-tuning the underlying VLM architecture. The proposed method attains an accuracy of 88.9% on the LIBERO-PRO benchmark, significantly outperforming pure-RGB baselines while demonstrating strong robustness against perturbations. Furthermore, through policy distillation, it surpasses OpenVLA under few-shot settings and achieves zero-fine-tuning cross-embodiment generalization across different robotic arms, validating the effectiveness of this perception-augmented paradigm.

0 citationsRead paper

RACaP: Agentic Reasoning, Acting, and Coding as Policies for Evolvable Robot Learning

Sep 24, 2026

This study addresses the high runtime latency and tight coupling between policy reuse and task decision-making inherent in existing "code-as-policy" approaches. We propose RACaP, a framework that shifts code generation to an offline evolutionary phase, deploying frozen, typed policy APIs via a ReAct architecture at runtime to decouple physical mechanism reuse from decision-making. Furthermore, RACaP integrates curriculum learning with rejection sampling fine-tuning to drive autonomous policy evolution, leveraging multimodal memory for failure recovery without source code modification. Evaluated on the LIBERO benchmark, our method significantly outperforms baselines, achieving 45.0% success on LIBERO-PRO while accelerating inference by 13.2× and substantially reducing physical invocation overhead.

0 citationsRead paper
Recent publications

Latest Papers

RoboAware: Learning to Coordinate Embodied Skills from Counterfactual Outcomes

Oct 08, 2026

This study addresses the challenge faced by embodied agents in anticipating which policy family is optimal under current physical states when composing modular skills. To this end, it proposes a hierarchical MDP framework based on the P5 architecture to unify skill semantics, and introduces State-Conditioned Counterfactual Branching (SCB) to compensate for missing observations during training data generation. Furthermore, an Execution-Aware Learning (EAL) mechanism is designed to integrate Monte Carlo Tree Search with Q-learning for value function distillation, enabling dynamic coordination between frozen policies and code generation. Experimental results demonstrate that the proposed approach achieves an overall success rate of 77.0% across 100 tasks and attains state-of-the-art performance of 90.0% on both RoboSuite and RoboTwin benchmarks, significantly outperforming existing baselines.

0 citationsRead paper

Are Frontier VLM Agents Ready to Be Robot Generalists? An Empirical Study with the Embodied Agent Arena

Sep 30, 2026

This study addresses the significant gap between the localized perception capabilities of frontier vision-language models (VLMs) and their ability to execute complete tasks as general-purpose robots. To bridge this divide, we construct Embodied Agent Arena, a benchmark comprising thousands of cases, and introduce the GeoProbe evaluation framework. By employing hybrid validation through Blender rendering and real-world scenes, our approach decouples perception accuracy from task completion via minimal testing, enabling a systematic assessment of seven prominent VLMs in embodied scenarios. This work explicitly delineates the boundary between perceptual precision and functional deployment, revealing fundamental shortcomings of current models in coordinated, goal-directed actions. Ultimately, these findings identify critical challenges that must be addressed to advance toward general-purpose robotic agents.

0 citationsRead paper

Seg3DParts: Segmentation-Grounded Controllable Part-Level 3D Generation

Sep 29, 2026

This study addresses the challenges of occlusion-induced structural ambiguity and insufficient inter-part coherence in single-image part-level 3D generation. To overcome these limitations, this work proposes a controllable 3D mesh generation framework that leverages segmentation as an explicit anchor. Methodologically, it pioneers the use of segmentation as an explicit localization signal for generation and introduces global context exchange alongside structured cross-part attention mechanisms. Combined with a newly constructed large-scale dataset, PartObjectNet, the proposed approach achieves high-quality part assembly without requiring post-processing alignment. Experimental results demonstrate that this method significantly outperforms existing approaches in geometric quality, part coherence, and controllable decomposition.

0 citationsRead paper

Robo-Harness K1: Harnessing Robot-Use Agents via Perception Augmentation

Sep 24, 2026

This study addresses the heavy reliance of Vision-Language-Action (VLA) models on extensive demonstrations and the high cost of pure-RGB control by proposing the Robo-Harness K1 framework. Its core innovation lies in encapsulating 3D perception as a tool interface, enabling agents to select universal actions by querying geometric evidence such as depth and anchor points, thereby achieving robotic manipulation without modifying or fine-tuning the underlying VLM architecture. The proposed method attains an accuracy of 88.9% on the LIBERO-PRO benchmark, significantly outperforming pure-RGB baselines while demonstrating strong robustness against perturbations. Furthermore, through policy distillation, it surpasses OpenVLA under few-shot settings and achieves zero-fine-tuning cross-embodiment generalization across different robotic arms, validating the effectiveness of this perception-augmented paradigm.

0 citationsRead paper

RACaP: Agentic Reasoning, Acting, and Coding as Policies for Evolvable Robot Learning

Sep 24, 2026

This study addresses the high runtime latency and tight coupling between policy reuse and task decision-making inherent in existing "code-as-policy" approaches. We propose RACaP, a framework that shifts code generation to an offline evolutionary phase, deploying frozen, typed policy APIs via a ReAct architecture at runtime to decouple physical mechanism reuse from decision-making. Furthermore, RACaP integrates curriculum learning with rejection sampling fine-tuning to drive autonomous policy evolution, leveraging multimodal memory for failure recovery without source code modification. Evaluated on the LIBERO benchmark, our method significantly outperforms baselines, achieving 45.0% success on LIBERO-PRO while accelerating inference by 13.2× and substantially reducing physical invocation overhead.

0 citationsRead paper