Institution profile

AMAP

Industry researchasia · cn
Official website
Research library14linked papers
Opportunities0open roles
Selected work

Representative Papers

SimpleTouch: Can Vision-Language-Action Models Master Contact-Rich Manipulation Without Tactile Policy Pretraining?

Oct 02, 2026

This study addresses the data and computational bottlenecks arising from integrating tactile sensing into large-scale pretraining for Vision-Language-Action (VLA) models by proposing SimpleTouch. Built upon the π0.5 architecture, this method eliminates complex multi-stage alignment pipelines by freezing the tactile encoder and incorporating multi-horizon latent action prediction, enabling contact-rich manipulation through single-stage action-supervised fine-tuning. Experimental results demonstrate that SimpleTouch achieves an average success rate of 77.5% on the UniVTAC benchmark using only 50 demonstrations—outperforming FTP-π0.5 by 32.3%—and attains a 71.3% success rate in real-world tasks. These findings significantly surpass existing baselines, validating the efficiency of directly fine-tuning pretrained representations for tactile-dependent robotic manipulation without extensive retraining.

0 citationsRead paper

In-context Robot Learning Made Simple: A Democratized Recipe for Manipulation Tasks

Sep 29, 2026

This study addresses the ambiguity of visual demonstrations and the absence of explicit task definitions in robotic in-context learning by formally defining this problem for the first time and proposing the SimpleICL framework. Methodologically, it introduces a visual prompt encoder alongside a low-cost data collection protocol to clarify learning objectives, enabling manipulation task inference without large-scale pretraining. Experiments demonstrate that the proposed framework achieves superior performance in both simulated and real-world environments, revealing critical discriminative properties such as action and semantic sensitivity. By fully open-sourcing the datasets and training pipelines, this work establishes a minimalist and reproducible research paradigm for robotic in-context learning.

0 citationsRead paper

Structured-Prior-Guided Diffusion Inpainting with Physical Consistency for Traffic Sign Augmentation

Sep 02, 2026

"This study addresses the challenges of long-tail distribution and the scarcity of rare traffic signs by proposing a structured prior-guided diffusion framework for generating physically consistent images. The method injects semantic, appearance, and geometric priors through three orthogonal paths and enforces physical consistency via color and edge structure losses. By leveraging JSON text prompts, front-view vector templates, and affine-aligned vector templates, with Stable Diffusion 1.5 as the backbone, this approach outperforms seven state-of-the-art methods in terms of reconstruction fidelity, physical consistency, and semantic controllability. Notably, it achieves a 91.1% OCR accurate match rate, significantly enhancing the detection performance of rare categories."

0 citationsRead paper
Recent publications

Latest Papers

SimpleTouch: Can Vision-Language-Action Models Master Contact-Rich Manipulation Without Tactile Policy Pretraining?

Oct 02, 2026

This study addresses the data and computational bottlenecks arising from integrating tactile sensing into large-scale pretraining for Vision-Language-Action (VLA) models by proposing SimpleTouch. Built upon the π0.5 architecture, this method eliminates complex multi-stage alignment pipelines by freezing the tactile encoder and incorporating multi-horizon latent action prediction, enabling contact-rich manipulation through single-stage action-supervised fine-tuning. Experimental results demonstrate that SimpleTouch achieves an average success rate of 77.5% on the UniVTAC benchmark using only 50 demonstrations—outperforming FTP-π0.5 by 32.3%—and attains a 71.3% success rate in real-world tasks. These findings significantly surpass existing baselines, validating the efficiency of directly fine-tuning pretrained representations for tactile-dependent robotic manipulation without extensive retraining.

0 citationsRead paper

In-context Robot Learning Made Simple: A Democratized Recipe for Manipulation Tasks

Sep 29, 2026

This study addresses the ambiguity of visual demonstrations and the absence of explicit task definitions in robotic in-context learning by formally defining this problem for the first time and proposing the SimpleICL framework. Methodologically, it introduces a visual prompt encoder alongside a low-cost data collection protocol to clarify learning objectives, enabling manipulation task inference without large-scale pretraining. Experiments demonstrate that the proposed framework achieves superior performance in both simulated and real-world environments, revealing critical discriminative properties such as action and semantic sensitivity. By fully open-sourcing the datasets and training pipelines, this work establishes a minimalist and reproducible research paradigm for robotic in-context learning.

0 citationsRead paper

Structured-Prior-Guided Diffusion Inpainting with Physical Consistency for Traffic Sign Augmentation

Sep 02, 2026

"This study addresses the challenges of long-tail distribution and the scarcity of rare traffic signs by proposing a structured prior-guided diffusion framework for generating physically consistent images. The method injects semantic, appearance, and geometric priors through three orthogonal paths and enforces physical consistency via color and edge structure losses. By leveraging JSON text prompts, front-view vector templates, and affine-aligned vector templates, with Stable Diffusion 1.5 as the backbone, this approach outperforms seven state-of-the-art methods in terms of reconstruction fidelity, physical consistency, and semantic controllability. Notably, it achieves a 91.1% OCR accurate match rate, significantly enhancing the detection performance of rare categories."

0 citationsRead paper