Geometry Conditioning in an Embodied SLM: Training Controls and Robustness Diagnostics in a 0.8B Hybrid Model

📅 2026-09-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过不同几何条件训练0.8B混合语言模型,以解决物体操纵任务中的物理状态输入影响问题,但未发现训练时几何对齐的可靠优势。
📝 Abstract
We study how physical-state inputs affect a 0.8B hybrid language model adapted for manipulation with 6.2M trainable parameters. Six conditions are trained on three LIBERO-Spatial tasks and evaluated over three seeds and 540 held-out rollouts. Conditioning recurrent decay gates on geometric increments yields 28.9% success, compared with 36.7% when those increments are shuffled during training and 24.4% without explicit object/goal geometry. Both geometry policies receive correct inputs at evaluation. A token adapter using the same increments scores 27.8%; differences vary across seeds and remain inconclusive. Token-clock conditioning scores 11.1%, including one seed that fails to converge. In separate robustness tests, a state-only relative-coordinate policy retains 7/10 success under frame relabeling, whereas all four tested visual policies fall to at most 3/20 after a 5 cm object displacement. These results show no reliable advantage from training-time geometric alignment under this recipe and illustrate the gap between coordinate invariance and physical-layout generalization. Episode records, seed-level analyses, and figure-generation code accompany the paper.
Problem

Research questions and friction points this paper is trying to address.

physical-state inputs
hybrid language model
manipulation
geometric increments
robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Geometry Conditioning
Hybrid Language Model
Manipulation Tasks
Recurrent Decay Gates
Physical-state Inputs
🔎 Similar Papers
No similar papers found.
H
Hao Li
H
Haofei Sun
L
Lin He
University of Tennessee, Knoxville