You Don't Need To Stay in The Loop: An Agentic Robotics Loop for Robot-Policy Improvement

📅 2026-08-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the frequent failures of toolchains in robotic policy training and the absence of reliable evaluation and recovery mechanisms. It proposes AgenticRobotics—a backend-agnostic agentic control plane powered by large language models that dynamically orchestrates ephemeral worker nodes to establish a recoverable and verifiable “train–evaluate–refine” transactional loop. Key innovations include an evidence-gated promotion mechanism, commit-key-based crash recovery, and a signed skill repository with a verified tool registry. The system enables zero-loss, zero-duplication execution and valid decision-making at arbitrary points under unattended operation, reducing erroneous promotions from 0.005–0.021 to 0.001 and successfully detecting all six classes of artifact tampering.
📝 Abstract
Coding agents such as Claude Code and Codex close the software loop: a main agent manages the loop, subagents analyze and execute, tools do the work. We port this architecture to robot-policy improvement, where one difference dominates the design: robotic tools---trained policies, training pipelines, data collection---fail routinely, so a tool's quality must be measured, recorded at every call, and expired when the artifact behind it changes. AgenticRobotics is a backend-independent control plane in which an LLM controller drives disposable workers through durable train--evaluate--improve transactions: an immutable objective, controller-owned measurement, commit-keyed crash recovery, an evidence-graded skill library, and a tool registry with a standardized, recorded call surface. The title is an operational claim, not a selection claim: the operator can leave because promotion is evidence-gated, state is recoverable, and capability quality is derived from records---not because the loop picks better checkpoints than a human; on the one lineage we measured, it does not. The gates measurably buy false-promotion control (0.001 per run hardened versus 0.005--0.021 shipped), anytime-valid decisions under optional stopping, zero lost or duplicate effects under kill injection, and six of six artifact-tampering classes caught by a signed verifier.
Problem

Research questions and friction points this paper is trying to address.

robot-policy improvement
tool reliability
automation loop
failure handling
quality measurement
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Robotics
LLM-controlled automation
evidence-gated promotion
tool registry with recorded call surface
crash recovery via commit keys
🔎 Similar Papers
No similar papers found.