🤖 AI Summary
This study investigates whether introducing a learned command adapter onto a frozen locomotion policy yields observable and recoverable performance gains. To this end, the authors propose an adapter necessity auditing framework that integrates closed-loop system identification, counterfactual reasoning, cluster refitting, and constraint violation analysis to disentangle deployment gain, state allocation gain, and global operational gain, thereby supporting GO/NO-GO/ABSTAIN deployment decisions. Experiments on the Go2 platform reveal only 0.55% recoverable allocation gain; direct querying yields a NO-GO verdict, while VGCC and MPC-based queries result in ABSTAIN, indicating that the value of an adapter must be grounded in observable evidence rather than prior assumptions.
📝 Abstract
Adding a learned adapter to a frozen, command-conditioned locomotion policy is worthwhile only if the interface exposes improvements that are both real and recoverable from deployment-time observations. We introduce an adapter necessity audit that separates global operating-point gain,same-state counterfactual headroom, deployment gain over a cross-fitted fixed action, and state-allocation gain over a frequency-matched randomized policy. Source-cluster learner refits map these quantities and constraint violations to a GO/NO-GO/ABSTAIN decision. Closed-loop command- response identification provides optional decision features. On Go2, an archived scale-prefix diagnostic finds 5.2% same-state headroom but only 0.55% recovered allocation gain. Our confirmatory audit evaluates direct, scale, heading, and yaw interventions on twenty independent clusters for each of three query distributions induced by direct control, VGCC, and MPC, using 200 full learner refits. At 1% deployment and allocation thresholds and a 5% violation tolerance, direct queries return NO-GO, while VGCC and MPC queries ABSTAIN. VGCC has the largest mean deployment gain (1.34%), but its allocation lower bound is 0.09% and its violation upper bound is 6.25%. A deployment-representative twenty-cluster H1 audit also returns NO-GO, whereas a learner-level synthetic control returns GO. The audit therefore tests whether observable signal justifies state-dependent adaptation rather than presuming that an adapter is valuable.