Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search

📅 2026-07-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Although contrastive critics effectively rank actions, their objective functions are prone to selecting out-of-distribution actions during expressive policy search due to embedding norm drift and cosine binding. This work is the first to systematically identify the root cause as a decoupling between ranking capability and value calibration, diagnosing the issue through support set decomposition and disentangling training from readout. Leveraging TD-Q baselines with Bellman-consistent training, OGBench navigation tasks, and simulator rollouts, the study reveals substantial one-step selection costs in PointMaze and Q* tasks, while controllers in AntMaze and HumanoidMaze exhibit self-correction. Notably, even after inference-time normalization, original embeddings retain weak ranking ability. The findings elucidate how training objectives and inference normalization jointly shape ranking performance.
📝 Abstract
Good action rankings do not make a contrastive critic safe to maximize. These critics increasingly act as value-like objectives for best-of-$K$ selection, planning, and critic-guided generation. Unbounded bilinear scores can let large embedding norms inflate off-support values, but cosine bounding does not remove the failure. A controlled support decomposition attributes most raw bilinear regret to norm drift. Cosine and hybrid critics nevertheless select off-support actions from most pools and incur comparable regret. Contrastive scores are weakly calibrated or inverted in the top score decile across four OGBench navigation tasks, and they fail to order fixed-query actions by value. Bellman-trained TD-Q succeeds, including in a parameter-matched function-class control. Realized costs depend on the task: simulator rollouts reveal single-step selection costs on PointMaze and the exact-$Q^*$ toy but well-powered nulls on AntMaze and HumanoidMaze, where the controller can self-correct. A training/readout decomposition traces the lost ordering to the cosine training objective; raw-trained embeddings retain weak ordering after inference-time normalization. Candidate maximization can therefore exploit false positives caused by norm drift, score saturation, or in-support misranking. Contrastive critics remain useful compatibility rankers on navigation and manipulation tasks, but action selection requires a value-calibrated scalar.
Problem

Research questions and friction points this paper is trying to address.

contrastive critics
off-support actions
value calibration
norm drift
action ranking
Innovation

Methods, ideas, or system contributions that make the work stand out.

contrastive critics
norm drift
value calibration
off-support action selection
bilinear scoring
💼 Related Jobs
No related jobs found.
A
Ayushman Singh
Stanford University
S
Siddharth Aphale
Stanford University