ContractRL: Shielded Group-Relative Policy Optimization for Auditable Tool-Call Repair

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the frequent failures of structured tool calling caused by schema violations, where full regeneration incurs prohibitive costs and hinders auditability. We propose a contract-constrained sequential repair protocol that formulates JSON repair as a bounded decision process. Specifically, we leverage RFC-6902 operation masks to filter invalid actions for precise local correction, and introduce a novel contract-constrained group relative objective function that compresses the action space and enhances auditability via an immutable repair history and a residual budget mechanism. The policy is optimized through reinforcement learning with typed feedback. Experiments demonstrate that our approach achieves a semantic success rate of 0.9362 while requiring only 34.4 tokens on average, significantly outperforming baseline and supervised fine-tuning methods in efficiency.
📝 Abstract
Structured tool calls often fail after only a small number of fields violate a schema or an execution contract. Regenerating the complete object enlarges the action surface and makes repeated repair difficult to audit. We introduce ContractRL, a contract-constrained sequential repair protocol that models verifier-guided JSON repair as a bounded decision process. At each step the policy observes the candidate, typed verifier feedback, JSON Pointer, immutable repair history, and remaining budget; a contract-derived action mask filters malformed or prohibited RFC-6902 operations before a deterministic validator performs the transition. We specify a contract-constrained group-relative objective for patch, retry, and abstention decisions while keeping canonical targets and semantic labels outside the online state until trace freeze. Under identical verifier information, ContractRL attains 0.9362 semantic success with 34.4 generated tokens, compared with 0.9076 and 44.9 tokens for Patch-SFT and 0.9148 and 137.2 tokens for full regeneration over 192 cases per seed and five seeds. Policy optimization improves semantic success from 0.9186 for supervised ContractRL to 0.9375. A separate three-seed paired evaluation against Patch-SFT yields a semantic difference of +0.0396 (95% CI $[+0.0137,+0.0662], p=0.0039$). Feedback, action-mask, budget, and schema-shift analyses connect these gains to localized correction, while adversarial and multi-turn evaluations characterize the remaining failure modes.
Problem

Research questions and friction points this paper is trying to address.

Tool-Call Repair
Structured Generation
Auditable Repair
JSON Schema Validation
Contract-Constrained Optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Contract-Constrained Sequential Repair
Group-Relative Policy Optimization
Action Masking
Auditable Tool-Call Repair
Bounded Decision Process
💼 Related Jobs
No related jobs found.
M
Miaobo Hu
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China
S
Shuhao Hu
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
Xiaobo Guo
Xiaobo Guo
Dartmouth College
machine learningdeep learningnatural language processingsocia mediapropagantion
Xin Wang
Xin Wang
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences
Biomedical Engineering
Bokun Wang
Bokun Wang
Texas A&M University
Machine LearningArtificial IntelligenceMultimodal Machine Learning
Y
Yina Sa
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
D
Daren Zha
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
J
Jun Xiao
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China