FORGE: Form-Optimal Routing of Grounded Evidence for Frozen LLM Agents

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance degradation and high token costs in frozen large language models caused by suboptimal evidence formatting and misallocated reasoning budgets. To this end, it proposes FORGE, a framework introducing the first joint action-space routing mechanism tailored for frozen models. By integrating a factorized router, an entropy-regularized utility objective, a closed-form Boltzmann policy, and Group Relative Policy Optimization (GRPO), FORGE jointly optimizes retrieved evidence formats and chain-of-thought depth without requiring access to model weights. Experiments across five benchmarks and eight backbone architectures demonstrate that the proposed method simultaneously improves accuracy while reducing token consumption by 42–45%. Furthermore, FORGE exhibits zero-shot cross-host transferability.
📝 Abstract
In agentic AI systems, frozen foundation models are increasingly deployed as closed-weight API endpoints, making downstream adaptation possible only through the inputs and inference procedures surrounding the model. As a result, for each input query, two coupled decisions largely determine both answer quality and token cost: what evidence to provide and how much reasoning budget to allocate. Fixed defaults along these axes are often suboptimal, misallocating support form or reasoning depth on roughly 80% of queries in our analysis. To address this challenge, we propose FORGE, a unified framework for adapting frozen models through per-query routing over a joint action space that spans both support form and thinking depth. Under an entropy-regularized, cost-aware utility objective, we derive a closed-form Boltzmann routing target and instantiate the policy as a lightweight 269K-parameter factorized router. The routing policy is trained around the frozen host, without any weight access, through a three-stage pipeline: offline arm enumeration, supervised Kullback-Leibler (KL) distillation from the Boltzmann target, and Group Relative Policy Optimization (GRPO) refinement with host feedback. Across 5 knowledge-intensive benchmarks and 8 frozen backbones ranging from 7B to 671B parameters, FORGE improves accuracy at 42-45% lower token cost on both main hosts, transfers zero-shot across hosts at lower token cost, and composes with intrinsic thinking budgets where available.
Problem

Research questions and friction points this paper is trying to address.

frozen LLMs
evidence routing
reasoning budget
token cost optimization
agentic AI
Innovation

Methods, ideas, or system contributions that make the work stand out.

Frozen LLM Routing
Joint Action Space
Lightweight Factorized Router
GRPO Refinement
Cost-aware Optimization
🔎 Similar Papers
No similar papers found.