Canopy: Exploiting Piecewise Smooth Tree Priors for Multi-Fidelity Bandits

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing tree-based optimization methods for LLM inference, which rely on global smoothness assumptions and struggle to efficiently handle piecewise-smooth objectives. We propose CANOPY, a multi-fidelity tree-arm algorithm that eliminates the need for predefined smoothness schedules. By leveraging online bias certificate aggregation for detection and randomized path probing, CANOPY adaptively localizes discontinuous regions and dynamically allocates evaluation resources. The method is accompanied by theoretical guarantees established through fixed-budget regret analysis. Experimental results demonstrate that CANOPY improves Top-10 recall by 2.9×, increases the number of resolved tasks on SWE-bench by 1.6×, and reduces first-token latency by 3.6×, highlighting its effectiveness in optimizing complex LLM inference workloads.
📝 Abstract
Many LLM inference problems, including model routing, prefix-cache management, prompt trimming, and test-time search, can be viewed as optimization over a tree. This structure arises naturally from autoregressive generation: every prefix defines a node, and its continuations form a subtree below it. Internal nodes of the tree provide cheap but biased estimates of a region's value, while leaf evaluations are expensive but accurate. Hierarchical bandit methods can exploit this structure, but typically require a specific smoothness schedule to be specified in advance, even though real objectives are often only piecewise smooth and their optima may lie near sharp boundaries. We introduce CANOPY, a multi-fidelity tree bandit that learns where the smoothness prior is valid rather than assuming it globally. CANOPY uses cheap random-path probes to construct an online certificate of local aggregation bias, then directs expensive leaf evaluations toward cells where the certificate detects a smoothness violation. We prove fixed-budget and regret guarantees whose additional cost is additive in the number of discontinuities, recovering the smooth-tree rate when no violations are present and approaching structure-blind search as violations become dense. Across routing, top-$k$ identification, test-time search, caching, and prompt trimming, CANOPY consistently improves matched-budget performance, including $2.9\times$ higher top-10 recall on a 1000-model pool, $1.6\times$ more SWE-bench Verified issues resolved than best-of-$N$, and $3.6\times$ lower median time-to-first-token with prefix caching.
Problem

Research questions and friction points this paper is trying to address.

Multi-Fidelity Bandits
Tree-Structured Optimization
Piecewise Smoothness
LLM Inference
Hierarchical Bandits
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Fidelity Bandits
Piecewise Smooth Tree Priors
Online Certificate
Local Aggregation Bias
LLM Inference Optimization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.