🤖 AI Summary
This study addresses the joint optimization challenge of executor allocation and output verification in heterogeneous LLM task graphs by proposing an online decision framework. This work is the first to jointly model execution and asynchronous verification, employing mixed-integer linear programming for coordinated scheduling. Furthermore, it introduces an information-gain-based reward mechanism that dynamically updates model quality estimates to balance immediate costs against long-term benefits. Theoretical analysis establishes that the proposed algorithm achieves a sublinear regret bound under stochastic service quality. Experimental evaluations across four inference benchmarks demonstrate that the method maintains competitive accuracy while reducing average cost and latency by at least 3.17×.
📝 Abstract
Complex reasoning queries can be decomposed into directed acyclic task graphs and distributed across heterogeneous LLMs, reducing latency through parallelism and enabling smaller models to solve complex tasks. In practice, however, the suitability of an LLM for a given subtask may be a priori unknown, and execution alone does not reveal output correctness. We propose JOVE, an online framework that jointly assigns executor LLMs and selects intermediate outputs for paid verification. Verification runs asynchronously and is used to improve future allocations, so the system must balance spending on execution now against learning for later. We study how to optimize this trade-off under a long-term budget and a per-query latency constraint, with stochastic, initially unknown LLM service quality, invocation costs, and execution times. JOVE makes execution and verification decisions by solving a sequence of per-query mixed-integer linear programs. Online learning updates task-dependent estimates of LLM quality based on verification feedback, while an information-gain bonus incorporates the value of learning into allocation decisions. Under a natural set of assumptions, we establish sublinear quality-learning regret for JOVE. Across four reasoning benchmarks, JOVE achieves competitive accuracy against standard inference baselines while reducing average cost and latency by at least 3.17 times.