Score
Design, build, and evaluate inference pipelines and gating modules that route inputs or computation paths based on estimated uncertainty type or level, including uncertainty estimators, dynamic inference gates, selective computation routers, and rejection/deferral mechanisms. Analyze routing policies and their effects on retaining ambiguous in-distribution examples, rejecting out-of-distribution inputs, and balancing trade-offs among accuracy, latency, and computational cost.
This work addresses key challenges in model routing—namely, the trade-off between performance and cost, difficulty handling ambiguous queries, and limited adaptability to diverse loss–cost configurations. The authors propose an uncertainty-aware routing mechanism that decomposes total predictive uncertainty into reducible and irreducible components for the first time. Without requiring model retraining, the method dynamically selects actions: invoking a weak model under low uncertainty, a strong model when reducible uncertainty is high, and abstaining when irreducible uncertainty dominates. The framework flexibly adapts to arbitrary cost and loss functions via tunable hyperparameters and provides a theoretical regret bound relative to the optimal task-specific router. Experiments demonstrate that, particularly when reducible and irreducible uncertainties are weakly correlated, the approach significantly outperforms baselines on both synthetic and real-world datasets while effectively balancing system performance and computational cost.
This paper addresses the deployment of multi-model inference pipelines on resource-constrained edge devices. We propose an end-to-end adaptive configuration framework that, for the first time, explicitly incorporates device resource constraints into joint optimization decisions. Our method integrates residual feature extraction, LSTM-based workload forecasting, and policy-gradient reinforcement learning to jointly optimize QoS guarantees (e.g., latency and throughput), operational cost, and real-time adaptability. Evaluated on a real Kubernetes-based edge cluster, the framework achieves a 27% reduction in average inference latency, a 31% increase in throughput, a 22% decrease in deployment cost, and over 42% faster configuration decision-making for complex pipelines—outperforming state-of-the-art baselines. The core contributions are: (i) a resource-aware joint optimization model that unifies hardware constraints with pipeline scheduling and scaling decisions; and (ii) a lightweight, learning-driven configuration mechanism enabling efficient, online adaptation under dynamic edge conditions.
This work addresses the challenge of cascading failures in multi-step reasoning, where a single erroneous step can compromise the entire solution. Existing model routing approaches treat all reasoning steps uniformly, leading to suboptimal efficiency. To overcome this, we propose TRIM, the first method enabling dynamic, step-level routing: leveraging a process reward model to identify high-uncertainty, critical steps, TRIM selectively delegates only these error-prone steps to a large language model under a computational budget, while simpler steps are handled by a smaller, more efficient model. This strategy effectively interrupts error propagation. Experiments demonstrate that TRIM achieves up to 5× higher cost efficiency than prior methods on MATH-500, with advanced routing strategies matching baseline performance using 80% fewer large-model tokens. On challenging benchmarks like AIME, TRIM delivers up to 6× cost efficiency gains while maintaining strong generalization.
This work addresses the challenge of balancing energy supply and stochastic fluctuations in scheduling heterogeneous large inference models. It proposes a variance-aware routing strategy that optimizes model selection and execution under critical operating conditions to minimize system energy consumption. For the first time, a variance-absorption mechanism is integrated into model routing design, leveraging second-order performance characteristics within critical energy consumption intervals to establish a theoretical foundation for energy-aware scheduling. By combining computational scaling laws, stochastic process modeling, and critical systems theory, the study reveals fundamental principles governing how system performance is constrained by energy variability, thereby offering quantifiable design guidelines for energy-efficient large-model scheduling.
This work addresses the inefficiency and accuracy trade-off in existing encrypted neural network inference systems, which rely on fixed models for all inputs, resulting in high latency and cost. The paper proposes the first end-to-end encrypted dynamic routing framework that adaptively selects the optimal model size for Transformer inference directly on ciphertext, based on input features. By integrating secure multi-party computation (MPC), an MPC-overhead-aware encrypted router, a co-optimized model pool, and quantization strategies, the framework unifies secure routing, inference, and protocol execution while preserving the confidentiality of both data and models. Experimental results demonstrate that the approach reduces inference latency by up to 1.95× compared to state-of-the-art methods with negligible accuracy loss, offering a practical solution for scalable and secure AI inference.
This work reveals a novel attack surface in edge-cloud collaborative dual-path distributed inference systems: malicious “oscillating burst” traffic can induce resource contention in the slow path, causing benign requests to time out and be dropped, thereby triggering an “accuracy collapse” wherein the system degrades to low-precision fast-path outputs. Notably, this attack requires no access to the model or data and operates solely through network-level interference, significantly degrading perceptual performance. Evaluated on a multi-object tracking simulation platform under autonomous driving scenarios, approximately 4,000 burst requests increased the p99 latency for benign users from 92 ms to 2 seconds, reduced average HOTA by 7.0 points, and caused nearly 50% accuracy loss on rare classes such as stop signs.
This work addresses the challenge that existing single-model attacks struggle to exploit path selection mechanisms in dynamic inference pipelines to amplify system overhead. It formally introduces the problem of “adversarial path selection,” modeling inter-component couplings to steer inputs toward high-computation paths, thereby significantly increasing resource consumption. The proposed attack strategy integrates vulnerability-aware path ranking with adaptive loss weighting and accommodates practical deployment constraints such as batching and buffer limits. Experiments demonstrate that under white-box and gray-box settings, the attack inflates computational costs by up to 2407× FLOPs (419× latency) and 58× FLOPs (17× latency), respectively. Even against system-level defenses, it can reduce throughput to as low as 0.006 inputs/second or induce 96.7% data loss, underscoring the severe threat posed by path-level adversarial exploitation.
This work addresses a critical limitation in existing dynamic routing methods, which conflate irreducible ambiguity with recoverable risk that can be mitigated by ensembling more experts. The authors formalize routing as an information-value allocation problem, introducing a novel mechanism that generates simultaneous upper-bound risk certificates via counterfactual risk estimation. By greedily allocating computational budget based on marginal risk reduction per unit cost, the method decides whether to answer or abstain. This approach is the first to explicitly disentangle the two types of uncertainty and, when integrated with a LoRA-based mixture-of-experts architecture, provides theoretical guarantees on risk certification and optimal resource allocation. Experiments demonstrate significant accuracy gains under identical computational budgets, effective high-coverage risk control, and superior performance over current MoE-LoRA baselines under distribution shifts, tail latency, and risk-coverage trade-offs.
This work addresses the overestimation of deployable router performance in existing large language model (LLM) routing evaluation methods, which suffer from selection bias and information leakage. The authors propose the first selection-valid diagnostic framework that explicitly disentangles three distinct objectives: opportunity, attainability, and realized gain, and constructs post-selection valid confidence intervals. Their approach innovatively integrates a signal information sandwich theorem, Bayesian optimal gain analysis, greedy pool construction under submodular coverage, and multiple hypothesis testing correction. Experiments across four benchmarks reveal that the actual attainable gain constitutes only 7.5%–14.4% of the theoretical opportunity, and while the strongest routers outperform any fixed model, the majority of potential gains remain unrealizable in practice.