instance-level model routing

Designs and evaluates systems that decide, per input instance, whether to route the request to a cheaper model, to an early-exit branch, or to a more expensive model before generating the full output. This includes building decision rules or classifiers that jointly assess draft and target outputs to predict adequacy and manage cost–accuracy tradeoffs to minimize inference cost while preserving required accuracy.

instance-levelmodelrouting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.84
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

论文解决了LLM评估管道中选择和调度评委的问题,通过角色条件分配方法估计评委角色并制定策略,以优化评委组合。

allocation problemconditional informationjudge panel

Decision-Theoretic Approaches in Learning-Augmented Algorithms

Jan 29, 2025
SA
Spyros Angelopoulos
🏛️ Sorbonne University | International Laboratory on Learning Systems

Evaluating learning-augmented online algorithms under uncertainty remains challenging, as conventional metrics focus narrowly on worst-case prediction errors, neglecting both prediction accuracy and risk sensitivity. Method: We propose a dual-track evaluation framework grounded in decision theory, jointly incorporating distance-based prediction error quantification (deterministic aspect) and risk-sensitive modeling (stochastic aspect). By embedding decision-theoretic loss functions into online algorithm analysis, we integrate prediction error modeling with risk-controllable optimization, designing novel learning-augmented algorithms for contract scheduling and 1-max search. Contribution/Results: Our approach achieves provable robustness to prediction errors, performance guarantees with tight bounds, and explicit risk controllability. It is the first to unify prediction accuracy, worst-case robustness, and risk preference within a single theoretical framework—establishing a systematic evaluation paradigm and design principle for learning-augmented online algorithms.

Decision TheoryMachine LearningOptimal Balance

This paper challenges the unverified implicit assumption in the predict-then-optimize paradigm that “higher prediction accuracy necessarily yields better downstream decisions,” particularly in multiclass classification settings. Method: We propose a controllable, interpretable multiclass prediction simulation framework that explicitly models error types and distributions, enabling systematic analysis of how classification errors affect decision quality in constrained optimization. Contribution/Results: Experiments on job scheduling and other combinatorial optimization tasks reveal a nonlinear relationship between prediction error and decision performance: improving prediction accuracy does not guarantee improved solution quality—and can even degrade decisions when error patterns shift. Our findings question the conventional coupling logic between prediction and optimization, providing theoretical foundations and practical guidance for designing, evaluating, and calibrating classifiers specifically tailored to decision objectives.

Assessing Predict-Then-Optimize performance in machine scheduling problemsEvaluating how prediction error affects optimization solution qualitySimulating multiclass classifier predictions for experimental analysis

This study addresses whether the precipitous decline in machine evaluation costs, driven by generative AI’s reduction of production costs, may trigger a Jevons paradox in organizational evaluation consumption. Building upon the TypeSafe AI Jev model, this work pioneers the introduction of rebound economics into the domain of machine evaluation. By integrating probabilistic decision models with sociological frameworks of organizational authority, it systematically delineates the functional distinctions and cost asynchronies among prediction, evaluation, judgment, and delegation. Furthermore, this paper proposes the Jevons hypothesis for machine evaluation, elucidating the boundary conditions under which inexpensive machine evaluation substitutes human labor or generates novel demand. Ultimately, it reveals the profound mechanisms through which low-cost evaluation reshapes the allocation of organizational decision rights and the foundational basis of commitment within institutions.

generative AIJevons paradoxmachine evaluation

EERO: Early Exit with Reject Option for Efficient Classification with limited budget

Feb 06, 2024
FV
Florian Valade
🏛️ LAMA | Université Gustave Eiffel | Université De Pau Et Des Pays De L’adour

Under constrained computational budgets, balancing classification accuracy and inference efficiency remains challenging. Method: This paper proposes a协同 mechanism of early exit and rejection, the first to formulate early exiting as a rejection-aware multiclass classification problem. It integrates Bayesian risk minimization with head-level budget consumption modeling to enable end-to-end inference under hard budget constraints. The approach employs exponential weighted probability calibration, confidence-aware exit policies, and budget-aware aggregation, evaluated on ResNet-18 and ConvNeXt. Contribution/Results: Experiments on CIFAR and ImageNet demonstrate significant improvements in accuracy under fixed budget constraints, effectively mitigating the “overthinking” phenomenon—where excessive computation yields diminishing returns—while maintaining high accuracy and substantially improving inference efficiency.

Ensuring budget constraints while maintaining classification accuracyManaging computational budget for early exit strategies in classificationSelecting optimal exit heads using multiple classifiers with reject option

Latest Papers

What's happening recently
View more

This work addresses the optimal allocation of post-training compute resources for reinforcement learning under a fixed FLOP budget. It introduces the first accounting framework that explicitly decomposes post-training computation into rollout/search, policy updates, and reward model evaluation, systematically quantifying the trade-offs among model scale, search intensity, number of learning steps, and feedback quality. Using GRPO with LoRA fine-tuning on the Qwen2.5 model family and combining rule-based and PRM rewards, the authors conduct large-scale ablation studies under a unified compute budget. Their findings reveal that the optimal allocation is highly sensitive to model size, total budget, reward type, and evaluation objective; notably, larger models incur higher per-inference costs, yielding fewer updates or rollouts within the same FLOP budget, thereby uncovering nonlinear coupling in compute allocation.

compute allocationFLOP budgetfoundation models

This study addresses the deployment costs and performance misreporting associated with System-1 decision models in LLM agents. To this end, it constructs a paired evaluation framework encompassing eleven decision points and 7,283 cases, introducing a self-auditing mechanism to rectify pipeline errors. By integrating byte-level consistency verification, zero-shot routing evaluation, and RAG gating techniques, the work achieves full-process transparency. The findings reveal that although the Jev model outperforms Laya on nine of eleven tasks, neither surpasses a random baseline. Following error correction, the actual savings amount to only 4.3%, exposing severe design confounding and data leakage risks. Ultimately, this research establishes a rigorous benchmark for assessing the genuine efficacy of System-1 models.

LLM agent harnessesmodel evaluationpipeline auditing

This study addresses the lack of systematic performance comparisons among System One models, supervised classifiers, and large language models (LLMs) in automated decision-making. We construct a unified evaluation framework to assess six categories of System One models, classifiers, and generative LLMs under matched conditions, incorporating both zero-shot reasoning and fine-tuned checkpoints. Furthermore, we propose a likelihood-based comparative method conditioned on option keys. Our analysis reveals the condition-dependent nature of model rankings and establishes corresponding design principles. Notably, we find that decision models outperform zero-shot classifiers in unlabeled scenarios, and that a two-stage strategy can match the accuracy of Jev at merely 43% of the computational cost. These findings provide critical benchmarks and actionable design paradigms for advancing automated decision-making systems.

automated decision gatesbenchmarkingclassification

This study addresses the stagnation of enterprise AI initiatives in regulated financial institutions due to the absence of quantifiable evaluation criteria. Focusing on six document-intensive workflows, it systematically compares AI system performance across four model families and three tool configurations, distinguishing between demonstration and production environments. For the first time, it links deployment feasibility with human review rates. The authors propose a production-grade evaluation framework encompassing accuracy, reproducibility, traceability, and informative confidence, integrating multi-model comparison, confidence signals, source citation, and self-verification mechanisms. Experiments reveal that 56.1% of the 72 evaluated configurations meet production readiness thresholds. Incorporating source citation and confidence estimation reduces human review requirements to 49%, and adding self-verification further lowers this to 44%, albeit at the cost of reduced error tolerance.

AI deploymentconfidence calibrationproduction readiness

Hot Scholars

YZ

Yanyong Zhang

University of Science and Technology of China ; Rutgers University (Adjunct Visiting Professor)
SensingCyber-Physical SystemsMulti-Modal PerceptionEfficient AI Systems
BY

Bing Yin

Amazon.com
NLPInformation RetrievalDeep LearningKnowledge Graphs
SG

Shivam Garg

Senior Researcher, Microsoft Research
YQ

Yanyun Qu

Xiamen University
Computer Vision
AI

Alexander Ihler

University of California, Irvine
Artificial IntelligenceMachine LearningApplied Statistics