query difficulty estimation

Designs, builds, and evaluates models and algorithms that estimate the difficulty of individual queries (in real time or offline), stratify queries into difficulty tiers, and model difficulty-adjustment dynamics to inform routing, adaptive processing, or workload-allocation decisions.

querydifficultyestimation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenges microservices face in dynamic environments—such as load fluctuations, network variations, and failures—which hinder the coordination of scaling, routing, and repair strategies. The work presents the first taxonomy for adaptive microservice management tailored to dynamic settings, systematically reviewing 84 systems and 13 evaluation artifacts across four dimensions: control placement, dynamic modeling, adaptation strategies, and evaluation evidence. It identifies critical limitations in existing approaches, particularly incomplete modeling of dynamics and insufficient evaluation fidelity, underscoring the importance of high-fidelity evaluation for realizing performance gains. The paper further outlines promising future directions, including cross-layer coordination, telemetry-driven control abstractions, and safe learning-based control, offering a structured roadmap for subsequent research.

adaptive managementdynamic computing environmentsmicroservices

Hardness, Structural Knowledge, and Opportunity: An Analytical Framework for Modular Performance Modeling

Sep 13, 2025
OG
Omid Gheibi
🏛️ Carnegie Mellon University | University of South Carolina

Configuration space explosion complicates performance impact modeling, while gray-box approaches rely on structural knowledge (e.g., module execution graphs) to improve model accuracy—yet the mechanisms by which structural features (e.g., number of modules or configuration options) and structural knowledge influence modeling difficulty and optimization potential remain unclear. Method: We formally define “modeling hardness” and “improvement opportunity,” establishing an analytical framework and matrix to quantify the interplay among system structural complexity, structural knowledge level, and modeling benefit. Controlled experiments on synthetic systems integrate module execution graph analysis with gray-box modeling. Contribution/Results: We identify module count and configuration option count as dominant determinants of modeling hardness. Under high hardness, strong structural knowledge significantly increases improvement opportunity. Structural knowledge primarily enhances ranking accuracy, whereas hardness predominantly degrades prediction accuracy. Our findings provide theoretical foundations and strategic guidance for allocating structural knowledge investment according to specific modeling objectives.

Analyzing the varying impact of knowledge and hardness on different performance metricsInvestigating how structural aspects and knowledge affect modular performance modeling opportunitiesQuantifying modeling hardness driven by module and option counts

This work addresses a critical limitation in traditional recommendation systems that route queries between lightweight heuristics and large language models solely based on task difficulty, ignoring disparities in error cost and business value. To overcome this, the authors propose a value-weighted routing mechanism that makes unsupervised routing decisions by jointly estimating task difficulty and item-level business value within a fully synthetic retail assortment simulation environment. The framework incorporates decision logging and monitoring modules to uncover category-level biases obscured by aggregate metrics. Through slow-path budget control and seasonal parameter tuning, the system achieves a 60% recall rate on high-value items while improving overall accuracy from 94.3% to 98.3%, demonstrating enhanced robustness under simulated Black Friday traffic surges.

cost-aware routinglarge language modelsLLM routing

Latest Papers

What's happening recently
View more

This work addresses a critical limitation in existing large language model (LLM) routing strategies, which rely solely on model-level labels to determine substitutions while neglecting the influence of a model’s role within multi-call workflows and its deployment context on actual performance. To remedy this, the authors propose a role-conditioned substitution principle that decouples replacement decisions into “whether to replace” and “effect evaluation” via predicate-action decomposition. Through a controlled solve-merge-validate pipeline, they systematically investigate the conditional dependencies governing effective substitutions. Empirical results from multi-call LLM workflow experiments—including input-matching interventions, allocation ablations, and cross-model comparisons across Qwen and GPT families—demonstrate that substitution efficacy is jointly determined by the model’s procedural role and deployment context. Notably, in mixed Qwen/GPT configurations, sparse, role-aware replacements reduce RMSE from 4.818 to 1.538, significantly outperforming indiscriminate full-model upgrades.

deployment contextmodel substitutionmulti-call LLM workflows

Database management system (DBMS) configuration tuning is notoriously expensive due to the need to execute full workloads. This work proposes an efficient tuning framework that partitions the tuning process into time slices, adaptively sampling a representative subset of queries in each slice and dynamically refining subsequent evaluation strategies based on runtime profiling. The optimal configuration identified through this compressed workload is ultimately validated on the original workload. This approach is the first to systematically leverage workload compression to enhance tuning efficiency, substantially reducing overhead. Experimental results demonstrate that, compared to the state-of-the-art method, the proposed framework reduces tuning time by up to 73.5% while achieving performance improvements of up to 16.2%.

Configuration OptimizationDatabase TuningPerformance Tuning

This work addresses the high computational cost of large language models or reward functions in modern AI workflows by proposing a novel optimization framework that reduces expensive invocations without altering the underlying models or pipelines. For the first time, it systematically integrates Approximate Query Processing (AQP) from database research and Proxy Models (PMs) into AI workflow optimization. The approach employs declarative modeling combined with online aggregation and adaptive termination strategies from AQP, alongside a lightweight decision tree–based PM pre-filtering mechanism. This dual strategy significantly cuts down costly calls while preserving result quality. Experiments on TPC-DS and LLM post-training tasks demonstrate that AQP reduces invocation counts by 85–90% with estimation errors under 10%, while PM achieves up to 19× speedup and 60–70% fewer calls, with accuracy degradation limited to within 10%.

AI workflowsapproximate query processingexpensive model invocations

This work addresses the challenge of efficiently constructing comprehensive yet concise regression test suites under limited query budgets by proposing a dynamic evaluation set construction method grounded in a capability taxonomy. The approach introduces the notion of “capability signatures” and maps customer queries to the platform’s capability space through a hybrid classifier—integrating deterministic parsing with large language model–based semantic reasoning—augmented by an Intelligence Quotient (IQ) scorer for capability invocation quality, a rule-based cascade integrator, and a conservative replacement mechanism. This framework dynamically determines whether to retain, replace, or escalate queries for human review, while supporting an evolvable classification schema that maintains broad capability coverage and significantly reduces query redundancy. Empirical evaluation on the Microsoft 365 Copilot platform demonstrates its effectiveness in optimizing declarative agent evaluation sets.

agent extensibilitybenchmark compressioncapability taxonomy

Hot Scholars

WZ

Wentao Zhang

Institute of Physics, Chinese Academy of Sciences
photoemissionsuperconductivitycupratehtsc
HM

Haitao Mi

Principal Researcher, Tencent US
Large Language Models
LZ

Linfeng Zhang

DP Technology; AI for Science Institute
AI for Sciencemulti-scale modelingmolecular simulationdrug/materials design
FW

Furu Wei

Distinguished Scientist, Microsoft Research
Natural Language ProcessingArtificial IntelligenceGeneral AIGenerative AI
XL

Xuebo Liu

Associate Professor of Computer Science, Harbin Institute of Technology, Shenzhen
Large Language ModelsNatural Language ProcessingMachine Translation