Predicting Task Difficulty Without Rollouts

📅 2026-08-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work proposes a method to accurately predict agent task difficulty directly from task descriptions without requiring time-consuming simulations, thereby supporting benchmark calibration and curriculum design. Addressing the limitations of AUC-based metrics in difficulty assessment, the approach introduces token-level entropy as a predictive signal and incorporates difficulty residual modeling to effectively identify contaminated or infeasible tasks within environments. Through textual entropy analysis, fine-grained task feature extraction, and validation across multiple domains, the method significantly enhances the reliability of difficulty prediction across 17 diverse agent scenarios—spanning coding, mathematical reasoning, and web navigation—and further aids in uncovering flaws in environment design.
📝 Abstract
Task difficulty dictates an agent's likelihood of success, and estimating it without rollouts means forecasting this directly from a task description before executing costly simulations in stateful environments. Reliable estimates would therefore allow environment designers to calibrate evaluation benchmarks and construct progressive training curricula. This becomes increasingly important as agents move into long-horizon domains, where empirical trial-and-error is a severe computational bottleneck. Prior work on early prediction is limited to static tasks or isolated coding environments, often relying on narrow features and inaccurate evaluation metrics. We study \textit{ex ante} difficulty prediction across 17 agentic benchmarks spanning coding, mathematics, machine learning, web navigation, function calling, and other domains. We show that AUC can mask poor difficulty estimates, identify token-level entropy as a useful predictive signal, and show how residuals between expected and observed difficulty can expose hidden environment flaws such as contamination and infeasibility.
Problem

Research questions and friction points this paper is trying to address.

task difficulty prediction
ex ante estimation
stateful environments
agent benchmarks
difficulty calibration
Innovation

Methods, ideas, or system contributions that make the work stand out.

difficulty prediction
ex ante estimation
token-level entropy
agent benchmarks
environment diagnostics
💼 Related Jobs
No related jobs found.
S
Stefan Krsteski
Andromede AI
C
Charlotte Meyer
Andromede AI