TokenCast: Forecasting Token Consumption During LLM Agent Execution

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unpredictable token consumption caused by context inflation during the execution of large language model (LLM) agents. To tackle this issue, we propose a segmented and composable cost modeling approach that constructs composable cost representations by integrating context growth records with a dynamic evidence fusion algorithm. Through cumulative estimation and dynamic updating mechanisms, the method achieves precise real-time prediction without requiring additional LLM invocations. Experimental results demonstrate that our approach reduces the average prediction error by 14.5% and saves 21.3% in token consumption under budget constraints, while maintaining a single-prediction latency of only 32.8 milliseconds. These findings indicate that the proposed method effectively balances computational efficiency with predictive accuracy for resource-aware agent deployment.
📝 Abstract
When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The agent chooses its next steps based on tool feedback and intermediate results, while the growing context steadily inflates the input size of every subsequent call. The total consumption of a task is therefore hard to predict before execution and the prediction must be revised as the run unfolds. In this paper, we propose TokenCast, which learns a composable cost representation for each execution segment, recording its own consumption and the context growth it introduces. Composing adjacent segments yields a cumulative estimate that captures the extra input cost incurred when context from earlier segments is re-read by every later call. As execution unfolds, newly observed evidence refreshes the forecast, requiring no additional LLM calls and incurring a mean cumulative prediction time of 32.8 ms per run on SWE-bench Verified. Across 4 task suites and 6 agent models, TokenCast's mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations. In offline budget-control replay, TokenCast uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion. The code is available at https://github.com/DEFENSE-SEU/TokenCast.
Problem

Research questions and friction points this paper is trying to address.

LLM agent
token consumption forecasting
context growth
dynamic prediction
budget control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Token Forecasting
LLM Agent
Composable Cost Representation
Dynamic Prediction
Budget Control
🔎 Similar Papers
No similar papers found.