Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of long-horizon agents maintaining task consistency amid context extension, varying task complexity, and new data integration. To this end, it proposes the Long-Transduction diagnostic framework and evaluation protocol, which pioneers a controlled-variable methodology to independently quantify the isolated effects of local complexity, input formatting, and context length on models’ read-and-modify capabilities, benchmarked across seven open-source models. The findings reveal critical failure modes: extending context to 128K tokens, altering input formats, and increasing task complexity degrade performance by 62.8%, 36.5%, and 39.9%, respectively. These results provide an empirical foundation for optimizing the reliability of long-horizon agents.
πŸ“ Abstract
Long-horizon agentic workflows require models to sustain repeated state-dependent actions all while the context grows, sub-task complexity changes, and new data arrives. Each situation represents an independent axis along which an agent may fail. An agent reconciling a long ledger, for example, must repeatedly read its state, update the correct record, and preserve alignment across thousands of outputs. A model may accept the entire ledger yet lose its place or stop applying the operation consistently as generation proceeds. We introduce Long-Transduction, a controlled diagnostic that tests a model's ability to stay on task during long generation while continuously reading, mutating, and outputting input-context dependent operations such as arithmetic, sorting, variable lookups, and table transformations. Long-Transduction evaluation independently varies local task complexity, input data formatting, and context length isolate failures along each axis. We evaluate seven open-weight models, finding a 62.8\% decrease when scaling context length from 4-128K, a 36.5\% decrease when varying input format, and a 39.9\% decrease by increasing local task complexity. Together, these failures represent critical liabilities in long-horizon agentic workflows.
Problem

Research questions and friction points this paper is trying to address.

Long-horizon agent
Context length
Task reliability
Agentic workflows
State-dependent actions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Long-Transduction
Long-Horizon Agent
Context Length
Diagnostic Evaluation
Agentic Reliability
πŸ”Ž Similar Papers
No similar papers found.