π€ AI Summary
This study addresses the challenge of long-horizon agents maintaining task consistency amid context extension, varying task complexity, and new data integration. To this end, it proposes the Long-Transduction diagnostic framework and evaluation protocol, which pioneers a controlled-variable methodology to independently quantify the isolated effects of local complexity, input formatting, and context length on modelsβ read-and-modify capabilities, benchmarked across seven open-source models. The findings reveal critical failure modes: extending context to 128K tokens, altering input formats, and increasing task complexity degrade performance by 62.8%, 36.5%, and 39.9%, respectively. These results provide an empirical foundation for optimizing the reliability of long-horizon agents.
π Abstract
Long-horizon agentic workflows require models to sustain repeated state-dependent actions all while the context grows, sub-task complexity changes, and new data arrives. Each situation represents an independent axis along which an agent may fail. An agent reconciling a long ledger, for example, must repeatedly read its state, update the correct record, and preserve alignment across thousands of outputs. A model may accept the entire ledger yet lose its place or stop applying the operation consistently as generation proceeds. We introduce Long-Transduction, a controlled diagnostic that tests a model's ability to stay on task during long generation while continuously reading, mutating, and outputting input-context dependent operations such as arithmetic, sorting, variable lookups, and table transformations. Long-Transduction evaluation independently varies local task complexity, input data formatting, and context length isolate failures along each axis. We evaluate seven open-weight models, finding a 62.8\% decrease when scaling context length from 4-128K, a 36.5\% decrease when varying input format, and a 39.9\% decrease by increasing local task complexity. Together, these failures represent critical liabilities in long-horizon agentic workflows.