🤖 AI Summary
This study addresses the unclear organization of agent trajectories and the difficulty of jointly training long-context understanding with atomic capabilities. We propose TrajLong, a framework that compiles trajectories into densely supervised long-context tasks, focusing on three atomic capabilities: evidence grounding, aggregation, and state maintenance. Based on these, we establish a principle for designing intermediate training data driven by shared capability requirements. Conducting intermediate training and supervised fine-tuning (SFT) on Qwen3 models, we validate our approach through controlled ablations and capability-level analysis. Across 18 benchmarks, our method significantly outperforms both raw and masked trajectory baselines. These results reveal the intrinsic connection between long-context processing and agentic capabilities, effectively enhancing agents' comprehensive reasoning performance.
📝 Abstract
LLM agents for coding, search, and workplace tasks increasingly rely on long-context capabilities to effectively aggregate and reason over extended interaction histories. Recent work has incorporated agent trajectories into mid-training stage, drawing on their naturally long and interaction-rich structure. Yet how to organize the information within these trajectories into effective mid-training supervision remains underexplored. In this work, we investigate the relationship between long-context and agent atomic capabilities and introduce TrajLong, a novel framework that compiles trajectories into long-context training tasks with dense supervision, targeting three representative atomic capabilities: evidence grounding, cross-evidence aggregation, and temporal state maintenance. We mid-train Qwen3-14B-Base and Qwen3-30B-A3B-Base with data compiled by TrajLong, followed by supervised fine-tuning. Experiments on 6 long-context and 12 agent benchmarks demonstrate broad performance gains, with controlled ablations showing improvements over raw and masked trajectory baselines. Capability-level analyses further reveal task-dependent associations between long-context and agent atomic capabilities. These findings suggest that the shared capability demands of long-context reasoning and agent execution provide a principled basis for designing mid-training data to develop downstream agent capabilities.