🤖 AI Summary
This study addresses the limited adaptive capacity of large language model (LLM) agents when environmental assumptions become invalid, highlighting a significant disparity between task competence and environmental adaptability. To investigate this, we propose AGNI, a framework that defines an environmental novelty benchmark through automated trajectory analysis and assumption extraction, while systematically evaluating agent adaptability by injecting environmental perturbations. Our findings reveal that LLM agents struggle to autonomously diagnose the underlying causes of environmental shifts. Furthermore, we demonstrate that targeted post-training can substantially enhance their overall performance across both novel and familiar tasks. This work underscores the critical need for improved diagnostic and adaptive mechanisms in autonomous agents operating within dynamic environments.
📝 Abstract
LLM agents increasingly solve long-horizon tasks by autonomously interacting with their environment. In doing so, their strategies rely on assumptions about that environment: which resources and tools exist, where they are located, and how they behave. When these assumptions no longer hold, reliable agents must detect the change and adapt while pursuing the same goal. We study this adaptation capability through environmental novelty: a change that keeps the task objective fixed while invalidating an assumption underlying an otherwise successful trajectory. We introduce AGNI, an automated pipeline that extracts trajectory-relevant assumptions, injects targeted environmental changes, and validates that the resulting novel tasks remain solvable. Across three terminal benchmarks, AGNI produces diverse novelties spanning resources, interfaces, constraints, and execution semantics. Evaluating multiple LLM agents reveals a substantial adaptation gap between base and novel tasks. Trajectory analysis suggests that agents often encounter evidence of the change but fail to diagnose its cause and revise their strategy. Finally, post-training for environmental novelty improves adaptation to held-out novel tasks while also improving performance on base tasks. Our results highlight a gap between task competence and adaptive capability and motivate environmental variation as a core dimension of agent training and evaluation.