🤖 AI Summary
This study addresses the challenges of error accumulation and safe self-improvement faced by large language models in long-horizon physical control tasks. To overcome these limitations, this work proposes a dual-timescale agent framework in which a fast timescale verifies and corrects actions through deterministic physics simulations, while a slow timescale consolidates failure experiences into persistent contextual principles. By decoupling reasoning from execution mechanisms, the framework enables constrained self-improvement without violating physical constraints. The proposed approach integrates LLM reasoning, structured interface verification, and long-term memory consolidation. When applied to irrigation scheduling, the system achieves a 51% water reduction compared to historical baselines. Furthermore, ablation studies validate the effectiveness of each individual module within the architecture.
📝 Abstract
Large language model (LLM) agents increasingly combine reasoning, tool use, and action, but most evidence comes from episodic tasks with relatively immediate feedback and reset failures. Long-running physical control operates in a different regime: actions alter future states, errors compound across decisions, and an agent must improve from experience without being allowed to rewrite the physical rules that make execution safe. We study this regime through irrigation, where daily decisions interact with soil-water dynamics over entire growing seasons. We present Mimir, a physics-grounded LLM agent organized around two repair timescales. At the fast timescale, a structured physical interface and deterministic simulator turn an LLM output into a proposal that we numerically check, revise, and subject to bounded deterministic action selection before execution. At the slow timescale, recurrent failure patterns are consolidated into persistent contextual principles that condition future proposals, while the physical model, evaluator, and execution constraints remain immutable. Under a common retrospective evaluator across multiple sites, crops, and years, Mimir attains the lowest reported aggregate control cost among the evaluated references and uses about 51% less irrigation than the historical schedule replay. The ablation study show higher control cost when forward simulation, verified revision, or persistent context is removed; model-scale and model-family studies show no monotonic gain from increasing LLM size. The resulting lesson show that persistent physical agents can combine semantic reasoning with bounded, evidence-driven self-improvement while reserving physical truth and actuator authority for explicit numerical mechanisms.