Learn Now, Use Next, Trust Later: Prequential Test-Time Learning for LLM Agents

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of delayed knowledge acquisition and unreliable single-experience reuse in the deployment of LLM-based agents by proposing StepLearn, a novel framework that decouples immediate utilization from persistent trust. Specifically, it introduces a non-parametric progressive testing-based learning mechanism that enables dynamic updating and safe reuse of external knowledge through stepwise hypothesis generation and cross-episode prospective validation, all without modifying model parameters. Experimental evaluations on WebArena-Lite and ALFWorld demonstrate that StepLearn significantly outperforms baseline methods, exhibiting advantages even on the first attempt and achieving a maximum improvement of 12.7 percentage points in success rate.
📝 Abstract
Adapting large language model agents during deployment requires not only retaining past experience, but also turning new observations into timely guidance. Many test-time learning methods, however, acquire knowledge from completed episodes. Feedback from an ongoing interaction may therefore not be distilled into knowledge soon enough to help the next decision. Acquiring knowledge at the granularity of individual transitions could reduce this delay, but raises a separate challenge: a rule that is useful within one episode may not be reliable enough to guide future episodes. Waiting for validation can forfeit immediate benefits, whereas unrestricted reuse can propagate accidental or misattributed guidance. We introduce StepLearn, a nonparametric framework that separates immediate use from persistent trust. It turns informative transitions into hypotheses that can guide the next step, while requiring prospective validation before reuse across episodes. Their predicted effects are checked against subsequent observations outside the source episodes, and only sufficiently supported hypotheses become available for persistent guidance. This process updates external knowledge while keeping all model parameters fixed. Over five rounds on WebArena-Lite and ALFWorld, StepLearn achieves average success rates of 59.9% and 84.0% with GPT-5-mini, and 57.8% and 88.1% with Qwen3.5-35B-A3B, respectively. It outperforms EvoTest, the strongest baseline, by 2.2-12.7 percentage points across the four settings. Learning dynamics further shows that these gains are not restricted to the final repetition, with advantages already present on first task attempts in most settings.
Problem

Research questions and friction points this paper is trying to address.

test-time learning
LLM agents
knowledge acquisition
transition-level learning
experience reuse
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-Time Learning
LLM Agents
Nonparametric Framework
Prequential Validation
External Knowledge Update
🔎 Similar Papers
No similar papers found.