Falsifiable Commitment Planning for Self-Correcting Web Agents

📅 2026-07-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Long-horizon web agents often deviate from user instructions due to state drift, skill obsolescence, or outdated planning assumptions, leading to task failure. This work proposes FCPAgent, a framework that models plan steps as Falsifiable Commitment Units (FCUs), explicitly specifying subgoals, confirming and falsifying evidence, and associated confidence levels. Robust execution is achieved through a “plan–test–repair” loop, which incorporates scope-aware repair strategies to precisely localize and correct errors in execution, skills, or planning. Additionally, FCPAgent introduces a hybrid commitment testing module that combines lightweight evidence matching with large language model–based diagnostics, enabling dual verification before and after action execution. Evaluated on the WebArena benchmark, FCPAgent improves average success rates by 13.8% over the strongest baseline, with particularly pronounced gains in long-horizon tasks.
📝 Abstract
Long-horizon web agents often go off track before final failure: a trajectory can remain locally plausible even after the current state, reused skill, or plan assumption no longer supports the user instruction. Existing agents can plan, reflect, or reuse experience, but their plans rarely specify the evidence under which an active step should still be trusted. We propose FCPAgent, a falsifiable commitment planning framework for robust long-horizon web agents. FCPAgent represents each plan step as a Falsifiable Commitment Unit (FCU): a subgoal grounded in a reusable skill, together with confirming evidence, falsifying evidence, and a confidence score. Execution is organized as a plan-test-repair loop. The hybrid commitment testing module checks candidate actions before they modify the browser and checks observations after execution; for efficiency, it combines lightweight evidence matching with LLM-based diagnostic verification. When evidence falsifies a commitment, scope-aware repair localizes the contradiction to the execution, skill, or planning level and revises the smallest adequate part. On WebArena, FCPAgent achieves a 13.8% relative improvement in average success over the strongest baseline, with especially large gains on long-horizon tasks.
Problem

Research questions and friction points this paper is trying to address.

long-horizon web agents
falsifiable commitment
plan reliability
execution deviation
trustworthy planning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Falsifiable Commitment
Plan-Test-Repair Loop
Scope-Aware Repair
Long-Horizon Web Agents
Hybrid Commitment Testing
🔎 Similar Papers
No similar papers found.