Learning the Loop, Not Just the Page: Execution-Grounded Loop Learning for Web Generation

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing web generation methods that lack intermediate diagnosis and repair, rendering evaluator feedback signals unactionable. We propose WebLoop, a framework that jointly optimizes generation, critique, and revision roles by treating looped learning itself as an optimization objective for the first time. By training an execution-free critic via complementary signals and introducing consequence-aware credit assignment, WebLoop effectively resolves the challenge of unactionable criticism. The method builds upon GRPO with discriminative and utility signals, leveraging Qwen3.5 for multimodal transfer, and demonstrates that performance gains stem from the looping mechanism rather than merely increasing revision iterations. Experiments show that WebLoop achieves 41.5 points on WebRise and 38.9% accuracy on WebGen-Bench, outperforming baselines by 11.3 and 15.4 points respectively, while exhibiting strong scalability.
📝 Abstract
Functional Web generation is increasingly optimized with executable rewards, yet existing methods largely focus on the quality of the final page and leave the process of diagnosing and repairing imperfect implementations underexplored. We identify a central challenge in this setting: the Generator and Refiner produce executable artifacts with direct environment rewards, whereas the intermediate Critic influences downstream behavior without a directly executable outcome. We introduce WebLoop, an execution-grounded framework that jointly learns generation, critique, and refinement within a shared policy. WebLoop trains an execution-free Critic with complementary signals for requirement-level discriminability and downstream helpfulness, first establishing reliable diagnosis and then introducing consequence-aware credit, while all three roles are jointly optimized with group-relative policy learning. With Qwen3.5-9B, WebLoop reaches 41.5 Overall on WebRise and 38.9% accuracy on WebGen-Bench, improving the base model by 11.3 and 15.4 points, respectively. The gains transfer to first-pass generation, persist at 27B scale, and generalize from text-only training to multimodal inputs. Controlled analyses further show that the improvement cannot be explained by an additional refinement pass alone, highlighting the importance of learning the Critic and the loop itself.
Problem

Research questions and friction points this paper is trying to address.

Web generation
Executable rewards
Critic optimization
Iterative refinement
Execution-grounded learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Execution-Grounded Learning
Joint Policy Optimization
Critic Training
Web Generation
Group-Relative Policy Learning