JET: Judge-Guided Evolution at Test Time for Agent Programs

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of evaluating program modifications during test-time agent evolution due to the absence of ground-truth rewards. We propose JET, a method that evolves executable judges on source task trajectories and transfers them in a frozen manner to target domains, guiding unsupervised program adaptation without requiring external evaluators or model weight updates. The core innovation lies in introducing a cross-task transfer mechanism for executable judges, integrating program evolution with trajectory-based judgment logic construction. Experiments demonstrate that JET achieves a 13% average reward improvement and a 36% relative increase in exact success rate under WebShop cold-start settings, while also validating the existence of search bottlenecks in the PushT task.
πŸ“ Abstract
An agent's executable program governs how it uses tools, processes observations, and responds to failures. Evolving this program at test time can help adaptation, but deciding which changes to retain is difficult when true rewards are unavailable. Execution traces provide evidence of agent behavior, yet interpreting that evidence requires a judge that remains useful as tasks and candidate programs change. We introduce Judge-Guided Evolution at Test Time (JET), which evolves an executable judge on labeled source trajectories, then freezes and transfers it to guide target-side program evolution. The judge supplies scores and diagnostic feedback without target evaluator access or model-weight updates. On unseen WebShop tasks, JET achieves approximately 13% higher mean reward than fixed-rubric guidance when evolution begins from an unevolved program (cold start) and 4% higher when it begins from one already optimized on source tasks (warm start), with a 36% relative improvement in cold-start exact success. An exact-judge control on PushT, where the judge reconstructs the scoring rule from observations, shows that without judge error, program search becomes the bottleneck. Analyses identify useful reward-prediction logic in the evolved code and show that better final selection alone cannot explain the gains. These results support executable judge transfer for program adaptation under evaluator-preserving task shifts.
Problem

Research questions and friction points this paper is trying to address.

test-time evolution
agent programs
reward signal absence
program adaptation
executable judge
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-time evolution
Executable judge transfer
Agent programs
Reward prediction
Program adaptation