Humanize: Judgement Engineering for Agentic Coding

πŸ“… 2026-10-06
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the β€œeasy to generate, hard to complete” dilemma and defect escape caused by weak judges in agent-based code generation. We propose a judgment-engineering-based multi-agent orchestration framework that ensures reliability through 72 mechanical gates governing boundary decisions across planning, implementation, review, and learning phases. By introducing cross-vendor independent review, deterministic hook routing, and modeling the workflow as a Markov chain for joint sampling, the approach significantly reduces defect survival rates while integrating a kernel knowledge base with profiling-driven feedback loops. The resulting project has garnered 1,468 GitHub stars, achieved top-three placements in MLSys competitions, earned perfect scores on IOI, IMO, and IPhO benchmarks, attained full accuracy on PutnamBench, and ranked first on Lean-Eval.
πŸ“ Abstract
Agentic coding makes code generation cheap, but reliable completion remains difficult: the agent that writes the code is a weak judge of whether it is done. We present Humanize, a multi-agent orchestration workflow for agentic coding built around judgement engineering: explicit, mechanically enforced decisions at the boundaries between planning, implementation, review, and learning. A human approves a plan contract, a builder agent implements it in rounds, and a reviewer agent from another vendor decides completion; deterministic hooks, not a model, route work between these roles and enforce 72 mechanical gates. Viewed as a Markov chain over repository states, alternating builder and reviewer samples jointly from two models, so a defect survives only if both miss it. We study Humanize through its deployment, 118 public postmortems of real loops, and its applications. Over 68 versions in 108 days, it gathered 1,468 GitHub stars. Applications include a 567-file gem5 build-system migration under upstream review; Kernel Design Agents, which extend the loop with a kernel knowledge base and profiling feedback and placed in the top three of all three Full-Agent tracks of the MLSys 2026 FlashInfer contest; and, through Humanize Olympiad Agents (HOA), full scores in IOI 2026, IMO 2026, IPhO 2026, and IBO 2024, 418.5/437 in IChO 2026 (gold-medal). Humanize also achieves 672/672 on PutnamBench and ranks first (251/303) on Lean-Eval's leaderboard even competiting with professional mathematicians. The postmortems show that independent review catches unsupported builder claims, but stopping remains a key weakness. In reports that separate rounds by phase, two thirds of rounds occurred after implementation was accepted. This evidence is observational, not a controlled comparison of workflows.
Problem

Research questions and friction points this paper is trying to address.

Agentic Coding
Judgement Engineering
Code Completion
Task Termination
Innovation

Methods, ideas, or system contributions that make the work stand out.

Judgement Engineering
Multi-agent Orchestration
Agentic Coding
Markov Chain
Mechanical Gates
Sihao Liu
Sihao Liu
UCLA
Computer ArchitectureVLSICPUFPGA
Ligeng Zhu
Ligeng Zhu
Nvidia
Machine LearningEfficient Deep Learning
Z
Zijian Zhang
NVIDIA
D
Dongyun Zou
MIT
Z
Zhengyang Zhang
Tsinghua University
C
Changye Li
Tsinghua University
S
Song Bian
NVIDIA
S
Song Han
NVIDIA
Tony Nowatzki
Tony Nowatzki
University of California, Los Angeles
Computer Architecture