Auto: The AGI Compiler

📅 2026-07-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high computational cost, slow inference speed, and lack of determinism in large language model (LLM) agents stemming from token-by-token reasoning. The authors propose an AGI compilation paradigm that records and analyzes agent behavior to identify deterministic segments, which are then extracted as verifiable programs or distilled into expert models. These components are compiled into WebAssembly cognitive binaries equipped with capability declarations and performance guarantees, and executed within a sandboxed environment. This approach enables, for the first time, the automatic conversion of agent experiences into permanent skills with near-zero marginal cost and supports quantifiable uncertainty estimation. Experiments on AUTO-BENCH show that 87.1% of behavioral segments exhibit observational determinism; under distribution shift, per-query inference cost drops from 59 to 2 micro-dollars (a 6.4× speedup), achieving 96.9% accuracy with zero errors.
📝 Abstract
Every LLM agent run re-derives its behavior token by token on a frontier model: brilliant, expensive, slow, and unbounded. We present Auto, a compiler that records live agent behavior, measures which parts are secretly deterministic, extracts them into verified programs or distilled specialists, and emits cognition binaries: WebAssembly artifacts whose manifests carry measured guarantees and whose declared capabilities are physically enforced by the sandbox. A tiered runtime executes compiled behavior behind conformally calibrated guards; guard trips deopt to the reference agent, and the captured trace recompiles back down, so nothing is figured out twice. We use "AGI compiler" in one narrow, testable sense: a system that autonomously converts novel experience into permanent, verified, near-free skill while measuring what it does not know. On AUTO-BENCH, a benchmark we introduce and pre-register, 87.1% of 560 recorded frontier-agent spans are witnessed-deterministic (three of the four censused task families measure 100.0%). On a 300-item stream with three scheduled distribution shifts, the closed loop compiles three artifact generations and drives marginal cost from 59 to 2 micro-dollars per item (6.4x end-to-end) at 96.9% parity on witnessed inputs with zero errors. The same stream also quantifies the failure modes: a loose guard silently mislabels 48.9% of compiled answers, and an unfaithful deopt reference causes the verification gate to refuse recompilation. Calibration and reference fidelity, not model capability, decide whether cheap stays correct. Code: https://github.com/RightNow-AI/auto
Problem

Research questions and friction points this paper is trying to address.

AGI compiler
deterministic behavior
cognition binaries
distribution shifts
verification
Innovation

Methods, ideas, or system contributions that make the work stand out.

AGI compiler
deterministic extraction
cognition binaries
deoptimization
verified distillation
J
Jaber Jaber
RightNow AI
O
Osama Jaber
RightNow AI