AI Harness: Certification under Proposal-Conditioned Information for Foundation-Model Agents

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the failure of safety certification during runtime interventions in foundation model agents, caused by neglecting the correlation between proposals and historical context. It is the first to reveal that this correlation constitutes a necessary condition for certification, analyzing how state folding compromises certifiability and establishing information preservation boundaries. By integrating bounded robust interfaces, viability kernel theory, and standard safety reachability fixed-point algorithms, this work demonstrates that observing current proposals restores robust viability and decouples incompatible intervention modes. Through controlled experiments, it reproduces the predictive barriers arising from telemetry removal or privilege restriction, validating the correctness of the theoretically defined boundaries and achieving a precise characterization of robustly safe interventions.
📝 Abstract
Foundation-model agents are often modeled as policies over an observed state. In deployed systems, however, a runtime may intervene only after the model has emitted a semantic proposal, making the proposal both an action candidate and a decision-time observation generated by a history-conditioned process. We show that collapsing this structure into a state-only proposal envelope can preserve proposal coverage while destroying certifiability. In a finite robust interface, the viability kernel of the collapsed model is contained in the physical projection of the history-augmented kernel, and the collapse is lossless exactly when every proposal-conditioned collapsed fiber retains a common robust-safe intervention. This gap can be maximal even with constant-size proposal and history alphabets. The same common-action condition yields a dual result: observing the current proposal can restore robust feasibility when it separates latent modes requiring incompatible interventions. We extend these one-step results over time using exact finite beliefs and standard safety and reachability fixed points, separating indefinite operational viability from finite worst-case verified progress. Controlled model-in-the-loop tests reproduce the predicted obstructions when telemetry or effect verification is removed or intervention authority is restricted. Thus, our contribution is not a new fixed-point calculus, but a characterization of when proposal--history correlation at the model--tool boundary is necessary for certification.
Problem

Research questions and friction points this paper is trying to address.

foundation-model agents
certification
robust safety
proposal-conditioned information
viability kernel
Innovation

Methods, ideas, or system contributions that make the work stand out.

Foundation-model agents
Certifiability
Proposal-history correlation
Viability kernel
Robust intervention
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.