Nous: Learning and Certifying Memory Decisions Before Source Calibration

📅 2026-09-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses how agents can efficiently learn decisions and certify performance improvements when source reliability is unknown. Based on a Hidden Markov Model formulation, it reveals that the sample complexity of decision learning is substantially lower than that of source estimation, establishing a quadratic sample advantage. Furthermore, this work proposes a finite-sample certification mechanism integrating witness regions with conditional replication to construct an auditable policy revision framework. Evaluations on the MultiWOZ dataset demonstrate that a balanced batching strategy significantly improves positive certification rates under weak auditing while effectively preserving policy gains. Collectively, these contributions provide both theoretical foundations and practical solutions for efficient, auditable agent decision-making in uncertain environments.
📝 Abstract
Agent memory systems update state decisions from reports whose reliability may be unknown. Existing analyses of source estimation do not determine when a policy can be learned or its improvement certified without identifying the reporting channel. We study these three tasks using the same observed records. For a specified hidden Markov family with continuous source uncertainty, decision learning and powered certification have quadratic sample complexity, whereas fixed-precision source estimation has quartic complexity. We characterize a sharp identified interval for policy gain under an unknown shared-background channel and derive finite-sample certificates under bounded history dependence and conditional copying. Independently trained witness regions support general history spaces, and disagreement-conditioned auditing improves power for sparse revisions. For dependent histories, prediction-count-preserving batches cancel the unknown reporting background and admit conditional certificates. A MultiWOZ 2.4 evaluation uses text-processing policies on 1,000 human-written test dialogues with simulated audits. Balanced batches retain 2.26 percentage points of the full candidate's 7.34 percentage-point mean gain and obtain more positive certificates under weak audits. These results establish task-specific information requirements and provide an auditable policy-revision framework for Nous.
Problem

Research questions and friction points this paper is trying to address.

agent memory systems
source calibration
decision learning
policy certification
hidden Markov model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agent memory systems
Source calibration
Finite-sample certificates
History dependence
Auditable policy revision