🤖 AI Summary
This study addresses the tendency of AI research agents to misjudge fragile candidates and conflate benchmark improvements with theorem-level claims due to insufficient evidence governance. To mitigate this, we propose EPOCH, an architecture that establishes an evidence-governed discovery loop through explicit task contracts, typed memory, active falsification, and independent replay mechanisms, thereby rigorously binding the search process to claim strength and effectively distinguishing transient optimizations from verifiable scientific breakthroughs. Experimental results demonstrate that EPOCH achieves state-of-the-art performance on the AlgoTune and Math14 benchmarks while yielding substantive advances in algorithms, constructions, and proofs across ten discovery tasks, including AgentHPO. These findings indicate that the proposed framework significantly enhances the credibility and reliability of AI-driven scientific research.
📝 Abstract
AI research agents are increasingly used to search over programs, mathematical constructions, and proofs. However, existing systems typically optimize evaluator feedback without adequately governing how that feedback is interpreted, challenged, and reused. As a result, promising but fragile candidates can be promoted as discoveries, while benchmark improvements, finite certificates, and theorem-level claims are too easily conflated. We introduce EPOCH, an evidence-governed architecture designed to close this gap. EPOCH implements an evidence-governed discovery loop by combining explicit task contracts, typed memory, active falsification, admission checks, and independent replay, so that each candidate is evaluated against the strength and scope of the claim it supports. EPOCH achieves state-of-the-art aggregate performance on AlgoTune, substantially exceeding the strongest baseline in mean normalized score (0.65 vs. 0.53), and attains the highest mean score on the internal Math14 suite (0.57). It further shows favorable held-out behavior under official-test replay and leads the descriptive aggregate on AgentHPO. Across ten discovery problems, EPOCH delivers substantial task-specific advances, including improved executable constructions, optimized algorithms, counterexamples, and proof-supported results. These advances demonstrate its ability to convert search into concrete progress across mathematical and computational domains. Together, the results suggest that evidence governance is a necessary step toward AI research agents that produce not only stronger solutions, but also more trustworthy scientific discoveries.