Embedded Assessments for Frontier AI

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
论文提出通过嵌入式评估方法,让独立评估者深入AI公司内部系统和实践,以更全面地评估前沿AI模型的风险。
📝 Abstract
Third-party evaluations for frontier AI have mostly tested models through external interfaces before deployment. But the risks from frontier AI models depend on how their developers use and govern them internally. Recently, CEOs of frontier AI companies have committed to hosting embedded assessments. These assessments would give independent evaluators employee-like access to a developer's internal systems, staff, and documentation. First, we argue that this can enable deeper and more flexible assessments of risks that depend on internal systems and practices, while providing access under stronger security controls. Then, we examine seven design questions about scope, information gathering, duration, timing, terms of engagement, disclosure, and escalation. We recommend that frontier AI developers begin hosting embedded assessments now, covering at least three areas central to managing risks from internal AI use: internal agent monitoring, internal agent security controls and permissions, and model alignment. To enable meaningful third-party scrutiny, assessments should be continuous, evaluators should publish detailed reports at least quarterly, and clear escalation mechanisms should be established. These recommendations are intended as a starting point, with further steps needed to realize the full potential of embedded assessments.
Problem

Research questions and friction points this paper is trying to address.

embedded assessments
frontier AI
internal systems
risk management
Innovation

Methods, ideas, or system contributions that make the work stand out.

embedded assessments
internal systems
security controls
🔎 Similar Papers