🤖 AI Summary
This study addresses the challenge of simultaneously committing to an upper risk bound and a development deadline when the safe production rate is unknown. To this end, it proposes a "learn-then-replay" strategy that allocates residual compute to safety alignment and implements state-wise risk calibration. Theoretically, we demonstrate that eliminating inefficient research suffices to achieve both objectives, whereas fixed compute allocation can only guarantee deadlines without controlling risk. Notably, the proposed approach approximates the theoretical optimum by requiring merely 3.3% productivity beyond the necessary boundary, with the additional safeguarding costs borne by the client. Ultimately, this framework provides a verifiable scheduling paradigm for navigating the safety-progress trade-off in AI development.
📝 Abstract
Can a regulator promise both a risk ceiling and a development deadline when safety productivity is unknown? A ceiling below the final model's unprotected hazard requires a minimum stock of safety knowledge, so both promises hold only if weak research can be ruled out. Learning first and then replaying development, with spare compute in safety, comes close to that minimum. In a calibration with only state risk, constant safety yield, and full knowledge transfer, it needs only 3.3 percent more productivity than the necessary bound. Customer services pay for the guarantee. Rules that fix compute allocation fix dates, not risk.