π€ AI Summary
This work addresses a critical limitation in mainstream AI research, which prioritizes average-case performance yet struggles to mitigate catastrophic failures arising from sparse evidence, adversarial dynamics, and heavy-tailed risks in alignment and safety contexts. To bridge this gap, the paper introduces ECAISAβthe first cognitive normative framework specifically designed for AI safety and alignment research. ECAISA systematically identifies five core cognitive gaps and establishes a multi-layered governance architecture comprising eight principles, a three-tier scoring scale, a four-level disclosure ladder, and seven anti-gaming mechanisms. By integrating preregistered bibliometric baselines with structured synthesis methodologies, the framework substantially enhances auditability, independent verifiability, and control over information hazards, thereby rectifying key deficiencies of current scientific paradigms in safety-critical scenarios.
π Abstract
Mainstream AI research emphasises capability growth and tolerates low failure rates when average-case performance is high. AI safety and alignment research has a different mission: to ensure that catastrophic failures never occur, under sparse evidence, adversarial dynamics, and fat-tailed risk. We argue that the two domains differ along two analytically independent axes---{\it capability profile}, demonstrating the absence of hazardous behaviours rather than the presence of positive capabilities, and {\it risk profile}, bounding worst-case outcomes under fat-tailed uncertainty rather than optimising average-case performance---and that mainstream epistemic practices are inadequate on both. Building on a structured synthesis grounded in a preregistered bibliometric baseline, we identify five cross-cutting gap dimensions in current alignment research, including the near-absence of institutionalised independent verification. To address these gaps, we propose {\sc ECAISA}, an Epistemic Code for AI Safety and Alignment comprising eight principles, a three-level scoring rubric, a four-level disclosure ladder that reconciles transparency with information-hazard and commercial-confidentiality constraints, a tiered applicability scheme, an information-hazard adjudication procedure, and seven anti-gaming mechanisms. {\sc ECAISA} does not certify that any AI system is safe; it constrains how safety-relevant research claims are documented, checked, and relied upon, with auditability rather than certification as its governance target.