Radical AI Interpretability

📅 2026-06-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the critical challenge of reliably inferring an AI system’s beliefs, desires, and semantic content from its computational architecture—a key requirement for safe alignment and deception detection. Drawing on radical interpretation in philosophy and mechanistic interpretability in machine learning, the paper proposes the “holographic attribution” principle: beliefs, desires, and propositional structures must be jointly constrained to ensure coherent and reliable attitude ascriptions. Building on this principle, the authors develop an integrative framework that yields actionable testing methodologies and establishes evaluative criteria bridging representationalist and interpretivist approaches. This framework provides existing interpretability techniques with a theoretically grounded and empirically verifiable foundation, uniting philosophical rigor with practical applicability in AI safety research.
📝 Abstract
We develop a framework for interpreting AI systems as agents, drawing on the philosophical tradition of radical interpretation and the tools of mechanistic interpretability. The core question is: given the computational facts about a system, how do we solve for its beliefs, desires, and meanings? This matters increasingly for safety. We want to be able to trust the systems we deploy, whether by understanding their goals or, more modestly, by reliably detecting deception. Interpretability researchers are building tools to read beliefs and desires off a model's internals, but there is no settled account of when such a tool has succeeded. This book supplies one. We propose criteria on both representationalist and interpretationist approaches, and tie each to tests current interpretability methods can carry out. A central lesson is that these attributions cannot be made piecemeal. Beliefs, desires, and the propositional structure they presuppose are jointly constrained, and a method that fixes one while measuring the others inherits whatever distortions that introduces. This holism becomes pressing for AI systems, which may not share the interpreter's concepts. However, it also provides leverage: a system's attitudes constrain its propositional structure, that structure constrains which attitudes can be attributed, and mechanistic interpretability can help us measure both.
Problem

Research questions and friction points this paper is trying to address.

AI interpretability
radical interpretation
beliefs and desires
mechanistic interpretability
AI safety
Innovation

Methods, ideas, or system contributions that make the work stand out.

radical interpretation
mechanistic interpretability
belief-desire attribution
semantic holism
AI safety