Weight Oracles: Reading Neural Network Weights with Language Models

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing neural interpretability methods that rely on specific inputs, making it difficult to directly read model weights for detecting hidden capabilities such as backdoors. This work proposes the Weight Oracles paradigm, which fine-tunes language models to directly read raw weights of target networks for security auditing, enabling anomaly diagnosis without behavioral testing. By pioneering a direct weight-reading auditing mechanism integrated with staged curriculum learning and an external deterministic Chain-of-Computation, this approach bridges simulated forward propagation and zero-shot backdoor detection. Experimental results demonstrate that the model achieves 99% accuracy in simulating forward propagation on unseen targets and attains an AUROC of 0.93 for attention-routing backdoor detection, while maintaining robust performance across diverse threat distributions.
📝 Abstract
Interpretability methods for neural networks are predominantly reactive: they analyse activations produced during specific forward passes, requiring known inputs to find hidden capabilities such as backdoors. We propose Weight Oracles, fine-tuned language models that diagnose properties of a target network by reading its raw weights directly, without behavioural testing. We investigate this paradigm in two phases. Phase I establishes feasibility: through a staged curriculum and an external chain-of-computation that delegates parameter-free operations to deterministic code, an explainer LLM learns to simulate the forward pass of small transformers from their weights, achieving 99% holdout accuracy on unseen targets. Phase II repurposes this infrastructure for safety auditing. We train an oracle on natural language diagnostic questions about weight anomalies using only benign pathologies as training signal, and evaluate it zero-shot on backdoors absent from training. The oracle achieves AUROC 0.93 on attention-routed backdoors and 0.81 across a diversified threat distribution including stealth and adversarially regularized variants. Hand-crafted statistical detectors are sharp on the threat models they implicitly target but collapse on threat-model shift, while the oracle remains uniformly competent across attack types. Scaling to realistic model sizes remains the principal open challenge.
Problem

Research questions and friction points this paper is trying to address.

neural network interpretability
weight reading
backdoor detection
safety auditing
language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Weight Oracles
Mechanistic Interpretability
Chain-of-Computation
Zero-shot Backdoor Detection
Safety Auditing
🔎 Similar Papers
No similar papers found.