Linear Probe Accuracy Scales with Model Size and Benefits from Multi-Layer Ensembling

📅 2026-04-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the fragility of single-layer linear probes in detecting "deliberate" erroneous outputs from language models, where the optimal layer varies across models and tasks. To enhance robustness, the authors propose an ensemble of multi-layer linear probes that leverages the geometric property of deception directions progressively rotating across model layers. Systematic evaluation across models ranging from 0.5B to 176B parameters demonstrates that the multi-layer ensemble improves AUROC by 29% on the Insider Trading task and by 78% on the Harm-Pressure Knowledge task. Furthermore, the study quantifies—for the first time—the scaling behavior of probe performance with model size, revealing that AUROC increases by approximately 5% per tenfold increase in parameter count (R = 0.81).

Technology Category

Natural Language Processing: Safety and RobustnessMachine Learning: Large Multimodal Models (LMMs)Search and Optimization: Learning to Search

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: LLM based quality controls for crowd workGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Linear probes can detect when language models produce outputs they "know" are wrong, a capability relevant to both deception and reward hacking. However, single-layer probes are fragile: the best layer varies across models and tasks, and probes fail entirely on some deception types. We show that combining probes from multiple layers into an ensemble recovers strong performance even where single-layer probes fail, improving AUROC by +29% on Insider Trading and +78% on Harm-Pressure Knowledge. Across 12 models (0.5B--176B parameters), we find probe accuracy improves with scale: ~5% AUROC per 10x parameters (R=0.81). Geometrically, deception directions rotate gradually across layers rather than appearing at one location, explaining both why single-layer probes are brittle and why multi-layer ensembles succeed.
Problem

Research questions and friction points this paper is trying to address.

linear probe
deception detection
model scaling
multi-layer ensembling
reward hacking
Innovation

Methods, ideas, or system contributions that make the work stand out.

linear probing
multi-layer ensembling
deception detection
model scaling
AUROC
🔎 Similar Papers
No similar papers found.
E
Erik Nordby
Georgia Institute of Technology, Atlanta, Georgia, USA
T
Tasha Pais
Independent Researcher
A
Aviel Parrack
Stanford University, Stanford, California, USA