Correctness-Optimized Residual Activation Lens (CORAL): Transferrable and Calibration-Aware Inference-Time Steering

📅 2026-02-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the prevalent issue of miscalibration in large language models following instruction tuning and preference alignment, which undermines their reasoning reliability. The authors propose a lightweight, inference-time intervention method that directly optimizes output correctness without requiring retraining. Specifically, they introduce a weight-decay-regularized MLP probe to extract distributed correctness signals from residual activations and apply a regularization mechanism to modulate model behavior. This approach enables effective cross-task transferability. Experimental results demonstrate that the method improves average accuracy by 10% and reduces Expected Calibration Error (ECE) by 50% across multiple 7B-scale models. Furthermore, on four unseen benchmark datasets, it achieves an average accuracy gain of 14% and a 49% reduction in ECE, highlighting its robustness and generalization capability.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationNatural Language Processing: Safety and RobustnessComputer Vision: Large Vision Models

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Large language models (LLMs) exhibit persistent miscalibration, especially after instruction tuning and preference alignment. Modified training objectives can improve calibration, but retraining is expensive. Inference-time steering offers a lightweight alternative, yet most existing methods optimize proxies for correctness rather than correctness itself. We introduce CORAL (Correctness-Optimized Residual Activation Lens), a regularized inference-time steering method that captures distributed correctness signals from model internal activations using weight-decay MLP probes. We evaluate CORAL across three 7B-parameter models and find that it consistently improves accuracy by 10\% and expected calibration error (ECE) by 50\% on average. We additionally demonstrate that these gains transfer without retraining to the complete published test sets of four held-out benchmarks (ARC-Challenge, HellaSwag, Math-MC, OpenBookQA), averaging 14\% accuracy improvements and 49\% ECE improvements. Our results support the hypothesis that distributed information in model internals can be extracted using regularized probes when individual neurons are insufficient. CORAL thus provides a compute-efficient, transferable, and calibration-aware approach to improve MCQA performance during inference.
Problem

Research questions and friction points this paper is trying to address.

miscalibration
large language models
inference-time steering
multiple-choice question answering
expected calibration error
Innovation

Methods, ideas, or system contributions that make the work stand out.

inference-time steering
model calibration
correctness optimization
transferable probing
residual activation lens
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
M
Miranda Muqing Miao
Department of Computer and Information Science, University of Pennsylvania, Philadelphia, USA
Young-Min Cho
Young-Min Cho
University of Pennsylvania
Artificial IntelligenceNatural Language ProcessingAI Agents
L
L. Ungar
Department of Computer and Information Science, University of Pennsylvania, Philadelphia, USA