Faithful Activation Verbalization: Reducing Hallucinations in LLM Representation Interpretation

πŸ“… 2026-09-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the issues of incomplete descriptions and hallucinations in activation interpretation for large language models by proposing the AVPO framework. This method introduces an explicit intermediate readout mechanism that transforms hidden representations into natural language through two-stage text reconstruction. It leverages Direct Preference Optimization (DPO) to jointly optimize semantic recoverability and lexical fidelity, while employing a frozen question-answering model for objective evaluation. Experimental results demonstrate that AVPO significantly improves information recovery rates and effectively suppresses detail fabrication, yielding more faithful and reliable interpretations of internal model representations.
πŸ“ Abstract
Activation verbalization methods such as Activation Oracle and Natural Language Autoencoders decode hidden representations of large language models into human-readable natural language. However, existing methods can produce incomplete or hallucinated descriptions, making their activation verbalizations difficult to trust and use reliably in practice. To this end, we introduce AVPO, a two-stage framework that first reconstructs source text from a hidden activation and then evaluates the resulting text with a separate frozen question-answering model, yielding an explicit and inspectable intermediate readout. We further optimize the inverter with direct preference optimization (DPO), using rewards that capture both semantic recoverability and lexical fidelity. Across six text families, AVPO improves gist- and detail-level information recovery over the strongest baseline by up to 17.1 and 9.3 percentage points, respectively. Crucially, the gains arise from preference optimization rather than fine-tuning on selected reconstructions alone, enabling compact cross-model inverters to surpass donor-matched question-conditioned verbalizers while improving both semantic recoverability and lexical fidelity. Moreover, out-of-distribution case study shows that AVPO better recovers high-level semantics while fabricating fewer details.
Problem

Research questions and friction points this paper is trying to address.

activation verbalization
hallucination
large language models
representation interpretation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Activation Verbalization
Direct Preference Optimization (DPO)
Hallucination Reduction
Semantic Recoverability
Lexical Fidelity
πŸ”Ž Similar Papers
No similar papers found.