Understanding Confabulation and Rethinking Reconstruction in Activation Explanations

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high hallucination rates and insufficient readability in Natural Language Autoencoder (NLA) explanations by proposing the Flow-NLA framework. Departing from the conventional single-point reconstruction paradigm, this method leverages the diffusion likelihood bound to probabilistically model compatible activation distributions, thereby replacing point-wise optimization with distribution-level optimization to enhance explanation generation quality. Additionally, a standardized evaluation system is established. Experimental results demonstrate that Flow-NLA preserves predictive utility across multiple large language models while significantly suppressing hallucinations and improving explanation readability. This work provides a novel paradigm for generating high-quality NLA explanations.
📝 Abstract
Natural Language Autoencoders (NLAs) produce unsupervised text explanations of a model's activations: a verbalizer describes an activation and a reconstructor learns to recover it from this text. Under the established point-reconstruction NLA training recipe, explanations become more useful for predicting model behavior while also increasingly introducing unsupported details and exhibiting writing defects. To assess these changes separately, we introduce a standardized evaluation framework for unstructured NLA explanations, measuring information recoverable from explanations, contextual support for their claims, and writing quality. To address confabulation and writing defects, we move beyond predicting a single activation: explanations can distinguish distributions of possible activations even when their means and optimal point-reconstruction rewards are identical. We introduce Flow-NLA, which models the distribution of activations compatible with an explanation and trains the verbalizer using a diffusion likelihood bound. Across Qwen, Gemma, and Apertus, this richer signal retains the utility gains of point reconstruction while curbing the growth of confabulation and writing defects, opening up a direction for improving activation-derived training to encourage more informative, supported, and readable explanations. Code and evaluation prompts will be made publicly available upon acceptance.
Problem

Research questions and friction points this paper is trying to address.

Natural Language Autoencoders
Confabulation
Activation Explanations
Point Reconstruction
Writing Defects
Innovation

Methods, ideas, or system contributions that make the work stand out.

Natural Language Autoencoders
Flow-NLA
Diffusion Likelihood Bound
Confabulation
Activation Distribution Modeling
🔎 Similar Papers
2024-07-30Conference on Empirical Methods in Natural Language ProcessingCitations: 0