🤖 AI Summary
This study addresses the high hallucination rates and insufficient readability in Natural Language Autoencoder (NLA) explanations by proposing the Flow-NLA framework. Departing from the conventional single-point reconstruction paradigm, this method leverages the diffusion likelihood bound to probabilistically model compatible activation distributions, thereby replacing point-wise optimization with distribution-level optimization to enhance explanation generation quality. Additionally, a standardized evaluation system is established. Experimental results demonstrate that Flow-NLA preserves predictive utility across multiple large language models while significantly suppressing hallucinations and improving explanation readability. This work provides a novel paradigm for generating high-quality NLA explanations.
📝 Abstract
Natural Language Autoencoders (NLAs) produce unsupervised text explanations of a model's activations: a verbalizer describes an activation and a reconstructor learns to recover it from this text. Under the established point-reconstruction NLA training recipe, explanations become more useful for predicting model behavior while also increasingly introducing unsupported details and exhibiting writing defects. To assess these changes separately, we introduce a standardized evaluation framework for unstructured NLA explanations, measuring information recoverable from explanations, contextual support for their claims, and writing quality. To address confabulation and writing defects, we move beyond predicting a single activation: explanations can distinguish distributions of possible activations even when their means and optimal point-reconstruction rewards are identical. We introduce Flow-NLA, which models the distribution of activations compatible with an explanation and trains the verbalizer using a diffusion likelihood bound. Across Qwen, Gemma, and Apertus, this richer signal retains the utility gains of point reconstruction while curbing the growth of confabulation and writing defects, opening up a direction for improving activation-derived training to encourage more informative, supported, and readable explanations. Code and evaluation prompts will be made publicly available upon acceptance.