🤖 AI Summary
This study addresses the lack of precise distribution evaluation metrics for non-autoregressive language models, where existing likelihood bounds fail to accurately quantify discrepancies between generated and real data distributions. We propose SOL, a distance metric that leverages a fixed Transformer to extract hidden states for constructing empirical measures, and compares text distributions via a dual-sliced Wasserstein distance. Theoretically, we prove that SOL constitutes a strict mathematical metric under Transformer injection conditions. Empirically, SOL effectively detects distributional failures, recovers expected model trends, and provides stable sample estimates. By successfully re-evaluating multiple models trained on OpenWebText, this work fills the critical gap in quantitatively assessing distributional fit for non-autoregressive models.
📝 Abstract
Evaluating text generation requires measuring how well the generated distribution matches the data distribution. For autoregressive models, this is done by the perplexity. Diffusion and flow-based language models can only provide a likelihood bound, whose tightness differs between model families. Sample-based substitutes such as generative perplexity with entropy do not consider the distribution fit. We propose SOL, a distance between text distributions. Each sequence is represented by the empirical measure of its hidden states under a fixed transformer and the distributions of these measures are compared by the double sliced Wasserstein distance. We prove that SOL is a metric if the transformer is injective. Experiments show that SOL detects distributional failures, recovers expected model trends, and provides stable sample-based estimates. We put forward SOL to fill the gap in the current evaluation protocol used for non auto-regressive models. As a first step we use SOL to re-evaluate a variety of models trained on OpenWebText.