Where Models Converge and Humans Diverge: A Coverage Framework for Distributional Pluralism in Open-Ended Generation

๐Ÿ“… 2026-08-05
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
While large language models (LLMs) produce plausible outputs in open-ended generation tasks, their output distributions are often narrow, lacking the diversity and cultural breadth characteristic of human writing. Existing evaluation methods struggle to systematically quantify such distributional discrepancies. This work proposes the first evaluation framework grounded in the empirical distribution of human-authored texts, introducing two novel metricsโ€”LLM Coverage (LLM-Cov) and In-Boundary Rate (IBR)โ€”that disentangle content plausibility from distributional breadth to measure the cultural accessibility of model generations. Experimental results demonstrate that current LLMs predominantly generate outputs clustered near the center of human response space, revealing a significant limitation in cultural coverage. The proposed framework thus offers a measurable pathway toward enhancing generative diversity.
๐Ÿ“ Abstract
When a large language model (LLM) writes Harry Potter fanfiction, it reliably produces fundamental elements of the Hogwarts universe, such as recognizable places and characters. Human-written Harry Potter fanfictions, however, typically include these fundamentals and much more, incorporating stylistically irregular content and relationship-diverse plotlines. This gap between LLM and human writing has been noted across a variety of domains. LLMs tend to produce "average" writing, while human writing contains more diverse content that covers a broader distribution. Existing work has shown the existence of this distributional "gap", but no work has proposed a systematic way to measure it. Our paper proposes a human-grounded framework that uses the empirical distribution of human writing on a topic to measure the distributional breadth of LLM-generated content on that same topic. We propose two metrics, LLM Coverage (LLM-Cov) and In-Boundary Rate (IBR), that separate the plausibility of LLM content from its distributional breadth. Across ideation and narrative tasks, we find that current LLMs produce plausible but narrow content that concentrates near the center of the human response space. Our framework can enable researchers to better assess the distributional breadth of LLM-authored content, which we term its "cultural reach".
Problem

Research questions and friction points this paper is trying to address.

distributional pluralism
large language models
human writing
content diversity
cultural reach
Innovation

Methods, ideas, or system contributions that make the work stand out.

distributional pluralism
LLM Coverage
In-Boundary Rate
cultural reach
open-ended generation
๐Ÿ”Ž Similar Papers
No similar papers found.