Sample What You Say: Aligning Language Models to Sample the Distributions They State

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation that language models, despite being capable of correctly articulating a target distribution, struggle to sample from it accurately. To overcome this, we propose a witness advantage mechanism based on Maximum Mean Discrepancy (MMD), integrated with Group Relative Policy Optimization (GRPO) to align model sampling behavior. By introducing a witness function, this mechanism resolves the learning signal degradation caused by intra-group relative centering in standard GRPO, enabling closed-form computation of per-sample independent advantages so that generated samples precisely match the target distribution. Our approach significantly reduces the total variation distance between the generated and target distributions while effectively preserving the model's general capabilities.
📝 Abstract
Language models are increasingly used to sample from a specified distribution, for instance, to simulate survey respondents or generate synthetic data. Instruction-tuned models can state such a distribution correctly and still fail to sample from it. Prompting and changes to decoding reduce this mismatch only partly, which motivates training with policy optimization. Group relative policy optimization (GRPO) is a natural fit for this problem because it already samples a group of rollouts per prompt, and the group's empirical distribution can be compared with the target. However, scoring the group as a whole gives every rollout the same reward. Group-relative centering then sets all advantages to zero, and the model receives no learning signal. To give each rollout its own signal, we introduce the witness advantage, a per-rollout advantage derived from maximum mean discrepancy (MMD). It trains a model to match a target distribution over a finite set of outcomes. The MMD between the model's distribution and the target has a witness function that measures how over- or under-produced each outcome is. Each rollout's advantage estimates the negative witness at its outcome, so a rollout is rewarded for an outcome the group under-produces and penalized for one it over-produces. The witness advantage is computed in closed form from the group's outcome counts, and we use it as the reward in GRPO. On unseen target distributions, training with the witness advantage substantially reduces the total variation distance to the target while largely preserving the model's general capabilities.
Problem

Research questions and friction points this paper is trying to address.

language models
distribution alignment
sampling
instruction-tuned models
distribution mismatch
Innovation

Methods, ideas, or system contributions that make the work stand out.

Witness Advantage
Maximum Mean Discrepancy (MMD)
Group Relative Policy Optimization (GRPO)
Distribution Alignment
Language Model Sampling
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.