Don't Commit Alone: Joint Token Commitment in Diffusion Large Language Models

📅 2026-07-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Diffusion-based large language models decode multiple tokens in parallel during denoising, often neglecting inter-positional dependencies and thereby introducing factorization errors. This work proposes CoCommit, a method that introduces a token-gating coordination mechanism following conventional token selection. By briefly delaying commitment and reusing the final layer of the backbone network for localized forward passes, CoCommit enables collaborative decoding among selected positions, better approximating the joint distribution. Notably, this approach achieves multi-token joint commitment for the first time without additional model parameters, relying solely on token guidance and partial recomputation to effectively mitigate conditional total correlation error. Evaluated on LLaMA2.1-mini with LoRA adapters, CoCommit consistently improves accuracy across all six benchmarks, with particularly pronounced gains on reasoning and exact-answer tasks.
📝 Abstract
Diffusion large language models (dLLMs) commit multiple tokens per denoising step by decoding each selected position independently from the shared context; when those positions are dependent, the resulting factorization error is captured by conditional total correlation, which confidence-based selection cannot observe from marginals alone. We propose CoCommit, a marker-gated coordination pass that briefly defers commitment: after the usual bundle selection, a learned marker announces the commit set and the backbone's last-$n$ layers are re-applied so marked positions coordinate -- approximating joint-mode decoding -- before greedy argmax writes tokens. The method reuses existing weights with one extra partial forward pass and no auxiliary model. On LLaDA2.1-mini with LoRA adapters and matched greedy inference, joint commitment improves accuracy on all six benchmarks we evaluate, with the largest gains on reasoning and exact-answer tasks.
Problem

Research questions and friction points this paper is trying to address.

diffusion large language models
token commitment
factorization error
conditional total correlation
joint decoding
Innovation

Methods, ideas, or system contributions that make the work stand out.

joint token commitment
diffusion LLMs
conditional total correlation
marker-gated coordination
CoCommit
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.