Blackboard Intelligence Can Surpass Autoregressive on Globally Constrained Problems

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent limitation of autoregressive language models, which are constrained by causal generation trajectories and thus struggle with problems involving complex global constraints. To overcome this bottleneck, we propose a reasoning paradigm grounded in the blackboard architecture that leverages the arbitrary-order prediction capability of diffusion language models, specifically LLaDA-8B. By employing average confidence as a proxy metric for global coherence, our approach integrates test-time search and refinement mechanisms to transcend conventional sequential generation limitations. Experimental evaluations on tasks such as ZebraLogic demonstrate that the proposed method substantially outperforms both autoregressive baselines of comparable scale and state-of-the-art frontier large language models. These results validate the significant advantages of diffusion language models in handling reasoning tasks governed by global constraints.
📝 Abstract
Next-token prediction has driven remarkable progress in large language models, yet a growing body of evidence suggests that they can struggle on problems governed by complex global constraints. In this work, we focus on this regime and ask whether some of these limitations arise from the inference interface induced by next-token prediction itself. We study this question through blackboard intelligence: an inference-time perspective in which a model works on a fixed, revisable canvas and searches over candidate solution states rather than committing to a causal, left-to-right trajectory. We instantiate this idea with diffusion language models, whose any-order prediction interface naturally exposes predictions over partially filled solution states. Our key observation is that mean confidence, a simple model-internal quantity available from the standard masked diffusion objective, provides a useful proxy for global coherence and can guide inference-time search and revision. Empirically, across ZebraLogic, Nurse Rostering, and Job-Shop Scheduling, Blackboard consistently improves inference while holding the fine-tuned LLaDA-8B-Instruct checkpoint fixed and substantially outperforms same-scale autoregressive baselines, reaching 90.4% accuracy on ZebraLogic-Hard, 76.4% exact feasibility on Nurse Rostering, and 80.2% optimality on JSSP. Stronger autoregressive search and refinement also fail to close the gap on ZebraLogic-Hard, while Blackboard surpasses tested frontier LLMs there and on JSSP despite their substantially greater scale and strong test-time reasoning. We open-source our codebase at https://github.com/jwoosang1/blackboard-intelligence.
Problem

Research questions and friction points this paper is trying to address.

next-token prediction
global constraints
autoregressive models
blackboard intelligence
diffusion language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Blackboard Intelligence
Diffusion Language Models
Global Constraints
Mean Confidence
Inference-time Search