Towards Spec Learning: Inference-Time Alignment from Preference Pairs

📅 2026-06-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Traditional alignment of large language models relies on handcrafted prompts or costly preference-based fine-tuning, which is inefficient and lacks interpretability. This work proposes Spec Learning, a novel framework that, for the first time, automatically generates human-readable and transparent natural language specifications from a small set of user preference pairs and brief instructions. During inference, these specifications conditionally guide model behavior without requiring any parameter updates. By circumventing black-box fine-tuning, Spec Learning outperforms Direct Preference Optimization (DPO) on preference-intensive, domain-specific datasets while providing interpretable behavioral rules. This approach significantly enhances both the efficiency and explainability of the alignment process.
📝 Abstract
Steering a large language model (LLM) toward a desired behavior typically relies on an iterative process of hand-crafting a prompt based on a careful inspection of the model's responses. This is an involved, brittle, and error-prone process. Preference-based fine-tuning is a more rigorous but often prohibitively expensive solution. We propose spec learning, a framework that relies on a brief user instruction and a small set of preference judgments. These are compiled into specifications in the form of natural-language prompts for an LLM. Specifications condition LLMs at inference time, and no parameter updates to the underlying models are required. We show that the responses generated based on the compiled specifications often outperform direct preference optimization (DPO) on datasets from specialized domains whose preference signal is dense. Unlike opaque weight updates, the resulting specifications are human-readable and double as interpretable and transparent written embodiments of the preference signal that produced them.
Problem

Research questions and friction points this paper is trying to address.

preference alignment
inference-time steering
specification learning
large language models
prompt engineering
Innovation

Methods, ideas, or system contributions that make the work stand out.

spec learning
preference alignment
inference-time conditioning
interpretable specifications
parameter-free steering
🔎 Similar Papers
No similar papers found.
D
Dhriti Krishnan
Department of Computer Science, Carnegie Mellon University
T
Tejas Goyal
Department of Computer Science, Carnegie Mellon University
J
Jaromir Savelka
Department of Computer Science, Carnegie Mellon University