ABC-Align: Prediction-Powered Alignment with Adaptive Bias Control

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity of human preference data in language model alignment and the systematic biases inherent in AI-generated pseudo-labels. To mitigate these challenges, it proposes an adaptive bias control mechanism that leverages large-scale pseudo-labels to reduce estimation variance while performing lightweight bias correction using limited human annotations. A plug-in estimator is introduced to automatically modulate the correction intensity, achieving an optimal bias-variance trade-off. This approach is compatible with mainstream alignment algorithms, including RLHF, DPO, and GRPO, and is integrated within a prediction-driven semi-supervised learning framework for efficient alignment. Empirical results demonstrate that the proposed method significantly outperforms existing semi-supervised baselines in data-scarce scenarios, offering a more robust and efficient solution for large language model alignment.
📝 Abstract
Language model post-training is often bottlenecked by the need for human-collected preference data, which is expensive and difficult to scale. Reinforcement learning from AI feedback (RLAIF) style approaches that leverage pseudo labels offer an abundant alternative but introduce systematic biases that degrade downstream alignment. Recent general-purpose semi-supervised methods correct for teacher bias using a small set of human-labeled examples, but suffer from high variance especially when human annotations are scarce. To this end, we propose ABC-Align, leveraging abundant pseudo label signal to minimize variance and applying a lightweight, adaptive correction grounded in the human-labeled subset. The correction strength is tuned automatically during training using plug-in estimates of the relevant bias--variance quantities. On LLM alignment with RLHF, DPO, and GRPO where human feedback is scarce, we empirically demonstrate that ABC-Align achieves superior performance over prior semi-supervised baselines in a series of experiments on an increasing scale. Our code is available at https://github.com/SewoongLab/abc-align .
Problem

Research questions and friction points this paper is trying to address.

Language model alignment
Pseudo labels
Systematic bias
Semi-supervised learning
Preference data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Bias Control
Prediction-Powered Alignment
Pseudo Labels
Semi-supervised Learning
Bias-Variance Tradeoff
🔎 Similar Papers
2024-06-05arXiv.orgCitations: 1