Where Do Test-Time Scaling and Training Fall Short in Individual Stance Prediction?

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant limitations of existing test-time scaling and training methods in individual stance prediction. We construct STANCE-BENCH, a benchmark that reveals four major failure modes of large language models in this task, including consensus errors across the generation, selection, and learning stages. Building upon Qwen3-8B, we propose an enhanced framework that integrates direct scoring with historical evidence support, employing supervised fine-tuning, reinforcement learning, and test-time scaling strategies for targeted optimization. The proposed approach effectively overcomes existing deficiencies, achieving a Macro F1 score of 21.83% on the test set—a substantial improvement over the 19.27% baseline. This work provides a reliable evaluation framework and a principled optimization pathway for advancing individual stance prediction.
📝 Abstract
Test-time scaling and post-training have improved LLM performance in coding and mathematical reasoning, but their effectiveness for individual stance prediction remains unclear. We study this question by predicting a person's stance in a new discussion from their history. We evaluate widely used test-time scaling strategies and post-training methods, such as supervised fine-tuning and reinforcement learning, and identify four failure modes across generation, selection, and learning: (1) incorrect consensus, where repeated samples agree on the wrong stance; (2) selection failure, where generation covers the observed stance but selection misses it; (3) response overfitting, where supervised fine-tuning improves imitation but harms prediction; and (4) early plateau, where reinforcement learning shows modest initial gains followed by limited further improvement. We expose these failures using STANCE-BENCH, which contains 2499 prediction tasks from 500 Hacker News users. Guided by this analysis, we explore a simple approach that combines direct scores for all candidate stances with explicit assessments of support from the individual's history. On the 781-task test set, this approach achieves 21.83 discussion-specific Macro F1 with Qwen3-8B, compared with 19.27 for direct scoring. Our results motivate evaluating candidate generation, final selection, and person-specific evidence use separately.
Problem

Research questions and friction points this paper is trying to address.

individual stance prediction
test-time scaling
post-training
large language models
failure modes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Individual Stance Prediction
Test-Time Scaling
Failure Modes
STANCE-BENCH
Post-training
🔎 Similar Papers
No similar papers found.