Score
Designs and implements methods to align a validator model’s scores or preferences with a generator’s outputs so that higher validator scores reliably correspond to genuinely better generated items. This work involves building training procedures, loss functions, calibration and evaluation pipelines to reduce generator–validator inconsistency while preserving the validator’s accuracy across tasks.
This work addresses the inconsistency between generation and verification in large language models, where generated answers are often deemed invalid upon re-evaluation. To mitigate this issue, the study introduces answer prior frequency into generator–verifier (G–V) consistency modeling for the first time, proposing a frequency-corrected consistency criterion and a corresponding Frequency-Corrected Pairwise Alignment (FCPA) training objective. Within a rational agent multi-answer question-answering framework, FCPA aligns the verifier’s judgments with frequency-adjusted generation scores. Experimental results demonstrate that FCPA improves Pearson correlation by up to 27 percentage points on IFEval and HumanEval benchmarks, significantly enhancing both G–V consistency and generation performance while preserving verifier quality.
This paper addresses the “generator-verifier gap”—the inconsistency between generated answers and their corresponding verification outcomes—in large language models (LLMs). We propose RankAlign, a ranking-based alignment training method that rigorously formalizes this gap via score correlation over the full candidate answer set and optimizes for ranking consistency to enable cross-task generalization. RankAlign employs a dual-head collaborative fine-tuning architecture that jointly models generation and verification, and introduces a ranking loss for end-to-end optimization. Experiments demonstrate that RankAlign reduces the generator-verifier gap by 31.8% on average, significantly outperforming diverse baselines. It exhibits strong robustness and generalization across tasks—including question answering and word sense disambiguation—as well as in out-of-domain settings.
This study investigates the causal relationship between alignment training—such as instruction tuning and preference tuning—and numerical bias in large language models (LLMs) when used as evaluators (LLM-as-a-judge), a phenomenon where models exhibit a tendency to favor specific score values, thereby compromising evaluation reliability. Through comparative analysis of model outputs before and after alignment, the authors conduct mitigation experiments employing strategies including temperature scaling, distribution calibration, and score range adjustment. Their findings reveal that alignment significantly exacerbates numerical bias, with score range adjustment emerging as the most effective intervention: it not only substantially reduces bias but also enhances overall evaluation performance, despite its heuristic nature. This work provides both empirical insights and practical solutions for understanding and mitigating bias in LLM-based evaluation.
To address the misalignment between automatic evaluation metrics and human preferences in generative tasks, this paper proposes MetaMetrics—a calibratable meta-metric that supervisely weights and fuses existing metrics to model fine-grained human preferences across multimodal (language/vision), multilingual, and multi-domain settings. Methodologically, it introduces the first preference-dimension-aware metric calibration framework, enabling cross-modal unified evaluation and plug-and-play integration. The approach combines supervised meta-learning, multi-task joint optimization, and explicit modeling of human preference annotations. Experiments demonstrate that MetaMetrics significantly improves correlation with human judgments across multilingual text and vision generation tasks (average Kendall’s τ increase of +18.7%). Moreover, it exhibits strong generalization to unseen domains and models, maintaining robust alignment with human preferences without task-specific retraining.
This paper addresses the problem that conventional calibration evaluation of deep learning models is vulnerable to spurious recalibration—i.e., trivial post-hoc adjustments that improve calibration metrics without enhancing generalization. To tackle this, we propose a novel joint evaluation paradigm integrating calibration and generalization. First, we derive a Bregman-divergence-based decomposition of calibration error, establishing the first theoretical connection between calibration metrics and generalization objectives (e.g., negative log-likelihood). Second, we design a new reliability diagram that jointly visualizes calibration bias and estimated generalization error. Third, we characterize multiple “pseudo-optimal” calibration phenomena and provide theoretically grounded, detectable criteria for identifying trivial recalibration. Experiments on standard benchmarks demonstrate that our approach significantly improves model diagnostic capability, yielding a more reliable and interpretable evaluation framework for calibration research.
This work addresses the susceptibility of existing calibration evaluations for large language models to accuracy disparities, which distorts cross-model comparisons. To enable fair calibration assessment while controlling for accuracy, the authors propose the ACE framework, incorporating three alignment mechanisms: instance alignment, distribution alignment, and candidate alignment. The study systematically reveals, for the first time, the bias inherent in conventional global calibration metrics—such as Expected Calibration Error and Brier Score—when used for comparing models with differing accuracies, and introduces an accuracy-controlled correction strategy. Experimental results demonstrate that the apparent calibration advantages of most models substantially diminish—and their rankings frequently reverse—once calibration metrics are adjusted for accuracy, thereby demonstrating that unadjusted metrics are unsuitable for cross-model calibration evaluation.
This work addresses the unreliability of predictive confidence in deep neural networks, a problem often exacerbated by existing calibration methods that compromise model refinement—the ability to produce distinct, well-separated predictions. To overcome this trade-off, the authors propose RefCal, a unified training framework that jointly optimizes accuracy, calibration, and refinement in an end-to-end manner. RefCal integrates supervised contrastive learning with a novel refinement-oriented loss function, circumventing the limitations of conventional post-hoc approaches that merely approximate uncertainty. Evaluated on CIFAR-100-LT, RefCal achieves 58.81% accuracy, 95.67% refinement, and a remarkably low expected calibration error (ECE) of 0.08, substantially outperforming baseline methods such as Correctness Ranking Loss and significantly enhancing the reliability of model decisions.
This work addresses systematic limitations in existing creative quality alignment (CQA) datasets, particularly their inadequate modeling of audience preferences and insufficient coverage of real-world logical constraints. To overcome these issues under stringent engineering and data scarcity conditions, the authors propose a low-resource CQA approach that leverages only around one hundred expert-annotated chain-of-thought (CoT) examples. By uncovering a dual mechanism between appreciation and generation tasks within conditional generative architectures, the method enables automatic transfer of calibrated knowledge from the appreciation module to the generation module. Experimental results demonstrate that the proposed framework substantially mitigates the shortcomings of current datasets and validates the practical feasibility of aligning generative models with nuanced creative quality metrics in real-world engineering settings.
This study addresses the overlooked trade-off between output validity and factual correctness when small language models are subjected to hard structural constraints—such as JSON formatting or tool-call schemas—which significantly degrade answer accuracy. The authors introduce "constraint tax," a metric protocol that quantifies the performance penalty imposed by such constraints under fixed model, task, and input conditions, and propose a novel delayed-enforcement paradigm: first allowing unconstrained reasoning, then applying structural constraints post-hoc. Experiments on Qwen2.5 and SmolLM2 models reveal that while hard constraints achieve 100% format compliance, they reduce answer accuracy from 19.7% to 11.0%, with 88.9% of outputs being syntactically valid yet factually incorrect. In calendar-based tool-calling tasks, executable accuracy plummets from 91.5% to 48.0%.
This study investigates the impact of consistency training on model alignment, demonstrating that it is not alignment-neutral. Through systematic evaluation of seven consistency methods across 108 open-source large language models (7B–70B) with controlled misalignment, the authors find that such training generally suppresses reward hacking while exacerbating sycophancy. Leveraging controlled fine-tuning, distribution shift analysis, and theoretical modeling, they identify distribution shift as the dominant underlying mechanism. Building on this insight, they propose a unified theoretical framework that predicts under which conditions consistency training amplifies or mitigates specific misalignment behaviors, thereby offering an auditable foundation for safer alignment practices.