Score
Designs and evaluates methods that consume a small set of labeled affinity comparisons as in‑context demonstrations and, without gradient updates, infer and produce ranked orderings or pairwise affinity predictions among candidate items. Builds prompting or model‑inference procedures and associated evaluation to exploit few‑shot context and generalize comparative affinity relationships to unseen examples.
Example selection in in-context learning (ICL) is highly sensitive, yet existing methods lack a unified optimization objective and suffer from fragmented theoretical foundations. Method: We propose the first unified, representation-based evaluation metric for ICL—comprising two computable, downstream-accuracy–correlated dimensions: *affinity* (semantic similarity between examples and the query) and *diversity* (representational dissimilarity among examples). Our approach extracts intermediate-layer representations from pre-trained language models, constructs a similarity matrix and a diversity measure, and jointly optimizes the example set accordingly. Contribution/Results: Across multiple benchmarks, our metric exhibits strong correlation with task accuracy (average Spearman ρ > 0.85). Moreover, it provides the first unified explanatory framework that reproduces and interprets the efficacy of prominent selection strategies—including KNN, BERT-KNN, and Self-Adaptive—thereby establishing an interpretable, generalizable theoretical foundation for ICL example selection.
This work addresses the problem of efficiently selecting the top-k in-context demonstration examples for few-shot learning. We propose the first linear-time demonstration selection algorithm based on gradient estimation in the input embedding space. Our method’s core innovation lies in applying a first-order Taylor approximation of the model output with respect to input embeddings to quantify the influence of each individual demonstration on the target prediction—without requiring full forward passes. By integrating randomized subset sampling with influence-score aggregation, we avoid exhaustive inference over all candidate demonstrations. Evaluated on six benchmark datasets, our approach achieves an average prediction error below 1%, accelerates inference by 37.7× over exhaustive search, and improves average accuracy by 11% over current state-of-the-art methods on a 34B-parameter language model.
To address the ambiguity and subjectivity in evaluating language model responses to fuzzy queries (e.g., subjective or open-ended questions), this paper proposes a contextualized evaluation protocol that embeds structured context—such as synthesized user identities, query intents, and utility criteria—into the assessment process. Methodologically, it integrates context synthesis modeling, multi-dimensional human evaluation design, cross-dimensional quality analysis, and bias-sensitivity quantification. Our study is the first to systematically uncover mainstream models’ implicit preference for WEIRD (Western, Educated, Industrialized, Rich, Democratic) contexts and reveal pronounced asymmetry in their contextual adherence capabilities. Experiments demonstrate that the protocol reverses relative model win rates, mitigates superficial stylistic biases, and yields fine-grained behavioral insights—thereby substantially enhancing evaluation objectivity, interpretability, and diagnostic utility.
Conventional drug–target affinity prediction models suffer from biased evaluation: random test-set splitting artificially enriches high molecular similarity samples, thereby masking model failure in low-similarity generalization scenarios. Method: We propose a similarity-aware evaluation framework featuring a novel, controllable similarity-distribution data splitting strategy, formulated as a differentiable optimization problem and solved efficiently via gradient descent. Using fingerprint-based molecular similarity metrics, we conduct multi-model benchmarking across four standard datasets. Results: Our framework reveals severe performance degradation—averaging 30–50% drops—in low-similarity regimes, exposing critical limitations of existing methods. It significantly enhances evaluation fidelity and model interpretability, establishing a new paradigm for reliable deployment of affinity prediction models.
Existing in-context learning (ICL) methods suffer from sensitivity to demonstration selection and poor generalization of static, single-dimensional selection strategies. To address this, we propose Iterative Demonstration Selection (IDS), a novel framework that dynamically integrates semantic relevance and diversity. IDS introduces, for the first time, an iterative feedback mechanism grounded in zero-shot chain-of-thought (Zero-shot-CoT) reasoning paths, enabling task-adaptive, multi-dimensional demonstration filtering. It further incorporates iterative context reconstruction and reasoning-path-driven semantic matching, augmented by majority-voting ensemble to enhance stability. Evaluated across diverse tasks—including logical reasoning, question answering, and topic classification—IDS consistently outperforms state-of-the-art ICL demonstration selection approaches. It demonstrates superior robustness to input perturbations and stronger cross-task generalization capability, establishing a new benchmark for adaptive, reasoning-aware demonstration selection in few-shot ICL.
This work addresses the high sensitivity of in-context learning to prompt examples and the prohibitive cost of searching for optimal example combinations. To this end, the authors propose DiSP, a novel framework that advances a “judgment over search” paradigm: queries are stratified by difficulty through sampling and discrimination strategies, a lightweight router is trained to predict query difficulty, and dedicated discriminators are assigned to each difficulty tier. During inference, DiSP employs a budget-aware “accept-and-stop” decision policy, outputting a risk-diagnosis label upon failure. Evaluated on five classification benchmarks, DiSP achieves up to a 3.4% absolute accuracy gain over strong baselines and accelerates end-to-end inference by up to 23×.
This work addresses the challenge of efficiently and reliably evaluating newly released machine learning models on unlabeled data without incurring costly annotations or repeated fine-tuning. The authors propose MetaEvaluator, the first model-agnostic, label-free, and training-free evaluation framework that eliminates the need for per-model retraining. Leveraging meta-learning, MetaEvaluator learns a transferable evaluator initialization from a pool of reference models, enabling rapid unsupervised performance estimation for models of unseen architectures and modalities. Extensive experiments demonstrate that MetaEvaluator consistently and accurately predicts model performance across diverse datasets and model types, substantially reducing evaluation costs and facilitating large-scale benchmarking on unlabeled data.
This work addresses the challenge of efficiently selecting high-quality in-context examples under limited prompt budgets to improve model accuracy while controlling computational overhead. We propose Meta-Sel, a lightweight supervised meta-learning approach that, for the first time, applies supervised meta-learning to example selection. Meta-Sel trains an interpretable scoring function—combining TF-IDF cosine similarity and length compatibility ratio—on a constructed meta-dataset, enabling fast, deterministic, and auditable example ranking without fine-tuning or additional large model invocations. Experiments across four intent classification datasets and five open-source large language models demonstrate that Meta-Sel consistently achieves superior performance, with particularly notable gains for smaller models, all while maintaining minimal selection overhead.
This paper addresses the problem of model-free inference of global rankings from noisy pairwise comparisons (e.g., tennis match outcomes), where both the latent object strengths and the functional mapping from strengths to win probabilities are unknown. To overcome the limitations of parametric models—such as Bradley–Terry—which impose strong prior assumptions on the link function (e.g., logistic), we propose the first Bayesian nonparametric framework that jointly infers latent strength parameters and an unknown response function. Our method integrates variational adaptive function modeling with MCMC sampling and incorporates empirical calibration to enhance robustness. Evaluated on real-world datasets spanning sports and academic citation networks, our approach significantly outperforms baselines relying on prespecified link functions, demonstrating superior ranking accuracy and generalization stability even under model misspecification.
This work addresses the challenge of efficiently selecting the most informative pairwise comparisons under limited annotation budgets to improve alignment in preference-based large language model post-training. Framing comparison selection as a sampling design problem within the Direct Preference Optimization (DPO) framework, this study establishes the first theoretical connection between comparison pair sampling and policy suboptimality, deriving matching upper and lower bounds. Building on this analysis, the authors propose an explicit sampling criterion based on the Fisher information matrix to guide data acquisition. Experimental results demonstrate that the proposed method significantly outperforms existing heuristic strategies on both synthetic benchmarks and real-world language model post-training tasks, achieving substantially higher sample efficiency.