Score
Designs and implements algorithms and systems that select a variable-size subset of candidate items by dynamically choosing a probability threshold or sparsity level instead of a fixed-count cutoff. Builds methods to estimate item attention/selection probabilities, adjust the retrieval budget at runtime, and analyze correctness and error bounds for the adaptive top-p selection procedure.
Large language models are susceptible to selection bias in adaptive prompting and program search, leading to an overestimation of the winning candidate’s performance under real-world deployment. This work proposes the SIREN protocol, which enables unbiased performance inference for the full tuning-to-deployment pipeline under a fixed tuning budget by freezing the candidate set, decoupling selection and evaluation data, and incorporating an entry-wise Gaussian multiplier bootstrap. SIREN is the first method to simultaneously support accurate estimation of program-level performance curves on limited-budget grids and construct confidence intervals for both within-budget and cross-budget comparisons. Empirical results demonstrate that conventional winner-reporting practices exhibit substantial optimistic bias, whereas SIREN closely approximates the true evaluation target under finite-sample conditions, offering reliable guidance for deployment decisions.
This paper addresses the hypothesis selection problem: given i.i.d. samples from an unknown distribution and a set of candidate distributions, the goal is to select, with high probability, a hypothesis whose total variation (TV) distance to the true distribution is at most $C cdot mathrm{OPT} + varepsilon$, where $mathrm{OPT}$ is the minimum TV distance achievable within the candidate set. We present the first algorithm achieving the optimal approximation factor $C = 3$, optimal sample complexity, and near-linear running time $ ilde{O}(n/(deltavarepsilon^2))$, significantly improving upon prior work. Furthermore, under settings where $mathrm{OPT}$ is known or preprocessing is allowed, we design more efficient algorithms with only weak dependence on the error and confidence parameters. Technically, our approach integrates precise TV distance estimation, refined probabilistic analysis, sampling optimization, and a novel high-probability error control mechanism—thereby resolving, for the first time, the long-standing open problem of simultaneously attaining optimal approximation factor and optimal runtime.
The Plackett-Luce model overlooks cognitive constraints and fails to capture “consider-then-choose” behavior under large option sets. Method: We propose a “consider-then-choose” framework centered on the identifiability of unobserved item consideration probabilities. Assuming known utilities but unknown consideration probabilities, we develop a convex optimization–based constrained propagation algorithm that systematically tightens bounds on consideration probabilities. Contribution/Results: We establish the first theoretical identifiability results for consideration probabilities in top-k rankings—characterizing both their relative ordering and absolute upper/lower bounds. Empirical validation on psychological experiment data demonstrates joint estimation of utility parameters and consideration probability bounds, confirming both the theoretical guarantees and practical efficacy of the derived bounds.
Traditional retrieval systems often exhibit near-random selectivity at high recall levels, limiting the performance of downstream large language model (LLM) tasks. This work proposes the Bits-over-Random (BoR) metric, which introduces an opportunity-correction mechanism grounded in information theory and models the random baseline using the hypergeometric distribution to quantify the true selectivity of retrieval results. Experiments reveal that BM25 and SPLADE achieve BoR ≈ 0 at K=100, indicating a practical loss of selectivity. In contrast, BoR effectively discriminates system performance across BEIR, SciFact, and MS MARCO benchmarks, approaching theoretical upper bounds and demonstrating broad applicability and practical guidance—particularly in deep retrieval and LLM tool selection scenarios within Retrieval-Augmented Generation (RAG) frameworks.
This paper studies preference learning from choice feedback over a dynamic item set, aiming to identify either the optimal item or the complete preference ranking with minimal samples and high confidence. We propose two algorithms—Nested Elimination (NE) and Nested Partitioning (NP)—that provide the first non-asymptotic, instance-dependent sample complexity bounds for arbitrary strict preference structures; NE achieves asymptotic optimality in the information-theoretic worst case, while NP attains constant-factor optimality. Our approach integrates multi-dimensional random walk modeling, divide-and-conquer strategies, and an information-theoretic analysis framework. Rigorous theoretical analysis is complemented by extensive experiments on both synthetic and real-world datasets, demonstrating superior efficiency and robustness. The core contribution is the establishment of the first algorithmic framework for dynamic preference learning that simultaneously ensures practical applicability and theoretical optimality.
本文提出一种基于置信区域的筛选框架,用于解决模拟系统可接受性问题,保证高概率筛选出所有或每个可接受系统,并支持并行化。
This work addresses the group testing problem of identifying $k$ defective items among $n$ by introducing an adjustable-threshold model, where each test returns positive if the number of defectives in the tested group meets or exceeds a pre-specified threshold. Building on this framework, the authors design efficient test matrices and decoding algorithms, complemented by information-theoretic analysis. They establish tight achievability and converse bounds under fixed thresholds and demonstrate that, when the maximum threshold is unbounded or sufficiently large, recovery succeeds with high probability near the information-theoretic limit. Notably, in the dense regime where $k = \Theta(n^\theta)$ with $\theta \to 1$, the upper and lower bounds on the number of tests coincide, and in the unbounded-threshold setting, the testing rate approaches the theoretical maximum of one.
论文提出ASC和K-ASC算法,有效解决了大规模推荐系统中Swing分数计算效率低的问题,通过结合随机算法处理高、低度查询项,显著提高了计算速度。
This study addresses the limitation of approximate rejection sampling, which relies on known distributional properties to set thresholds. To overcome this, we propose Uniform Race, a parameter-free algorithm that constructs a scoring mechanism by combining importance weights with uniform random variables. By selecting the candidate with the maximum score, the method achieves an exact approximation of the target distribution without requiring preset thresholds, simultaneously attaining the optimal error upper bound across all fixed thresholds and uniquely determining selection probabilities under specific conditions. Theoretical analysis demonstrates that Uniform Race yields an exponential reduction in total variation distance error compared to baseline methods. Furthermore, experiments on mathematical reasoning tasks with large language models empirically validate both the efficiency and accuracy of the proposed approach.
This work proposes a novel approach to integrate multi-source expert prior knowledge into best subset selection. Addressing the limitations of purely data-driven feature selection, the method embeds expert assessments of feature relevance—aggregated in the form of Poisson binomial distributions, pairwise win probabilities, or normalized average ranks—as log-odds penalty terms within a mixed-integer optimization (MIO) objective function via a maximum a posteriori (MAP) framework. This constitutes the first theoretically principled and analytically tractable Bayesian–MIO joint formulation, which naturally reduces to classical best subset regression in the absence of expert information. Theoretical derivations and algorithmic implementation have been completed, with empirical results forthcoming.