Score
Designs and evaluates algorithms and selection policies that rank and choose candidate queries, questions, or attributes to present next so as to maximize expected information gain or task utility while minimizing redundancy, cost, or labeling effort. Works include building informativeness scoring functions, adaptive prioritization and stopping rules, and sequential or batch selection strategies for query-driven data collection or interaction.
This work addresses the challenge of applying traditional multi-winner voting rules in large-scale or attention-constrained settings, where eliciting complete preference rankings from voters is impractical. To overcome this limitation, the authors propose a structured-query framework for multi-winner elections that approximates an optimal committee by querying voters’ preferences over subsets of candidates within a limited budget. They formally define a cognitive cost function and axiomatic evaluation criteria, and introduce a query strategy based on recursively partitioning the candidate set. Experimental results demonstrate that this approach significantly outperforms alternative querying mechanisms across various election models and multi-winner rules—such as k-Borda—achieving high committee selection accuracy while substantially reducing the information acquisition cost.
This paper addresses the classical secretary problem—online selection of the best candidate based solely on relative rank information. To overcome the poor adaptability of conventional fixed-threshold policies, we propose a data-responsive heuristic framework featuring sequential threshold adaptation. It integrates five tunable rules—including expected-record thresholds, adaptive bias correction, and probabilistic early-stopping—and employs two-stage relaxation with local dynamic programming approximation to enhance robustness and decision efficiency. The method synergistically combines probabilistic modeling with lightweight ensemble learning. Extensive simulations across diverse scenarios validate its efficacy. Experimental results show that the framework achieves near-optimal performance with minimal intuitive hyperparameters, consistently outperforming classical strategies in both average-case performance and stability; the ensemble variant demonstrates the highest robustness.
This study addresses the challenge of identifying bias in search query suggestions, which is hindered by sparse contextual information and the limited number of suggestions provided per query round, resulting in an insufficient data foundation for robust analysis. To overcome these limitations, the authors propose a recursive algorithmic interrogation method that systematically constructs a query suggestion tree to uncover secondary and deeper-level suggestions, thereby moving beyond the conventional reliance on only top-tier recommendations. This approach substantially expands the dataset of suggestions associated with political figures and enables, for the first time, a thorough investigation into latent thematic group biases. Consequently, the method significantly enhances both the comprehensiveness and accuracy of bias detection in search suggestion systems.
This work addresses the limitations of existing large language model–based search agents, whose retrieval accuracy and reasoning performance are often compromised by the generation of low-quality intermediate queries. To mitigate this issue, the paper proposes a process reward–driven dual-level credit assignment mechanism coupled with a selective query optimization strategy. Furthermore, a three-stage curriculum learning framework—progressing from imitation to generalization—is designed to guide the agent in internalizing the ability to generate high-quality queries. Empirical evaluations demonstrate that the proposed approach significantly outperforms current state-of-the-art methods across multiple benchmarks, achieving notable improvements in both query quality and search efficiency.
This work addresses the challenge of effectively handling complex queries in retrieval-augmented generation (RAG) systems, where single-step retrieval often proves insufficient and existing reinforcement learning approaches suffer from unstable training due to combinatorial search spaces and difficulty in reward design. To overcome these limitations, the authors propose the Adaptive Complex Query Optimization (ACQO) framework, which integrates adaptive query decomposition, a ranking-scoring fusion module, and curriculum-based reinforcement learning. ACQO dynamically determines when to decompose queries and how deeply to perform multi-path retrieval, enabling stable and efficient optimization. The framework is compatible with diverse retrieval architectures and achieves state-of-the-art performance across three complex query benchmarks, significantly outperforming existing methods while simultaneously improving computational efficiency and generalization capability.
This work addresses the challenges of non-unique solutions and noise-induced temporary infeasibility in structured ranking and selection problems by proposing a unified framework, ENDS. The framework integrates answer-level acceptance sets, a constrained generalized likelihood ratio stopping rule, and a novel answer–trap decomposition mechanism, yielding a max-max-min eigenvalue characterization and a general information-directed sampling principle. By dynamically constructing acceptance sets, explicitly detecting traps, and incorporating cost-aware sampling, ENDS is broadly applicable to diverse settings such as multi-fidelity ranking and Condorcet winner identification. Empirical results demonstrate that the method achieves superior performance across a range of pure exploration tasks, confirming its generality and practical utility.
Large language models are susceptible to selection bias in adaptive prompting and program search, leading to an overestimation of the winning candidate’s performance under real-world deployment. This work proposes the SIREN protocol, which enables unbiased performance inference for the full tuning-to-deployment pipeline under a fixed tuning budget by freezing the candidate set, decoupling selection and evaluation data, and incorporating an entry-wise Gaussian multiplier bootstrap. SIREN is the first method to simultaneously support accurate estimation of program-level performance curves on limited-budget grids and construct confidence intervals for both within-budget and cross-budget comparisons. Empirical results demonstrate that conventional winner-reporting practices exhibit substantial optimistic bias, whereas SIREN closely approximates the true evaluation target under finite-sample conditions, offering reliable guidance for deployment decisions.
This work addresses the issue that stochastic re-ranking strategies may induce worst-case fluctuations in retrieval effectiveness prior to re-ranking. It presents the first theoretical framework to quantify the maximum absolute change in Discounted Cumulative Gain (DCG)—termed “re-ranking risk”—caused by such randomness. This risk is determined by the positional distribution of relevant documents in the initial retrieval list. By integrating probability theory with ranking metric analysis, the study performs an extremal analysis of DCG variation and derives a computable upper bound for this risk. Experiments on the TREC Fairness 2022 dataset demonstrate strong alignment between theoretical predictions and observed DCG fluctuations, thereby filling a critical theoretical gap in pre-re-ranking risk assessment.
Traditional scoring functions exhibit inherent theoretical limitations in balancing utility and fairness for ranking tasks. This work rigorously proves, for the first time, that such functions cannot span the entire Pareto frontier of utility–fairness trade-offs. The universality of this limitation is demonstrated through counterexamples constructed across multiple settings, including deterministic versus stochastic scenarios and single-query versus multi-query contexts. To address this gap, the paper introduces a semi-greedy post-processing algorithm that efficiently approximates the ideal solution within a general formal fairness framework. Experimental results show that the proposed method substantially outperforms existing scoring mechanisms and achieves performance close to that of exhaustive post-processing within computationally feasible bounds.