Score
Design, build, and evaluate methods that produce and manage sets of candidate items or actions by generating, sampling, scoring, and pruning proposals — including rapid or scan-order generation, candidate sampling strategies, and candidate pruning or filtering pipelines. Implement and analyze cheap proxy-based screening (including thermo-inspired proxies) and prioritization schemes to cheaply rank or discard candidates, support in-loop policy-side filtering, and choose sampling strategies that control recall, cost, and variance during training.
This study addresses the lack of empirical evaluation regarding the error costs and institutional compatibility of AI systems in the initial screening of high-risk research proposals. It introduces a novel assessment framework that explicitly incorporates error asymmetry—particularly the critical impact of irreversible false negatives—and institutional governance requirements. The authors compare a rule-based TF-IDF approach with large language model (LLM) semantic classification in the context of national-level grant proposal screening. Results demonstrate that the TF-IDF method substantially outperforms the LLM, achieving a recall of 78.95% (versus 45.82%) and generating only 68 false negatives (versus 175), thereby significantly reducing the erroneous rejection of high-quality proposals. The findings underscore that transparency and auditability should take precedence over model complexity, offering a new paradigm for institutionally aligned AI-assisted peer review.
This study investigates the impact of Initial Screening Order (ISO) on fairness and selection quality in human candidate screening. We identify two distinct screening behaviors—best-k (optimizing for top candidates) and good-k (satisfying a quality threshold)—and propose the first computational model capturing time-varying inconsistency in human reviewers and its interaction with ISO. Our evaluation framework integrates position-bias analysis, quantification of both group-level and individual-level fairness, and behavioral simulation. Experimental results demonstrate that, under the good-k paradigm, ISO preserves group fairness but substantially degrades individual fairness and reduces top-k selection quality. This work provides the first systematic characterization of the latent bias mechanisms induced by ISO, establishing a theoretical foundation and a generalizable methodology for designing bias-mitigating screening strategies.
Bayesian inference under computationally expensive and nonsmooth likelihoods remains challenging due to prohibitive evaluation costs and the absence of reliable gradients. Method: This paper proposes a subset-driven delayed-acceptance MCMC framework comprising: (1) a data-driven surrogate model evaluated on random subsets—eliminating reliance on inaccurate gradients or Taylor approximations; (2) a computation-aware adaptive controller that jointly tunes proposal scale and subset size; and (3) a hierarchical delayed-acceptance mechanism that rapidly filters candidates via the coarse surrogate and rigorously validates them on the full dataset. Results: On real-world high-cost inference tasks—including disease modeling—the method significantly reduces sampling error under fixed computational budgets, balancing exploration efficiency and posterior accuracy. It outperforms state-of-the-art baselines (e.g., standard DA-MCMC, HINTS) in both convergence speed and estimation accuracy.
This work addresses the limitations of traditional outcome-based reinforcement learning in retrieval-augmented generation, where sparse rewards and ambiguous credit assignment often lead models to arrive at correct answers through flawed reasoning paths—a phenomenon known as process hallucination. To mitigate this, the authors propose ProRAG, a novel framework that integrates online policy exploration with learnable step-level process rewards. ProRAG employs a four-stage pipeline: policy warm-up, Monte Carlo Tree Search (MCTS)-based process reward modeling, reward-guided reasoning refinement, and reinforcement training with a dual-granularity advantage mechanism. This approach enables precise feedback on each action within long-horizon reasoning trajectories, effectively decoupling local behaviors from global outcomes. Evaluated across five multi-hop reasoning benchmarks, ProRAG significantly outperforms both outcome-driven and process-aware baselines, demonstrating particularly strong gains on complex, long-chain reasoning tasks.
This work addresses the computational infeasibility of traditional sequential feature selection methods in ultra-high-dimensional settings and the inability of univariate ranking approaches to capture feature interactions. To overcome these limitations, the authors propose a Stochastic Sequential Search (SSS) framework that dramatically reduces computational overhead by evaluating only a small subset of candidate features at each step through a budget-constrained sampling mechanism. The method innovatively integrates an online learning–driven dependency-aware statistic, temperature-controlled Softmax sampling with a uniform exploration floor, and embeds these components into a stochastic variant of Sequential Floating Forward Selection (SFFS), termed sSFFS. Experiments demonstrate that sSFFS achieves performance on par with or superior to full SFFS and state-of-the-art ranking methods on Madelon (500D), Gisette (5,000D), and Reuters (10,105D) datasets, while incurring substantially lower computational costs—marking the first scalable sequential search approach that effectively models feature interactions in high dimensions.
This work addresses the high cost of ground-truth evaluation in chemical and materials design, where existing machine learning surrogate models often lack reliability guarantees. Departing from conventional reliance on prediction accuracy metrics such as R²—which can paradoxically increase the risk of worst-case selections—the study proposes “rank preservation” as a core criterion for surrogate validation. It formally introduces the concept of “selection tax” and derives its theoretical upper and lower bounds. A safety certification framework for surrogates is established through selection-aware auditing, rank correlation analysis, and multi-task ground-truth validation. Experiments demonstrate that the proposed audit statistics achieve Spearman correlations of 0.80–0.99 with actual search performance, substantially outperforming R² (as low as 0.33). Certified screening strategies based on this framework reduce evaluation costs by up to 25-fold.