Score
Designs and implements algorithms and procedures to compute k-th order statistics (quantile selection) and value-at-risk (VaR) measures from samples or distribution descriptions, including exact computations for discrete distributions. Focuses on efficient selection and partitioning techniques that avoid full sorting and improve time complexity (e.g., linear-time selection or expected sublinear performance in structured cases).
This work addresses the high computational cost of traditional methods for evaluating risk measures—such as Value-at-Risk (VaR) and Conditional Value-at-Risk (CVaR)—under large-scale discrete random variables. The authors propose two novel algorithms, QuickVaR and QuickDivergence, which enable efficient computation of VaR and a broad class of φ-divergence-based risk measures that include CVaR. By integrating the Quickselect algorithm with polyhedral optimization techniques, the proposed approach achieves expected linear time complexity for discrete distributions—the first such result in this setting. Empirical evaluations demonstrate that the new algorithms deliver orders-of-magnitude speedups on large-scale scenarios, substantially improving computational efficiency. Implementations of both methods are publicly available and integrated into the RiskMeasures.jl library.
Precise quantile computation in large-scale distributed environments faces fundamental challenges: it traditionally requires global sorting, incurs prohibitive communication overhead, and struggles to balance accuracy with efficiency. To address this, we propose GK Select—an algorithm that pioneers the use of the Greenwald-Khanna (GK) sketch for *exact* quantile computation. GK Select operates by extracting candidate values within the GK error bound, performing partition-wise linear scans, and applying a tree-based reduction—achieving exact results without full data shuffling in a constant number of communication rounds. Theoretically, its time complexity matches that of the GK sketch, while its space complexity is O(1/ε). Experiments on a 30-core AWS EMR cluster demonstrate that GK Select outperforms Spark’s built-in global-sorting approach by 10.5× in throughput and achieves latency comparable to approximate methods—thereby substantially overcoming the scalability bottleneck of exact quantile computation.
This study addresses the limitations of independent and identically distributed (iid) sampling in achieving adequate space-filling and coverage properties for multivariate distribution simulation. We extend quantile stratified sampling to multivariate normal and other multivariate distributions by introducing an ordered log-density evaluation mechanism, which effectively mitigates sample clustering in high-dimensional spaces. Experimental results demonstrate that the proposed method significantly outperforms traditional iid sampling in terms of spatial uniformity and coverage completeness. Consequently, this approach provides a more representative sampling strategy for efficient simulation of complex multivariate distributions, thereby enhancing the reliability of downstream statistical inference.
This paper addresses the slow convergence and high variance of traditional independent and identically distributed (IID) sampling in statistical simulation, particularly for skewed and heavy-tailed target distributions. We propose a one-dimensional quantile-based stratified sampling method: deterministic strata are constructed via exact quantiles of the target distribution, integrated with importance sampling and stratified design. We establish, for the first time, a rigorous theoretical framework for this approach and elucidate its variance-reduction mechanism—rooted in quantile mapping and within-stratum control variates. Both theoretical analysis and Monte Carlo experiments demonstrate that the method achieves 30–65% lower relative error than IID sampling across multiple non-regular test functions, with variance convergence rate of (O(n^{-3/2})). This represents a substantial improvement over the standard (O(n^{-1})) rate of IID sampling, especially in high-skewness and heavy-tailed regimes requiring high-precision simulation.
This paper investigates the asymptotic behavior of the comparison-cost residual ρₙ = Sₙ/n − S for QuickVal—a quantile-finding variant of quickselect—where Sₙ denotes the total comparison cost on a sample of size n and S is its limiting random variable. For general cost functions, we establish, for the first time, distributional convergence of √n ρₙ at the √n scaling, proving it converges in distribution to a scale mixture of a centered Gaussian variable, and further obtain Lᵖ (p ≥ 1) and moment convergence. In the unit-cost case (α = 0, i.e., QuickMin), we derive an exact closed-form expression for the L²-norm of ρₙ and its asymptotic equivalence. Our methodology integrates stochastic algorithm modeling, probabilistic analysis, Lᵖ and almost-sure convergence theory, and asymptotic distribution derivation. The key contributions are: (i) a novel standardized convergence framework for QuickVal residuals, and (ii) precise L²-characterization for the pivotal unit-cost special case.
This study addresses the limitation of existing portfolio theory, which is constrained by Simaan’s (1993) three-fund separation assumption and thus struggles to characterize more general weighted selection elliptical distributions. To overcome this, the work proposes stochastically representing weighted selection elliptical distributions as the sum of an affine combination and an independent directional elliptical component, while employing first-order stochastic dominance analysis techniques to construct a unified framework. The primary contribution lies in transcending the traditional three-fund restriction by rigorously deriving the first-order stochastic dominance separation conditions for q+2 funds. This result substantially broadens the applicability of fund separation theory, providing a more generalized theoretical foundation for portfolio optimization under complex distributional assumptions.
This study addresses the limitations of traditional risk measures—such as Value-at-Risk (VaR)—which rely on fixed confidence levels and thus fail to accommodate dynamic risk preferences. While λ-quantiles offer variable confidence levels, their computation suffers from discontinuities, multiple roots, and convergence difficulties. To overcome these challenges, this work proposes the Λ-Newton-Bis hybrid algorithm, which integrates Newton’s method with bisection to ensure global convergence while achieving local quadratic convergence. By incorporating interval analysis, the algorithm effectively handles discontinuities and multiple roots. The paper further introduces two novel solution strategies that, for the first time, efficiently embed λ-quantiles into portfolio optimization frameworks. Numerical experiments demonstrate that the proposed approach significantly outperforms existing methods in terms of convergence, robustness, and computational efficiency.
This study investigates how information content in order statistics evolves with increasing sample size, with a focus on their capacity to aggregate information in auction and voting settings. By leveraging Blackwell informativeness, hazard rate functions, and log-supermodularity, the paper establishes the first precise link between the informational precision of order statistics and the tail properties of the underlying distribution. The main contributions include proving that, except for the exponential distribution, enlarging the sample size almost always enhances the information conveyed by a fixed-rank order statistic; demonstrating that central order statistics become fully informative in large samples; and showing that intermediate ranks are advantageous under incomplete information. These results unify and extend existing theories of information aggregation in auctions and voting, offering a novel analytical framework grounded in order statistics.
Selecting the lag order of vector autoregressive (VAR) models remains challenging in small-sample or high-dimensional settings: AIC tends to overselect, while BIC and HQ require large samples for consistency. This paper proposes the Mean Squared Information Criterion (MIC), a novel lag-order selection method grounded in the theoretical stability of mean squared error (MSE) loss—specifically, its convergence to a plateau when the fitted order is at least as large as the true order. MIC achieves consistency under mild regularity conditions and exhibits robustness to small samples, high dimensionality, and model misspecification—addressing key limitations of classical criteria. Theoretical analysis and extensive simulations demonstrate that MIC consistently outperforms AIC, BIC, and HQ across diverse configurations. In empirical application, MIC significantly improves short-term forecasting accuracy for New York City’s COVID-19 case trajectories. An open-source R package, *micvar*, implements automated lag selection and forecasting, facilitating reproducible and scalable VAR modeling.
This work addresses the challenge of constructing real-time review queues for risk-scoring streams in financial crime investigations by proposing a label-free, adaptive thresholding mechanism. The approach leverages online adaptive kernel density estimation (KDE), dynamically satisfying queue capacity constraints through tail-mass curves and identifying stable thresholds via persistent density minima “snapshots” detected across multiple bandwidths. Integrated with sliding windows, exponential forgetting, and priority-based queue management, the system supports multi-queue routing and real-time processing. Experimental results demonstrate that the method strictly adheres to capacity limits across synthetic, concept-drifting, and multimodal data streams, significantly reduces threshold jitter, and achieves per-event update complexity of O(G) with constant memory usage.