Two-Level Softmax Sampling Done Right: Correcting Bias from Size Imbalance and Dispersion
This study addresses the systematic bias inherent in conventional two-level sampling methods for large-scale Softmax sampling, which arises from neglecting cluster size imbalance and dispersion heterogeneity. To mitigate this, we propose two correction algorithms, S-2LS and SD-2LS, that rigorously quantify and rectify these biases through probabilistic analysis, achieving unbiased sampling while preserving sublinear time complexity. This work provides the first theoretical elimination of size and dispersion biases in standard two-level sampling, yielding provably superior approximations with negligible computational overhead. Extensive experiments across five large-scale datasets demonstrate that the proposed methods substantially enhance the accuracy of Softmax distribution approximation.