Institution profile

Spotify

Industry researcheurope · se
Official website
Research library66linked papers
Opportunities0open roles
Selected work

Representative Papers

$σ$Transfer: Uncertainty Transfer from Small to Large Networks under $μ\mathrm{P}$

Oct 08, 2026

This study addresses the prohibitive computational cost of selecting prior precision in Laplace approximations for large neural networks. By deriving a prior covariance rescaling mechanism grounded in maximal update parameterization ($\mu$P), this work proposes $\sigma$Transfer, a hyperparameter-free method that zero-shot transfers prior precision from small to large models, with theoretical guarantees that posterior stability converges as network width increases. Empirically, $\sigma$Transfer achieves an approximately 5000-fold speedup on MNIST with only a 0.002 increase in negative log-likelihood (NLL). Furthermore, it enables efficient precision transfer across models scaling from 1B to 7B parameters, yielding NLL degradation below $10^{-4}$. These results establish a scalable, low-cost paradigm for uncertainty estimation in large-scale models.

0 citationsRead paper

When Can You Ship on Evals Alone? Trial-Level Surrogacy for Offline Evaluation of LLM Systems

Oct 07, 2026

This study addresses the limitation that offline evaluation cannot determine whether modifications to large language models are suitable for direct deployment, while also failing to distinguish between trial-level and unit-level surrogacy. To this end, this work formally establishes the logical independence between trial-level and Prentice surrogacy for the first time. It derives closed-form decision boundaries that correct for noise under a bivariate normal model, incorporates variance-subtraction covariance recovery techniques, and analyzes risks associated with judge bias and distributional drift. Empirical validation on the Upworthy dataset demonstrates a weak correlation between offline evaluations and online outcomes, confirming that offline metrics alone are insufficient to replace A/B testing in practice.

0 citationsRead paper

Two-Level Softmax Sampling Done Right: Correcting Bias from Size Imbalance and Dispersion

Oct 07, 2026

This study addresses the systematic bias inherent in conventional two-level sampling methods for large-scale Softmax sampling, which arises from neglecting cluster size imbalance and dispersion heterogeneity. To mitigate this, we propose two correction algorithms, S-2LS and SD-2LS, that rigorously quantify and rectify these biases through probabilistic analysis, achieving unbiased sampling while preserving sublinear time complexity. This work provides the first theoretical elimination of size and dispersion biases in standard two-level sampling, yielding provably superior approximations with negligible computational overhead. Extensive experiments across five large-scale datasets demonstrate that the proposed methods substantially enhance the accuracy of Softmax distribution approximation.

0 citationsRead paper

Variance-Optimal Off-Policy Evaluation with Conjunct Effect Modeling

Oct 06, 2026

This study addresses the problem of excessive action-level weight variance in off-policy evaluation for contextual multi-armed bandits by proposing the VOCEM estimator. This method constructs an unbiased interpolation family between the OffCEM and doubly robust estimators, deriving a closed-form optimal interpolation coefficient to minimize estimation variance while theoretically guaranteeing that the resulting variance is strictly lower than that of both endpoint estimators. By integrating joint effect modeling with variance optimization theory, VOCEM achieves significant improvements over existing baselines across 23 experimental configurations on both synthetic and real-world benchmarks. These results demonstrate that the proposed approach effectively enhances the stability and robustness of off-policy evaluation.

0 citationsRead paper

Learning to Re-Draft: A Variational Stackelberg Game for Discrete Diffusion

Sep 28, 2026

This study addresses the challenge of optimizing semantic correction and redrafting capabilities in discrete diffusion models, which is hindered by fixed or stochastic corruption processes. We propose the VSDD framework, formulating training as a Stackelberg game wherein the denoiser acts as the follower and the corruption strategy as the leader, thereby learning a semantics-aware corruption process driven by performance gains rather than reconstruction difficulty. By integrating variational inference with score function estimation, our method efficiently solves for Markovian corruption processes via single-step gradient approximation. Experimental results demonstrate that the proposed framework substantially improves molecular generation validity, reduces text perplexity, and enhances offline recommendation metrics.

0 citationsRead paper
Recent publications

Latest Papers

$σ$Transfer: Uncertainty Transfer from Small to Large Networks under $μ\mathrm{P}$

Oct 08, 2026

This study addresses the prohibitive computational cost of selecting prior precision in Laplace approximations for large neural networks. By deriving a prior covariance rescaling mechanism grounded in maximal update parameterization ($\mu$P), this work proposes $\sigma$Transfer, a hyperparameter-free method that zero-shot transfers prior precision from small to large models, with theoretical guarantees that posterior stability converges as network width increases. Empirically, $\sigma$Transfer achieves an approximately 5000-fold speedup on MNIST with only a 0.002 increase in negative log-likelihood (NLL). Furthermore, it enables efficient precision transfer across models scaling from 1B to 7B parameters, yielding NLL degradation below $10^{-4}$. These results establish a scalable, low-cost paradigm for uncertainty estimation in large-scale models.

0 citationsRead paper

When Can You Ship on Evals Alone? Trial-Level Surrogacy for Offline Evaluation of LLM Systems

Oct 07, 2026

This study addresses the limitation that offline evaluation cannot determine whether modifications to large language models are suitable for direct deployment, while also failing to distinguish between trial-level and unit-level surrogacy. To this end, this work formally establishes the logical independence between trial-level and Prentice surrogacy for the first time. It derives closed-form decision boundaries that correct for noise under a bivariate normal model, incorporates variance-subtraction covariance recovery techniques, and analyzes risks associated with judge bias and distributional drift. Empirical validation on the Upworthy dataset demonstrates a weak correlation between offline evaluations and online outcomes, confirming that offline metrics alone are insufficient to replace A/B testing in practice.

0 citationsRead paper

Two-Level Softmax Sampling Done Right: Correcting Bias from Size Imbalance and Dispersion

Oct 07, 2026

This study addresses the systematic bias inherent in conventional two-level sampling methods for large-scale Softmax sampling, which arises from neglecting cluster size imbalance and dispersion heterogeneity. To mitigate this, we propose two correction algorithms, S-2LS and SD-2LS, that rigorously quantify and rectify these biases through probabilistic analysis, achieving unbiased sampling while preserving sublinear time complexity. This work provides the first theoretical elimination of size and dispersion biases in standard two-level sampling, yielding provably superior approximations with negligible computational overhead. Extensive experiments across five large-scale datasets demonstrate that the proposed methods substantially enhance the accuracy of Softmax distribution approximation.

0 citationsRead paper

Variance-Optimal Off-Policy Evaluation with Conjunct Effect Modeling

Oct 06, 2026

This study addresses the problem of excessive action-level weight variance in off-policy evaluation for contextual multi-armed bandits by proposing the VOCEM estimator. This method constructs an unbiased interpolation family between the OffCEM and doubly robust estimators, deriving a closed-form optimal interpolation coefficient to minimize estimation variance while theoretically guaranteeing that the resulting variance is strictly lower than that of both endpoint estimators. By integrating joint effect modeling with variance optimization theory, VOCEM achieves significant improvements over existing baselines across 23 experimental configurations on both synthetic and real-world benchmarks. These results demonstrate that the proposed approach effectively enhances the stability and robustness of off-policy evaluation.

0 citationsRead paper

Learning to Re-Draft: A Variational Stackelberg Game for Discrete Diffusion

Sep 28, 2026

This study addresses the challenge of optimizing semantic correction and redrafting capabilities in discrete diffusion models, which is hindered by fixed or stochastic corruption processes. We propose the VSDD framework, formulating training as a Stackelberg game wherein the denoiser acts as the follower and the corruption strategy as the leader, thereby learning a semantics-aware corruption process driven by performance gains rather than reconstruction difficulty. By integrating variational inference with score function estimation, our method efficiently solves for Markovian corruption processes via single-step gradient approximation. Experimental results demonstrate that the proposed framework substantially improves molecular generation validity, reduces text perplexity, and enhances offline recommendation metrics.

0 citationsRead paper