Institution profile

Ecole d’Ingénieurs SJTU-ParisTech

Academic institutioneurope · fr
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Two-Level Softmax Sampling Done Right: Correcting Bias from Size Imbalance and Dispersion

Oct 07, 2026

This study addresses the systematic bias inherent in conventional two-level sampling methods for large-scale Softmax sampling, which arises from neglecting cluster size imbalance and dispersion heterogeneity. To mitigate this, we propose two correction algorithms, S-2LS and SD-2LS, that rigorously quantify and rectify these biases through probabilistic analysis, achieving unbiased sampling while preserving sublinear time complexity. This work provides the first theoretical elimination of size and dispersion biases in standard two-level sampling, yielding provably superior approximations with negligible computational overhead. Extensive experiments across five large-scale datasets demonstrate that the proposed methods substantially enhance the accuracy of Softmax distribution approximation.

0 citationsRead paper

Privacy in Personalized AI Is a System Property, Not Just a Model Property

Sep 29, 2026

This project addresses the limitation of single-model analyses in capturing system-level privacy risks within personalized AI by departing from traditional component-level paradigms to establish privacy as an emergent system property. Through systematic architectural analysis and privacy risk modeling, it identifies four distinct risk channels and constructs a multidimensional evaluation framework encompassing interaction trajectories, information flows, indirect leakage, and utility-privacy trade-offs. Ultimately, this work delivers actionable, systematized privacy auditing standards and an assessment methodology that effectively bridges critical gaps in existing audit approaches at the system level.

0 citationsRead paper

Music Playlist Captioning at Scale with Large Language Models

Jun 21, 2026

This work addresses the lack of interpretability in personalized playlists on music streaming platforms by deploying, for the first time at industrial scale, a large language model (LLM)-based automatic captioning system within Deezer’s Daily Mix recommendation service. The proposed approach integrates multi-source heterogeneous data and leverages a controllable generation mechanism to produce semantically rich and personalized natural language descriptions. Following deployment, the system yielded significant gains in user engagement, demonstrating that semantic explanations play a pivotal role in enhancing both the interpretability of recommendations and overall user experience. This study establishes an effective paradigm for controllable text generation with LLMs in real-world applications, offering practical insights into bridging the gap between algorithmic personalization and human-understandable justifications.

0 citationsRead paper
Recent publications

Latest Papers

Two-Level Softmax Sampling Done Right: Correcting Bias from Size Imbalance and Dispersion

Oct 07, 2026

This study addresses the systematic bias inherent in conventional two-level sampling methods for large-scale Softmax sampling, which arises from neglecting cluster size imbalance and dispersion heterogeneity. To mitigate this, we propose two correction algorithms, S-2LS and SD-2LS, that rigorously quantify and rectify these biases through probabilistic analysis, achieving unbiased sampling while preserving sublinear time complexity. This work provides the first theoretical elimination of size and dispersion biases in standard two-level sampling, yielding provably superior approximations with negligible computational overhead. Extensive experiments across five large-scale datasets demonstrate that the proposed methods substantially enhance the accuracy of Softmax distribution approximation.

0 citationsRead paper

Privacy in Personalized AI Is a System Property, Not Just a Model Property

Sep 29, 2026

This project addresses the limitation of single-model analyses in capturing system-level privacy risks within personalized AI by departing from traditional component-level paradigms to establish privacy as an emergent system property. Through systematic architectural analysis and privacy risk modeling, it identifies four distinct risk channels and constructs a multidimensional evaluation framework encompassing interaction trajectories, information flows, indirect leakage, and utility-privacy trade-offs. Ultimately, this work delivers actionable, systematized privacy auditing standards and an assessment methodology that effectively bridges critical gaps in existing audit approaches at the system level.

0 citationsRead paper

Music Playlist Captioning at Scale with Large Language Models

Jun 21, 2026

This work addresses the lack of interpretability in personalized playlists on music streaming platforms by deploying, for the first time at industrial scale, a large language model (LLM)-based automatic captioning system within Deezer’s Daily Mix recommendation service. The proposed approach integrates multi-source heterogeneous data and leverages a controllable generation mechanism to produce semantically rich and personalized natural language descriptions. Following deployment, the system yielded significant gains in user engagement, demonstrating that semantic explanations play a pivotal role in enhancing both the interpretability of recommendations and overall user experience. This study establishes an effective paradigm for controllable text generation with LLMs in real-world applications, offering practical insights into bridging the gap between algorithmic personalization and human-understandable justifications.

0 citationsRead paper