Institution profile

Université Paris Nanterre

Academic institutioneurope · fr
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Langevin-Informed Transfer Learning: Replacing Target Samples by Black-Box Feedback

Oct 01, 2026

This study addresses the challenge of extracting low-dimensional slow dynamics in stochastic dynamical systems where target trajectories are inaccessible and only biased samples are available. To this end, we propose the LITL framework, which relies solely on black-box feedback. By leveraging Dirichlet representation learning, the method captures the spectral structure and projected drift of the target infinitesimal generator, while introducing spherical variants to optimize the guided control of normalized latent representations, thereby achieving Langevin dynamics reconstruction via spectral operator learning. Experimental results demonstrate that the proposed approach successfully recovers physical timescales and spherical symmetry from biased simulations, establishes a dynamical structure for generative models, and enables post-hoc guided control within the latent space.

0 citationsRead paper

HyperLogLog for probabilists

Jul 24, 2026

This study addresses the problem of efficiently approximating the number of distinct elements in large-scale data streams. Revisiting the HyperLogLog algorithm from a probabilistic perspective, it establishes—for the first time—the non-asymptotic, explicit exponential concentration inequality for the estimator using elementary probability methods. This work breaks away from the traditional asymptotic analysis framework that relies on Poissonization and Mellin transforms, instead deriving rigorous finite-sample error bounds. The resulting bounds significantly enhance the theoretical reliability and error controllability of HyperLogLog in practical applications.

0 citationsRead paper

On the scaling relationship between cloze probabilities and language model next-token prediction

Feb 19, 2026

This study investigates how the scale of language models influences their ability to predict human cloze behavior, including eye movements and reading times. By evaluating next-word prediction performance across models of varying sizes against human cloze responses and cognitive metrics, the research finds that although larger models still systematically underestimate the probability of actual human responses, they significantly improve semantic alignment with human predictions due to enhanced memory and semantic modeling capabilities. Consequently, these models rely less on low-level lexical co-occurrence statistics. The work demonstrates that increasing model scale positively enhances cognitive modeling capacity, clarifying both the advantages and limitations of current large language models in predicting human language comprehension.

0 citationsRead paper
Recent publications

Latest Papers

Langevin-Informed Transfer Learning: Replacing Target Samples by Black-Box Feedback

Oct 01, 2026

This study addresses the challenge of extracting low-dimensional slow dynamics in stochastic dynamical systems where target trajectories are inaccessible and only biased samples are available. To this end, we propose the LITL framework, which relies solely on black-box feedback. By leveraging Dirichlet representation learning, the method captures the spectral structure and projected drift of the target infinitesimal generator, while introducing spherical variants to optimize the guided control of normalized latent representations, thereby achieving Langevin dynamics reconstruction via spectral operator learning. Experimental results demonstrate that the proposed approach successfully recovers physical timescales and spherical symmetry from biased simulations, establishes a dynamical structure for generative models, and enables post-hoc guided control within the latent space.

0 citationsRead paper

HyperLogLog for probabilists

Jul 24, 2026

This study addresses the problem of efficiently approximating the number of distinct elements in large-scale data streams. Revisiting the HyperLogLog algorithm from a probabilistic perspective, it establishes—for the first time—the non-asymptotic, explicit exponential concentration inequality for the estimator using elementary probability methods. This work breaks away from the traditional asymptotic analysis framework that relies on Poissonization and Mellin transforms, instead deriving rigorous finite-sample error bounds. The resulting bounds significantly enhance the theoretical reliability and error controllability of HyperLogLog in practical applications.

0 citationsRead paper

On the scaling relationship between cloze probabilities and language model next-token prediction

Feb 19, 2026

This study investigates how the scale of language models influences their ability to predict human cloze behavior, including eye movements and reading times. By evaluating next-word prediction performance across models of varying sizes against human cloze responses and cognitive metrics, the research finds that although larger models still systematically underestimate the probability of actual human responses, they significantly improve semantic alignment with human predictions due to enhanced memory and semantic modeling capabilities. Consequently, these models rely less on low-level lexical co-occurrence statistics. The work demonstrates that increasing model scale positively enhances cognitive modeling capacity, clarifying both the advantages and limitations of current large language models in predicting human language comprehension.

0 citationsRead paper