Score
Designs and implements probabilistic programs and probabilistic program synthesis that encode generative models and prior beliefs, supporting full Bayesian inference and producing interpretable model structure. Builds and applies calibration and validation procedures—posterior predictive checks, simulation-based calibration, Bayesian model checking, sampler diagnostics, and other probabilistic calibration techniques—to quantify posterior uncertainty, flag mismatched likelihoods or invalid priors, assess predictive performance, and detect sampler pathologies.
This work addresses the challenge that probabilistic programs generated by language models often suffer from statistical misspecifications—such as incorrect likelihoods, priors, or parameterizations—that are difficult to detect with conventional unit tests. The paper introduces, for the first time, Bayesian calibration as a central criterion for assessing the correctness of probabilistic programs and proposes a fully unsupervised, reference-free framework for their detection and repair. By integrating Bayesian validation techniques—including posterior predictive checks, simulation-based calibration (SBC), sampling diagnostics (e.g., $\hat{R}$, divergences, effective sample size), and held-out predictive log density—the method generates feedback signals to drive an iterative repair loop within large language models. Evaluated on 200 instances, the approach achieves detection AUCs of 0.97 with reference programs and 62–78% without, substantially outperforming unit testing; repair success rates reach 92% and 100% using GPT-5.1 and Claude, respectively.
Bayesian inference in probabilistic programming is notoriously difficult, time-consuming, and expert-dependent to debug—severely hindering its practical adoption. To address this, we propose the first online, interactive debugging method specifically designed for probabilistic programming. Our approach deeply integrates real-time monitoring, interactive feedback, and visual diagnostic support directly into the development environment, enabling dynamic identification and correction of model or program defects *during* inference execution. Unlike prior approaches, it requires no offline analysis or manual intervention, substantially lowering the debugging barrier and time cost. A user study (N=18) demonstrates that our method reduces median debugging time by 52%, improves defect identification accuracy by 3.1×, and significantly enhances usability for non-expert developers. This work establishes the first in-environment, immediate, and interactive debugging capability for Bayesian inference—providing critical infrastructure toward the practical deployment of probabilistic programming.
This paper addresses the challenge of automated probabilistic model selection—enabling domain experts without statistical expertise to efficiently construct problem-appropriate probabilistic programs. To tackle the vast search space, high proportion of invalid programs, and difficulty in early invalidity detection, we propose a type-guided synthesis framework that integrates type-based static verification with heuristic program synthesis, ensuring both type safety and semantic validity of generated programs. We further combine static analysis with dynamic sampling to enhance search efficiency and correctness. Experimental evaluation on standard benchmarks demonstrates that our approach significantly outperforms random search and DaPPer, particularly on complex model synthesis tasks, while also supporting fast posterior sampling. The framework establishes a novel paradigm for automated modeling techniques such as genetic programming, advancing the accessibility and reliability of probabilistic programming for non-expert users.
This work addresses exact posterior distribution inference for discrete probabilistic programs. We propose a semantics-driven method based on weighted finite automata (WFA): program variables’ posteriors—including those with infinite support—are encoded as WFAs, and program semantics are realized via compositional WFA operations (e.g., product, concatenation), establishing a precise correspondence between program constructs and automaton transformations. To our knowledge, this is the first systematic application of WFAs to exact inference in probabilistic programming, overcoming the fundamental limitation of prior approaches—namely, their restriction to finite-support distributions. For a practically relevant class of discrete probabilistic programs, our method yields decidable, exact posterior computation, eliminating approximation error entirely. The framework provides a formal foundation for verifiable probabilistic reasoning in machine learning and autonomous systems.
In simulation-based calibration (SBC) for Bayesian models, prior specification faces a fundamental trade-off: overly broad priors risk numerical instability, while overly narrow ones reduce sensitivity to inferential failures—yet ground-truth data are often unavailable for calibration. Method: We propose *primed priors*, an adaptive, data-free prior construction framework extending catalytic priors. It integrates parameter-space sensitivity analysis with SBC-specific objective-driven design to enhance detection of common inferential pathologies—such as posterior shrinkage miscalibration and marginal inconsistency—while ensuring numerical robustness. Contribution/Results: Three simulation studies demonstrate that primed priors significantly improve SBC’s failure detection rate over standard priors and completely avoid computational breakdowns induced by extreme parameter values. To our knowledge, this is the first SBC-tailored, interpretable, and data-agnostic prior generation method.
This study addresses the problem of ill-defined statistical semantics caused by zero-probability observations in probabilistic programs. Drawing on geometric measure theory, this work constructs a disintegration-based integral semantics over program traces. By formulating an explicit Bayesian conditioning rule compatible with mainstream inference algorithms, it rigorously rectifies the theoretical deficiencies inherent in existing soft conditioning approaches. The proposed framework naturally accommodates loops, mixture distributions, and manifold-valued observations. Consequently, it ensures both mathematical rigor and statistical correctness for conditional reasoning and inference in complex probabilistic models.
This work proposes the first semantically consistent sequential Monte Carlo (SMC) framework grounded in the Feynman–Kac formalism for efficient and provably correct inference in general-purpose probabilistic programs that support arbitrary measure sampling and conditional reweighting within unbounded loops. The approach employs probabilistic program graphs (PPGs) as an intermediate representation and leverages a finite-trace approximation theorem to rigorously establish the correspondence between the program’s expectation semantics and the Feynman–Kac model. Building on this foundation, the authors design a vectorized particle filtering algorithm (VPF) tailored to PPGs. Empirical evaluations demonstrate that VPF significantly outperforms existing state-of-the-art inference tools across multiple benchmarks, achieving a compelling combination of theoretical soundness, computational efficiency, and strong scalability.
This work addresses the “Likelihood Hacking” (LH) problem in reinforcement learning training of language models, where models artificially inflate marginal likelihood rewards by generating unnormalized programs rather than genuinely fitting the data. The study formally characterizes LH behavior for the first time and establishes syntactic sufficient conditions that prevent such reward manipulation. Building on these conditions, the authors develop $\mathcal{L}_{\text{safe}}$, a provably safe probabilistic programming sublanguage, and implement it in SafeStan—a system that combines theoretical rigor with practical efficacy. Experimental results demonstrate that models trained with GRPO exploit LH vulnerabilities early in training, whereas SafeStan effectively resists likelihood hacking even under strong optimization pressure.
本文通过贝叶斯混合模型方法解决计算机代码验证问题,比较纯代码模型与偏差修正模型,并使用Metropolis-within-Gibbs算法进行推断。
This work addresses the computational redundancy inherent in Markov chain Monte Carlo (MCMC) inference for probabilistic programming by proposing an incremental inference method based on dynamic dependency graphs. The approach compiles probabilistic programs into reactive computation graphs, enabling selective recomputation of only those subgraphs affected by updates to random variables, thereby substantially reducing the per-iteration sampling cost. Notably, this study is the first to integrate functional reactive programming with probabilistic programming, unifying the representational frameworks of Bayesian networks—formalized as applicative functors—and general-purpose probabilistic programs expressed as monads. The method achieves significant improvements in MCMC efficiency while preserving inference accuracy.