Score
Designs and implements compilers and program transformations that take probabilistic program representations and produce explicit, deterministic probability density functions or density-evaluating code while preserving the original program’s probabilistic semantics. Builds passes and emitters that ensure the resulting densities are correct, numerically well-formed, and suitable for downstream inference (e.g., differentiable, incrementalizable, or compatible with nonparametric model techniques).
This work addresses the scalability challenges of probabilistic program inference on large-scale data, which often suffers from expensive recomputation of intermediate results. The authors propose a modular approach that compiles probabilistic programs into deterministic density-function programs and leverages compositional incremental optimization based on incremental λ-calculus, enabling Monte Carlo inference to efficiently reuse intermediate computations. By decoupling density evaluation from incremental optimization, the method supports nonparametric models and employs denotational logical relations to verify correctness in a stepwise manner. Experimental results demonstrate that the system achieves asymptotic speedups with respect to data size across a range of models and inference algorithms, with a Julia prototype confirming its practical efficacy.
This paper addresses exact Bayesian inference for discrete probabilistic programs. We propose a semantics modeling framework based on weighted automata. Our method formally maps program statements—including `observe` instructions and branching/looping control flow—to weighted automata over exchangeable alphabets, encoding prior distributions as weighted automata over natural-number vectors, and performing symbolic posterior computation via automata operations. The approach guarantees fully exact, non-approximate posterior inference for a class of discrete probabilistic programs with rich control structures. We prove its correctness with respect to standard operational semantics and provide theoretical foundations via probability-generating functions. Experimental evaluation demonstrates substantial improvements in both symbolic reasoning capability and expressive power compared to existing methods.
This work addresses probabilistic programs with user-annotated sampling statements and while loops (e.g., Gen, Turing, Pyro), where random variables may be generated dynamically. We present the first static factorization method supporting such dynamic variable generation. Our approach extends operational semantics, constructs a probability-annotated control-flow graph, and integrates static dependency analysis with program slicing to yield an exact graphical model of the implicit program density—formally equivalent to a verifiable Bayesian network. The method guarantees theoretical correctness and overcomes the longstanding limitation of traditional graphical models, which require statically fixed variable structures. Empirical evaluation demonstrates that our representation substantially reduces gradient estimation variance and accelerates convergence in single-site Metropolis–Hastings and sequential Monte Carlo inference, achieving performance competitive with or superior to state-of-the-art techniques.
This paper addresses the challenge of verifying trustworthiness in probabilistic programs. Method: We propose the first computational framework that formally defines “trust” as statistical consistency between observed output frequencies and a target probability distribution. Our approach introduces an extended typed λ-calculus featuring runtime experiment operators and a static confidence-type system, enabling joint verification of dynamic sampling and distributional compliance. We establish its computability and adherence to Kolmogorov’s axioms through rigorous probabilistic semantics, along with formal proofs of progress and termination. Contribution/Results: This work achieves the first statically verifiable notion of probabilistic consistency, unifying semantic correctness with statistical reliability guarantees. It provides a theoretically sound and practically implementable foundation for trust verification in probabilistic computation.
This work addresses the high GPU computational cost incurred by large language models when repeatedly sampling programs for code generation and mathematical reasoning. The authors propose a novel test-time framework that explicitly models the token-level output probabilities of the model as a probabilistic program, compactly representing an exponential number of deterministic programs in a structured form. By leveraging lightweight CPU-based probabilistic inference, the method enables efficient sampling without requiring additional calls to the large language model. This approach achieves significant performance improvements across benchmarks in code generation, code understanding, and mathematical reasoning while substantially reducing computational overhead.
This work addresses the challenge of implementing reverse-mode automatic differentiation for programs featuring algebraic effects such as finite discrete probabilistic choice. Building upon the Compositional Homomorphic Automatic Differentiation (CHAD) framework, it formulates automatic differentiation as a semantics-preserving program transformation and constructs a backward-pass mechanism tailored to the finite atomic distribution monad. The correctness of this construction is established using logical relations from category theory. This study presents the first systematic extension of reverse-mode automatic differentiation to effectful languages with discrete outputs, introducing a reusable differentiation scheme applicable to a broad class of algebraic effects—including nondeterminism, exceptions, and writer effects. The approach not only enables correct reverse differentiation of programs with finite discrete probabilistic structure but also lays a foundational theoretical groundwork for differentiating more general effectful languages.
This work proposes the first semantically consistent sequential Monte Carlo (SMC) framework grounded in the Feynman–Kac formalism for efficient and provably correct inference in general-purpose probabilistic programs that support arbitrary measure sampling and conditional reweighting within unbounded loops. The approach employs probabilistic program graphs (PPGs) as an intermediate representation and leverages a finite-trace approximation theorem to rigorously establish the correspondence between the program’s expectation semantics and the Feynman–Kac model. Building on this foundation, the authors design a vectorized particle filtering algorithm (VPF) tailored to PPGs. Empirical evaluations demonstrate that VPF significantly outperforms existing state-of-the-art inference tools across multiple benchmarks, achieving a compelling combination of theoretical soundness, computational efficiency, and strong scalability.
This study addresses the problem of accurately computing the output distributions of small-scale programs—such as those processing GPS or inertial sensor data—that involve random inputs. To this end, it introduces cylindrical algebraic decomposition (CAD) into probabilistic program analysis for the first time, combining symbolic and numerical integration to effectively handle conditional branches and nonlinear operations. The approach is grounded in a rigorous semantic model of probabilistic programs and has been validated on both floating-point arithmetic benchmarks and representative programs from open-source sensor libraries, demonstrating its feasibility and effectiveness in deriving exact output distributions.
This work addresses the loss of dimensional semantics in traditional compilation, where type systems discard such information prior to code generation, leading to ad hoc numerical representations and memory management that struggle to balance efficiency, determinism, and verifiability. To overcome this, the paper introduces a Dimensional Type System (DTS) that propagates dimensional annotations as compile-time metadata throughout MLIR’s multi-stage lowering pipeline. DTS enables joint optimization of representation selection and deterministic memory management within a unified semantic graph. Grounded in finitely generated Abelian group constraints, the system supports polynomial-time decidable, complete, and principal type inference. A coeffect system unifies escape analysis and memory allocation, while also revealing the closure of dimensional algebra under automatic differentiation. Experiments demonstrate that DTS enables design-time verifiable memory strategies, representation fidelity, cache locality estimation, and coeffect-based AD verification, significantly enhancing compilation reliability and performance in resource-constrained settings.