Score
Designs and implements probabilistic sequence or variable models and inference procedures that can compute and sample conditional distributions for arbitrary subsets of tokens or variables given any other subset, enabling single-pass evaluation of arbitrary conditional likelihoods. Builds model architectures and decoding algorithms (e.g., AC-GPT–style integrations with causal transformers) that provide efficient autoregressive decoding and flexible sampling from past, future, or mixed contexts.
This paper addresses the challenge of modeling uncertainty by systematically establishing a pedagogical and theoretical framework for probabilistic graphical models (PGMs). To tackle the intractability of representing and reasoning over high-dimensional joint distributions, it unifies directed graphs (Bayesian networks) and undirected graphs (Markov random fields) to compactly encode variable dependencies, integrating probability theory with graph theory. The work develops a comprehensive methodology encompassing parameter learning, structure learning, and exact/approximate inference—including variable elimination, belief propagation, and variational inference. Its primary contribution is a tripartite PGM pedagogical paradigm—representation, learning, and inference—that rigorously aligns graph structure with probabilistic semantics. Through algorithmic design and concrete case studies, the framework enhances model interpretability and practical utility in prediction and decision-making tasks, thereby providing foundational support for uncertainty reasoning in machine learning and AI.
Standard causal Transformers struggle to efficiently model and sample arbitrary conditional distributions, such as text that depends simultaneously on both past and future context. This work proposes a lightweight architectural modification that enables flexible conditioning—including mixed temporal contexts—within a single forward pass, while preserving the standard left-to-right training objective and autoregressive decoding mechanism. The approach seamlessly integrates into existing large language model training pipelines, allowing direct fine-tuning without disrupting conventional autoregressive generation. Empirical results demonstrate that the method significantly outperforms current baselines on tasks requiring arbitrary conditional modeling, all while maintaining competitive performance in standard autoregressive text generation.
This work proposes a probabilistic inference framework that integrates inductive biases to address the challenges of uncertainty quantification in deep sequential models. While traditional Bayesian approaches struggle with prior specification and inference accuracy in large-scale networks, the proposed method establishes a theoretical connection between Transformer attention mechanisms and sparse Gaussian processes, enabling scalable approximate Bayesian inference. It introduces cross-domain inducing points derived from HiPPO operators to support long-range historical modeling in online learning settings. Furthermore, self-supervised signals are leveraged to enrich the probabilistic structure of latent variables in sequence generation. The resulting approach significantly enhances the uncertainty quantification capability, probabilistic expressiveness, and scalability of deep sequential models, all while maintaining competitive predictive performance.
This work addresses exact posterior distribution inference for discrete probabilistic programs. We propose a semantics-driven method based on weighted finite automata (WFA): program variables’ posteriors—including those with infinite support—are encoded as WFAs, and program semantics are realized via compositional WFA operations (e.g., product, concatenation), establishing a precise correspondence between program constructs and automaton transformations. To our knowledge, this is the first systematic application of WFAs to exact inference in probabilistic programming, overcoming the fundamental limitation of prior approaches—namely, their restriction to finite-support distributions. For a practically relevant class of discrete probabilistic programs, our method yields decidable, exact posterior computation, eliminating approximation error entirely. The framework provides a formal foundation for verifiable probabilistic reasoning in machine learning and autonomous systems.
This paper addresses two critical limitations in exchangeable sequence modeling: (i) the inability to disentangle epistemic from aleatoric uncertainty, and (ii) the lack of theoretical guarantees—particularly strict exchangeability—in existing architectures. We systematically analyze how inference mechanisms and structural inductive biases affect posterior uncertainty quantification. We show that standard single-step autoregressive modeling conflates uncertainty types, and current exchangeable Transformers violate strict permutation invariance. To resolve these issues, we propose a multi-step autoregressive generative framework grounded in Bayesian posterior inference and causal masking analysis, along with a novel architectural design principle ensuring provable exchangeability. Through rigorous theoretical analysis and controlled synthetic experiments, we demonstrate that our architecture achieves significantly improved uncertainty calibration. It consistently outperforms baselines on downstream decision-making tasks—including active learning and contextual bandits—thereby exposing structural inefficiencies and redundant computation inherent in prevailing models.
To address the challenge of balancing autoregressive generation efficiency with flexible set-conditional modeling in Transformer-based probabilistic models for joint distribution forecasting, this paper proposes Causal Autoregressive Caching (CAC). CAC decouples context encoding from conditional set updates via a dynamic cache buffer, enabling single-shot conditional encoding, multi-step cache reuse, and batched autoregressive decoding—thereby unifying efficient joint log-likelihood evaluation and joint set-conditional/autoregressive training. Innovatively, it introduces cache-context attention and target-dependent dynamic modeling, achieving substantial inference speedups without compromising conditional modeling flexibility—a first in the literature. Experiments across synthetic functions, EEG time series, cognitive modeling tasks, and tabular data demonstrate that CAC matches state-of-the-art baseline accuracy while accelerating joint sampling by up to 20×.
This work addresses the challenge of extracting structured knowledge from large language models (LLMs) in a probabilistically coherent manner to support complex variable reasoning. It proposes Large Language Gibbs (LLG), a method that embeds the LLM’s conditional distribution as a transition kernel within a Gibbs sampling framework. By iteratively resampling individual variables rather than generating sequences autoregressively, LLG avoids biases induced by sequential dependencies. This approach constitutes the first effective integration of LLMs with Markov chain Monte Carlo (MCMC), yielding a stationary distribution consistent with all local conditional distributions. Empirical results demonstrate that LLG excels in tasks including synthetic distribution sampling, consistency-aware reasoning, and Bayesian structure learning, establishing a new paradigm for structured probabilistic inference beyond single-pass generation.
This work proposes the first semantically consistent sequential Monte Carlo (SMC) framework grounded in the Feynman–Kac formalism for efficient and provably correct inference in general-purpose probabilistic programs that support arbitrary measure sampling and conditional reweighting within unbounded loops. The approach employs probabilistic program graphs (PPGs) as an intermediate representation and leverages a finite-trace approximation theorem to rigorously establish the correspondence between the program’s expectation semantics and the Feynman–Kac model. Building on this foundation, the authors design a vectorized particle filtering algorithm (VPF) tailored to PPGs. Empirical evaluations demonstrate that VPF significantly outperforms existing state-of-the-art inference tools across multiple benchmarks, achieving a compelling combination of theoretical soundness, computational efficiency, and strong scalability.
This work addresses the limitations of traditional autoregressive models, which rely solely on next-token prediction and struggle to capture sequence-level properties, often resulting in local overfitting and poor global structure modeling. Moreover, controllable generation typically demands costly sampling or architectural modifications. To overcome these challenges, the authors propose the Conditional Attribute Transformer, which jointly optimizes next-token prediction and sequence attribute estimation conditioned on the current token within a single forward pass. This approach uniquely enables token-level attribute attribution, counterfactual attribute evaluation, and attribute-guided generation—all within a unified model without additional sampling or structural changes. Experiments demonstrate state-of-the-art performance on sparse-reward tasks, improved language modeling at large scales, attribute estimation orders of magnitude faster than sampling-based methods, and effective guidance across diverse language generation tasks.
Traditional statistical inference often fails when models or parameters are selected in a data-driven manner. This work systematically investigates selective inference, focusing on a conditional inference framework that conditions on the selection event to restore the nominal coverage of confidence intervals. We clarify the scientific interpretation of this framework, unify several existing approaches under its umbrella, and demonstrate its application to canonical settings such as inference for the “winner,” region-specific means in regression trees, and differences between clusters. Through simulations and analyses of single-cell RNA sequencing data, we show that the proposed methodology yields valid and reliable statistical inference in practical scenarios.