arbitrary conditional modeling

Designs and implements probabilistic sequence or variable models and inference procedures that can compute and sample conditional distributions for arbitrary subsets of tokens or variables given any other subset, enabling single-pass evaluation of arbitrary conditional likelihoods. Builds model architectures and decoding algorithms (e.g., AC-GPT–style integrations with causal transformers) that provide efficient autoregressive decoding and flexible sampling from past, future, or mixed contexts.

arbitraryconditionalmodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.25
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Standard causal Transformers struggle to efficiently model and sample arbitrary conditional distributions, such as text that depends simultaneously on both past and future context. This work proposes a lightweight architectural modification that enables flexible conditioning—including mixed temporal contexts—within a single forward pass, while preserving the standard left-to-right training objective and autoregressive decoding mechanism. The approach seamlessly integrates into existing large language model training pipelines, allowing direct fine-tuning without disrupting conventional autoregressive generation. Empirical results demonstrate that the method significantly outperforms current baselines on tasks requiring arbitrary conditional modeling, all while maintaining competitive performance in standard autoregressive text generation.

arbitrary conditionalscausal Transformersconditional sampling

This work proposes a probabilistic inference framework that integrates inductive biases to address the challenges of uncertainty quantification in deep sequential models. While traditional Bayesian approaches struggle with prior specification and inference accuracy in large-scale networks, the proposed method establishes a theoretical connection between Transformer attention mechanisms and sparse Gaussian processes, enabling scalable approximate Bayesian inference. It introduces cross-domain inducing points derived from HiPPO operators to support long-range historical modeling in online learning settings. Furthermore, self-supervised signals are leveraged to enrich the probabilistic structure of latent variables in sequence generation. The resulting approach significantly enhances the uncertainty quantification capability, probabilistic expressiveness, and scalability of deep sequential models, all while maintaining competitive predictive performance.

approximate inferenceBayesian inferencedeep sequence models

Probabilistic Programming Meets Automata Theory: Exact Inference using Weighted Automata

Dec 15, 2025
DG
Dominik Geißler
🏛️ TU Berlin | RWTH Aachen University

This work addresses exact posterior distribution inference for discrete probabilistic programs. We propose a semantics-driven method based on weighted finite automata (WFA): program variables’ posteriors—including those with infinite support—are encoded as WFAs, and program semantics are realized via compositional WFA operations (e.g., product, concatenation), establishing a precise correspondence between program constructs and automaton transformations. To our knowledge, this is the first systematic application of WFAs to exact inference in probabilistic programming, overcoming the fundamental limitation of prior approaches—namely, their restriction to finite-support distributions. For a practically relevant class of discrete probabilistic programs, our method yields decidable, exact posterior computation, eliminating approximation error entirely. The framework provides a formal foundation for verifiable probabilistic reasoning in machine learning and autonomous systems.

Encoding infinite-support distributions via automata-theoretic constructionsExact inference of posterior distributions in probabilistic programsUsing weighted automata for analyzing discrete probabilistic programs

Architectural and Inferential Inductive Biases For Exchangeable Sequence Modeling

Mar 03, 2025
DM
Daksh Mittal
🏛️ Columbia Business School

This paper addresses two critical limitations in exchangeable sequence modeling: (i) the inability to disentangle epistemic from aleatoric uncertainty, and (ii) the lack of theoretical guarantees—particularly strict exchangeability—in existing architectures. We systematically analyze how inference mechanisms and structural inductive biases affect posterior uncertainty quantification. We show that standard single-step autoregressive modeling conflates uncertainty types, and current exchangeable Transformers violate strict permutation invariance. To resolve these issues, we propose a multi-step autoregressive generative framework grounded in Bayesian posterior inference and causal masking analysis, along with a novel architectural design principle ensuring provable exchangeability. Through rigorous theoretical analysis and controlled synthetic experiments, we demonstrate that our architecture achieves significantly improved uncertainty calibration. It consistently outperforms baselines on downstream decision-making tasks—including active learning and contextual bandits—thereby exposing structural inefficiencies and redundant computation inherent in prevailing models.

Addresses limitations in modeling exchangeable sequences with autoregressive models.Explores effective inferential and architectural biases for uncertainty quantification.Identifies gaps in Transformer architectures for ensuring sequence exchangeability.

Efficient Autoregressive Inference for Transformer Probabilistic Models

Oct 10, 2025
CH
Conor Hassan
🏛️ Aalto University | University of Helsinki

To address the challenge of balancing autoregressive generation efficiency with flexible set-conditional modeling in Transformer-based probabilistic models for joint distribution forecasting, this paper proposes Causal Autoregressive Caching (CAC). CAC decouples context encoding from conditional set updates via a dynamic cache buffer, enabling single-shot conditional encoding, multi-step cache reuse, and batched autoregressive decoding—thereby unifying efficient joint log-likelihood evaluation and joint set-conditional/autoregressive training. Innovatively, it introduces cache-context attention and target-dependent dynamic modeling, achieving substantial inference speedups without compromising conditional modeling flexibility—a first in the literature. Experiments across synthetic functions, EEG time series, cognitive modeling tasks, and tabular data demonstrate that CAC matches state-of-the-art baseline accuracy while accelerating joint sampling by up to 20×.

Achieving efficient joint distribution sampling in transformer probabilistic modelsBalancing autoregressive generation with flexible set-conditioning capabilitiesEliminating expensive re-encoding during autoregressive inference steps

Latest Papers

What's happening recently
View more

This work addresses the challenge of extracting structured knowledge from large language models (LLMs) in a probabilistically coherent manner to support complex variable reasoning. It proposes Large Language Gibbs (LLG), a method that embeds the LLM’s conditional distribution as a transition kernel within a Gibbs sampling framework. By iteratively resampling individual variables rather than generating sequences autoregressively, LLG avoids biases induced by sequential dependencies. This approach constitutes the first effective integration of LLMs with Markov chain Monte Carlo (MCMC), yielding a stationary distribution consistent with all local conditional distributions. Empirical results demonstrate that LLG excels in tasks including synthetic distribution sampling, consistency-aware reasoning, and Bayesian structure learning, establishing a new paradigm for structured probabilistic inference beyond single-pass generation.

conditional distributionslarge language modelsMCMC

This work proposes the first semantically consistent sequential Monte Carlo (SMC) framework grounded in the Feynman–Kac formalism for efficient and provably correct inference in general-purpose probabilistic programs that support arbitrary measure sampling and conditional reweighting within unbounded loops. The approach employs probabilistic program graphs (PPGs) as an intermediate representation and leverages a finite-trace approximation theorem to rigorously establish the correspondence between the program’s expectation semantics and the Feynman–Kac model. Building on this foundation, the authors design a vectorized particle filtering algorithm (VPF) tailored to PPGs. Empirical evaluations demonstrate that VPF significantly outperforms existing state-of-the-art inference tools across multiple benchmarks, achieving a compelling combination of theoretical soundness, computational efficiency, and strong scalability.

Feynman-Kac ModelsOperational SemanticsParticle Filtering

This work addresses the limitations of traditional autoregressive models, which rely solely on next-token prediction and struggle to capture sequence-level properties, often resulting in local overfitting and poor global structure modeling. Moreover, controllable generation typically demands costly sampling or architectural modifications. To overcome these challenges, the authors propose the Conditional Attribute Transformer, which jointly optimizes next-token prediction and sequence attribute estimation conditioned on the current token within a single forward pass. This approach uniquely enables token-level attribute attribution, counterfactual attribute evaluation, and attribute-guided generation—all within a unified model without additional sampling or structural changes. Experiments demonstrate state-of-the-art performance on sparse-reward tasks, improved language modeling at large scales, attribute estimation orders of magnitude faster than sampling-based methods, and effective guidance across diverse language generation tasks.

attribute estimationautoregressive sequence modelsglobal structure

Traditional statistical inference often fails when models or parameters are selected in a data-driven manner. This work systematically investigates selective inference, focusing on a conditional inference framework that conditions on the selection event to restore the nominal coverage of confidence intervals. We clarify the scientific interpretation of this framework, unify several existing approaches under its umbrella, and demonstrate its application to canonical settings such as inference for the “winner,” region-specific means in regression trees, and differences between clusters. Through simulations and analyses of single-cell RNA sequencing data, we show that the proposed methodology yields valid and reliable statistical inference in practical scenarios.

conditional inferencedata-driven selectionpost-selection inference

Hot Scholars

YZ

Yue Zhang

Postdoc, UNC Chapel Hill
NLPVLNMulti-modal
ZH

Zhicheng He

Huawei Noah's Ark Lab
recommender systemnatural language processingnetwork embedding
HL

Han Lin

CS PhD Student, UNC
Image/Video GenerationMultimodal LearningSelf-Supervised LearningLLMs
JC

Jaemin Cho

PhD Student at UNC Chapel Hill
Multimodal LearningNatural Language ProcessingMachine Learning
ZW

Zun Wang

UNC Chapel Hill
Multimodal AlGenerative AlEmbodied Al