Score
Designs and implements mechanisms that condition models at inference time on provided contextual examples or spatial priors, including techniques that inject per-location (spatial) priors into inputs, features, or latent states to steer outputs. Builds and evaluates methods for transferring those in-context priors and for enabling data-efficient adaptation for dense, per-location prediction tasks.
This work addresses the challenge of enabling predictive systems to dynamically adapt their behavior based on contextual information for personalized inference. To this end, it proposes a unified framework that maps context into adaptation parameters for prediction and, for the first time, establishes a mathematical equivalence between explicit parameter adaptation and implicit expert routing under kernel ridge regression. The framework theoretically unifies diverse methodologies—including varying-coefficient models, local regression, prompt engineering, retrieval-augmented approaches, and mixture-of-experts—under fixed features and squared loss. Key contributions include deriving a general formulation for context-adaptive inference, proposing practical design principles and evaluation metrics such as adaptation efficiency and routing stability, and highlighting critical open problems concerning identifiability and robustness under distributional shifts.
Traditional Bayesian inference is computationally expensive and difficult to scale, while existing in-context learning approaches lack the flexibility to adapt to new priors at test time, resulting in poor robustness under distributional shift. This work proposes a multitask in-context learning framework that explicitly encodes prior information as a prefix in the Transformer input sequence, enabling adaptive prediction across diverse prior families. By integrating hierarchical Bayesian reasoning with sequence modeling, the method supports efficient generalization to both unseen priors and high-dimensional latent-structured priors. Experiments demonstrate that the approach achieves predictive performance comparable to an ideal Bayesian inference engine while accelerating inference by several orders of magnitude, with practical efficacy validated on real-world spatiotemporal temperature forecasting tasks.
This paper addresses the challenge of posterior inference in large-scale nonparametric regression. We propose a spatially adaptive distributed Gaussian process (GP) approximation method. The approach partitions the input space into disjoint subsets, fits independent GP posteriors on each subset using a Matérn kernel and an integrated Brownian motion prior, and incorporates a Bayesian prior on the length scale to enhance regularization of local smoothness. A novel weighted spatial aggregation scheme is then introduced to fuse these sub-posteriors into a global approximation of the full-data posterior. Theoretically, we establish that the resulting approximate posterior achieves a convergence rate that automatically adapts to the local smoothness of the true regression function—matching the minimax optimal rate attainable under the full-data setting. Empirically, our method significantly outperforms existing distributed GP approaches on both synthetic and real-world datasets, while effectively capturing heterogeneous local regularity and overcoming the smoothness rigidity inherent in standard GP models.
To address the reliance on manual intervention and high-performance hardware for automated analysis, model evaluation, and uncertainty quantification in large-scale geospatial data, this paper introduces the first Bayesian predictive stacking framework tailored for geospatial transfer learning. The method integrates Bayesian modeling, predictive stacking, streaming minibatch training, and spatial statistical inference to enable continual learning propagation and full-scale inference. It achieves low-overhead, scalable, and interpretable joint uncertainty quantification for multivariate geospatial predictions on commodity GPUs. Experiments on climate science datasets—sea surface temperature and vegetation index—demonstrate over 60% faster inference and significantly improved accuracy, while eliminating dependence on high-end hardware and manual hyperparameter tuning. This framework advances practical, resource-efficient uncertainty-aware geospatial modeling.
To address computational bottlenecks in latent spatial field inference and pointwise prediction in geostatistical models, this paper proposes Bayesian Predictive Stacking (BPS). BPS analytically aggregates posterior distributions of regression coefficients and spatial processes across multiple hyperparameter configurations, circumventing iterative algorithms such as MCMC and enabling fully parallelized inference. Its key innovation lies in a unified stacking framework that jointly combines predictive means and posterior densities, underpinned by infill asymptotic theory that guarantees statistical consistency. Experiments demonstrate that BPS achieves predictive accuracy comparable to full-sample Bayesian inference while drastically reducing computational time. The method is thus highly efficient, robust to hyperparameter specification, and scalable to large spatial datasets.
This study addresses the unclear necessity and practical benefits of external context for hypotheses in model-based reinforcement learning. We propose a "predictive sufficiency" metric that quantifies the contribution of context to state prediction, effectively distinguishing historically recoverable information from genuine requirements to guide mechanism matching. Methodologically, we construct a practical criterion based on standard trained agents that evaluates context effectiveness without requiring a reference policy, integrating risk decomposition analysis with classification algorithms into a comprehensive evaluation framework. Our findings reveal that this metric remains invariant across MDP classes and demonstrate that external context yields no substantial gains in implicitly identifiable scenarios. These results provide both theoretical foundations and empirical guidance for the design of context mechanisms in reinforcement learning.
This study addresses the disconnect between spatial reasoning benchmarks and navigation performance by proposing the Spatial-NPD framework. We first construct the Spatial-Nav-100K dataset and subsequently employ a two-stage fine-tuning strategy combined with large language model training. Through knowledge distillation, a teacher model guides a lightweight student model to achieve efficient navigation without explicit spatial reasoning. The resulting 8B-parameter model attains state-of-the-art performance on benchmarks such as HM3D while maintaining minimal computational overhead, requiring only 45 GPU hours for training and achieving an inference speed of 148 ms per step. This work demonstrates that high-performance embodied navigation can be effectively realized through distilled implicit spatial understanding rather than costly explicit reasoning pipelines.
This study addresses the challenge of poor cross-regional generalization in spatiotemporal point process models under sparse event history conditions. The authors propose integrating AlphaEarth geographic embeddings as static exogenous spatial context into a log-Gaussian Cox process, enabling substantial improvements in cross-regional event prediction performance using only information available prior to prediction. The approach demonstrates, for the first time, that static spatial context can yield 2–6× performance gains even with extremely short historical windows (1–2 weeks) and maintains consistent improvements of 10%–20% with longer histories (20–104 weeks). Evaluated on EMS event prediction across eight held-out regions, the method exhibits strong robustness and transferability, highlighting its effectiveness in real-world sparse-data scenarios.
This study addresses the challenge of enabling general-purpose robots to infer novel tasks from demonstrations and translate them into physical actions under fixed parameters. It presents a systematic review of in-context learning for robotics, categorizing interaction interfaces into four types and examining the roles of training, correspondence, and memory mechanisms in manipulation and navigation. The work synthesizes techniques including context-conditioned policies, geometric demonstration transfer, world-model-based control, and skill agents. Its core contributions lie in clarifying transfer assumptions across different interfaces, bridging methodological design with evaluation practices, and elucidating mechanisms that preserve instructional requirements amid environmental variations. Furthermore, it proposes a research agenda centered on recursive self-improvement, establishing a foundation for compositional task acquisition and faithful physical transfer.
This work addresses the issue in trajectory prediction where Winner-Take-All (WTA) training yields uninformative posterior probabilities over modes, hindering effective mode pruning. The study identifies hard assignment in WTA as the root cause of mode oversplitting and instability, and unifies existing approaches under a Gaussian Mixture Model (GMM) framework. To rectify mode probabilities without retraining, the authors propose two lightweight post-processing strategies: test-time posterior-weighted fusion and a single-step EM-based soft responsibility update. Evaluated across multiple WTA-based architectures, these methods significantly enhance the informativeness of posterior probabilities, improve mode ranking accuracy, and boost overall prediction performance.