Test-Time Adaptation of Reasoning Strategies with Bayesian Nonparametric Memory

πŸ“… 2026-10-05
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the high inference costs of large language models and their lack of mechanisms to transfer test-time insights to subsequent queries. To overcome these limitations, this work proposes a structured Bayesian cheat sheet based on a Hierarchical Dirichlet Process Gaussian Mixture Model (HDP-GMM). By clustering behavioral embeddings, the module facilitates cross-domain sharing and dynamic recombination of domain-specific weights. Furthermore, it integrates posterior predictive retrieval with collapsed Gibbs sampling to enable low-cost online adaptive updates. Experimental results demonstrate that the proposed approach significantly outperforms existing methods on benchmarks such as AIME'25. Notably, it maintains effective performance improvements even in cold-start scenarios, thereby validating the value of structured metacognitive reasoning for enhancing model efficiency and adaptability.
πŸ“ Abstract
While modern large language models (LLMs) have been trained to reason through verbalized chains-of-thought, the generation cost grows substantially due to suboptimal paths to reach the final answer. Furthermore, as new insights are discovered while observing various input queries (e.g. through self-reflection), limited mechanisms exist for carrying forward these findings to be applied to subsequent problems. One can view the list of such strategies or behaviors as a growing cheatsheet, with elements retrieved from this memory module at inference-time. In this work, we consider structured cheatsheets, with learned clusters of behaviors. We introduce a Hierarchical Dirichlet Process Gaussian Mixture Model (HDP-GMM) over behavior embeddings, which shares components across domains while allowing domain-specific mixing weights, and uses the posterior predictive to retrieve relevant behaviors for a query; we call this a $\textit{Bayesian Cheatsheet}$. This mechanism allows for cheap adaptation in an online test-time training (TTT) setting, softly updating the mixture's sufficient statistics following each sample and enabling the creation of new components when the synthesized behaviors are sufficiently novel. We demonstrate that Bayesian Cheatsheet achieves clear performance gains relative to existing memory modules across reasoning benchmarks such as AIME'25, Omni-MATH, and PhysReason, even in the cold-start setting. We show that the Bayesian Cheatsheet is an adaptively reorganizing memory module, as behaviors can be re-assigned to components through a single step of collapsed Gibbs sampling. Our findings highlight the value of Bayesian-inspired memory modules for effective test-time adaptation and the role of structure in metacognitive reasoning.
Problem

Research questions and friction points this paper is trying to address.

Test-Time Adaptation
Reasoning Strategies
Large Language Models
Memory Module
Chain-of-Thought
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-Time Adaptation
Bayesian Nonparametric Memory
Hierarchical Dirichlet Process
Reasoning Strategies
Large Language Models