π€ AI Summary
This study addresses the challenge of achieving efficient test-time adaptation and compute scaling for large language models (LLMs) using only query signals. To this end, it proposes a distributed conditional hypernetwork that dynamically adapts model parameters by sampling weights rather than tokens to predict LoRA update distributions. Furthermore, the method introduces a novel test-time scaling mechanism enabled by differentiable Monte Carlo approximation and distributional parameterization. Empirical results demonstrate that this approach outperforms deterministic baselines, supports multi-model sampling, and enables weight updates to transfer across queries. Ultimately, this work establishes an efficient new paradigm for test-time computation in LLMs.
π Abstract
Hypernetworks have recently shown success in dynamically adapting the parameters of Large Language Models (LLMs) at runtime based on signals such as task descriptions or additional demostrations. Here we ask: how much adaptation signal can be obtained using only the input query to an LLM?. To answer this, we study query-conditioned Hypernetworks for LoRA estimation. Further, we introduce distributional Hypernetworks, able to produce not only point estimates of parameter adaptors, but also a distribution over possible LoRAs. For this we propose a simple end-to-end loss using a differentiable Monte Carlo approximation and explore multiple distribution parametrizations including regression and convex combination variants. Results show that even using the mean of the learned distribution can outperform deterministic hypernetworks. Crucially, the learned distribution enables a different form of test-time scaling: instead of spending additional compute only by sampling more token sequences from a fixed model, we sample weight updates, yielding multiple adapted models for the same query. Performance improves as more weight samples are considered and remains stronger than corresponding token-sampling adaptation baselines. Finally, we find that generated updates can transfer across queries, suggesting that the hypernetwork learns reusable structure in how the model should adapt. Together, these results show that query-conditioned distributions over weight updates can support both adaptation and test-time scaling.