Learning to Predict Distributions over Weight Updates for Test-Time Adaptation

πŸ“… 2026-10-01
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of achieving efficient test-time adaptation and compute scaling for large language models (LLMs) using only query signals. To this end, it proposes a distributed conditional hypernetwork that dynamically adapts model parameters by sampling weights rather than tokens to predict LoRA update distributions. Furthermore, the method introduces a novel test-time scaling mechanism enabled by differentiable Monte Carlo approximation and distributional parameterization. Empirical results demonstrate that this approach outperforms deterministic baselines, supports multi-model sampling, and enables weight updates to transfer across queries. Ultimately, this work establishes an efficient new paradigm for test-time computation in LLMs.
πŸ“ Abstract
Hypernetworks have recently shown success in dynamically adapting the parameters of Large Language Models (LLMs) at runtime based on signals such as task descriptions or additional demostrations. Here we ask: how much adaptation signal can be obtained using only the input query to an LLM?. To answer this, we study query-conditioned Hypernetworks for LoRA estimation. Further, we introduce distributional Hypernetworks, able to produce not only point estimates of parameter adaptors, but also a distribution over possible LoRAs. For this we propose a simple end-to-end loss using a differentiable Monte Carlo approximation and explore multiple distribution parametrizations including regression and convex combination variants. Results show that even using the mean of the learned distribution can outperform deterministic hypernetworks. Crucially, the learned distribution enables a different form of test-time scaling: instead of spending additional compute only by sampling more token sequences from a fixed model, we sample weight updates, yielding multiple adapted models for the same query. Performance improves as more weight samples are considered and remains stronger than corresponding token-sampling adaptation baselines. Finally, we find that generated updates can transfer across queries, suggesting that the hypernetwork learns reusable structure in how the model should adapt. Together, these results show that query-conditioned distributions over weight updates can support both adaptation and test-time scaling.
Problem

Research questions and friction points this paper is trying to address.

Test-Time Adaptation
Hypernetworks
Large Language Models
Test-Time Scaling
Weight Updates
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hypernetworks
Test-Time Adaptation
Distributional LoRA
Test-Time Scaling
Monte Carlo Approximation
πŸ”Ž Similar Papers
No similar papers found.