Universal Test-Time Training

📅 2026-10-04
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing test-time training (TTT) architectures, where memory is private to individual layers and recurses solely along the temporal dimension without cross-depth sharing. To overcome this, we propose uTTT, a framework that introduces a cross-layer shared memory mechanism enabling all layers to read from and write to a unified memory. This design facilitates dual recursion across both temporal and depth dimensions, instantiated in Mixture-of-Experts (MoE) and Dense variants. By decoupling memory from network depth, uTTT supports deep writing and shallow reading. Experimental results demonstrate that uTTT improves language modeling accuracy by over two percentage points and yields significant PSNR gains in novel view synthesis. Furthermore, it consistently outperforms layer-private models under equivalent computational budgets, establishing a more effective paradigm for test-time adaptation.
📝 Abstract
Recent Test-Time Training (TTT) architectures compress context into fast weights that are updated online and queried as memory. Existing TTT designs keep this memory private to each layer: it recurs only over time, and depth merely indexes L separate memories. We argue that memory ownership need not be tied to depth, and introduce Universal Test-Time Training (uTTT), in which all layers read and write one shared memory while retaining layer-specific backbone parameters. The shared memory thus recurs over two dimensions, time and depth, with chunks and layers as their units: a write by a deep layer in one chunk can be read by a shallow layer in the next. We instantiate this idea as uTTT-MoE and uTTT-Dense. uTTT-MoE routes each token head to a few experts in a pool shared by all layers; uTTT-Dense applies the whole shared memory at every layer without routing. In language modeling, uTTT-MoE reaches 15.5 and 27.9 RULER accuracy at 124M and 760M, 2.6 and 2.1 points above its layer-private counterpart at equal state and active compute, the highest among tested bounded-state models, with per-token loss matching or beating full attention. In novel view synthesis, sharing at fixed per-layer compute gains 0.92 dB in view-23 object PSNR in routed models and 0.76 dB in dense models.
Problem

Research questions and friction points this paper is trying to address.

Test-Time Training
Shared Memory
Layer-Private Memory
Language Modeling
Novel View Synthesis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-Time Training
Shared Memory
Mixture of Experts
Cross-layer Recurrence
Universal TTT