Elastic Selective Spectral Hybrids for Train-Once, Export-Many Budgeted Inference

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing elastic architectures in adaptively preserving critical information when deploying language models under varying computational budgets. We propose ESSH, an architecture that integrates a selective spectral mixer with sliding window attention, introducing input-dependent decay gating and dual-rate capacity mapping for dynamic memory management. Furthermore, it employs damped rotational modes to construct recurrent units and utilizes full-model distillation to support multi-capacity joint training and parallel decoding, enabling single-training, multi-capacity export. Experimental results demonstrate that a 1.53B-parameter ESSH model achieves 1.37 ms/token on the B300 platform, yielding a 2.14–2.8× speedup over Mamba-series baselines while realizing a smooth trade-off between inference quality and computational cost.
📝 Abstract
Deploying a language model under different computing and latency budgets calls for compact models with different quality-cost trade-offs. Towards this end, elastic spectral state space models provide ordered and temporally decomposed channels that can be truncated, but they are associated with linear time-invariant filters that cannot selectively preserve relevant past information or forget irrelevant information as the context evolves. To address this, we introduce the Elastic Selective Spectral Hybrid (ESSH), which realizes each Hankel spectral channel as an independent recurrent unit using fitted damped rotation modes. It also features an input-dependent decay and write/read gates that make temporal retention and state update input-dependent while preserving channel-wise truncation and a structured recurrence for efficient execution. ESSH combines these selective spectral mixers with sliding-window attention and jointly trains multiple capacities by reducing spectral-channel count and feed-forward width at different rates through a two-rate capacity map with full-model distillation. The resulting models support chunked parallel training and fused recurrent decoding while avoiding computation for discarded channels. At full capacity, ESSH achieves language-modeling quality comparable to similarly sized independently trained models, while smaller exports exhibit a smooth quality-cost trade-off. We validate the effectiveness of the proposed framework using language understanding, retrieval, cross-domain text, and DNA experiments, by assessing quality retention and the trade-off against independently trained and elastic baselines. At the 1.53B model configuration, fused batch-one decoding takes 1.37 ms per token on a B300, providing a 2.14-2.80x speedup over the tested Mamba-2 and Mamba-3 implementations and 3.03x over Transformer++ at matched parameter counts.
Problem

Research questions and friction points this paper is trying to address.

elastic inference
budgeted deployment
selective state space models
language models
quality-cost trade-off
Innovation

Methods, ideas, or system contributions that make the work stand out.

Elastic Selective Spectral Hybrid
Input-dependent gating
Two-rate capacity map
Sliding-window attention
Train-once export-many
🔎 Similar Papers
No similar papers found.