Score
Designs, implements, modifies, and analyzes transformer-based sequence models and their components—including attention mechanisms, layer blocks, decoder variants, and internal representations—and engineers their training and inference pipelines. Works on architecture adaptations (e.g., pruning, memory/latency reductions), conditioning and interfacing mechanisms (variable-length tokens, positional/length embeddings, cross-length alignment), and optimization for inference (plug-and-play use, efficient implementations) to meet specified performance, resource, and compatibility constraints.
Traditional Transformers incur prohibitively high computational costs and resource consumption during large-scale training and deployment. Method: This paper systematically constructs an efficient architecture framework for large language models (LLMs), proposing a unified taxonomy that integrates linearized sequence modeling, sparse attention mechanisms, Mixture-of-Experts (MoE) architectures, diffusion-based language modeling, and multimodal transfer techniques. It innovatively unifies sparsification, linearization, and hybrid modeling paradigms to jointly ensure theoretical interpretability and engineering practicality. Contribution/Results: We present the first comprehensive architectural landscape of efficient LLMs—spanning training, inference, and multimodal extension—providing a systematic blueprint for scalable foundation model design under resource constraints. Our framework significantly reduces computational and memory overhead, enabling the practical deployment of high-performance, low-resource AI systems.
This work clarifies the functional boundaries and necessity of core Transformer components—tokenization, embedding/un-embedding, masking, positional encoding, and padding—addressing widespread conceptual ambiguity in their mechanistic roles. Targeting ML engineers, we propose an incremental, invertibility-based analytical framework: using binary (0/1) sequences as probes, we systematically introduce each component via a “zero-one construction” and empirically validate its irreplaceability in the encode-decode pipeline. Implemented lightweightly in PyTorch, our framework supports manual attention matrix construction, explicit positional embedding injection, and interpretable mask design. Experiments demonstrate significantly improved conceptual accuracy among learners. Notably, we provide the first empirical verification that, in the absence of self-attention, positional encoding combined with padding alone suffices for basic length-aware tasks.
This study systematically investigates how positional encoding affects the expressive power, generalization, and long-sequence extrapolation capability of Transformers. To address this, we propose a unified theoretical framework that, for the first time, incorporates linear bias methods—including ALiBi—into the formal modeling of positional encodings. We further introduce a novel orthogonal encoding scheme based on wavelet and Legendre polynomial transforms, and rigorously establish its superiority via function approximation theory and Rademacher complexity analysis. Empirical evaluation on synthetic sequence tasks demonstrates that our encoding reduces generalization error by 37% and improves extrapolation length by 2.1× compared to sinusoidal encoding, significantly enhancing out-of-distribution generalization to unseen sequence lengths. Our core contributions are (i) a unified analytical paradigm for positional encoding, grounded in learning theory, and (ii) a theoretically justified, orthogonal encoding design that achieves state-of-the-art empirical performance.
This study investigates how sequence modeling architectures affect the foundational capabilities of pretrained language models, revealing significant degradation in state-based architectures (e.g., RNNs, Mamba) under constrained-domain pretraining and out-of-distribution evaluation. To address this, we propose “full-sequence arbitrary selection capability” as a core architectural design principle and instantiate it via a lightweight Top-1 element/block selection mechanism. Through ablation studies, cross-distribution evaluation, and joint efficiency-capability analysis, we demonstrate that this capability strongly correlates with foundational competencies—including long-range dependency modeling and symbolic reasoning. Crucially, the Top-1 block selection architecture fully restores Transformer-level foundational capabilities with negligible computational overhead. Our work provides both theoretically grounded principles and empirically validated pathways for designing efficient, capable sequence modeling architectures. (149 words)
This work investigates how Transformers model hierarchical positional dependencies in tree-structured data. We propose a hierarchical filtering generative model that enables controlled modulation of positional dependencies across multiple scales, and integrate attention map analysis, hierarchical probing, and encoder-only training to systematically characterize the underlying modeling mechanisms. We make the first discovery that Transformer encoder layers progressively capture long-range hierarchical dependencies with depth: shallow layers encode local adjacency relations, while deeper layers specialize in global tree topology; moreover, each layer approximately reconstructs correlation patterns at a specific scale. Empirical results demonstrate that this architecture achieves performance approaching exact Bayesian inference on trees for root-node classification and masked language modeling tasks. Our findings establish a verifiable, scale-separated computational mechanism for interpretable AI, grounded in principled hierarchical representation learning.
The functional mechanisms underlying layer-wise operations in Transformer models remain poorly understood, particularly regarding the necessity and interchangeability of layer ordering. Method: This work proposes an empirical analysis framework based on freezing large language models (LLMs) and systematically conducts three types of architectural interventions: layer ablation, layer reordering, and parallel layer execution. Contribution/Results: We discover that middle layers exhibit strong functional uniformity and order invariance—enabling safe skipping, arbitrary reordering, or concurrent execution—thereby challenging the conventional assumption of strict layer-order dependency. Across diverse downstream tasks, skipping or parallelizing middle layers reduces inference latency by up to 30% on average, with accuracy degradation under 1%. This study is the first to empirically characterize the functional heterogeneity spectrum across Transformer layers, providing an interpretable, evidence-based foundation for model lightweighting, architectural compression, and novel variant design.
This work proposes a general method to automatically decompile concise and interpretable RASP programs from Transformer models that exhibit strong performance on length generalization tasks, thereby verifying whether these models truly implement generalizable algorithmic logic. By integrating Transformer reparameterization, causal intervention analysis, and RASP program synthesis with simplification techniques, the approach identifies minimal subprograms sufficient to explain model behavior. Applied across multiple algorithmic and formal language tasks, the method successfully recovers simple RASP programs whose execution matches the models’ predictions, offering the first direct and interpretable evidence of the computational mechanisms internally implemented by Transformers.
This work investigates the trade-off between expressivity and computational efficiency in sequence modeling and proposes a novel hybrid architecture that integrates Transformers with state space models (SSMs). Through theoretical analysis, it establishes—for the first time—that pure Transformers or SSMs inherently suffer from fundamental limitations in parameter or memory requirements on certain tasks. To overcome this bottleneck, the authors construct a provably effective hybrid model. Experiments demonstrate that the proposed small-scale hybrid model outperforms non-hybrid counterparts with up to six times more parameters on tasks such as selective copying and associative recall, achieving significantly lower memory consumption while exhibiting superior length generalization and out-of-distribution robustness.
This work investigates the number of distinct output sequences a Transformer model can generate given a prompt and uncovers the fundamental reasons behind its failures in simple tasks such as copying and memorization. Through rigorous theoretical analysis, the study establishes—for the first time—that the maximum length of accessible sequences grows linearly with prompt length, while the fraction of accessible sequences decays exponentially beyond a critical length, a phenomenon that persists even with unlimited context and computational resources. Combining upper-bound derivations, asymptotic analysis, and experiments across multiple architectures, the proposed theory remains tightly aligned with empirical observations across model scales, with error factors below 10, thereby offering a precise characterization of the expressive limitations inherent to Transformers.
This work challenges the prevailing view of large language models as mere “stochastic parrots” by introducing the Sequence-level Interactive Dynamic Parallel Processing (SIDPP) framework, which conceptualizes Transformers as systems that dynamically generate transformation parameters from input prompts to perform concept-to-concept mappings. The framework incorporates an output-weight interconnection mechanism that reveals a strong prompt sensitivity—where dynamic processing capacity intensifies with longer prompts—and suggests a potential correspondence with human cortical language processing. Experimental results demonstrate that such dynamic processing can contribute comparably to, or even surpass, static processing in model performance. These findings not only open new avenues for model interpretability and controllability but also provide theoretical foundations for developing compact, efficient architectures and advancing our understanding of human language cognition.
本文提出了一种通过合并模块压缩输入序列的方法,以减少Transformer模型的计算成本,同时保持准确性。