Score
Designs and implements encoder networks based on transformer architectures that accept variable-size, unordered input sets and produce permutation-invariant, compact latent representations; this includes set transformer variants and transformer-adapted encodings. Builds and evaluates models and processing pipelines that generate features suitable for retrieval and downstream tasks while supporting efficient, lightweight inference and scalability across varying set sizes.
This work investigates the theoretical connections between Transformers and Graph Neural Networks (GNNs), challenging the conventional view of Transformers as purely sequence-based models. Method: We formalize the Transformer as a message-passing GNN operating on a fully connected token graph, where self-attention implements dynamic, content-aware neighborhood aggregation and positional encodings implicitly encode structural priors. We reinterpret Transformer computation within a unified message-passing framework and introduce the concept of the “hardware lottery”—highlighting that its efficiency stems not only from architectural design but also from hardware-level optimizations for dense matrix operations in modern accelerators. Contribution/Results: (1) We establish the first rigorous theoretical bridge between NLP and graph representation learning; (2) we identify the dual origins of Transformer expressivity—structural modeling capacity and hardware alignment; and (3) we propose a new paradigm for interpretable modeling and hardware-aware neural architecture design.
Existing efforts to encode algorithms directly into Transformer parameters are fragmented and lack systematicity, hindering both learning and reuse. Method: This paper introduces, for the first time, a unified “algorithm encoding recipe” framework that systematically integrates arithmetic implementations in feed-forward layers with dynamic data routing via self-attention, yielding a modular and composable Transformer construction methodology. Contribution: It bridges a critical gap between interpretable modeling and controllable architectural design of Transformers. The framework enables rigorous computational complexity analysis, formally verifiable architecture specification, and mechanistic interpretation of internal operations. By significantly lowering the barrier to algorithmic encoding—providing accessible onboarding for novices and structured, reusable building blocks for experts—it advances research and deployment of controllable, interpretable Transformers.
This work addresses the high computational overhead in semantic classification caused by processing raw bytes or fully decoding compressed media. We propose TEMPEST, the first method to directly feed the intrinsic byte-stream structure of compressed files into Transformers—bypassing decoding and full media reconstruction. TEMPEST introduces lightweight, compression-aware tokenization and encoding strategies grounded in the statistical and structural properties of compressed data, enabling end-to-end semantic representation learning. Evaluated across multimodal domains (images, audio) and diverse compression formats (JPEG, MP3, etc.), TEMPEST achieves on-par classification accuracy with state-of-the-art methods while reducing token count by 62% on average, significantly lowering memory footprint and FLOPs. Its core innovation lies in treating compressed-domain byte streams as inherently semantic-rich, compact representations—establishing a new paradigm for efficient, cross-modal, decoding-agnostic representation learning.
Scaling Transformer models is prohibitively expensive due to fixed-parameter linear projection layers; architectural modifications necessitate full retraining. Method: We propose TokenFormer, the first architecture introducing *parameter tokenization*, which models model parameters as learnable tokens and replaces all linear layers with token-parameter self-attention—unifying parameter and input token representations in a shared latent space. Contribution/Results: Our method enables zero-shot, progressive parameter expansion without retraining, overcoming classical scaling bottlenecks. Without altering network topology, we scale model parameters from 124M to 1.4B while matching the performance of fully trained baselines, achieving substantial training cost reduction. The code and models are publicly released.
Existing neural network weight representation methods are constrained by architecture and scale, limiting generalization across heterogeneous architectures and datasets. This paper proposes the SNE encoder—the first approach to produce unified, set-level representations of neural networks regardless of architecture or parameter count, enabling cross-architecture and cross-dataset network property prediction. Our method introduces three key innovations: (1) a Logit Invariance constraint that jointly models computational hierarchy and weight-space symmetry; (2) a tunable, hierarchical encoding pipeline comprising padding, chunking, and encoding stages; and (3) formal definition and solution of the novel task of cross-dataset/cross-architecture prediction. Evaluated on standard benchmarks, SNE significantly outperforms existing baselines, demonstrating strong generalization capability and explicit architecture independence.
This work addresses the limitation of Transformers in lacking an explicit knowledge storage mechanism, which hinders efficient retention and retrieval of learned information. To overcome this, the authors propose a chaptered sparse memory bank, where learnable memory tokens are queried by the Transformer via cross-attention. Inspired by Mixture-of-Experts, a dynamic chapter routing strategy selectively activates relevant subsets of memory, enabling scalable knowledge access while maintaining computational efficiency. The approach expands memory capacity to 262K tokens—introducing a new scaling dimension beyond model parameters—without incurring prohibitive computational overhead. Experiments demonstrate that, under matched FLOPs, the proposed model outperforms standard Transformers in both pretraining and instruction fine-tuning tasks, while also exhibiting substantially improved knowledge retention and robustness against catastrophic forgetting in continual learning scenarios.
This work addresses the lack of rigorous theoretical analysis regarding the expressive power of Transformers, particularly in approximating general function classes. By establishing an explicit approximation relationship between Transformers and Maxout networks, the study constructs a theoretical framework that reveals key structural properties: the self-attention layer can implement max-type operations, while the feedforward layer performs token-wise affine transformations. For the first time, the paper connects the representational capacity of Transformers to classical approximation theory for feedforward networks, proving that Transformers can approximate Maxout networks with comparable complexity, thereby inheriting the universal approximation capability of ReLU networks. Furthermore, it quantifies how the number of linear regions grows exponentially with depth, elucidating the role of depth in enhancing expressive power.
This work addresses the parameter redundancy in embedding dimensions of large language models, which incurs substantial computational and memory costs during scaling. The authors propose L-Transformer, the first architecture to enable differentiable spectral decomposition of the embedding space. By leveraging the third-order tensor L-product, token embeddings are reshaped into spectral tensor slices, and attention and feed-forward operations are performed in the transform domain. This yields a Tensor Transformer composed of p independent spectral sub-transformers, introducing an inductive bias at the frequency level that supports slice-dependent frequency scaling to enhance generalization while remaining compatible with standard training pipelines. Experiments show that on IMDB and AG News, the encoder achieves up to 75% parameter reduction (with p=4) while maintaining competitive accuracy, and fully recovers baseline performance at BERT-base width.
This work investigates how deep Transformers achieve adaptive in-context reasoning under constraints of communication, locality, and depth. By modeling the network as a mean-field interacting system, the authors introduce “function vectors” to represent internal states and develop a distributed inference theory that reveals a nontrivial relationship between network depth and the hierarchical structure of latent variables. Within this theoretical framework, they integrate a constrained linear attention architecture and validate its predictions on in-context regression tasks, demonstrating that deep Transformers possess expressive adaptive reasoning capabilities surpassing those of prior models.
This study investigates the geometric limits of feature representations in Transformer language models, focusing on the upper bound of nearly orthogonal directions. Building upon assumptions of linear representability and superposition, the work analyzes the boundary of cosine similarity distributions in embedding matrices and introduces a tolerable orthogonality deviation ε. Centered on ε, it proposes a refined representational capacity formula that reveals exponential sensitivity of capacity to ε. The analysis further uncovers that large models tend to strengthen orthogonality constraints rather than merely expanding dimensionality. Combining cosine similarity analysis, a modified Johnson–Lindenstrauss lemma, and modeling of near-orthogonal packing efficiency, the method validates two distinct orthogonality patterns across dozens of open-source models, reducing capacity prediction error by two orders of magnitude and establishing, for the first time, a quantifiable metric for representational capacity.