Score
Design and implement transformer-based encoders — including GPT-2–style autoregressive and causal self-attention variants — that extract hierarchical, sequence-level features and model long-range dependencies. Build and evaluate feature-extraction pipelines and attention analyses that produce compact, transferable encoder weights and diagnostic feature maps to support representation learning and downstream tasks.
Existing feature transformation methods predominantly rely on computationally intensive encoder-decoder architectures, resulting in excessive parameter counts, slow inference, and poor scalability. This paper proposes a lightweight generative feature transformation framework that reformulates feature transformation as an autoregressive sequence reconstruction task by rearchitecting the GPT architecture, jointly optimizing embedding-space continuity and downstream task performance. Crucially, we introduce a gradient-ascent-guided learnable transformation mechanism that eliminates the explicit encoder, reducing model parameters by 62% on average and accelerating inference by 2.3×. Evaluated across multiple benchmark datasets, our method achieves state-of-the-art or competitive performance on diverse downstream tasks—including classification and regression—while demonstrating strong generalization and deployment efficiency.
This work clarifies the functional boundaries and necessity of core Transformer components—tokenization, embedding/un-embedding, masking, positional encoding, and padding—addressing widespread conceptual ambiguity in their mechanistic roles. Targeting ML engineers, we propose an incremental, invertibility-based analytical framework: using binary (0/1) sequences as probes, we systematically introduce each component via a “zero-one construction” and empirically validate its irreplaceability in the encode-decode pipeline. Implemented lightweightly in PyTorch, our framework supports manual attention matrix construction, explicit positional embedding injection, and interpretable mask design. Experiments demonstrate significantly improved conceptual accuracy among learners. Notably, we provide the first empirical verification that, in the absence of self-attention, positional encoding combined with padding alone suffices for basic length-aware tasks.
This work addresses the challenge of moving beyond correlational analyses to establish causal relationships in identifying truly influential features within Transformer models. It proposes a five-stage causal analysis framework—comprising probe design, feature extraction, causal validation, robustness testing, and deployment integration—to systematically evaluate features in GPT-2 small on the Indirect Object Identification (IOI) task. By integrating activation patching, sparse autoencoders, feature ablation, and an NLA-inspired variance attribution method, the study reveals a negative correlation between feature selectivity and causal efficacy, and uncovers a substantial gap between detection robustness and causal robustness. While replicating the IOI circuit, it finds that only a subset of 15 highly selective features exhibit genuine causal effects, and model performance degrades significantly under distributional shift. A cost-aware monitoring strategy achieves 99.1% resource savings.
Domain experts face significant challenges in customizing Transformer architectures due to their monolithic design and high implementation complexity. Method: This paper introduces TransModular, the first plug-and-play modular framework for Transformers, enabling low-code composition and seamless integration of attention mechanisms, feed-forward networks, normalization layers, and positional encodings. It uniquely supports flexible hybridization of four distinct positional encoding strategies and pioneers deep integration of neural architecture search (NAS) into the modular Transformer design pipeline. Fully compatible with the PyTorch ecosystem, TransModular generalizes across domains, including genomic sequence modeling. Contribution/Results: Experiments demonstrate successful reproduction of the original Transformer under resource constraints, improved machine translation performance, and 95.2% accuracy in cell-type classification on single-cell gene expression data. TransModular substantially lowers the barrier to architecture customization and accelerates domain-driven AI innovation.
This work investigates how Transformers model hierarchical positional dependencies in tree-structured data. We propose a hierarchical filtering generative model that enables controlled modulation of positional dependencies across multiple scales, and integrate attention map analysis, hierarchical probing, and encoder-only training to systematically characterize the underlying modeling mechanisms. We make the first discovery that Transformer encoder layers progressively capture long-range hierarchical dependencies with depth: shallow layers encode local adjacency relations, while deeper layers specialize in global tree topology; moreover, each layer approximately reconstructs correlation patterns at a specific scale. Empirical results demonstrate that this architecture achieves performance approaching exact Bayesian inference on trees for root-node classification and masked language modeling tasks. Our findings establish a verifiable, scale-separated computational mechanism for interpretable AI, grounded in principled hierarchical representation learning.
This study addresses the challenge of accurately discovering time-lagged causal structures from multivariate time series without explicit causal constraints. The work proposes leveraging the autoregressive Transformer’s inherent sensitivity—specifically, the gradient of its predictions with respect to historical inputs—as a natural encoder of causal relationships, from which a causal graph is extracted via gradient attribution aggregation. It is rigorously demonstrated for the first time that standard Transformers possess intrinsic causal learning capabilities without requiring additional causal objectives or architectural modifications. Notably, causal discovery accuracy improves significantly as data heterogeneity increases. The method substantially outperforms existing causal discovery algorithms under challenging conditions, including nonlinear dynamics, long-range dependencies, and nonstationarity, with particularly pronounced gains in highly heterogeneous datasets.
This work addresses the limited interpretability of internal representations in large language models, a challenge exacerbated by existing post-hoc methods like sparse autoencoders (SAEs), which incur substantial computational overhead. The authors propose ParityTransformer, which integrates a Deep Parity Bottleneck (DPB) into a GPT-2–scale model to natively induce layer-wise, wide-dimensional, and sparse intermediate representations within the Transformer architecture for the first time. By combining a parameter-free algebraic dictionary—ensuring deterministic incoherence—with a multi-level mixture-of-experts structure to efficiently enforce sparsity, the approach significantly reduces both computational and memory costs. Experiments demonstrate that ParityTransformer matches SAE performance on sparse probing tasks while outperforming it in feature absorption, intervention steering, and fine-grained causal mediation, confirming that the model can directly leverage interpretable features for reasoning.
Existing analyses of Transformers often focus on individual attention heads or layers, failing to capture the model’s global behavior and lacking a unified representation. This work proposes TensorLens, the first framework that models the entire Transformer as an input-dependent linear operator. By introducing a high-order attention interaction tensor, TensorLens holistically encodes all components—including attention mechanisms, feed-forward networks, activation functions, normalization layers, and residual connections—into a single, coherent structure. The resulting representation is end-to-end, theoretically grounded, and highly expressive. Experimental results demonstrate that TensorLens captures richer features compared to existing attention aggregation methods, offering a powerful new tool for model interpretability and structural understanding.
Standard supervised training often struggles to learn effective query-key attention patterns in Transformer-based sequence classification tasks, particularly failing to induce a preference for neighboring positions. This work demonstrates through systematic ablation studies and simplified theoretical analysis that self-pretraining (SPT), driven by a masked reconstruction objective, enables the model to acquire such localized attention structures from random initialization, substantially improving optimization dynamics. Without relying on external data, SPT significantly outperforms purely supervised training on benchmarks such as the Long-Range Arena. The performance gains are primarily attributed to the model’s enhanced ability to learn interactions among nearby tokens, highlighting proximity-aware attention as a key mechanism underlying SPT’s effectiveness.
This work addresses the phenomenon of "attention concentration" in Transformer models, wherein a disproportionate amount of attention is allocated to uninformative or specific tokens, thereby undermining model interpretability, destabilizing training and inference, and exacerbating hallucination issues. The paper presents the first comprehensive survey of this phenomenon, introducing a three-dimensional classification framework—comprising foundational utilization, mechanistic explanation, and mitigation strategies—to systematically organize the evolving research landscape. By synthesizing recent findings on anomalous attention behaviors, the study constructs a structured knowledge base that clarifies core concepts and key challenges. It offers both theoretical insights and practical pathways for understanding and mitigating attention concentration, and further supports community advancement by releasing a curated list of relevant publications.