Score
Designs, implements, and evaluates encoder components and end-to-end encoder–decoder architectures, specifying layer layouts, interfaces to decoders, and criteria for correctness and quality. Builds and optimizes encoder inference pipelines—including runtime performance tuning, caching strategies, and specialized encoder types such as visual encoders and vocoders—and measures their latency, throughput, and resource use during integration with decoders.
To address high first-token latency and low throughput of small models (≤1B) on edge devices, this work revisits the encoder-decoder architecture for its efficiency advantages under resource constraints. We propose a task-adaptive knowledge distillation framework that enables lightweight encoder-decoder student models to effectively inherit capabilities from large decoder-only teachers while preserving their intrinsic properties—single-pass input encoding and decoupled understanding and generation. This is the first systematic validation of such architectures on asymmetric-sequence tasks. Integrated with RoPE, cross-platform optimization (GPU/CPU/NPU), and fused visual encoders, our approach achieves a 47% reduction in first-token latency, a 4.7× throughput improvement, and an average +6-point gain in task-specific performance—particularly pronounced in tasks with large input-output distribution mismatch.
Existing efforts to encode algorithms directly into Transformer parameters are fragmented and lack systematicity, hindering both learning and reuse. Method: This paper introduces, for the first time, a unified “algorithm encoding recipe” framework that systematically integrates arithmetic implementations in feed-forward layers with dynamic data routing via self-attention, yielding a modular and composable Transformer construction methodology. Contribution: It bridges a critical gap between interpretable modeling and controllable architectural design of Transformers. The framework enables rigorous computational complexity analysis, formally verifiable architecture specification, and mechanistic interpretation of internal operations. By significantly lowering the barrier to algorithmic encoding—providing accessible onboarding for novices and structured, reusable building blocks for experts—it advances research and deployment of controllable, interpretable Transformers.
This work clarifies the functional boundaries and necessity of core Transformer components—tokenization, embedding/un-embedding, masking, positional encoding, and padding—addressing widespread conceptual ambiguity in their mechanistic roles. Targeting ML engineers, we propose an incremental, invertibility-based analytical framework: using binary (0/1) sequences as probes, we systematically introduce each component via a “zero-one construction” and empirically validate its irreplaceability in the encode-decode pipeline. Implemented lightweightly in PyTorch, our framework supports manual attention matrix construction, explicit positional embedding injection, and interpretable mask design. Experiments demonstrate significantly improved conceptual accuracy among learners. Notably, we provide the first empirical verification that, in the absence of self-attention, positional encoding combined with padding alone suffices for basic length-aware tasks.
This work addresses critical efficiency bottlenecks in training and inference of mainstream encoder-decoder large speech models (e.g., Whisper, Seamless), identifying two root causes: padding waste induced by fixed-length sequence sampling and computational redundancy from autoregressive decoding. To mitigate these, we propose a dynamic-length batch sampling strategy that reduces padding overhead by over 50%, and introduce the first decoder-to-encoder parameter reallocation architecture—achieving a 3× inference speedup without accuracy degradation. Leveraging GPU performance profiling and RTFx (inverse real-time factor) evaluation, our approach enables a 5× increase in effective batch size, cutting training time by 50% under fixed compute budget or reducing GPU requirements to one-quarter. All code and models are publicly released.
This study investigates the impact of decoder scaling strategies—specifically depth versus width expansion—on the performance of neural routing solvers. Building upon an encoder-decoder architecture, the authors systematically construct twelve models ranging from 1M to 150M parameters and evaluate their efficacy on vehicle routing problems across three dimensions: parameter efficiency, data efficiency, and computational efficiency. The findings reveal that model performance cannot be reliably predicted by parameter count alone, with depth expansion consistently outperforming width expansion. Based on these insights, the work proposes a “depth-first” design principle for decoders, which significantly enhances both solution quality and resource utilization efficiency in neural combinatorial optimization.
This work addresses the high computational cost of conventional neural architecture search (NAS) performance predictors, which often rely on expensive fine-tuning or intricate architecture representations. The authors propose Code-Oriented Language Model Embeddings (COLE), a method that directly uses raw PyTorch class definition code as input and leverages a frozen off-the-shelf language model to extract architecture embeddings. Coupled with a lightweight regression head, COLE constructs an efficient performance predictor without requiring NAS-specific fine-tuning. Evaluated on the NAS-Bench-201 benchmark, COLE achieves within 1% of the optimal architecture’s accuracy while reducing the evaluation budget by 34% compared to path-based encoding. Furthermore, experiments on CIFAR-100 demonstrate its strong generalization capability and superior search efficiency.
Existing approaches struggle to elucidate the interaction mechanisms and differential contributions of individual visual encoders in multi-encoder vision-language models prior to training, hindering efficient architectural design. This work addresses this gap by conducting from-scratch training of 31 encoder subsets under a unified framework on the Cambrian-1 benchmark. We propose a Capacity–Necessity dual-axis decomposition framework, revealing that an encoder’s standalone performance (Capacity) does not equate to its necessity within joint training. Our analysis demonstrates that optimal ensembles are not merely combinations of high-Capacity encoders and introduces the effective rank of pre-projection layers as a novel predictor of collaborative encoder performance. Experiments show that combining a single high-Capacity anchor encoder with one complementary encoder nearly matches the performance of the full five-encoder model, with diminishing returns from additional encoders.
This study addresses the challenges of catastrophic forgetting in supervised fine-tuning (SFT) and the unclear internal mechanisms underlying instruction-following capabilities in language models. By integrating information-theoretic analysis, geometric metrics, and optimization trajectory inspection across models ranging from 1B to 32B parameters, the work reveals—for the first time—that instruction alignment exhibits architectural locality: representations in intermediate layers (20%–80% depth) remain stable, while those in the final layers are highly sensitive. Building on this insight, the authors propose Mid-Block Efficient Tuning, which fine-tunes only critical intermediate layers. This approach achieves up to a 10.2% improvement over standard LoRA on GSM8K (using OLMo2-7B) while substantially reducing the number of trainable parameters.
This work addresses the efficiency bottlenecks in large vision-language model inference, which are dominated by high-resolution visual features, quadratic-complexity attention mechanisms, and memory bandwidth constraints—collectively forming a “vision-token-dominated” bottleneck. The study proposes the first end-to-end optimization framework that decouples inference into three distinct stages: encoding, prefilling, and decoding. It systematically analyzes stage-specific bottlenecks and their inter-stage couplings, advocating for a balanced trade-off between visual fidelity and computational efficiency through strategies such as information density modulation, long-context attention management, and memory-bound mitigation. Key contributions include a structured taxonomy of optimization techniques, an open-source and maintainable literature repository, insights into combinatorial optimization potential, and a roadmap highlighting four promising directions: function-aware hybrid compression, modality-aware decoding, streaming state management, and hardware-software co-designed staged serving.