factorized dense routing

Designs and implements routing and mixing modules that approximate fully-global dense interactions across tensor dimensions by factorizing dense routing into hierarchical or separable contractions. Builds 2D-to-3D routing mechanisms and hierarchical tensor-contraction schemes that provide a global receptive field while reducing compute and memory compared with naïve dense routing, targeting sub-quadratic computational cost.

factorizeddenserouting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing sketching-based methods for tensor network contraction are limited to acyclic structures and struggle with general topologies containing cycles, while exact contraction suffers from prohibitive computational complexity. This work proposes two novel sketching-based approximate contraction approaches: the first is a sketching algorithm capable of handling arbitrary tensor network topologies, including those with cycles, marking the first such method in the literature; the second is a new technique tailored for acyclic networks, whose time and space complexity scale only polynomially with the number of contraction steps. By overcoming the structural limitations of conventional sketching techniques, this study significantly reduces computational and memory costs for acyclic networks and provides an efficient, general-purpose framework for approximate contraction of large-scale tensor networks.

approximationcomputational complexitycyclic tensor networks

This work addresses the high computational and communication overhead incurred when multiple users compute nonlinearly separable functions in distributed environments. To overcome the limitations of conventional approaches that rely on linear separability assumptions, the paper proposes an efficient computation framework based on sparse tensor representations. By introducing a fixed-support SVD-based sparse tensor decomposition method combined with a multidimensional sub-tensor partitioning strategy, the framework jointly optimizes task allocation and communication patterns. This integrated approach significantly reduces system resource consumption and achieves substantial improvements over state-of-the-art methods in both computational efficiency and communication cost.

Computation-communication tradeoffDistributed computingNon-linearly separable functions

This study investigates the impact of decoder scaling strategies—specifically depth versus width expansion—on the performance of neural routing solvers. Building upon an encoder-decoder architecture, the authors systematically construct twelve models ranging from 1M to 150M parameters and evaluate their efficacy on vehicle routing problems across three dimensions: parameter efficiency, data efficiency, and computational efficiency. The findings reveal that model performance cannot be reliably predicted by parameter count alone, with depth expansion consistently outperforming width expansion. Based on these insights, the work proposes a “depth-first” design principle for decoders, which significantly enhances both solution quality and resource utilization efficiency in neural combinatorial optimization.

decoder scalingmodel depthmodel width

Attention Is All You Need For Mixture-of-Depths Routing

Dec 30, 2024
AG
Advait Gadhikar
🏛️ Bosch Center for Artificial Intelligence | CISPA Helmholtz Center for Information Security | University of Stuttgart | University of Tübingen

Traditional Mixture-of-Depth (MoD) models suffer from training instability, high computational complexity, and deployment overhead due to dedicated routing layers. To address these issues, this paper proposes a parameter-free, attention-driven dynamic depth routing mechanism. Instead of introducing auxiliary routing networks, our method reuses the self-attention maps from the preceding layer as token-wise routing signals for the current layer—enabling zero-parameter overhead and seamless plug-and-play integration with pretrained Vision Transformers (ViTs). The core innovation lies in leveraging the inherent self-attention mechanism for dynamic, input-adaptive computation allocation, thereby balancing inference efficiency and model capacity. On ImageNet, our approach achieves up to 2% higher top-1 accuracy than standard MoD baselines and outperforms ViT counterparts with comparable FLOPs. Furthermore, in transfer learning scenarios, it accelerates convergence by up to 2×.

Deep LearningModel OptimizationResource Allocation

This work addresses a critical limitation in existing mixture-of-LoRA models, where imbalanced routing weights often lead to the activation of only a few adapters, thereby constraining model expressiveness. To overcome this, the authors propose ReMix, a novel approach that eliminates learnable routing weights and instead introduces a non-learnable, balanced routing mechanism. By leveraging the REINFORCE leave-one-out (RLOO) gradient estimator from reinforcement learning, ReMix constructs an unbiased, non-differentiable routing policy that ensures equal contribution from all activated LoRA modules. Under the constraint of identical active parameter counts, ReMix significantly outperforms current state-of-the-art parameter-efficient fine-tuning methods, effectively breaking through the performance bottleneck inherent in conventional mixture-of-LoRA architectures.

expressive powerlow-rank adaptersMixture-of-LoRAs

Latest Papers

What's happening recently
View more

This study addresses the challenge that agent performance is constrained by model–framework compatibility, where training data cannot exhaustively cover the entire combinatorial space. To this end, we propose a joint routing method that leverages CP tensor decomposition to model such compatibility. By integrating decoupled representation learning with interaction term co-optimization, our approach accurately infers unobserved model–framework combinations from limited samples. Experimental results demonstrate that the proposed method surpasses baseline approaches in routing accuracy by 7.3% and achieves a substantial improvement of 15.8 percentage points on unobserved combinations. These findings indicate that our framework significantly enhances both the cost-effectiveness and generalization capability of agent selection.

Agentic SystemsCompatibilityModel-Harness Routing

This study addresses the computational waste caused by padding variable-length inputs in deep learning and the poor composability of packing operations. We propose a PyTorch-native multi-sawtooth tensor abstraction that internalizes variable-length structures as intrinsic tensor properties, enabling broadcasting, transformations, and reductions to automatically preserve logical axes and sample boundaries. By binding packed values to partitions and logical dimensions, this approach achieves seamless end-to-end support spanning automatic differentiation to compiled execution. Experimental results demonstrate that on A100 GPUs, the proposed method accelerates BERT by 2.74× to 3.39× and FCN by 1.97×, while reducing the peak memory consumption of Pairformer by approximately 86%.

deep learningmulti-ragged tensorspacked operations

This study addresses the limitations of sparse depth scaling and capacity coupling in Mixture-of-Depths (MoD) architectures by proposing X-MoD. The method decouples token sparsity from anchor stride, enabling total parameter growth while maintaining constant activation capacity. It formulates sparse depth routing as a conditional architecture design problem for the first time, establishing an interpretable scaling law to decompose performance gains. Furthermore, dense anchors, variance-scaled gating, and depth-wise token balancing techniques are introduced to optimize training. Pretraining experiments demonstrate that this framework accurately predicts validation loss across diverse configurations and significantly outperforms both Dense and MoE baselines.

Conditional ComputationMixture-of-DepthsScaling Laws

Hot Scholars

SW

Shuo Wang

University of Science and Technology of China
Computer VisionMultimedia
SZ

Shaofeng Zhang

Southern University of Science and Technology
Learn to Optimize
XH

Xinting Hu

Max Planck Institute for Informatics
Multimodal ReasoningContinual LearningSemi-Supervised Learning
ML

Minji Lee

Assistant Professor, The Catholic University of Korea
Machine LearningNeuroscienceBrain-Computer Interface