DanLing NestedTensor: Composable Multi-Ragged Tensors for Deep Learning

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the computational waste caused by padding variable-length inputs in deep learning and the poor composability of packing operations. We propose a PyTorch-native multi-sawtooth tensor abstraction that internalizes variable-length structures as intrinsic tensor properties, enabling broadcasting, transformations, and reductions to automatically preserve logical axes and sample boundaries. By binding packed values to partitions and logical dimensions, this approach achieves seamless end-to-end support spanning automatic differentiation to compiled execution. Experimental results demonstrate that on A100 GPUs, the proposed method accelerates BERT by 2.74× to 3.39× and FCN by 1.97×, while reducing the peak memory consumption of Pairformer by approximately 86%.
📝 Abstract
Variable-size inputs are common in deep learning, but dense batching allocates a shared envelope and spends computation on padding. The cost multiplies across varying axes: an explicit pair state allocates $BN_{\max}^2$ positions instead of $\sum_i N_i^2$. Packing removes that waste, but composing packed operations still requires the logical axes and sample boundaries a flat buffer no longer exposes. We present DanLing NestedTensor, a PyTorch tensor abstraction that makes multi-ragged structure a property of the tensor itself. Packed values carry tensor-backed partitions and logical dimension order, so broadcasting creates ragged axes, feature transformations retain them, and reductions consume them. The same representation carries through autograd and both eager and compiled execution. On an A100, the geometric-mean speedup over same-mode padding is 2.74$\times$ eager and 3.39$\times$ compiled across four BERT scales, and 1.97$\times$ eager across four FCN backbones. A four-block Pairformer-style workload runs 2.40-4.32$\times$ faster than a padded reference using native PyTorch kernels across square length regimes in eager execution, with peak allocation falling from 38.08 to 5.41 GiB on its high-variation batch. The tensor interface lets model code built from its supported operators compose efficient variable-size computation without managing offsets at any call site. Code will be released publicly upon publication.
Problem

Research questions and friction points this paper is trying to address.

variable-size inputs
multi-ragged tensors
padding waste
packed operations
deep learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

NestedTensor
Multi-Ragged Tensors
Variable-Size Computation
Memory Efficiency
PyTorch Abstraction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhiyuan Chen
DanLing Team