Score
Designs, implements, and evaluates sparse data encodings and representation-selection mechanisms that monitor runtime sparsity and dynamically switch formats or encodings; includes the policies and system support to choose representations and adapt algorithms or communication patterns to minimize storage, computation, and communication overhead.
Existing sparse matrix optimization methods often coarsely classify structured sparsity (e.g., clustered non-zeros) as either fully dense or fully sparse, leading to redundant zero computations in fixed-block formats (e.g., BCSR) or substantial overhead in variable-block approaches due to unknown loop bounds at compile time. This work proposes a region-aware, multi-stage compilation framework that automatically identifies high-benefit variable-size blocks, statically infers dynamic loop bounds, and generates customized vectorized code—balancing efficiency and adaptability. Key techniques include sparse partition analysis, domain-specific code generation, loop vectorization, and compile-time scheduling specialization. Evaluated on the SuiteSparse dataset, our approach achieves 1.07×, 2.73×, and 1.9× higher single-threaded SpMV performance over Intel MKL, CSR5, and Partial Strided Codelets, respectively; parallel scalability further enhances throughput.
The computational and energy overhead of deep learning inference is increasingly prohibitive, yet sparsity—a key optimization avenue—remains underutilized in production systems. Method: Targeting performance engineers, this work systematically surveys structured and unstructured sparsity exploitable in DNN inference and proposes an end-to-end engineering methodology—from sparse model representation to efficient sparse kernels (SpMM/SDDMM). We implement and benchmark multiple sparse computation schemes on CPU and GPU platforms, integrating support for mainstream frameworks, toolchains, and datasets. Contribution/Results: We present the first production-grade sparse inference reference framework encompassing hardware adaptation, kernel optimization, and deployment validation. Experiments demonstrate 2–5× inference speedup and substantial energy efficiency gains across representative models, establishing a reproducible, scalable practical paradigm for industrial deployment of sparse deep learning.
This work addresses the loss of dimensional semantics in traditional compilation, where type systems discard such information prior to code generation, leading to ad hoc numerical representations and memory management that struggle to balance efficiency, determinism, and verifiability. To overcome this, the paper introduces a Dimensional Type System (DTS) that propagates dimensional annotations as compile-time metadata throughout MLIR’s multi-stage lowering pipeline. DTS enables joint optimization of representation selection and deterministic memory management within a unified semantic graph. Grounded in finitely generated Abelian group constraints, the system supports polynomial-time decidable, complete, and principal type inference. A coeffect system unifies escape analysis and memory allocation, while also revealing the closure of dimensional algebra under automatic differentiation. Experiments demonstrate that DTS enables design-time verifiable memory strategies, representation fidelity, cache locality estimation, and coeffect-based AD verification, significantly enhancing compilation reliability and performance in resource-constrained settings.
To address the high memory overhead and computational inefficiency of conventional sparse matrix formats (e.g., CSC, COO) when representing highly redundant sparse matrices—characterized by a small number of distinct non-zero values—this paper proposes two novel column-compressed formats: VCSC and IVCSC. VCSC integrates intra-column value deduplication and run-length counting into the CSC layout, enabling compact storage. IVCSC further enhances locality by introducing variable-length delta encoding and segmented metadata. Both formats preserve O(1) random access complexity. Evaluated on five real-world datasets, VCSC achieves an average compression ratio 1.8× higher than CSR/CSC, while IVCSC attains 2.5×, significantly outperforming state-of-the-art alternatives.
Existing sparse and structured tensor computation frameworks suffer from rigid control-flow abstractions and fragmented structural support, hindering efficient exploitation of intrinsic properties such as sparsity, symmetry, and blocking. This paper introduces Finch—a novel domain-specific language that unifies arbitrary control flow (e.g., loops, conditionals, breaks) with diverse tensor structures (sparse, symmetric, blocked) via a joint control-flow–data-structure representation, enabling automatic structural specialization. Finch integrates structure-aware code generation, sparse tensor algebra compilation (e.g., SpMV, SpGEMM), and metadata-driven runtime execution. Evaluated on sparse matrix multiplication, image processing, and graph analytics, Finch achieves substantial performance gains over state-of-the-art frameworks. It significantly improves utilization of structural zeros, redundant values, and non-zero clusters—demonstrating superior efficiency in leveraging inherent tensor structure.
This work addresses the challenge of unifying diverse structured sparsity patterns for efficient model compression and acceleration. The authors propose S³, an algebraic framework that formally integrates three core components—View (tensor reshaping), Block (atomic pruning units), and Scope (sparsity decision range)—to express a wide spectrum of sparsity patterns, ranging from fine-grained N:M sparsity to coarse-grained channel pruning, within a single formalism. Notably, S³ enables cross-tensor collaborative sparsification. Building upon this framework, the authors incorporate Optimal Brain Damage and Surgeon algorithms to develop structured variants of OBS/OBD. These methods significantly outperform current state-of-the-art second-order heuristic approaches in terms of output reconstruction accuracy.
本文针对稀疏矩阵向量乘法(SpMV)性能问题,提出了一种新的分层CSR格式(HCSR),并通过在RISC-V处理器上使用RVV 1.0内联函数实现,显著提升了计算速度。
This work addresses the performance degradation in dynamic sparse attention mechanisms, where per-token selection of key-value cache entries leads to fragmented working sets and poor cache locality, thereby increasing last-level cache (LLC) misses and reducing decoding throughput. To mitigate this issue, the authors propose a service-oriented cache architecture featuring a lightweight indexer that analyzes key-value access patterns, coupled with an LLC reservation mechanism and a fine-grained, token-level LRU replacement policy. This design effectively alleviates cache fragmentation and improves data locality. Experimental evaluation across multiple open-source backbone models demonstrates that the proposed approach significantly reduces LLC stall misses and enhances online inference throughput for dynamic sparse attention models.
本文提出了一种新的代码生成策略,通过扩展TACO编译器的表示方法来处理稀疏张量收缩中的冲突数据布局问题,避免了显式转置,从而提高了计算效率。
This study addresses the computational and memory bottlenecks of dynamic programming in the exact optimization of synonymous coding sequences. To overcome these limitations, this work proposes a multi-loop recursive acceleration algorithm based on candidate sparsification, enabling joint optimization of folding energy and codon usage over weighted codon automata. We rigorously prove the equivalence between the sparse recurrence and its dense counterpart, and introduce an endpoint ownership mechanism to support lock-free parallel construction of candidate sets. Integrated with the Turner 2004 thermodynamic model, our approach reduces the candidate retention rate for natural proteins to approximately 3%. Consequently, it achieves over a 20-fold speedup and nearly a 28-fold reduction in memory consumption compared to conventional dense algorithms.