GroupMask: Layer-Adaptive Group-wise Sparsity for Semi-Structured LLM Pruning

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance bottleneck in semi-structured pruning caused by fixed sparsity ratios across layers, proposing GroupMask, a framework that for the first time validates the effectiveness of layer-wise adaptive allocation under group-level sparsity. The method employs a lightweight hypernetwork to generate selectors and trains with frozen weights by integrating Gumbel-Sigmoid parameterization, the straight-through estimator (STE), self-distillation, and budget regularization, thereby achieving efficient large language model compression while preserving regular structures. Experimental results demonstrate that LLaMA-2-7B pruned at 50% sparsity attains a perplexity of 8.30 and a zero-shot accuracy of 0.496, significantly outperforming uniform allocation strategies and multiple baseline methods.
📝 Abstract
Semi-structured pruning compresses large language models (LLMs) while keeping a regular sparse structure, but the prevailing N:M pattern fixes the same local sparsity ratio in every layer. Layer-adaptive sparsity allocation improves unstructured pruning, yet it has been reported to be less effective under N:M sparsity, leaving open whether adaptive allocation is of limited value for semi-structured pruning in general or only under the fine-grained N:M pattern. We examine this question with group-level sparsity, which partitions each weight matrix into regular groups, retains or prunes each group as a whole, and allows each layer's sparsity ratio to vary under a global budget. We propose GroupMask, which generates the group selectors of all layers with a lightweight hypernetwork, relaxes them with a Gumbel-Sigmoid parameterization and a straight-through estimator, and learns them through sparsity-budget regularization and self-distillation while keeping the pretrained weights frozen. On LLaMA-2-7B at 50% sparsity with the same $1\times256$ group size, learned layer-adaptive allocation reduces WikiText-2 perplexity from 10.02 to 8.30 and raises the average zero-shot accuracy from 0.455 to 0.496 relative to a uniform per-layer ratio. GroupMask obtains the lowest WikiText-2 perplexity on LLaMA-2-7B and the highest average zero-shot accuracy with Alpaca calibration among the evaluated baselines on five LLaMA and Qwen models. Our code is available at https://github.com/ZhengaoLi/GroupMask.
Problem

Research questions and friction points this paper is trying to address.

Semi-structured pruning
Large language models
Layer-adaptive sparsity
Model compression
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semi-Structured Pruning
Layer-Adaptive Sparsity
Hypernetwork
Gumbel-Sigmoid
Self-Distillation