IrekoGPT: Turning Structured Pruning into Post-Hoc Slimmable LLMs

πŸ“… 2026-09-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the inability to dynamically adjust the width of pretrained large language models during inference. We propose a post-hoc slimming method built upon SliceGPT that reformulates structured pruning as the construction of nested subnetworks via retained projection matrices, enabling dynamic exposure of varying model widths at inference without re-pruning. Furthermore, we integrate hierarchical multi-scale calibration with gradient-free ridge regression to correct error propagation in downstream layers. Experimental results demonstrate that our approach significantly outperforms naive PCA-based slimming baselines on Llama and Qwen architectures, yielding particularly pronounced performance improvements under high compression ratios.
πŸ“ Abstract
We introduce IrekoGPT, a post-hoc method for converting pretrained LLMs into slimmable models whose width can be adjusted at inference time. Building on SliceGPT, we retain its projection matrices without pruning them, allowing a single model to expose nested subnetworks at different widths. We improve robustness by calibrating each layer across multiple compression ratios, and correct downstream linear layers through gradient-free ridge regression. Across Llama and Qwen models, preliminary results show improvements over naive PCA-based slimming, with the largest gains at high compression. Code is available at https://github.com/aimagelab/IrekoGPT
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Structured Pruning
Slimmable Networks
Post-Hoc Compression
Model Width Adjustment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Post-Hoc Slimmable LLMs
Structured Pruning
Projection Matrices
Gradient-Free Ridge Regression
Nested Subnetworks
πŸ’Ό Related Jobs
No related jobs found.