π€ AI Summary
This study addresses the inability to dynamically adjust the width of pretrained large language models during inference. We propose a post-hoc slimming method built upon SliceGPT that reformulates structured pruning as the construction of nested subnetworks via retained projection matrices, enabling dynamic exposure of varying model widths at inference without re-pruning. Furthermore, we integrate hierarchical multi-scale calibration with gradient-free ridge regression to correct error propagation in downstream layers. Experimental results demonstrate that our approach significantly outperforms naive PCA-based slimming baselines on Llama and Qwen architectures, yielding particularly pronounced performance improvements under high compression ratios.
π Abstract
We introduce IrekoGPT, a post-hoc method for converting pretrained LLMs into slimmable models whose width can be adjusted at inference time. Building on SliceGPT, we retain its projection matrices without pruning them, allowing a single model to expose nested subnetworks at different widths. We improve robustness by calibrating each layer across multiple compression ratios, and correct downstream linear layers through gradient-free ridge regression. Across Llama and Qwen models, preliminary results show improvements over naive PCA-based slimming, with the largest gains at high compression. Code is available at https://github.com/aimagelab/IrekoGPT