Importance-Aware Feature Sparsification for Wireless Split Learning

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the communication bottleneck caused by high-dimensional feature transmission and the accuracy degradation under non-independent and identically distributed (non-IID) data in wireless split learning. We propose an importance-aware, class-balanced sparsification method. Specifically, the server computes feature importance based on gradients and broadcasts it to clients, which reuse this vector to retain critical channels with zero additional overhead. Furthermore, Grad-CAM is incorporated to enable task-aware, class-balanced evaluation for mitigating label skew. We theoretically derive a non-asymptotic convergence bound that supports parallel splitting and Transformer architectures. Experimental results demonstrate that the proposed method significantly outperforms baselines under strongly non-IID scenarios, effectively reducing communication costs while maintaining high accuracy.
📝 Abstract
Wireless split learning (SL) reduces on-device computation by offloading upper layers to a server, yet transmitting high-dimensional intermediate features at each iteration remains a major communication bottleneck. Existing methods select features at the client side using task-agnostic criteria such as magnitude, statistics, or clustering, which increases client-side processing and often degrades accuracy under non-independent and identically distributed (non-i.i.d.) client data. We propose importance-aware class-balanced sparsification (ICS), a lightweight approach in which the server ranks feature channels using Grad-CAM-based scores obtained from the true-class logit during backpropagation. The per-class scores are aggregated into a class-balanced, label-agnostic importance vector that mitigates head-class bias under label skew, and each client reuses this vector in the next round to retain the top-$N$ feature channels, incurring no additional client-side forward or backward passes. We further derive a non-asymptotic convergence bound that isolates the sparsification-induced error and characterizes how the sparsification ratio and mini-batch size jointly affect convergence under a fixed communication budget, and we analyze the communication and computational overhead of ICS against representative baselines. Beyond sequential CNN-based SL, we extend ICS to parallel split learning and to transformer-based models. Experiments show that ICS consistently outperforms the baselines, with larger gains under severe non-i.i.d. partitions.
Problem

Research questions and friction points this paper is trying to address.

Wireless Split Learning
Communication Bottleneck
Feature Sparsification
Non-IID Data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Wireless Split Learning
Feature Sparsification
Grad-CAM
Non-IID Data
Convergence Bound
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.