Communication-Aware Model Distributed Inference via Latent Representation Compression

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the trade-off between model accuracy and communication cost, along with stringent quality-of-service (QoS) constraints, in distributed edge inference. We propose a joint optimization framework based on latent representation compression. By integrating convex optimization with a water-filling strategy for resource allocation, we further design a stochastic dual descent algorithm grounded in Lyapunov analysis. This algorithm relies solely on causal channel state information to achieve a bounded optimality gap while satisfying long-term delay constraints. Both simulations and experiments conducted on real-world edge devices demonstrate that the proposed framework exhibits strong robustness and scalability in dynamic environments.
📝 Abstract
We study optimization of distributed model inference over resource-constrained edge resources. We propose a framework that optimizes the trade-off between model accuracy and communication costs by controlling latent representation compression to meet strict Quality of Service (QoS) throughput targets. For settings with known channel state information (CSI), we derive a closed-form optimal solution for single tasks and reduce the multi-task problem to a convex optimization program characterized by a per-link water-filling strategy. We extend these to handle unpredictable environments via a stochastic dual descent algorithm that relies only on causal channel estimates. We provide Lyapunov-based proofs demonstrating that our approach strictly satisfies long-term delay constraints while achieving a bounded optimality gap. Our results offer a robust, scalable blueprint for maximizing the performance of pipelined AI tasks in dynamic, resource-constrained distributed systems. We verify the effectiveness of our proposed framework through simulations and experiments with real edge devices.
Problem

Research questions and friction points this paper is trying to address.

Distributed Inference
Edge Computing
Latent Representation Compression
Communication Cost
Quality of Service
Innovation

Methods, ideas, or system contributions that make the work stand out.

Distributed Inference
Latent Representation Compression
Convex Optimization
Stochastic Dual Descent
Edge Computing
🔎 Similar Papers
No similar papers found.