Score
Designs, builds, and evaluates compression and communication systems that encode raw data, intermediate features, or compact metadata to minimize bandwidth or storage while preserving semantic, task‑relevant information for downstream inference or control. This includes creating redundancy‑aware feature encoders/decoders that allocate bits by task importance, remove spatial and temporal redundancy (e.g., across channels or sequential frames), and analyze tradeoffs between compression rate, latency, and task performance.
To address high redundancy in multimodal network traffic and bandwidth waste caused by heterogeneous, modality-specific compressors in resource-constrained IoT scenarios, this paper proposes the first byte-level, cross-modal, single-model universal compression framework. The method uniformly models video, audio, images, and text as byte sequences; employs a Transformer to predict byte-level probability distributions; generates sparse-rank representations; and integrates adaptive lossless entropy coding for end-to-end compression. By abandoning the conventional paradigm of deploying separate compressors per modality, the framework significantly reduces system complexity and data entropy. It supports scalable model sizes to accommodate heterogeneous edge–cloud devices. Experimental evaluations on IoT nodes and servers demonstrate over 50% compression ratio improvement, more than 50% reduction in bandwidth consumption, and substantial throughput gains.
Existing speech codecs do not explicitly disentangle semantic hierarchies, making it challenging to simultaneously preserve perceptual quality and downstream task performance at ultra-low bitrates (e.g., <1.5 kbps). To address this, we propose the first decoupled framework for semantic speech compression, introducing hierarchical semantic representations—explicitly separating and differentially encoding phonetic, prosodic, emotional, and speaker-related features—derived from generative speech models into the codec architecture. Leveraging a semantic communication paradigm and multi-granularity reconstruction, our method achieves or surpasses state-of-the-art performance of codecs such as EnCodec on automatic speech recognition, emotion analysis, and speaker verification, while operating at 2–4× lower bitrates. Crucially, intelligibility and naturalness are preserved. This work establishes a novel paradigm for ultra-low-bitrate semantic speech communication.
This work addresses lossy image compression for multi-task scenarios, proposing a framework that jointly optimizes reconstruction fidelity, perceptual quality, and classification accuracy. We establish, for the first time, an information-theoretic rate–distortion–classification (RDC/RPC) triadic model and derive its closed-form solution. Theoretically, we prove that under RPC constraints, classification performance and perceptual fidelity are not fundamentally trade-offs, and reveal the critical regulatory role of source noise in task-oriented compression. Leveraging the information bottleneck principle, we unify generative and discriminative objectives and derive optimal rate bounds for binary and Gaussian sources. Experiments demonstrate that our deep compression network achieves Pareto-optimality across PSNR, LPIPS, and classification accuracy—providing both theoretical foundations and a practical paradigm for task-driven compression.
Existing semantic communication research overemphasizes transmission fidelity while neglecting the fundamental fact that AI task performance is determined by model training—leading to a “fidelity-constraint paradox.” Method: This paper proposes a goal-oriented semantic communication paradigm that models and proactively estimates how communication-induced distortions—introduced by semantic compression—affect AI task accuracy. It innovatively extends rate-distortion theory to semantic communication by quantifying distributional shifts between original and distorted data, thereby establishing a distortion–accuracy mapping. The framework integrates distribution-shift modeling, semantics-aware compression, and joint communication-computation optimization under network constraints (e.g., bandwidth, latency). Contribution/Results: Experiments demonstrate that the proposed method significantly outperforms fidelity-driven baselines, improving task accuracy by 12.6%–28.3% under identical resource budgets, while guaranteeing empirically validated accuracy for downstream AI tasks.
This work addresses the instability of existing context compression methods in long-context scenarios, which often overlook the impact of data distribution on compression efficacy. From a data-centric perspective, the study systematically investigates how the input data and the intrinsic knowledge distribution of large language models jointly influence compression quality. The authors propose a semantic integrity evaluation framework based on autoencoders and introduce an input entropy metric under a frozen decoder setting. They reveal, for the first time, a negative correlation between input entropy and compression quality, and demonstrate that distributional discrepancies between encoder and decoder inputs significantly diminish compression gains. Building on these insights, they formulate targeted data-side optimization strategies, offering both theoretical grounding and practical guidance for improving context compression performance.
General-purpose compression algorithms often struggle to simultaneously achieve high compression ratios and high throughput with low overhead, whereas specialized compressors, while offering superior performance, incur high development and maintenance costs and suffer from limited applicability. This work proposes a novel “graph-based” compression framework that, for the first time, models the compression process as a modular composition of encoders and decoders represented by a directed acyclic graph. By integrating a self-describing format with a universal decoder (OpenZL), the framework unifies the generality of generic methods with the performance of specialized ones. The approach substantially reduces the cost of developing and deploying domain-specific compressors, outperforming mainstream general-purpose compressors in both compression ratio and speed across multiple real-world datasets. It remains competitive with deep learning–based methods while operating orders of magnitude faster, and internal adoption at Meta has reduced development cycles from months to days.
This work addresses the challenges of enhancing robustness against channel noise, reducing resource consumption, and controlling encoding complexity in semantic and task-oriented communication. It proposes an end-to-end semantic coding architecture that integrates manifold-constrained hyperconnectivity (mHC) with an entropy bottleneck (EB). By leveraging multi-residual streams and a doubly stochastic mixing matrix, the method enriches representation diversity and improves training stability without increasing model parameters or computational overhead. The entropy bottleneck enables explicit rate control while preserving optimal code length. Experimental results demonstrate that, under AWGN, Rayleigh/Rician fading channels, and imperfect channel state information, the proposed approach significantly outperforms baseline models based on residual networks and unconstrained hyperconnectivity, achieving superior semantic fidelity, task performance, and convergence stability at the same number of channel uses.
This study addresses the dual-fidelity rate–distortion trade-off for structured semantic sources in task-oriented communication. By restricting to deterministic intra-class encoding mappings and integrating conditional mean decoding with the Lloyd–Max stationarity conditions, the paper derives fundamental rate–distortion characteristics. It innovatively constructs the first dual-purpose feasibility band tailored to such semantic sources, revealing upper and lower bounds on achievable communication rates under non-single-letter fidelity criteria and their dependence on partition cardinality, thereby extending the Shannon–Kolmogorov (SK) framework. Theoretically, it proves that the intra-class SK rate is no less than the maximum of the corresponding Shannon rate–distortion function. In an aggregation validation scenario, a feasibility band of width $\log_2(K_{\max}/K_{\min})$ bits is obtained, with empirical validation provided through a smart grid economic dispatch case study.
To address the bandwidth bottleneck in AI model split inference caused by intermediate feature transmission, this paper proposes a lightweight, lossless feature compression method based on Z-score normalization. The core innovation lies in the first integration of Z-score standardization into a feature compression framework to explicitly preserve global statistical properties—namely, mean and variance—enabling an end-to-end differentiable encoding architecture. Unlike the conventional scaling scheme in the MPEG FCM standard draft, our approach yields a more compact and hardware-efficient implementation. Extensive evaluation across multiple vision tasks demonstrates an average bitrate reduction of 17.09%, with up to 65.69% savings in object tracking, while maintaining zero accuracy degradation in downstream tasks.
This work addresses the limitation of traditional channel decoding, which neglects the statistical and semantic structure of source data and thus fails to provide differentiated protection for semantically critical information. To overcome this, the paper proposes a Semantic Error Control Coding (SECC) framework that, for the first time, integrates semantic priors learned by foundation models—such as large language models—into the entire channel encoding and decoding pipeline. While preserving the algebraic structure of conventional channel codes, SECC enables adaptive redundancy allocation and error correction guided by semantic importance. By combining maximum a posteriori estimation, semantic-driven candidate search, and error detection mechanisms, SECC achieves several decibels of coding gain over text sources in AWGN channels, significantly outperforming the normal approximation bound in terms of bit error rate.