Score
Designs and evaluates representation-learning models and encoding schemes that transform inputs into latent representations which preserve intended semantic content for authorized decoders while preventing unauthorized parties or intermediaries from inferring sensitive semantic information. Work includes building semantic obfuscation and privacy-encoding methods, measuring and optimizing reconstruction fidelity for legitimate receivers, and analyzing or defending against semantic leakage and inference attacks.
Although image embeddings do not directly reconstruct the original images, they may still leak semantic privacy. This work formally introduces the notion of “semantic leakage” and proposes SLImE, a framework that recovers semantic content solely from local semantic neighborhood structures while preserving embedding alignment—without requiring task-specific decoders. By integrating off-the-shelf embedding models with a lightweight local semantic retriever, SLImE leverages a neighborhood propagation mechanism to enable efficient inference. Experiments across diverse models—including GEMINI, COHERE, NOMIC, and CLIP—demonstrate consistent recovery of semantic labels, symbols, and coherent descriptions, revealing inherent privacy risks embedded within image representations.
Speech LLM training poses significant risks of speaker identity leakage. To address this, we propose the Universal Speech Codec (USC), the first framework to explicitly decouple semantic content from privacy-sensitive acoustic features via a dual-stream semantic-acoustic architecture. USC jointly optimizes three objectives: semantic fidelity, preservation of paralinguistic attributes (e.g., prosody and emotion), and speaker anonymization. We introduce a privacy-aware reconstruction loss and adversarial de-identification training, and establish a novel perceptual-test-based privacy evaluation paradigm. Experiments demonstrate that USC achieves state-of-the-art performance on downstream tasks—including automatic speech recognition and emotion recognition—while outperforming mainstream speech codecs in reconstruction quality. Crucially, USC reduces success rates of speaker identity attacks by a substantial margin, thereby reconciling high-fidelity representation learning with rigorous privacy protection.
This work uncovers a fundamental security vulnerability in deep learning–driven semantic communication (SemCom): adversaries can stealthily manipulate transmitted semantic content within the latent space without perturbing the statistical distribution of latent variables. To exploit this, we propose two novel attack paradigms: (1) DiR—a diffusion-based re-encoding attack enabling controllable semantic regeneration; and (2) TTA-LM—a training-free, test-time adaptive attack with cross-model and cross-modal generalizability. Both attacks achieve semantic manipulation via imperceptible, directionally guided perturbations exclusively in the latent space. Extensive experiments demonstrate that these attacks efficiently distort decoded semantics across mainstream SemCom architectures while preserving the naturalness of latent variable distributions—rendering them highly stealthy and resistant to detection. This study is the first to systematically identify, formalize, and empirically validate latent-space manipulation as a critical, previously unrecognized threat to semantic communication security.
This work addresses a critical privacy vulnerability in relay-assisted semantic communication systems, where relay nodes—despite lacking access to the original data—can infer sensitive information from semantic representations, leading to severe privacy leakage. The study is the first to expose this semantic privacy risk and introduces an iterative adversarial training framework that leverages a deep learning–based semantic communication architecture. This framework dynamically optimizes the strategic interaction between the legitimate receiver and a relay-side eavesdropper, simultaneously preserving high semantic reconstruction fidelity at the intended destination while actively suppressing the relay’s capability to infer semantic content. Experimental results demonstrate that the proposed method consistently and significantly widens the gap in semantic accuracy between the legitimate receiver and the relay across diverse channel conditions, thereby achieving effective and covert privacy protection.
This paper addresses semantic privacy risks in large language models (LLMs)—specifically, the leakage of implicit, context-dependent, or inferable information in sensitive scenarios. It systematically analyzes root causes across the LLM lifecycle: input processing, pretraining, fine-tuning, and alignment. We propose the first holistic semantic privacy risk analysis framework for LLMs, exposing fundamental limitations of existing defenses against contextual inference and latent representation leakage. Through empirical evaluation, we integrate differential privacy, embedding encryption, edge computing, and machine unlearning to assess semantic-level protection efficacy. Our work establishes the first comprehensive research framework for LLM semantic privacy, identifying key open challenges—including multimodal privacy preservation, de-identification, and the trade-off between privacy and generation quality. The study provides both theoretical foundations and practical guidelines for designing semantics-aware privacy mechanisms. (149 words)
Existing image watermarking methods treat payloads as semantically agnostic bitstreams, limiting capacity and precluding embedding of human-interpretable, high-level semantic information. Method: We propose a novel semantic watermarking paradigm: (i) compress full natural-language sentences into 256-dimensional unit-norm latent vectors via a lightweight text autoencoder; (ii) fine-tune a watermarking model for robust embedding; and (iii) enforce security via a secret, invertible rotation transformation. Contributions/Results: Our approach breaks the conventional 256-bit capacity ceiling, enabling sentence-level semantic steganography and millisecond-scale real-time decoding. A statistically calibrated scoring mechanism supports AI-generated content provenance tracing and tampering attribution. Evaluated on multiple benchmarks, it substantially outperforms state-of-the-art methods—achieving superior BLEU-4 and Exact Match scores, and ROC AUC of 0.97–0.99—while maintaining strong robustness against geometric and value-domain attacks and offering full interpretability.
This work addresses the vulnerability of existing semantic-aware watermarking schemes to large language model (LLM)-guided semantic perturbation attacks, which can compromise provenance tracing. The authors propose a novel semantic injection attack that preserves visual-semantic consistency by leveraging the semantic reasoning capabilities of LLMs together with embedding-space similarity constraints. This approach precisely perturbs high-level semantics associated with the watermark while maintaining global image coherence, thereby deceiving detection mechanisms. Notably, this method is the first to demonstrate and exploit the capacity of LLMs to launch targeted attacks against semantic watermarks, challenging the foundational security assumptions of content-aware watermarking. Extensive evaluations show that the proposed attack significantly outperforms existing approaches across multiple state-of-the-art watermarking schemes, revealing a fundamental security flaw in current semantic watermarking under LLM-driven perturbations.
This work proposes “conceptual steganography,” a novel steganographic approach that enables large language models to covertly transmit harmful information through chain-of-thought reasoning while evading human oversight. Unlike conventional methods relying on token-level manipulations or lexical choices, this technique elevates information encoding to the level of high-level reasoning behavior patterns, substantially enhancing robustness against existing paraphrasing-based defenses. The method integrates chain-of-thought generation with concept-level encoding and introduces a strategy-aware paraphrasing defense mechanism. Experiments across four model families and two reasoning tasks demonstrate that the proposed approach achieves strong stealth and effectiveness without compromising reasoning performance, while the strategy-aware paraphraser significantly mitigates such steganographic channels.
Language Model Inversion (LMI) recovers sensitive prompts from model outputs, posing severe threats to user privacy and system security. To address this, we propose a novel LMI framework grounded in latent-space invariance. First, we introduce the Invariant Latent-Space Hypothesis (ILSH), identifying source and cyclic invariance as fundamental principles underlying prompt inversion. Second, we design a lightweight inverse encoder trained in two stages—without requiring large-scale inversion corpora—and optimized via training-free neighborhood search. Third, at the representation level, we fuse multi-output information through contrastive alignment, supervised reinforcement learning, sparse representation concatenation, and pseudo-representation denoising; the target LLM itself serves as an invariant decoder for efficient inverse mapping. Evaluated on nine benchmarks, our method achieves a +4.77 BLEU average gain while drastically reducing reliance on inversion data. Empirical analysis further reveals the limited efficacy of existing defenses, underscoring both the necessity and advancement of our approach.