🤖 AI Summary
This work addresses a critical security gap in existing vector databases, which lack native support for embedding integrity, anomaly detection, and provenance authentication, rendering them vulnerable to attackers with write access who can exploit subtle perturbations—such as orthogonal rotations or scaling—to steganographically exfiltrate sensitive data through embeddings. We present the first systematic demonstration and validation of steganographic attacks at the embedding layer in Retrieval-Augmented Generation (RAG) systems. To counter this threat, we propose VectorPin, an embedding-level integrity protection mechanism that cryptographically binds each embedding to its source content via Ed25519 digital signatures, enabling verifiable provenance. Experiments show that minor orthogonal rotations can evade conventional distribution-based detectors, whereas VectorPin reliably identifies any post-embedding tampering across diverse models, corpora, and seven vector database configurations, thereby fundamentally mitigating this class of attacks.
📝 Abstract
Modern retrieval-augmented generation (RAG) systems convert sensitive content into high-dimensional embeddings and store them in vector databases that treat the resulting numerical artifacts as opaque. Major vector-store products do not provide native controls for embedding integrity, ingestion-time distributional anomaly detection, or cryptographic provenance attestation. We show this opens a class of steganographic exfiltration attacks: an attacker with write access to the ingestion pipeline can hide payload data inside embeddings using simple post-embedding perturbations (noise injection, rotation, scaling, offset, fragmentation, and combinations thereof) while preserving the surface-level retrieval behavior the RAG system exposes to legitimate users.
We evaluate these techniques across a synthetic-PII corpus on text-embedding-3-large, four locally hosted open embedding models, a cross-corpus replication on BEIR NFCorpus and a Quora subset (over 26,000 chunks combined), seven vector-store configurations, an adaptive-attacker variant of the detector evaluation, and a paraphrased-query retrieval benchmark. Distribution-shifting perturbations are often caught by simple anomaly detectors; small-angle orthogonal rotation defeats distribution-based detection across every (model, corpus) pair tested. A disjoint-Givens rotation encoder gives a closed-form per-vector capacity ceiling of floor(d/2) * b bits, but real embedding manifolds impose a capacity-detectability trade-off, and the retrieval-preserving operating point sits well below it.
We propose VectorPin, a cryptographic provenance protocol that pins each embedding to its source content and producing model via an Ed25519 signature over a canonical byte representation. Any post-embedding modification breaks signature verification. Embedding-level integrity is a deployable, standardizable control that closes this attack class.