🤖 AI Summary
This study addresses the security and intellectual property risks in black-box retrieval systems arising from unverifiable model identities. We propose TellTail, a fingerprinting attack framework that innovatively exploits the low transferability of embedding models as distinctive fingerprint signatures. By integrating adversarial query optimization with retrieval overlap analysis, TellTail enables retriever identification across diverse scenarios, ranging from ranking lists to generated answers. Extensive experiments involving 53 retrievers demonstrate that our approach achieves 100% identification accuracy under full-ranking settings. Notably, it maintains robust performance even under highly constrained conditions, attaining 94.3% accuracy for Top-3 rankings and 92.5% when only generated answers are exposed. Furthermore, we successfully validate the practical applicability of TellTail on real-world Retrieval-Augmented Generation (RAG) systems such as OpenWebUI, highlighting its effectiveness in identifying proprietary retrievers deployed in production environments.
📝 Abstract
Dense embedding models are core to modern text retrieval, enabling systems ranging from web search to retrieval-augmented generation (RAG). Yet, retrievers are usually deployed within opaque systems, exposing only ranked results, cited sources, or generated answers. This opacity prevents users from verifying which retrievers providers serve and may create a false sense of robustness against retrieval attacks. We show that attackers can infer retrievers'identity only through queries. We introduce TellTail, a fingerprinting attack for identifying retrievers behind black-box systems across a spectrum of access levels. Despite sharing many properties, retrievers can be steered to emit distinct results. When retrieved passages are exposed, TellTail compares retrieval overlap to deduce the underlying model. When only the final generated response is exposed, TellTail optimizes model-specific queries that induce a chosen retrieval behavior on the target retriever but transfer poorly to others. Interestingly, the poor transferability that limits attacks elsewhere is exactly what makes them useful for fingerprinting. We evaluate TellTail using 53 retrievers under three increasingly restrictive settings. TellTail perfectly identifies the deployed retriever from full retrieval rankings; in 94.3% of attempts with unordered top-3 results; and in 92.5% of cases from language-model-generated answers alone. Furthermore, fingerprinting succeeds against a popular RAG system (OpenWebUI). Overall, TellTail shows that retrieval-based systems leak the identity of their embedding model---exposing intellectual property and enabling reconnaissance ahead of model-specific attacks such as corpus poisoning.