๐ค AI Summary
This study investigates the trade-off between download cost and privacy leakage in private information retrieval (PIR) from multiple colluding servers for DNA data storage. Leveraging information-theoretic analysis and mutual information metrics, we derive theoretical lower bounds under multi-collusion scenarios and propose a novel randomized subset single-server architecture applicable to an arbitrary number of servers. The primary contributions include constructing a tight scheme for two files with proven bound optimality and designing a general-purpose random sampling framework for arbitrarily scaled datasets. This work establishes a rigorous quantitative balance between download efficiency and privacy protection, thereby providing theoretical and methodological foundations for efficient, secure private data retrieval in DNA-based storage systems.
๐ Abstract
As DNA-based data storage evolves, protecting user privacy during data retrieval has become increasingly important. We study sequential random sampling DNA private information retrieval (SRS DNA PIR) with multiple colluding random sampling servers, where the database is partitioned into servers of equal size. We investigate the tradeoff between the download cost, defined as the expected number of queries, and the privacy leakage, measured by mutual information. We derive lower bounds on this tradeoff, including a bound given by an optimization problem. This bound is tight when each server stores two files, and we construct schemes that attain it. For servers of any size, we construct schemes that apply a single-server scheme to a randomly selected subset of servers.