🤖 AI Summary
This study addresses the challenges of redundant reads, retrieval latency, and elevated sequencing costs arising from random sampling in distributed DNA storage. Drawing upon coding theory and probabilistic combinatorial optimization, it investigates the coverage depth required for full message recovery, deriving the exact recovery time distribution and its expectation for arbitrary linear codes. The primary contributions include proving the optimality of MDS codes and establishing a lower bound on total read cost, as well as resolving the unique minimization conjecture for simplex codes. Furthermore, this work delineates the container size regimes under which distributed sampling effectively reduces latency or cost, thereby elucidating the theoretical conditions governing the trade-off between parallelism and cost efficiency.
📝 Abstract
Random sampling in DNA sequencing produces repeated reads, increasing retrieval latency and sequencing cost. We study the coverage-depth problem for full-message recovery in distributed DNA storage under noiseless uniform sampling, where strands are partitioned among $M$ containers and one strand is independently sampled with replacement from each container per round. For arbitrary linear codes and ordered partitions, we derive exact formulas for the recovery-time distribution and expectation. We prove that MDS codes, whenever they exist, are optimal for every fixed partition, and establish a universal lower bound on the expected total read cost together with its equality conditions. For MDS codes, we identify container-size regimes that yield genuine savings in total reads and regimes that provide only parallelism without changing the asymptotic sequencing cost. For simplex codes, we prove that the $q$-ary simplex code is, up to isomorphism, the unique single-container minimizer among codes with the same parameters, resolving a recent conjecture by Bertuzzo, Ravagnani, and Yaakobi. We further construct a partition attaining the minimum total read cost and derive bounds for intermediate and balanced partitions. These results clarify when distributed sampling reduces latency alone and when it also reduces sequencing cost.