🤖 AI Summary
This study addresses how coding redundancy allocation affects dual-file retrieval time in DNA data storage. Based on systematic linear codes, the authors analyze the growth process of the column space of generator matrices from a geometric perspective. They innovatively introduce the concept of "mixed columns" and sharpen projection bounds to control their influence, combining finite field coding theory with random sampling analysis to track subspace dimension evolution. When the total information dimension satisfies specific conditions, this work rigorously proves the conjecture that the sum of reciprocals of expected dual-file retrieval times does not exceed one. By extending prior results for codes without mixed columns to the general case, the proposed approach optimizes coding design bounds for equal-sized files.
📝 Abstract
In DNA-based storage systems, data are retrieved by sequencing molecules sampled from a DNA pool. We study how the way coding redundancy is shared between files affects their expected retrieval times. Our focus lies on the case of two files that are encoded by a systematic linear code over an arbitrary finite field. We consider the conjecture that the sum obtained by dividing each file dimension by its expected retrieval time is at most one whenever the dimension of at least one file is more than one. For this, we develop a geometric view of the retrieval process. As molecules are sampled, we follow the growing span of the corresponding columns of the generator matrix and track how much of this span comes from each file. Among the samples that enlarge the overall span, this lets us compare those that make progress toward recovering both files with those that make progress toward neither. We call a column mixed if the corresponding encoded symbol combines information from both files. In this work, we sharpen a projection bound and use it to control the effect of mixed columns. We prove the conjecture whenever the total information dimension is at least twice the number of mixed columns plus two. For equal-sized files, this allows up to one fewer mixed column than the dimension of either file and extends the previous result for codes with no mixed columns.