🤖 AI Summary
This study addresses the long-standing open problem of computing the capacity and constructing optimal codes for single-tag labeling of DNA sequences. By modeling the labeling process as a deterministic channel, this work establishes, for the first time, an equivalence between the labeling capacity and the zero-error capacity of star graphs, conducting a systematic analysis that integrates information-theoretic, graph-theoretic, and combinatorial coding techniques. The authors fully derive the zero-error capacities of all star graphs and precisely characterize the labeling capacity for all single-tag scenarios over arbitrary finite alphabets. Furthermore, they present a universal coding construction scheme that achieves this capacity and delineate the capacity bounds along with their extremal structures for fixed tag lengths. These contributions collectively provide a complete theoretical resolution to the fundamental limits of single-tag DNA labeling.
📝 Abstract
DNA labeling has attracted increasing attention in biomedical applications, including molecular imaging, diagnostics, and genomic analysis. In a DNA labeling process, a set of DNA sequence patterns, referred to as labels, is designed according to the requirements of a specific application. For each DNA sequence, the labeling process generates an output sequence that records the positions of the labels. DNA sequences with different labeling outputs can therefore be distinguished through the labeling process. To quantify this capability, the labeling capacity is defined as the exponential growth rate of the maximum number of DNA sequences that can be distinguished through the labeling process as the sequence length tends to infinity [2]. To date, the labeling capacities of several cases in the single-label setting have been determined. In this paper, we formulate the labeling process as a deterministic channel and show that its zero-error capacity is equal to the labeling capacity. For a single label, the corresponding channel can be represented by a star graph. Thus, characterizing the labeling capacity of a single label is equivalent to determining the zero-error capacity of the corresponding star graph. We derive the zero-error capacities of all star graphs, thereby providing a complete characterization of the labeling capacities for all single-label cases. Furthermore, we develop a general method for constructing capacity-achieving codes. These results apply to labeling problems over arbitrary finite alphabets and are not restricted to the DNA alphabet. Finally, for a fixed label length, we exactly characterize the range of achievable labeling capacities and identify all single-label structures that attain the minimum and maximum capacities.