🤖 AI Summary
This work addresses the multi-constraint DNA code design problem for DNA-based data storage, requiring simultaneous satisfaction of eight constraints: reverse, reverse-complement, GC-content (40%–60%), Hamming distance lower bound, non-correlation, thermodynamic stability, absence of homopolymers, and secondary structure suppression. We propose an algebraic construction framework based on the finite ring ℤ₄[i]/⟨2⟩, employing an isometric mapping from codewords over the ring to DNA sequences. To enhance coding performance, we introduce a joint metric combining the Gau distance and the non-homopolymer distance, enabling the construction of non-repeating DNA codes with high Hamming distance. We establish novel algebraic bounds and a unified metric system for multi-constraint cooperative optimization, significantly improving codebook size and error-correction capability. This framework provides both theoretical foundations and efficient constructive methods for scalable, robust DNA storage coding.
📝 Abstract
DNA strings and their properties are widely studied since last 20 years due to its applications in DNA computing. In this area, one designs a set of DNA strings (called DNA code) which satisfies certain thermodynamic and combinatorial constraints such as reverse constraint, reverse-complement constraint, $GC$-content constraint and Hamming constraint. However recent applications of DNA codes in DNA data storage resulted in many new constraints on DNA codes such as avoiding tandem repeats constraint (a generalization of non-homopolymer constraint) and avoiding secondary structures constraint. Therefore, in this chapter, we introduce DNA codes with recently developed constraints. In particular, we discuss reverse, reverse-complement, $GC$-content, Hamming, uncorrelated-correlated, thermodynamic, avoiding tandem repeats and avoiding secondary structures constraints. DNA codes are constructed using various approaches such as algebraic, computational, and combinatorial. In particular, in algebraic approaches, one uses a finite ring and a map to construct a DNA code. Most of such approaches does not yield DNA codes with high Hamming distance. In this chapter, we focus on algebraic constructions using maps (usually an isometry on some finite ring) which yields DNA codes with high Hamming distance. We focus on non-cyclic DNA codes. We briefly discuss various metrics such as Gau distance, Non-Homopolymer distance etc. We discuss about algebraic constructions of families of DNA codes that satisfy multiple constraints and/or properties. Further, we also discuss about algebraic bounds on DNA codes with multiple constraints. Finally, we present some open research directions in this area.