SparseDesign: Scaling Exact Coding-Sequence Design

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the computational and memory bottlenecks of dynamic programming in the exact optimization of synonymous coding sequences. To overcome these limitations, this work proposes a multi-loop recursive acceleration algorithm based on candidate sparsification, enabling joint optimization of folding energy and codon usage over weighted codon automata. We rigorously prove the equivalence between the sparse recurrence and its dense counterpart, and introduce an endpoint ownership mechanism to support lock-free parallel construction of candidate sets. Integrated with the Turner 2004 thermodynamic model, our approach reduces the candidate retention rate for natural proteins to approximately 3%. Consequently, it achieves over a 20-fold speedup and nearly a 28-fold reduction in memory consumption compared to conventional dense algorithms.
📝 Abstract
Exact optimization of synonymous coding sequences under a joint folding-energy and codon-usage objective is limited by expensive dynamic-programming splits and large working sets. \textsc{SparseDesign} applies candidate sparsification to the multiloop recurrence of a Turner~2004 dangle-0 solver over a weighted codon automaton. A direct branch is retained only when it strictly improves on every partitionable or endpoint-unpaired realization of the same endpoint states. We prove equivalence to the dense recurrence in real arithmetic, under an explicit scalar branch-interface assumption. With $N$ automaton states, edge set $E$ and $Z$ retained candidates, multiloop work is $O(N^2+N|E|+NZ)$; worst-case time remains cubic for bounded-width automata and total memory remains quadratic. Endpoint ownership permits parallel candidate construction without locks. While synthetic stress families can benefit little from sparsification and exhibit near-quadratic candidate growth, natural proteins show substantial candidate-count reductions. In our 7,600-task campaign, the 2,000-protein human-table panel has median retention of only 3.53\% at $\lambda=0$ and 2.15\% at $\lambda=4$, corresponding to approximately 28.3-fold and 46.4-fold reductions relative to all feasible direct intervals. The primary performance experiments use an AMD EPYC 7313 server. For human Dp427c (11,031 nt, $\lambda=0$), 16-thread packed \textsc{SparseDesign} achieves five-run medians of 236.54 seconds wall-clock time and 14.43 GiB peak RSS. Compared with the single-thread local dense LinearDesign fork on the same server (4,912 seconds, 402.10 GiB RSS), this gives a 20.8-fold wall-clock speedup and a 27.9-fold peak-memory reduction. On a Core i9-14900KF commodity PC with 64 GiB RAM, the same input, layout and thread count achieve 126.42 seconds and 14.43 GiB RSS.
Problem

Research questions and friction points this paper is trying to address.

coding-sequence design
synonymous optimization
folding energy
codon usage
dynamic programming
Innovation

Methods, ideas, or system contributions that make the work stand out.

candidate sparsification
coding-sequence design
multiloop recurrence
weighted codon automaton
parallel optimization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hao Lin
Department of Mechanical and Aerospace Engineering, Rutgers University, Piscataway, NJ, USA
Jingjin Yu
Jingjin Yu
Associate Professor of Comp. Sci., Rutgers Univ. at New Brunswick, Roboticist
Algorithmic Foundations for Robotics