Score
Designs, implements, and evaluates computational models, architectures, and end‑to‑end pipelines that consume or produce ordered data (token sequences or time‑indexed measurements). This work covers seq2seq and generative sequence models, long‑sequence and lifelong/temporal modeling, multimodal sequence fusion, sequence prediction and analysis pipelines, and methods for sequence compression, training, and evaluation.
To address the fundamental mismatch between teacher-forced training and autoregressive inference—along with associated complexities in state management and error-proneness in sequence modeling—this paper introduces a unified, streamable neural network layer API. The core innovation is an explicit state interface coupled with a `step()` method, enabling identical numerical behavior under both parallel (training-time) and incremental (inference-time) execution modes within the same layer. By abstracting temporal states—including KV caches, convolutional buffers, and RNN hidden states—the API supports declarative, composable layer design in JAX and TensorFlow 2. An open-source library provides a comprehensive suite of streamable primitive layers and composition utilities. This framework significantly lowers the barrier to developing and deploying streaming models while ensuring production-grade correctness, efficiency, and maintainability.
Long-sequence modeling faces fundamental challenges including limited context length, difficulty in capturing long-range dependencies, and low efficiency in online learning. To address these, this work systematically reviews the resurgence of state-space models (SSMs) and recurrent computation, proposing a novel local forward-computation paradigm tailored for real-world online learning—thereby circumventing the temporal backtracking constraints inherent in standard backpropagation through time (BPTT). We introduce the first unified taxonomy encompassing both deep SSMs and large-context Transformers. Our framework integrates structured linear attention, enhanced RNN architectures, local recurrence mechanisms, and online optimization algorithms. The study rigorously clarifies the theoretical representational advantages and practical sequential reasoning benefits of recurrent modeling over alternatives. Collectively, this work delivers a scalable technical roadmap for low-latency, highly extensible long-sequence modeling.
This work addresses the lack of systematic evaluation of sequence models’ ability to capture diverse temporal dependencies—such as short- and long-range, decaying, and oscillatory patterns. We propose the first synthetic benchmark framework based on controllable, parameterized memory functions. By explicitly designing memory kernel functions, our framework generates synthetic tasks with continuous-time complexity, enabling fine-grained, interpretable, and theoretically grounded analysis of model memory characteristics. We evaluate mainstream architectures—including RNNs, Transformers, and State Space Models (SSMs)—under a unified benchmark across multiple dimensions. Our experiments not only validate existing theoretical predictions but also uncover, for the first time, implicit architectural preferences for specific memory patterns and their precise failure boundaries. The results provide reproducible, interpretable, and quantitative guidance for selecting appropriate sequence modeling paradigms.
To address the high memory consumption, low computational efficiency, and substantial communication overhead incurred by multi-dimensional Transformers when processing long sequences, this paper proposes Dynamic Sequence Parallelism (DSP), a novel parallelization paradigm. DSP transcends conventional single-dimension sequence parallelism by enabling adaptive switching of parallelization dimensions during computation, coupled with a communication-aware tensor resharding mechanism that supports module-level distributed execution with minimal constraints and maximal flexibility. By tightly integrating multi-dimensional sequence modeling with distributed training optimization, DSP achieves significant improvements over the state-of-the-art embedded sequence parallelism: throughput increases by 32.2%–10×, communication volume decreases by over 75%, and measured communication overhead accounts for less than 25% of total execution time.
Existing gene sequence alignment methods lack systematic, cross-platform evaluation. Method: We introduce the first benchmark platform specifically designed for multi-platform sequencing data (Illumina, Oxford Nanopore Technology, and PacBio), systematically evaluating 11 state-of-the-art aligners—including exact, heuristic, and learning-enhanced algorithms. Contribution/Results: Our end-to-end empirical analysis quantifies, for the first time, the high sensitivity of alignment performance to both sequencing data quality and hyperparameter configurations. We propose a standardized four-dimensional evaluation framework assessing accuracy, speed, memory footprint, and noise robustness. Results reveal widespread deficiencies in robustness and resource efficiency across current tools. To support reproducibility and methodological advancement, we open-source a fully documented, end-to-end benchmarking pipeline on GitHub—providing an evidence-based foundation for algorithm selection, comparative analysis, and future alignment method development.
Efficient deployment of large language models (LLMs) in resource-constrained settings necessitates synergistic model compression, yet the optimal sequencing and interaction effects of knowledge distillation (KD), structured pruning, and low-bit quantization remain unclear. Method: We systematically investigate all permutations of these three techniques via controlled experiments on Qwen2.5-3B, evaluating their impact on both model performance and compression ratio. Contribution/Results: We identify pruning–KD–quantization (P-KD-Q) as the optimal cascade: structured pruning first preserves architectural redundancy; KD subsequently recovers accuracy lost during pruning; and quantization is applied last to avoid irreversible information loss from early low-bit approximation. This sequence achieves a 3.68× compression ratio while preserving strong language understanding and instruction-following capabilities. Our findings establish a reproducible, generalizable pipeline for LLM lightweighting—offering the first empirical evidence that compression order critically governs trade-offs between efficiency and fidelity.
This study systematically evaluates the practical benefits of pretraining in DNA language models for downstream genomic tasks and investigates the effectiveness of Byte Pair Encoding (BPE) tokenization. By comparing Transformer-based architectures (e.g., DNABERT2) with convolutional models (e.g., ConvNova) across a range of genomic fine-tuning benchmarks, the work provides the first empirical analysis of whether pretraining is necessary and how BPE compares to traditional k-mer representations. The findings indicate that the performance gains conferred by pretraining are limited and that BPE does not consistently outperform k-mer tokenization across all tasks. These results offer critical empirical insights for the design of foundational models in genomics, challenging prevailing assumptions about the universal advantages of large-scale pretraining and subword tokenization in this domain.
This work addresses the inefficiency of traditional Dynamic Time Warping (DTW) in aligning long sequences due to its inherently serial computation, which hinders effective GPU parallelization. The study presents the first systematic exploration of GPU-oriented parallel alternatives to DTW, introducing four novel algorithms. The first three employ rectangular block-based approximations to accelerate computation, while the fourth, termed ParDTW, achieves exact alignment through diagonal-wise parallelization. ParDTW integrates block matrix processing with a diagonal scheduling strategy, preserving full alignment accuracy while delivering 15–100× speedup over existing methods on long sequences. This breakthrough substantially overcomes the performance limitations of conventional DTW, establishing ParDTW as an efficient and practical solution for large-scale sequence alignment tasks.
Current biomolecular sequence design methods lack unified, reproducible evaluation standards, hindering fair and rigorous performance comparison. To address this, we introduce BioSeqEval—a modular, open-source Python evaluation library that systematically integrates three model-agnostic metric categories: sequence-based, embedding-based, and property-based—representing the first such comprehensive framework. It supports one-shot and iterative design evaluation across diverse sequence modalities, including small molecules, DNA, RNA, peptides, and proteins. The library incorporates state-of-the-art pretrained embedding models, machine learning–based property predictors, efficient sequence alignment tools, and interactive visualization modules for diagnostic analysis. Empirical evaluation demonstrates that BioSeqEval significantly enhances evaluation standardization, cross-method comparability, and methodological transparency. It exhibits strong flexibility and robustness across multiple benchmark design tasks, enabling reproducible, interpretable, and scalable assessment of generative sequence models.
This work addresses the limitations of traditional autoregressive models, which rely on discrete tokenization and struggle to accurately model continuous values—often leading to functional failures in precision-sensitive tasks such as semiconductor circuit design. To overcome this, the authors propose AGDC, a unified framework that enables end-to-end autoregressive generation of hybrid discrete-continuous sequences for the first time. Built upon a Transformer architecture, AGDC integrates classification-based prediction with diffusion modeling and introduces a dynamic EOS logit adjustment mechanism alongside a length regularization loss. Evaluated on a newly curated high-precision semiconductor layout dataset, ContLayNet (334K samples), and SVG-based graphics tasks, AGDC significantly outperforms both discretization-based and fixed-structure baselines, breaking through the precision bottleneck and enabling high-fidelity, variable-length vector data generation.