MuSAlS: A Fast Multiple Sequence Alignment Approach Using Hierarchical Clustering

๐Ÿ“… 2026-01-21
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work proposes a novel approach that integrates hierarchical clustering with exact dynamic programming to address the poor computational efficiency and limited scalability of multiple sequence alignment (MSA) on large-scale sequence data. By constructing a guide tree based on Levenshtein distance and employing a bottom-up strategy to invoke the Needlemanโ€“Wunsch algorithm, the method achieves high alignment accuracy while substantially improving runtime performance. Implemented efficiently in Rust, the proposed framework matches the accuracy of state-of-the-art tools yet significantly reduces computation time. This study presents the first effective fusion of exact dynamic programming with scalable hierarchical clustering, making it well-suited for large-scale genomic and metagenomic analyses.

Technology Category

Machine Learning: Scalability of ML SystemsData Mining & Knowledge Management: Scalability, Parallel & Distributed SystemsSearch and Optimization: Metareasoning and Metaheuristics

Application Category

Graph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsSearch and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesSemantics and Knowledge: Scalable techniques for the creation, curation, publication, maintenance, and consumption of large, Web-based, structured, reusable, knowledge graphs and ontologies
๐Ÿ“ Abstract
Motivation: The multiple sequence alignment (MSA) problem has been extensively studied, with numerous approaches developed over recent years. With the rapid growth of sequence data, there is an increasing need for fast and accurate MSA tools that scale effectively to large datasets. Building on our previous work on CLAM, we are able to use exact dynamic programming (Needleman-Wunsch) while scaling to large datasets. We introduce MuSAlS (Multiple Sequence Alignment at Scale), a fast and scalable de novo MSA aligner. MuSAlS uses hierarchical clustering to construct a guide tree based on the Levenshtein distance metric, enabling efficient and accurate alignment through a bottom-up approach. Results: MuSAlS achieves competitive accuracy compared to state-of-the-art methods while significantly improving runtime performance. This makes it a valuable tool for researchers analyzing large-scale genomic and metagenomic datasets, addressing the growing demand for scalable bioinformatics solutions. Availability and Implementation: MuSAlS is implemented in the Rust programming language, and available at https://github.com/URI-ABD/clam
Problem

Research questions and friction points this paper is trying to address.

Multiple Sequence Alignment
Scalability
Large-scale genomic data
Computational efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multiple Sequence Alignment
Hierarchical Clustering
Levenshtein Distance
Scalable Bioinformatics
Exact Dynamic Programming
๐Ÿ”Ž Similar Papers
๐Ÿ’ผ Related Jobs
No related jobs found.
E
Emily G. Light
Department of Computer Science and Statistics, University of Rhode Island
M
Morgan E. Prior
Department of Computer Science, Tufts University
Noah M. Daniels
Noah M. Daniels
Associate Professor, Computer Science, University of Rhode Island
Computational BiologyMachine LearningAlgorithms
N
Najib Ishaq
Department of Computer Science and Statistics, University of Rhode Island