LvD: A New Algorithm for Computing the Likelihood of a Phylogeny

📅 2026-01-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work proposes LvD (Likelihood via Decomposition), a novel algorithm that overcomes the O(n) worst-case time complexity bottleneck of traditional phylogenetic likelihood computation methods—such as Felsenstein’s pruning algorithm—when performing updates on large trees. By introducing an innovative tree decomposition strategy, LvD achieves O(log n) worst-case time complexity for likelihood updates and enables O(log n) per-site parallel computation. The method is fully compatible with all standard nucleotide substitution models and demonstrates substantial speedups on both simulated and empirical datasets, with performance gains increasing as tree balance improves. This advancement represents the first approach to break the longstanding O(n) barrier in phylogenetic likelihood evaluation, offering significant scalability for large-scale phylogenomic analyses.

Technology Category

Machine Learning: Matrix & Tensor MethodsSearch and Optimization: Distributed SearchConstraint Satisfaction and Optimization: Distributed CSP/Optimization

Application Category

Graph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsWeb Mining and Content Analysis: Models for Web evolution
📝 Abstract
There are few, if any, algorithms in statistical phylogenetics which are used more heavily than Felsenstein's 1973 pruning method for computing the likelihood of a tree. We present LvD, (Likelihood via Decomposition), an alternative to Felsenstein's algorithm based on a different decomposition of the underlying phylogeny. It works for all standard nucleotide models. The new algorithm allows updates of the likelihood calculation in worst case $O(\log n)$ time with $n$ taxa, as opposed to worst case $O(n)$ time for existing methods. In practice this leads to appreciable improvements in likelihood calculations, the extent of speed-up depending on how balanced or unbalanced the trees are. We explore implications for parallel computing, and show that the approach allows likelihoods to be computed in $O(\log n)$ parallel time per site, compared to (worst case) $O(n)$ time. We implemented and applied the algorithm to large numbers of simulated and empirical data sets and showed that these theoretical advances lead to a significant practical speed-up, although the extent of the improvement depends on how balanced the phylogenies already are.
Problem

Research questions and friction points this paper is trying to address.

phylogeny likelihood
computational efficiency
tree inference
statistical phylogenetics
algorithmic scalability
Innovation

Methods, ideas, or system contributions that make the work stand out.

phylogenetic likelihood
LvD algorithm
tree decomposition
logarithmic time complexity
parallel computing
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.