🤖 AI Summary
Serverless architectures are inherently ill-suited for long-running, memory-intensive bioinformatics workloads—particularly RNA-seq alignment—due to strict time and memory limits.
Method: This work presents the first successful migration of the memory-intensive STAR aligner to a cloud-native serverless environment, leveraging an AWS ECS–based containerized deployment framework. Key innovations include STAR-specific container optimization, coordinated memory–I/O tuning, and fine-grained batch scheduling to overcome serverless constraints.
Contribution/Results: Evaluated on 17 TB of real RNA-seq data, the serverless pipeline reduces execution time by 32% and operational cost by 41% compared to conventional VM-based deployments in small-batch scenarios. The study demonstrates that serverless computing is both feasible and cost-effective for scalable, elastic bioinformatics analysis, establishing a novel paradigm for cloud-native genomics.
📝 Abstract
The application of serverless computing for alignment of RNA-sequences can improve many existing bioinformatics workflows by reducing operational costs and execution times. This work analyzes the applicability of serverless services for running the STAR aligner, which is known for its accuracy and large memory requirement. This presents a challenge, as serverless services were designed for light and short tasks. Nevertheless, we successfully deploy a STAR-based pipeline on AWS ECS service, propose multiple optimizations, and perform experiment with 17 TBs of data. Results are compared against standard virtual machine (VM) based solution showing that serverless is a valid alternative for small-scale batch processing. However, in large-scale where efficiency matters the most, VMs are still recommended.