🤖 AI Summary
This study addresses the fragmentation of omics data, the difficulty of cross-modal analysis, and limited reproducibility inherent in traditional file-based workflows by proposing vcf2db, a sample-centric relational framework for multi-omics integration. The approach models individual omics layers as independent yet linkable modules, enabling cross-layer SQL queries via shared sample identifiers while supporting flexible extension to new omics layers without modifying the underlying schema. Validation using 1000 Genomes Project data demonstrates that this framework significantly outperforms existing tools in coordinate filtering and annotation query performance. Furthermore, it successfully achieves seamless integration of a transcriptomic layer, providing an efficient and scalable solution for unified multi-omics data access.
📝 Abstract
Background: Rapid growth of high-throughput molecular data demands systems for efficient retrieval, integration, and scalability across omics layers. Traditional file-based workflows hinder cross-modal analysis and reproducibility because of fragmented storage and ad-hoc querying. Few existing tools for genomic variation data prioritize modular multi-omics integration. We developed vcf2db, a sample-centric relational framework modeling each omics modality as a distinct but linkable component centered on biological samples. This study evaluates whether this design delivers competitive genomic retrieval while enabling extension to additional molecular layers. Results: We implemented a proof-of-concept genomic schema and ingestion pipeline for annotated VCF data using the European subset of the 1000 Genomes Project (502 samples, 25 million variants). We benchmarked it against three established VCF-oriented tools on seven retrieval tasks: coordinate filtering, annotation-driven queries, genotype extraction, and aggregation. Under controlled conditions, vcf2db performed strongly on selective queries, often outperforming other systems for coordinate and annotation filters, and remained usable for genotype retrieval. Aggregation-heavy tasks were less efficient, indicating optimization targets. We also validated modular extensibility by adding a synthetic transcriptomic layer without modifying genomic tables, linking layers via shared sample identifiers. Conclusion: vcf2db supports cross-layer retrieval directly as SQL queries anchored on shared sample identifiers, enabling integrated multi-omics access that is difficult with file-based approaches.