arXiv Science⌕ Search

arXiv · 2610.01482

Bridging the Omics Divide: A Modular Relational Approach to Multi-Layer Biological Data Management

Abstract

Background: Rapid growth of high-throughput molecular data demands systems for efficient retrieval, integration, and scalability across omics layers. Traditional file-based workflows hinder cross-modal analysis and reproducibility because of fragmented storage and ad-hoc querying. Few existing tools for genomic variation data prioritize modular multi-omics integration. We developed vcf2db, a sample-centric relational framework modeling each omics modality as a distinct but linkable component centered on biological samples. This study evaluates whether this design delivers competitive genomic retrieval while enabling extension to additional molecular layers. Results: We implemented a proof-of-concept genomic schema and ingestion pipeline for annotated VCF data using the European subset of the 1000 Genomes Project (502 samples, 25 million variants). We benchmarked it against three established VCF-oriented tools on seven retrieval tasks: coordinate filtering, annotation-driven queries, genotype extraction, and aggregation. Under controlled conditions, vcf2db performed strongly on selective queries, often outperforming other systems for coordinate and annotation filters, and remained usable for genotype retrieval. Aggregation-heavy tasks were less efficient, indicating optimization targets. We also validated modular extensibility by adding a synthetic transcriptomic layer without modifying genomic tables, linking layers via shared sample identifiers. Conclusion: vcf2db supports cross-layer retrieval directly as SQL queries anchored on shared sample identifiers, enabling integrated multi-omics access that is difficult with file-based approaches.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Alessandro Balestrucci, Donald Friggieri, Andrea Gariboldi, Panagiotis Alexiou. 2026-10-01. Bridging the Omics Divide: A Modular Relational Approach to Multi-Layer Biological Data Management. https://arxiv.org/abs/2610.01482

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

VectorMaton: Efficient Vector Search with Pattern Constraints via an Enhanced Suffix Automaton

Approximate nearest neighbor search (ANNS) has become a cornerstone in modern vector database systems. Given a query vector, ANNS retrieves the closest vectors from a set of base vectors. In real-world applications, vectors are often accompanied by additional information, such as sequences or structured attributes, motivating the need for fine-grained vector search with constraints on this auxiliary data. Existing methods support attribute-based filtering or range-based filtering on categorical and numerical attributes, but they do not support pattern predicates over sequence attributes. In relational databases, predicates such as LIKE and CONTAINS are fundamental operators for filtering records based on substring patterns. As vector databases increasingly adopt SQL-style query interfaces, enabling pattern predicates over sequence attributes (e.g., texts and biological sequences) alongside vector similarity search becomes essential. In this paper, we formulate a novel problem: given a set of vectors each associated with a sequence, retrieve the nearest vectors whose sequences contain a given query pattern. To address this challenge, we propose VectorMaton, an automaton-based index that integrates pattern filtering with efficient vector search, while maintaining an index size comparable to the dataset size. Extensive experiments on real-world datasets demonstrate that VectorMaton consistently outperforms all baselines, achieving up to 10x higher query throughput at the same accuracy and up to 18x reduction in index size.

cs.DB↗

Recommendation Systems for Exploratory Data Tasks

A large class of data-centric tasks is exploratory, where users iteratively steer workflows, refining subjective goals as new insights emerge. These Exploratory Data Tasks (EDTs) are performed by millions of users with varying levels of expertise to understand unfamiliar data, discover trends, and identify evidence that informs critical decision-making. However, a key challenge in EDTs is the enormous space of possible actions that one can take at each step: users struggle to choose among thousands of joins, transformations, and aggregations, causing "exploration paralysis". Because EDT workflows are interconnected, each choice impacts subsequent exploration, and suboptimal choices can lead to inefficiency, missed insights, confirmation bias, and incomplete coverage. This calls for intelligent recommendations that efficiently guide users toward optimal EDT actions. We envision recommendation as a core capability of data systems, proactively guiding users toward promising actions and thereby lowering the barrier to exploratory data tasks. EDT recommendation is challenging because the action space is combinatorial and actions are data-dependent, which require costly materialization. Moreover, recommendation often involves bundles or sequences of actions across interdependent tasks, requiring coordination across tasks. In this paper, we present our vision of EDT recommendation systems along two axes: single-task vs. multi-task settings and single-action vs. multi-action recommendations. We outline a research agenda that progresses from recommending individual EDT actions to constrained bundles and sequences of actions, and ultimately to coordinated recommendations across interconnected EDTs. We identify research directions for incorporating various contexts (user, data, task, and ecosystem), addressing efficiency challenges, and coordinating across tasks.

cs.DB↗

HakiCC: LLM-Driven Multi-Agent Design and Optimization of Concurrency Control Protocols

Large language models (LLMs) have recently been applied in systems research as a tool to reduce human-intensive engineering effort through cost-efficient automation. Decades of research have produced a rich landscape of concurrency control (CC) protocols, each encoding distinct trade-offs in correctness, throughput, and abort behavior. However, most applications in practice default to 2PL or OCC, because selecting and adapting a protocol to a specific application requires expert knowledge that is rarely available to application designers. This is a wasted opportunity, as an application-specific CC protocol can yield significant performance advantages over a generic baseline, but designing one requires deep expertise in CC protocol design. In this paper, we propose HakiCC, an LLM-driven multi-agent pipeline that automatically designs, verifies, and optimizes concurrency control protocols tailored to a given target application. HakiCC provides a two-stage pipeline. In Stage 1, a multi-agent system takes a workload description as input and generates an application-specific CC protocol implementation, which is iteratively repaired and verified for conflict-serializability. In Stage 2, the verified protocol is further optimized for that application through an LLM-driven evolutionary loop targeting correctness and throughput. We evaluate HakiCC on TPC-C and AuctionMark as target workloads, producing and reporting ten application-specific CC protocols. All ten are conflict-serializable after Stage 1; Stage 2 improves throughput for every protocol, with average gains of +50.6% for TPC-C protocols and +92.2% for AuctionMark protocols.

cs.DB↗