arXiv Science⌕ Search

arXiv · 2610.02534

TREMOR: Template Matching for Large Seismic Data Collections

Abstract

Seismic station networks continuously record the ground velocity at several locations on earth in the form of waveform data series, which seismologists analyze to detect various kinds of geophysical events, including earthquakes. Some of these events are of particular interest: they are called templates and are used to search the seismic data collections for matching, similar events. This is known as template matching, and is a fundamental task in seismology, serving as the backbone for various seismic analyses. However, template matching requires extensive processing times, especially for seismic collections that exceed the memory capacity of a single machine. This poses a significant challenge to seismologists and is becoming worse as the seismological datasets continue to grow in size. In this paper, we introduce TREMOR, a distributed data series processing framework for template matching, designed to efficiently handle large waveform collections. We apply TREMOR to two representative real-world seismic use cases for template matching and, through an extensive experimental evaluation, we demonstrate its efficiency, with TREMOR being up to 12x faster than the best competing method, while returning the same, exact results.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Manos Chatzakis, Rodrigo Flores-Allende, Mikael Freire, Yoann Cano, Léonard Seydoux, Themis Palpanas. 2026-10-01. TREMOR: Template Matching for Large Seismic Data Collections. https://arxiv.org/abs/2610.02534

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Query Performance Tuning with Optimal Exploration of Optimizer Cost Model Parameter Space

Modern query optimizers use analytical cost models to estimate the cost of a given query plan. Such cost models are typically functions of a set of "cost units" that specify unit CPU cost when processing a row or unit IO cost when accessing a disk page. These cost units are traditionally viewed as platform-dependent constants, that is, they require a one-shot calibration when a database is deployed on a hardware/software platform, but are fixed afterward regardless of the query being optimized for. Some very recent work has taken a different perspective by viewing these cost units as tunable parameters that we call "cost model parameters (CMPs)" in this paper. However, so far there is no approach that offers any optimality guarantee for the tuning results. We present a new approach to systematically explore the query plan space spanned by the CMPs and find the best plan in terms of execution time. Compared to alternative exploration approaches that use random search (RS) or Bayesian optimization (BO), our new approach is deterministic and, therefore, avoids the undesirable instability that is inevitable when applying RS or BO. Moreover, it is guaranteed to find all candidate plans in the query plan space without suffering from the overhead of an exhaustive enumeration. We also present a set of optimization techniques to reduce the overall evaluation time spent on executing the candidate plans found, a factor that is often overlooked by previous work but is critical from a practical point of view. Experimental evaluation on top of PostgreSQL and Microsoft SQL Server demonstrates the efficacy of tuning the CMPs, which can find query plans that are orders of magnitude faster in execution time than the ones found by RS or BO.

cs.DB↗

RaBitQ-SSD: Split Codes and Pipelined I/O for SSD-Resident Vector Search

SSD-resident approximate nearest-neighbor search is essential when vector collections exceed DRAM capacity. The challenge is to reduce SSD reads and hide I/O latency through concurrent reads and overlap with computation. However, for graph-based search, progressive candidate discovery limits advance I/O planning while IVF search reads inverted lists in full, incurring unnecessary SSD reads. In this paper, we present RaBitQ-SSD, an extension of IVF-RaBitQ for SSD-resident vector search. Specifically, we propose Split-RaBitQ, which retains a configurable prefix of each binary code in DRAM, allowing in-memory code storage below one bit per dimension. Using this partial representation, it provides an unbiased distance estimator with a probabilistic error bound. Once the in-memory coarse quantizer identifies candidate lists, this bound supports pruning individual candidates within them, enabling finer-grained SSD access. We also design an asynchronous search pipeline that coordinates candidate pruning with SSD read scheduling to reduce unnecessary reads while overlapping I/O with computation. On datasets ranging from $5$ million to $1$ billion vectors, RaBitQ-SSD delivers up to $1.74\times$ the throughput of graph-based baselines at $90\%$ recall, while reducing SSD page reads by up to $3.8\times$. Its on-SSD indexes are up to $7.0\times$ smaller and $10.0\times$ faster to build than DiskANN. We also build and search an index of $10$ billion vectors on a single machine with one SSD, achieving over $3{,}000$ queries per second at $90\%$ recall.

cs.DB↗

Tracking State Footprints: How Agents Can Transact

Can AI agents transact? We argue that they must: as multi-agent systems (MASs) increasingly write code, deploy infrastructure, modify databases, and call web services, lost updates or stale reads can be catastrophic. Although MASs increasingly execute plans in parallel, current orchestrators do not track the state that agents read and write. As a result, concurrency anomalies manifest even in simple coding tasks. We frame MAS coordination as a data management problem and propose to describe agents by their state footprint: the state they read and write across their own local context and state, as well as the state of the orchestrator and external systems. We posit that MASs require guarantees similar to those of databases, but providing them raises new challenges and opportunities: unlike database transactions, agents do not read from a fixed schema or an isolated snapshot, and cannot be replayed deterministically upon failure. They can, however, resolve conflicts semantically instead of aborting, enabling new forms of concurrency control and conflict resolution. Towards agents that can transact, we outline a vision for next-generation agent orchestrators and transactional interfaces for external systems to participate in agentic transactions.

cs.DB↗