arXiv · 2106.01595
Position Heaps for Cartesian-tree Matching on Strings and Tries
Abstract
The Cartesian-tree pattern matching is a recently introduced scheme of pattern matching that detects fragments in a sequential data stream which have a similar structure as a query pattern. Formally, Cartesian-tree pattern matching seeks all substrings $S'$ of the text string $S$ such that the Cartesian tree of $S'$ and that of a query pattern $P$ coincide. In this paper, we present a new indexing structure for this problem called the Cartesian-tree Position Heap (CPH). Let $n$ be the length of the input text string $S$, $m$ the length of a query pattern $P$, and $σ$ the alphabet size. We show that the CPH of $S$, denoted $\mathsf{CPH}(S)$, supports pattern matching queries in $O(m (σ+ \log (\min\{h, m\})) + occ)$ time with $O(n)$ space, where $h$ is the height of the CPH and $occ$ is the number of pattern occurrences. We show how to build $\mathsf{CPH}(S)$ in $O(n \log σ)$ time with $O(n)$ working space. Further, we extend the problem to the case where the text is a labeled tree (i.e. a trie). Given a trie $T$ with $N$ nodes, we show that the CPH of $T$, denoted $\mathsf{CPH}(T)$, supports pattern matching queries on the trie in $O(m (σ^2 + \log (\min\{h, m\})) + occ)$ time with $O(N σ)$ space. We also show a construction algorithm for $\mathsf{CPH}(T)$ running in $O(N σ)$ time and $O(N σ)$ working space.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Akio Nishimoto, Noriki Fujisato, Yuto Nakashima, Shunsuke Inenaga. 2021-08-14. Position Heaps for Cartesian-tree Matching on Strings and Tries. https://arxiv.org/abs/2106.01595
Cite the original work for its findings. Save a collection to share your selection of sources.