arXiv Science⌕ Search

arXiv · 2610.08576

Subset selection for matrices by volume sampling

Abstract

We address the Subset selection problem for matrices, where the goal is to select a subset $\mathcal{S}$ of $k$ column indices from a \enquote{short-and-fat} matrix $X \in \mathbb{R}^{m \times n}$, such that the sampled submatrix $X_{\mathcal{S}}$ has $\|X_{\mathcal{S}}^†X\|_F$ as small as possible. Our approach is centered on volume sampling, which attains the tightest known bound on this objective in expectation. As our primary contribution, we propose a new deterministic algorithm, Forward derandomized volume sampling (FDVS), which provably attains this bound and has asymptotic complexity $O(nkm)$. In contrast, the complexity of all previously known algorithms with this guarantee scales at least quadratically in $n$ when $k \ll n$, making FDVS particularly attractive for very wide matrices. In addition, we systematically structure the landscape of volume sampling methods: we propose a simple $O(nm^2)$ algorithm for exact forward volume sampling, show that a known deterministic method is in fact a derandomization of Reverse iterative volume sampling, and derive a modification of the latter that avoids repeated SVD downdating, along with a fast greedy forward variant. The algorithms are verified and compared in numerical experiments involving optimal experimental design, sensor placement, and DLRA-DEIM.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ivan Kozyrev, Alexander Osinsky. 2026-10-06. Subset selection for matrices by volume sampling. https://arxiv.org/abs/2610.08576

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Interface Energy and Phase Transformations: A Comparative Analysis of Cahn-Hilliard and CALPHAD-based Models in Ternary Substitutional Alloys

Diffusional phase transformations in alloys are modeled either by phase-field methods of Cahn-Hilliard type, in which a double-well free energy and a gradient term impose an interface energy, or by reactive diffusion models, in which the driving forces follow from the CALPHAD methodology, the convex hull of the molar free energy, and the interface carries no energy. Their numerical schemes are tied to their free energies: phase-field schemes break down when the interface energy is set to zero, and reactive diffusion codes rely on finite differences. We derive, from the thermodynamic extremal principle, a single finite element scheme that covers both limits: the stationarity of the Rayleighian functional (free energy rate plus one half of the dissipation, with mass conservation imposed by multipliers) with respect to the rates, fluxes and multipliers at a frozen state. Fluxes and chemical affinities are primary unknowns, so no derivative of the affinity is required. The time-discrete scheme is the corresponding incremental minimization; it conserves mass exactly and satisfies a discrete energy-dissipation inequality for every convex molar free energy, including the convex hull without regularization. We make explicit that this last case is a degenerate parabolic problem of Stefan type, whose degeneracy can be removed, without altering the free energy, only by a moving mesh; since our aim is the coupling to continuum mechanics and damage on a fixed mesh, we operate at the boundary of degeneracy and show that mass conservation, the energy balance and the second law remain intact there. Thermodynamic consistency is proved for multi-component systems, a scheme relating Onsager and diffusion coefficients is proposed, and the role of interface energy is studied in binary and ternary examples, the latter with vacancies as a non-conserved component.

math.NA↗

A posteriori existence for the Keller-Segel model via a finite volume - finite element scheme

We derive two forms of conditional a posteriori error estimates for a finite volume scheme approximating the parabolic-elliptic Keller-Segel system. The estimates control the error in the $L^\infty(0,T, L^2(Ω))$-norm and exhibit linear convergence in the mesh size, as observed in numerical experiments. Crucially, we show that, as long as the condition of the error estimate is satisfied, a weak solution exists. This means, as long as the numerical solution has good properties, we can rigorously infer existence of an exact solution.

math.NA↗

A nonlocal model for heterogeneous material flow on conveyor belts

In this paper, a finite volume approximation scheme is used to solve a nonlocal macroscopic material flow model in two space dimensions, accounting for the presence of boundaries in the nonlocal terms. Based on a previous result for the scalar case, we extend the setting to a system of heterogeneous material on bounded domains. We prove the convergence of the approximate solutions constructed using the Roe scheme with dimensional splitting. We consider a regularized version of the flux function and establish BV bounds for a Roe-type scheme. Numerical tests show a good agreement with microscopic simulations.

math.NA↗