arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,351 records · Page 75Linked to original sources

Repurposing Obsolete Representations for Post-Deployment Adaptation

Deep neural networks are increasingly deployed in long-lived systems, where task requirements may change after training. In such settings, part of the original output space may become obsolete: a class, prediction region, or learned behaviour may no longer be valid. Existing approaches either leave the obsolete behaviour intact or require fine-tuning, which can be expensive. We propose Deep Repurposing (DR), a post-hoc framework for adapting models under task obsolescence. DR estimates the latent geometry of obsolete and retained regions, removes obsolete-supporting components, and reallocates retained-compatible evidence through an analytic repair map without gradient updates. This yields repaired predictions and representations in which obsolete regions no longer act as valid outputs, while useful obsolete structure can support the retained task. Across multiple task settings, DR removes obsolete behaviour while preserving retained utility. More importantly, across classification benchmarks, DR matches or exceeds competing unlearning and editing baselines in retained accuracy, eliminates obsolete predictions, and adapts up to $60\times$ faster than competing unlearning methods.

cs.LG↗

Constraining the density distribution of the cold Circumgalactic Medium of quasars at z>2 through Lya, HeII and Ha emission

Extended Lya emission is routinely detected around high-redshift quasars, revealing large reservoirs of cool circumgalactic medium (CGM). Its resonant nature complicates interpretation of gas density, ionization state, and kinematics, making non-resonant tracers such as HeII and Ha essential. Extending previous analytical models, we use CLOUDY photo-ionization simulations to predict HeII/Ha in quasar-illuminated gas over a broad parameter range, including log-normal density distributions. The ratio constrains density fluctuations, commonly quantified by the clumping factor C. We compare the models with extended Lya, HeII, and Ha observations around quasars at z~2.1-4.5. At z~2.2, we observed 13 quasars with Keck/KCWI and targeted the systems with the most prominent extended Lya and HeII emission with Keck/MOSFIRE, providing the first statistical quasar-CGM sample with combined Lya, HeII, and Ha measurements. Within matched apertures out to ~35 pkpc, stacked spectra give Lya/Ha=9.1, HeII/Ha=0.54, and HeII/Lya=0.06. Single-density models require very high CGM densities (n>10 cm^-3) or implausibly soft AGN spectra to reproduce the low HeII/Ha ratio. Broad log-normal density distributions instead reproduce it at lower median densities but require a substantial emissivity-weighted high-density tail, corresponding formally to C~10^3-10^4 in our parameterization. We also analyze 26 VLT/MUSE quasars at 3.1<z<4.5, where Ha is inaccessible from the ground, finding 14 systems with co-spatial extended Lya and HeII emission. The detected, emission-weighted HeII/Lya component shows no evidence for strong evolution between z~2.2 and z~3.7, though this does not imply an invariant underlying density PDF. These constraints provide a basis for comparison with theoretical models and hydrodynamical simulations of the physical processes shaping quasar CGM at high redshift.

astro-ph.GA↗

From Modal Memory to Topological Relaxation in a Driven Holographic Superfluid

We study how driving history affects the nonlinear topological relaxation of a holographic superfluid. Different histories are constructed to reach the same present driving conditions and then follow the same future protocol. Although they relax to the same final winding sector, their late-time dynamics remain strongly history-dependent: the residence time of the penultimate winding sector changes by about a factor of \(3.7\). This variation is organised by the modal imbalance carried into the common amplification stage. Histories with different shapes but matched modal imbalance produce nearly identical relaxation times. The same qualitative behaviour is found for a second pair of competing modes.

hep-th↗

Streaming algorithms for robust max-min diversification

Given a set of $n$ points $X$ in a metric space and an integer $k$, max-min diversification aims to select $k$ points of $X$ maximizing their minimum pairwise distance. This objective function is however highly vulnerable to noisy points. In[Amagata, AAAI23], a robust formulation is proposed which addresses this vulnerability by excluding solutions containing any of $z$ outliers, defined as the $z$ points in $X$ with the largest nearest-neighbor distances. That paper also presents a coreset-based streaming algorithm for the new formulation, based on a suitable inlier-outlier separation assumption. However, we identify three shortcomings in the algorithm by [Amagata, AAAI23]: its coreset construction requires an offline computation over $X$, which needs memory linear in $n$, in stark contrast with the typical goals of stream processing; the one-pass procedure used to extract the solution from the coreset may return fewer than $k$ points (hence, an unfeasible solution) because it permanently discards points too far from the current solution; and its outlier-exclusion guarantee is only probabilistic and weakens as the coreset size shrinks. In contrast, we present a deterministic coreset-based algorithm that, under a natural inlier-outlier separation assumption (similar to the one used in [Amagata, AAAI23]), returns exactly $k$ inliers which are a $(2+\varepsilon)$-approximate solution, for any $\varepsilon>0$, thus only $\varepsilon$ above the best polynomial-time sequential approximation, even without outliers. Its one-pass streaming implementation adapts obliviously to the dataset's doubling dimension $D$ and, for wide ranges of $k$, $z$, $\varepsilon$, and $D$, it uses memory independent of $n$. For sufficiently long streams, its amortized update time is proportional to the coreset size, thus also independent of $n$.

cs.LG↗

Disk-Like Matter Outflow from a Kerr-Type Wormhole: Possible Observational Manifestations and Magnetic Effects

We present a phenomenological study of a hypothetical compact object with a traversable 2 Kerr-type wormhole geometry connecting two asymptotically distinct regions. One mouth of the wormhole is assumed to reside in an active galactic nucleus (AGN) environment with ongoing accretion, while the second mouth is located in a comparatively quiescent galactic nucleus without accretion activity. In such a configuration, matter accreted near the AGN-side mouth may traverse the wormhole throat and emerge from the opposite mouth. Because the inflowing matter possesses angular momentum, the resulting outflow is expected to be predominantly equatorial, producing a disk-like structure fundamentally different from a conventional accretion disk around a black hole. Several dynamical regimes are considered depending on the outflow velocity relative to the local escape and orbital velocities. Possible observational manifestations, thermal properties, recombination signatures, and magnetic field effects are discussed. Special attention is given to the possibility of jet formation in the absence of a conventional accretion disk. We explicitly distinguish the weak-field, large-radius Newtonian classification used for the outer flow from the strong-field region near the throat, where the dynamics should be treated in a general-relativistic and magnetohydrodynamic framework. We further discuss the specific energy and angular momentum inherited from the inner edge of the accretion flow, the possible modification of these quantities by magnetic stresses and energy-extraction processes, and the effects of an intrinsic dipolar magnetic field of the wormhole.

astro-ph.HE↗

Rethinking Probability-Based Reinforcement Learning From Posterior Concentration

Verifier-free reinforcement learning with probability-based rewards offers a promising way to train LLMs on general reasoning tasks where external verifiers are unavailable. Yet the reliability of these rewards, especially in long-horizon reasoning, remains underexplored. This work identifies a length-dependent failure mode of probability rewards, which we call the Posterior Concentration Phenomenon (PCP). We show that the probability of a reference answer conditioned on a reasoning trace often collapses to a low-variance interval as the trace becomes lengthy. This phenomenon results in nearly indistinguishable rewards, which, under GRPO-based settings, makes probability-based policy optimization unstable and inefficient. Motivated by this, we propose Reinforcement Learning with Concentration-aware Posterior Rewards (RLCPR), a verifier-free RL framework to explicitly account for PCP for better optimization stability and token efficiency. It has two components: uncertainty-aware data sampling, which reduces concentration-prone rollouts before generation, and concentration-aware regularization, which penalizes unnecessarily long traces when posterior rewards collapse. Extensive experiments show that, alongside higher token efficiency, RLCPR outperforms the state-of-the-art verifier-free RL baseline by up to 4.0% on six of seven benchmarks, including general-domain and mathematical reasoning challenges.

cs.AI↗

Tight Transition Time Bounds for Separable Logistic Regression at the Edge of Stability

We study logistic regression on linearly separable data under gradient descent with a large constant stepsize $η$. Such dynamics may exhibit a characteristic Edge of Stability phenomenon, in which the loss initially oscillates before transitioning to a stable phase of monotone decrease. Existing work provides a tight $Θ(1)$ bound in dimension $d=2$ as $η\to \infty$ and conjectures a bound independent of $η$ in arbitrary dimensions $d\geq 2$. In this paper, we disprove this conjecture by showing that, for every fixed sample size $n\geq 2$ and sufficiently small margin $γ$, the worst-case transition time is $$Θ\!\left((\logη)^{\min\{n-2,d-2\}}\right)$$ uniformly over $d\geq2$. The key challenge in establishing a tight bound is that the sample contributing most strongly to the gradient can change repeatedly across iterations. To address this issue, we control such changes by induction on dimension and sample size, and construct matching hard instances.

cs.LG↗

Free groups amenably act on unital simple AF-algebras

We show that free groups admit amenable actions on unital simple AF-algebras. This settles, for free groups, the last remaining case of the existence problem for amenable actions of non-amenable groups on classifiable simple C*-algebras. The proof relies on the Baire category theorem. This result was obtained with substantial assistance from GPT-6 Astra.

math.OA↗

NextMe-800: Anticipating Personal Behavior from Months of Egocentric Video

We often plan ambitiously yet act habitually and wonder, in retrospect, whether we would have planned differently had we known what we would actually do. Hindsight offers a valuable perspective on past decisions, although we often wish we could have simulated hindsight at the moment of choosing. If a system could generate plausible trajectories from one's personal history, such previews might help people formulate more realistic plans and make better informed decisions. We introduce NextMe-800, an approximately 800-hour first-person dataset from one volunteer over 126 days with 1 Hz images, gaze, and audio, captioned at five hierarchical abstraction levels from atomic actions to major activities. We formulate personalized action anticipation as open-vocabulary K-step sequence prediction and construct NextAct, a 1,500-point benchmark combining NextMe-800 with the multi-person EgoLife dataset. Using an embedding-based soft edit distance as the metric, we evaluate how well different models can anticipate personal behavior across abstraction levels and prediction horizons. NextMe-800 and NextAct provide a months-long resource and evaluation framework for studying how far ahead personal behavior can be anticipated from egocentric observation.

cs.AI↗

Unbounded separation between definite and indefinite causal order in finite-dimensional quantum metrology

Indefinite causal order is known to offer enhancements in quantum metrology, notably including an unbounded advantage in the measurement of a geometric phase of a harmonic oscillator. This advantage, however, is specific to the infinite dimensional setting, and its finite dimensional analogue remains elusive, with recent findings suggesting that advantages in finite dimensions may be fundamentally limited to bounded constant factors. Here we show that, in fact, arbitrarily large advantages arise for finite dimensional systems in the finite sample regime. Specifically, we establish an unbounded separation between definite and indefinite causal order in the estimation of a geometric phase associated to two sets of $N$ displacements generated by discrete position and momentum operators on a $d$-dimensional quantum system: for any given constant $R$, there exist values of $N$ and $d=Ω(N^2)$ such that a strategy with indefinite order uses an initial probe with $R$ times less energy than the probe required by every strategy with definite order achieving the same mean squared error, whenever the number of measurement shots $ν$ is bounded as $ν=\mathcal{O}(\exp(πd/16)/\mathrm{poly}(d))$. In other words, indefinite order offers an energy saving that grows arbitrarily large with the parameters of the problem. To prove this result, we establish an approximate Weyl relation for discrete Gaussian wavepackets, which is of independent technical interest.

quant-ph↗

Learning to structure data from user-generated thematic corpora

Thematic corpora, such as social media communities, contain unstructured text describing data that could be made structured. These include, for example, personal attributes, behaviors, and experiences mentioned in social media data. Extracting structured data is challenging as relevant attributes are often implicit, domain-dependent, and unknown in advance. We propose a fully automated, iterative framework for discovering and extracting domain-specific attribute schemas without a predefined ontology. Using large language models (LLMs), the framework induces candidate attributes, sequentially consolidates semantically overlapping attributes, and assigns a structural type. These enable creating an ontology and populating it with values from the corpus. The framework also enables the use of smaller LLMs for value extraction with estimable accuracy loss compared to large LLMs. We evaluate the framework on 5 health-related Reddit communities. Discovered attributes achieved 61% agreement with human-identified attributes, close to the 62% agreement between independent annotators. In most cases, the algorithm converges to a stable attribute set in fewer than 10 iterations. Structural type assignment achieves 82% accuracy, and value extraction reaches an F1 score of 0.8 compared to human annotations. Across four LLM families, smaller instruction-tuned models show statistically significant improvements in extraction performance with model scale when evaluated against a high-capacity reference LLM, supporting informed accuracy-cost trade-offs. These results show that attributes comparable to those identified by humans can be discovered automatically, enabling the creation of high-quality structured datasets economically and at scale. By removing the need for predefined ontologies, iterative model-driven schema induction offers a practical and scalable foundation for mining thematic corpora.

cs.IR↗

Inference after data-driven control-unit selection in difference-in-differences with estimated covariance

In difference-in-differences (DiD), researchers may use pre-treatment trends to select a control group for which the parallel-trends assumption appears plausible, with the aim of estimating the average treatment effect on the treated (ATT). Our earlier paper,Nakano and Hoshino (2016), and the present paper jointly provide the first selective-inference approach to the ATT that explicitly accounts for this control selection. We generalize our exact Gaussian procedure with known covariance to allow the covariance matrix to be estimated from the same individual-level data used for control selection and DiD estimation. We use this estimate to compute the variance, conditioning direction, residual, and truncation set. With fixed numbers of regions and periods, we establish uniform conditional coverage for selection events with probabilities bounded away from zero, and marginal coverage of the selected target without that restriction. We allow unequal regional sample sizes, heterogeneous covariances, ties in population fit, and regional sample shares that converge to zero. We establish asymptotic equivalence between the plug-in and known-covariance interval endpoints and derive rates for interval length. For staggered adoption, the control pools may differ across cohorts and periods, controls may be not yet treated, observations may be reused, and treatment effects may be heterogeneous. We also construct inference conditional on unions of selection paths that leave the reported parameter unchanged, together with simultaneous confidence bands for finitely many event-time effects. Under parallel trends and the other identifying conditions, the coverage results apply to the ATT. We give sufficient sampling conditions for individual panels and independent repeated cross-sections.

econ.EM↗

A distributional modelling approach with application to electricity price forecasting

The increasing volatility of electricity prices driven by renewable energy integration, market shocks, and regulatory changes has reinforced the need for forecasting methods that go beyond point predictions and accurately describe the full conditional price distribution. This paper applies the Generalised Additive Models for Location, Scale and Shape (GAMLSS) framework to forecast Spanish day-ahead electricity prices using hourly data from 2020 to 2024. Alternative specifications based on Normal, Johnson's SU (JSU), and Sinh-Arcsinh (SHASH) distributions are considered, allowing the location, scale, and shape parameters to vary with market fundamentals, including electricity demand, renewable generation, seasonal effects, and regulatory and geopolitical risk factors. Forecasts are generated using a rolling-window approach and evaluated through the mean absolute error (MAE), pinball loss, and Diebold-Mariano tests. The results show that flexible distributional specifications improve forecasting performance relative to a naive benchmark and the standard normal specification. While SHASH and JSU specifications provide the lowest point forecasting errors, the hourly analysis reveals substantial intraday variation in relative performance across specifications. JSU specification with all four parameters driven by covariates achieves the best probabilistic forecasting performance, particularly in the tails of the distribution. Diebold-Mariano tests confirm the statistical significance of these improvements. These findings highlight the importance of modelling time-varying shape distributional parameters and demonstrate the value of GAMLSS models for forecasting and risk management in increasingly volatile electricity markets.

econ.EM↗

Zero- Versus Infinite-Temperature Damping in Variational Quantum Circuits: Feature Scale, Sampling Cost, and Frame Gauge

The Pauli twirl of amplitude damping (AD) is generalized amplitude damping at infinite temperature: it keeps the contraction of AD and removes its non-unital term, so comparing the two in variational circuits isolates the zero-temperature bias, which acts mainly through the scale of the features. For random parameters, features under AD settle on a floor, which at strong damping is set by the last layer, has a closed form, and at fixed $T_1$ falls with temperature as $\tanh(\hbarω/2k_BT)$; under the twirls they shrink by a constant factor per layer, up to eight qubits. A trainable output scale removes most of the resulting accuracy differences, leaving AD ahead of its twirls by at most about three percentage points in our simulations; what it removes reappears as a cost in measurement shots: trained and tested with $10^3$ shots per image, a four-qubit classifier under AD at $p=0.3$ stays within 1.5 points of noiseless accuracy, while the twirled classifiers lose up to 33. At weaker damping the separation depth grows roughly as $(np)^{-1}\ln(1/p)$. The damping direction is a gauge when the damping follows complete entangling layers and the circuit boundaries are trainable; in an eigensolver it becomes physical inside a decomposed two-qubit gate.

quant-ph↗

A monodromy relation for the Cartwright-Steger fibration

We study the genus-$19$ Albanese fibration of the Cartwright--Steger surface and its order-three symmetry. Passing to the orbifold quotient of the elliptic base gives a relation among the two handle monodromies and the three Dehn twists about the vanishing cycles. The two real vanishing paths give cycles whose real fixed points lie on different ovals of the invariant fiber. For a cyclic triple cover, we express the intersection of a curve with its image under the deck transformation as a signed crossing count on the quotient. We determine the two local branch values of a degree-$72$ bicanonical pencil near a node and study a symmetric bicanonical pencil. A general pencil also gives a finite degree-$72$ map from a blowup of the surface to the product of the elliptic curve and a projective line. The monodromy translates of the nodal vanishing classes span the first homology of a smooth fiber.

math.GT↗

LESS: Lightweight Evolutionary Supernet Search in Minutes

Low-cost NAS must both explore high-performing architectures and identify them reliably, yet reducing evaluation cost often weakens the fidelity of candidate comparisons. Training-free methods reduce evaluation cost by replacing learned task feedback with proxy signals measured at initialization. We introduce LESS (Lightweight Evolutionary Supernet Search), a data-driven method that combines a brief fair hard-path warm-up with discrete search under a single CMA-ES distribution. Each proposal is evaluated as its decoded hard genotype after six candidate-conditioned supernet updates. On NAS-Bench-201, LESS achieves \(93.189\pm0.467\%\) CIFAR-10 test accuracy in 409.1 seconds, coming within 0.04 percentage points of FairNAS using approximately \(1/24\) of its source-reported search time. Matched controls show that calibration improves selected validation accuracy by \(0.577\) percentage points while changing best-visited accuracy by only \(0.054\) points, indicating that its primary effect is to reduce selection regret. The frozen configuration transfers without tuning to CIFAR-100 and ImageNet16-120 with \(69.615\pm1.139\%\) and \(43.720\pm1.697\%\) accuracy. Applied without tuning to the larger DARTS space, LESS achieves \(96.95\pm0.14\%\) on CIFAR-10 and \(82.43\pm0.80\%\) on CIFAR-100, with each search completing in approximately 43.5 minutes on a single GPU. Together, these results show that short, balanced, data-dependent updates enable competitive neural architecture search across datasets and search spaces within minutes.

cs.NE↗

A power-commutator presentation for the group of truncated polynomials under substitution

Given a prime $p$, the Nottingham group $G(p)$ is the group of formal power series $x+a_2x^2+a_3x^3+\cdots$, $a_i\in{\bf Z}/p{\bf Z}$, under substitution. For $n\geq 1$, we write $G_n(p)$ for the quotient of $G(p)$ by its normal subgroup $K_n=\{x+a_{n+1}x^{n+1}+a_{n+2}x^{n+2}+\cdots\,|\, a_i\in {\bf Z}/p{\bf Z}\}$, namely the group of truncated polynomials $x+a_2x^2+\cdots+a_nx^n$, $a_i\in {\bf Z}/p{\bf Z}$, under substitution, which is a $p$-group of order $p^{n-1}$ generated by $x+x^2,\dots,x+x^n$. In this paper, we determine the power-commutator presentation of $G_{15}(p)$ relative to $x+x^2,\dots,x+x^{15}$. This, in turn, automatically gives the power-commutator presentation of $G_{n}(p)$ relative to $x+x^2,\dots,x+x^{n}$ for any $1\leq n<15$, as well as the initial segments of the power and commutator relations in $G_{n}(p)$ corresponding to $x+x^2,\dots,x+x^{15}$ for any $n>15$.

math.GR↗

Constraining Particle Stiffness in Asteroid Regolith from Early-Time High-Speed Penetrator Dynamics under Local Granular Variability

Mechanical properties of rubble-pile asteroid regolith remain poorly constrained because contact responses depend on both material properties and local particle configuration. This study examines whether early-time high-speed penetrator dynamics can constrain particle stiffness, represented by the particle Young's modulus E, within a controlled discrete-element model. A bidirectionally coupled EDEM-Adams discrete-element-multibody model is used to screen six parameters: Young's modulus, Poisson's ratio, coefficient of restitution, static friction coefficient, rolling friction coefficient, and adhesion strength. Five Young's-modulus levels are then tested at the same 17 spatial locations in one settled granular bed, giving 85 simulations with location-matched comparisons across modulus levels. Within the 75-90 m s^-1 speed range, varying Young's modulus yields the clearest and most systematic differences in probe deceleration, whereas the other parameters have weaker effects over the tested ranges. Validation by withholding all modulus cases at one location in turn indicates that the early response is more informative for distinguishing broad modulus ranges than adjacent modulus levels. Mesoscale analysis links the response differences to concentrated interface load sharing and spatially extended three-dimensional high-capacity contact paths as Young's modulus increases. The results provide a numerical basis for narrowing the particle Young's-modulus range from high-speed penetration responses despite local granular variability. Quantitative application to real asteroid regolith requires further experimental validation.

astro-ph.EP↗