arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,135 records · Page 63Linked to original sources

Reflex-Guard: A Low-Latency Guardrail for LLM Prompt Safety Using Dense Semantic Embeddings

Large Language Models (LLMs) in real-world applications often face the risks of specially crafted prompts designed to bypass the safety controls. Existing guardrail methods, such as LLM-as-a-judge and cloud-based safety APIs are able to detect unsafe content. However, they often add a delay of about 250-900 ms to each request. This delay is too high for real-time applications, when the system usually needs to respond in less than 100 ms. Furthermore, routing user prompts through external moderation endpoints raises significant data privacy concerns. This paper introduces Reflex-Guard, a lightweight guardrail that runs locally. It uses jailbreak-aware preprocessing, compact sentence-transformer embeddings, and seven fast binary classifiers. Together, these components enable high-accuracy prompt safety filtering with much lower latency than existing solutions. Through systematic evaluation on a strategically balanced dataset of 30,568 samples drawn from five complementary sources, we demonstrate that Reflex-Guard achieves 95.9% recall on harmful prompts at 37.6 ms end-to-end latency. It is faster than existing baselines, including Llama Guard 2 at 255 ms and SafeDecoding at 723 ms. It can detect 100% of GCG suffix attacks and Base64-encoded prompts using the default threshold. However, DrAttack structured prompts required lowering the threshold to 0.03 for optimal detection, as they produced a distinct probability distribution. Reflex-Guard achieves Reflex Efficiency Score (RES) scores up to 16.79, significantly outperforming Llama Guard 2 (11.90) and SafeDecoding (9.80). This analysis offers practical deployment advice and shows that different attack types occupy distinct regions in the embedding probability space.

cs.CR↗

Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection

Online log anomaly detection is critical for maintaining the reliability of large-scale computing systems. Although recent language model-based log anomaly detectors achieve strong detection performance, their confidence estimates remain poorly calibrated. We show that these detectors frequently assign excessive confidence to incorrect predictions, particularly for anomalous logs under severe class imbalance. Moreover, confidence on erroneous predictions remains persistently high even when conventional calibration metrics indicate good calibration, creating a critical reliability gap for operational monitoring systems. To address this issue, we propose Log Reconstruction and Distance (LoRD), a lightweight post-hoc calibration framework for reliable log anomaly detection. LoRD learns prediction-route-specific reliability models from latent representations of correctly classified validation samples and estimates prediction reliability through route-wise reconstruction distances. Based on the estimated reliability, LoRD selectively recalibrates high-risk predictions to suppress overconfident errors while preserving reliable predictions. Extensive experiments on four large-scale log benchmark datasets and multiple language model-based detectors demonstrate that LoRD consistently improves confidence reliability and substantially reduces overconfident anomaly-related errors without sacrificing anomaly detection performance.

cs.LG↗

Rational Solids with Equal Surface Area and Volume

We classify pairs consisting of a right square pyramid and a right square prism, each having rational base side length and rational height, for which the two solids have equal surface area and equal volume. We prove that, up to scaling by a common positive rational factor, there is a unique such pair. The problem reduces to determining the rational points on a genus 2 bielliptic curve whose Jacobian has rank 2. We accomplish this using a quadratic Chabauty computation followed by the Mordell--Weil sieve. The same analysis gives an analogous classification for right circular cones and right circular cylinders.

math.NT↗

Beyond Forgetting: Diagnosing and Harnessing Shared Reasoning in Continual RLVR

Reinforcement learning with verifiable rewards (RLVR) commonly post-trains reasoning models on multiple tasks, while rerunning multitask RLVR (MTRL) as new tasks are added makes capability expansion costly. We therefore study continual RLVR, which updates the existing model as each task arrives. The central question is whether a model updated this way can perform as well as a jointly trained model. To answer this question, we introduce Continual Reasoning Gym, a continual-RLVR environment that organizes text and visual reasoning tasks into five task sequences. In this setting, we identify two key observations: Sequential RLVR exhibits modest forgetting, yet its final performance remains below that of MTRL. To understand the latter, we decompose final performance and show that forgetting accounts for only part of the gap. To explain the former, we identify shared reasoning: transferable reasoning structure allows training on one task to support others on average. We therefore introduce Continual Prompt Replay (CPR), which harnesses shared reasoning to improve learning on the arriving and future tasks by replaying previous-task prompts and regenerating their responses with the current policy. On average, only CPR reaches MTRL-level performance.

cs.LG↗

Metrics That Write Themselves: Evolving an Evaluator from Its Own Blind Spots

Agents improve quickly against a reliable automatic metric and stall without one, and the applications that need them most, report generation among them, are the ones nobody knows how to score. Can the metric write itself? Saying what makes an answer good is hard; pointing at something wrong with one is easier, so the metric we evolve is a pool of small Python operators that each flag a candidate for one named defect, or abstain, and vote. Asking a model for operators directly does not work: 183 candidates realise only 96 distinct behaviours, from one narrow region of an enormous space. EvalCEGAR instead borrows counterexample-guided abstraction refinement from program verification. It reads the pool as an abstraction and searches for a collision, two answers the operators score identically, one correct and one not. That pair, not a prompt, is the authoring request, and when a collision defeats every attempt the loop widens what an operator may read rather than resampling. On MBPP+ and HumanEval+, a sandbox whose hidden unit tests give exact ground truth, the loop writes a 55-line operator that closes 15.4% of the gap between flagging nothing and a perfect filter on 428 unseen tasks (+0.0065, p=0.0010) at a quarter of our best hand-written operator's flags. On the benchmark it never saw it matches that operator's effect exactly on a third of the flags. Six of eight runs admit such an operator and all six help out of sample; our 15 hand-written operators applied together as one filter lose accuracy. An LLM judge on the same information ties that delta on a nearly disjoint set of candidates, and charges a model call per candidate forever where the operator charges none.

cs.AI↗

Cluster Assignments in Soft Targets Shape Speech Representations: Evidence from S-JEPA

Cluster-based prediction is widely used in self-supervised speech learning. A soft target preserves a distribution over clusters rather than a single label. This distribution specifies both the probability values and which clusters receive them. Comparisons between soft targets and hard labels do not separate the contributions of these two aspects to the learned representation. We study this in S-JEPA, a recent high-performing self-supervised speech model trained with soft Gaussian mixture model (GMM) targets. We compare its original targets with counterfactual targets that preserve the most likely cluster and all probability values but change which remaining clusters receive the other probabilities. Across three training seeds, the original soft distribution is recovered more accurately from Encoders trained with the original than counterfactual targets. Because this could reflect target matching alone, we also test low-level acoustic and phonetic information. Both are more accessible from Encoders trained with the original targets. This suggests that cluster assignments affect acoustic and phonetic properties of the learned representation, not just recovery of the training target.

cs.LG↗

Information-Computation Inversion in Pseudo-Marginal MCMC

Observation refinement changes posterior uncertainty and the likelihood calculation in pseudo-marginal MCMC. We compare their combined effect through finite-run squared-error risk. A sufficient inversion condition relates information gain to accepted event flow and coarse-kernel contraction. A bootstrap construction realizes inversion at every fixed particle count. We then couple cross-event proposals and refresh same-event proposals independently. A swap identity establishes invariance; continuation identities describe subsequent risk. In the same finite model, a rational certificate proves inversion against optimized constant mixtures and repair by selective allocation over an initialization class at a common action-price budget. Reaction-network experiments measure CPU costs. Under finite-pool initialization, selective allocation reduces event mean-squared error by 40.5% against a tuned mixture at 25 post-initialization CPU seconds. Paired transcription observations show how increased particle effort can raise finite-budget error.

stat.ME↗

Supergroup Gauged Linear Sigma Models and their Physical Mathematics

We construct 2d $\mathcal{N}=(2,2)$ gauged linear sigma models with $\mathrm{U}(1|1)^N$ supergauge group possibly with superpotential. Despite being nonunitary, one can still study their space of supersymmetric states and explore their applications to mathematics. In particular, we find a relation between a nonlinear sigma model on a Calabi-Yau complete intersection of hypersurfaces in a super-Grassmannian and a supergauged Landau-Ginzburg orbifold, which can reduce to a regular Calabi-Yau/Landau-Ginzburg correspondence for complete intersections. This defines a super-Grassmannian/supergroup generalization of the correspondence proved by Clader [1] and Zhao [2]. Similarly, we find a relation between a nonlinear sigma model on a Calabi-Yau hypersurface in a product of super-Grassmannians and a hybrid NLSM/supergauged Landau-Ginzburg orbifold, which can reduce to a regular hybrid Calabi-Yau/Landau-Ginzburg correspondence for hypersurfaces in product space. This defines a super-Grassmannian/supergroup generalization of the correspondence proved by Fan-Jarvis-Ruan [3]. We also find that Calabi-Yau supervector bundles over a super-Grassmannian can undergo a physically related mild topology change which is reducible to a regular Atiyah-type flop transition. This defines a super-Grassmannian generalization of a birational equivalence of Calabi-Yau vector bundles in mathematics. Similarly, we find that a Calabi-Yau complete intersection of quadrics in a super-Grassmannian can also undergo a physically related topology change which is reducible to a regular conifold transition. This defines a super-Grassmannian generalization of a homological projective duality for Calabi-Yau quadrics by Kuznetsov-Perry [4] in mathematics.

hep-th↗

CLaST: Context-aware Contrastive VAE for Probabilistic Time Series Forecasting

Probabilistic forecasting models are widely used for time series forecasting in domains such as energy systems, finance, medicine, and transportation. In recent years, deep generative models have shown strong results on probabilistic forecasting, yet many conventional approaches struggle to capture internal temporal dependencies, leading to latent representations with limited expressive power. To address this limitation, we propose \textit{CLaST}, a VAE framework for probabilistic multivariate time series forecasting. Unlike existing generative models, CLaST learns embeddings that preserve contextual similarity between observations through our contrastive loss function. Experiments across nine widely adopted benchmarks demonstrate that CLaST consistently surpasses strong baseline methods. In short-term forecasting tasks, our approach achieves improvements of up to $16.4\%$ in CRPS and $14.4\%$ in NMAE over the second-best method. Furthermore, in long-term prediction CLaST attains superior overall performance, exceeding the second-best method by up to $48.6\%$ and $25.1\%$ in CRPS and NMAE, respectively.

cs.LG↗

Self-Normalizing Denominators in Rational Covariance Estimators

Many estimators are ratios of coprime polynomials in a sample covariance matrix, and their accuracy depends on the relative fluctuation of the sample denominator. Under Gaussian sampling in fixed dimension, we call a nonconstant polynomial denominator self-normalizing if the first-order variance of its relative error does not depend on the population covariance. We prove that these denominators are exactly the flag powers, nonzero constant multiples of products of positive integer powers of nested generalized variances. Equivalently, the denominator's sample-to-population ratio has a covariance-independent finite-sample law, which we determine explicitly. Sufficiency is classical; the new converse shows that a first-order variance condition forces an exact sampling law. We show that relative stability, meaning bounded first-order relative variance, characterizes uniform tightness of scaled relative errors over positive-definite covariances. It permits replacing the sample denominator by its population value in the limit theory of the ratio. Self-normalization is its rigid core. We locate these classes in applications, where regression on predecessors in a fixed order yields only constant or self-normalizing denominators, instrumental-variable formulas yield relatively unstable ones, and nonparametric identifiability does not guarantee relative stability.

math.ST↗

Beyond Attention Masks: Instruction Anchoring for Efficient In-Context Diffusion Generation

In-context diffusion transformers concatenate instruction, target, and reference tokens into a single sequence for joint attention. Reference-side computation must therefore be repeated at every denoising step, with the cost growing rapidly as more references are added. Decoupling reference tokens from the target enables exact key-value reuse across denoising steps, but prevents the references from attending to the instruction, degrading instruction following and reference fidelity. This trade-off cannot be resolved through attention-mask design alone. We introduce AnchorCache, a parameter-free token-layout and attention-mask co-design that inserts static text anchors. These anchors condition the reference representations on the instruction during cache construction, after which the resulting reference keys and values can be reused exactly across denoising steps. To recover the quality initially lost through this structural conversion, we apply teacher-forced velocity distillation followed by a short on-policy stage that queries the teacher at student-visited states. To our knowledge, this is the first use of on-policy distillation for architectural recovery in diffusion models. Across benchmarks spanning image, speech, and video generation, AnchorCache matches full-attention quality. Its efficiency gains increase with the reference-context size, reaching a 6.40x speedup in diffusion transformer inference.

cs.CV↗

Assessing the Impact of High-Resolution Imaging on Statistical Validation of TESS Planet Candidates

High-resolution imaging is widely used to constrain false-positive scenarios in exoplanet validation, but it is a finite follow-up resource that reaches only a subset of candidates, and its population-level impact on validation outcomes has not been quantified through controlled removal experiments. Using an automated pipeline built on TRICERATOPS, we compute the false-positive probability (FPP) of 443 TESS planet candidates. For the 264 planet candidates with high-resolution imaging observations, we compute FPP with and without the corresponding contrast curves, allowing us to quantify the impact of the additional data. We find that 72% of 68 contrast-curve bearing validated planets would fail validation without their adopted contrast curves. The fraction requiring imaging decreases with increasing planet size, from 100% below $1.7~R_\oplus$ to $33\%$ above $4~R_\oplus$: within our sample and TRICERATOPS-based analysis, the availability of high-resolution imaging directly limits the yield of small-planet validation and the supply of validated targets for atmospheric characterization. Our analysis statistically validates 64 new TESS planets with sizes spanning 0.94 to 7.83 $R_\oplus$ across hosts of spectral type M through F. Four of these are highly amenable to JWST observations based on the transmission and emission spectroscopy metrics, and each achieves validation only with its imaging constraint.

astro-ph.EP↗

Learning Generalizable Behaviors for Terminal Agents

Terminal agents are a compelling application of large language models (LLMs), with the potential to integrate deeply into users' daily workflows. Reinforcement learning (RL) is a key technique for improving their capabilities, making scalable training environments a central challenge. Since public real-user interaction data are scarce, synthetic environments provide a practical alternative, but often suffer from domain gaps and limited fidelity, leading to poor generalization. Existing work mainly scales the quantity and diversity of synthetic environments, while reward-signal quality and the mechanisms governing generalization remain under-explored. We study how RL improves terminal agents and propose the Agentic Compositional Generalization hypothesis: rather than teaching new domain-specific skills from scratch, RL primarily shapes high-level decision-making behaviors that compose and route low-level skills acquired during pre-training and supervised fine-tuning (SFT). This account is consistent with our empirical results and suggests that verifier quality, which determines which behaviors are reinforced, is more important than simply increasing environment quantity or diversity. Motivated by this insight, we propose River, a simple training recipe that improves reward quality by filtering low-quality environments and augmenting outcome rewards with process-level behavior regularization. Using this recipe, our RL-trained agent achieves the best performance among evaluated open-source RL-trained 8B models across four terminal-agent benchmarks. River also generalizes across model families, scales, agent harnesses, and RL objectives. Using fewer than 30% of the TMax training environments, River improves RL gains by 106% and 30% on average for models ranging from 2B to 27B on Terminal-Bench-Lite and Terminal-Bench-v2.1, respectively.

cs.LG↗

$Δ$ resonance contributions to QED radiative corrections in neutron and inverse beta decay

We incorporate the $Δ(1232)$ resonance into pion-induced QED radiative corrections to neutron decay and inverse beta decay (IBD). Within the framework of heavy-baryon chiral perturbation theory with explicit $Δ$ degrees of freedom, we compute additional contributions and study their impact on IBD cross sections and on the renormalization of the nucleon isovector vector and axial-vector charges. $Δ$ resonance does not renormalize the vector charge. For the axial-vector charge, including the $Δ$ resonance improves convergence and reduces the QED radiative correction to the experiment-over-lattice-QCD ratio $g_A/\left(g^\mathrm{QCD}_A g_V\right)$. $Δ$ resonance increases the pion-induced QED radiative corrections to IBD by a factor $1.2$-$1.3$.

hep-ph↗

Extreme-ultraviolet spectroscopy using quantum logic: a feasibility study for the 1S-2S transition in singly-ionized helium

Extreme-ultraviolet (XUV) spectroscopy represents an important new direction in precision physics, with potential applications ranging from the metrology of fundamental constants to tests of physics beyond the Standard Model. However, the application of quantum control methods for precision spectroscopy remains an open challenge in the XUV range. Here we present a novel quantum logic (QL) spectroscopy method for precision spectroscopy of weak XUV transitions, and numerically validate its feasibility for the $1S-2S$ transition at 40.81\,eV in singly-ionized helium (He$^{+}$). We propose a scheme based on a single He$^{+}$ ion co-trapped with a Be$^{+}$ ion in a Paul trap, and He$^{+}$ excitation with pairs of frequency-comb (FC) laser pulses upconverted to the XUV via High-Harmonic Generation (HHG). We investigate a nondestructive QL scheme to detect $1S-2S$ excitation, and compare its performance with a destructive readout based on state-selective ionization. Phase coherence of the XUV light is modelled and an optical cavity is used to filter the FC pulses prior to HHG. We model the motional excitation dynamics of trapped ions outside the Lamb-Dicke regime, and numerically validate a scheme we proposed in \cite{Grundeman} to cancel the first-order Doppler broadening and the recoil shift by synchronizing the ion's secular period with the time delay between the two excitation pulses. We show that precision spectroscopy of the $1S-2S$ transition in He$^{+}$ at the 10 kHz level is feasible, for improved tests of quantum electrodynamics (QED), a measurement of the Rydberg constant $R_{\infty}$ independent of hydrogen measurements, or an improved determination of the alpha particle and helion charge radii. The proposed method may also be applied to XUV spectroscopy of other ions outside the Lamb-Dicke regime.

physics.atom-ph↗

The Sharp Tail of Uniform Stability

Uniform stability controls how much one training example can change the loss at any test point. A new logarithmic-free upper bound shows that a $γ$-uniformly stable algorithm with loss in $[0,L]$ has generalization gap at most $O \left(γ\log(1/δ) +L\sqrt{\frac{\log(1/δ)}{n}}\right)$ with probability $1-δ$. Whether an actual bounded-loss learning algorithm can realize the linear dependence on $\log(1/δ)$ has remained open. The known construction realizes it only for auxiliary weakly dependent random variables whose pointwise range grows with $n$. The known learning lower bound holds only at constant probability. We close this gap. For every $n$, stability level $γ$, and loss bound $L$, we construct one deterministic $γ$-uniformly stable learning problem whose tail satisfies, simultaneously for $1\le p\le c n$, $\mathbb P \left( R(A_S)-R_S(A_S) \ge c'\min \left\{L,γp+L\sqrt{p/n}\right\} \right)\ge e^{-p}.$ The construction is ordinary bounded absolute-loss regression with constant labels. Its key is a multiscale collection of rare Rademacher features. A coordinatewise ramp is stable in sup norm, while an odd symmetrized maximum converts a unique extreme feature into a gap of order $γp$ without violating the loss bound. Geometrically spaced ramps put all confidence levels into the same problem. Together with the logarithmic-free upper bound, this determines the optimal high-probability and moment dependence of uniform stability up to universal constants.

cs.LG↗

Two Dimensions Govern Agnostic Multiclass Transductive Learning

In transductive classification, an adversary fixes a labeled population, one label is hidden uniformly, and the learner sees all remaining labels. For binary classes, agnostic transductive and PAC learning have the same minimax rate. Whether this extends to multiclass learning was open, especially for unbounded label spaces where uniform convergence can fail. We resolve the question up to logarithmic factors. For every multiclass class $\mathcal H$ with DS dimension $d_{DS}$ and Natarajan dimension $d_{\mathrm N}$, the optimal agnostic transductive excess error satisfies $\widetildeΘ\left(\frac{d_{DS}}{n}+\sqrt{\frac{d_{\mathrm N}}{n}}\right).$ The result holds for arbitrary label spaces. The two terms are both necessary. A DS pseudo-cube gives the realizable $d_{DS}/n$ obstruction, while a Natarajan cube with repeated points and fair labels gives the agnostic $\sqrt{d_{\mathrm N}/n}$ obstruction. The upper bound uses a random-reservation principle. The learner deliberately ignores a constant fraction of the visible labels, which makes the true test point uniform in a large unseen block. We combine realizable compression, a label-space reduction, and inside-menu agnostic compression across this finite-population split. A new without-replacement multiplicative-weights lemma preserves the fast $d_{DS}/n$ term. Consequently, agnostic multiclass PAC and transductive learning obey the same two-dimension law up to logarithmic factors.

cs.LG↗

A Taxonomy of Construction Task Activities for Robot Workers

Recent vision-language-action models offer a path toward robots with broader capabilities than conventional task-specific systems. Deploying such systems in construction, however, requires a precise inventory of worker activities and the capabilities needed to execute them. We present TARCAT, an occupation-grounded taxonomy derived from 91 O*NET task statements across seven high-employment construction occupations and 30 instructional videos. TARCAT defines 41 primitives spanning intellectual, social, and physical categories and provides a mechanism for composing parameterized primitive sequences into reusable skills. This human-interpretable vocabulary supports the specification of robot requirements and enables coding agents to retrieve and extend skill libraries for task execution. We also demonstrate selected primitives on a DOBOT CR3 arm with a CRAFT hand. TARCAT thus provides a common vocabulary for analyzing construction work and developing general-purpose construction robots. Annotations are available at https://github.com/AICPS/TARCAT-Taxonomy.

cs.RO↗