arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 235 records · Page 13Linked to original sources

SCoCaT: Success Conditioned Constrained Reinforcement Learning for Spacecraft Docking

Termination-based constrained reinforcement learning is attractive for safety-critical robotic deployments: it avoids online optimization at inference, scales easily to many constraints via a single scalar per constraint, and is simpler to implement than commonly used Lagrangian methods. Instead of pricing violations through summed cost penalties, this approach makes violations structurally unprofitable by shortening the effective horizon for each violation. We identify a structural failure mode of this method class on terminal-navigation tasks: reaching a precise goal configuration while satisfying safety constraints that tighten along the final approach. When the goal sits inside the region close to where the constraints become active, the survival-weighted objective makes dwelling outside the goal region strictly preferable to entering, producing high constraint compliance with low task completion. We formalize this pathology and show that a minimal augmentation to off-the-shelf RL algorithms like PPO resolves this ``feasibility collapse''. We empirically demonstrate that adding a dense per-step success signal via an auxiliary value critic improves the task completion rate while maintaining safety-critical constraint compliance. Validation across two representative spacecraft platforms: a 6U-CubeSat spanning the mass and degree-of-freedom envelope of operational proximity operations, and a floating platform testbed for zero-shot sim-to-real transfer in our laboratory, supports the generality of these findings.

cs.RO↗

Hopf quotients of the infinite-dimensional Gaussian pyramid

We study the infinite-dimensional Gaussian pyramid and its quotients by the global sign flip and the $U(1)$-Hopf action. We resolve affirmatively a long-standing problem posed by Tomohiro Fukaya around 2014: these three limiting geometries are pairwise non-similar, meaning that no positive rescaling makes any two of them coincide.

math.MG↗

Centering Drives Normalization Gains: Price-Offset Nuisances in Cross-Sectional Return Prediction

Cross-sectional return prediction from raw intraday bars is sensitive to each instrument price level, an additive nuisance under a return-ranking hypothesis. We test whether removing this offset, rather than rescaling amplitudes or changing the encoder, explains gains on a point-in-time CSI~300 five-minute panel. We evaluate eight parameter-matched encoders with and without RevIN normalization; a parameter-free ladder then separates identity, scale-only, centering, last-value referencing, differencing, and standardization across all fields and restricted channels. Centering drives the reliable effect, while scale-only normalization does not help. All eight paired effects are positive and survive Holm correction on raw rank IC, after style residualization, and after further residualizing on short-term reversal. Among six stronger encoders, gains of 0.0376-0.0567 exceed the 0.0109 spread of normalized IC (0.0830-0.0939). Price-only standardization retains 93--101% of the all-field gain. These results place the main effect in transformed price-channel offset removal rather than amplitude scaling or encoder choice.

cs.CE↗

LZ nuclear recoil event from inelastic singlet-doublet scalar dark matter

We study the possibility of explaining the recently reported high-energy nuclear recoil event by the LUX-ZEPLIN (LZ) collaboration within the framework of singlet-doublet scalar dark matter (DM). Considering DM to be the CP even mass eigenstate formed out of a scalar doublet and a real scalar singlet, both being odd under an unbroken $Z_2$ symmetry, we find the parameter space of the model consistent with correct relic abundance and direct-detection limits on elastic scattering rates. A large part of the parameter space can also lead to inelastic up-scattering of DM with a rate and recoil energy consistent with the recent LZ event. Depending upon the singlet-doublet mixing, the model allows a much wider range of currently allowed parameter space compared to the pure scalar doublet DM limit, which is disfavored due to non-observations of high-energy neutrinos from solar-captured DM annihilation by the IceCube experiment. Presence of $Z_2$-odd right-handed neutrinos also leads to other interesting phenomenology related to the origin of light neutrino masses and leptogenesis.

hep-ph↗

Textures as a phase-transition probe for quantum spin chains

The idea of quantum texture has been recently proposed and used as a tool for quantifying coherences and for quantum gate identification. In this work we offer a study on its usage to quantum phase transitions, demonstrating the rugosity metric as a simple tool for effective phase-transition probing. We establish the link between rugosity in the computational basis and the hierarchy of spin correlators, and analyze rugosities defined in the global ground-state and in ground-states belonging to different magnetization sectors (to which we refer to as global vs symmetry-resolved rugosities) to study the phase diagram of the Heisenberg XXZ model. We find distinct rugosity signatures at both transition points. In particular, a sharp feature appears at $Δ=1$ already for small systems, revealing a pronounced sensitivity of the correlation hierarchy encoded by the texture to this point. Since the BKT transition coincides with the isotropic $SU(2)$ point of the XXZ model, this behavior may reflect a particular sensitivity of rugosity to the structure of the spin-correlation hierarchy at isotropy.

cond-mat.str-el↗

A Foundation Model for Large-Scale Wireless Network Planning , Operation and Optimization

Wireless cellular networks form the connective tissue of human society, sustained by a continuous physical dialogue between engineered infrastructure and its surroundings. Radio signals emitted from base stations traverse terrain, diffract around buildings and scatter through streets before reaching billions of users. Together, these interactions produce the city-wide radio environment on which every network decision rests. Shaping this environment through deployment and optimization determines the connectivity societies rely on, yet learning it effectively at city scale and generalizing across diverse cities and deployments remain open challenges. Here we answer positively by introducing ChaRT, a foundation model that learns transferable radio representations from measurement reports generated by deployed cellular networks. These reports provide abundant multi-cell, multi-beam observations without dedicated campaigns, forming a scalable data foundation for city-scale learning. ChaRT embeds beam-level angular structure, network hierarchy and propagation-regime diversity in its architecture, and is pretrained through context-aware masked beam modelling and self-distillation with channel-model-constrained augmentation. We pretrain ChaRT on over one billion reports comprising 18.2 billion beam-level observations from 3,503 cells in one city. With a single set of weights, ChaRT reconstructs radio environments in unseen cities and transfers to radio map construction, new-site prediction and network parameter tuning. With only 1% of labelled data, it supports user localization, beam prediction, propagation scenario classification and estimation of the signal-to-interference-plus-noise ratio. The learned representation further enables beamspace clustering for reusable radio-grid construction. These results establish ChaRT as a transferable foundation for network-wide intelligence.

eess.SP↗

Combating Instruction Conflict via Energy-Driven Latent Conflict Detection

Large Language Models (LLMs) are increasingly deployed with hierarchical instructions, yet they remain vulnerable to conflicts in which user directives override system-level constraints. Existing defense mechanisms predominantly focus on static input inspection and therefore fail to detect Response Drift, a phenomenon in which the model's final response violates system-level constraints despite seemingly compliant inputs. To bridge this gap, we introduce ELCD, a response-level latent conflict detector for post-generation, pre-delivery verification. Given the full generated output, ELCD constructs a composite hidden-state representation by concatenating the final-token embedding with the mean-pooled response embedding. It then optimizes a pairwise margin ranking objective to separate compliant and drifting responses in latent space. Extensive experiments across five mainstream LLMs ranging from 1.5B to 14B parameters demonstrate that ELCD significantly outperforms competitive baselines. Notably, it improves the PR-AUC on Llama-2-7B by approximately 30 percentage points and reduces the False Positive Rate at 95% TPR (FPR95) on Mistral-7B to 2.67%. These results suggest that ELCD provides a promising approach for latent instruction-conflict detection in open-weight or self-hosted LLM deployments.

cs.CL↗

The topology of Gromov--Hausdorff space

We prove that the space of isometry classes of nonempty compact metric spaces, equipped with the Gromov--Hausdorff distance, is homeomorphic to the real separable infinite-dimensional Hilbert space. We construct a continuous assignment of full-support probability measures that is equivariant under isometries and finite-dimensional local approximations that control all pairwise distances. These approximations yield the absolute retract property for all metrizable spaces. We also prove that any countable family of continuous maps from compact metrizable spaces can be approximated, with respect to a prescribed open cover, by maps whose images form a discrete family.

math.MG↗

Pushing the Boundaries of Streaming Multi-Speaker ASR: A Systematic Study of Architectural Trade-offs

Streaming multi-speaker ASR is a challenging task that must balance accuracy, latency, and efficiency while handling overlapping speech and maintaining coherent long-context modeling over extended conversations in an online fashion. We present a unified framework that categorizes streaming multi-speaker ASR into four architectural strategies based on how diarization and ASR are integrated. Using a shared pair of open-source streaming ASR and diarization models as a common foundation, we derive four multi-speaker ASR systems that differ in whether they employ multiple model instances, fine-tuning, or both. We evaluate these systems across multi-speaker accuracy, single-speaker accuracy degradation, memory footprint, and training complexity. Through this systematic architectural analysis, we clarify the design space for streaming multi-speaker ASR and provide practical guidance for selecting the most suitable approach under diverse deployment constraints.

eess.AS↗

Construction of Multi-sequences With High Nonlinear Complexity via Narrow Ray Class Fields

Nonlinear complexity is a fundamental criterion in the evaluation of pseudorandom sequences. The construction of multi-sequences with high nonlinear complexity is both theoretically and practically important in cryptography. Motivated by prior constructions of multi-sequences with high nonlinear complexity in [IEEE Trans. Inf. Theory, 60(10), 2014] and [IEEE Trans. Inf. Theory, 63(12), 2017], we provide a unified framework via narrow ray class fields and the cyclic descent introduced by Guruswami and Xing in [J. Combin. Theory Ser. A 129 (2015) ]. Then we can generate new multi-sequences with high nonlinear complexity over function fields with arbitrary genera.

cs.IT↗

Noncommutative sharp Hausdorff-Young inequality

We prove the sharp Hausdorff--Young inequality on the quantum Euclidean space. Our result implies the sharp Hausdorff--Young constants for the Weyl transform, as well as that for Heisenberg groups. The key ingredient is a novel flow related to the mixed-norm of noncommutative Gabor transform. This, meanwhile, implies a new proof of the classical sharp Hausdorff--Young inequality. We then apply the sharp Hausdorff--Young inequality to establish the sharp Young inequality for noncommutative convolution in the range $1\le p,q\le2\le r\le\infty$, with $1/p+1/q=1+1/r$. After appropriate rescaling and trace normalization, this convolution coincides with beam-splitter convolution for bosonic systems, yielding the corresponding Young inequalities with optimal constants in the same range.

math.FA↗

Stability Framework for the Singularity of the Euler Equations on $\mathbb{R}^3$

In a recent numerical study, we found a high-precision singular profile for the Euler equations on the unbounded domain $\mathbb{R}^3$. The present manuscript complements that study by establishing in detail a preliminary framework for proving (nonlinear) stability of the approximate self-similar profile, reducing the analysis to a large but finite collection of explicit estimates and computable constants. Conditional on rigorous certification of the estimates and constants appearing in the argument, and on the candidate profile satisfying the required nonlinear stability conditions, the framework closes the stability proof and, crucially, allows the resulting stable rescaled profile to be reconstructed as an admissible solution in the original variables that becomes singular in finite physical time. With the overall stability and reconstruction mechanisms formulated, the remaining work within this approach is largely quantitative: determining whether the explicit constants and margins can be rigorously certified with sufficient positive margin and, where necessary, sharpening selected analytic estimates.

math.AP↗

Self-Similar Singularity of the Euler Equations on $\mathbb{R}^3$

We provide evidence of a finite-time singularity in the 3D Euler equations on the unbounded domain. Using a physics-informed neural network (PINN) with a self-similar ansatz, we find an approximate singular profile for the Euler system at the critical blowup rate of $0.5$ and certify it using a spline representation. The transport field associated with the obtained profile has local outgoing property throughout the domain that suggests linear damping, a key stabilizing mechanism for the candidate profile. We also establish a framework for proving nonlinear stability of the approximate self-similar profile, reducing the analysis to a large but finite collection of explicit estimates and computable constants.

math.AP↗

Human Agreement and Return Association Are Not Interchangeable Criteria

Financial NLP has a standard workflow: validate a sentiment tool against human labels, then trust it to extract market signal. This assumes the two evaluations measure the same thing. We test that assumption in a setting where both can be measured at once: a corpus of securities class actions (2002-2025) linking 70,500 X messages to abnormal stock returns, with a single-annotator human labelled gold sample. Running five instruments (VADER, Loughran-McDonald, FinBERT, Twitter-RoBERTa, and an LLM annotator) through one identical pipeline, we find that the relationship between construct and predictive validity depends on the sampling convention and score representation. Under conventional method-specific sampling, human agreement aligns more closely with graded same-day associations than with one-day leads. On a fixed-n panel, however, agreement has similar graded rank correlations at both horizons, while the coarse ordering remains weak. Benchmark agreement therefore establishes semantic validity but does not by itself determine predictive rankings. In a conversation that is 17.6% spam, message volume predicts neither market damage nor settlement size.

cs.AI↗

Fold First, Detect Directly: Communication Symbol Detection Without Unfolding for Low-Bitrate Modulo-ADCs

Modulo-folding analog-to-digital converters (MF-ADCs) enable low-dynamic-range quantizers to sample high-amplitude signals without clipping. However, downstream processing traditionally relies on waveform unfolding algorithms, which require high oversampling rates and are highly sensitive to noise. For communication receivers, where the goal is symbol detection rather than signal reconstruction, unfolding is a redundant intermediate step. In this paper, we propose an unfolding-free maximum-likelihood symbol-detection framework for oversampled MF-ADCs in the presence of joint channel and quantization noise. By leveraging modulo wrap cancellation, we derive an exact, single-term Mahalanobis-distance metric that operates directly on folded observations. To handle long sequences, we introduce a parallelized block search algorithm that reduces computational complexity to scale linearly with sequence length. Simulations show our detector significantly outperforms existing unfolding baselines and approaches unclipped conventional ADC performance.

eess.SP↗

The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation

Cooperative AI agents are evaluated against other AIs, yet human cooperation relies on implicit conventions -- shared protocols for reading meaning beyond the literal message -- which AI-AI benchmarks may not capture. We propose the convention gap, the difference between the failure probability predicted from the literal content of communication and the observed failure rate, as a metric of implicit communication. In the card game Hanabi, the finite deck and deterministic hint constraints make this posterior exactly computable. We replayed about 101,000 play actions from three public datasets of human-human (an online Hanabi platform), AI-AI (HOAD), and human-AI (HanabiData) games. The gap was +26.2 percentage points (pp) in human pairs, -0.7 pp in AI pairs, and +16.4 pp in human-AI pairs, and was concentrated on plays of cards that had received no hints (+46 pp in human pairs). Within human-AI play, the literal information available to humans was similar across the three AI partners (mean predicted failure 38-41%), but human failure rates ranged from 14.4% to 34.4% and the gap from +24.1 to +6.2 pp; the partner eliciting the largest gap produced the fewest human failures. Game score carried different information: it depended on each corpus's roster composition, whereas the gap separated human from AI play at the agent level. As a known-answer check, Off-Belief Learning agents, whose convention content is controlled by construction, gave a gap of +1.6 pp at the convention-free level, rising monotonically to +21.7 pp. These results suggest that convention compatibility, rather than AI-AI performance, may predict an AI's effectiveness with human partners.

cs.AI↗

Asymptotic $q,t$-Fuss--Catalan numbers for type $B$

Let $W=W(B_n)$ act diagonally on $\mathfrak{h}\oplus\mathfrak{h}^*$, let $S=\mathbb{C}[\mathfrak{h}\oplus\mathfrak{h}^*]$, let $J\subset S$ be the ideal generated by the $W$-alternating polynomials and $\mathfrak{m}_S$ is the maximal ideal of the origin. For sufficiently large $m$ we compute $q,t$-Fuss-Catalan polynomial $Cat^{(m)}(B_n;q,t):=Hilb(\frac{J^m}{\mathfrak{m}_S J^m})_{det-part}$ and imply $Cat^{(m)}(B_n;1,1)=\binom{n(m+1)}{n}$. For proofs, we work with the $Γ$-equivariant Hilbert scheme $Y_n=nΓ$-$Hilb(\mathbb{C}^2)$, $Γ=μ_2$ and Haiman-type Koszul complex that defines the punctual locus of $Y_n$. Our formula for $Cat^{(m)}(B_n;q,t)$ is derived from a localization computaion for the Haiman-type Koszul complex.

math.CO↗

SignMimic: Robust High-Quality Sign Language Motion Generation via Human-Shape-Oblivious Pose Transfer Guidance

We study the challenge of sign language video mimicking: given a driving video and a single reference frame, synthesize a video where the target signer reproduces the source motion while preserving identity and linguistic form. Prior pipelines entangle rigid motion, non-rigid deformation, and view-dependent completion in a monolithic generator, causing handshape drift and spatio-temporal instability. We present SignMimic, which (i) applies a TNet-based model to study SE(3) rigid canonicalization to stabilize global pose, (ii) performs non-rigid adaptation in a canonical space to preserve fine-grained articulators (hands/face) and coarticulation via NIF2D, and (iii) uses Pose-MAE-style completion before conditional video diffusion. This factorization injects geometric and linguistic priors, yielding shape and spatio-temporal consistency. On several large-scale datasets (ASL 50K, How2Sign, CSL News), SignMimic achieves state-of-the-art-level performance on video quality, identity similarity, and frame continuity while also achieving minimal loss when performing back translation (SLT) on generated videos. Ablations confirm the role of rigid canonicalization, non-rigid adaptation, and completion. Code is available at https://anonymous.4open.science/r/UniSignMimicTurbo-6088; model checkpoints and video examples will be released.

cs.CV↗