arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

Sim-to-Real Aware End-to-End Learning Environment for Micromobility

While end-to-end autonomous driving systems show promise, their application to micromobility vehicles is hindered by simulators failing to capture specific kinematics, such as differential drives and omni-wheels. This paper pro- poses a sim-to-real-aware, vehicle-specific end-to-end learning environment for the WHILL Model CR on AWSIM and ROS 2. To minimize the sim-to-real gap, physical parameters are optimized via Bayesian optimization using real-world data, reducing trajectory errors across various driving scenarios. Additionally, this study introduces a synchronized architecture tailored for the stable training of world model-based agents. An end-to-end policy trained with DreamerV3 exhibited learning progress and achieved task completion in a simulated obstacle avoidance setting. Furthermore, this policy demonstrated direct sim-to-real transfer to the physical vehicle, enabling the vehicle to navigate around a cardboard box in a real-world corridor replica without fine-tuning. This paper provides a practical foundation for sim-to-real micromobility policy studies.

cs.RO↗

Tessellated Isotropic Elastic Lattice Spring Model for Quasi-Brittle Fracture

Quasi-brittle fracture is prevalent in concrete, rock, ceramics, composites, and masonry, and its simulation faces a trade-off among accuracy, efficiency, and simplicity. The classical Lattice Spring Model (LSM) captures cracking via bond breakage without remeshing, but its elements are empirical and limited to a few tessellable shapes with fixed Poisson's ratios. We propose a tessellated Isotropic Elastic Lattice Spring Model (IELSM) that discretizes continua into polygonal elements with axial springs and a volumetric constraint, achieving isotropic elasticity on arbitrary polygonal tessellations. Macroscopic isotropy reduces to governing equations whose solvability gives a theoretical criterion for element admissibility, proving conventional tessellable elements and extending to arbitrary regular N-gons and concave elements. Exploiting boundary interpolation compatibility with finite elements, IELSM is assembled by direct node sharing, without interface elements or kinematic constraints. Coupled with an isotropic damage model, a pure bending test and four fracture benchmarks show that the coupling preserves displacement accuracy, yields crack paths and load-displacement curves agreeing with experiments and outperforming standard FEM, and is insensitive to mesh refinement. IELSM can also be restricted to damage-prone regions, with the remainder modeled by finite elements via node sharing. For the benchmarks, this reduces nodes by 41.2-80.9% and CPU time by 34.2-85.2% versus full-domain IELSM. The framework advances IELSM element construction from empirical trial and error to theoretical determination and simplifies coupling to node sharing, offering a balanced route for complex quasi-brittle fracture analysis.

math.NA↗

A diffusion-free multi-layer neural-network method for multidimensional nonlinear hyperbolic equations

This paper is devoted to the non-diffusive computation of solutions to $N$-dimensional hyperbolic systems of conservation laws (HCL) and more specifically the accurate approximation of simple waves. We develop the LeafNet algorithm, which relies on: i) a reformulation of HCL as a coupled system involving smooth solutions, and ii) an accurate physics informed algorithms approximating smooth functions. An error estimate analysis and two-dimensional numerical experiments are proposed to illustrate the presented strategy.

math.NA↗

Cross-Country Code-Mixing for Generative Recommendation

Cross-country recommendation on modern e-commerce platforms is typically deployed with disjoint user and item ID spaces across markets, removing the shared anchors that conventional cross-domain methods rely on. Generative recommendation (GR) mitigates this by mapping items into a shared token space and training a unified model, but existing approaches keep behavior sequences strictly country-specific, so knowledge transfer occurs only at the parameter level and remains absent at the data level. Inspired by code-switching corpora in multilingual natural language processing, we propose CMRec, a cross-country GR framework that injects cross-country supervision at the data level via dual-constrained, context-aware code-mixing. CMRec first learns a shared semantic codebook from multi-modal content and behavioral co-occurrence across countries. It then uses this codebook to synthesize mixed-country sequences via token-level substitutions that satisfy both static (content) and dynamic (e.g., price, audience, popularity) constraints. Finally, it introduces a context-aware loss that reweights mixed samples according to their plausibility in the current sequence. Experiments on two real-world multi-country datasets and an online A/B test show that CMRec substantially improves recommendation quality in data-sparse countries while preserving performance in data-rich countries, achieving +1.77% advertising revenue and +2.64% orders on a large-scale e-commerce platform.

cs.IR↗

AquaMend: Minimal Re-probing and Conditional Rollback for Latent-Belief Failures in Embodied Agents

Physical changes or sensing errors can invalidate embodied agents' task-relevant beliefs. AquaMend compares re-probing, rollback, and supported continuation on a probe-belief-action graph under an expected-loss objective covering sensing, physical recovery, and uncorrected failures. A joint posterior guides a one-step policy with conditional detection-power screening. The per-belief three-way optimum requires independence, separability, and fully resolving probes; the general policy has no global optimality guarantee. Across 32 paired scenarios in a self-constructed simulation benchmark, AquaMend recovers in 28/32 cases and reduces mean complete loss by 21.6% versus restart. Its paired loss difference from decision-theoretic troubleshooting (DTT) is not statistically significant after Holm correction. Against the all-candidate ablation, online decision time decreases by 12.3% overall but increases by 3.4% in the uncovered late stage.

cs.RO↗

Same Bit Width, Different Outcomes: Post-Training Quantization of Text-to-Speech Across Architectures

Post-training quantization (PTQ) reduces the cost of on-device text-to-speech (TTS), but published evaluations cover one system or method. We evaluate PTQ across TTS architectures under one protocol with three core models, weight and activation ablations of eight more, and two held-out models quantized blind. Four-bit per-channel weights reduce UTMOS, a predicted mean opinion score, by 2.8 on Supertonic and 0.07 on Kokoro, and per-tensor scaling can cause severe degradation even at 8 bits. The same bit width yields different outcomes, because the sensitive component is model-specific and not reliably predicted from the model class. A staged ablation procedure identifies it, and per-layer GPTQ can restore it to within 0.1 UTMOS. Real int8 and int4 kernels reproduce the simulated ordering at hardware-dependent cost. On a Mac mini, a 4-bit weight kernel runs Supertonic at 0.60x the fp32 latency while int8 is slower, so each configuration requires validation on the target runtime.

eess.AS↗

Well-Posedness for KdV-Type Equations with Second-Order Derivative Nonlinearities

We study a class of complex-valued KdV-type equations with cubic second-order derivative nonlinearities and arbitrary complex coefficients. For sufficiently small initial data in $L^2$, we prove local well-posedness in $H^s(\mathbb R)$ for $s\ge3/4$. The key ingredient is a family of dyadic resolution spaces $Z_k=X_k+Y_k$, designed to overcome the logarithmic divergence arising from high$*$low$*$low interactions in which both derivatives fall on the highest-frequency factor. For the completely integrable third-order Kaup--Newell flow, we establish global-in-time $H^s(\mathbb R)$ bounds for $0\le s<1$ and global well-posedness in $H^s(\mathbb R)$ for every $s\ge3/4$, with no smallness assumption on the initial data.

math.AP↗

ReVNM: Learning-Based Visual Navigation from a Remote Camera

Visual Navigation Models (VNMs) enable robots to navigate from egocentric visual observations without geometric localization and planning, but long-range navigation still requires pre-built maps. This paper presents the Remote Visual Navigation Model (ReVNM), which uses a single remote surveillance camera to serve as both an observation source and an implicit environmental map for visual navigation. While the use of remote cameras could eliminate the need for pre-built maps as well as onboard vision processing, their limited field of view instead of egocentric observations makes it hard to achieve collision-free navigation. The lack of existing data with diverse remote viewpoints, which are crucial for training robust VNMs, further complicates the challenge. In this work, we propose a learning-by-synthesis approach to address this two-fold challenge. Our ReVNM extends a state-of-the-art VNM architecture with an exocentric-to-egocentric (exo2ego) module that predicts an egocentric depth observation from remote-camera observations. This helps the VNM to plan a path while considering obstacles in front of the robot. Trained only on randomly generated worlds with diverse obstacle layouts and camera viewpoints, ReVNM can generalize well to real robot navigation without additional fine-tuning. Experiments in both simulation and real-world environments confirmed the effectiveness of the proposed approach.

cs.RO↗

Geometric stability for complex Monge-Ampere equations

Let $X$ be a compact Kahler manifold. The analytic stability theorem of Kolodziej for complex Monge-Ampere equation states that for any Kahler metrics $ω$ and $ω'$ in the same cohomology class, if their volume measures are bounded in $L^p(X)$ (for some $p>1$) and close in $L^1(X)$, then their Kahler potentials are close in $L^\infty(X)$. In this paper, we establish the geometric stability for complex Monge-Ampère equations that $L^1$-closeness of volume measures implies $L^\infty$-closeness for the induced distance functions by $ω$ and $ω'$. Consequently, we prove that any non-smooth Kahler current with volume measure bounded in $L^p$ (for some $p>1$) and Ricci current bounded below induces a unique metric space, which turns out to be a compact RCD space homeomorphic to $X$ itself.

math.DG↗

Spectral Graph Neural Networks with Hermite Polynomials: A Comprehensive Study

We study spectral graph neural networks built from Hermite polynomials and propose HermNet, a simple model that combines a nodewise predictor with normalized Hermite propagation. Its sparse recurrence requires neither eigendecomposition nor a learned basis. We distinguish the basic model from optional coordinate calibration, response normalization and Gaussian derivative regularization. Hermite and other complete polynomial bases span the same degree-bounded filter space, but their coordinates can produce different optimization behavior under limited training budgets. We analyze this behavior through spectral signal energy, label sampling, changes in learned features and the bias--variance trade-off of regularization. Controlled synthetic experiments identify a regime in which plain HermNet outperforms matched polynomial-basis alternatives, including with a jointly trained nonlinear predictor. Curvature regularization further improves HermNet when the same functional penalty is available to every comparator. Fixed-predictor controls support the advantage under short training budgets, but longer training removes the plain-model lead. Matched real-data comparisons show accuracy deficits, and architectural and numerical studies identify further limits. Together, the analysis and experiments clarify when Hermite propagation is useful and how calibration and regularization affect its performance.

cs.LG↗

Relative periodic orbits in spatially extended Kolmogorov turbulence: from global to localized recurrence

Recurrent invariant solutions provide a dynamical framework for interpreting turbulence, but their computation becomes increasingly difficult in spatially extended flows, where close whole-domain recurrences are rare. We study two-dimensional Kolmogorov flow in a [Lx,Ly] = [2pi,12pi] domain with Re = 10 using a distributed reduced-order model composed of a patch-based autoencoder and a neural ordinary differential equation. Near- recurrence candidates are refined by latent multiple shooting, decoded to physical space, screened under the full Navier-Stokes dynamics and supplied to a full-state Newton- Krylov solver. Fourteen candidates converge to relative periodic orbits (RPOs), with periods 3.65 < T <16.45. The catalogue contains both domain-filling states and RPOs with recurrent activity localised to a restricted cross-stream region. The weakly modulated surroundings generally remain finite-amplitude and distinct from the laminar solution. Selected localised RPOs persist under continuation in Reynolds number, and recurrent states can be reconverged after truncating substantial portions of their surroundings. In addition, the model trained at Ly = 12pi generates convergent RPO seeds at Ly = 6pi and 8pi without retraining. These results demonstrate that distributed learned dynamics can provide useful initial conditions for exact coherent states searches in extended domains and reveal recurrent solutions with strongly localised temporal dynamics.

physics.flu-dyn↗

Geometric Holder estimates for complex Monge-Ampere equations

This paper is the continuation of our earlier work on Holder estimates for Kahler potentials on compact normal Kahler spaces. We establish a geometric counterpart for analytic Holder regularity of Kolodziej. If a singular metric $ω_ϕ$ has Holder continuous potentials with respect to a smooth background metric, then its distance function is Holder continuous with respect to a smooth background distance. The result holds on normal Kahler spaces and yields compactness of the metric completion as well as its identification with the underlying complex space if $ω_ϕ$ is a Kahler current. Under this positivity assumption, our estimates establish the Holder equivalence between the intrinsic canonical Kahler metrics and the extrinsic smooth metrics on Kahler-RCD spaces, particularly on smoothable Kahler-Einstein spaces.

math.DG↗

Antichain polynomials of products of chains and minuscule posets

This paper studies the antichain polynomials of $[k]\times P$, where $P$ is a connected minuscule poset. We give a formula for the number of antichains, counted by size, of an arbitrary poset. Using this formula, we present necessary and sufficient conditions for the palindromicity of antichain polynomials for two infinite families of connected minuscule posets. We show that, for every connected minuscule poset $P$, if the antichain polynomial of $[k]\times P$ is palindromic, then it has only real and strictly negative zeros. This result, in particular, gives an affirmative answer to Ding-Dong's conjecture about $γ$-positivity of the antichain polynomial of $[k]\times P$. By constructing a bijection between antichains and labeled Dyck paths, we also give a layer-refined enumeration of the antichains in $[2]\times[m]\times[n]$. We establish a relation between such antichains and Clar covers of the hexagonal flakes $O(2,m,n)$, thus prove a conjectured determinantal formula for a family of Zhang-Zhang polynomials. Real-rootedness and stability results are also obtained when the shortest chain has length at most two. Finally, we present infinitely many connected Peck posets whose antichain polynomials are not unimodal, disproving the log-concavity conjecture of Ding and Dong.

math.CO↗

CrossSafe: Towards Cross-Embodiment Latent Safety Filters

Cross-embodiment learning has shown that a single model, such as a vision-language-action (VLA) model, can learn state representations and manipulation skills that can be applied across heterogeneous robots to accomplish various tasks. We hypothesize that the same holds for safety enforcement. The reasoning required to satisfy a safety constraint, such as detecting an obstacle, recognizing that it should be avoided, and selecting a safe abstract action, is largely shared across robots. What differs across embodiments is how the abstract safe action is realized: morphology, kinematics, and dynamics determine which actions are safe and feasible. Consequently, the same action can be safe for one robot and unsafe for another. This is especially important for generalist manipulation policies that operate in a common end-effector action space without explicitly capturing how safety depends on the robot's morphology and kinematics. We propose embodiment-conditioned safety filtering, in which a Hamilton-Jacobi reachability-based value function and its corresponding safety-maximizing policy are shared across robots. Using a morphology-aware latent representation of the robot and its environment, we perform Hamilton-Jacobi reachability analysis directly in latent space so that the learned safety concepts can generalize across embodiments while remaining explicitly conditioned on each robot's morphology and kinematics. We evaluate our approach across five bimanual robot embodiments and five manipulation tasks with whole-body collision-avoidance constraints. Our results show that a single policy, jointly trained across five manipulation tasks and four embodiments, exhibits zero-shot generalization to a held-out embodiment, reducing the nominal policy's collision rate. They also show that training using more embodiments improves generalization.

cs.RO↗

Uniqueness and stability of nonlinear filtering equations with unbounded random coefficients

We study a multidimensional nonlinear filtering model whose coefficients depend on a given observation-adapted predictable process and whose observation drift may grow linearly in both the state and the random input. Due to the unboundedness of the observation drift, a global reference measure is not available. To overcome this hurdle, a localized entropy argument is adapted to prove the stopped likelihood to be a uniformly integrable martingale at each control-energy stopping level. The stopped Zakai equation, and hence, the stopped filtering equation is derived. The global filtering equation is then established by de-localization. The uniqueness of the solution to the stopped Zakai equation is obtained by a duality backward stochastic partial differential equation. This uniqueness then propagates to that of the global filtering equation through the stopped ones. Finally, a stability result is established in $W_1$-distance of measures.

math.PR↗

Ricci entropy, RCD structures and Kahler spaces

This paper is the final installment in our series on the geometric theory of complex Monge-Ampere equations. We study singular Kahler metrics on compact normal Kahler spaces with klt singularities whose volume densities lie in $L^p$ for some $p>1$. We show that a lower bound for the Ricci current in the sense of pluripotential theory is equivalent to a synthetic Ricci lower bound in the sense of RCD theory. As a consequence, every Kahler class on a compact normal Kahler space with klt singularities admits an RCD structure, which makes differential geometric analysis available on singular complex spaces. As an application, we show that the fundamental group of a compact Kahler Calabi-Yau space with klt singularities is almost Abelian. We further establish compactness results for singular Kahler-Einstein spaces.

math.DG↗

Generative crystallographic phasing through invariant relationships

Crystal structure determination requires the phases of scattered waves -- yet diffraction measures only their intensities. Direct methods exploit phase invariants but become less reliable as diffraction information diminishes. Learned phase prediction has lowered the resolution barrier, yet remains primarily confined to centrosymmetric crystals with binary phases. We introduce PhiGen, a generative reformulation of traditional direct methods that learns origin-independent phase relationships for binary and continuous phasing. Across 210 space groups, including groups absent from training, it recovered high-quality maps for 99.0% of centrosymmetric structures and invariant-consistent phases for 92.8% of non-centrosymmetric structures. From simulated 3 Å zeolite powder data, the network recovered framework maps for 84.2% of held-out structures, versus 1.0% for Superflip. For experimental ZSM-25 and TNU-9, generated phases seeded high-resolution phase extension. These results suggest a route to structure determination from low-resolution, incomplete, and overlapped diffraction data.

cond-mat.mtrl-sci↗

Personalized Korean Lipreading as Visual Speech Recognition: Transfer, Census and Adaptation on OLKAVS

We present a personalized Korean visual speech recognition (VSR) system and quantify, on the nine-camera OLKAVS corpus, the gap between the population-level benchmark score and an individual user's error. A video-only Conformer initialized from English-trained weights attains 9.95 - 12.19% character error rate (CER) under the corpus protocol against the published 26.64, and 19.00 - 21.52 on unseen wording. Per speaker, CER spans 1.0 to 52.2%, with seen wording lowering CER by 7.0 - 9.0 points and professional delivery and spontaneous speech raising it by 8.5 - 10.5 and 12.7 points. A low-rank adapter with 4.6% of the parameters, trained on 4 to 29 minutes of the user's frontal video, lowers the CER of twelve high-error speakers by 2.13 to 3.58 points, transfers to every camera without loss, and keeps 85% of the full fine-tuning gain at 12% of its cost to other speakers. Cameras above the mouth plane add about six CER points as a constant offset that training on all views keeps small.

eess.AS↗