arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49Linked to original sources

Poster: FedWM-Guard: Thwarting Imagination Poisoning in Federated World Model-based Autonomous Driving

Federated learning (FL) can improve world model (WM)-based autonomous driving (AD) without centralizing raw private vehicle data, but it also turns model aggregation into a safety-critical integrity boundary. We introduce a new threat in federated WM-AD, namely \emph{imagination poisoning}: compromised vehicles submit bounded WM updates that preserve benign short-horizon predictions yet corrupt long-horizon rollouts (e.g., trigger-conditioned) during training, thereby misleading a downstream planner. We present \emph{FedWM-Guard}, to the best of our knowledge, the first defense to characterize planner-facing rollouts in federated WM-AD, screen authenticated updates in hidden-canary scenarios, audit predicted futures against later observations, and invoke a WM-independent safety shield when persistent inconsistency is detected. Unlike parameter-space defenses, it scores what an update makes the model \emph{imagine}, not only how the update looks. We also outline how we plan to evaluate it under non-IID (not independent and identically distributed) data, adaptive attacks, and benign distribution shift. This work highlights an unexplored domain, federated WM-AD, and its threat surface and potential countermeasures.

cs.CR↗

Towards Deployable Underwater Vessel Classification

We propose a compact underwater acoustic classification framework combining multi-representation feature engineering, temporal statistical pooling, and compact convolutional architectures designed for acoustic time-frequency and cochlear representations. We investigate multiple conventional and auditory-inspired representations and first evaluate lightweight classifiers and Conventional Neural Networks (CNNs) on ShipsEar dataset. On the provided split, a two-layer CNN achieves a macro F1 of 0.9918, while a Radial Basis Function Support Vector Machine (RBF-SVM) reaches 0.9883. However, source-recording provenance cannot be reconstructed, preventing verification of recording-independent generalisation. We therefore evaluate on DeepShip dataset using recording-level partitioning before segmentation. Under this protocol, a 157K-parameter compact CNN achieves a test macro F1 of 0.7226, while an 11.17M-parameter ResNet18 provides no improvement in validation performance under the matched setting. These results demonstrate the importance of representation-aware feature and model design, together with rigorous recording-level evaluation, for classification performance and deployability in compact underwater acoustic systems.

cs.SD↗

X-Rec Technical Report

Recent advances in generative modeling have reshaped recommender systems by formulating recommendation as a next-item generation problem. Existing retrieval approaches primarily follow two paradigms: user-to-item (U2I) methods represent user context using one or a few deterministic embeddings, which limits the ability to capture diverse and multi-mode interests, while semantic-ID-based autoregressive (SID-AR) methods model more expressive distributions but suffer from quantization errors and the low throughput of sequential decoding. To address these limitations, we propose X-Rec to directly learn the recommendation distribution in the continuous item embedding space through flow matching and generate embedding triggers for approximate nearest neighbor retrieval. X-Rec incorporates three key designs to make this formulation effective and efficient. First, we introduce anchor conditioning to decompose generation into coarse semantic-region selection and fine-grained refinement. Second, we adopt Riemannian flow matching to align generative trajectories with the hyperspherical geometry of item embeddings. Third, we design a late-interaction diffusion Transformer that restricts repeated velocity-field estimation to the final Transformer layer. On a streaming benchmark, X-Rec substantially outperforms U2I baselines, matches the retrieval quality of SID-AR methods, and delivers 3.46x higher inference throughput than SID-AR. X-Rec has also been deployed as a new retrieval source for a specific vertical content on TikTok, where two consecutive launches have yielded significant improvements in both vertical engagement (+4.1484%) and general engagement (+0.0111%).

cs.IR↗

Right Choice of Classification Algorithms Based on Reinforcement Learning for Prediction of Non-Alcoholic Fatty Liver

There are many complex issues in the world of artificial intelligence. Some of these problems are solved using other artificial intelligence methods, which are called artificial intelligence for artificial intelligence. Finding an appropriate classifier algorithm is a time-consuming task. For this reason, an algorithm that can automatically learn the choice of classification algorithms is very important. Classification algorithms are useful in predicting various diseases. Also, Primary Biliary Cirrhosis is one of the most well-known diseases that have been predicted by classification algorithms. This research's most significant achievement and novelty is the automatic increase in learning through a scoring method of reinforcement learning is called square learning (SL). In this research, an algorithm is presented that learns to automatically select the appropriate classification algorithm to predict Primary Biliary Cirrhosis. In this article, with inspiration from four evaluation metrics in classification algorithms, a new reinforcement learning method by the name of Fourth Degree Learning has been presented. In this research, we increased the performance of the classification algorithms used in this method from 63% of accuracy and achieved 98% accuracy.

cs.AI↗

ScalarLens: Numerical Embeddings with Stable Coordinates and Contextual Responses for CTR Prediction

Numerical embeddings for click-through rate (CTR) prediction are built on a convenient but restrictive premise: a scalar has one representation. This premise conflates where a value lies with what it means for the current sample. On the Criteo validation split, the same numerical interval carries residual click evidence with opposite signs across categorical and numerical contexts, even after additive main effects are removed. Production pipelines compound this mismatch because externally normalized features require transformations and statistics to remain synchronized between training and serving. We introduce ScalarLens, a numerical embedding that preserves what a value is while adapting how it should be interpreted. A monotone local mesh constructs a stable coordinate from the focal scalar alone; bounded low-rank dynamics then produce a contextual response without moving that coordinate or replacing categorical tokens and the CTR backbone. In a 1,539-run primary evaluation covering 19 representations, three datasets, nine backbones, and three seeds, ScalarLens ranks first in 25 of 27 settings on original numerical scales and second in the remaining two. Matched ablations show that scale correction, additional local capacity, and generic conditioning do not reproduce the gain. A controlled study further recovers categorical, numerical, and mixed response mechanisms under context shift while the focal coordinate remains exactly invariant. A complete rerun under shared standardization retains significant advantages over DEER, DAES, and NaryDis, showing that the result is not explained by tolerance to raw scales alone. ScalarLens therefore recasts numerical embedding as a measurement problem: coordinates belong to values, while predictive responses belong to values in context.

cs.IR↗

Predicting Emerging Topics from Outliers: A Prospective Study of Weak Signals in Embedding Space

Some documents that embedding-based topic models initially classify as noise later become founding members of emerging topics. At publication time, however, they appear as scattered points in embedding space and are difficult to distinguish from ordinary noise without the benefit of hindsight. We study whether such anticipatory outliers can be predicted prospectively, using only information available when a document first appears. We derive labels from the subsequent trajectories of outlier documents, distinguishing those that anticipate new topics from those that reinforce existing topics or remain isolated, and estimate label confidence through agreement across multiple embedding models. On two French news corpora, anticipatory outliers prove predictable at publication time. Under cross-validation, $F_1$ rises from about 0.77 over the full eligible population to above 0.90 on high-consensus subsets, and remains at 0.76-0.80 under a strictly chronological evaluation. Predictive performance is driven mainly by geometric features capturing each outlier's position in embedding space.

cs.CL↗

HistoRAG: A Citation-Grounded Question Answering Assistant for Teaching with Scanned Local History and Heritage Archives

Teachers who prepare lessons on local history and cultural heritage work from material that is hard to use. The primary sources are scanned books without a text layer, and the supporting records are administrative catalogs released as spreadsheets. A general chatbot answers such questions fluently but without a verifiable source, which is the property a teacher needs most. This paper presents HistoRAG, a question answering assistant that answers from one regional collection and cites a volume and a page for every fact. HistoRAG transcribes each page with a vision language model and keeps a line level confidence from the token probabilities. It builds three stores from the same collection: a hybrid text index, a relational catalog database, and a knowledge graph extracted only from entity dense passages. A lightweight router sends each question to the stores it needs, so that counting questions reach the database and relational questions reach the graph. We build a benchmark of 516 questions over a collection of 36 scanned volumes and the heritage catalogs of the same region, covering statistical, factual, temporal, and multi-hop questions. HistoRAG answers more questions correctly than passage retrieval baselines and than a graph based retrieval system, at a far smaller cost per question. The assistant runs behind a chat interface, so a teacher can check any statement against the page it came from.

cs.SE↗

$p$ and $rp$-Spectral Element Methods for Two-dimensional Elliptic Boundary Layer Problems

We propose $p$ and $rp$-spectral element methods for elliptic boundary layer problems on two dimensional rectangular domains. To resolve boundary layers, we use $p$-version with a fixed mesh and an $rp$-version with a boundary layer mesh consisting of thin needle-like elements near the boundary layer and coarse elements away from the layer. Stability estimates are derived using non-conforming spectral element functions. The numerical scheme for both the methods is based on minimising a least-squares functional in appropriate Sobolev norms. We construct robust preconditioners to manage the condition number of the normal equations derived from the least-squares formulation. The method is able to approximate boundary layers at a rate $O\left(\frac{\log W}{W^2}\right)$ for the $p$-version and at the rate $O(εα^{2W})$, uniformly in $ε$, for the $rp$-version, where $0<ε\leq1$ is the boundary layer parameter, $α<1$ is a constant and $W$ denotes the degree of the approximating polynomial. Numerical results are provided for model elliptic boundary layer problems for a range of boundary layer parameters. Simulation results demonstrate the efficiency and robustness of the method in capturing boundary layers on rectangular domains.

math.NA↗

An Automated Georeferencing Technique for Multi-Temporal Stope Point Clouds for Downstream Geotechnical Analysis

The increasing use of UAV laser scanning in underground mines has enabled frequent acquisition of 3D point clouds from challenging environments such as stopes, generating large volumes of multi-temporal spatial data throughout successive excavation stages. However, in GNSS-denied underground environments, independently acquired stope point clouds are generated within local scanner reference frames and require registration and georeferencing before integration with mine reference data for downstream geotechnical analysis, monitoring, and mine planning. This process is commonly performed manually by aligning individual stope scans with mine reference drives, making repeated georeferencing time-consuming and potentially limiting the utilisation of routinely acquired data. This study proposes the 3D Tag-based Automated Registration and Georeferencing Technique (3D-TARGeT), an automated framework using low-cost, generic, non-unique rectangular tags to establish spatial correspondence between stope point clouds and the mine reference coordinate system. The framework combines automated tag identification, geometric tag matching, and rigid transformation estimation. It was evaluated as a proof of concept using four multi-temporal point-cloud scans of an underground mine stope, with the proposed tags simulated under representative scanning conditions. 3D-TARGeT achieved consistent centimetre-level georeferencing accuracy, with median cloud-to-cloud distance and root mean square error below 0.03 m across all scans, while substantially outperforming widely used automatic point-cloud registration techniques. Overall, 3D-TARGeT provides an accurate and robust approach for automating stope point-cloud georeferencing, reducing reliance on manual alignment and facilitating multi-temporal datasets for downstream geological and geotechnical applications.

cs.CV↗

The Entropy Triangle Method (ETM): A novel framework for the prevention of cardiac arrhythmia with a review of more than 10,000 patients

One of the most important problems in medicine is to facilitate prediction. In this study, we propose entropy triangle method, a novel framework for predicting heart rhythms using a novel machine learning technique. This framework includes three steps: feature engineering, entropy triangle oversampling, and disease prediction. The dataset used in this study is a 12-lead electrocardiogram (ECG) arrhythmia research database with 10,646 patients. This dataset contains 11 different heart rhythms (5 sinus rhythms and 6 non-sinus rhythms). In this article, we introduce two firsts in machine learning and medicine that can predict non-sinus rhythm with over 85% accuracy. Our experimental results show, among others, that the most accurate classifier based on entropy triangles and the most useful oversampling are the supported vector classifiers and oversampling techniques for shark scent.

cs.AI↗

Local unmarked length spectrum rigidity for hyperbolic surfaces

Let $(M,g)$ be a closed negatively curved surface. If $g$ is strictly $\tfrac 19$-pinched and has the same unmarked length spectrum as a hyperbolic metric, we show that $g$ is hyperbolic. As a consequence, we show that any hyperbolic metric on a surface admits a $C^2$-neighborhood in the space of metrics in which it is characterized by its unmarked length spectrum, up to isometry.

math.DS↗

When Honesty is Not Enough in AI Debate

Scalable oversight aims to verify the behaviour of agents whose capabilities exceed those of their overseers. AI debate has been proposed as an oversight solution in which competing agents help a resource-limited verifier assess claims that it cannot reliably evaluate unaided. Much of its promise rests on incentivizing honest arguments that lead to correct verdicts. Yet a correct verdict need not uniquely determine the arguments used to support it. Agents may retain discretion over which correct claims to present, how to frame them, and in what order to disclose them. This residual freedom can allow agents to shape what the verifier learns beyond the task-relevant conclusion, pursuing latent objectives without compromising verdict correctness. To study this phenomenon, we introduce the framework strategic interactive oversight (SIO), which treats oversight jointly as a verification mechanism and a strategic communication channel. Within this framework, we formalise the notion of task-admissible latent optimisation, which entails the pursuit of latent objectives while maintaining a prescribed task performance. As proof-of-concept, we instantiate SIO in the establish protocol debate with cross-examination and quantify a tradeoff between task success and information disclosure about a hidden variable. The trade-off identifies a strategic window in which substantial disclosure remains compatible with task admissibility. Towards mitigation, we reduce admissible bias by expanding the cross-examiner's role to mitigate persistent disclosure over finite interaction horizons. Our results highlight the need to evaluate oversight not only by the correctness of its verdicts, but also by the information conveyed through its transcripts.

cs.AI↗

Antisymmetric breathing in altermagnetic skyrmions

A skyrmion in an altermagnet with \(d\) wave symmetry consists of two elliptical sublattice textures with perpendicular long axes. For each sublattice component $η=A,B$, we define an effective skyrmion radius $R_η=\sqrt{a_ηb_η}$, where $a_η$ and $b_η$ are the distances from the skyrmion center to its boundary along $y$ and $x$, respectively. For the anisotropic exchange parameters studied, the equilibrium radius at zero field is smaller than in a reference with parallel sublattice textures and otherwise identical parameters. A magnetic field perpendicular to the film expands one sublattice skyrmion and contracts the other, generating a radius difference $\dR=R_A-R_B$ and a net magnetic moment. The exchange modulation used to represent uniaxial strain also produces a nonzero radius difference at zero field. Because the strong and weak exchange directions are interchanged between the two sublattices, a common directional change of the exchange couplings increases one radius and decreases the other. Unlike the magnetic coupling, this mechanism vanishes when the two sublattice exchange tensors become identical. The radius difference also supports an antisymmetric breathing mode. After a short field pulse, it oscillates in quadrature with the uniform helicity, the common rotation of the wall magnetization within the film plane, at \(49.5\,\mathrm{GHz}\) for the reference parameters.

cond-mat.mes-hall↗

ASIRF: An Agentic Framework for Context-Dependent Sensitive Information Redaction

Sensitive information is defined by domain and intent, not a universal category, yet redaction systems such as privacy filters and named-entity recognizers fix a taxonomy at training time, requiring retraining for each new domain. We introduce ASIRF (Agentic Sensitive Information Redaction Framework), which retrieves domain-specific definitions based on the input's domain from a flexible knowledge base at inference time, needing no retraining to adapt. Two architectures, a three-call multi-agent pipeline and a single-agent variant, are evaluated across ten small open-weight models and eight datasets, including out-of-distribution fictional domains, against the OpenAI Privacy Filter (OPF) as a trained-classifier baseline. With only a few dozen expert-authored definitions per domain and no training data, ASIRF's recall exceeds OPF's in 68 of 80 model-domain combinations (85 percent), by at least one of the two architectures, with shortfalls confined mostly to OPF's training-distribution domains.

cs.AI↗

LC3EM: Long-Range Context Extrapolation Enhanced Entropy Model for Coordinate-based Overfitting Image Codecs

Coordinate-based overfitting image codecs have attracted increasing attention for their low decoding complexity and independence from cross-image generalization. However, representative approaches such as COOL-CHIC face an inherent entropy-modeling trade-off: lightweight models have limited capacity, while more expressive ones incur additional bitrate overhead from transmitting image-specific parameters. Inspired by the prediction mechanism in traditional codecs, we propose a new entropy-modeling strategy that introduces complementary prediction modes with region-adaptive soft mode selection, rather than relying on a single learned predictor to model diverse types of redundancy. Based on this concept, we develop a Long-Range Context Extrapolation Enhanced Entropy Model (LC3EM), which can be integrated into coordinate-based overfitting codecs. Specifically, a parameter-free Neighborhood-based Linear Extrapolation Mode (NLEM) complements the tiny MLP-based local predictor to exploit long-range contextual redundancy and strongly directional structures. A Minimum-Entropy-Inspired Continuous Mode Selection strategy is designed to adaptively fuse these two complementary modes, while requiring the transmission of only the parameters of a single additional linear layer. Moreover, to alleviate the mismatch between training-time relaxed and actual discrete quantization, we introduce a lightweight iterative latent rounding refinement stage to improve compression performance. Experiments demonstrate consistent improvements across diverse benchmarks, particularly on highly regular computer-generated images. When integrated with COOL-CHIC 4.0, the proposed method achieves BD-rate gains of -3.43\% and -7.69\% on the SIQAD and API datasets, respectively. With COOL-CHIC 5.0 as the backbone, the corresponding gains are -2.88\% and -3.15\%, respectively. The code will be made publicly available soon.

eess.IV↗

ImCorr: Sub-pixel Semantic Correspondence via Implicit Feature Decoding

The strong performance that modern semantic correspondence methods achieve at standard thresholds plateaus sharply at fine-grained thresholds. We argue that this plateau stems not from the representational capacity of backbone features, but from a grid-tied readout. Patch-based vision transformers tokenize images onto discrete grids, introducing two forms of quantization error: querying nearest patch features instead of exact keypoints on the source side, and the absence of grid features representing precise ground-truth locations on the target side. We quantify this quantization ceiling across all 499,188 keypoints in SPair-71k: under the standard 448x448, patch-14 setting, 84.9% of ground-truth keypoints have no grid feature representing their precise location at PCK@0.01. This is a structural limitation at the representation level, independent of the matching strategy. We address this with ImCorr: Sub-pixel Semantic Correspondence via Implicit Feature Decoding, which formulates correspondence estimation over a continuous feature field queryable at arbitrary continuous coordinates. A FiLM-conditioned decoder is trained to embed sub-pixel positional information into the feature field. Querying the field directly at exact keypoint coordinates theoretically eliminates representation-level quantization error on the source side, while decoding onto a grid denser than the backbone grid substantially reduces quantization error on the target side. On SPair-71k and AP-10K (intra-species, cross-species, and cross-family), ImCorr improves performance at fine-grained thresholds (PCK@0.01-0.05), achieving a 6.2 percentage point gain over the prior state of the art at PCK@0.01 on SPair-71k. These results demonstrate that representational continuity is an effective solution for precise semantic correspondence. Code is available at https://github.com/YusungChoi/ImCorr.

cs.CV↗

Continuous Online Fault Detection for Mobile Robots via Adaptive Edge Models

Mobile robots require robust, real-time fault detection capable of continuous adaptation on constrained edge hardware. While deep time-series models excel at unsupervised anomaly detection, their computational cost prohibits high-frequency onboard execution. This paper bridges this gap via a Teacher-Student distillation framework. An offline foundation model (TSPulse) generates pseudo-labels from unlabeled time series augmented with fault injections. A lightweight MiniRocket Student, adapted with a Recursive Least Squares estimator, approximates this complex decision boundary to execute real-time inference onboard. Evaluations on the TSB-AD benchmark and a physical mobile robot demonstrate the Student achieves a 4.30 ms CPU inference latency. During real-world domain shifts, online adaptation enables the Student to recover from unseen mechanical degradation, improving VUS-PR scores from 0.26 to 0.75 without catastrophic forgetting. Crucially, an uncertainty-guided active learning strategy minimizes operator cognitive load, requesting sparse interventions only when encountering novel fault distributions. These results validate the deployment of state-of-the-art anomaly detection on resource-constrained robotics through offline-to-online distillation.

cs.RO↗

Closed Response Calculus for SLE Weldings and Weil--Petersson Kähler Geometry

For $0<κ\leq4$, we construct a closed response calculus for $\mathrm{SLE}_κ$ weldings and a canonical Dirichlet form. The divergence covariance combines the Weil--Petersson and Velling--Kirillov forms. The integrated response gives exact changes of measure and canonical Liouville--capacity increments on conformally removable weldings.

math.PR↗