arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 865 records · Page 48Linked to original sources

Cross-Model Autoscaling for Shared LLM Serving

Multi-model LLM serving is moving toward shared MaaS clusters, where co-hosted models compete for a fixed GPU budget while each model experiences time-varying demand and must satisfy its own latency SLO. Existing LLM autoscalers remain largely model-local: their signals expose local runtime activity or delayed latency outcomes, and their scale-up or provisioning decisions do not directly determine how shared capacity should be allocated across competing models. We present the Token-service-share Rebalancing Engine (TRE), a control-plane framework for hot-switched multi-model LLM serving. TRE introduces Token Service Share (TSS), a calibrated, demand-normalized signal that estimates effective token service per active or queued request and yields a comparable health score across heterogeneous models and SLO classes. Guided by TSS, TRE coordinates bounded receiver--donor capacity movement under a fixed GPU budget: it separates fast rescue from slower rebalancing and incrementally reallocates active replicas toward models with the largest calibrated service deficits. We implement TRE on a Kubernetes-based hot-switch serving stack without modifying the inference scheduler. Across seven LLM serving traces, TRE reduces P95 end-to-end latency by 11.9--79.0\% and P99 latency by 12.5--72.6\% compared with a state-of-the-art KV-cache-based reactive autoscaler running on the same hot-switch runtime. The gains hold on both targeted stress probes and production-derived conversation/code traces, where TRE reduces P95/P99 latency by 50.8/63.7\% and 79.0/72.6\%, respectively. These results show that effective hot-switched autoscaling requires not only fast replica actuation, but also calibrated service-deficit signals and coordinated cross-model capacity arbitration. Our code and artifacts are available at https://github.com/zxzx9898/Token-service-share_Rebalancing_Engine.

cs.DC↗

A framework for linking literature-based knowledge integration and infrastructure-supported knowledge integration: Opportunities and challenges from a case study

Integrating knowledge across disciplines is central to sustainability research, yet most evidence-synthesis methods rely on findings as reported in publications, limiting verification and reuse of underlying data and workflows. We develop a conceptual framework linking literature-based and infrastructure-supported knowledge integration, using a systematic review case study to examine when integration can extend beyond reported findings. We reviewed 37 studies on climate change, violent conflict, and household food security. Literature-based synthesis enabled integration across all included studies, whereas access to reusable outputs was limited: over half provided no data availability statement, 27% reported availability upon request, but reusable data and workflows were available for only 8%. To explore infrastructure-supported integration, we used the TIB Knowledge Loom to represent studies with accessible data and code as machine-readable outputs, and produced a knowledge gap map (KGM) from manually extracted and Loom-derived data, comparing manual and infrastructure-supported synthesis. Where outputs were reusable, synthesis could be produced directly from data and workflows rather than from publications. These findings show that literature-based synthesis can be complemented by infrastructure-supported integration where outputs are accessible and usable, and that advancing knowledge integration depends not only on infrastructures but on making data, code, and workflows accessible, executable, and reusable.

cs.DL↗

Binary-pulsar Timing: A 3D Solar-System Accelerometer and Unknown-Source Probe

The acceleration of the Solar System barycenter (SSB) offers a unique precision probe of the local gravitational environment, enabling searches for otherwise invisible gravitating sources such as Planet Nine and nearby (primordial) black holes. Using the binary-pulsar timing measurements, we develop an analysis framework that combines the line-of-sight differential accelerations of 26 binary pulsars to jointly reconstruct the three-dimensional acceleration of the SSB and place directional upper limits. Across three smooth Galactic-potential baselines, the residual acceleration is consistent with zero. We obtain directional 95\% upper limits of $0.198$, $0.262$, and $0.693~μ\mathrm{as}\,\mathrm{yr}^{-1}$ over 50\%, 75\%, and all sampled directions, respectively. These limits are tighter by factors of 3.5, 3.6, and 1.6 than a matched Gaia EDR3 benchmark and improve previous pulsar constraints by more than an order of magnitude. We further extend the framework to the full SSB--pulsar two-endpoint response, enabling direct position-dependent mass constraints on unknown gravitational sources. A $10\,M_\odot$ object is excluded within 0.39, 0.34, and 0.21 pc over the same sky fractions. Our unknown-mass limits surpass the tidal-equivalent INPOP19a reference beyond 0.2 pc by about an order of magnitude over most of the sky at 1 pc.

astro-ph.GA↗

Edge AI on Constrained Devices for Binary Sleep-Wake Classification in Dynamic Environments

This paper presents an Edge AI-based system for detecting sleep and wake states in non-stationary mobile environments using resource-constrained embedded hardware. Conventional approaches relying on accelerometer-based activity metrics are highly susceptible to motion and vibration artifacts and are limited by strict compute and energy budgets of wearable and IoT devices. To address these challenges, a multimodal pipeline is designed and implemented on an ESP32-S3 microcontroller. The system combines inertial sensing for head movement analysis and visual pose classification. A dual-core architecture with FreeRTOS enables parallel execution of real-time data acquisition and on-device inference. Sleep detection follows a two-stage strategy: low-movement detection over a temporal window, followed by visual validation of poses. Experimental results show accuracies of 96.5% for motion-based detection and 89% for pose classification, yielding robust binary sleep-wake classification. Field tests confirmed feasibility in representative mobile scenarios. The results demonstrate that privacy-preserving, local sleep detection is achievable on edge hardware through careful co-design, while highlighting limitations in sensing intrusiveness, dataset scale, and system integration.

cs.LG↗

Enhancing 5G NTN VSAT RACH in GNSS-Denied environments

The 3GPP 5G non-terrestrial networks (NTN) technology fundamentally relies on a global navigation satellite system (GNSS) fix at the user equipment (UE) to pre-compensate for user-link delay and Doppler shifts prior to any transmission, including the random-access procedure for initial access. In GNSS-denied or interfered environments, this dependency triggers preamble collisions and connection dropouts, undermining satellite-based global connectivity. To address this vulnerability, this paper introduces a novel single-satellite positioning framework designed for higher-frequency (Ku- or Ka-band) 5G NTN very small aperture terminals (VSATs) equipped with a phased-array antenna. The multi-step scheme derives an initial coarse position from angle of arrival (AoA) observables, extracts frequency of arrival (FoA) and time of arrival (ToA) measurements from low Earth orbit (LEO) downlink reference signals, and hybridizes them within a Levenberg-Marquardt navigation filter. Systematic evaluations across static grids and realistic dynamic urban trajectories detail the incremental gains of transitioning from single-epoch AoA error minimization to multi-epoch angular tracking, followed by the integration of Doppler and ranging observables. Grounded in a realistic error budget, the proposed framework reduces initial positioning uncertainty from tens of kilometres down to a few kilometres. This level of positioning accuracy, combined with the GNSS-resilient features introduced in 3GPP Release 20 NR NTN specifications, enables robust initial access in compromised operational environments with minimal relaxation of 3GPP standards.

eess.SP↗

Optical spectropolarimetry of extreme H$α$ line profiles in seven active galactic nuclei

Some active galactic nuclei (AGNs) are known to show extremely asymmetric broad Balmer lines, with their main peak redshifted or blueshifted by thousands of km s$^{-1}$, severely challenging our understanding of the standard AGN paradigm. We wanted to explore the causes of such asymmetric features by carrying out a spectropolarimetric study of a sample of bright AGNs with known extreme Balmer line profiles. We present optical spectropolarimetry of seven bright (V $<$ 17.5) Seyfert-1 galaxies obtained with the Very Large Telescope (VLT) FOcal Reducer/low dispersion Spectrograph 2 (FORS2) instrument from March 2012 to January 2013. The main broad Balmer lines (H$α$ and H$β$, and H$γ$ occasionally) were sampled with a resolving power ranging from 350 to 500. We performed a detailed spectral decomposition of each H$α$ profile to isolate kinematic substructures within the broad line region (BLR) and measured the polarization degree and angle of each component, providing constraints on the inner geometry and scattering properties. The broad H$α$ components exhibit various scattering geometries. Their polarization as a function of full width at half maximum (FWHM) reveals two distinct behaviors, corresponding to either a BLR of roughly constant scale height or a stratified structure with decreasing height at higher velocities. This supports a disk-wind geometry linking equatorial and polar regions.

astro-ph.GA↗

HarnessPAI: An Evolving Harness for Physical AI

Physical AI aims to build embodied agents that perceive the world, understand and reason about it, and decide how to act. Yet the field has focused primarily on the last component: the action model that maps observations to low-level controls. The prevailing training recipe can erode the perceptual and reasoning capabilities needed for robust behavior, leaving even strong action models vulnerable to scene perturbations and long-horizon tasks. We introduce HarnessPAI, a model- and embodiment-agnostic Harness framework for Physical AI that treats code as the executable and evolvable interface that organizes the underlying action primitive. The framework separates two timescales: within a rollout, it executes open-loop at the program level, with a fixed program guiding and checking execution; across rollouts, it evolves closed-loop, using execution feedback to revise the program and distill failures into reusable skills. Across desktop robot arms, household robots, a robot vacuum, and a legged walking agent, HarnessPAI improves on both pure action models and code-as-policy baselines without retraining the underlying model: a 61.6-point gain over $π_{0.5}$ on LIBERO-PRO and a 27.2-point gain over WorldDreamer on RoboCasa atomic tasks. Once a program is selected, rollout execution requires no online high-level LLM deliberation. Beyond execution, the converged program is also a cheap and reliable expert-data collector, and fine-tuning $π_{0.5}$ on collected expert data lifts success rate on LIBERO-PRO by 38.8 points. Our results suggest that the frontier of Physical AI depends not only on stronger action models, but also on executable harnesses that integrate perception, task understanding and reasoning, and action execution into a unified, verifiable, and feedback-driven system. Website: https://darwin-agent.github.io/HarnessPAI

cs.RO↗

IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking

Banking assistants must use account-specific information to answer requests and, in many cases, take actions through tools. Evaluating only the final response misses important errors. An assistant may ask for information it already has, rely on stale context, select the wrong account, or write an invalid value after stating the correct one. We introduce IndicBankBench, a 799-case benchmark for Indian retail banking spanning five operational domains, a capability/refusal domain, and twenty primary axes. Cases are evaluated at four stages: safety, action and tool use, response adequacy, and advisory quality. Tool use and most safety checks are deterministic. A narrow resolver handles only ambiguous confirmation-before-write cases, while a separate LLM judge evaluates semantic response adequacy. We run every case three times and report strict pass^3, which requires success on all trials. Across the eleven evaluated models, strict reliability ranges from 43.7% to 58.2%, whereas at-least-once success ranges from 60% to 74%. This gap shows that at-least-once success can overstate dependable banking behavior. The case-level diagnostics also distinguish systems that ask unnecessary questions from those that act but fail to reconcile customer context or fully resolve the request. We release the cases, mock environment, and evaluation harness.

cs.AI↗

Hidden magnetic order within the pressure induced superconducting dome of UTe2

Unconventional superconductivity typically occurs near magnetic instabilities, and the corresponding spin fluctuations are widely believed to play a crucial role in mediating electron pairing. UTe$_2$ is a promising candidate for exhibiting multiple spin-triplet superconducting phases when tuning with applied pressure and magnetic fields, but the nature of the magnetism driving these unconventional pairing states is undetermined. Our measurements of UTe$_2$ under applied pressures and magnetic fields reveal the presence of a magnetic order hidden within the pressure-induced superconducting dome, which vanishes together with the superconductivity once there is sufficiently high pressure to induce the three-dimensional antiferromagnetic phase. Extrapolation of the phase boundary of the hidden magnetic order, which is most likely antiferromagnetic in nature, points to a zero-temperature quantum critical point that coincides with the maximum transition temperature of the pressure-induced superconducting dome, suggesting that it could corresponds to the parent magnetic phase of the critical antiferromagnetic spin fluctuations driving the triplet superconductivity. These findings advance the understanding of the interplay of magnetism and superconductivity in an exemplar candidate triplet superconductor, which is necessary for revealing the microscopic origin of the different unconventional superconducting phases.

cond-mat.supr-con↗

Spot, Separate, and Enhance: Fully Generative Approach for Audio Mixing

We introduce Spot, Separate, and Enhance (SSE), the first multimodal, user-guided generative model for audio remixing and enhancement. SSE enhances video content by rebalancing the audio, removing unwanted audio sources, and reducing reverberation, guided by both video and textual descriptions. To support its training and evaluation, we propose DegradedMix, a new dataset built on the audio remixing benchmark MuddyMix. We also adopt evaluation metrics from generative modeling, which better capture the creative nature of remixing than standard reconstruction-based metrics. SSE outperforms existing baselines in both controllability and remixing quality, as shown by extensive experiments. Project page: https://sse-ai.notion.site

cs.SD↗

Sharp stability for the second-order weighted Heisenberg Uncertainty Principle with explicit stability constants and optimizers

By using spherical harmonicas decomposition and Gaussian-type Poincaré inequalities, we establish several sharp stability estimates for the following second-order weighted Heisenberg Uncertainty Principle \begin{equation*} \int_{\mathbb{R}^N} \!\frac{|Δu|^2} {|x|^{2α}} \mathrm{d}x \int_{\mathbb{R}^N} |x|^{2α+2}|\nabla u|^2\mathrm{d}x \geq \frac{(N+4α+2)^2}{4} \left(\int_{\mathbb{R}^N} |\nabla u|^2 \mathrm{d}x\right)^2, \end{equation*} and \begin{equation*} \int_{\mathbb{R}^{N}} \frac{|Δu|^{2}} {|x|^{2α}} \mathrm{d}x +\int_{\mathbb{R}^{N}} \left|x\right|^{2α+2} |\nabla u|^{2}\mathrm{d}x \ge\left(N+4α+2\right) \int_{\mathbb{R}^{N}} |\nabla u|^{2}\mathrm{d}x. \end{equation*} We also provide the explicit value and the necessary and sufficient condition for attainability of the sharp stability constants. Moreover, when $α=0$, our results reduce into those of [\emph{Calc. Var. Partial Differential Equations} \textbf{64} (2025), Paper No. 129], [\emph{J. Funct. Anal.} \textbf{290} (2026), Paper No. 111321] and [arXiv:2510.00453].

math.AP↗

Representation World Model: Learning States, Transition and Executable Plans in Representation

We propose the Representation World Model (RWM), which learns states, transitions, and executable plans directly in representation space. Unlike existing world models that typically learn latent representations together with explicit dynamics models and perform planning through search, optimization, or policy-based prediction, RWM directly incorporates planning into the learned representation geometry. RWM learns the representation geometry by applying inverse-dynamics supervision locally along latent paths constructed from endpoint representations, requiring these paths to preserve task-relevant state and transition information. At inference, planning is performed by directly constructing a latent path between the current and goal representations, with inverse dynamics used to recover the corresponding actions, without recursive rollouts or action-space search. Experiments on continuous-control benchmarks demonstrate the effectiveness of RWM for direct planning, while results on robotic manipulation further show its potential to extend to more complex embodied control tasks. These results suggest that planning directly in representation space provides a promising alternative to conventional world-model planning.

cs.RO↗

Orchestrating AI-Assisted Code Remediation: Socio-Technical Bottlenecks in a Large Industrial Repository

Background: Code degradation in large, long-lived codebases is costly to remediate through manual refactoring and opportunistic clean-ups. LLM-based coding assistants can perform mechanical remediation at scale, but their impact on industrial workflows is underexplored. Objective: We investigate how massive AI-assisted code remediation affects build-on-commit continuous integration (CI), code review, and team coordination in a large industrial repository, and which socio-technical bottlenecks constrain such remediation when source editing becomes cheap through AI assistance. Method: We report on a 15-day exploratory single-case field study in which an experienced developer used a command-line AI coding buddy to remediate widespread issues in a closed-source industrial C++ repository. We triangulate Gerrit metadata with a developer diary and team chat, analyzed through descriptive statistics and qualitative coding. Results: AI-assisted remediation rapidly generated hundreds of commits touching thousands of lines, saturating CI and reviewer attention. Naïve per-file commits overloaded build-on-commit CI; Switching to directory-based batching and capping the number of files per change restored throughput, but still required explicit review solicitation, negotiation of acceptable commit granularity, and iterative follow-up to resolve build and static-analysis failures. Conclusion: When mechanical editing is cheap, CI capacity, review effort, and change orchestration become primary bottlenecks. Sustainable AI-assisted remediation in very large repositories requires deliberate control of commit, review, and CI batch granularity and treating semantic change sets, such as ``fix all instances of warning X'', as first-class units of work that can be sliced differently for developers, reviewers, and CI.

cs.SE↗

Through Human Eyes and Machine Eyes: Understanding View Mismatch in Video See-Through Extended Reality

Video see-through extended reality (VST XR) systems commonly use headset screenshots or captured frames as proxies for the user's first-person visual context. However, the system-captured view and the user's effective visible field do not necessarily coincide: a screenshot records a rectangular machine-readable frame, whereas the user's effective visible region can be more constrained and non-rectangular. This paper studies this human-system view mismatch in VST XR. We formalize the relationship between the system-captured region and the human-visible region by defining their co-visible, system-only, and human-only regions. \rev{We then conduct a pilot-level boundary measurement on Meta Quest 3, revealing a clear mismatch between the rectangular screenshot frame and the approximate human-visible boundary. Building on this model, we analyze how view mismatch can affect screenshot-based XR sensing and downstream vision-language model tasks. Through four representative case studies, we illustrate potential risks and failure modes including prompt injection, privacy leakage, human-invisible information bias, and missing human-visible information. Our results show that view mismatch is not only a geometric artifact, but can also introduce security, privacy, and reliability concerns for AI-integrated VST XR systems.

cs.HC↗

Metal-Poor Stars as Metal-Rich Impostors via Planet Engulfment

Planet engulfment leads to planetary mass accretion and angular momentum deposition into the stellar envelope. This study uses MESA models to investigate the long-term impact of engulfment on stellar metallicity and rotational velocity, as well as the effect of rotational mixing on metal enrichment. Our models show that engulfment occurring during the late main-sequence (MS) stage yields significantly higher surface enrichment than at the ZAMS. Rotational mixing generally suppresses enrichment; however, metal-poor stars can sustain super-solar abundances for several Gyr, whereas high-metallicity stars remain largely unaffected. Although engulfment causes an instantaneous spin-up, this angular momentum injection is transient, and the star quickly re-converges to its standard rotational evolution. Thus, MS spin-up cannot serve as a long-term observational indicator. Because convective envelope thickness and MS lifetime dictate the magnitude and duration of enrichment, identifying engulfment signatures requires balancing the long-lived but diluted signals in 0.7 M_sun stars against the strong yet short-lived enrichment in 1.2 M_sun stars. Consequently, 1.0 M_sun stars offering both substantial enrichment and a relatively long MS lifetime represent optimal targets for observational searches.

astro-ph.SR↗

A Human-Like Pedestrian Model for Automated Driving Simulations

Automated vehicles must be able to interact with pedestrians safely and efficiently across diverse traffic situations. Although driving simulators offer a scalable testbed for learning such capabilities, existing theory-inspired pedestrian models are narrow in scope and limited to go/no-go crossing decisions in single-lane settings. While data-driven approaches can predict pedestrian behavior in complex situations, they lack sufficient observations in rare, safety-critical scenarios. Here, we propose an approach to training pedestrian models in simulators so that learned policies generate demonstrably human-like behavior in realistic, complex traffic scenarios, including multiple lanes, heavy traffic, and dangerous driving styles. Our technical contribution is a novel definition of pedestrian-vehicle interaction as a partially observable Markov decision process (POMDP) with theory-grounded perceptual, cognitive, and motor constraints. It accounts for the highly adaptive nature of human behavior in traffic and simulates how people adjust their responses according to perceived danger, time pressure, and the complexity of the situation. When trained via deep reinforcement learning (RL) with domain randomization in a simulator, the model reproduces the broadest range of empirical findings shown so far on human crossing behavior, including gap acceptance, yielding acceptance, hesitation, and evasive speed adjustment. We show that learned policies transfer to unseen traffic environments, and can be further adapted to local traffic norms with finetuning. Together, these results establish a blueprint for simulator-ready pedestrian models that can support the development and evaluation of automated driving systems.

cs.HC↗

Anthropomimetic Soft Robotic Forearm with Independently Articulated Carpal Bones Enabling Human-Like Adaptive Stiffness Modulability

The human wrist exhibits adaptive stiffness modulability: joint stiffness anisotropy can be actively regulated through muscle co-contraction. This functionality is essential for stable manipulation, yet the underlying morphological factors remain unclear. To identify these factors, we developed an anatomically accurate anthropomimetic soft robotic forearm comprising eight independently movable carpal bones interconnected by ligaments, 22 actuated muscles, and compliant fingertips. We measured wrist joint stiffness under four muscle activation patterns across three skeletal configurations: anatomically normal carpal bones, a fused proximal carpal row, and a geometric ellipsoidal skeleton. The stiffness ellipse exhibited low stiffness along the dart-throwing motion (DTM) direction when finger muscles were activated, but high stiffness along the same direction when wrist and finger muscles were activated simultaneously. These results agree with previously reported human measurements, demonstrating that precise anatomical replication reproduces human-like stiffness modulability. Fusing the proximal carpal row eliminated the low DTM-direction stiffness under finger muscle activation, while the geometric ellipsoidal skeleton showed poor stiffness ellipse reorientation across all conditions. Carpal bone motion analysis revealed significantly opposing coupling patterns between wrist and finger muscles at the proximal carpal row, accompanied by a consistent but non-significant trend at the midcarpal joint, providing a mechanical explanation for this modulation. These findings demonstrate that carpal bone morphology plays a dominant role in human wrist stiffness modulation and provide design principles for humanoid robot wrists.

cs.RO↗

Duality for matrix space questions

We present a new proof of a classical theorem of Dieudonné: if a linear space of $n\times n$ matrices consists entirely of singular matrices, then its dimension is at most $n^2-n$. Our proof is based on a surprising ``duality'' argument: we prove this universal upper bound by exhibiting a single matrix space that serves as a lower bound for a related problem. Interestingly, this approach only works for certain fields, but we use model-theoretic arguments to obtain the same result for all fields. We hope that this approach can be generalized to provide new duality-based proofs of other classical theorems on matrix spaces, and give some preliminary results in this direction.

math.RA↗