arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,189 records · Page 66Linked to original sources

Observing Colossal and Tunable Near-field Thermal Radiation with A Highly Sensitive Annular Micro-thermocouple Junction

Near-field radiative heat transfer between two objects has been theoretically predicted and experimentally demonstrated to exceed far-field blackbody limit across nanoscale vacuum gap distance enabled by evanescent surface waves coupling, while significant enhancement by more than 100 times usually requires sub-50-nm vacuum gaps around room temperature. This is challenging for parallel-plate configuration with millimeter sample sizes due to intrinsic wafer bow and contaminant particles. Sphere-plate configuration with a microsphere attached to bimaterial cantilevers or micro-thermocouple tips has been used to experimentally demonstrate near-field radiative heat transfer down to 30-nm gaps, but it is much less developed because of the challenges in the sophisticated sensor fabrication, low sensitivity and weak signals. In this work, we overcome these challenges by an annular micro-thermocouple junction fabricated at the end of a glass fiber with straightforward thin-film deposition to achieve high measurement accuracy with large Seebeck coefficient 25 uV/K and thermal resistance 8.7e6 K/W. With a silica microsphere attached underneath the micro-thermocouple junction, we report experimental observation of colossal near-field radiation heat transfer over blackbody limit down to 10-nm gap up to 2600 times with quartz and 900 times with doped silicon. Upon phase transition of VO2 thin film emitter, tunable near-field heat transfer up to 430-fold enhancement is experimentally demonstrated at 15-nm gap with 64% reduction. Experimental data agrees well with rigorous modeling based on fluctuational electrodynamics and Derjaguin approximation, and underlying mechanism is understood by energy transmission calculations. The results will advance the experimental study and fundamental understanding of energy transport at nanoscale gaps.

cond-mat.mes-hall↗

Predictive Semantic Safety: From Visual Physical Reasoning to Safety-Critical Control

Physical interactions can create future hazards that are not apparent from the robot's current geometric surroundings. We present a framework termed Predictive Semantic Safety (PSS), which connects visual physical reasoning to backup-based safety filtering. A vision-language model (VLM) predicts physical events and their timing or directly predicts object displacements. An explicit motion model converts event hypotheses into object trajectories. Split conformal prediction calibrates position errors jointly across specified objects, observation times, and future times; geometric shape bounds convert the resulting position regions into predicted object occupancy. PSS evaluates a prescribed backup maneuver against this occupancy and derives input-affine constraints for minimally modifying the nominal input while preserving backup feasibility under the robot dynamics and input limits. MuJoCo experiments with a Unitree Go1 consider falling fixtures, impact-driven support loss, and contact propagation. PSS achieves a safe episode rate of 99.3%, compared with 43.3% for a Backup Control Barrier Function baseline that only uses current obstacle geometry.

cs.RO↗

R Coronae Borealis fadings: a dusty gas cloud eclipse model

R Coronae Borealis stars (RCBs) are hydrogen-deficient supergiants that undergo deep fading events due to transient extinction by dust. First observed over two centuries ago, that behaviour remains poorly understood. Our aim is to investigate the possibility that the fading events are eclipses by orbiting, dusty gas clouds. We construct a simple physical model of the dust wind that is driven from the cloud surface by the stellar radiation field, and we calculate the extinction and scattered light from wind and cloud. We consider a broad range of dust compositions. For the model that best matches the data, we sketch out implications for the RCB phenomenon more broadly. Our calculations show that conventional dust materials like graphite and silicate produce eclipses whose ingress is too slow and whose morphology too symmetric to account for RCB fadings. But dust made of solid H2 yields a good match: the eclipses are deep and strongly asymmetric, with rapid onset and slow recovery, and event timescales are about right if the clouds' periastra are within a few tens of AU. Such close approaches imply tidal stripping/disruption of each cloud, with hydrogen-deficient debris accreting onto the star via a disk. Starlight incident on the disk creates a metal-ion plasma that emits a powerful bremsstrahlung continuum, peaking in the mid-IR. We conclude that eclipses by orbiting, dusty clouds could be the cause of RCB fading events, but only if the dust is made of solid H2, in which case both the unusual, hydrogen-deficient nature of the stars and their mid-IR excess follow naturally. It is, however, challenging to account for the high eclipse rates that are observed; we propose a scenario in which the progenitor star is a wide binary inside a massive halo of clouds, suggesting a connection to the Galactic "missing mass" problem.

astro-ph.GA↗

FORGE: Form-Optimal Routing of Grounded Evidence for Frozen LLM Agents

In agentic AI systems, frozen foundation models are increasingly deployed as closed-weight API endpoints, making downstream adaptation possible only through the inputs and inference procedures surrounding the model. As a result, for each input query, two coupled decisions largely determine both answer quality and token cost: what evidence to provide and how much reasoning budget to allocate. Fixed defaults along these axes are often suboptimal, misallocating support form or reasoning depth on roughly 80% of queries in our analysis. To address this challenge, we propose FORGE, a unified framework for adapting frozen models through per-query routing over a joint action space that spans both support form and thinking depth. Under an entropy-regularized, cost-aware utility objective, we derive a closed-form Boltzmann routing target and instantiate the policy as a lightweight 269K-parameter factorized router. The routing policy is trained around the frozen host, without any weight access, through a three-stage pipeline: offline arm enumeration, supervised Kullback-Leibler (KL) distillation from the Boltzmann target, and Group Relative Policy Optimization (GRPO) refinement with host feedback. Across 5 knowledge-intensive benchmarks and 8 frozen backbones ranging from 7B to 671B parameters, FORGE improves accuracy at 42-45% lower token cost on both main hosts, transfers zero-shot across hosts at lower token cost, and composes with intrinsic thinking budgets where available.

cs.LG↗

Improving Large Language Models for Code through Runtime Program-State Reasoning

Large language models receive limited explicit training in reasoning about runtime program states. We study whether training models to reason about runtime program states improves downstream software-engineering capabilities. We introduce two complementary program-state reasoning tasks. Buggy input-output reasoning requires a model to generate a concrete input that exposes a behavioral difference between a buggy program and a hidden correct implementation and to predict the resulting execution behavior. Precondition-postcondition reasoning requires an agent to symbolically characterize a bug-triggering precondition, predict the expected postcondition, explain their causal connection, and instantiate this reasoning as an executable regression test. By incorporating these two tasks into a staged post-training pipeline, we develop Comet-9B, a 9B language model based on Qwen3.5-9B Base. We evaluate the resulting checkpoints on repository-level patch generation, regression-test generation, and security PoC generation. Adding both program-state reasoning tasks to supervised fine-tuning (SFT) on issue resolution improves success rates by 7.25 percentage points on SWE-bench Pro and 9.70 points on SWT-Bench Verified. Sequential reinforcement learning on the two tasks yields further gains of 7.25, 26.79, and 4.67 percentage points on SWE-bench Pro, SWT-Bench Verified, and CyberGym, respectively. Despite having only 9B parameters, Comet-9B achieves a score comparable to the reported GPT-5.2 result on SWE-bench Pro and matches the reported success rate of a GPT-4o-based agent on SWT-Bench Verified.

cs.AI↗

CoeF-SFL: Preserving Collaborative Server-Client Learning with Enhanced Communication Efficiency

Split Federated Learning (SFL) enables resource-constrained clients to participate in collaborative training, but vanilla SFL exchanges smashed data and gradients at every batch, which incurs significant communication overhead. Recent methods reduce this overhead with an auxiliary network at the client-side cut layer. However, we identify that this approach makes the client optimize a local objective that differs from the end-to-end objective, which fundamentally limits the collaborative training between the client and the server. We propose Compensated Feedback based SFL (CoeF-SFL), a communication-efficient framework that retains the end-to-end objective without any auxiliary network. In CoeF-SFL, the client and the server exchange the smashed data and the gradients once per round and reuse them during local training. Since this reuse makes the gradients stale on the client side, we compensate them with a curvature-based correction in the activation space and develop two variants. CoeF-D approximates the Hessian with a diagonal gradient outer product, while CoeF-J exploits the tractable Jacobian-based Hessian of a surrogate loss that upper-bounds the true loss. We provide the theoretical background of each method, characterizing its compensation. Across vision and language tasks, model capacities, cut layers, and data distributions, CoeF-SFL significantly outperforms auxiliary-network-based methods under the same communication frequency, and the improvement is most substantial on vision tasks. Code is available at https://anonymous.4open.science/r/CoeF-SFL-2686/README.md

cs.AI↗

FutureDuet: Decoupling Observation Access from Future Supervision in World Action Models

World Action Models (WAMs) augment robot action generation with future visual supervision. Existing WAMs commonly fuse main and wrist observations into one visual stream and train both with the same future-video objective, despite their different visual dynamics. A stable main camera reveals scene-level task evolution, whereas wrist cameras move with the end effector, mixing local interaction changes with viewpoint shifts and self-occlusion. These contrasting predictive demands suggest that the two views may benefit from different future objectives. We introduce FutureDuet, which retains both views for control, while allowing each visual stream to receive a different future objective. For the main view, future RGB models task evolution, while interaction masks and robot skeletons focus supervision on task objects and robot motion. For the wrist stream, future latent prediction models short-horizon interaction changes without requiring pixel-level reconstruction. ActionDiT jointly reads the resulting Task State and Interaction State, combining scene-level progress with close-range interaction evidence. All auxiliary prediction modules are training-only, adding no inference overhead. FutureDuet achieves 94.2% clean and 94.1% randomized success on RoboTwin50 and 99.2% average success on LIBERO. The improvements are most pronounced on six RoboTwin50 tasks that require precise interaction, averaging gains of 9.2% and 12.8% over Fast-WAM in clean and randomized settings. Controlled studies further show complementary gains from separating the wrist pathway and designing future supervision separately for the two views.

cs.RO↗

SyncRA: Learning Temporal Correspondence in Omni-Modal Models

Recent omni-modal models demonstrate strong perception of audio and visual inputs, yet often struggle to connect what they hear with what they see at the same moment. This weakness in temporal correspondence can cause models to associate spoken cues with the wrong visual scenes, producing plausible answers grounded in incorrect audio-visual pairings. We diagnose this problem through controlled temporal swaps, revealing that model answers do not reliably follow changes in these pairings. To address it, we propose Synchrony-Guided Representation Alignment (SyncRA), a lightweight method for strengthening temporal correspondence between audio and vision. Specifically, SyncRA contrasts intermediate audio-visual representations within each video, aligning matching moments while separating mismatched ones to capture local temporal correspondence within a shared global context. The objective derives supervision directly from existing input timing, requiring no additional annotations and leaving inference unchanged. We evaluate SyncRA across four open omni-modal models spanning different sizes and architectures on five public video benchmarks. SyncRA consistently outperforms answer-only fine-tuning across all model-benchmark combinations, while substantially improving the ability to track changing audio-visual pairings in controlled evaluations. These results demonstrate that lightweight, targeted supervision can effectively strengthen temporal correspondence and translate into broad improvements in audio-visual question answering.

cs.CV↗

Next-to-Next-to-Leading Order QCD Corrections to $η_t \to HZ$

Motivated by recent LHC observations of a threshold enhancement in the $t\bar t$ system consistent with toponium-like dynamics, we compute the next-to-next-to-leading order (NNLO) QCD correction to the hard short-distance coefficient for $η_t\to HZ$, where $η_t$ denotes the pseudoscalar color-singlet configuration $t\bar t({}^1S_0^{[1]})$. Within the NRQCD factorization framework, we retain the full dependence on the Higgs- and $Z$-boson masses and include both the two-loop virtual corrections and the real double-gluon-emission channel. At next-to-leading order (NLO), the finite-mass result remains close to its large-mass limit. The NNLO coefficient develops logarithms of both the renormalization scale $μ_R$ and the NRQCD factorization scale $μ_Λ$. For $μ_R=μ_Λ=m_t$, the NLO and NNLO terms suppress the leading-order width by about $29\%$ and $19\%$, respectively; the two-loop contribution is about two thirds of the one-loop effect and of the same sign, so the perturbative series converges slowly and the NNLO prediction amounts to about $52\%$ of the leading-order width. The resulting coefficient provides a necessary ingredient for future threshold studies of $η_t\to HZ$ and for assessing the sensitivity of this channel to the top-Higgs interaction. The accuracy of the result is further supported by three checks: reproduction of the known large-mass limit at one loop, independence of the extracted coefficient on the $γ_5$ prescription, and consistency of its scale logarithms with the renormalization-group structure.

hep-ph↗

Cavity magnonics and bound states in the continuum with Bragg and anti-Bragg mirrors

A periodic resonant emitter array coupled to a one-dimensional waveguide can act as a mirror. A typical example is a Bragg mirror, where emitters are spaced by multiples of half wavelengths and collectively enhance light reflection. Two such mirrors form an effective cavity that hosts bound states in the continuum (BICs). However, generating and detecting the spatial profiles of such BICs has proven challenging. Here, we experimentally demonstrate BICs in cavity magnonics using two periodic ferrimagnetic-sphere arrays and a probe sphere in a dual-open-waveguide architecture. We realize both Bragg and anti-Bragg cavities, whose mirrors have different lattice constants. In the Bragg cavity, there are degenerate supermodes formed by mirror spheres. We show that a single dark supermode coherently couples to the probe magnon in the cavity region, forming two polaritons whose splitting scales with both the probe-sphere size and the number of mirror spheres. By contrast, the anti-Bragg cavity has bright and dark supermodes in a bandgap, substantially changing the magnon-cavity interaction. Moreover, by moving the probe sphere, we perform position-dependent detection of the cavity field, highlighting the role of BICs in these cavities. Our flexible experimental setup with a scanning probe opens possibilities to detect other exotic states created by light-matter interaction, to interface with superconducting circuits in hybrid quantum networks, and to study non-Hermitian physics with Bragg and anti-Bragg cavities.

quant-ph↗

When Harness Beats Scale, and When Reading Beats Both

We describe our system for DocSem, the document-grounded quantitative reasoning shared task at DocInsights 2026, and analyze why it succeeded on labeled data and failed on the test set. The pipeline pairs hybrid block retrieval with Program-of-Thoughts (PoT) generation executed in a sandboxed interpreter, self-consistency sampling, and entity enrichment from chunk-level knowledge graphs. On our held-out split, application architecture moved the metrics far more than model scale did: PoT added 0.282 joint accuracy to a compact 7B model but at most 0.005 to a 72B model, and a 27B model with the full harness matched the 72B (0.884 vs.\ 0.873) at roughly 2.7$\times$ fewer parameters and a quarter of the CO$_2$. We read this through a distinction between world knowledge, which scales steeply with parameters, and language knowledge, which scales gently, and show that structured-output training makes a compact model harness-ready rather than merely small. On the raster, watermarked test PDFs the same system collapsed to 13.58\% joint (rank 149 of 163); a controlled re-rendering of the validation set reproduces the OCR half of the collapse while bounding what the simulation misses. Auditing the physical nature of evaluation inputs precedes architecture, and the leaderboard's bimodality is consistent with reading quality, not reasoning, having separated the field.

cs.CL↗

Rate-Distortion Adaptive Primitive Selection for Omnidirectional Gaussian Splatting

Learned image codecs (LICs) achieve high reconstruction quality, but their decoding speed is often insufficient for immersive virtual reality (VR). Gaussian splatting (GS) codecs render much faster, yet still lag in reconstruction quality and typically decide primitive allocation without considering the coding cost of each primitive. We introduce OIC-GS, an omnidirectional GS codec with a new hierarchical HEALPix primitive grid representation. Gaussian primitives are anchored at predefined spherical locations, eliminating explicit coordinate coding. Finer levels refine their coarser ancestors, naturally supporting coarse-to-fine reconstruction and layered transmission. The predefined grid also enables efficient viewport decoding by selecting only view-relevant primitives. We further introduce a lightweight entropy model for quantized primitives and optimize the codec under a spherical rate-distortion objective. Primitives with insufficient rate-distortion benefit are automatically removed when their quantized opacity becomes zero, allowing OIC-GS to adapt both primitive density and level of detail without a fixed primitive budget. A single bitstream supports full-sphere, viewport-dependent, and progressive decoding. The first viewport reaches final quality after decoding only 52% of the bitstream, and is then rendered at 1,270 FPS. On a 100-image omnidirectional benchmark, OIC-GS outperforms all evaluated GS codecs, reducing WS-PSNR BD-rate by 49.6% over GaussianImage++ and 68.6% over SGI, which uses a learned entropy model.

cs.CV↗

A Critical Value for the Viscosity Coefficient: Triggering Oscillatory Traveling Waves in the Pseudo-parabolic Fisher-KPP Equation

The pseudo-parabolic term $u_{xxt}$ serves as a canonical example of higher-order viscosity that provides effective regularization while capturing the refined physical mechanism. How it alters the dynamics of prototype equations remains a fundamental issue, which is not yet fully understood. The goal of this paper is to show that this term induces a sharp transition in traveling wave structure, via studying the pseudo-parabolic Fisher-KPP equation $$ u_t - τu_{xxt} = D u_{xx} + u(1-u). $$ We completely characterize how the parameter ratio $τ/D$ acts as a critical switch: it preserves the monotonic structure of classical traveling waves when $τ/D \leq 1$, yet it triggers a qualitative shift to oscillatory traveling waves when $τ/D > 1$. Numerical simulations support these findings. The threshold $τ/D=1$ thus captures the dual role of the pseudo-parabolic term: it acts as a structure-preserving viscosity when small, but when large it reflects dominant capillary hysteresis, generating the oscillations that explain the saturation overshoot, a phenomenon that contradicts classical diffusion models yet is widely observed in two-phase flow.

math.AP↗

The Shape of Speed: Impacts of Partition Geometry and Rank Density in Distributed Quantum Circuit Simulations

In distributed quantum circuit simulation, a poorly shaped partition can halve performance before computation begins. Evaluation on Fugaku across 764 validated configurations (twelve algorithms, thirteen torus partition geometries, and six rank densities for 39-qubit simulations on 1,024 nodes) shows that partition geometry dominates runtime. All twelve algorithms run 1.73-2.31x slower on flat partitions than on near-cubic ones despite identical data transfer, proving the slowdown stems from network delivery rather than communication volume. This penalty scales with the 3D torus partition aspect ratio (runtime $\propto a^{0.39}$, $r = 0.72$). Rank density is secondary, cutting runtime by 11% at 16 ranks per node only on compact geometries. Ultimately, requesting a near-cubic partition with 16 ranks per node roughly halves time-to-solution relative to flat partitions, which also consume 1.82x more energy. A simulator-free all-to-all microbenchmark confirms a similar geometry penalty for collective-dominated workloads.

cs.ET↗

Spexis: Speculative Lookahead Scheduling for LLM Inference

Spexis is a multi-GPU LLM inference framework that improves the efficiency of pipeline and tensor parallelism through speculative parallelism. Rather than using speculative decoding only to accelerate token generation, Spexis runs speculation in parallel with normal execution, introducing a new parallelism axis without increasing KV-cache memory usage. This improves memory efficiency and helps mitigate the bottlenecks of multi-GPU inference. Spexis further uses lookahead scheduling to predict speculation quality and future memory pressure, allowing it to reduce wasted speculation, KV-cache eviction, and recomputation. Built on top of vLLM, Spexis largely improves serving performance across a range of GPU configurations, achieving speedups of up to 34% over a baseline that uses the optimal combination of pipeline and tensor parallelism. Spexis's source code is publicly available at https://github.com/mlsys-seo/spexis.

cs.LG↗

From Static to Dynamic: On-Policy Distillation from Image to Video Diffusion Models

On-policy distillation (OPD) specializes pretrained video diffusion models through teacher supervision along the student's own generation trajectory. Although large video models are natural teachers, developing specialized video experts can require costly video data and training, while querying them incurs substantially higher latency than querying image experts. More readily available and cheaper to query, image experts offer a cost-effective alternative, particularly for largely temporal-agnostic capabilities such as aesthetics and OCR that admit frame-level supervision. However, heterogeneous image and video latent spaces prevent direct supervision of intermediate student states, while image experts lack cross-frame motion supervision, making temporal consistency vulnerable to frame-level improvements. In this paper, we propose MILD, a Motion-Preserving Image-to-Video Latent Distillation framework that transfers specialized image expertise while preserving pretrained video dynamics. MILD uses a learnable linear connector that aligns student latent states and predicted updates with those of image experts, enabling supervision transfer across heterogeneous latent spaces. We further constrain image-guided corrections around the pretrained student's predictions to preserve video dynamics and incorporate an optical-flow-based motion reward to improve motion quality and temporal consistency. Across specialized image experts and multiple video-student backbones, our method consistently outperforms video-teacher OPD baselines, with further studies demonstrating effective transfer across connector designs and heterogeneous architectures. These results establish image-to-video distillation as an effective route to improving video generation by drawing on the diverse and evolving capabilities of the image-generation ecosystem.

cs.CV↗

PersMem: Internalizing Personality into Dual-Pathway Memory for LLM Agents

The profile of a role-playing agent usually depends on the pre-defined personality in a system prompt, whereas its memory processing pipeline, including prioritisation of stored memories and subsequent retrieval, remains independent of this personality. This separation causes the agent's memory processing to be inconsistent with the pre-defined personality, and makes it difficult to validate whether agent behaviours follow this personality. In this paper, we propose Personality-Integrated Memory (PersMem), which integrates personality into the agent's memory processing pipeline, making it consistently personality-dependent. PersMem processes memory using four steps, where the personality is mapped to operation-specific parameters controlling: (i) affective appraisal annotating emotion states of the user input; (ii) retention of previously stored memories along with the current input; (iii) passive affect-driven memory retrieval exploring memories similar to user input in semantics and personality-guided emotions; and (iv) active goal-driven memory retrieval that refines and selects passively retrieved memories for the reply. Consequently, consistency with the pre-defined personality can be examined by inspecting memory-processing traces during human-agent interactions. We evaluate these personality-dependent differences in attachment and Big Five settings. PersMem exceeds the chance baseline for four-way attachment classification by 23.1 percentage points. In Big Five dialogue comparisons, PersMem achieves 67.5% accuracy, 6.7 percentage points above a baseline using uniformly sampled memories. On CoSER, PersMem achieves an average score of 66.13, with scores of 69.33 for Character Fidelity and 84.33 for Storyline Quality. Together, these results show that PersMem produces distinguishable personality-related memory-processing patterns.

cs.AI↗

Emergence, Not Bandwidth: Physical Coupling and the Limits of Learned Multi-Agent Communication

Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message should encode, what compression costs over a horizon, and when a learned protocol is unique enough for a teammate to read. We answer them for rate-limited Dec-POMDPs, then measure how far reinforcement learning falls short of the optimum. Our theorems fix what is achievable independently of any learner, so a gap between an engineered and a learned sender at the same bit budget is an optimization fact, not an information-theoretic one. We instantiate this on three MuJoCo arenas spanning zero, partial and rigid physical coupling, charging every condition exactly 2 bits per decision, and create the discriminating regime by closing a physical side channel within one arena, holding bodies, task and reward fixed. Communication value is governed by coupling: under rigid coupling through a shared object, no channel beats silence (+0.001 +/- 0.001, p = 0.982, n = 25), since proprioception already carries that information; without coupling, every condition solves the task; under partial coupling, the engineered 2-bit sender reaches an interquartile mean of 1.000 but the learned one reaches 0.482, indistinguishable from silence (p = 0.400, n = 25). With a shared alphabet, bandwidth cannot explain the gap. Warm-starting from an engineered receiver localizes the failure: the same channel reaches 0.857 versus 0.562 cold-started (p < 0.001), so it is neither representational nor one of maintenance; reinforcement learning fails to discover the protocol. Cross-play shows learned protocols are individually meaningful but mutually unintelligible: self-play 0.980 collapses to 0.144 across seeds, and our best constructed alignment leaves at least 77% of that gap. All headline results use 25 seeds per arena and seven published baselines at matched rate.

cs.LG↗