arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,495 records · Page 83Linked to original sources

Permutation-Robust Decision Modeling with Candidate-Independent Block-Causal Attention

Decision models often score a variable-sized set of candidate actions encoded in a single sequence. This setting is increasingly relevant for System 1 components inside generative systems, where candidates may be proposed or ordered differently across runs. Standard causal cross-encoding is expressive, but it can make a candidate's score depend on serialization order rather than on the underlying decision problem. We introduce candidate-independent block-causal attention, which preserves causal computation within the shared context and each candidate while blocking cross-candidate information flow and resetting candidate positions. We compare this architecture with standard causal attention and complementary invariant baselines across Gemma 3 1B, Qwen3 1.7B, and Qwen3 4B backbones. Candidate-independent attention consistently reduces permutation sensitivity while retaining competitive decision quality; ablations indicate that candidate isolation is the primary source of the effect, with position resetting completing the intended symmetry. A larger Qwen3-4B study further examines the behavior of the proposed architecture with substantially more training data. Code is available at the \href{https://github.com/guyAmit/ci-decision-models}{\textcolor{blue}{project repository}}, and the \href{https://huggingface.co/Guy-Amit/qwen3-4b-ci-decision-4096-poc}{\textcolor{blue}{Qwen3-4B model artifact}} is available on Hugging Face.

cs.LG↗

Open-Source Live-Reconfigurable Multi-Mode Wearable Ultrasound

Wearable ultrasound enables continuous deep-tissue monitoring, and a single programmable probe can operate in multiple complementary modes, such as structural A-mode and Doppler flow measurement. However, each operating mode requires dedicated measurement parameters and peripheral states, with no single configuration serving all modes on resource-constrained devices. Time multiplexing of operating modes introduces reconfiguration latency that lowers the effective mode repetition rate. To address this limitation, we present an open-source, transition-aware control stack for low-latency, in-session reconfiguration of the 32-channel TinyProbe wearable platform. Operating modes are described as hardware configurations, and host-side shadow registers track the peripheral states, enabling transition-specific register updates. Transition sequences are executed either by the host (over Wi-Fi 6) or by a firmware loop on the probe MCU. We validate the stack on a pulsatile-flow phantom by interleaving blocks of 25 to 100 pulsed-wave Doppler shots at 1.43 kHz PRF with single 16-channel A-mode acquisitions, changing channel configurations at every transition. Compared to full reconfiguration, the overhead per transition decreases from 30.2 ms to 11.6 ms (host-scheduled) and 3.1 ms (MCU-scheduled). For 75-shot Doppler blocks, the multi-mode repetition rate reaches 16.0 Hz (MCU-scheduled), 90.1% of the theoretical maximum of 17.7 Hz. Concurrent reconstruction of a Doppler spectrogram and a lumen-diameter trace demonstrates the functionality of time-multiplexed flow and structural monitoring.

eess.SP↗

U-Sonic: An Open-Source 8-Channel Ultrasound Transmit IP in a 130 nm RISC-V SoC

Miniaturized ultrasound (US) probes require programmable and synchronized transmit (TX) excitation across multiple elements, while existing compact platforms often rely on limited microcontroller (MCU) pulse generators or closed-source fixed-function pulser devices. We present U-Sonic, an open-source digital US TX peripheral integrated into a 32-bit RISC-V system-on-chip (SoC). The implemented SoC integrates 8 pulser cores, while the parameterized architecture supports up to 16 channels. Each core generates single- or dual-tone bursts with programmable period, duty cycle, pulse count, polarity, and idle level, together with optional inverted stop pulses for active damping. A shared memory-mapped Open Bus Interface (OBI) enables synchronous start and stop of arbitrary channel subsets and supports composite bipolar, gated, and three-level excitation schemes. Functional correctness was verified in Verilator against a Python golden model over 4379 checked cycles across directed and randomized configurations, and confirmed on a Terasic DE10-Lite field-programmable gate array (FPGA). The design was synthesized and placed-and-routed in IHP 130 nm. The post-layout area in kilo gate equivalents (kGE), scales as 1.65 kGE plus 1.66 kGE per channel. The 8-channel instance occupies 14.9 kGE, corresponding to approximately 14.3% of the 104 kGE SoC. The register-transfer level (RTL), register descriptions, verification collateral, and software support are released as open source.

cs.AR↗

Where Chaos Pauses: Sliding-Window Frequency Ratios Reveal Transient Resonances in Relativistic Orbits

A long-time Poincaré section flattens a chaotic orbit into one static blur, erasing the order in which the orbit visits its own local structures. We introduce the sliding-window frequency ratio (SWFR), built from nothing but radial and polar turning events, to recover that order along individual relativistic orbits. Each window spans a fixed number of radial cycles and simply counts the polar events inside it; because event times are kept, any candidate interval can be sliced out and replotted as its own Poincaré section. No spectrum, no reference center, no basis functions. Integrable Kerr benchmarks recover the known frequency ratios, and for regular motion with uniformly bounded count deviations the counting-error bound falls off as $O(W^{-1})$. The same counts then pay off in chaos. In magnetized Kerr spacetime a single charged orbit dwells near $3/5$ and later near $4/7$, and exactly those intervals open into fivefold and sevenfold section structures; a second orbit dwells near $1/2$ with two lobes. In Schwarzschild--Melvin spacetime a photon holds $4/5$ across a fivefold pattern. In every case the section covers exactly the interval SWFR selected: integer counts and phase-space geometry agree that a globally chaotic orbit is paying a temporary visit to a resonance.

gr-qc↗

Hob-VL: A Benchmark for Visually Grounded Boolean Reasoning

Reliable visual reasoning requires composing multiple visual observations and returning consistent answers to logically equivalent questions. We introduce Hob-VL, a benchmark for visually grounded Boolean reasoning. Hob-VL comprises two tasks: (1) evaluating whether a Boolean rule holds in an image, and (2) identifying the (unique) object satisfying a Boolean description. Hob-VL contains 6,000 human-verified balanced Yes/No questions, each defined by a Boolean combination of ten visual statements, across 1,000 generated scenes and 46 diverse labeled photographs, along with 1,000 object-identification questions over the same photographs. Our question families are deliberately constructed to challenge reasoning through misleading local cues and nested logical operations, and include symbolic and structured natural-language presentations. Across eight model configurations with thinking disabled or minimized, Boolean accuracy ranges from 48.52% to 50.57%, while the identification accuracy reaches at most 43.0%. A thinking-enabled GLM configuration achieves uneven gains while retaining substantial errors and inconsistencies. Hob-VL exposes these failures through executable reference answers and matched evaluations.

cs.CV↗

From Simulation to Fabrication: Realizing Silicon Crystalline Undulators with Silicon Nitride Stressor Layer Patterning

A crystalline undulator is a crystal exhibiting periodic deformations that cause channeled particles to oscillate, generating coherent electromagnetic waves. In this study, stressor layer patterning has been employed to induce the desired deformations on a silicon substrate. Finite element method simulations were performed to optimize the geometric parameters of the undulator. The primary focus was on achieving sinusoidal deformation with sub-millimeter period and amplitude exceeding 1~nm, utilizing the silicon (110) plane for its superior channeling properties. Key parameters, such as substrate thickness and undulator period, were meticulously refined to ensure uniform deformation while minimizing the impact of higher harmonics, which could degrade performance. Based on the simulation results, a crystalline undulator, 160 $μ$m thick and consisting of 10 periods with a period length of 334 $μ$m, has been successfully fabricated. The final device demonstrates structural integrity and a uniform deformation extending up to 20 $μ$m from the surface. Specifically designed for use with 5-30 GeV particle beams, the undulator is capable of generating gamma photons in the 5-15 MeV range. This work effectively integrates advanced simulation techniques with precise fabrication methods, demonstrating the feasibility of crystalline undulators as high-performance devices for generating gamma radiation.

physics.acc-ph↗

Structural stability of systems and cycle covers in random graphs

Structural system theory studies which network topologies can sustain a prescribed system property such as controllability or stability. When the topology is itself random, the relevant question becomes probabilistic: how likely is a graph drawn from a stochastic model to sustain the property? Such probabilities measure the abundance and robustness of the property across topologies, and indicate whether systems requiring it can be reliably deployed in uncertain environments. We address this question for asymptotic stability of linear systems in the directed graphon setting. We consider two graph-theoretic properties. The first is $\mathcal N$, and it requires that for every $k\leq n$ some $k$-vertex induced subdigraph of $D$ admits a cycle cover, and the second is $\mathcal S$, which requires that these subdigraphs can be chosen so that their node sets form a nested sequence $V_1\subset\cdots\subset V_n=V(D)$ starting from a single vertex with a loop. We have shown that $\mathcal N$ is necessary and $\mathcal S$ is sufficient for structural stability. We sample $D$ from a directed step-graphon $W$. Our main results give necessary and sufficient conditions for $\Pr(\mathcal N)\to 1$ and $\Pr(\mathcal S)\to 1$ as $n\to\infty$. In more detail, to a step-graphon $W$ with skeleton digraph $S$ on $q$ nodes and concentration vector $x^*$ we associate a cycle polytope $\vec{\mathcal X}(S)\subseteqΔ_q$. The conditions are then formulated in terms of the position of $x^*$ within $\vec{\mathcal X}(S)$, the dimension of the polytope, the loop density of $W$ and, for $\mathcal S$, an ordering condition on the cycles of the skeleton. Together these results identify, for directed step-graphons, the regime in which a sampled topology is overwhelmingly likely or unlikely to sustain stable dynamics.

math.OC↗

Solvable Algebraic Generalized Ricci Solitons in low dimensions

We prove new structural results for algebraic generalized Ricci solitons on solvable Lie algebras and we discuss the existence of examples in low dimensions. We establish that all non-trivial solvable algebraic generalized Ricci solitons must be expanding and give necessary conditions for the existence of such solitons on one-dimensional extensions of abelian Lie algebras. Finally, we classify algebraic generalized Ricci solitons on three-dimensional Lie algebras, and algebraic generalized Ricci solitons on four-dimensional solvable unimodular Lie algebras, up to isometry and scaling.

math.DG↗

From Balanced Growth to Innovation Cycles:A Solow-Type Semi-Endogenous Growth Model

We develop a growth model with an R\&D sector in which the elasticity of innovation with respect to existing knowledge can be negative. We investigate the existence and uniqueness of a balanced growth path (BGP) and derive closed-form growth factors, showing that productivity and output growth are semi-endogenous and population-driven. We also establish the global stability under a general production function. Under testable parameter restrictions, the economy converges to the BGP. However, when the knowledge elasticity is sufficiently low, a Jacobian-based condition implies instability, and an N-period innovation cycle can emerge. Comparative statics are conducted to explore the impact of several factors, including research efficiency, input elasticity, and population growth. Numerical simulations map out transitions and mark the points at which stability is lost. They make clear when innovation frictions endanger balanced growth and provide practical guidance for R\&D and demographic policy.

math.DS↗

Architectural Degradation: How to Measure and to Remediate

Context. Architectural degradation undermines software maintainability, evolvability, and quality. However, existing research remains fragmented across measurement approaches, metrics, tools, and remediation strategies, limiting our understanding of how these elements relate across the degradation lifecycle. Aim. We consolidate the state of the art on architectural degradation by examining how researchers measure it, which metrics and tools support its assessment, and how existing approaches address remediation. Method. We conducted a Multivocal Literature Review of 284 peer-reviewed and grey-literature studies. We supported screening, data extraction, and classification with a locally executed LLM-assisted pipeline combining Retrieval-Augmented Generation, multi-model validation, and human adjudication. We then analyzed the resulting taxonomies and their cross-dimensional relationships. Results and Conclusions. We identified 277 measurement approaches, 357 metrics, 238 tools, and 395 remediation approaches. Research strongly concentrates on static and structural analysis, structural metrics, and detection-oriented tools. In contrast, remediation spans heterogeneous code-level, architectural, and organizational interventions and shows substantially less consolidation. Overall, the field has developed a mature diagnostic apparatus but has made less progress in connecting degradation detection with effective remediation. Our results provide a structured view of the available techniques and identify the diagnosis-remediation gap as a key direction for future research.

cs.SE↗

ReCo: Response-Consistent Locomotion with Policy-Aware MPC for Legged Manipulation

Continuous legged manipulation requires accurate end-effector tracking while the base keeps walking. Combining reinforcement learning (RL) with model predictive control (MPC) suits this task: the learned policy provides robust locomotion, while MPC coordinates the base and arm to compensate for tracking errors. However, MPC can compensate only for base motion that it can predict, and a learned policy's command response varies with gait phase, contact, and payload. We present ReCo, a framework that couples response-consistent locomotion with policy-aware MPC for legged manipulation. Response shaping trains the policy to respond to commands consistently and repeatably across randomized dynamics. An identified closed-loop response model then lets MPC jointly plan locomotion commands and arm motion. On the simulation benchmark, ReCo reduces position and orientation root-mean-square error (RMSE) by 28.7% and 27.4% relative to the best baseline for each metric. Real-world experiments demonstrate onboard continuous legged manipulation with coordinated base and arm motion.

cs.RO↗

From the LZ Event to Grand Unification: Heavy Higgsinos and Proton Decay

The recently reported LUX-ZEPLIN (LZ) event has renewed interest in inelastic Higgsino dark matter. While a weak-scale Higgsino provides a simple interpretation of the event, such a scenario is in tension with the IceCube bound from dark-matter capture and annihilation in the Sun. This tension can be avoided for a much heavier Higgsino, with a mass of order $10^5$ GeV and a neutral-state mass splitting of $\simeq 400$ keV, which in turn points to gaugino masses around $10^7$ GeV. Such a heavy supersymmetric spectrum is often thought to be unfavorable for supersymmetric grand unified theories (GUTs), since it tends to worsen gauge coupling unification and lower the unification scale, potentially leading to excessively rapid proton decay. In this work, we revisit this expectation in minimal supersymmetric SU(5), taking into account the GUT-scale threshold corrections. We show that the resulting unification conditions allow the SU(5) gauge bosons to remain sufficiently heavy when the adjoint-Higgs self-coupling is small, thereby suppressing dimension-six proton decay. At the same time, the color-triplet Higgs mass can remain near the conventional GUT scale, $M_{H_C}\sim 10^{16}$ GeV, particularly when the gluino is significantly heavier than the wino. Remarkably, in this region the dimension-five decay mode $p\to K^+\barν$ can have a lifetime within the reach of next-generation proton-decay searches, including Hyper-Kamiokande, JUNO, and DUNE. We also discuss supersymmetry-breaking scenarios that can give rise to the required supersymmetric mass spectrum.

hep-ph↗

Oneira: From Open-Ended Generation to Open-World Interaction in Video World Models

Generative video world models can now synthesize open-ended environments that agents can navigate and interact with in simple ways. Yet open-ended generation does not imply full interaction: as a generated world expands, newly created content through navigation should expand what the agent can act upon, and as the agent changes the world, those changes should become persistent parts of the environment rather than transient visual effects. We characterize these two requirements as Open-World Interactivity, where newly generated or encountered entities are incorporated into the actionable world, and Persistent State, where interaction outcomes are committed to the world state and continue to influence subsequent observations and interactions. We present Oneira, an interactive video world model that closes the loop between generation and interaction through an explicit, extensible world state managed by a coding agent. Given the current observation and an action or high-level goal, the agent reads the world state, grounds the relevant entities, plans the interaction, and writes its outcome back into a world state table. When exploration reveals new objects, the agent incorporates them from generated observations, allowing the interaction space to expand with the generated world. Meanwhile, previously induced state changes are carried across video segments, making the consequences of interaction persistent parts of subsequent world evolution. The updated world state is rendered along the camera action trajectory into a coarse conditioning video, from which a video generator fills in the appearance, motion, and interaction details not represented in the state. Experiments show that Oneira enables direct and consistent interaction with newly generated objects, while preserving the effects of prior interactions over long horizons. Project page: https://madaoer.github.io/projects/oneira

cs.CV↗

Vanishing orders of Dirichlet solutions to the Schrödinger equation in dimensions three and higher

Let $B_3=B(0,3)\subset\mathbb{R}^3$. For every sufficiently large integer $k$, we construct a nonzero real function $u_k\in C^2(\overline{B_3})$ and a real potential $V_k\in L^\infty(B_3)$ such that $Δu_k=V_ku_k$, $u_k|_{\partial B_3}=0$, $\mathrm{ord}_0u_k=k$, and $\|V_k\|_\infty\le Ck^{3/2}$. Consequently, for every sufficiently large $N$, there is such a Dirichlet solution with potential norm smaller than $N$ and vanishing order at least $cN^{2/3}$. The construction extends directly to higher dimensions and to spheres. Together with previous results, this example indicates that $C(1+\|V\|_\infty^{2/3})$ is the sharp bound for the vanishing order in the real-valued case. It also indicates that the bound conjectured independently by Kukavica \cite{Kukavica1998} and Kenig \cite{Kenig2006} is not attainable in general.

math.AP↗

Can LLMs Reliably Annotate Bioassay Metadata to Improve Data Readiness?

The emergence of foundation models for molecular property prediction requires a high degree of AI data readiness, including reliable metadata annotation. However, both public repositories and industrial screening databases suffer from missing, inconsistent, or conflated assay annotations. In this work, we quantify the extent of missing annotations in PubChem for the BioAssay Ontology (BAO) assay format and physical detection method fields and investigate whether open-source and proprietary large language models (LLMs) can reliably predict and audit metadata annotations directly from the assay text. In our assessment, we found that the annotation coverage across PubChem's $\sim$2 million bioassays is critically sparse, 36\% lacking an assay format, 89\% a BioAssay type, and >99.9\% any BAO-mapped assay format or detection technology term. This motivates the need for automated test-metadata curation. Using evaluation sets derived from PubChem and ChEMBL, we assess the agreement of seven open-source and proprietary LLMs with existing silver labels. Recall is at least 0.96 for biochemical and cell-based assay formats, with a similar pattern for detection technology, although disagreements increase on under-represented classes. Manual inspection shows that many of these disagreements trace back to inconsistencies between silver sources rather than to LLM error. Moreover, in a qualitative study with a senior industrial curator, LLM-generated evidence prompted the expert to revise some of their own labels, showing LLMs can flag potentially mislabeled assays. Across the study, performance differences between proprietary and open-source models were small. Together, these results suggest LLMs can support the large-scale annotation and auditing of assay metadata, though per-class reliability estimates and targeted human review remain necessary before such labels enter downstream ML pipelines.

cs.CL↗

A negative Schur coefficient for products of two chains

We prove that the product of chains $\mathbf{m}\times\mathbf{n}$ is not Schur positive whenever $n\ge4$ and $m\ge3n-1$. Writing $m=n+k$, we exhibit a negative coefficient indexed by $(2n+k-1,2n+k-3,\ldots,k+5,k,k,4)$ and evaluate it explicitly as $(n-2)!$ times a polynomial of degree three in $k$. Our argument refines the chain-partition method of Li, Qiu, Yang, and Zhang: rank capacity forces $n-2$ long chains, and the remaining three chains are counted by separating a middle interval with fixed row coordinates from two bounded boundary regions. Together with their theorem and the known cases of widths two and three, this shows that these products are not Schur positive for $n\ge3$, $m\ge n+5$, and for $n=2$, $m\ge8$.

math.CO↗

Agents Are Systems, Not Models: Rethinking Agentic Evaluation

Agent evaluations increasingly go beyond a single success rate, reporting metrics such as cost, consistency, and robustness. Yet they typically treat the agent itself as fixed. In practice, an agent is a configurable system: users decide what to tell it, how long to let it run, and which model to use, and each of these choices can change how well and how consistently it performs. We study these choices on a new benchmark of four scientific tasks, where a coding agent must find and correctly operate a published specialist model. We investigate five parts of the agent's configuration: task information, reasoning, self-verification, time budget, and backbone model. We find substantial run-to-run variability, with approximately 54% of the outcome variance coming from repeating the same configuration rather than changing it. Across configurations, the information provided to the agent has the largest effect, exceeding both time budget and model size, while also reducing cost and improving calibration. Configuration choices also interact: additional time helps only when the agent has sufficient information or a capable enough model to use it. Finally, a trajectory-based taxonomy of agent behavior reveals that prompting an agent to verify its answer has little effect on its verification behavior, whereas providing a dedicated verification tool changes that behavior substantially. These results suggest that agents should be evaluated as configurable systems themselves, and that some desired behaviors are more effectively implemented in the system than requested through prompting. We release the benchmark and more than 18,000 agent trajectories.

cs.AI↗

Exposing the Cost of Deep Learning Audio Development

The environmental impact of deep learning has attracted increasing attention over the past decade. Existing studies mainly focus on the energy and carbon emissions of model training and inference, while the whole development phase is often overlooked. Yet, architecture prototyping and intensive experiments are conducted during this stage, which is highly energy-demanding. In this article, we propose a methodology to estimate these costs, based on activity logs from the Grid5000 shared computing platform used by the LORIA laboratory. As a case-study, we focus on audio projects developed in the Multispeech research team. We evaluate the overall energy cost of four projects, and we compare them to those of training the reported models. Our results show that the energy required for the development phase is 3 to 256 times greater than that required to train the best-performing model alone. These results advocate for a more systematic reporting of energy consumption across the entire life cycle of deep learning-based audio projects.

cs.LG↗