arXiv Science⌕ Search

arXiv subjects

:

Publications and source records attributed to :.

At least 37 records · Page 2Linked to original sources

MOSS Transcribe Diarize Technical Report

Speaker-Attributed, Time-Stamped Transcription (SATS) aims to transcribe what is said and to precisely determine the timing of each speaker, which is particularly valuable for meeting transcription. Existing SATS systems rarely adopt an end-to-end formulation and are further constrained by limited context windows, weak long-range speaker memory, and the inability to output timestamps. To address these limitations, we present MOSS Transcribe Diarize, a unified multimodal large language model that jointly performs Speaker-Attributed, Time-Stamped Transcription in an end-to-end paradigm. Trained on extensive real wild data and equipped with a 128k context window for up to 90-minute inputs, MOSS Transcribe Diarize scales well and generalizes robustly. Across comprehensive evaluations, it outperforms state-of-the-art commercial systems on multiple public and in-house benchmarks.

cs.SD↗

RhinoVLA Technical Report

Vision-Language-Action (VLA) models have shown strong potential for robotic manipulation, but real-time deployment on edge hardware remains challenging. In this work, we identify VLM visual and context tokens as a major source of deployment latency: for GEMM-dominated projection operators, computation grows linearly with the number of input tokens when model dimensions are fixed. Motivated by this observation, we propose RhinoVLA, a deployment-oriented VLA model co-designed with the Huixi R1 edge SoC. RhinoVLA adopts a token-efficient Qwen3-VL backbone and a continuous Action Expert, reducing the VLM-side token and computation burden while preserving pretrained multimodal capability. To support cross-robot learning, RhinoVLA further introduces a unified interface that combines View Registry, 72D physical state-action slot space, and robotinstance LoRA, allowing heterogeneous robot observations and action schemas to be aligned under a shared policy. On the deployment side, RhinoVLA is optimized through hardware-aware compilation, mixed-precision execution, and parallel visual encoding. Experiments show that RhinoVLA achieves downstream performance comparable to π0.5 at a similar parameter scale, while reaching 11.69 Hz end-to-end inference on Huixi R1, meeting the 10 Hz real-time closedloop control target. The project will be open-sourced at https://github.com/HuixiAI/RhinoVLA.

cs.RO↗

Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation

3D simulation platforms are critical for autonomous driving because they enable end-to-end policy evaluation, thereby reducing development costs and improving safety. In recent years, neural simulation has become predominant, with methods such as NuRec playing a central role; however, these methods remain relatively slow and typically require per-scene tuning. In this work, we present Instant NuRec, a feed-forward neural reconstruction model that turns a short multi-view driving log into a fully simulatable 3D Gaussian Splatting (3DGS) world in a single forward pass. The model accepts multi-view input from a calibrated camera rig and emits a layered output consisting of static and dynamic 3DGS layers, a sky cubemap, and per-camera ISP corrections, while providing native support for non-pinhole camera models via 3DGUT. It reconstructs a 10-20-second multi-camera scene in roughly 1.5 seconds and achieves a PSNR on the Waymo Open Dataset that is 2.01 dB above the strongest evaluated baseline. Instant NuRec is deeply integrated into NuRec and is compatible with AlpaSim for closed-loop simulation.

cs.GR↗

Vilya-1: An all-atom foundation model for macrocycle structure prediction and design

Macrocyclic peptides are an increasingly important therapeutic modality, but existing computational methods for modeling their structures and properties are limited in scope and do not generalize well across the synthetically accessible chemical space. In this work, we introduce Vilya-1, a deep learning model that addresses two central challenges in macrocycle design: sampling biologically relevant conformations across arbitrary chemistries and predicting key developability properties such as membrane permeability. Vilya-1 operates on a uniform all-atom representation and is trained on heterogeneous structural datasets spanning diverse topologies and chemical classes. Across a broad set of macrocycles composed of canonical and non-canonical residues, Vilya-1 substantially improves geometric accuracy relative to physics-based methods, co-folding networks, and deep-learning conformer generators, while maintaining broad chemical coverage that extends to small molecules. Vilya-1 also supports generative applications, enabling the design of novel macrocycles with tailored chemical, structural, and property profiles. Together, these capabilities establish Vilya-1 as a foundation model for accelerating the development of next-generation macrocycle therapeutics.

cs.LG↗

Sub-Torque-Balance Upper Limits on Continuous Gravitational Waves from Scorpius X-1

We present the results of a search for continuous gravitational waves from the low-mass X-ray binary Scorpius X-1 using LIGO data from the first part of the fourth LIGO-Virgo-KAGRA observing run. By applying the resampling version of the cross-correlation pipeline to search for signal frequencies $f_0$ between $25$ and $200\un{Hz}$ (corresponding to neutron star spin frequencies of $12.5$ to $100\un{Hz}$ for GW due to triaxiality, or $\sim15-20$ to $\sim120-150\un{Hz}$ for GW due to $r$-modes), we set upper limits below the standard torque balance level, independent of neutron star spin inclination, for $50\un{Hz}\lesssim f_0\lesssim200\un{Hz}$. While uncertainties in the modelling of torque and equation of state limit the strength of our inference, our results nonetheless argue against torque balance in this spin range for a neutron star described by a hadronic equation of state. The most sensitive upper limits on the gravitational wave amplitude $h_0$, at the upper end of the frequency band searched, approach $5\times10^{-26}$ marginalized over inclination angle and $2\times10^{-26}$ assuming the most favorable inclination. The marginalized upper limits correspond to a sensitivity depth of $70-75\un{Hz}^{-1/2}$, improving sensitivity considerably over previous searches. Expressed as constraints on the triaxial deformation of the neutron star, the limits correspond to an ellipticity of $3\times10^{-5}$ if the GW frequency $f_0$ is $75\un{Hz}$ and $3\times10^{-6}$ if $f_0=200\un{Hz}$, approaching deformations which could be supported by ordinary nuclear matter. Outliers from the search were ruled out as potential signals by a combination of hierarchical followup and analysis of additional data from later in the observing run.

astro-ph.HE↗

Searches for charged-lepton-flavor violation in $χ_{bJ}(1P)$ decays

We report the first searches for charged-lepton-flavor violation in decays of $χ_{bJ}(1P)$ ($J=0, 1,$ and $2$) to a pair of charged leptons using 158 million $Υ(2S)$ decays collected with the Belle detector in $e^+e^-$ collisions at the KEKB collider. No significant signal is observed, and we set upper limits on the branching fractions for $χ_{bJ}(1P)$ decays to $e^\pmμ^\mp$ at the level of $10^{-6}$ and to $e^\pmτ^\mp$ or $μ^\pmτ^\mp$ at the level of $10^{-5}$. Limits on $χ_{b0}(1P)$ decays are translated into bounds on the corresponding Wilson coefficients of scalar operators that mediate charged-lepton-flavor violation.

hep-ex↗

First measurement of the masses of the $Υ_1(1D)$ and $Υ_3(1D)$ states and the energy dependence of the cross sections for $e^+e^-\toΥ_J(1D)η$ and $e^+e^-\toΥ_J(1D)π^+π^-$

We study the processes $e^+e^-\toΥ_J(1D)η$ and $e^+e^-\toΥ_J(1D)π^+π^-$ at center-of-mass energies $\sqrt{s}$=(10.73 -- 11.02) GeV using a $142.5\,\mathrm{fb}^{-1}$ data sample, including 122~fb$^{-1}$ near the $Υ$(10860) peak ($\sqrt{s}$ = 10.866 GeV), collected with the Belle detector at the KEKB asymmetric-energy $e^+e^-$ collider. From the peak sample, the products of Born cross section times branching fraction are obtained for $σ_{\rm Born}(e^+e^-\toΥ_J(1D)η)$ or $σ_{\rm Born}(e^+e^-\toΥ_J(1D)π^+π^-)$ and ${\cal B}(Υ_J(1D)\toχ_{b1}γ)$ or ${\cal B}(Υ_J(1D)\toχ_{b2}γ)$ for each $Υ_J(1D)$ state. The corresponding branching fractions for $Υ(10860)$ decays are also obtained. The significances of the $Υ_1(1D)$, $Υ_2(1D)$, and $Υ_3(1D)$ signals are 4.8$σ$, ${>}10σ$, and 3.0$σ$, respectively, including systematic uncertainties. The mass for $Υ_2(1D)$ is measured to be $(10167.0\pm 1.0\pm 0.2)$ MeV/$c^2$, where the first and second uncertainties are statistical and systematic. The mass splittings $Δm_{12}=m(Υ_2(1D))-m(Υ_1(1D))$ and $Δm_{23}=m(Υ_3(1D))-m(Υ_2(1D))$ are $(11.8\pm1.5\pm0.4)$ MeV/$c^2$ and $(7.6\pm2.4\pm0.6)$ MeV/$c^2$, respectively.~We determine the energy dependence of the cross sections for $e^+e^-\toΥ_J(1D)η$ and $e^+e^-\toΥ_J(1D)π^+π^-$ for the $Υ_1(1D)$, $Υ_2(1D)$, and $Υ_3(1D)$ states, combined.

hep-ex↗

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation

Inference efficiency is typically pursued by shrinking the model: distillation, pruning, quantization, and sparse routing each lower per-token cost while treating token count as fixed. But output length has been inflating, and it is precisely the component the standard toolkit leaves untouched. Here, we argue that brevity is the missing inference-efficiency lever, and that pretraining data curation is a practical way to pull it: a model trained on concise, correct data learns to answer in fewer tokens; i.e. it has a lower Cost-of-Pass. We apply our VLM curation pipeline to the MAmmoTH-VL single-image subset, and compare models trained on our curated data, the standard MAmmoTH-VL data, and external open-weight frontier VLMs. On a controlled 20-evaluation set and 14 VLMs at 1B-4B activated parameters, we hold output length fixed with a per-model regression, separating brevity from quality, and price models in FLOPs per correct answer. Curation buys a 35x Cost-of-Pass advantage over the most verbose 4B comparator (Qwen3.5-4B) within $\sim$1 pp of accuracy (0.41 vs 14.58 TFLOPs per correct answer; 0.691 vs 0.704 mean accuracy). Curation also buys a +17.55-percentage-point matched-length accuracy gain over the uncurated baseline that grows with model scale (from +16.7 pp at 1B to +21.2 pp at 4B). This brevity improvement concedes no quality: generic verbosity buys no accuracy at any capability or scale, and the window where reasoning-structured verbosity still earns its tokens shrinks from 4 of 8 capability groups at 2B to 1 of 8 at 4B. Per example, the concise model even reaches correct answers the verbose reasoning model misses, marking reasoning as a distinct curation target rather than something brevity gives up. Inference efficiency in this regime is a tokens-per-correct problem, and brevity is the lever that targets it directly.

cs.LG↗

Study of $B^+ \to μ^+ ν_μ$ decays at Belle and Belle II

We report a measurement of the branching fraction for the leptonic decay $B^+\toμ^+ν_μ$. This work presents the first $B^+\toμ^+ν_μ$ result using Belle~II data, an updated Belle measurement that supersedes the previous result, and their combination, which yields the most precise search to date. The analysis is based on $1076\,\mathrm{fb}^{-1}$ of $e^+e^-$ collision data collected at a center-of-mass energy of $10.58\,\mathrm{GeV}$ with the Belle and Belle~II detectors at the KEKB and SuperKEKB colliders, respectively. We measure $\mathcal{B}(B^+\toμ^+ν_μ)=(4.4\pm1.9\pm 1.0)\times10^{-7}$, where the first uncertainty is statistical and the second systematic. The observed significance relative to the background-only hypothesis is 2.4 standard deviations. We set a 90\% confidence level upper limit of $\mathcal{B}(B^+\toμ^+ν_μ)<6.7\times10^{-7}$ using a frequentist approach and a 90\% credibility level upper limit of $\mathcal{B}(B^+\toμ^+ν_μ)<7.2\times 10^{-7}$ using a Bayesian approach. These are the most stringent limits to date. The result is interpreted as an exclusion region in the parameter space of type~II and type~III two-Higgs-doublet models. We search for stable sterile neutrinos with masses $m_N\in[0,1.5]\,\mathrm{GeV}$. No signal is observed, and the resulting exclusion on the squared mixing parameter $|U_{μN}|^2$ provides improvement over previous limits. We report a measurement of the partial branching fraction of semileptonic $B\to X_u\ellν_\ell$ decays with $p_μ^B>2.2\,\mathrm{GeV}$, obtaining $Δ\mathcal{B}(B\to X_u\ellν_\ell)=(2.72\pm0.05\pm0.29)\times10^{-4}$. We present a model-dependent study of weak annihilation decays using the muon momentum spectrum. We observe a signal of 2.4 standard deviations above the background-only hypothesis in regions where the distribution resembles that of $B\to X_u\ellν_\ell$ decays.

hep-ex↗

Cosmos 3: Omnimodal World Models for Physical AI

We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-transformers architecture. By supporting highly flexible input-output configurations, Cosmos 3 seamlessly unifies critical modalities for Physical AI -- effectively subsuming vision-language models, video generators, world simulators, and world-action models into a single framework. Our evaluation demonstrates that Cosmos 3 establishes a new state-of-the-art across a diverse suite of understanding and generation tasks, demonstrating omnimodal world models as scalable, general-purpose backbones for embodied agents. Our post-trained Cosmos 3 models were ranked as the best open-source Text-to-Image and Image-to-Video models by Artificial Analysis, and the best policy model by RoboArena at the time the technical report was written. To accelerate open research and deployment in Physical AI, we make our code, model checkpoints, curated synthetic datasets, and evaluation benchmark available under the Linux Foundation's OpenMDW-1.1 License at https://github.com/nvidia/cosmos and https://huggingface.co/collections/nvidia/cosmos3. The project website is available at https://research.nvidia.com/labs/cosmos-lab/cosmos3.

cs.CV↗

First evidence of $X(3872)\toπ^0χ_{c0}(1P)$ and search for $X(3915)\toπ^0χ_{c1}(1P)$

We search for the pionic transitions $X(3872)\toπ^0χ_{cJ}(1P)$ $(J = 0,~1,~2)$ and $X(3915)\toπ^0χ_{c1}$ in $B^+\to π^0χ_{cJ}K^+$ decays using the Belle and Belle~II data samples collected at the $Υ(4S)$ resonance, corresponding to integrated luminosities of $711~\mathrm{fb}^{-1}$ and $492~\mathrm{fb}^{-1}$, respectively. We report the first evidence for the decay $X(3872)\toπ^0χ_{c0}$ with a significance of $3.4σ$, including systematic uncertainties. We measure the product of branching fractions ${\cal B}(B^+\to X(3872)K^+)\times{\cal B}(X(3872)\toπ^0χ_{c0})=(20.0\pm6.8\pm2.3)\times10^{-6}$ and the branching fraction ratio ${\cal B}(X(3872)\toπ^0χ_{c0})/{\cal B}(X(3872)\toπ^+π^-J/ψ)=2.3\pm0.8\pm0.4$, where the first and second uncertainties are statistical and systematic, respectively. The upper limits at 90\% credibility on the products of branching fractions for the $π^0χ_{c1}$ and $π^0χ_{c2}$ modes are $7.5\times10^{-6}$ and $15.3\times10^{-6}$, respectively. The corresponding upper limits on the branching fraction ratios relative to the $π^+π^-J/ψ$ decay are $0.9$ and $1.8$. The measured branching fractions for $X(3872)\toπ^0χ_{cJ}$ are consistent with several theoretical predictions based on the hadronic molecular interpretation of the $X(3872)$. No significant signal is seen for the $X(3915)\toπ^0χ_{c1}$ decay, and we set the 90\% credibility upper limit of ${\cal B}(B^+\to X(3915)K^+)\times{\cal B}(X(3915)\toπ^0χ_{c1})<6.6\times10^{-6}$, while the decays for $J=0$ and 2 are forbidden by parity conservation.

hep-ex↗

Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

We introduce Nemotron 3 Ultra, a 550 billion total and 55 billion active parameter Mixture-of-Experts Hybrid Mamba-Attention language model. We pre-trained Nemotron 3 Ultra on 20 trillion text tokens, then extended the context length to 1M tokens, and post-trained using Supervised Fine Tuning (SFT), Reinforcement Learning (RL), and Multi-teacher On-Policy Distillation (MOPD). Nemotron 3 Ultra is our most capable model yet, employing multiple key technologies - LatentMoE, Multi Token Prediction (MTP), NVFP4 pre-training, multi-environment RLVR, MOPD, and reasoning budget control. Nemotron 3 Ultra achieves up to ~6x higher inference throughput as compared to state-of-the-art publicly available LLMs while attaining on-par accuracy. The state-of-the-art accuracy, high inference throughput, and 1M token context length make Nemotron 3 Ultra ideal for long-running autonomous agentic tasks. We open-source the base, post-trained, and quantized checkpoints, along with the training data and recipe on HuggingFace.

cs.CL↗

Observation of the decays $B^{+} \to Σ_{c}(2455)^{++} \barΞ_{c}^{\prime-}$ and $B^{0} \to Σ_{c}(2455)^{0} \barΞ_{c}^{\prime0}$

We report the first observation of the decays $B^{+} \to Σ_{c}(2455)^{++} \barΞ_{c}^{\prime-}$ and $B^{0} \to Σ_{c}(2455)^{0} \barΞ_{c}^{\prime0}$, with significances of $6.4\,σ$ and $5.3\, σ$, respectively, including systematic uncertainties. This analysis is based on data samples containing $771.6 \times 10^{6}$ $Υ(4S)$ decays collected with the Belle detector at the KEKB collider and $520.6 \times 10^{6}$ $Υ(4S)$ decays collected with the Belle~II detector at the SuperKEKB collider. The branching fractions are measured to be $\mathcal{B}(B^+ \to Σ_c(2455)^{++} \barΞ_c^{\prime -}) = (1.68 \pm 0.31 \pm 0.12^{+1.49}_{-0.54}) \times 10^{-3}$ and $\mathcal{B}(B^0 \to Σ_c(2455)^{0} \barΞ_c^{\prime 0}) = (1.28 \pm 0.32 \pm 0.10^{+0.30}_{-0.21}) \times 10^{-3}$, where the first and second uncertainties are statistical and systematic, respectively, and the third arises from the uncertainties in the absolute branching fractions of $\barΞ_{c}^{-}$ and $\barΞ_{c}^{0}$ decays. This result represents the first observation of $B$-meson decays into a pair of charmed baryon-antibaryon states belonging to the same $SU(3)$ flavor sextet.

hep-ex↗

National Mapping and Testing of Astronomical Sites in Ethiopia (NMTASE)

This work aims to choose potential astronomical sites that can be candidates for a new astronomical optical observatory in Ethiopia, in addition to the Entoto Observatory and Lalibela sites. For our primary investigation, the six basic criteria, namely the altitude of the mountains, artificial light pollution, cloud coverage, humidity, wind speed, and wind direction, were taken into account. Consequently, using the multi-criteria statistical Decision analysis (MCDSA) techniques, 21 high-potential places are selected and presented for further investigation out of 367 mountains. Among these 21 selected places, three sites, Bauhit, Meseraia, and T'at'a are the most suitable places for optical astronomy in Ethiopia. Those selected mountains are mapped and presented to study the future of the astronomical seeing effect. This study may contribute to the protection of those potential astronomical sites and their dark skies and the development of astrotourism for the sustainable development of modern astronomy in Ethiopia and in the East African region.

astro-ph.IM↗

Open-H-Embodiment: A Large-Scale Dataset for Enabling Foundation Models in Medical Robotics

Autonomous medical robots hold promise to improve patient outcomes, reduce provider workload, democratize access to care, and enable superhuman precision. However, autonomous medical robotics has been limited by a fundamental data problem: existing medical robotic datasets are small, single-embodiment, and rarely shared openly, restricting the development of foundation models that the field needs to advance. We introduce Open-H-Embodiment, the largest open dataset of medical robotic video with synchronized kinematics to date, spanning more than 50 institutions and multiple robotic platforms including the CMR Versius, Intuitive Surgical's da Vinci, da Vinci Research Kit (dVRK), Rob Surgical BiTrack, Virtual Incision's MIRA, Moon Surgical Maestro, and a variety of custom systems, spanning surgical manipulation, robotic ultrasound, and endoscopy procedures. We demonstrate the research enabled by this dataset through two foundation models. GR00T-H is the first open foundation vision-language-action model for medical robotics, which is the only evaluated model to achieve full end-to-end task completion on a structured suturing benchmark (25% of trials vs. 0% for all others) and achieves 64% average success across a 29-step ex vivo suturing sequence. We also train Cosmos-H-Surgical-Simulator, the first action-conditioned world model to enable multi-embodiment surgical simulation from a single checkpoint, spanning nine robotic platforms and supporting in silico policy evaluation and synthetic data generation for the medical domain. These results suggest that open, large-scale medical robot data collection can serve as critical infrastructure for the research community, enabling advances in robot learning, world modeling, and beyond.

cs.RO↗

Measurement of time-dependent $CP$ violation parameters in $B^{0} \to K_{S}^{0} π^{0} γ$ decays at Belle and Belle II

We perform a measurement of time-dependent $CP$ violation parameters in $B^{0} \to K_{S}^{0} π^{0} γ$ decays using a dataset of approximately $772 \times 10^6$ and $521 \times 10^6$ $Υ(4S)$ decays collected by the Belle and Belle II experiments, respectively. The measured parameters for the combined dataset in the $K^{*0}(892)$ dominated region ($M_{K_{S}^{0} π^{0}} \in [0.8,1.0] \mathrm{GeV}/c^2$) are $S = 0.09 \pm 0.16 \pm 0.02$ and $C = -0.09 \pm 0.08 \pm 0.04$. For the non-$K^{*0}(892)$ region ($M_{K_{S}^{0} π^{0}} \in [1.0,1.8] \mathrm{GeV}/c^2$), the corresponding values are $S = -0.32 \pm 0.33 \pm 0.09$ and $C = -0.07 \pm 0.17 \pm 0.08$. The first quoted uncertainties are statistical, while the second ones are systematic. These results are consistent with Standard Model predictions and more precise than previous measurements.

hep-ex↗

On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters

Parameter-efficient fine-tuning (PEFT) is usually treated as a cheaper alternative to full fine-tuning. We study a broader role: small trainable adapters as persistent local state on top of strong shared foundation models. In this framing, the base model provides shared competence while adapters carry instance-specific behavior such as preferences, skills, tool habits, and memory-like updates. We organize the problem around three scaling axes: Scale Up, where stronger shared priors make small local updates more useful; Scale Down, where we study how small adapters can be while remaining reliable; and Scale Out, where many persistent adapted instances coexist. MinT provides one infrastructure example for managing adapter identity, revision, provenance, evaluation, and serving residency. Together, the results suggest that PEFT can be a compact substrate for persistent personal models rather than only a budget substitute for full fine-tuning.

cs.LG↗

Directed Nano-antennas for Laser Fusion

Why do we use nano-antennas for fusion? In three sentences: The present laser induced fusion plans use extreme mechanical shock compression to get one hotspot and then ignition. Still fusion burning spreads slower than expansion, and mechanical instabilities may also develop. With nano-antennas in radiation dominated systems, simultaneous ignition can be achieved in the whole target volume and there is no time left for mechanical instabilities. Ignition is achieved with protons accelerated in the direction of the nanoantennas that are orthogonal to the direction of laser irradiation. Present laser fusion methods are based on extreme and slow mechanical compression with an ablator surface on the fuel target pellet to increase compression and eliminate penetration of laser electromagnetic energy into the target. This arises from a mistaken assumption, [1] that the detonation normal 4-vector should have vanishing time-like component, and this assumption eliminates the possibility to rapid or even simultaneous, radiation dominated detonations, (which are well known in the burning (or hadronization) of Quark Gluon Plasma).

physics.plasm-ph↗