arXiv Science⌕ Search

arXiv subjects

:

Publications and source records attributed to :.

At least 163 records · Page 9Linked to original sources

Rethinking Reflection in Pre-Training

A language model's ability to reflect on its own reasoning provides a key advantage for solving complex problems. While most recent research has focused on how this ability develops during reinforcement learning, we show that it actually begins to emerge much earlier - during the model's pre-training. To study this, we introduce deliberate errors into chains-of-thought and test whether the model can still arrive at the correct answer by recognizing and correcting these mistakes. By tracking performance across different stages of pre-training, we observe that this self-correcting ability appears early and improves steadily over time. For instance, an OLMo2-7B model pre-trained on 4 trillion tokens displays self-correction on our six self-reflection tasks.

cs.CL↗

Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control

We introduce Cosmos-Transfer, a conditional world generation model that can generate world simulations based on multiple spatial control inputs of various modalities such as segmentation, depth, and edge. In the design, the spatial conditional scheme is adaptive and customizable. It allows weighting different conditional inputs differently at different spatial locations. This enables highly controllable world generation and finds use in various world-to-world transfer use cases, including Sim2Real. We conduct extensive evaluations to analyze the proposed model and demonstrate its applications for Physical AI, including robotics Sim2Real and autonomous vehicle data enrichment. We further demonstrate an inference scaling strategy to achieve real-time world generation with an NVIDIA GB200 NVL72 rack. To help accelerate research development in the field, we open-source our models and code at https://github.com/nvidia-cosmos/cosmos-transfer1.

cs.CV↗

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

General-purpose robots need a versatile body and an intelligent mind. Recent advancements in humanoid robots have shown great promise as a hardware platform for building generalist autonomy in the human world. A robot foundation model, trained on massive and diverse data sources, is essential for enabling the robots to reason about novel situations, robustly handle real-world variability, and rapidly learn new tasks. To this end, we introduce GR00T N1, an open foundation model for humanoid robots. GR00T N1 is a Vision-Language-Action (VLA) model with a dual-system architecture. The vision-language module (System 2) interprets the environment through vision and language instructions. The subsequent diffusion transformer module (System 1) generates fluid motor actions in real time. Both modules are tightly coupled and jointly trained end-to-end. We train GR00T N1 with a heterogeneous mixture of real-robot trajectories, human videos, and synthetically generated datasets. We show that our generalist robot model GR00T N1 outperforms the state-of-the-art imitation learning baselines on standard simulation benchmarks across multiple robot embodiments. Furthermore, we deploy our model on the Fourier GR-1 humanoid robot for language-conditioned bimanual manipulation tasks, achieving strong performance with high data efficiency.

cs.RO↗

Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM

Recent advancements in code large language models (LLMs) have demonstrated remarkable capabilities in code generation and understanding. It is still challenging to build a code LLM with comprehensive performance yet ultimate efficiency. Many attempts have been released in the open source community to break the trade-off between performance and efficiency, such as the Qwen Coder series and the DeepSeek Coder series. This paper introduces yet another attempt in this area, namely Ling-Coder-Lite. We leverage the efficient Mixture-of-Experts (MoE) architecture along with a set of high-quality data curation methods (especially those based on program analytics) to build an efficient yet powerful code LLM. Ling-Coder-Lite exhibits on-par performance on 12 representative coding benchmarks compared to state-of-the-art models of similar size, such as Qwen2.5-Coder-7B and DeepSeek-Coder-V2-Lite, while offering competitive latency and throughput. In practice, we achieve a 50\% reduction in deployment resources compared to the similar-sized dense model without performance loss. To facilitate further research and development in this area, we open-source our models as well as a substantial portion of high-quality data for the annealing and post-training stages. The models and data can be accessed at~\url{https://huggingface.co/inclusionAI/Ling-Coder-lite}.

cs.LG↗

Measurement of the Branching Fraction of $Λ_c^+ \to p K_S^0 π^0$ at Belle

We report a precise measurement of the ratio of branching fractions $\mathcal{B}(Λ_c^+\to p K_S^0 π^0)/\mathcal{B}(Λ_c^+\to p K^- π^+)$ using 980 fb$^{-1}$ of $e^+e^-$ data from the Belle experiment. We obtain a value of $\mathcal{B}(Λ_c^+\to p K_S^0 π^0)/\mathcal{B}(Λ_c^+\to p K^- π^+)=0.339\pm 0.002\pm 0.009$, where the first and second uncertainties are statistical and systematic, respectively. This Belle result is consistent with the previous measurement from the CLEO experiment but has a fivefold improvement in precision. By combining our result with the world average $\mathcal{B}(Λ_c^+\to p K^- π^+)$, we obtain the absolute branching fraction $\mathcal{B}(Λ_c^+\to p K_S^0 π^0)=(2.12\pm 0.01\pm 0.05 \pm 0.10)\%$, where the uncertainties are statistical, systematic, and the uncertainty in the absolute branching fraction scale $\mathcal{B}(Λ_c^+\to p K^- π^+)$, respectively. This measurement can shed light on hadronic decay mechanisms in charmed baryon decays.

hep-ex↗

Detection of RS Oph with LST-1 and modelling of its HE/VHE gamma-ray emission

The recurrent nova RS Ophiuchi (RS Oph) underwent a thermonuclear eruption in August 2021. In this event, RS Oph was detected by the High Energy Stereoscopic System (H.E.S.S.), the Major Atmospheric Gamma Imaging Cherenkov (MAGIC), and the first Large-Sized Telescope (LST-1) of the future Cherenkov Telescope Array Observatory (CTAO) at very-high gamma-ray energies above 100 GeV. This means that novae are a new class of very-high-energy (VHE) gamma-ray emitters. We report the analysis of the RS Oph observations with LST-1. We constrain the particle population that causes the observed emission in hadronic and leptonic scenarios. Additionally, we study the prospects of detecting further novae using LST-1 and the upcoming LST array of CTAO-North. We conducted target-of-opportunity observations with LST-1 from the first day of this nova event. The data were analysed in the framework of cta-lstchain and Gammapy, the official CTAO-LST reconstruction and analysis packages. One-zone hadronic and leptonic models were considered to model the gamma-ray emission of RS Oph using the spectral information from Fermi-LAT and LST-1, together with public data from the MAGIC and H.E.S.S. telescopes. RS Oph was detected at $6.6σ$ with LST-1 in the first 6.35 hours of observations following the eruption. The hadronic scenario is preferred over the leptonic scenario considering a proton energy spectrum with a power-law model with an exponential cutoff whose position increases from $(0.26\pm 0.08)$ TeV on day 1 up to $(1.6\pm 0.6)$ TeV on day 4 after the eruption. The deep sensitivity and low energy threshold of the LST-1/LST array will allow us to detect faint novae and increase their discovery rate.

astro-ph.HE↗

Measurements of higher-order cumulants of multiplicity and net-electric charge distributions in inelastic proton-proton interactions by NA61/SHINE

This paper presents the energy dependence of multiplicity and net-electric charge fluctuations in p+p interactions at beam momenta 20, 31, 40, 80, and 158 GeV/c. Results are corrected for the experimental biases and quantified with the use of cumulants and factorial cumulants. Cumulant ratios are an essential tool in the search for the critical point of strongly interacting matter in heavy ion collisions. Measurements performed in p+p interactions provide a vital baseline estimation in these studies. The measured signals are compared with the string hadronic models EPOS1.99 and FTFP-BERT.

hep-ex↗

Nonperturbative aspects of the electromagnetic pion form factor at high energies

The structure of hadronic form factors at high energies and their deviations from perturbative quantum chromodynamics provide insight on nonperturbative dynamics. Using an approach that is consistent with dispersion relations, we construct a model that simultaneously accounts for the pion wave function, gluonic exchanges, and quark Reggeization. In particular, we find that quark Reggeization can be investigated at high energies by studying scaling violation of the form factor.

hep-ph↗

Unveiling the Infrared Excess of SIPS J2045-6332: Evidence for a Young Stellar Object with Potential Low-Mass Companion

The Disk Detective project, a citizen science initiative, aims to identify circumstellar discs around stars by detecting objects with infrared (IR) excess using data from the Wide-field Infrared Survey Explorer (WISE). In this study, we investigate SIPS J2045-6332, a potential brown dwarf with significant IR excess in WISE and 2MASS bands, initially identified by project volunteers. Despite early indicators of a circumstellar disc, discrepancies between observed brightness and expected Spectral Energy Distribution (SED) models suggested unusual properties. To explore potential explanations, we created SED templates for spectral types M9 to L4 and compared them with SIPS J2045-6332's photometric data, revealing an excess brightness that points to either an unresolved low-mass companion or a young, inflated primary star. Further analysis of infrared spectral features and surface gravity indicators supports a youthful classification, estimating the object's age at 26-200 million years. Observations also suggest the presence of a mid L-type companion at a projected distance of 6.7 AU. This study highlights SIPS J2045-6332 as an intriguing system with unique IR characteristics and recommends follow-up observations with high-resolution telescopes to confirm the companion hypothesis and further characterize the system.

astro-ph.SR↗

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

We introduce Phi-4-Mini and Phi-4-Multimodal, compact yet highly capable language and multimodal models. Phi-4-Mini is a 3.8-billion-parameter language model trained on high-quality web and synthetic data, significantly outperforming recent open-source models of similar size and matching the performance of models twice its size on math and coding tasks requiring complex reasoning. This achievement is driven by a carefully curated synthetic data recipe emphasizing high-quality math and coding datasets. Compared to its predecessor, Phi-3.5-Mini, Phi-4-Mini features an expanded vocabulary size of 200K tokens to better support multilingual applications, as well as group query attention for more efficient long-sequence generation. Phi-4-Multimodal is a multimodal model that integrates text, vision, and speech/audio input modalities into a single model. Its novel modality extension approach leverages LoRA adapters and modality-specific routers to allow multiple inference modes combining various modalities without interference. For example, it now ranks first in the OpenASR leaderboard to date, although the LoRA component of the speech/audio modality has just 460 million parameters. Phi-4-Multimodal supports scenarios involving (vision + language), (vision + speech), and (speech/audio) inputs, outperforming larger vision-language and speech-language models on a wide range of tasks. Additionally, we experiment to further train Phi-4-Mini to enhance its reasoning capabilities. Despite its compact 3.8-billion-parameter size, this experimental version achieves reasoning performance on par with or surpassing significantly larger models, including DeepSeek-R1-Distill-Qwen-7B and DeepSeek-R1-Distill-Llama-8B.

cs.CL↗

Measurement of the inclusive branching fractions for $B_s^0$ decays into $D$ mesons via hadronic tagging

We report measurements of the absolute branching fractions $\mathcal{B}(B_s^0 \to D_s^{\pm} X)$, $\mathcal{B}(B_s^0 \to D^0/\bar{D}^0 X)$, and $\mathcal{B}(B_s^0 \to D^{\pm} X)$, where the latter is measured for the first time. The results are based on a 121.4\,fb$^{-1}$ data sample collected at the $Υ(10860)$ resonance by the Belle detector at the KEKB asymmetric-energy $e^+ e^-$ collider. We reconstruct one $B_s^0$ meson in $e^+e^- \to Υ(10860) \to B_s^{*} \bar{B}_s^{*}$ events and measure yields of $D_s^+$, $D^0$, and $D^+$ mesons in the rest of the event. We obtain $\mathcal{B}(B_s^0 \to D_s^{\pm} X) = (68.6 \pm 7.2 \pm 4.0)\%$, $\mathcal{B}(B_s^0 \to D^0/\bar{D}^0 X) = (21.5 \pm 6.1 \pm 1.8)\%$, and $\mathcal{B}(B_s^0 \to D^{\pm} X) = (12.6 \pm 4.6 \pm 1.3)\%$, where the first uncertainty is statistical and the second is systematic. Averaging with previous Belle measurements gives $\mathcal{B}(B_s^0 \to D_s^{\pm} X) = (63.4 \pm 4.5 \pm 2.2)\%$ and $\mathcal{B}(B_s^0 \to D^0/\bar{D}^0 X) = (23.9 \pm 4.1 \pm 1.8)\%$. For the $B_s^0$ production fraction at the $Υ(10860)$, we find $f_s = (21.4^{+1.5}_{-1.7})\%$.

hep-ex↗

Competitive Programming with Large Reasoning Models

We show that reinforcement learning applied to large language models (LLMs) significantly boosts performance on complex coding and reasoning tasks. Additionally, we compare two general-purpose reasoning models - OpenAI o1 and an early checkpoint of o3 - with a domain-specific system, o1-ioi, which uses hand-engineered inference strategies designed for competing in the 2024 International Olympiad in Informatics (IOI). We competed live at IOI 2024 with o1-ioi and, using hand-crafted test-time strategies, placed in the 49th percentile. Under relaxed competition constraints, o1-ioi achieved a gold medal. However, when evaluating later models such as o3, we find that o3 achieves gold without hand-crafted domain-specific strategies or relaxed constraints. Our findings show that although specialized pipelines such as o1-ioi yield solid improvements, the scaled-up, general-purpose o3 model surpasses those results without relying on hand-crafted inference heuristics. Notably, o3 achieves a gold medal at the 2024 IOI and obtains a Codeforces rating on par with elite human competitors. Overall, these results indicate that scaling general-purpose reinforcement learning, rather than relying on domain-specific techniques, offers a robust path toward state-of-the-art AI in reasoning domains, such as competitive programming.

cs.LG↗

Search for $C\!P$ violation in $D^+_{(s)}\to{}K_{S}^{0}K^{-}π^{+}π^{+}$ decays using triple and quadruple products

We perform the first search for $C\!P$ violation in ${D_{(s)}^{+}\to{}K_{S}^{0}K^{-}π^{+}π^{+}}$ decays. We use a combined data set from the Belle and Belle II experiments, which study $e^+e^-$ collisions at center-of-mass energies at or near the $Υ(4S)$ resonance. We use 980 fb$^{-1}$ of data from Belle and 428 fb$^{-1}$ of data from Belle~II. We measure six $C\!P$-violating asymmetries that are based on triple products and quadruple products of the momenta of final-state particles, and also the particles' helicity angles. We obtain a precision at the level of 0.5% for $D^+\to{}K_{S}^{0}K^{-}π^{+}π^{+}$ decays, and better than 0.3% for $D^+_{s}\to{}K_{S}^{0}K^{-}π^{+}π^{+}$ decays. No evidence of $C\!P$ violation is found. Our results for the triple-product asymmetries are the most precise to date for singly-Cabibbo-suppressed $D^+$ decays. Our results for the other asymmetries are the first such measurements performed for charm decays.

hep-ex↗

The Breeze 2 Herd of Models: Traditional Chinese LLMs Based on Llama with Vision-Aware and Function-Calling Capabilities

Llama-Breeze2 (hereinafter referred to as Breeze2) is a suite of advanced multi-modal language models, available in 3B and 8B parameter configurations, specifically designed to enhance Traditional Chinese language representation. Building upon the Llama 3.2 model family, we continue the pre-training of Breeze2 on an extensive corpus to enhance the linguistic and cultural heritage of Traditional Chinese. In addition to language modeling capabilities, we significantly augment the models with function calling and vision understanding capabilities. At the time of this publication, as far as we are aware, absent reasoning-inducing prompts, Breeze2 are the strongest performing models in Traditional Chinese function calling and image understanding in its size class. The effectiveness of Breeze2 is benchmarked across various tasks, including Taiwan general knowledge, instruction-following, long context, function calling, and vision understanding. We are publicly releasing all Breeze2 models under the Llama 3.2 Community License. We also showcase the capabilities of the model running on mobile platform with a mobile application which we also open source.

cs.CL↗

Open-Source Retrieval Augmented Generation Framework for Retrieving Accurate Medication Insights from Formularies for African Healthcare Workers

Accessing accurate medication insights is vital for enhancing patient safety, minimizing errors, and supporting clinical decision-making. However, healthcare professionals in Africa often rely on manual and time-consuming processes to retrieve drug information, exacerbated by limited access to pharmacists due to brain drain and healthcare disparities. This paper presents "Drug Insights," an open-source Retrieval-Augmented Generation (RAG) chatbot designed to streamline medication lookup for healthcare workers in Africa. By leveraging a corpus of Nigerian pharmaceutical data and advanced AI technologies, including Pinecone databases and GPT models, the system delivers accurate, context-specific responses with minimal hallucination. The chatbot integrates prompt engineering and S-BERT evaluation to optimize retrieval and response generation. Preliminary tests, including pharmacist feedback, affirm the tool's potential to improve drug information access while highlighting areas for enhancement, such as UI/UX refinement and extended corpus integration.

cs.IR↗

Yi: Open Foundation Models by 01.AI

We introduce the Yi model family, a series of language and multimodal models that demonstrate strong multi-dimensional capabilities. The Yi model family is based on 6B and 34B pretrained language models, then we extend them to chat models, 200K long context models, depth-upscaled models, and vision-language models. Our base models achieve strong performance on a wide range of benchmarks like MMLU, and our finetuned chat models deliver strong human preference rate on major evaluation platforms like AlpacaEval and Chatbot Arena. Building upon our scalable super-computing infrastructure and the classical transformer architecture, we attribute the performance of Yi models primarily to its data quality resulting from our data-engineering efforts. For pretraining, we construct 3.1 trillion tokens of English and Chinese corpora using a cascaded data deduplication and quality filtering pipeline. For finetuning, we polish a small scale (less than 10K) instruction dataset over multiple iterations such that every single instance has been verified directly by our machine learning engineers. For vision-language, we combine the chat language model with a vision transformer encoder and train the model to align visual representations to the semantic space of the language model. We further extend the context length to 200K through lightweight continual pretraining and demonstrate strong needle-in-a-haystack retrieval performance. We show that extending the depth of the pretrained checkpoint through continual pretraining further improves performance. We believe that given our current results, continuing to scale up model parameters using thoroughly optimized data will lead to even stronger frontier models.

cs.CL↗

Initial measurement of reactor antineutrino oscillation at SNO+

The SNO+ collaboration reports its first spectral analysis of long-baseline reactor antineutrino oscillation using 114 tonne-years of data. Fitting the neutrino oscillation probability to the observed energy spectrum yields constraints on the neutrino mass-squared difference $Δm^2_{21}$. In the ranges allowed by previous measurements, the best-fit $Δm^2_{21}$ is (8.85$^{+1.10}_{-1.33}$) $\times$ 10$^{-5}$ eV$^2$. This measurement is continuing in the next phases of SNO+ and is expected to surpass the present global precision on $Δm^2_{21}$ with about three years of data.

hep-ex↗

Qwen2.5 Technical Report

In this report, we introduce Qwen2.5, a comprehensive series of large language models (LLMs) designed to meet diverse needs. Compared to previous iterations, Qwen 2.5 has been significantly improved during both the pre-training and post-training stages. In terms of pre-training, we have scaled the high-quality pre-training datasets from the previous 7 trillion tokens to 18 trillion tokens. This provides a strong foundation for common sense, expert knowledge, and reasoning capabilities. In terms of post-training, we implement intricate supervised finetuning with over 1 million samples, as well as multistage reinforcement learning. Post-training techniques enhance human preference, and notably improve long text generation, structural data analysis, and instruction following. To handle diverse and varied use cases effectively, we present Qwen2.5 LLM series in rich sizes. Open-weight offerings include base and instruction-tuned models, with quantized versions available. In addition, for hosted solutions, the proprietary models currently include two mixture-of-experts (MoE) variants: Qwen2.5-Turbo and Qwen2.5-Plus, both available from Alibaba Cloud Model Studio. Qwen2.5 has demonstrated top-tier performance on a wide range of benchmarks evaluating language understanding, reasoning, mathematics, coding, human preference alignment, etc. Specifically, the open-weight flagship Qwen2.5-72B-Instruct outperforms a number of open and proprietary models and demonstrates competitive performance to the state-of-the-art open-weight model, Llama-3-405B-Instruct, which is around 5 times larger. Qwen2.5-Turbo and Qwen2.5-Plus offer superior cost-effectiveness while performing competitively against GPT-4o-mini and GPT-4o respectively. Additionally, as the foundation, Qwen2.5 models have been instrumental in training specialized models such as Qwen2.5-Math, Qwen2.5-Coder, QwQ, and multimodal models.

cs.CL↗