arXiv Science⌕ Search

arXiv subjects

:

Publications and source records attributed to :.

At least 145 records · Page 8Linked to original sources

Cosmos World Foundation Model Platform for Physical AI

Physical AI needs to be trained digitally first. It needs a digital twin of itself, the policy model, and a digital twin of the world, the world model. In this paper, we present the Cosmos World Foundation Model Platform to help developers build customized world models for their Physical AI setups. We position a world foundation model as a general-purpose world model that can be fine-tuned into customized world models for downstream applications. Our platform covers a video curation pipeline, pre-trained world foundation models, examples of post-training of pre-trained world foundation models, and video tokenizers. To help Physical AI builders solve the most critical problems of our society, we make Cosmos open-source and our models open-weight with permissive licenses available via https://github.com/nvidia-cosmos/cosmos-predict1.

cs.CV↗

Essential-Web v1.0: 24T tokens of organized web data

Data plays the most prominent role in how language models acquire skills and knowledge. The lack of massive, well-organized pre-training datasets results in costly and inaccessible data pipelines. We present Essential-Web v1.0, a 24-trillion-token dataset in which every document is annotated with a twelve-category taxonomy covering topic, format, content complexity, and quality. Taxonomy labels are produced by EAI-Distill-0.5b, a fine-tuned 0.5b-parameter model that achieves an annotator agreement within 3% of Qwen2.5-32B-Instruct. With nothing more than SQL-style filters, we obtain competitive web-curated datasets in math (-8.0% relative to SOTA), web code (+14.3%), STEM (+24.5%) and medical (+8.6%). Essential-Web v1.0 is available on HuggingFace: https://huggingface.co/datasets/EssentialAI/essential-web-v1.0

cs.CL↗

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

We introduce MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model. MiniMax-M1 is powered by a hybrid Mixture-of-Experts (MoE) architecture combined with a lightning attention mechanism. The model is developed based on our previous MiniMax-Text-01 model, which contains a total of 456 billion parameters with 45.9 billion parameters activated per token. The M1 model natively supports a context length of 1 million tokens, 8x the context size of DeepSeek R1. Furthermore, the lightning attention mechanism in MiniMax-M1 enables efficient scaling of test-time compute. These properties make M1 particularly suitable for complex tasks that require processing long inputs and thinking extensively. MiniMax-M1 is trained using large-scale reinforcement learning (RL) on diverse problems including sandbox-based, real-world software engineering environments. In addition to M1's inherent efficiency advantage for RL training, we propose CISPO, a novel RL algorithm to further enhance RL efficiency. CISPO clips importance sampling weights rather than token updates, outperforming other competitive RL variants. Combining hybrid-attention and CISPO enables MiniMax-M1's full RL training on 512 H800 GPUs to complete in only three weeks, with a rental cost of just $534,700. We release two versions of MiniMax-M1 models with 40K and 80K thinking budgets respectively, where the 40K model represents an intermediate phase of the 80K training. Experiments on standard benchmarks show that our models are comparable or superior to strong open-weight models such as the original DeepSeek-R1 and Qwen3-235B, with particular strengths in complex software engineering, tool utilization, and long-context tasks. We publicly release MiniMax-M1 at https://github.com/MiniMax-AI/MiniMax-M1.

cs.CL↗

Magistral

We introduce Magistral, Mistral's first reasoning model and our own scalable reinforcement learning (RL) pipeline. Instead of relying on existing implementations and RL traces distilled from prior models, we follow a ground up approach, relying solely on our own models and infrastructure. Notably, we demonstrate a stack that enabled us to explore the limits of pure RL training of LLMs, present a simple method to force the reasoning language of the model, and show that RL on text data alone maintains most of the initial checkpoint's capabilities. We find that RL on text maintains or improves multimodal understanding, instruction following and function calling. We present Magistral Medium, trained for reasoning on top of Mistral Medium 3 with RL alone, and we open-source Magistral Small (Apache 2.0) which further includes cold-start data from Magistral Medium.

cs.CL↗

The Critical Importance of Software for HEP

Particle physics has an ambitious and broad global experimental programme for the coming decades. Large investments in building new facilities are already underway or under consideration. Scaling the present processing power and data storage needs by the foreseen increase in data rates in the next decade for HL-LHC is not sustainable within the current budgets. As a result, a more efficient usage of computing resources is required in order to realise the physics potential of future experiments. Software and computing are an integral part of experimental design, trigger and data acquisition, simulation, reconstruction, and analysis, as well as related theoretical predictions. A significant investment in computing and software is therefore critical. Advances in software and computing, including artificial intelligence (AI) and machine learning (ML), will be key for solving these challenges. Making better use of new processing hardware such as graphical processing units (GPUs) or ARM chips is a growing trend. This forms part of a computing solution that makes efficient use of facilities and contributes to the reduction of the environmental footprint of HEP computing. The HEP community already provided a roadmap for software and computing for the last EPPSU, and this paper updates that, with a focus on the most resource critical parts of our data processing chain.

hep-ex↗

Bench to the Future: A Pastcasting Benchmark for Forecasting Agents

Forecasting is a challenging task that offers a clearly measurable way to study AI systems. Forecasting requires a large amount of research on the internet, and evaluations require time for events to happen, making the development of forecasting benchmarks challenging. To date, no forecasting benchmark provides a realistic, hermetic, and repeatable environment for LLM forecasters. We introduce Bench To the Future (BTF), a "pastcasting" benchmark with hundreds of high-quality questions for which the resolution is already known. Each question is accompanied by a large offline corpus of tens of thousands of relevant web pages, enabling a way to elicit realistic "forecasts" on past events from LLMs. Results suggest that our pastcasting environment can produce results comparable to those based on forecasts using the internet on at-the-time unresolved questions. We show results benchmarking agent and chain-of-thought forecasting approaches using several LLMs, including the recently-released Claude 4 models, and demonstrate BTF's ability to track steady forecasting capability progress over time. We intend this to be a living benchmark, with new questions added continually to account for increasing training data cutoff dates. We invite researchers to contact us at hello@futuresearch.ai to utilize our benchmark or tooling for their own research.

cs.CL↗

MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining

We present MiMo-7B, a large language model born for reasoning tasks, with optimization across both pre-training and post-training stages. During pre-training, we enhance the data preprocessing pipeline and employ a three-stage data mixing strategy to strengthen the base model's reasoning potential. MiMo-7B-Base is pre-trained on 25 trillion tokens, with additional Multi-Token Prediction objective for enhanced performance and accelerated inference speed. During post-training, we curate a dataset of 130K verifiable mathematics and programming problems for reinforcement learning, integrating a test-difficulty-driven code-reward scheme to alleviate sparse-reward issues and employing strategic data resampling to stabilize training. Extensive evaluations show that MiMo-7B-Base possesses exceptional reasoning potential, outperforming even much larger 32B models. The final RL-tuned model, MiMo-7B-RL, achieves superior performance on mathematics, code and general reasoning tasks, surpassing the performance of OpenAI o1-mini. The model checkpoints are available at https://github.com/xiaomimimo/MiMo.

cs.CL↗

Measurement of the Positive Muon Anomalous Magnetic Moment to 127 ppb

A new measurement of the magnetic anomaly $a_μ$ of the positive muon is presented based on data taken from 2020 to 2023 by the Muon $g-2$ Experiment at Fermi National Accelerator Laboratory (FNAL). This dataset contains over 2.5 times the total statistics of our previous results. From the ratio of the precession frequencies for muons and protons in our storage ring magnetic field, together with precisely known ratios of fundamental constants, we determine $a_μ = 116\,592\,0710(162) \times 10^{-12}$ (139 ppb) for the new datasets, and $a_μ = 116\,592\,0705(148) \times 10^{-12}$ (127 ppb) when combined with our previous results. The new experimental world average, dominated by the measurements at FNAL, is $a_μ(\text{exp}) =116\,592\,0715(145) \times 10^{-12}$ (124 ppb). The measurements at FNAL have improved the precision on the world average by over a factor of four.

hep-ex↗

Detection of the Geminga pulsar at energies down to 20 GeV with the LST-1 of CTAO

Geminga is the third gamma-ray pulsar firmly detected by imaging atmospheric Cherenkov telescopes (IACTs) after the Crab and the Vela pulsars. Most of its emission is expected at tens of GeV, and, out of the planned telescopes of the upcoming Cherenkov Telescope Array Observatory (CTAO), the Large-Sized Telescopes (LSTs) are the only ones with optimised sensitivity at these energies. We aim to characterise the gamma-ray pulse shape and spectrum of Geminga as observed by the first LST (hereafter LST-1) of the CTAO-North. Furthermore, this study confirms the great performance and the improved energy threshold of the telescope, as low as 10 GeV for pulsar analysis, with respect to current-generation Cherenkov telescopes. We analysed 60 hours of good-quality data taken by the LST-1 at zenith angles below 50$^\circ$. Additionally, a new Fermi-LAT analysis of 16.6 years of data was carried out to extend the spectral analysis down to 100 MeV. Lastly, a detailed study of the systematic effects was performed. We report the detection of Geminga in the energy range between 20 and 65 GeV. Of the two peaks of the phaseogram, the second one, P2, is detected with a significance of 12.2$σ$, while the first (P1) reaches a significance level of 2.6$σ$. The best-fit model for the spectrum of P2 was found to be a power law with $Γ= (4.5 \pm 0.4_{stat})^{+0.2_{sys}}_{-0.6_{sys}}$, compatible with the previous results obtained by the MAGIC. No evidence of curvature is found in the LST-1 energy range. The joint fit with Fermi data confirms a preference for a sub-exponential cut-off over a pure exponential, even though both models fail to reproduce the data above several tens of GeV. The overall results presented in this paper prove that the LST-1 is an excellent telescope for the observation of pulsars, and improved sensitivity is expected to be achieved with the full CTAO-North.

astro-ph.HE↗

Practical Efficiency of Muon for Pretraining

We demonstrate that Muon, the simplest instantiation of a second-order optimizer, explicitly expands the Pareto frontier over AdamW on the compute-time tradeoff. We find that Muon is more effective than AdamW in retaining data efficiency at large batch sizes, far beyond the so-called critical batch size, while remaining computationally efficient, thus enabling more economical training. We study the combination of Muon and the maximal update parameterization (muP) for efficient hyperparameter transfer and present a simple telescoping algorithm that accounts for all sources of error in muP while introducing only a modest overhead in resources. We validate our findings through extensive experiments with model sizes up to four billion parameters and ablations on the data distribution and architecture.

cs.LG↗

Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning

Physical AI systems need to perceive, understand, and perform complex actions in the physical world. In this paper, we present the Cosmos-Reason1 models that can understand the physical world and generate appropriate embodied decisions (e.g., next step action) in natural language through long chain-of-thought reasoning processes. We begin by defining key capabilities for Physical AI reasoning, with a focus on physical common sense and embodied reasoning. To represent physical common sense, we use a hierarchical ontology that captures fundamental knowledge about space, time, and physics. For embodied reasoning, we rely on a two-dimensional ontology that generalizes across different physical embodiments. Building on these capabilities, we develop two multimodal large language models, Cosmos-Reason1-7B and Cosmos-Reason1-56B. We curate data and train our models in two stages: Physical AI supervised fine-tuning (SFT) and Physical AI reinforcement learning (RL). To evaluate our models, we build comprehensive benchmarks for physical common sense and embodied reasoning according to our ontologies. Evaluation results show that Physical AI SFT and RL bring significant improvements. To facilitate the development of Physical AI, we make our code and pre-trained models available under the NVIDIA Open Model License at https://github.com/nvidia-cosmos/cosmos-reason1.

cs.AI↗

FullStack Bench: Evaluating LLMs as Full Stack Coders

As the capabilities of code large language models (LLMs) continue to expand, their applications across diverse code intelligence domains are rapidly increasing. However, most existing datasets only evaluate limited application domains. To address this gap, we have developed a comprehensive code evaluation dataset FullStack Bench focusing on full-stack programming, which encompasses a wide range of application domains (e.g., basic programming, data analysis, software engineering, mathematics, and machine learning). Besides, to assess multilingual programming capabilities, in FullStack Bench, we design real-world instructions and corresponding unit test cases from 16 widely-used programming languages to reflect real-world usage scenarios rather than simple translations. Moreover, we also release an effective code sandbox execution tool (i.e., SandboxFusion) supporting various programming languages and packages to evaluate the performance of our FullStack Bench efficiently. Comprehensive experimental results on our FullStack Bench demonstrate the necessity and effectiveness of our FullStack Bench and SandboxFusion.

cs.AI↗

Model-independent measurement of $D^0$-$\overline{D}{}^0$ mixing parameters in $D^0\rightarrow K^0_{S}π^+π^-$ decays at Belle and Belle II

We perform a model-independent measurement of the $D^0$-$\overline{D}{}^0$ mixing parameters using samples of $e^+e^-$-collision data collected by the Belle and Belle II experiments that have integrated luminosities of $951\ \text{fb}^{-1}$ and $408\ \text{fb}^{-1}$, respectively. Approximately $2.05\times10^6$ neutral $D$ mesons are reconstructed in the $D^0\rightarrow K^0_{S}π^+π^-$ channel, with the neutral $D$ flavor tagged by the charge of the pion in the $D^{*+}\rightarrow D^0π^+$ decay. Assuming charge-parity symmetry, the mixing parameters are measured to be $ x = (4.0\pm1.7\pm0.4)\times 10^{-3} $ and $ y = (2.9\pm1.4\pm0.3)\times 10^{-3}$, where the first uncertainties are statistical and the second systematic. The results are consistent with previous determinations.

hep-ex↗

Deep Research Bench: Evaluating AI Web Research Agents

Amongst the most common use cases of modern AI is LLM chat with web search enabled. However, no direct evaluations of the quality of web research agents exist that control for the continually-changing web. We introduce Deep Research Bench, consisting of 89 multi-step web research task instances of varying difficulty across 8 diverse task categories, with the answers carefully worked out by skilled humans. We provide a "RetroSearch" environment with a large frozen set of scraped web pages, and demonstrate that offline "RetroSearch" agents perform comparably to "live web" agents, enabling reliable evaluations of models over time. We provide robust agent tooling and scaffolding to benchmark major LLMs as they are released, including "thinking" models like o3 and Gemini 2.5 Pro. We include automated evaluations of the lengthy agent traces to report progress over time in hallucinations, tool use, and forgetting. Finally, we evaluate the major web research products branded as "Deep Research", "Deep Search", "Search", or "Research." Results are available on a public leaderboard at https://drb.futuresearch.ai/.

cs.AI↗

Seed1.5-Thinking: Advancing Superb Reasoning Models with Reinforcement Learning

We introduce Seed1.5-Thinking, capable of reasoning through thinking before responding, resulting in improved performance on a wide range of benchmarks. Seed1.5-Thinking achieves 86.7 on AIME 2024, 55.0 on Codeforces and 77.3 on GPQA, demonstrating excellent reasoning abilities in STEM and coding. Beyond reasoning tasks, the method demonstrates notable generalization across diverse domains. For instance, it surpasses DeepSeek R1 by 8% in win rate on non-reasoning tasks, indicating its broader applicability. Compared to other state-of-the-art reasoning models, Seed1.5-Thinking is a Mixture-of-Experts (MoE) model with a relatively small size, featuring 20B activated and 200B total parameters. As part of our effort to assess generalized reasoning, we develop two internal benchmarks, BeyondAIME and Codeforces, both of which will be publicly released to support future research. Model trial link: https://www.volcengine.com/experience/ark.

cs.CL↗

Helicity-dependent parton distribution functions at next-to-next-to-leading order accuracy from inclusive and semi-inclusive deep-inelastic scattering data

We present MAPPDFpol1.0, a new determination of the helicity-dependent parton distribution functions (PDFs) of the proton from a set of longitudinally polarised inclusive and semi-inclusive deep-inelastic scattering data. The determination includes, for the first time, next-to-next-to-leading order QCD corrections to both processes, and is carried out in a framework that combines a neural-network parametrisation of PDFs with a Monte Carlo representation of their uncertainties. We discuss the quality of the determination, in particular its dependence on higher-order corrections, on the choice of data set, and on theoretical constraints.

hep-ph↗

Search for lepton-flavor-violating $τ^- \to \ell^- K_s^0$ decays at Belle and Belle II

We present the results of a search for charged-lepton-flavor violating decays $τ^{-} \rightarrow \ell^{-}K_{S}^{0}$, where $\ell^{-}$ is either an electron or a muon. We combine $e^+e^-$ data samples recorded by the Belle II experiment at the SuperKEKB collider (428 fb$^{-1}$) with samples recorded by the Belle experiment at the KEKB collider (980 fb$^{-1}$) to obtain a sample of 1.3 billion $e^+e^-\toτ^+τ^-$ events. We observe 0 and 1 events and set $90\%$ confidence level upper limits of $0.8 \times 10^{-8}$ and $1.2 \times 10^{-8}$ on the branching fractions of the decay modes $τ^{-} \rightarrow e^{-}K_{S}^{0}$ and $τ^{-} \rightarrow μ^{-}K_{S}^{0}$, respectively. These are the most stringent upper limits to date.

hep-ex↗

Command A: An Enterprise-Ready Large Language Model

In this report we describe the development of Command A, a powerful large language model purpose-built to excel at real-world enterprise use cases. Command A is an agent-optimised and multilingual-capable model, with support for 23 languages of global business, and a novel hybrid architecture balancing efficiency with top of the range performance. It offers best-in-class Retrieval Augmented Generation (RAG) capabilities with grounding and tool use to automate sophisticated business processes. These abilities are achieved through a decentralised training approach, including self-refinement algorithms and model merging techniques. We also include results for Command R7B which shares capability and architectural similarities to Command A. Weights for both models have been released for research purposes. This technical report details our original training pipeline and presents an extensive evaluation of our models across a suite of enterprise-relevant tasks and public benchmarks, demonstrating excellent performance and efficiency.

cs.CL↗