arXiv ScienceSearch

arXiv subjects

Kecheng Xiao

Publications and source records attributed to Kecheng Xiao.

8 recordsLinked to original sources

The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The flagship M2 contains 229.9B total parameters with only 9.8B activated per token. Designed end-to-end for agentic deployment, the M2 series rests on three components: (i) agent-driven data pipelines producing large-scale, verifiable trajectories across agentic coding and agentic cowork, each grounded in an executable workspace and an artifact-aligned reward; (ii) Forge, a scalable agent-native RL system that adapts to long-horizon agent trajectories, paired with windowed-FIFO scheduling, prefix-tree merging, inference optimization, and a clean training-inference-agent decoupling that supports both white-box and black-box agents; (iii) the latest M2.7 checkpoint takes an early step toward self-evolution -- autonomously debugging training runs and modifying its own scaffold. Across M2 through M2.7, this combination translates a mini-activation footprint into frontier-tier performance on agentic coding, deep search, office-task, and reasoning benchmarks.

cs.AI

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

We introduce MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model. MiniMax-M1 is powered by a hybrid Mixture-of-Experts (MoE) architecture combined with a lightning attention mechanism. The model is developed based on our previous MiniMax-Text-01 model, which contains a total of 456 billion parameters with 45.9 billion parameters activated per token. The M1 model natively supports a context length of 1 million tokens, 8x the context size of DeepSeek R1. Furthermore, the lightning attention mechanism in MiniMax-M1 enables efficient scaling of test-time compute. These properties make M1 particularly suitable for complex tasks that require processing long inputs and thinking extensively. MiniMax-M1 is trained using large-scale reinforcement learning (RL) on diverse problems including sandbox-based, real-world software engineering environments. In addition to M1's inherent efficiency advantage for RL training, we propose CISPO, a novel RL algorithm to further enhance RL efficiency. CISPO clips importance sampling weights rather than token updates, outperforming other competitive RL variants. Combining hybrid-attention and CISPO enables MiniMax-M1's full RL training on 512 H800 GPUs to complete in only three weeks, with a rental cost of just $534,700. We release two versions of MiniMax-M1 models with 40K and 80K thinking budgets respectively, where the 40K model represents an intermediate phase of the 80K training. Experiments on standard benchmarks show that our models are comparable or superior to strong open-weight models such as the original DeepSeek-R1 and Qwen3-235B, with particular strengths in complex software engineering, tool utilization, and long-context tasks. We publicly release MiniMax-M1 at https://github.com/MiniMax-AI/MiniMax-M1.

cs.CL

MiniMax-01: Scaling Foundation Models with Lightning Attention

We introduce MiniMax-01 series, including MiniMax-Text-01 and MiniMax-VL-01, which are comparable to top-tier models while offering superior capabilities in processing longer contexts. The core lies in lightning attention and its efficient scaling. To maximize computational capacity, we integrate it with Mixture of Experts (MoE), creating a model with 32 experts and 456 billion total parameters, of which 45.9 billion are activated for each token. We develop an optimized parallel strategy and highly efficient computation-communication overlap techniques for MoE and lightning attention. This approach enables us to conduct efficient training and inference on models with hundreds of billions of parameters across contexts spanning millions of tokens. The context window of MiniMax-Text-01 can reach up to 1 million tokens during training and extrapolate to 4 million tokens during inference at an affordable cost. Our vision-language model, MiniMax-VL-01 is built through continued training with 512 billion vision-language tokens. Experiments on both standard and in-house benchmarks show that our models match the performance of state-of-the-art models like GPT-4o and Claude-3.5-Sonnet while offering 20-32 times longer context window. We publicly release MiniMax-01 at https://github.com/MiniMax-AI.

cs.CL

AMEIR: Automatic Behavior Modeling, Interaction Exploration and MLP Investigation in the Recommender System

Recently, deep learning models have been widely spread in the industrial recommender systems and boosted the recommendation quality. Though having achieved remarkable success, the design of task-aware recommender systems usually requires manual feature engineering and architecture engineering from domain experts. To relieve those human efforts, we explore the potential of neural architecture search (NAS) and introduce AMEIR for Automatic behavior Modeling, interaction Exploration and multi-layer perceptron (MLP) Investigation in the Recommender system. The core contributions of AMEIR are the three-stage search space and the tailored three-step searching pipeline. Specifically, AMEIR divides the complete recommendation models into three stages of behavior modeling, interaction exploration, MLP aggregation, and introduces a novel search space containing three tailored subspaces that cover most of the existing methods and thus allow for searching better models. To find the ideal architecture efficiently and effectively, AMEIR realizes the one-shot random search in recommendation progressively on the three stages and assembles the search results as the final outcome. Further analysis reveals that AMEIR's search space could cover most of the representative recommendation models, which demonstrates the universality of our design. The extensive experiments over various scenarios reveal that AMEIR outperforms competitive baselines of elaborate manual design and leading algorithmic complex NAS methods with lower model complexity and comparable time cost, indicating efficacy, efficiency and robustness of the proposed method.

cs.LG

AST/RO Observations of CO J=4-3 Emission from the N44 Complex in the Large Magellanic Cloud

We present Antarctic Submillimeter Telescope and Remote Observatory (AST/RO) observations of 12CO J=4-3 and C I emission in the N44 H II complex in the Large Magellanic Cloud. We detected strong 12CO J=4-3 emission toward the H II region called as N44BC, which is located on the rim of an expanding giant shell in the N44 region. Analysis with a photodissociation region (PDR) model showed that the 12CO J=4-3 emitting cloud is very dense, with n ~ 10^5 cm^-3. We also note that there is a high-velocity component associated with the 12CO J=4-3 emission. This probably originates from molecular material accelerated as a result of the motion induced by the expanding giant shell surrounding LH47 in the N44 complex. We found that the kinetic energy of this high-velocity gas observed in the CO J=4-3 emission toward the rim of the expanding H II shell is at least an order of magnitude higher than the kinetic energy derived for the H I and H II gas in this region.

astro-ph

Gas Density, Stability, and Starbursts Near the Inner Lindblad Resonance of the Milky Way

A key project of the Antarctic Submillimeter Telescope and Remote Observatory (AST/RO) reported by Martin et al. (2004) is the mapping of CO J=4-3 and J=7-6 emission from the inner Milky Way, allowing determination of gas density and temperature. Galactic center gas that Binney et al. (1991) identify as being on x_2 orbits has a density near 10^3.5 cm ^-3, which renders it only marginally stable against gravitational coagulation into a few Giant Molecular Clouds, as discussed by Elmegreen (1994). This suggests a relaxation oscillator mechanism for starbursts in the Milky Way, where inflowing gas accumulates in a ring at 150 pc radius for approximately 20 million years, until the critical density is reached, and the resulting instability leads to the sudden formation of giant clouds and the deposition of 4 x 10^7 solar masses of gas onto the Galactic center.

astro-ph

The AST/RO Survey of the Galactic Center Region. I. The Inner 3 Degrees

We present fully-sampled maps of 461 GHz CO (4-3), 807 GHz CO (7-6), and 492 GHz [CI] (3P1-3P0) emission from the inner 3 degrees of the Galactic Center region taken with the Antarctic Submillimeter Telescope and Remote Observatory (AST/RO) in 2001--2002. The data cover -1.3 < l < 2, -0.3 < b < 0.2 with 0.5 arcmin spacing, resulting in spectra in 3 transitions at over 24,000 positions on the sky. The CO (4-3) emission is found to be essentially coextensive with lower-J transitions of CO. The CO (7-6) emission is spatially confined to a far smaller region than the lower-J CO lines. The [CI] (3P1-3P0) emission has a spatial extent similar to the low-J CO emission, but is more diffuse. Bright CO (7-6) emission is detected in the well-known Galactic Center clouds Sgr A and Sgr B. We also detect CO (4-3) and CO (7-6) absorption from spiral arms in the galactic disk at velocities near 0 km s^-1 along the line of sight to the Galactic Center. Analyzing our CO (7-6) and CO (4-3) data in conjunction with J = 1 - 0 12CO and 13CO data previously observed with the Bell Laboratories 7-m antenna, we apply a Large Velocity Gradient (LVG) model to estimate the kinetic temperature and density of molecular gas in the inner 200 pc of the Galactic Center region. We show maps of the derived distribution of gas density and kinetic temperature as a function of position and velocity for the entire region. Kinetic temperature was found to decrease from relatively high values (>70K) at cloud edges to low values (<50K) in the interiors. Typical gas pressures in the Galactic Center gas are n(H_2) T_kin approx 10^5.2 K cm^-3. We also present an (l,b) map of molecular hydrogen column density derived from our LVG results.

astro-ph

Results from the AST/RO Survey of the Galactic Center Region

We have used the Antarctic Submillimeter Telescope and Remote Observatory (AST/RO), a 1.7m diameter single-dish submillimeter-wave telescope at the geographic South Pole, to determine the physical state of gas in the Galactic Center region and assess its stability. We present an analysis based on data obtained as part of an ongoing AST/RO key project: the large-scale mapping of the dominant cooling lines of the molecular interstellar medium in the Milky Way. These data are released for general use.

astro-ph