arXiv ScienceSearch

arXiv subjects

Jun Mei

Publications and source records attributed to Jun Mei.

13 recordsLinked to original sources

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents

LLM agents are rapidly being deployed in production, including coding assistants, customer-support chatbots, and scientific research assistants, yet they remain fundamentally static in enterprise deployment. The LLM weights, system prompts, tool repertoires, and in-context harnesses are frozen at deployment time, and any improvement requires a manual loop of human-curated data collection, offline fine-tuning, modification of the agentic paradigm, and re-deployment. Recent work on self-evolving agents, such as OpenClaw for individual users, indicates that the next leap in agent capability will come from agents that continually learn from their own experience. In this paper, we argue that this vision for self-evolving agent deployment is being held back for enterprise-level large-scale agentic service not by reinforcement learning (RL) algorithms but by agentic online RL systems. Specifically, current agentic RL systems and the surrounding observability software stack are inadequate along three essential aspects: (i) there is no standardized agent trajectory data protocol capable of carrying RL learning signals at step granularity across heterogeneous agent paradigms; (ii) there is no enterprise-grade comprehensive data proxy that converts real workloads into governed learning substrates; and (iii) there is no unified agent evolution control plane that automatically decides, based on trajectory statistics, when to update policy weights or evolve the in-context harness. The next generation of agentic RL systems must be co-designed around these three pillars, and we sketch concrete architectures, case studies, and counter-arguments. We instantiate one branch through AReaL2.0, reorganizing existing RL infrastructure into an agent-oriented online RL loop for policy weight updates from deployed workloads.

cs.DC

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale

Efficient and scalable agentic intelligence requires models that can deliver both low-latency responses and strong reasoning capabilities while remaining practical to train, serve, and deploy. In this report, we present Ling-2.6 and Ring-2.6, a family of models designed to address this challenge at scale. Ling-2.6 is optimized for instant response generation and high capability per output token, whereas Ring-2.6 is tailored for deeper reasoning and more advanced agentic workflows. Instead of training from scratch, we upgrade the Ling-2.0 base model through architectural migration pre-training and large-scale post-training. This upgrade is guided by a unified co-design of model architecture, optimization objectives, serving systems, and agent training environments, enabling improvements in both model capability and deployment efficiency. At the architectural level, we introduce a hybrid linear attention design that integrates Lightning Attention with MLA, improving the efficiency of long-context training and decoding. To further enhance token efficiency, we optimize capability per output token through Evolutionary Chain-of-Thought, Linguistic Unit Policy Optimization, bidirectional preference alignment, and shortest-correct-response distillation. For agentic capabilities, we propose KPop, a reinforcement learning framework designed to support stable training of Ring-2.6-1T on large-scale environment-grounded data. KPop improves training efficiency through asynchronous scheduling across coding, search, tool use, and workflow execution, enabling scalable learning from complex agent-environment interactions. Together, Ling-2.6 and Ring-2.6 provide a practical pathway toward efficient, scalable, and open agentic systems. We open-source all checkpoints in the 2.6 family to support further research and development in practical agentic intelligence.

cs.CL

Every Step Evolves: Scaling Reinforcement Learning for Trillion-Scale Thinking Model

We present Ring-1T, the first open-source, state-of-the-art thinking model with a trillion-scale parameter. It features 1 trillion total parameters and activates approximately 50 billion per token. Training such models at a trillion-parameter scale introduces unprecedented challenges, including train-inference misalignment, inefficiencies in rollout processing, and bottlenecks in the RL system. To address these, we pioneer three interconnected innovations: (1) IcePop stabilizes RL training via token-level discrepancy masking and clipping, resolving instability from training-inference mismatches; (2) C3PO++ improves resource utilization for long rollouts under a token budget by dynamically partitioning them, thereby obtaining high time efficiency; and (3) ASystem, a high-performance RL framework designed to overcome the systemic bottlenecks that impede trillion-parameter model training. Ring-1T delivers breakthrough results across critical benchmarks: 93.4 on AIME-2025, 86.72 on HMMT-2025, 2088 on CodeForces, and 55.94 on ARC-AGI-1. Notably, it attains a silver medal-level result on the IMO-2025, underscoring its exceptional reasoning capabilities. By releasing the complete 1T parameter MoE model to the community, we provide the research community with direct access to cutting-edge reasoning capabilities. This contribution marks a significant milestone in democratizing large-scale reasoning intelligence and establishes a new baseline for open-source model performance.

cs.CL

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs

We present Ring-lite, a Mixture-of-Experts (MoE)-based large language model optimized via reinforcement learning (RL) to achieve efficient and robust reasoning capabilities. Built upon the publicly available Ling-lite model, a 16.8 billion parameter model with 2.75 billion activated parameters, our approach matches the performance of state-of-the-art (SOTA) small-scale reasoning models on challenging benchmarks (e.g., AIME, LiveCodeBench, GPQA-Diamond) while activating only one-third of the parameters required by comparable models. To accomplish this, we introduce a joint training pipeline integrating distillation with RL, revealing undocumented challenges in MoE RL training. First, we identify optimization instability during RL training, and we propose Constrained Contextual Computation Policy Optimization(C3PO), a novel approach that enhances training stability and improves computational throughput via algorithm-system co-design methodology. Second, we empirically demonstrate that selecting distillation checkpoints based on entropy loss for RL training, rather than validation metrics, yields superior performance-efficiency trade-offs in subsequent RL training. Finally, we develop a two-stage training paradigm to harmonize multi-domain data integration, addressing domain conflicts that arise in training with mixed dataset. We will release the model, dataset, and code.

cs.CL

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reinforcement learning (RL) has become a dominant paradigm for training large language models (LLMs), particularly for reasoning tasks. Effective RL for LLMs requires massive parallelization and poses an urgent need for efficient training systems. Most existing large-scale RL systems for LLMs are synchronous, alternating generation and training in a batch setting where rollouts in each training batch are generated by the same model. This approach stabilizes RL training but suffers from severe system-level inefficiency: generation must wait until the longest output in the batch is completed before model updates, resulting in GPU underutilization. We present AReaL, a fully asynchronous RL system that completely decouples generation from training. Rollout workers in AReaL continuously generate new outputs without waiting, while training workers update the model whenever a batch of data is collected. AReaL also incorporates a collection of system-level optimizations, leading to substantially higher GPU utilization. To stabilize RL training, AReaL balances the workload of rollout and training workers to control data staleness, and adopts a staleness-enhanced PPO variant to better handle outdated training samples. Extensive experiments on math and code reasoning benchmarks show that AReaL achieves up to 2.77$\times$ training speedup compared to synchronous systems with the same number of GPUs and matched or improved final performance. The code of AReaL is available at https://github.com/inclusionAI/AReaL/.

cs.LG

Acoustic Metagrating Circulators: Nonreciprocal, Robust, and Tunable Manipulation with Unitary Efficiency

Nonreciprocal signal operation is highly desired for various acoustic applications, where protection from unwanted backscattering can be realized so that transmitting and receiving signals are processed in a full-duplex mode. Here we present the realization of a class of nonreciprocal circulators based on simply structured acoustic metagratings, which consist only of a few solid cylinders and a steady fluid flow with low velocity. These innovative metagratings are intelligently designed via a diffraction analysis of the linearized potential flow equation and a genetic-algorithm-based optimization process. Unitary reflection efficiency between desired ports of the circulators are demonstrated through full-wave numerical simulations, confirming nonreciprocal and robust circulation of the acoustic signal over a broad range of flow velocity magnitude and profile. Our design provides a feasible degree of tunability, including switching from reciprocal to nonreciprocal operation and reversing the handedness of the circulator, presenting a convenient but efficient approach for the realization of nonreciprocal acoustic devices from wavelength-thick metagratings. It may find applications in various scenarios including underwater communication, energy harvesting, and acoustic sensing.

physics.app-ph

Maximum A Posteriori Inference in Sum-Product Networks

Sum-product networks (SPNs) are a class of probabilistic graphical models that allow tractable marginal inference. However, the maximum a posteriori (MAP) inference in SPNs is NP-hard. We investigate MAP inference in SPNs from both theoretical and algorithmic perspectives. For the theoretical part, we reduce general MAP inference to its special case without evidence and hidden variables; we also show that it is NP-hard to approximate the MAP problem to $2^{n^\epsilon}$ for fixed $0 \leq \epsilon < 1$, where $n$ is the input size. For the algorithmic part, we first present an exact MAP solver that runs reasonably fast and could handle SPNs with up to 1k variables and 150k arcs in our experiments. We then present a new approximate MAP solver with a good balance between speed and accuracy, and our comprehensive experiments on real-world datasets show that it has better overall performance than existing approximate solvers.

cs.AI

Acoustic frequency filter based on anisotropic topological phononic crystals

There are growing efforts in constructing topological edge states in classical wave system. However, most of the work study the existence, creation and properties of the edge states, and the demonstration of application is highly desirable. Here, we present our design of a two-dimensional anisotropic phononic crystal that exhibits tunable topological phases. We further explore the contribution of anisotropy and show that the bandgap topology is also related to particular directions and frequency. Such frequency dependent behavior can be utilized as a frequency filter.

physics.app-ph

Pseudo-time-reversal symmetry and topological edge states in two-dimensional acoustic crystals

We propose a simple two-dimensional acoustic crystal to realize topologically protected edge states for acoustic waves. The acoustic crystal is composed of a triangular array of core-shell cylinders embedded in a water host. By utilizing the point group symmetry of two doubly degenerate eigenstates at the \Gamma point, we can construct pseudo-time-reversal symmetry as well as pseudo-spin states in this classical system. We develop an effective Hamiltonian model for the associated dispersion bands around the Brillouin zone center, and find the inherent link between the band inversion and the topological phase transition. With numerical simulations, we unambiguously demonstrate the unidirectional propagation of acoustic edge states along the interface between a topologically nontrivial acoustic crystal and a trivial one, and the robustness of the edge states against defects with sharp bends. Our work provides a new design paradigm for manipulating and transporting acoustic waves in a topologically protected manner. Technological applications and devices based on our design are expected in various frequency ranges of interest, spanning from infrasound to ultrasound.

cond-mat.mtrl-sci

Sound Absorption by Subwavelength Membrane Structures: A Generalized Perspective

Decorated membrane, comprising a thin layer of elastic film with small rigid platelets fixed on top, has been found to be an efficient absorber of low frequency sound. In this work we consider the problem of sound absorption from a perspective aimed at deriving upper bounds under different scenarios, i.e., whether the sound is incident from one side only or from both sides, and whether there is a reflecting surface on the back side of the membrane. By considering the negligible thickness of the membrane, usually on the order of a fraction of one millimeter, we derive a relation showing that the sum of the incoming sound waves' (complex) pressure amplitudes, averaged over the area of the membrane, must be equal to that of the outgoing waves. By using this relation, and without going to any details of the wave solutions, it is shown that the maximum absorption achievable from one-side incident is 50%, while the maximum absorption with a back reflecting surface can reach 100%. The latter was attained by the hybridized resonances. All the results are shown to be in excellent agreement with the experiments. This generalized perspective, when used together with the Green function formalism, can be useful in gaining insights and delineating the constraints on what are achievable in scatterings and absorption by thin film structures.

cond-mat.mtrl-sci

A Lumped Model for Rotational Modes in Phononic Crystals

We present a lumped model for the rotational modes induced by the rotational motion of individual scatterers in two-dimensional phononic crystals comprised of square arrays of solid cylindrical scatterers in solid hosts. The model provides a physical interpretation of the origin of the rotational modes, reveals the important role played by the rotational motion in the band structure, and reproduces the dispersion relations. The model increases the possibilities of wave manipulation in phononic crystals. In particular, expressions, derived from the model, for eigen-frequencies at high symmetry points unambiguously predict the presence of a new type of Dirac-like cone at the Brillouin center, which is found to be the result of accidental degeneracy of the rotational and dipolar modes.

cond-mat.mtrl-sci

Do Linear Dispersions of Classical Waves Mean Dirac Cones?

By using the \vec{k}\cdot\vec{p} method, we propose a first-principles theory to study the linear dispersions in phononic and photonic crystals. The theory reveals that only those linear dispersions created by doubly-degenerate states can be described by a reduced Hamiltonian that can be mapped into the Dirac Hamiltonian and possess a Berry phase of -\pi. Triply-degenerate states can also generate Dirac-like cone dispersions, but the wavefunctions transform like a spin-1 particle and the Berry phase is zero. Our theory is capable of predicting accurately the linear slopes of Dirac/Dirac-like cones at various symmetry points in a Brilliouin zone, independent of frequency and lattice structure.

cond-mat.mtrl-sci

Acoustic transmission enhancement through a periodically-structured stiff plate without any opening

We report both experimentally and theoretically that the enhanced acoustic transmission can occur in the subwavelength region through a thin but stiff structured-plate without any opening. This exotic acoustic phenomenon is essentially distinct from the previous related studies originated from, either collectively or individually, the interaction of the incident wave with openings in previous structures. It is attributed to the structure-induced resonant excitation of the non-leaky Lamb modes that exist intrinsically in the uniform elastic plate. Our finding should have impact on ultrasonic applications.

physics.class-ph