arXiv Science⌕ Search

arXiv · 2610.03243

A Kinetic Theory of the Gated Self-Evolving LLM Agent

Abstract

We find traces of fluid dynamics in the self-evolution of an LLM agent, and give the kinetic theory that predicts them. Gated self-evolution is the loop in which an agent rewrites its own skills under a validation gate. Self-evolution research has treated the agent as the unit; we study instead the individual instances inside it. Here the agent is DSH-plugin-based: it runs in production on DeepSeek Harness (DSH), and its plugins satisfy four architectural properties (permutation symmetry, reversibility, acyclicity, typed contracts), which license treating these instances as identical hard spheres; the theory is accordingly scoped to DSH-class plugin populations. On this scope the paper builds three theory layers. The rigorous layer, independent of any analogy, comprises an any-time hitting-time certificate bounding the expected rounds to any prescribed improvement, a resolution law that prices held-out validation budgets, and a separation theorem: the daemon must stay outside the population, because merging evaluator with evaluated voids the certificate. The kinetic layer is a master equation over the plugin x version x task grid with four operators (collision, reaction, external field, gate), where collision is co-activation. Its moment hierarchy, the step that turns a gas into fluid equations, generates the falsifiable statistical signatures. Throughout, the fluid reading is a bounded analogy: momentum is not conserved, so no Navier-Stokes limit exists. The measured layer runs on a faithful minimal instance, a large library of four-parameter skill plugins retrieved one per episode with a co-activation probe, in a one-model, one-task-family WebShop environment; every element maps to the DSH loop by architectural role. Population fluctuation scaling is density-gated: invisible at sparse edit density, it emerges at the predicted rate under tripled density, as directional evidence.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Haipeng Wang. 2026-10-02. A Kinetic Theory of the Gated Self-Evolving LLM Agent. https://arxiv.org/abs/2610.03243

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Budgeted-GS: Real-Time Large-Scale Gaussian Splatting via Factoring LOD

3D Gaussian Splatting achieves excellent visual quality with real-time rendering, but at the scale of entire cities it does not fit: a trained model carries millions of primitives and gigabytes of memory, and real-time rendering at high quality on a consumer GPU remains out of reach. We introduce Budgeted-GS, a post-hoc method that turns any trained 3DGS model into a factoring tree, a multi-resolution hierarchy of moment-matched aggregates. After a construction pass of a few seconds, a single quality parameter selects, for each view, the level of detail that fits the memory of the target device, so the same city-scale model serves GPUs with widely different memory capacities. When a new scene is to be trained, the same theory applies: instead of growing a full-sized model and compressing it afterwards, budget-centered training first measures how many primitives the scene needs and then trains the model directly at that size, avoiding the wasted effort of optimizing primitives that are later discarded. Both methods are grounded in a measurable capacity floor, a budget-error law derived from optimal transport in phase space; selection rules certified by recent covering theorems decide which primitives are redundant. The floor answers how many primitives a scene actually needs and how many can safely be given up. We validate the floor on 13 public scenes under a preregistered protocol, and exercise both methods from object scenes to an official city capture, rendering it at native 1920x1080, full SH, in real time on one consumer GPU.

cs.GR↗

The Shape of Speech: A Geometric Measure of Coarticulation for Speech-Driven 3D Facial Animation

Speech-driven 3D facial animation can reproduce recognizable mouth poses. However, it can simplify the motion between them, and that motion carries coarticulation, the way the sounds around each sound shape its articulation. We introduce a geometric measure of this trajectory shaping: lip-path length compared with the shortest route through the vowel, consonant and vowel positions of a speech segment. In contrast to the endpoint chord, this consonant-aware route accounts for obligatory transit and avoids degeneracy, while preserving invariance to uniform motion gain. The measure needs only a forced alignment, so it applies where no ground truth exists. We demonstrate it on four state-of-the-art methods, one per architectural family, real-time and offline. All four trace flatter lip trajectories than captured speech. Against frame-rate-matched ground truth, DiffPoseTalk, ARTalk and FaceFormer show clear deficits, equivalent on this measure to removing 15-60% of real speech's fast articulatory component. CodeTalker is marginal on the primary measure and clear on a companion measure. A pre-registered study with 97 viewers and 3,523 judgments underpins the measured direction: controlled damping of real motion lowers the score and is penalized, whereas exaggeration shows no detected penalty over the tested range. Viewers also prefer real speech in 73.4% of sentence comparisons and, in the aggregate, on single words. Together, the measure, its calibration and the study identify a perceptually relevant loss of trajectory shaping and a concrete target for improving synthesized articulation.

cs.GR↗

MotionPersona: Real-Time Locomotion Control across Personas, Bodies, and Styles

We present MotionPersona, a generative framework for character-aware locomotion control, in which the motion for a command depends on the captured persona, the body shape, and the character's style. Unlike style, which one performer can vary at will, persona and body shape are coupled in capture: each performer is observed in only one body. The captured data therefore cannot uniquely determine which motion characteristics should follow the persona and which should change with the body, leaving unseen persona-body combinations unconstrained. We capture 48 performers, aged 5 to 68, under the same nine styles and seven commands, 44 of them with persona annotation. From this repeated-measures design, we identify two robust associations between body shape and gait. These measurements guide a cross-body specification of which characteristics should change and which should be preserved. We implement this specification through a physically informed retargeting pipeline, producing cross-body training data while penalizing penetration and foot skating. On this data we train a single generative controller. A shape-aware VAE compresses each motion block into a few latent tokens and renders them on a conditioned target body under explicit geometric supervision; over these tokens, a latent flow-matching prior generates persona- and style-conditioned motion in two sampling steps. The controller covers all captured personas, a wide family of SMPL-X target bodies, and nine styles in one model, and runs at 27 ms per block on two threads of a laptop CPU. We verify the framework at every stage, following the same gait descriptors from captured to retargeted to generated motion and sweeping each axis in isolation. To our knowledge, this is the first real-time locomotion controller that carries part of a captured persona's performer-specific variation across independently selected body shapes and styles.

cs.GR↗