arXiv Science⌕ Search

arXiv · 2609.36593

Text2Sim: Agentic Physics-Based Simulation Generation with Distilled Expertise

Abstract

Creating diverse physical simulations remains labor-intensive because assets, layout, physical parameters, motion, control, and rendering must be designed and debugged jointly. We present Text2Sim, a simulation-specialized agentic pipeline that converts a text-only request into an executable, editable dynamic case. Built on Genesis, Text2Sim uses a hierarchical agentic structure that combines a Planner with specialized Writers, asset-generation tools, and an independent Critic. Compact skills (Debug Cards) distilled from graphics demonstrations provide role-specific physical guidance for execution-based repair. We evaluate physical quality, visual quality, and human preference on 42 held-out prompts spanning rigid, articulated, deformable, and cloth phenomena, with a paper-level split between experience construction and evaluation. We design automatic physical and visual scorers to evaluate the quality of the results, and Text2Sim achieves higher scores than all four state-of-the-art baselines on both metrics. In blinded user studies with these baselines, significantly more participants prefer Text2Sim than prefer the baselines, which is consistent with the results from our automatic scorers. The pipeline also supports a broad range of downstream applications; we select dataset construction and extension to multimodal input as two representative examples. We will release the code, the Debug Card library, and a dataset of generated cases, each pairing the text prompt and rendered video with the executable program, assets, physical parameters, controls, and recorded states.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Xiaoyu Xiong, Tsun-Hsuan Wang, Yi-Ling Qiao, Tao Du, Minchen Li. 2026-09-29. Text2Sim: Agentic Physics-Based Simulation Generation with Distilled Expertise. https://arxiv.org/abs/2609.36593

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Squeeze3D: Extreme Neural Compression with Latent Space Bridging

We propose Squeeze3D, a novel framework that leverages implicit prior knowledge learnt by existing pre-trained encoders and decoders to compress 3D data at extremely high compression ratios. Our approach bridges the latent spaces between a pre-trained encoder and a pretrained decoder model through trainable mapping networks. Any 3D asset represented as a mesh, point cloud, or radiance field is first encoded by the pre-trained encoder and then transformed (i.e. compressed) into a highly compact latent code by a mapping network. This latent code can effectively be used as an extremely compressed representation of the mesh, point cloud, or radiance field. A mapping network transforms the compressed latent code into the latent space of a powerful generative model; the decoder of this generative model then recreates the original 3D asset (i.e. decompression). Squeeze3D is trained entirely on generated synthetic data and does not require any 3D datasets. The Squeeze3D architecture can be flexibly used with existing pre-trained 3D encoders and existing generative models. It can flexibly support different formats, including meshes, point clouds, and radiance fields. Our experiments demonstrate that Squeeze3D achieves compression ratios of up to 2187$\times$ for textured meshes, 58.5$\times$ for point clouds, and more than 650$\times$ for radiance fields while maintaining visual quality comparable to many existing methods. Squeeze3D only incurs a small compression and decompression latency since it does not involve training object-specific networks to compress an object.

cs.GR↗

ViscoReg: Neural Signed Distance Functions via Viscosity Solutions

Implicit Neural Representations (INRs) that learn Signed Distance Functions (SDFs) from point cloud data represent the state-of-the-art for geometrically accurate 3D scene reconstruction. However, training these Neural SDFs involves enforcing the Eikonal equation, an ill-posed equation that also leads to unstable gradient flows. Numerical Eikonal solvers have relied on viscosity approaches for regularization and stability. Motivated by this well-established theory, we introduce ViscoReg, a novel regularizer for Neural SDF methods, and theoretically prove that it stabilizes training. Empirically, ViscoReg outperforms state-of-the-art approaches such as SIREN, DiGS, StEik, and HotSpot across most metrics on ShapeNet, the Surface Reconstruction Benchmark, 3D scene reconstruction and reconstruction from real scans. We also establish novel generalization error estimates for Neural SDFs in terms of the training error, using the theory of viscosity solutions. Our empirical and theoretical results provide confidence in the general applicability of our method.

cs.GR↗

SCAMP: Sparse-anchor Control is One Small Projection

Authoring with a text-to-motion generator needs sparse anchors: chosen joints, at chosen frames, at given positions. Meeting them currently costs a conditioning branch trained for the task, or hundreds of per-clip optimisation steps in the architecture's native variables. In any generator that decodes a continuous state through a frozen differentiable decoder, the anchors ask for little and leave most of the state free: a few hundred numbers against a state of tens of thousands. Every control method is a choice among the states that satisfy them, and the choices differ along the directions the anchors cannot see and the motion can. SCAMP makes the choice that moves none of them: damped Gauss-Newton in the space of the anchors, through the frozen decoder alone, training-free, with one dimensionless damping constant. Every increment is a combination of the rows of the anchors' Jacobian, so the correction is orthogonal to everything the anchors never see, and the system solved is the size of the request rather than of the state. Applied unchanged to seven published generators spanning diffusion, token and latent designs, it matches or exceeds in anchor error every released control method it is measured against, and closes the anchors on hosts that ship none. Confined to those rows, a correction can only take the shapes the decoder admits, so what it costs belongs to the decoder, and holding the solver fixed makes that cost measurable: it divides by decoder family, windowed decoders staying within a small multiple of the unconstrained generator's foot skating where analytic recoveries multiply it several times over. Built to that criterion, our own generator reaches 0.083 m anchor error at FID 0.102 in 0.50 s per clip. The decoder's temporal support is a design criterion for controllable motion generation.

cs.GR↗