arXiv Science⌕ Search

arXiv · 2610.04211

CurveCodec 2: Skeleton-agnostic animation compression with a learned entropy model

Abstract

Skeletal motion is stored as every joint's transform at every frame, yet most of it is implied by the body rather than by what the motion is about. Compression is one way to ask what a motion must still say once the body is known, and a production codec must answer it for any skeleton with a stated error bound. Our earlier codec, CurveCodec, matched the mean error of ACL, the production library of modern game engines, with a learned prior over sparse anchors, but not ACL's worst case, and it counted its payload as floats rather than bits. Here we ask where the redundancy of skeletal motion lies and which part of a codec a learned model should take over. Measurements give three answers. At production precision the largest saving comes from predicting each quantized curve from its own past, the second from choosing per joint, in closed loop through the hierarchy, which samples not to code. On the gaps such an encoder leaves, a nearest-neighbour oracle over millions of training samples is no better than linear interpolation, and no learned in-betweener we tried paid for itself. What a network does learn is the distribution of the residuals the codec must send. CurveCodec 2 codes every sub-track as a curve in the log map, quantized in closed loop and thinned to rate-distortion-selected keys, with residuals entropy-coded under a small learned model whose integer inference is bit-exact across platforms. Two contracts are verified on every decoded clip: ACL's own worst case per joint within a stated tolerance, or ACL's mean error per clip. On a held-out test side of 4,472 clips from 33 datasets, CurveCodec 2 needs 0.37x ACL's bytes at ACL's default precision of 0.01 cm under the worst-case contract and 0.22x at 0.1 cm under the mean contract, decodes on one CPU core, and transfers without retraining to a species absent from training. Project page: https://rubbly.cn/publications/curvecodec/

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mingyi Shi, Huancheng Lin, Xuelin Chen, Taku Komura. 2026-10-03. CurveCodec 2: Skeleton-agnostic animation compression with a learned entropy model. https://arxiv.org/abs/2610.04211

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Event-T2M: Event-level Conditioning for Complex Text-to-Motion Synthesis

Text-to-motion generation has advanced with diffusion models, yet existing systems often collapse complex multi-action prompts into a single embedding, leading to omissions, reordering, or unnatural transitions. In this work, we shift perspective by introducing a principled definition of an event as the smallest semantically self-contained action or state change in a text prompt that can be temporally aligned with a motion segment. Building on this definition, we propose Event-T2M, a diffusion-based framework that decomposes prompts into events, encodes each with a motion-aware retrieval model, and integrates them through event-based cross-attention in Conformer blocks. Existing benchmarks mix simple and multi-event prompts, making it unclear whether models that succeed on single actions generalize to multi-action cases. To address this, we construct HumanML3D-E, the first benchmark stratified by event count. Experiments on HumanML3D, KIT-ML, and HumanML3D-E show that Event-T2M matches state-of-the-art baselines on standard tests while outperforming them as event complexity increases. Human studies validate the plausibility of our event definition, the reliability of HumanML3D-E, and the superiority of Event-T2M in generating multi-event motions that preserve order and naturalness close to ground-truth. These results establish event-level conditioning as a generalizable principle for advancing text-to-motion generation beyond single-action prompts.

cs.GR↗

SECA: Strict-Elastic Compact Activation for Differentiable Elastoplastic Deformation

Return mapping keeps an exact elastic domain but switches its response tangent at yield, while smoothing laws with a sub-yield tail admit plastic flow below the threshold. We present SECA, a Strict-Elastic Compact Activation method for differentiable elastoplastic simulation. SECA prescribes a bounded $C^2$ activation over a finite positive-overstress interval and integrates it into a Perzyna-type flow law, so the implicit update returns exactly zero plastic increment for sub-yield trials while smoothing the onset. For scalar and proportional isotropic $J_2$ updates we derive an exact decomposition of the plastic-increment discrepancy into finite-width and viscosity contributions, and an explicit algebraic budget for the transition width. The update is embedded in finite-deformation loading histories with contact and complete unloading, and differentiated to released equilibria. Controlled indentation results stay closely matched. In the tested coarse cube-indentation run, where the transition is crossed without intermediate increments, SECA reduces whole-run Newton iterations by 47% and backtracking halvings by 93% against return mapping. A four-method comparison evaluates released-shape error and computational cost across loading partitions, and on a contact scene the path derivative drives a bounded control task to its target in three iterations. Further examples show elastic recovery and residual deformation under compression, tension and bending.

cs.GR↗

emg2face: Expressive Facial Animation with High-Density Surface EMG

Facial movements convey subtle and important information that is critical for human social communication. Optical methods for face capture are difficult or impossible to use when the face is occluded by head-mounted devices (HMDs), such as VR headsets. Even with a clear line of sight, such methods raise privacy concerns and require head-mounted capture rigs that offset cameras and lighting from the face. We show that high-density surface electromyography (HD-sEMG) provides a viable non-optical alternative that addresses these challenges. We measured 64 EMG channels, using two textile EMG grids, with 32 from the forehead (typically occluded by an HMD) and 32 from the side of the face. EMG data were digitized at 2048 Hz and filtered. Facial movements were simultaneously recorded and used to estimate 478 3D facial landmarks using MediaPipe's Face Landmarker. A major challenge in such multimodal recordings is synchronizing EMG and video data, which have different sampling frequencies and independent clocks. We developed a novel synchronization method using analog audio bursts that is capable of sub-millisecond synchronization. We also developed a staged fitting method that fits a recent high-resolution parametric head model (GNM), with 253 identity blendshapes and 383 expression blendshapes, to the MediaPipe landmarks as participants performed different facial expressions. We trained a deep neural network comprising per-grid spatial encoders followed by a dilated temporal convolutional network (TCN) to predict blendshape parameters from HD-sEMG signals at 100 Hz. Once trained, the network can predict expression blendshapes solely from HD-sEMG recordings. The output can be rendered using standard real-time blendshape animation methods. We demonstrate the methods using recordings from 25 participants, and direct expression transfer to a variety of human faces and non-human characters.

cs.GR↗