WATCH: World-aware Allied Trajectory and pose reConstruction for Camera and Human
Reconstructing global human motion from monocular video is fundamental to VR, graphics, and robotics, yet remains ill-posed due to depth ambiguity, motion ambiguity, and the entanglement of camera and human movements. Human-motion-centric methods achieve strong physical plausibility but leave two signals unused: camera orientation is processed through a fixed coordinate transformation with no independent supervision of its components, and camera velocity is discarded entirely despite being directly observable from SLAM. Camera-trajectory-centric methods use camera translation directly, but hard-decoding SLAM trajectories into human positions propagates depth errors and fails entirely under static cameras. We present WATCH (World-aware Allied Trajectory and pose reConstruction for Camera and Human). The key observation is that once camera orientation is made explicit, camera velocity becomes a natural additional input rather than an ambiguous one. We therefore decompose camera rotation into a network-estimated roll-pitch component and an analytically recoverable yaw, supervising each independently. This decomposition exposes a clean geometric interface through which camera velocity is incorporated as a learned spatial prior in the backbone, without the physically implausible artifacts that arise from hard-decoding. WATCH outperforms prior human-motion-centric methods on both static-camera (RICH) and dynamic-camera (EMDB) benchmarks in global trajectory accuracy, temporal smoothness, and physical plausibility, and remains robust when ground-truth camera is replaced with DPVO estimates.