TY - RPRT TI - ScalePRM: Training Process Reward Models by Scaling Verification Compute Without Ground Truth AU - Salman Rahman AU - Sruthi Gorantla AU - Arpit Gupta AU - Swastik Roy AU - Nanyun Peng AU - Yang Liu PY - 2026 UR - https://arxiv.org/abs/2512.03244 ID - 2512.03244 ER -