TY - RPRT TI - Program Learning with Verifiable Rewards: Symbolic Backpropagation for Post-Training LLMs AU - Vishvesh Bhat PY - 2026 UR - https://arxiv.org/abs/2608.28421 ID - 2608.28421 ER -