arXiv · 2609.36750
Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving
Abstract
Self-rewarding reinforcement learning (RL) enables large language models (LLMs) to self-evolve without human labels. Existing ensemble-based methods construct reward references from rollout groups and assign rewards accordingly. However, a response's reward representation also depends on its randomly sampled group context, i.e., the other responses in its group. Using only one group-context realization may miss desired reward signals and provide unreliable guidance for policy optimization. To address this issue, we propose Group-Marginalized Advantage Estimation (GMAE), which aggregates reward realizations across possible contexts into a response-level distribution and estimates expected advantages. Experiments across eight benchmarks and four base models demonstrate strong performance and cross-domain generalization. GMAE also exhibits stable learning, low extra cost, and good applicability across training datasets and RL backbones.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yiming Wang, Yikang Liu, Qingyuan Tian, Xingyu Chen, Zhuosheng Zhang, Zhaopeng Tu, Rui Wang. 2026-09-29. Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving. https://arxiv.org/abs/2609.36750
Cite the original work for its findings. Save a collection to share your selection of sources.