TY - RPRT TI - PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO AU - Jingquan Wang AU - Jun Yin AU - Xu Han AU - Yongsheng Mei AU - Jie Hao AU - Bin Guo PY - 2026 UR - https://arxiv.org/abs/2610.04132 ID - 2610.04132 ER -