TY - RPRT TI - Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents AU - Liming Pu AU - Xiaoxia Li AU - Yifu Liu AU - Teng Cao AU - Bin Yang PY - 2026 UR - https://arxiv.org/abs/2609.01245 ID - 2609.01245 ER -