arXiv ScienceSearch

arXiv subjects

Peng Pei

Publications and source records attributed to Peng Pei.

2 recordsLinked to original sources

GIFT: Reconciling Post-Training Objectives via Variational Finite-Temperature Gibbs Initialization

The prevailing post-training paradigm for Large Reasoning Models (LRMs)---Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL)-suffers from an intrinsic optimization mismatch: the rigid likelihood maximization in SFT induces distributional collapse, thereby exhausting the exploration space necessary for subsequent RL. Motivated by the Gibbs optimum of KL-regularized RL, we derive a token-level variational surrogate that makes the SFT target structurally compatible with the subsequent RL stage, and propose Gibbs Initialization with Finite Temperature (GIFT). Standard SFT emerges as a degenerate zero-temperature limit of this surrogate, while a finite temperature preserves structural diversity. Our experiments demonstrate that GIFT outperforms standard SFT and surpasses competitive baselines when utilized for RL initialization. Our code is available at https://github.com/zzy1127/GIFT.

cs.LG

EvoCUA-1.5: Online Reinforcement Learning for Multi-turn Computer-Use Agents

Computer-use agents must solve long-horizon tasks through repeated interaction with partially observable, multimodal desktop environments. Although imitation learning and offline trajectory refinement provide strong priors, static traces cannot cover the causal feedback loop of real computer use: each action changes the screen state, future action space, and recovery options. EvoCUA-1.5 extends self-evolving computer-use agents from offline experience learning to online reinforcement learning, where policies interact with executable sandbox environments and improve from verifiable task outcomes. Online RL in this setting requires more than directly reusing single-turn language-RL recipes. Multi-turn interaction introduces context-managed observations, sparse terminal rewards, variable-length trajectories, and slow environment feedback. EvoCUA-1.5 addresses these challenges with Step-Level Policy Optimization (STEPO), which preserves trajectory-level advantage balance after decomposition into step-level samples; policy-aware filtering and pass-rate calibration over verifiable synthesized tasks; Dynamic Tri-Adaptive Curriculum (DTAC), which combines learnable tasks, difficult positive replay, and controlled infeasible-task exposure; and a fully asynchronous RL infrastructure with staleness control and mini-group batching. Experiments show that these components improve training stability and downstream performance. EvoCUA-1.5 achieves 63.2\% success on OSWorld-Verified, outperforming comparable 32B/35B-scale open-weight baselines and even approaching models with significantly larger parameter counts. Overall, EvoCUA-1.5 provides a practical framework for scaling online RL in multi-turn computer-use agents.

cs.AI