arXiv ScienceSearch

arXiv subjects

Zhikun Xu

Publications and source records attributed to Zhikun Xu.

2 recordsLinked to original sources

Skill Reuse as Compression in Agentic RL

Large language model agents trained with reinforcement learning (RL) often learn brittle, task-specific shortcuts. We hypothesize that agents generalize better when their successful trajectories are structurally compressible, decomposed into a small set of reusable abstract patterns. To formalize this, we introduce ReuseRL, which grounds agentic RL in the Minimum Description Length (MDL) principle. ReuseRL extracts a shared skill dictionary from successful trajectories and augments the RL objective with a segmentation cost, explicitly penalizing idiosyncratic behaviors that encode poorly. We prove a PAC-Bayes bound guaranteeing that a dictionary extracted from successful trajectories has bounded expected description length on future successful behavior. Across ALFWorld, TextWorld-Cooking, and Countdown-Stepwise, ReuseRL improves in- and out-of-distribution success over vanilla GRPO and strong round-length baselines.

cs.LG

BOW: Training Language Models to Reason Over Plausible Next Words

Next-word prediction (NWP) trains language models against a single observed continuation, even though many contexts admit multiple plausible next words. Recent RL-based next-word reasoning methods make this tension explicit: they reward a model for producing a rationale that supports one context-conditioned continuation, which can turn a pre-existing preference into a confident, self-justifying trajectory. We introduce BOW, an RL framework that instead trains models to produce self-contained, neutral, and comprehensive descriptions of the plausible next-word space. The policy generates a next-word reasoning trajectory from the full context, but a frozen scorer computes the core reward from that trajectory alone, without separately receiving the context. BOW-Reg adds a lightweight breadth regularizer to this core reward to discourage premature collapse. On two model backbones, BOW remains competitive with the original models and often outperforms trained baselines across ten general reasoning benchmarks. On benchmarks testing ambiguous references and word meanings, BOW-Reg achieves the highest SharedRef correctness and the lowest HoWN-Simple single-sense collapse on both backbones. Human evaluation further shows that BOW-Reg produces broader next-word reasoning trajectories, while direct next-word-prediction evaluation shows that these trajectories remain predictive.

cs.CL