arXiv ScienceSearch

arXiv subjects

Hongyu Ge

Publications and source records attributed to Hongyu Ge.

8 recordsLinked to original sources

Speculative Rollback Correction for Quality-Diverse Web Agent Imitation

Training interactive web agents through imitation learning from expert trajectories has emerged as a highly effective approach. However, determining the optimal timing for expert intervention presents a critical challenge in this context. Delayed intervention often leads to the accumulation of early-stage errors, pushing the page state into an irrecoverable regime. Conversely, premature or excessive intervention causes the agent to become overly reliant on expert policies, trapping the model in local optima characterized by a single, rigid trajectory. We propose Speculative Rollback Correction (SRC), a branch-level imitation framework for resettable agent environments. Instead of requesting teacher labels at every visited state or correcting only after a completed trajectory, SRC uses fixed-horizon branch review: the student executes a short speculative segment before teacher review, and the teacher localizes the first harmful deviation only when local progress breaks. Rollback preserves useful prefixes, while successful rollouts are filtered by a hard verifier and retained in a lightweight quality-diversity archive. The resulting data supports next-action supervised fine-tuning on both localized corrections and verifier-passing trajectories. On WebArena-Infinity, SRC collects 977 verifier-passing trajectories and 9,183 next-action examples; fixed-horizon review improves the recovery-versus-query tradeoff over step-level review while retaining verifier-passing solution variants. Code is available at https://github.com/LongkunHao/SRC_gui_agent.

cs.LG

Where Hindsight Credit Can Reside: A Signed-Capacity View of Token Updates in RLVR

Reinforcement Learning with Verifiable Rewards (RLVR) improves the reasoning ability of Large Language Models (LLMs), but sparse outcome rewards make token-level credit assignment difficult. We study token-level credit as a reward-conditioned shift from the behavior policy to a hindsight posterior. In autoregressive RLVR, this shift can be expressed through Conditional Mutual Information (CMI), which shows that token entropy upper-bounds possible hindsight credit. Entropy, however, indicates capacity rather than update direction, so we introduce the Four Quadrant Decomposition to separate updates by reward polarity and token entropy. Controlled interventions show that these two factors jointly shape token updates. Sustained reasoning gains concentrate in signed high-entropy quadrants, whereas low-entropy updates saturate quickly. Based on this analysis, we propose Hindsight-Aware Policy Optimization (HAPO), a sign-preserving modification to GRPO that performs capacity-guided advantage reallocation. Experiments on mathematical reasoning benchmarks in two model settings show that HAPO achieves competitive performance among entropy-aware baselines.

cs.LG

GLM-5: from Vibe Coding to Agentic Engineering

We present GLM-5, a next-generation foundation model designed to transition the paradigm of vibe coding to agentic engineering. Building upon the agentic, reasoning, and coding (ARC) capabilities of its predecessor, GLM-5 adopts DSA to significantly reduce training and inference costs while maintaining long-context fidelity. To advance model alignment and autonomy, we implement a new asynchronous reinforcement learning infrastructure that drastically improves post-training efficiency by decoupling generation from training. Furthermore, we propose novel asynchronous agent RL algorithms that further improve RL quality, enabling the model to learn from complex, long-horizon interactions more effectively. Through these innovations, GLM-5 achieves state-of-the-art performance on major open benchmarks. Most critically, GLM-5 demonstrates unprecedented capability in real-world coding tasks, surpassing previous baselines in handling end-to-end software engineering challenges. Code, models, and more information are available at https://github.com/zai-org/GLM-5.

cs.LG

ClinKD: Cross-Modal Clinical Knowledge Distiller For Multi-Task Medical Images

Medical Visual Question Answering (Med-VQA) represents a critical and challenging subtask within the general VQA domain. Despite significant advancements in general VQA, multimodal large language models (MLLMs) still exhibit substantial limitations when handling multi-task VQA scenarios. These limitations manifest through erroneous spatial localization and misinterpretation of medical images, which primarily arise from two fundamental issues: inadequate image-text alignment and insufficient domain-specified knowledge for medical applications. To address these issues, we introduce the Cross-Modal Clinical Knowledge Distiller (ClinKD), an innovative framework designed to enhance image-text alignment and establish more effective medical knowledge transformation mechanisms, which enables MLLMs to perform better even when lacking prior medical knowledge. Our extensive experimental evaluations demonstrate that the ClinKD achieves state-of-the-art performance on several datasets which are challenging for Med-VQA task. The results indicate that our approach not only significantly improves image-text alignment but also effectively enables MLLMs to adapt to the medical knowledge. The source code for ClinKD is available at: https://github.com/overloadedHenry/ClinKD.

cs.CV

A theoretical investigation on the transport properties of armchair biphenylene nanoribbons

Armchair biphenylene nanoribbons are investigated by using density functional theory. The nanoribbon that contains one biphenylene subunit in a unit cell is a semiconductor with a direct band gap larger than 1 eV, while that containing four biphenylene subunits is a metal. The semiconducting nanoribbon has high electron mobility of 57174 cm2V-1s-1, superior to armchair graphene nanoribbons. Negative differential resistance behavior is observed in two electronic devices composed of the semiconducting and metallic nanoribbons. The on/off ratios are in the order of 10^3. All these indicate that armchair biphenylene nanoribbons are potential candidates for ultra-small logic devices.

cond-mat.mes-hall

A theoretical prediction on huge hole and electron mobilities of 6,6,18-graphdiyne nanoribbons

Two-dimensional 6,6,18-graphdiyne and the corresponding one-dimensional nanoribbons are investigated using crystal orbital method. Based on HSE06 functional, the one-dimensional confinement increases the band gaps. With band gaps larger than 0.4 eV, thirty-three 6,6,18-graphdiyne nanoribbons have larger majority carrier mobilities at room temperature than the highest value of armchair graphene nanoribbons. Unlike {\gamma}-graphdiyne, 6,6,18-graphdiyne nanoribbons have both huge hole and electron mobilities, depending on whether they are armchair or zigzag type. The huge mobilities are explained by crystal orbital analysis. The superior capabilities of 6,6,18-graphdiyne nanoribbons make them possible candidates for high speed electronic devices in complementary circuits.

cond-mat.mtrl-sci

Theoretical investigation on armchair graphene nanoribbons with oxygen-terminated edges

Armchair graphene nanoribbons with different proportions of edge oxygen atoms are investigated by using crystal orbital method based on density functional theory. All the nanoribbons are energetically favorable, although buckled edges are present. Isolated edge oxygen atoms cause semiconductor-metal transition via introducing edge states, while adjacent edge oxygen atoms not. For the graphene nanoribbons with all oxygen atoms on the edges, both band gap and carrier mobility alternate with respect to the ribbon width. The carrier mobilities are as 18%-65% large as those of the graphene nanoribbons with hydrogen-terminated edges. These values are as large as 103 cm2V-1s-1, which are still quite high for electronic devices. Crystal orbital analysis gives pictorial explanations to the phenomenon.

cond-mat.mtrl-sci

Theoretical investigation on electronic properties and carrier mobilities of armchair graphyne nanoribbons

Seven types of armchair graphyne nanoribbons are investigated with HSE06 functional. The quantum confinements in the graphyne nanoribbons open or increase the band gaps of the corresponding two-dimensional graphynes, which is crucial to high on/off ratio in electronic device operation. The major carrier mobilities of the graphyne nanoribbons with high percentage of sp hybridized carbon atoms are very large. The sparse linking pattern results in small number of frontier crystal orbitals and small deformation potential constants, which are responsible for the large carrier mobilities. Some graphyne nanoribbons have band gaps larger than 0.4 eV. Meanwhile, they have both high hole and electron mobilities. These benefit current complementary circuit with low power dissipation. Especially, the hole and electron mobilities of 14,14,18-graphyne nanoribbons are more than an order larger than those of the armchair graphene nanoribbons, indicating that they have potential applications in high speed electronic devices.

cond-mat.mtrl-sci