arXiv ScienceSearch

arXiv subjects

Hang Zhou

Publications and source records attributed to Hang Zhou.

At least 19 recordsLinked to original sources

NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness

Recursive self-improvement (RSI) requires a concrete mechanism through which an AI system observes its capabilities and converts that evidence into the next round of learning. We present NeoHorse-1, a family of agent-native models developed to explore this path through agentic post-training. Our system combines a heterogeneous model pool with intelligent routing, recording the predicted capability demand, selected service tier, and subsequent interaction for each user turn. These records are converted into training examples that preserve interleaved reasoning, tool calls, and harness context, and are admitted through structural validation, six-dimensional semantic evaluation, and subscene-level labeling. Routing signals organize supervised fine-tuning into a three-stage curriculum and extend to routing-guided on-policy distillation, where a teacher supervises student-generated responses under the same progression. Capability-guided allocation then converts evaluation feedback into the next training mixture, closing an evaluation-selection-update loop in which what the system learns to do shapes what it learns from next. Across eleven benchmarks covering harness-based agents, tool use, coding, and instruction following, post-training raises the macro-average from 58.94 to 64.87 at 4B and from 65.60 to 69.04 at 9B, substantially narrowing the aggregate gap between the post-trained 4B model and the 9B base model. NeoHorse-1 provides an initial prototype of this feedback-driven process and a path toward harness-mediated RSI across successive iterations.

cs.CL

Cache-Aware Joint Router Adaptation for Memory-Efficient MoE Inference

Mixture-of-Experts (MoE) models activate few experts per token, yet their full expert sets can exceed GPU memory and require repeated weight transfers during decoding. We formulate expert-cache management as a model-side algorithmic problem and propose cache-aware post-training that jointly adapts the MoE backbone and lightweight auxiliary routers while preserving the native inference-time Top-K rule. The update-only Temporal Router learns same-layer retention across tokens without proactive loading. The full Spatio-Temporal Router adds a Spatio Router that uses the causal predecessor's hidden state to refine the temporal cache before target-layer access. We evaluate both modes on Qwen3 and GPT-OSS across GSM8K, MATH, and CommonsenseQA. Temporal Router consistently improves hit rate and reduces expert-weight traffic over matched LM-only baselines. On Qwen3, the full mode improves adjusted hit rate by 1.15--18.03 points and reduces traffic by 4.6--53.3\% relative to the strongest evaluated prefetching baseline; GPT-OSS results are competitive but task-dependent. Auxiliary-only training preserves baseline accuracy but yields modest coverage gains; joint post-training achieves substantially higher coverage. Sensitivity analyses distinguish the effects of cache capacity, refinement budget, and cache-loss weight on coverage, traffic, and quality.

cs.CL

Encore: Infinite Audio-Video Generation with Adaptive Signal Routing

Existing audio-video generation methods produce well-synchronized clips but are limited to short durations, while long-video generation methods extend duration through chunk-based iterative synthesis yet lack audio entirely. Generating long audio-video jointly is fundamentally harder than either task alone: each chunk must simultaneously maintain video temporal coherence, audio temporal coherence, and cross-modal synchronization, whose conditioning signals enter the model through different pathways. In this work, we present Encore for long-form synchronized audio-video generation. Our key insight is to factor this challenge into: (1) local continuity handled by iterative generation with explicit cross-chunk context propagation, and (2) global consistency enforced via reference audio-video signals with shifted position embedding. Building on this design, we propose Adaptive Signal Routing (ASR), which introduces learnable attention biases within self-attention and learnable residual scales on cross-attention outputs, enabling the model to adaptively modulate the influence of each conditioning signal. Trained end-to-end for joint audio-video generation, Encore also supports infinite-length audio-to-video and video-to-audio synthesis at inference by conditioning on the ground-truth modality throughout the denoising process. Experiments on our extended VerseBench for long audio-video evaluation demonstrate that Encore significantly outperforms existing methods in both generation quality and temporal coherence. Code and data for this paper are at https://github.com/shaohua-pan/Encore.

cs.MM

Pandora cluster Lensing, AGN, and Transient Exploration (PLATE) from JWST Multi-Epoch Imaging. I. Discovery of a type II supernova candidate in a spiral galaxy at $z=0.7$

We report the discovery and multi-wavelength analysis of a $z\sim0.7$ transient PLATE-23a in the Abell 2744 field, as the first results of the Pandora Lensing, AGN, and Transient Exploration (PLATE) project. Using multi-epoch JWST NIRCam imaging spanning from 2022 to 2025, we detect PLATE-23a in 12 filters. Difference-imaging analysis reveals its rising and declining phases. The host galaxy of PLATE-23a is a barred spiral at $z=0.688$ with a stellar mass of $\sim 10^{10.4}M_{\odot}$ and a star formation rate of $\sim5.1M_{\odot}~\rm yr^{-1}$, placing it on the star-forming main sequence. Bayesian light-curve classification strongly favors a type IIP supernova (SN) origin. It lies $\sim 13\rm kpc$ (after lensing correction) from the galaxy center in a region of low local star formation, suggesting the progenitor may have migrated from a distant star-forming clump. Physical properties derived from blackbody modeling indicate a temperature decreasing from $\sim8010\rm K$ to $\sim6100\rm K$ at around 65 days after the explosion; the late-time SED is consistent with entering the radioactive decay phase at about one rest-frame year. This work demonstrates the power of deep, multi-epoch JWST observations for studying transients at cosmological distances.

astro-ph.GA

Retrieval-guided Twin Fusion with Similarity-aware Contrast for Molecule-Text Alignment

This paper studies the problem of molecule-text alignment, which aims to project molecules and their textual descriptions into a joint latent space for downstream tasks including molecule search and molecular property prediction. Previous approaches typically combine graph structure mining with contrastive learning to enhance joint representation learning. However, they typically neglect fine-grained semantic relationships between substructures and texts, leading to suboptimal performance on downstream tasks. Towards this end, we propose a novel approach named Retrieval-guided Twin Fusion with Similarity-aware Contrast (RISEN) for molecule-text alignment. The core idea of RISEN is to construct a latent twin molecule for each substructure with cross-modal retrieval for semantic enhancement. In particular, for each substructure query, we retrieve relevant textual descriptions and sample several molecules that share similar descriptions of substructures. Then, we aggregate their representations via attention pooling for a twin latent representation, which would be further fused with the original substructure for representation enrichment. In addition, we measure the similarity across substructures and texts, which would further guide cross-modal contrastive learning with soft thresholding. Extensive experiments on benchmark datasets validate the superiority of the proposed RISEN in comparison with existing baselines.

cs.LG

An emerging baryon cycle in a galaxy 500 million years after the Big Bang

The emergence of stellar feedback as a regulator of galaxy growth marks a fundamental transition in cosmic history. At early times, rapid gas accretion and collapse may induce intense star formation before feedback becomes effective, producing feedback-free starbursts. When and how such bursts subsequently develop into self-regulated baryon cycles remain observationally unknown. Here we show that Gz9p3, a merging galaxy at $z=9.311$, is caught in this transition only 500 million years after the Big Bang. Deep JWST spectroscopy reveals a substantial neutral-gas reservoir along its merger-driven tidal structure and a multiphase outflow. Fine-structure absorption provides the first direct measurement of the electron density of the cool outflowing gas at high redshift ($\approx\,17\,{\rm cm^{-3}}$), yielding a mass-loading factor among the highest yet measured for galaxies of comparable stellar mass. The emergence of such efficient feedback after an intense burst is consistent with the delayed onset of feedback expected in feedback-free starburst models. The cool outflowing gas is unlikely to escape the host halo, implying that much of this metal-enriched material may remain available for future recycling through the circumgalactic medium. Gz9p3 therefore provides an early view of a baryon cycle being established through the interplay of merger-driven gas redistribution, bursty star formation and stellar feedback, suggesting that feedback-regulated recycling was already shaping galaxy growth during the epoch of reionization.

astro-ph.GA

Agentic Routing: The Harness-Native Data Flywheel

Large language model agents are increasingly executed not by a single model call, but by an execution harness that manages observation, context, control, action, state, and verification. At the same time, frontier and open models are becoming structurally specialized: a model that is strong at code editing, long-context recovery, tool use, mathematical reasoning, or low-latency response may not dominate on the other axes. This makes model selection inside an agent a core systems problem rather than a per-query serving trick. Existing routing methods mostly optimize single-turn cost-quality trade-offs and therefore miss the execution state, intermediate failures, and feedback loops that make agents different from chat completion. We propose Harness-Native agentic routing, a step-level routing paradigm that selects either a single best-fit model for cost-effective execution or multiple complementary models for ensemble-style accuracy improvement, conditioned on the full harness state. The key insight is that every routing decision naturally produces a structured data record -- consisting of the query, harness state, model choice or model set, execution trace, outcome, and cost -- whose labels are supplied by the environment rather than by the router itself. These records form a harness-native data flywheel: execution traces train better routers and harness-native models, which improve cost-quality trade-offs and generate more traces under the same budget. We instantiate this idea in OpenSquilla with a four-layer routing stack, an open LightGBM cold-start ranker, and a staged router-model path that turns logged arena records into progressively stronger routing policies. The report studies singleton and multi-model routing on agentic benchmarks including DRACO and PinchBench, and argues that agentic routing is not merely cost control, but a data engine for agent-native training.

cs.CL

HieDG: A Hierarchical Discrete Geometry-Guided Framework for Multi-Animal Tracking

Multi-animal tracking (MAT) is critical for wildlife monitoring and behavioral analysis, yet remains challenging due to uniform appearance, high density, and irregular motion. Existing methods typically follow heuristic- or query-based paradigms: the former relies on handcrafted geometric associations without end-to-end optimization, whereas the latter enables joint optimization but relies heavily on appearance embeddings. In such conditions, continuous geometric embeddings can be unstable, as small coordinate perturbations may disproportionately alter cross-frame attention weights, degrading identity association performance. To address this limitation, we propose HieDG, a Hierarchical Discrete Geometry-guided tracking framework that reformulates geometric dynamics as structured discrete representations within a query-based tracker. Instead of directly using raw geometric signals, HieDG employs a two-stage residual codebook to discretize position, scale, and velocity cues, transforming unstable continuous geometry into structured, stable discrete tokens. These tokens are aligned with visual embeddings and integrated into the tracking queries to enhance identity consistency. Extensive experiments on animal-specific benchmarks (AnimalTrack, BFT, and BuckTales) demonstrate state-of-the-art association performance with significant improvements in HOTA, AssA, and IDF1. Additional evaluations on generic multi-object tracking benchmarks, including DanceTrack and SportsMOT, show competitive performance, indicating the broader applicability of discretized geometric modeling beyond animal-specific scenarios.

cs.CV

READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations

Audio Description aims to generate concise narrations of essential visual content in audio-visual media for blind and low-vision audiences. Existing methods either rely on prompting off-the-shelf multimodal models, which often mismatch AD style, or partially optimize training-based systems with next-token prediction, which under-explores model capacity and biases generation toward generic expressions. We present READ, the first reinforcement-learning (RL) framework for training-based AD generation. READ formulates AD as sequence-level optimization with reference-matching, length, and format rewards, and further introduces a dedicated coherence reward under context-aware supervision to promote narratively coherent descriptions. Experiments on MAD-Eval, CMD-AD, and TV-AD show that READ substantially outperforms prior methods across diverse evaluation metrics. Our results highlight RL as a promising paradigm for accurate and coherent AD generation. Our codes, models, and benchmark results will be publicly available.

cs.CV

Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks

General-purpose agents such as OpenClaw are increasingly used as autonomous tool users, but their coding ability is difficult to measure under SWE-bench: a generic agent does not by itself satisfy the clean Docker workspace, patch, and prediction contract required for scoring. We introduce Claw-SWE-Bench, a multilingual SWE-bench-style benchmark and adapter protocol that makes heterogeneous agent harnesses, or claws, comparable under fair settings including a fixed prompt, runtime budget, workspace contract, patch extraction procedure, and evaluator. The full benchmark contains 350 GitHub issue-resolution instances across 8 languages and 43 repositories, drawn from SWE-bench-Multilingual and SWE-bench-Verified-Mini after future-commit cleanup. We also release Claw-SWE-Bench Lite for faster validation, which is an 80-instance subset selected by a cost-aware, rank-aware procedure over 17 calibration columns. On the full benchmark, OpenClaw with a minimal direct-diff adapter scores only $19.1\%$ Pass@1, whereas the full adapter reaches $73.4\%$ with the same GLM 5.1 backbone, showing that adapter design is essential for enabling OpenClaw-style harnesses to perform coding tasks effectively. Across an OpenClaw $\times$ nine-model sweep and a five-claw $\times$ two-model sweep, model choice changes Pass@1 by $29.4$ pp and harness choice by $27.4$ pp under fixed models; systems with similar accuracy can differ substantially in total API cost. Claw-SWE-Bench therefore treats harness and cost accounting as first-class axes of SWE-style coding-agent evaluation, providing both a full benchmark and a low-cost reference set for reproducible comparison. The data is available at https://github.com/opensquilla/claw-swe-bench and https://huggingface.co/datasets/TokenRhythm/Claw-SWE-Bench.

cs.LG

Group theory of Raman effect in magnetic materials

Despite the wealth of experimental observations on Raman scattering in magnetic materials, the underlying selection rules have remained largely unexplored. In this work, we use Onsager reciprocity relation, other than the conventional corepresentation method, to deal with the mathematical structures of Raman tensors in magnetic groups. Using this approach, we generate Raman tensor tables for all magnetic point groups, and present a comprehensive understanding of the Raman selection rules in magnetic materials with direct product representations method. Our theoretical and numerical results match previous experiments well, and resolve a recent puzzle in the Raman spectroscopy of CrSBr. Moreover, we identify a common but overlooked phenomenon: the magneto-Raman vector can be orthogonal to the magnetic moment direction. Our method and associated Raman tensor tables will be helpful for the Raman studies in both experimental and theoretical domains.

cond-mat.mtrl-sci

What Should a Skill Remember? Quality--Cost Trade-offs in Cost-Aware Skill Rewriting for Language Model Agents

Large language model agents increasingly rely on skills: reusable procedural documents encoding workflows, tool use, implementation patterns, validation checks, and domain rules. Skill rewriting is often treated as prompt compression, but shorter skills can make agents more expensive by removing sparse operational anchors that prevent exploration, debugging, and recovery. We study skill rewriting through this economic lens. Our controlled framework profiles skill structure, rewrites skills using information-preservation strategies, and evaluates the rewrites under fixed task instructions, environments, and verifiers. Experiments on SkillsBench reveal distinct quality--cost trade-offs across strategies: API/code anchoring, workflow guarding, and rule/formula anchoring benefit different task families, with no universally dominant template. In the main held-out evaluation, the learned policy reduces total cost by 7.0% and downstream agent-token cost by 6.0%; in frozen cross-model transfer, the corresponding reductions average 14.7% and 13.7%, while verifier quality is preserved. These results position skill design as cost-aware operational knowledge engineering rather than prompt compression. Resources: https://github.com/1Reminding/Skill_EE.

cs.CL

FAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language Models

Diffusion Large Language Models (dLLMs) refine tokens iteratively but commit them irreversibly, leading to a "stability lag" where early decisions remain fragile even after being written. We reveal that Post-Training Quantization (PTQ) error easily flips these borderline decisions at the write frontier, which are then permanently locked in and amplified. To address this, we propose Frontier-Aware Instability-Reweighted Calibration (FAIR-Calib), a two-stage PTQ framework for dLLMs. Stage I probes a full-precision teacher to estimate a position prior that combines frontier hits and masked-stage reliability. Stage II performs off-policy, layer-wise calibration by minimizing a reweighted hidden-state MSE, effectively prioritizing the protection of fragile frontier states without requiring expensive end-to-end diffusion rollouts. We further theoretically justify our weighted objective as a surrogate for output KL divergence. Empirically, FAIR-Calib consistently outperforms state-of-the-art baselines on LLaDA and Dream (W4A4), significantly reducing frontier decision flips and suppressing post-commit mismatches across diverse benchmarks.

cs.LG

Testing lepton-flavor-violating decay of doubly charged Higgs bosons in type-II seesaw via photon fusion at the high-energy LHC

Tiny neutrino masses can be explained by the type-II seesaw mechanism, where a triplet scalar under $SU(2)_{L}$ is predicted. Collider searches for this exotic scalar have been extensively conducted, especially for its doubly charged component $\Delta^{\pm\pm}$. Utilizing the forward detectors at the Large Hadron Collider (LHC), we study the probing sensitivity for the elastic photon fusion production of the scalars $pp\to p(\gamma\gamma\to\Delta^{++}\Delta^{--})p$ followed by the lepton-flavor-violating (LFV) decay channels $\Delta^{\pm\pm}\to e^{\pm}\mu^{\pm}$. With a high center-of-mass energy of 100 TeV and several luminosity scenarios, we can extensively broaden the exclusion bounds in the parametric space of Br$(\Delta^{\pm\pm}\to e^{\pm}\mu^{\pm})$ versus the triplet scalar mass $m_{\Delta}$. Specifically, at the 100 TeV LHC with an integrated luminosity of 3 ab$^{-1}$, the mass exclusion limit at 95\% C.L. can reach around 1150 GeV with the assumption of inverted neutrino mass hierarchy.

hep-ph

Harnessing hidden quantum metric response in a 2D magnet via nonlocal photovoltaic effect

The quantum geometry of Bloch wavefunctions underpins a wealth of emergent phenomena in quantum materials. Its imaginary part, the Berry curvature, has long been recognized as a key source for hallmark effects such as quantum Hall and topological phenomena, etc. The real part of quantum geometry, the quantum metric, has recently garnered considerable attention due to predictions of a range of unconventional nonlinear and nonequilibrium responses. Such responses usually vanish in centrosymmetric systems, largely restricting relevant studies to non-centrosymmetric materials. Here we challenge this convention by revealing that the vanished quantum metric response can survive in a hidden form. Using a non-local photovoltaic scheme in a layered magnetic semiconductor, we spatially separate mutually compensating photocurrents and thereby detect such hidden quantum metric response. We demonstrate this effect across distinct magnetic states and down to the ultrathin limit. Moreover, we realize reconfigurable, nonvolatile and probabilistic photodetection enabled by the quantum metric response. These results not only fundamentally expand the material landscape for quantum geometric physics, but also open new gateway to harvest the quantum geometric contributions for state-of-the-art nonvolatile reprogrammable sensing and computing applications.

cond-mat.mes-hall

A Novel Segment-Based Tracking Algorithm for HLT under High-Occupancy and Complex Conditions

In the High-Level Trigger (HLT) of both electron-positron and hadron collision experiments, the tracking process for large-volume gaseous detectors typically consumes a latency of hundreds of milliseconds. Upgrades of existing experiments and the development of next-generation facilities demand enhanced HLT tracking performance: handling higher detector occupancy and suppressing latency. To address high occupancy conditions, a novel HLT tracking algorithm based on track segments is proposed. This method involves constructing a pattern bank comprising 11 pre-defined patterns, optimizing edge-matrix formation using position, momentum, and timing criteria, and merging stereo superlayer segments to improve track consistency. These measures significantly reduce the number of stored segments and the size of the edge matrix, thereby lowering the complexity of global tracking. Even at 25\% occupancy, the number of elements for global tracking is reduced to approximately 400-500, while the density of the edge matrix remains below 1\%. With the depth-first search within connected components, the simulation results show that the algorithm maintains stable performance with occupancy ranging from 5\% to 25\%, achieving a data compression ratio of approximately 50\% to 70\%. Validation against the STCF offline reconstruction algorithm confirms that the HLT algorithm preserves high signal-hit retention without introducing significant adverse effects on offline tracking efficiency or on the reconstruction. These results demonstrate that the proposed algorithm can retain high-quality signal hits across a broad range of occupancy levels, indicating a strong potential for further development and adaptation to even more challenging, high-luminosity experimental conditions.

physics.ins-det

Liouvillian spectral control for fast charging of quantum batteries

Quantum batteries, which use quantum systems to store and deliver energy, are promising for next-generation energy storage. However, optimizing charging strategies and understanding the interplay between dissipation and quantum coherence remain open challenges. Here, we investigate steady-state charging in an open quantum battery and demonstrate that the charging timescale depends on the spectral gap of the Liouvillian operator governing dissipative dynamics. As a minimal example, we examine a three-level quantum battery realized in a single trapped ${}^{40}\mathrm{Ca}^{+}$ ion, where energy from an engineered thermal photon reservoir is coherently transferred to a long-lived metastable storage state. We find that long-term dynamics are confined to a low-dimensional manifold of slow Liouvillian modes, with their spectral structure determining the relaxation rate to the charged steady state. By adjusting experimentally accessible parameters, such as reservoir occupation and coherent coupling strength, the non-Hermitian Liouvillian spectrum can approach an exceptional point. This increases the spectral gap and accelerates the approach to steady state. As a result, this mechanism significantly enhances asymptotic charging power without relying on many-body collectivity or steady coherence. Our findings offer fundamental insights into open quantum thermodynamics and provide a path to efficient energy storage and fast-charging solutions in emerging quantum technologies.

quant-ph

A Spatially Resolved HI Survey of Seyfert Galaxies: the Role of AGN Feedback in Shaping Atomic Gas Reservoirs

Active galactic nucleus (AGN) feedback is a key ingredient in galaxy evolution, yet its impact on the cold atomic gas reservoir -- the neutral hydrogen (HI) phase -- remains poorly constrained. We present the most extensive spatially resolved HI 21-cm survey of Seyfert AGN hosts to date, based on observations with the Giant Metrewave Radio Telescope (GMRT). Our high-resolution HI maps of eight Seyfert galaxies reveal detailed kinematics and surface density distributions of their atomic gas disks. We find that AGN-host galaxies exhibit a slightly shallower HI mass-size relation than the canonical relation or the SIMBA simulation predictions; however, the measured slope remains consistent with the canonical value within $2\sigma$ uncertainties. This result suggests that AGN feedback does not significantly disrupt the global extent or large-scale structure of atomic gas reservoirs. To investigate the internal HI kinematics in greater detail, we perform a 3D kinematic forward modeling of the HI disk in UGC 4503. Our analysis reveals an elevated intrinsic velocity dispersion of $\sigma = 14.9^{+6.1}_{-3.8}$ km/s and a reduced level of rotational support, with $V/\sigma = 14.28_{-4.17}^{+4.97}$, compared to large-sample star-forming spirals. These kinematic signatures, together with localized residuals in the velocity field, indicate that AGN-driven outflows or jets may inject or indirectly affect the turbulence in the atomic gas disk, potentially regulating the cold gas reservoir. Future GMRT observations, combined with optical integral-field spectroscopy from MaNGA, will enable quantitative constraints on the role of AGN feedback in regulating star formation efficiency across a larger and more representative galaxy sample.

astro-ph.GA