arXiv ScienceSearch

arXiv subjects

Yu Cao

Publications and source records attributed to Yu Cao.

At least 19 recordsLinked to original sources

Reflex-Informed Neuromuscular Reinforcement Learning for Muscle-Driven Locomotion

Muscle-driven locomotion provides a physically grounded approach to generating realistic human movement. However, achieving both physiological plausibility and adaptability to changes in musculoskeletal capacity and external disturbances remains a fundamental challenge. To address this limitation, we propose a Reflex-Informed Neuromuscular Reinforcement Learning framework for muscle-driven locomotion. Within this framework, a fixed phase-dependent reflex controller serves as the underlying neuromuscular control mechanism, while the reinforcement learning policy produces four biomechanically meaningful residual parameters to modulate key reflex gains and thresholds associated with hip swing, knee support, and ankle propulsion according to the current state. Experimental results demonstrate that the proposed framework generates physiologically plausible locomotion with improved kinematic accuracy and dynamic consistency, as well as better bilateral symmetry and stride-to-stride consistency under nominal walking conditions. The learned policy remains robust under muscle weakness and external perturbations without retraining.

cs.RO

Counter-examples for Tensorization Property of Strong Data Processing Inequality for Quantum Divergences

The data processing inequality is a fundamental property that describes the loss of information through noisy channels. A more refined description is characterized by the strong data processing inequality (SDPI). In classical information theory, the tensorization of strong data processing inequality holds for a whole family of $f$-divergences. However, its quantum counterpart is less known. The tensorization of SDPI was shown only for some special cases previously, and the general understanding about the tensorization property of SDPI for quantum divergences remains open. In this work, we report two negative results: the tensorization property fails for certain quantum chi-square divergences, and it also does not hold for the quantum relative entropy.

quant-ph

LACE: Large Language Model Aided Multi-Agent Framework for Agile RISC-V Instruction Extension

Domain-specific Instruction Set Architecture eXtensions (ISAX) are widely adopted in the RISC-V ecosystem to accelerate emerging workloads, but implementing and validating ISAXes across different cores remains slow and fragmented. Existing frameworks still require per-core interface adaptation, and differential testing often breaks once either the microarchitecture or the ISAX changes. We present LACE, an LLM-aided multi-agent workflow that translates natural-language ISAX intents into a compact two-level IR (operation-level and HDL task-level), performs retrieval-guided localized RTL edits over large repositories, and closes the loop with a compiler-agnostic riscv-formal checking flow (assuming RVFI availability or instrumentation). Across four embedded RISC-V cores, LACE raises pass@1 generation accuracy from near-zero to 72.8\% within our evaluation setup, while improving code localization and reducing integration rework. The code of LACE is available at https://github.com/UMN-ZhaoLab/LACE.

cs.AR

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning

Test-time reinforcement learning (TTRL) enables language models to self-evolve at inference time without labeled feedback. Existing methods rely on answer voting and therefore do not extend naturally to open-ended generation, where valid responses cannot be mapped to a shared canonical answer. Without external reward models or stronger judges, adaptation must instead construct reliable rewards from the model's own outputs. We introduce SERPO (Self-Evolving Rubric Policy Optimization), which replaces answer voting with a closed loop that co-evolves response evidence, query-specific rubrics, and policy parameters. Good-Normal-Bad (G-N-B) response evolution organizes maximally separated rollouts into ordered archives; rubric evolution retains criteria that discriminate these archives; probabilistic criterion scoring converts verdict-token likelihoods into reward signals; and policy evolution optimizes the actor with the resulting signals. New actor rollouts then refresh both the archives and rubrics, closing the three-way evolution loop. Across two model configurations, two in-domain benchmarks, and four OOD benchmarks, SERPO improves HealthBench and ResearchQA by up to 20.63 and 20.31 points over the corresponding base models, raises the six-benchmark macro-average by up to 8.06 points, and supports OOD transfer and continued cross-benchmark evolution.

cs.CL

From Compressing Complexity to Accommodating Complexity: How AI Transforms Standardization and Individualization

Why do societies composed of individuals pursuing individuality repeatedly generate highly standardized systems? This paper argues that the answer lies in the evolution of information processing capacity. Artificial intelligence represents a historical transition in this capacity, enabling social systems to accommodate forms of complexity that previously had to be compressed. Industrial standardization was not merely a consequence of capital preference or power relations, but an institutional arrangement for maintaining the manageability of large-scale systems under limited information-processing capacity by reducing the variety of the controlled system. The fundamental change in the AI era lies in the expansion of information processing capacity across three dimensions: perception, computation, and execution. This expansion shifts personalized production from physical adaptation toward information-based adaptation and enables a transition from discrete to continuous objectification of difference. This paper proposes "cognitive fixed cost" as an analytical concept to describe how the upfront concentration of cognitive labor transforms the cost structure of personalized production. It further argues that standardization has not disappeared but has moved from explicit constraints at the product level to implicit generation rules embedded in infrastructures, shifting the central contradiction from "whether to have commonality" to "who controls commonality." The evolution of civilizational production logic is not a movement from commonality to individuality, but from compressing complexity to accommodating complexity.

cs.CY

Citation Pathways in the AI Era: Interpretive Knowledge Nodes, Citation Compression Layers, and the Measurement Boundary of Scholarly Impact

This article identifies "citation pathway" as a long-neglected analytical dimension in scientometrics. Traditional evaluation metrics focus on measuring citation counts while paying insufficient attention to the intermediate nodes through which knowledge flows from its original source to the citing author. Building on an analysis of the normative structure of current reference systems, this article introduces two new concepts: Interpretive Knowledge Nodes (IKN) - academic papers that provide structured reorganizations of classic works - and Citation Compression Layers (CCL) - the intermediate layers that emerge when such knowledge products acquire stable publication identities and enter formal citation networks at scale. The central proposition is that AI has not changed citation rules themselves but has transformed the cost structure of producing citable knowledge intermediaries. Under conditions of full compliance, the network position effects of citation pathways may become a salient variable affecting the validity of impact measurement. Through a thought experiment involving a hypothetical journal R and a parsimonious "Citation Gravity conceptual model," this article substantiates this proposition and discusses the measurement boundaries of scholarly impact indicators, as well as the institutional risk of incentive misalignment under extreme scenarios.

physics.soc-ph

Correlation-consistent Gaussian basis sets for copper solids from material-constrained atomic optimization

Correlation-consistent Gaussian basis sets are central to systematic molecular quantum chemistry, but their direct use in periodic solids is often limited by severe linear dependence from diffuse atomically optimized primitives. This problem is particularly acute for metallic and metal-containing systems, where reliable complete-basis-set (CBS) extrapolation is needed for correlated-wavefunction benchmarks. We introduce material-constrained atomic optimization (MCAO), a basis-set optimization framework that preserves the atomic and correlation-consistent character of Gaussian basis sets while penalizing large overlap-matrix condition numbers in representative solids. As a proof of concept, we generate Dunning-style MCAO-cc-pVXZ basis sets (X = D, T, Q) for Cu with all-electron, scalar-relativistic all-electron, effective core potential (ECP), and pseudopotential treatments. The resulting basis sets remain numerically stable for Cu solids and surfaces while reproducing molecular Cu dimer energetics and plane-wave reference properties of bulk Cu. CBS-extrapolated random-phase approximation calculations further enable a controlled assessment of pseudopotential, relativistic, and basis-set errors in bulk Cu and CO adsorption on Cu(111), providing scalar-relativistic all-electron Gaussian-basis benchmarks for the CO adsorption puzzle.

physics.chem-ph

Efficient Bethe-Salpeter Equation Calculations Based on Numerical Atomic Orbitals and Norm-Conserving Pseudopotentials: Dual-${\boldsymbol k}$-Mesh Strategy

We present an efficient implementation of the Bethe--Salpeter equation (BSE) based on numerical atomic orbitals (NAOs) and norm-conserving pseudopotentials within the ABACUS+LibRPA framework. By exploiting the localized resolution-of-identity (LRI) technique, the screened Coulomb interaction is cast into a real-space, unit-cell-indexed form $W_{\mu\nu}(\boldsymbol R)$ that is inherently short-ranged and well localized. This spatial locality enables an efficient Fourier interpolation of the BSE kernel from the coarse $\boldsymbol k$-mesh used in the preceding $GW$ calculation to an arbitrarily dense $\boldsymbol k$-mesh on which the BSE Hamiltonian is assembled and diagonalized, thereby giving rise naturally to a dual-$\boldsymbol k$-mesh workflow. Building on this scheme, we systematically examine the convergence of the absorption spectra with respect to the NAO basis set, the auxiliary basis set, and the $\boldsymbol k$-point sampling. Benchmark calculations for both molecular and periodic systems collectively validate the accuracy of the present implementation and establish the dual-$\boldsymbol k$-mesh strategy as a practical and reliable approach for $GW$+BSE calculations.

cond-mat.mtrl-sci

Each Judge Its Own Yardstick: Discovering Per-VLM Taxonomies for Physical Video Evaluation

Maintaining physical consistency in video generators and world models increasingly relies on vision-language models (VLMs) as automated judges that provide reward signals, ranking decisions, and data-filtering criteria. Yet VLMs differ substantially in training data and architecture, encoding physical phenomena through distinct internal representations. A single global evaluation schema therefore gives every VLM the same axes of competence, regardless of what each can actually perceive. We propose JudgeFit, an iterative refinement procedure that discovers a per-VLM evaluation taxonomy. An initial taxonomy is constructed by prompting the target VLM to enumerate physics errors on a small set of videos and clustering the resulting descriptions. The taxonomy is then refined through a diagnostic step: we calibrate the VLM's per-dimension scores to human physical-commonsense ratings, diagnose which dimensions it scores unreliably or redundantly, and prompt an LLM to repair them, iterating until convergence. We further instantiate this procedure as a benchmark and apply it to 16 VLMs spanning eight model families. The refined taxonomy outperforms the global-schema baseline on held-out videos for every VLM tested, with a mean relative improvement of approximately 32%. Beyond aggregate accuracy, the per-VLM profiles expose model-specific blind spots that overall rankings cannot anticipate, with reliability patterns differing markedly across model families.

cs.CV

Grade Encoding and the Structural Representation of Student Academic Trajectories

Does the conversion of academic assessment from percentage scores to letter grades merely represent an adjustment in information precision, or does it systematically alter the underlying structure of student academic data? Drawing on 68 mathematics exam scores from 75 primary school students, this study employs Encoding Transformation simulation to compare structural differences in the same dataset under two encoding schemes across three dimensions: information loss, distance structure change, and clustering stability. Results indicate that letter-grade encoding compresses the mean pairwise distance in the trajectory feature space from 20.50 to 1.06 (a compression ratio of approximately 19:1); after standardization, the density gradient of the distance distribution is systematically flattened, with kurtosis decreasing by 0.54 and the coefficient of variation decreasing by 0.16; and the clustering structure becomes highly sensitive to minor sample perturbations-removing a single extreme observation causes the optimal K to jump from 4 to 8 and clustering reproducibility to plummet from 95% to 62%. Concurrently, a Compression Paradox is observed: letter-grade encoding increases the Silhouette coefficient while simultaneously reducing clustering stability. These findings demonstrate that letter-grade encoding systematically alters the structural representation of longitudinal academic data and exerts measurable effects on clustering stability.

cs.CY

LiteVSR: Lightweight Adaptation of Frozen Diffusion Transformers for Video Super-Resolution

Adapting large-scale pre-trained video generators for Video Super-Resolution (VSR) in novel domains remains computationally prohibitive. Methods that reformulate generation as direct Low-Quality to High-Quality mappings deviate from the original generative formulation, demanding extensive fine-tuning. ControlNet-style adapters lose their efficiency under modern Diffusion Transformers since the absence of encoder-decoder hierarchy forces duplication of the entire backbone. We observe that flow matching offers a principled alternative for cross-domain VSR adaptation. By predicting a constant velocity field across all timesteps, the adaptation task reduces to learning a fixed injection pattern rather than time-varying transformations. Building on this insight, we propose LiteVSR, a minimalist framework that performs VSR using a completely frozen Diffusion Transformer with a lightweight State-Aware Adapter. The adapter employs a dual-stream architecture that extracts static structural cues from the LQ input and dynamic cues from intermediate denoising states, aligning them through time-dependent cross-attention to enable adaptive transition from structural alignment to texture refinement as denoising proceeds. LiteVSR achieves competitive restoration quality with only 11.25% trainable parameters and 12 GPU-hours of training on a single A100, while maintaining fast sampling (down to a single step) compatibility.

cs.CV

Awareness of Technological Isomorphism: Integrating AI into Elementary Mathematics Teaching on Data and Prediction,A Case Study of the Compound Line Graph

The deep integration of Artificial Intelligence (AI) into elementary mathematics education necessitates a conceptual tool capable of explaining students' cognitive transition from disciplinary knowledge to AI understanding. This study proposes a novel core concept, "Awareness of Technological Isomorphism, " defined as a student's metacognitive realization that their own mathematical cognitive operations (e.g., observing trends, inducing patterns, and making predictions) share an underlying logical structure with AI technical operations (e.g., pattern recognition and predictive modeling). This awareness, in turn, facilitates cognitive transfer from disciplinary mathematics to AI comprehension. Underpinned by transfer learning and metacognitive theories, this study clarifies the distinct essence of this concept from traditional "computational thinking." We demonstrate the explanatory power of this framework in two ways: elucidating the mechanism of students' cognitive leap from mathematics to AI, and guiding instructors to identify "isomorphic interfaces" within disciplinary curricula. On this basis, a three-stage pedagogical pathway--spanning "Perception, Comprehension, and Creation"--is constructed alongside a corresponding evaluation rubric. This framework is empirically validated through a case study based on the "Compound Line Graph" lesson from a fifth-grade mathematics textbook in China, offering a highly replicable operational framework for the deep convergence of disciplinary instruction and AI literacy education.

cs.CY

Bridging the Last Mile of Circuit Design: PostEDA-Bench, a Hierarchical Benchmark for PPA Convergence and DRC Fixing

LLM-based agents are increasingly applied to the "last mile" of Electronic Design Automation (EDA): repairing residual sign-off Design Rule Check (DRC) violations and converging Power-Performance-Area (PPA) targets after tool runs. Existing EDA-LLM benchmarks, however, omit DRC fixing entirely and rely on flat hierarchies tied to a single toolchain. We introduce PostEDA-Bench, a hierarchical benchmark with 145 tasks across DRC-Essential, DRC-Reasoning, PPA-Mono, and PPA-Multi, supported by EDA toolchains with machine-checkable evaluation. Across eight commercial and open-source LLMs under multiple agent scaffolds, we find that agents handle synthetic DRC-Essential and single-objective PPA-Mono reasonably well but degrade sharply on the more practical DRC-Reasoning, where the best success rate is 36.66%, and PPA-Multi, where the best success rate is 20.00%; vision augmentation consistently enhances DRC-Bench; and trade-off reasoning, rather than knob knowledge, is the dominant PPA-Multi bottleneck.

cs.AR

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate

On-policy distillation (OPD) trains a student on its own trajectories under token-level teacher supervision, but existing methods are capped by a single-teacher capability ceiling: when the teacher errs, the student inherits the error. OPD also remains largely unexplored in agentic tasks, where per-step errors compound across long trajectories and destabilize training. We propose MAD-OPD (Multi-Agent Debate-driven On-Policy Distillation), which breaks this ceiling by recasting the distillation teacher as a deliberative collective of teachers that debate over the student's on-policy state; the debate produces an emergent collective intelligence that supplies token-level supervision, with each teacher's contribution weighted by its post-debate confidence. To extend OPD to agentic tasks, we also introduce On-Policy Agentic Distillation (OPAD), which adds step-level sampling to stabilize training under multi-step error compounding. We additionally derive a task-adaptive divergence principle, selecting JSD (Jensen-Shannon divergence) for agentic stability and reverse KL (Kullback-Leibler) divergence for code generation, and verify it both theoretically and empirically. Across six teacher-student configurations (Qwen3 and Qwen3.5; 1.7B-14B students, 8B-32B teachers) and five agentic and code benchmarks, MAD-OPD ranks first across all six configurations; on the 14B+8B$\to$4B setting it lifts the agentic average by $+2.4\%$ and the code average by $+3.7\%$ over the stronger single-teacher OPD.

cs.CL

Symplectic connection third-order Hall effect in a room-temperature ferromagnet

Third-order nonlinear Hall effects (THE) have recently attracted considerable experimental interest as powerful probes for quantum geometric properties in emergent quantum materials, encompassing quadrupole moments of quantum metric and Berry curvature. Here, we report a fundamentally new THE in room-temperature van der Waals ferromagnet Fe3GaTe2 from second-order Berry connection polarizability, which manifests a higher-order characterization of band geometry called symplectic connection. Our observations show that the third-order transverse response in Fe3GaTe2 is odd to magnetization, vanishes above the Curie temperature and remains independent of driving current directions. Scaling law analysis combined with first-principles calculations establishes this response as the symplectic-connection-induced THE. This discovery opens the door to probing high-order quantum geometric properties beyond Berry curvature and quantum metric through nonlinear transport, unveiling the potential of exploring nonlinear Hall phenomena in broad classes of magnets without breaking inversion symmetry. Moreover, the room-temperature manipulation of THE holds promises for device applications based on harnessing the quantum-geometric connection structure.

cond-mat.mes-hall

LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion

The recent success of inference-time scaling in large language models has inspired similar explorations in video diffusion. In particular, motivated by the existence of "golden noise" that enhances video quality, prior work has attempted to improve inference by optimising or searching for better initial noise. However, these approaches have notable limitations: they either rely on priors imposed at the beginning of noise sampling or on rewards evaluated only on the denoised and decoded videos. This leads to error accumulation, delayed and sparse reward signals, and prohibitive computational cost, which prevents the use of stronger search algorithms. Crucially, stronger search algorithms are precisely what could unlock substantial gains in controllability, sample efficiency and generation quality for video diffusion, provided their computational cost can be reduced. To fill in this gap, we enable efficient inference-time scaling for video diffusion through latent reward guidance, which provides intermediate, informative and efficient feedback along the denoising trajectory. We introduce a latent reward model that scores partially denoised latents at arbitrary timesteps with respect to visual quality, motion quality, and text alignment. Building on this model, we propose LatSearch, a novel inference-time search mechanism that performs Reward-Guided Resampling and Pruning (RGRP). In the resampling stage, candidates are sampled according to reward-normalised probabilities to reduce over-reliance on the reward model. In the pruning stage, applied at the final scheduled step, only the candidate with the highest cumulative reward is retained, improving both quality and efficiency. We evaluate LatSearch on the VBench-2.0 benchmark and demonstrate that it consistently improves video generation across multiple evaluation dimensions compared to the baseline Wan2.1 model.

cs.CV

Optimized growth of large-size, high quality $\text{ZrTe}_5$ single crystals enabling clear quantum oscillations in electrical transport

Quantum oscillation with nontrivial Berry phase is one of the characteristics of topological materials. As a Dirac semimetal candidate, zirconium pentatelluride ($\text{ZrTe}_5$) stands out as an intriguing material for investigating topological phase transitions and Dirac fermion physics; however, the extreme sensitivity of its electronic properties to stoichiometric variations and crystalline defects has hindered consistent experimental observation. Here, we report an optimized Te-flux synthesis method designed to produce centimeter-scale, high-quality single crystals meanwhile minimizing extrinsic carrier contamination. Comprehensive morphology, structural and chemical characterizations, including scanning electron microscopy, Laue backscattering and Rietveld refinement, confirm a high-purity $Cmcm$ phase with excellent crystallinity. Furthermore, magnetotransport measurements reveal a remarkably low Shubnikov-de Haas oscillation onset field ($B_{int} \approx 0.38$ T) with an ultra-high mobility of $5.58\times10^5$cm$^2$V$^{-1}$s$^{-1}$ and access to the the quantum limit at $B \approx 1.3$ T, attesting to the superior crystalline quality and the efficacy of this growth optimization. These results demonstrate that growth control is crucial for stabilizing intrinsic electronic behavior in $\text{ZrTe}_5$, establishing a robust platform for exploring topological phase transitions and exotic quantum phenomena in topological semimetals.

cond-mat.mtrl-sci

ParkingTwin: Training-Free Streaming 3D Reconstruction for Parking-Lot Digital Twins

High-fidelity parking-lot digital twins provide essential priors for path planning, collision checking, and perception validation in Automated Valet Parking (AVP). Yet robot-oriented reconstruction faces a trilemma: sparse forward-facing views cause weak parallax and ill-posed geometry; dynamic occlusions and extreme lighting hinder stable texture fusion; and neural rendering typically needs expensive offline optimization, violating edge-side streaming constraints. We propose ParkingTwin, a training-free, lightweight system for online streaming 3D reconstruction. First, OSM-prior-driven geometric construction uses OpenStreetMap semantic topology to directly generate a metric-consistent TSDF, replacing blind geometric search with deterministic mapping and avoiding costly optimization. Second, geometry-aware dynamic filtering employs a quad-modal constraint field (normal/height/depth consistency) to reject moving vehicles and transient occlusions in real time. Third, illumination-robust fusion in CIELAB decouples luminance and chromaticity via adaptive L-channel weighting and depth-gradient suppression, reducing seams under abrupt lighting changes. ParkingTwin runs at 30+ FPS on an entry-level GTX 1660. On a 68,000 m^2 real-world dataset, it achieves SSIM 0.87 (+16.0%), delivers about 15x end-to-end speedup, and reduces GPU memory by 83.3% compared with state-of-the-art 3D Gaussian Splatting (3DGS) that typically requires high-end GPUs (RTX 4090D). The system outputs explicit triangle meshes compatible with Unity/Unreal digital-twin pipelines. Project page: https://mihoutao-liu.github.io/ParkingTwin/

cs.CV