arXiv ScienceSearch

arXiv subjects

Duo Wang

Publications and source records attributed to Duo Wang.

At least 19 recordsLinked to original sources

N\'eel-Vector-Dependent Altermagnetic Spin Splitting in the One-Dimensional Limit

Altermagnetism is normally identified by nonrelativistic spin-split bands in a compensated collinear magnet. Under one-dimensional confinement, however, this diagnostic can disappear: a boundary-compatible sublattice-exchange operation may leave the only Bloch momentum unchanged, forcing the nonrelativistic spin-up and spin-down spectra to coincide. We show that nonrelativistic spin degeneracy can coexist with relativistic spin splitting in a compensated one-dimensional magnet. In a Lieb-based construction, confinement along the diagonal cancels the projected $d$-wave spin splitting while preserving the real-space sublattice-exchange motif. Spin-orbit coupling (SOC) then locks spin to the lattice, so the N\'eel-vector orientation selects a magnetic line group that can reveal spin-split bands. Using fully compensated $[110]$ Ta$_2$TeSeO nanoribbons as a prototype, we find spin-degenerate nonrelativistic bands and sizable SOC splitting in the inherited easy-axis domain, $\mathbf N\parallel[100]$ or $[010]$. For sufficiently wide ribbons that retain the parent easy-axis order, this splitting is an equilibrium property. The easy-axis ribbon also supports a right-moving-mode spin polarization of about $35\%$ at finite ideal conductance. Our results establish one-dimensional fully compensated altermagnetic functionality hidden inside apparently conventional antiferromagnetic bands.

cond-mat.mes-hall

xDailyBench: Benchmarking LLMs on Professional Consultation for Real-Life Problems

Large language models (LLMs) are increasingly used for everyday assistance, yet existing benchmarks only partially reflect the requests users naturally make in practice. Real-world requests are often open-ended, casually specified, and context-dependent, requiring models not only to follow explicit instructions but also to infer unstated needs from user background and situational context. We introduce xDailyBench, a benchmark of 248 carefully curated tasks spanning 51 scenarios across personal life, white-collar work, learning and research, and cross-domain activities. The tasks are grounded in requests that users have actually completed or genuinely intended to accomplish with AI, and are evaluated with fine-grained binary rubrics covering both explicit and implicit requirements. We evaluate 11 frontier models under standardized agentic settings. The best models achieve a task-level score of 75.6\%, while all models perform substantially worse on implicit than explicit requirements, with gaps no less than 9 percentage points. These results reveal implicit requirement inference as a persistent bottleneck for reliably satisfying real-world everyday user needs.

cs.AI

Aspire: Can Models Self-Evolve from Vague Goals?

Many important forms of human learning begin with a vague goal, such as "become a better physicist" or "improve at research." Learners must interpret the goal, identify capability gaps, decide how to learn, and determine whether they have actually improved. In contrast, existing work on LLM self-evolution typically begins with tasks and evaluation metrics specified by humans, reducing self-evolution to optimizing an explicit objective rather than deciding what and how to learn. We introduce ASPIRE, a benchmark for vague-goal-driven self-evolution. ASPIRE provides only a natural-language capability goal while downstream evaluation tasks remain hidden. The agent must operationalize the goal by choosing data and update methods, constructing training and validation signals, and deciding when to evaluate. ASPIRE supports both model-weight and agent-harness evolution in a unified interactive environment and evaluates the resulting systems on a hidden, expert-authored set of 520 items spanning six goals. Our experiments show that vague goals redirect search effort toward goal interpretation. Current agents routinely complete training and harness-editing loops, but weight-level gains remain sparse and unstable, and the strongest evolved harness remains below the engineered Qwen-Agent reference. Agents often train on mismatched data and trust narrow self-evaluations, so local gains fail to transfer to hidden evaluation and continued search and training can erase earlier improvements.

cs.CL

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually demand from AI systems. We introduce \textbf{StartupBench}, an E2E agent benchmark grounded in market-validated AI startup products. Rather than defining tasks from pre-defined assumptions about useful agent capabilities, we systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains. We translate these workflows into complete deliverable-oriented tasks and evaluate them with fine-grained rubrics capturing their complex requirements. Across representative models evaluated under a unified agent harness, even the strongest model successfully completes only approximately 30\% of StartupBench, despite making substantial partial progress on many tasks. Further analysis identifies aspects like complex instruction following and domain-specific expertise as major sources of failure. Our results reveal that many market-validated workflows remain beyond the reliable capabilities of current general-purpose agents, establishing StartupBench as an empirical measure of progress toward E2E completions of real-world user tasks.

cs.AI

ESR-HGNN: Eliminating Semantic Redundancy for Efficient Mini-batch HGNN Inference

Heterogeneous graph neural networks (HGNNs) are highly effective in processing heterogeneous graph data and have been widely adopted in critical domains. As real-world graph data continues to scale, performing direct inference on entire graphs becomes increasingly infeasible, making mini-batch methods the standard approach. However, in end-to-end HGNN inference, metapath-based mini-batch sampling constitutes a significant performance bottleneck due to the extensive random memory accesses induced by the irregular traversal of graph structures. Existing sampling paradigms suffer from excessive redundant traversals caused by inherent semantic redundancy, severely degrading sampling efficiency and, consequently, leading to suboptimal mini-batch inference performance. In this work, we propose a redundancy-aware HGNN sampling paradigm that leverages a metapath trie to reuse traversal paths, effectively eliminating redundant memory accesses. We then map it onto a multi-channel hardware sampling unit denominated ESR-HGNN. Furthermore, we introduce a reusability-driven metapath grouping technique that optimally clusters metapaths to maximize reusable traversal paths within hardware channels, enhancing efficiency in scenarios with semantic parallelism. Extensive experimental results demonstrate that ESR-HGNN achieves an average sampling performance improvement of one order of magnitude over CPU and GPU, accompanied by significant energy savings. Additionally, it delivers substantial speedup in end-to-end mini-batch inference when integrated with GPU and state-of-the-art HGNN inference accelerator.

cs.AR

Real-Time Design of Public Transport Lines: Reconciling Adaptivity and Efficiency

Demand-responsive transport (DRT) is typically routed by solving Dynamic Vehicle Routing Problems (DVRPs), where individual vehicle trajectories are adjusted on incoming requests. This limits demand consolidation and thus efficiency. On the other hand, Conventional Public Transport (CPT) bus systems are based on a network of lines and users find their routes on it, which provides high demand consolidation. However, such a network is built offline and cannot adapt to the demand. We propose a public transport management strategy that reconciles efficiency and adaptivity by dynamically designing a structured network of lines via a receding-horizon optimization approach. Using real-world trip requests, we show that we nearly double the fraction of served requests compared to DVRP-based routing, and we serve more requests than CPT with lower user trip times.

eess.SY

Direct Numerical Simulation of Fully Developed Turbulent Channel Flow Based on the Corrected Navier-Stokes Equations

Direct numerical simulations (DNS) of fully developed turbulent channel flows at Re_tau = 550 were performed to investigate the corrected Navier-Stokes (CNS) equations. Grounded in the fluid kinematics of Rortex, the CNS abandons the Stokes' isotropic hypothesis and applies the shearing-only constitutive relation by explicitly eliminating the controversial stretching terms in the stress tensor. Comparisons with the DNS data from the traditional Navier-Stokes (TNS) suggest that the CNS inherently rectifies the near-wall momentum transport. The removal of stretching-induced dissipation shifts the inner- and buffer-layer boundaries towards the wall and effectively suppresses the overshoot in the mean velocity profile in TNS. The turbulence statistics demonstrate a multiscale kinetic energy redistribution that intensifies the near-wall production-dissipation cycle. Furthermore, the topological delineation of instantaneous coherent structures, namely the discovery and definition of rotational/non-rotational interface (RNRI) via the velocity gradient tensor (VGT) discriminant (Delta = 0), confirms that the CNS is capable of capturing the highly complex and interwoven vortical structures. Ultimately, the spectral proper orthogonal decomposition (SPOD) unveils and elucidates that the shearing-only mechanisms intrinsically modulate the spatiotemporal energy cascade, promoting denser and more inclined vortices while enhancing turbulence intermittency by fragmenting the coherent packets. Overall, by isolating and detecting the shearing-only mechanism with physics purity, the DNS based on the CNS provides a more refined perception of the intrinsic dynamics of wall-bounded turbulence, offering a more physical soundness model to capture the interactions among the multiscale coherent structures and the improved capability in predicting the wall-bounded turbulence.

physics.flu-dyn

UME: A Unified Meta-Generalization Framework for Cross-Domain ETA

Accurate Estimated Time of Arrival (ETA) prediction on checkout page is crucial in instant logistics for enhancing user satisfaction, optimizing dispatching, and controlling operational costs. In international on-demand delivery platforms, where ETA data originates from diverse countries or regions with different patterns, multi-domain modeling is of great importance and has been widely adopted. However, existing methods still face three critical challenges in real-world deployment. First, current multi-domain models struggle to generalize to completely unseen domains, failing to achieve zero-shot prediction during the initial cold-start phase. Second, cross-domain feature spaces are often assumed to be consistent, whereas new domains commonly suffer from structural missingness of offline (statistical) features due to the lack of historical data. Third, such feature missingness often compels industrial systems to model mature and cold-start domains separately, hindering knowledge transfer and increasing maintenance overhead. To address these challenges, we propose \textbf{UME}, a \textbf{U}nified \textbf{M}eta-generalization framework for \textbf{E}TA. Specifically, UME integrates a unified dual-branch architecture with a novel meta-learning mechanism that employs a hypernetwork-based meta learner. By leveraging domain-level knowledge and instance-level context, the meta learner empowers three meta modules to dynamically modulate feature gating, expert attention, and final prediction, capturing cross-domain correlations and facilitating intra-domain adaptation. A knowledge distillation strategy is further introduce to enhance performance. UME has now been deployed in Meituan-keeta delivery platform (the largest international food delivery platform in China). Extensive offline experiments and online A/B tests demonstrate that UME significantly outperforms existing baselines.

cs.LG

Uni-LaViRA: Language-Vision-Robot Actions Translation for Unified Embodied Navigation

Embodied navigation requires an agent to map language and visual observations to a stream of spatial actions that drive a real robot through environments it has never seen. The dominant approach has been to scale vision-language-action (VLA) foundation models on ever-larger collections of robot trajectories. This paper argues that, for navigation specifically, generality can be obtained structurally, not only through data scale. The underlying decision structure of navigation reduces to a single Language-Vision-Robot Actions Translation. The language action emits semantic-level directional command and the vision action emits a pixel-level visual target. Both outputs lie inside the natural output manifold of pretrained multimodal large language models (MLLMs), so the task can be reasoned about by an agent rather than learned from robot data. Therefore, we present Uni-LaViRA, a unified agentic architecture that extends the same insight to four task families (VLN-CE, ObjectNav, EQA, and Aerial-VLN) and to four heterogeneous real robots (Wheeled, Quadruped, Humanoid robot, and a self-built UAV) in a zero-shot manner. Two agent-loop mechanisms make this unification practical. TODO List Memory (TDM) rewrites a structured checklist of pending sub-goals at every step, reciting the unfinished items back into the agent's most recent attention window. Second Chance Backtrack (SCB) rolls the robot back to the pre-error state and conditions the agent's next plan on the failed sub-trajectory, turning single-pass navigation into a self-correcting process. With zero training effort, Uni-LaViRA reaches 60.7% SR on VLN-CE R2R, 51.3% on VLN-CE RxR, 77.7% on HM3D-v2, 60.0% on HM3D-OVON, 54.7% on MP3D-EQA, and 40.0% on OpenUAV, matching or even surpassing recent training navigation foundation models that consume millions of samples and thousands of GPU-hours.

cs.RO

Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation

As Large Language Models (LLMs) exhibit plateauing performance on conventional benchmarks, a pivotal challenge persists: evaluating their proficiency in complex, open-ended tasks characterizing genuine expert-level cognition. Existing frameworks suffer from narrow domain coverage, reliance on generalist tasks, or self-evaluation biases. To bridge this gap, we present XpertBench, a high-fidelity benchmark engineered to assess LLMs across authentic professional domains. XpertBench consists of 1,346 meticulously curated tasks across 80 categories, spanning finance, healthcare, legal services, education, and dual-track research (STEM and Humanities). These tasks are derived from over 1,000 submissions by domain experts--including researchers from elite institutions and practitioners with extensive clinical or industrial experience--ensuring superior ecological validity. Each task uses detailed rubrics with mostly 15-40 weighted checkpoints to assess professional rigor. To facilitate scalable yet human-aligned assessment, we introduce ShotJudge, a novel evaluation paradigm that employs LLM judges calibrated with expert few-shot exemplars to mitigate self-rewarding biases. Our empirical evaluation of state-of-the-art LLMs reveals a pronounced performance ceiling: even leading models achieve a peak success rate of only ~66%, with a mean score around 55%. Models also exhibit domain-specific divergence, showing non-overlapping strengths in quantitative reasoning versus linguistic synthesis.. These findings underscore a significant "expert-gap" in current AI systems and establish XpertBench as a critical instrument for navigating the transition from general-purpose assistants to specialized professional collaborators.

cs.AI

R^2-HGP: A Double-Regularized Gaussian Process for Heterogeneous Transfer Learning

Multi-output Gaussian process (MGP) models have attracted significant attention for their flexibility and uncertainty-quantification capabilities, and have been widely adopted in multi-source transfer learning scenarios due to their ability to capture inter-task correlations. However, they still face several challenges in transfer learning. First, the input spaces of the source and target domains are often heterogeneous, which makes direct knowledge transfer difficult. Second, potential prior knowledge and physical information are typically ignored during heterogeneous transfer, hampering the utilization of domain-specific insights and leading to unstable mappings. Third, inappropriate information sharing among target and sources can easily lead to negative transfer. Traditional models fail to address these issues in a unified way. To overcome these limitations, this paper proposes a Double-Regularized Heterogeneous Gaussian Process framework (R^2-HGP). Specifically, a trainable prior probability mapping model is first proposed to align the heterogeneous input domains. The resulting aligned inputs are treated as latent variables, upon which a multi-source transfer GP model is constructed and the entire structure is integrated into a novel conditional variational autoencoder (CVAE) based framework. Physical insights is further incorporated as a regularization term to ensure that the alignment results adhere to known physical knowledge. Next, within the multi-source transfer GP model, a sparsity penalty is imposed on the transfer coefficients, enabling the model to adaptively select the most informative source outputs and suppress negative transfer. Extensive simulations and real-world engineering case studies validate the effectiveness of our R^2-HGP, demonstrating consistent superiority over state-of-the-art benchmarks across diverse evaluation metrics.

cs.LG

A Systematic Characterization of LLM Inference on GPUs

This work presents a systematic characterization of Large Language Model (LLM) inference to address fragmented understanding. Through comprehensive experiments, we establish a four-dimensional analytical framework: (1) Two-Phase Heterogeneity Observation; (2) Microarchitectural Root Cause Analysis; (3) System Scaling Principles; and (4) Emerging Paradigm Boundaries. Our investigation progresses systematically from observation to foresight: identifying performance phenomena, revealing hardware causes, validating system behavior, and exploring new paradigms. This study not only consolidates a reliable empirical foundation for existing research but also provides new discoveries and practical optimization guidance for LLM inference.

cs.AR

N\'eel-Vector-Orientation Induced Direction-Robust Spin Filtering in Two-Dimensional Altermagnets

Whether an antiferromagnet can host direction-robust spin-polarized transport without a conventional spin-selective band gap remains a central challenge in antiferromagnetic spintronics. Here we establish a gapless, direction-robust spin-filtering mechanism in a compensated two-dimensional altermagnetic Weyl semimetal that requires neither a spin-selective band gap nor a large velocity contrast between spin projections. Using Janus monolayer Ta$_2$TeSeO as a realistic platform, we combine symmetry analysis with first-principles calculations, full-Brillouin-zone Wannier interpolation, and semiclassical transport. Rotating the N\'eel vector removes a unitary-mirror constraint and shifts one Weyl-cone pair away from its parent high-symmetry line. For an in-plane N\'eel vector, the residual $C_{2z}\mathcal T$ symmetry forbids the independent $\sigma_y$ mass that would open a local gap, allowing the reconstructed cones to shift in momentum while remaining gapless. Breaking unitary $C_{2z}$ simultaneously lifts the energy equivalence of the remaining mirror-pinned Weyl cones. The resulting coexistence of a metallic spin-projected manifold and a low-DOS Weyl-derived manifold produces a predominantly DOS-driven conductance imbalance. At charge neutrality and 20~K, the longitudinal conductivity polarization for $\mathbf n\parallel x$ remains positive for every in-plane current direction and ranges from $76.4\%$ to $82.0\%$. The degenerate in-plane magnetic anisotropy facilitates reversible switching between symmetry-related spin-filtering states using strain or weak anisotropic fields. This N\'eel-vector-driven symmetry mechanism provides a general route to direction-robust gapless spin filtering in compensated altermagnets.

cond-mat.mes-hall

UniGTE: Unified Graph-Text Encoding for Zero-Shot Generalization across Graph Tasks and Domains

Generalizing to unseen graph tasks without task-specific supervision is challenging: conventional graph neural networks are typically tied to a fixed label space, while large language models (LLMs) struggle to capture graph structure. We introduce UniGTE, an instruction-tuned encoder-decoder framework that unifies structural and semantic reasoning. The encoder augments a pretrained autoregressive LLM with learnable alignment tokens and a structure-aware graph-text attention mechanism, enabling it to attend jointly to a tokenized graph and a natural-language task prompt while remaining permutation-invariant to node order. This yields compact, task-aware graph representations. Conditioned solely on these representations, a frozen LLM decoder predicts and reconstructs: it outputs the task answer and simultaneously paraphrases the input graph in natural language. The reconstruction objective regularizes the encoder to preserve structural cues. UniGTE is instruction-tuned on five datasets spanning node-level, edge-level, and graph-level tasks across diverse domains, yet requires no fine-tuning at inference. It achieves new state-of-the-art zero-shot results on node classification, link prediction, graph classification, and graph regression under cross-task and cross-domain settings, demonstrating that tight integration of graph structure with LLM semantics enables robust, transferable graph reasoning.

cs.LG

Towards Universal Material Property Prediction with Deep Learning and Single-Descriptor electronic Density

Owing to its high scalability and computational efficiency, machine learning methods have been increasingly integrated into various scientific research domains, including ab initio-based materials design. It has been demonstrated that, by incorporating modern machine learning algorithms, one can predict material properties with practically acceptable accuracy. However, one of the most significant limitations that restrict the widespread application of machine learning is its lack of transferability, as a given framework is typically applicable only to a specific property. The origin of this limitation is rooted in the fact that a material's properties are determined by multiple degrees of freedom -- and their complex interplay -- associated with nuclei and electrons, such as atomic type, structural symmetry, and the number and quantum states of the valence electrons, among others. The inherent complexity rules out the possibility of a single machine learning framework providing a full description of these critical quantities. In this paper, we develop a universal machine learning framework based solely on a physically grounded and theoretically rigorous descriptor -- electronic charge density. Our framework not only enables accurate prediction of eight different material properties (with R$^2$ values up to 0.94), but also demonstrates outstanding multi-task learning capability, as prediction accuracy improves when more target properties are incorporated into a single training process, thereby indicating excellent transferability. These results represent a significant step toward realizing the long-standing goal of a universal machine learning framework for the unified prediction of all material properties.

cond-mat.mtrl-sci

Networked Control and Mean Field Problems Under Diagonal Dominance: Decentralized and Social Optimality

In this article, we employ an input-output approach to expand the study of cooperative multi-agent control and optimization problems characterized by mean-field interactions that admit decentralized and selfish solutions. The setting involves $n$ independent agents that interact solely through a shared cost function, which penalizes deviations of each agent from the group's average collective behavior. Building on our earlier results established for homogeneous agents, we extend the framework to nonidentical agents and show that, under a diagonal dominant interaction of the collective dynamics, with bounded local open-loop dynamics, the optimal controller for $H_\infty$ and $H_2$ norm minimization remains decentralized and selfish in the limit as the number of agents $n$ grows to infinity.

math.OC

Wasserstein Distributionally Robust Adaptive Covariance Steering

We present a methodology for predictable and safe covariance steering control of uncertain nonlinear stochastic processes. The systems under consideration are subject to general uncertainties, which include unbounded random disturbances (aleatoric uncertainties) and incomplete model knowledge (state-dependent epistemic uncertainties). These general uncertainties lead to temporally evolving state distributions that are entirely unknown, can have arbitrary shapes, and may diverge unquantifiably from expected behaviors, leading to unpredictable and unsafe behaviors. Our method relies on an $\mathcal{L}_1$-adaptive control architecture that ensures robust control of uncertain stochastic processes while providing Wasserstein metric certificates in the space of probability measures. We show how these distributional certificates can be incorporated into the high-level covariance control steering to guarantee safe control. Unlike existing distributionally robust planning and control methodologies, our approach avoids difficult-to-verify requirements like the availability of finite samples from the true underlying distribution or an a priori knowledge of time-varying ambiguity sets to which the state distributions are assumed to belong.

eess.SY

Unveiling Insulating Ferro and Ferrimagnetism in Double-Double Perovskite Oxides

The emergence of ferro- and ferrimagnetic behavior in insulating materials is uncommon, largely due to Hund's rules. Utilizing symmetry analysis, first-principles methods, and classical Monte Carlo simulations, \textcolor{black}{we report technologically important insulating ferro and ferrimagnetic double-double perovskite oxides. Our study predicts LaA$^{\prime}$MnNiO$_6$ (A$^{\prime}$ = V, Cr, Mn, Co, and Ni) as promising candidates for spintronic and optical applications exhibiting band gaps between 1.3 eV and 1.9 eV. We explain the mechanisms driving band gap openings and magnetic exchange interactions in these ferro and ferrimagnetic compounds. Monte Carlo simulations, together with state-of-the-art orbital-decomposed exchange parameter analysis, reveal intriguing variations in magnetic transition temperatures (up to 242 K) and the corresponding exchange mechanisms in all LaA$^{\prime}$MnNiO$_6$ compounds.} In addition, we assess the thermodynamic and dynamic stability of these compounds to comment on the feasibility of these systems.

cond-mat.mtrl-sci