arXiv ScienceSearch

arXiv subjects

Ze Xu

Publications and source records attributed to Ze Xu.

14 recordsLinked to original sources

LWDrive: Layer-Wise World-Model-Guided Vision-Language Model Planning for Autonomous Driving

Vision-Language Models (VLMs) provide powerful semantic understanding and commonsense reasoning for End-to-End Autonomous Driving (E2E-AD) planning. However, trajectories directly generated by VLMs often encode only coarse driving intentions and remain insufficient for geometrically accurate, future-aware, and multi-view-grounded planning. To address these limitations, we develop the Layer-Wise World-Model-Guided Driving framework (LWDrive). LWDrive is a VLM planning framework that refines coarse trajectories through layer-wise world-model guidance. Instead of treating the VLM output as the final trajectory, LWDrive uses it as an intent-aware coarse plan, expands a diverse candidate space around it, and progressively refines the candidates through a Foresight Cascade Planner (FCP). Specifically, we introduce future-frame generation supervision to encourage the VLM to learn forward-looking scene representations, thereby injecting planning-relevant predictive dynamics into its internal hidden states. Built upon these world-model-supervised representations, FCP exploits VLM features across multiple layers and integrates historical temporal states, Action-Query representations, and current-frame multi-view Bird's-Eye-View (BEV) features to refine candidate trajectories in a coarse-to-fine manner. This design enables progressive correction of spatial positions and motion trends while grounding trajectory refinement with multi-view scene cues and preserving the high-level driving intention produced by the large model. Finally, a score head evaluates the refined candidates and selects the best trajectory as the final planning output. Experiments show that LWDrive achieves a score of 92.0 on the NAVSIM benchmark and 89.6 on NAVSIM-v2. Code and models will be made publicly available.

cs.CV

Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents

The paper introduces GUI-Owl-1.5, the latest native GUI agent model that features instruct/thinking variants in multiple sizes (2B/4B/8B/32B/235B) and supports a range of platforms (desktop, mobile, browser, and more) to enable cloud-edge collaboration and real-time interaction. GUI-Owl-1.5 achieves state-of-the-art results on more than 20+ GUI benchmarks on open-source models: (1) on GUI automation tasks, it obtains 56.5 on OSWorld, 71.6 on AndroidWorld, and 48.4 on WebArena; (2) on grounding tasks, it obtains 80.3 on ScreenSpotPro; (3) on tool-calling tasks, it obtains 47.6 on OSWorld-MCP, and 46.8 on MobileWorld; (4) on memory and knowledge tasks, it obtains 75.5 on GUI-Knowledge Bench. GUI-Owl-1.5 incorporates several key innovations: (1) Hybird Data Flywheel: we construct the data pipeline for UI understanding and trajectory generation based on a combination of simulated environments and cloud-based sandbox environments, in order to improve the efficiency and quality of data collection. (2) Unified Enhancement of Agent Capabilities: we use a unified thought-synthesis pipeline to enhance the model's reasoning capabilities, while placing particular emphasis on improving key agent abilities, including Tool/MCP use, memory and multi-agent adaptation; (3) Multi-platform Environment RL Scaling: We propose a new environment RL algorithm, MRPO, to address the challenges of multi-platform conflicts and the low training efficiency of long-horizon tasks. The GUI-Owl-1.5 models are open-sourced, and an online cloud-sandbox demo is available at https://github.com/X-PLUG/MobileAgent.

cs.AI

P-GenRM: Personalized Generative Reward Model with Test-time User-based Scaling

Personalized alignment of large language models seeks to adapt responses to individual user preferences, typically via reinforcement learning. A key challenge is obtaining accurate, user-specific reward signals in open-ended scenarios. Existing personalized reward models face two persistent limitations: (1) oversimplifying diverse, scenario-specific preferences into a small, fixed set of evaluation principles, and (2) struggling with generalization to new users with limited feedback. To this end, we propose P-GenRM, the first Personalized Generative Reward Model with test-time user-based scaling. P-GenRM transforms preference signals into structured evaluation chains that derive adaptive personas and scoring rubrics across various scenarios. It further clusters users into User Prototypes and introduces a dual-granularity scaling mechanism: at the individual level, it adaptively scales and aggregates each user's scoring scheme; at the prototype level, it incorporates preferences from similar users. This design mitigates noise in inferred preferences and enhances generalization to unseen users through prototype-based transfer. Empirical results show that P-GenRM achieves state-of-the-art results on widely-used personalized reward model benchmarks, with an average improvement of 2.31%, and demonstrates strong generalization on an out-of-distribution dataset. Notably, Test-time User-based scaling provides an additional 3% boost, demonstrating stronger personalized alignment with test-time scalability.

cs.CL

EcomBench: Towards Holistic Evaluation of Foundation Agents in E-commerce

Foundation agents have rapidly advanced in their ability to reason and interact with real environments, making the evaluation of their core capabilities increasingly important. While many benchmarks have been developed to assess agent performance, most concentrate on academic settings or artificially designed scenarios while overlooking the challenges that arise in real applications. To address this issue, we focus on a highly practical real-world setting, the e-commerce domain, which involves a large volume of diverse user interactions, dynamic market conditions, and tasks directly tied to real decision-making processes. To this end, we introduce EcomBench, a holistic E-commerce Benchmark designed to evaluate agent performance in realistic e-commerce environments. EcomBench is built from genuine user demands embedded in leading global e-commerce ecosystems and is carefully curated and annotated through human experts to ensure clarity, accuracy, and domain relevance. It covers multiple task categories within e-commerce scenarios and defines three difficulty levels that evaluate agents on key capabilities such as deep information retrieval, multi-step reasoning, and cross-source knowledge integration. By grounding evaluation in real e-commerce contexts, EcomBench provides a rigorous and dynamic testbed for measuring the practical capabilities of agents in modern e-commerce.

cs.AI

Motivic multiplicativity of complete intersections

For a smooth projective variety endowed with a Chow-K\"unneth (abbr. CK) decomposition, we introduce the notions of motivic multiple twist-multiplicativity and multiplicativity defect to measure the obstruction to the compatibility of multiple intersection products with the given CK decomposition. These notions extend the more restrictive notion of multiplicativity introduced by Shen-Vial. We establish their basic properties and derive natural upper bounds for the motivic multiplicativity defects of curves, surfaces, and ample subvarieties of varieties with trivial Chow groups. We then explicitly determine the motivic 2-fold multiplicativity defect of any smooth Fano or Calabi-Yau complete intersection in a smooth weighted projective space, thereby strengthening a result of Fu in the Calabi-Yau case. In particular, we prove that any smooth Fano or Calabi-Yau hypersurface admits motivic 0-multiplicativity. This generalizes the corresponding result for cubic hypersurfaces, proved independently by Diaz and Fu-Laterveer-Vial, and confirms a conjecture of Voisin in the Calabi-Yau case. As a consequence, certain relative powers of the associated universal families satisfy the Franchetta property. We also obtain several further applications.

math.AG

SKYLENAGE Technical Report: Mathematical Reasoning and Contest-Innovation Benchmarks for Multi-Level Math Evaluation

Large language models (LLMs) now perform strongly on many public math suites, yet frontier separation within mathematics increasingly suffers from ceiling effects. We present two complementary benchmarks: SKYLENAGE-ReasoningMATH, a 100-item, structure-aware diagnostic set with per-item metadata on length, numeric density, and symbolic complexity; and SKYLENAGE-MATH, a 150-item contest-style suite spanning four stages from high school to doctoral under a seven-subject taxonomy. We evaluate fifteen contemporary LLM variants under a single setup and analyze subject x model and grade x model performance. On the contest suite, the strongest model reaches 44% while the runner-up reaches 37%; accuracy declines from high school to doctoral, and top systems exhibit a doctoral-to-high-school retention near 79%. On the reasoning set, the best model attains 81% overall, and hardest-slice results reveal clear robustness gaps between leaders and the mid-tier. In summary, we release SKYLENAGE-ReasoningMATH and report aggregate results for SKYLENAGE-MATH; together, SKYLENAGE provides a hard, reasoning-centered and broadly covering math benchmark with calibrated difficulty and rich metadata, serving as a reference benchmark for future evaluations of mathematical reasoning.

cs.CL

Vapor-mediated wetting and imbibition control on micropatterned surfaces

Wetting of micropatterned surfaces is ubiquitous in nature and key to many technological applications like spray cooling, inkjet printing, and semiconductor processing. Overcoming the intrinsic, chemistry- and topography-governed wetting behaviors often requires specific materials which limits applicability. Here, we show that spreading and wicking of water droplets on hydrophilic surface patterns can be controlled by the presence of the vapor of another liquid with lower surface tension. We show that delayed wicking arises from Marangoni forces due to vapor condensation, competing with the capillary wicking force of the surface topography. Thereby, macroscopic droplets can be brought into an effective apparent wetting behavior, decoupled from the surface topography, but coexisting with a wicking film, cloaking the pattern. We demonstrate how modulating the vapor concentration in space and time may guide droplets across patterns and even extract imbibed liquids, devising new strategies for coating, cleaning and drying of functional surface designs.

cond-mat.soft

Origin of Increased Curie Temperature in Lithium-Substituted Ferroelectric Niobate Perovskite: Enhancement of the Soft Polar Mode

The functionality of ferroelectrics is often constrained by their Curie temperature, above which depolarization occurs. Lithium (Li) is the only experimentally known substitute that can increase the Curie temperature in ferroelectric niobate-based perovskites, yet the mechanism remains unresolved. Here, the unique phenomenon in Li-substituted KNbO3 is investigated using first-principles density functional theory. Theoretical calculations show that Li substitution at the A-site of perovskite introduces compressive chemical pressure, reducing Nb-O hybridization and associated ferroelectric instability. However, the large off-center displacement of the Li cation compensates for this reduction and further enhances the soft polar mode, thereby raising the Curie temperature. In addition, the stability of the tetragonal phase over the orthorhombic phase is predicted upon Li substitution, which reasonably explains the experimental observation of a decreased orthorhombic-to-tetragonal phase transition temperature. Finally, a metastable anti-phase polar state in which the Li cation displaces oppositely to the Nb cation is revealed, which could also contribute to the variation of phase transition temperatures. These findings provide critical insights into the atomic-scale mechanisms governing Curie temperature enhancement in ferroelectrics and pave the way for designing advanced ferroelectric materials with improved thermal stability and functional performance.

cond-mat.mtrl-sci

Split representation of adaptively compressed polarizability operator

The polarizability operator plays a central role in density functional perturbation theory and other perturbative treatment of first principle electronic structure theories. The cost of computing the polarizability operator generally scales as $\mathcal{O}(N_{e}^4)$ where $N_e$ is the number of electrons in the system. The recently developed adaptively compressed polarizability operator (ACP) formulation [L. Lin, Z. Xu and L. Ying, Multiscale Model. Simul. 2017] reduces such complexity to $\mathcal{O}(N_{e}^3)$ in the context of phonon calculations with a large basis set for the first time, and demonstrates its effectiveness for model problems. In this paper, we improve the performance of the ACP formulation by splitting the polarizability into a near singular component that is statically compressed, and a smooth component that is adaptively compressed. The new split representation maintains the $\mathcal{O}(N_e^3)$ complexity, and accelerates nearly all components of the ACP formulation, including Chebyshev interpolation of energy levels, iterative solution of Sternheimer equations, and convergence of the Dyson equations. For simulation of real materials, we discuss how to incorporate nonlocal pseudopotentials and finite temperature effects. We demonstrate the effectiveness of our method using one-dimensional model problem in insulating and metallic regimes, as well as its accuracy for real molecules and solids.

physics.comp-ph

Adaptively compressed polarizability operator for accelerating large scale \textit{ab initio} phonon calculations

Phonon calculations based on first principle electronic structure theory, such as the Kohn-Sham density functional theory, have wide applications in physics, chemistry and material science. The computational cost of first principle phonon calculations typically scales steeply as $\mathcal{O}(N_e^4)$, where $N_e$ is the number of electrons in the system. In this work, we develop a new method to reduce the computational complexity of computing the full dynamical matrix, and hence the phonon spectrum, to $\mathcal{O}(N_e^3)$. The key concept for achieving this is to compress the polarizability operator adaptively with respect to the perturbation of the potential due to the change of the atomic configuration. Such adaptively compressed polarizability operator (ACP) allows accurate computation of the phonon spectrum. The reduction of complexity only weakly depends on the size of the band gap, and our method is applicable to insulators as well as semiconductors with small band gaps. We demonstrate the effectiveness of our method using one-dimensional and two-dimensional model problems.

math.NA

Algebraic cycles on a generalized Kummer variety

We compute explicitly the Chow motive of any generalized Kummer variety associated to any abelian surface. In fact, it lies in the rigid tensor subcategory of the category of Chow motives generated by the Chow motive of the underlying abelian surface. One application of this calculation is to show that the Hodge conjecture holds for arbitrary products of generalized Kummer varieties. As another application, all numerically trivial 1-cycles on arbitrary products of generalized Kummer varieties are smash-nipotent.

math.AG

A remark on the Abel-Jacobi morphism for the cubic threefold

Let $X$ be a smooth cubic threefold and $J(X)$ be its intermediate Jacobian. We show that there exists a codimension 2 cycle $Z$ on $J(X)\times X$ with $Z_{t}$ homologically trivial for each $t\in J(X)$, such that the morphism $\phi_{Z}: J(X)\rightarrow J(X)$ induced by the Abel-Jacobi map is the identity. This answers positively a question of Voisin in the case of the cubic threefold.

math.AG

Remarks on Murre's conjecture on Chow groups

For certain product varieties, Murre's conjecture on Chow groups is investigated. In particular, it is proved that Murre's conjecture (B) is true for two kinds of four-folds. Precisely, if $C$ is a curve and $X$ is an elliptic modular threefold over $k$ (an algebraically closed field of characteristic 0) or an abelian variety of dimension 3, then Murre's conjecture (B) is true for the fourfold $X\times C.$

math.AG

On Hard Lefschetz Conjecture on Lawson Homology

Friedlander and Mazur proposed a conjecture of hard Lefschetz type on Lawson homology. We shall relate this conjecture to Suslin conjecture on Lawson homology. For abelian varieties, this conjecture is shown to be equivalent to a vanishing conjecture of Beauville type on Lawson homology. For symmetric products of curves, we show that this conjecture amounts to the vanishing conjecture of Beauville type for the Jacobians of the corresponding curves. As a consequence, Suslin conjecture holds for all symmetric products of curves with genus at most 2.

math.AG