arXiv ScienceSearch

arXiv subjects

Li Zhu

Publications and source records attributed to Li Zhu.

At least 19 recordsLinked to original sources

Vibrational spectroscopy identifies the bond asymmetry of hexagonal diamond

Bulk hexagonal diamond has been synthesized by independent routes, but its structure remains contested: the two recent refinements disagree even on the sign of the difference between its two inequivalent bond lengths, 238~mÅ apart, and both depart from an earlier 2003 refinement. Here we test the competing structures with first-principles lattice dynamics. Relaxed hexagonal diamond has an interlayer bond \emph{longer} than the intralayer bonds by 24~mÅ in both functionals, an effect of its eclipsed conformation that scales with polytype hexagonality. The bright zone-center $A_{1g}$ mode gauges the interlayer bond at $\approx\!-2{,}100$~\icm~Å$^{-1}$, and neither refined coordinate reproduces the full pattern of measured modes. The only structure matching the twinned sample's three bands requires tens-of-gigapascals confining stress and lattice constants excluded by its own diffraction. Raman spectroscopy and diffraction jointly select a small positive bond asymmetry: inverting the spectrum of the phase-pure sample gives $\OB-\OA=24\pm3$~mÅ (95\% interval), and two determinations on separate samples give $+33\pm8$ and $+60\pm45$~mÅ. The 1{,}529~\icm{} feature cannot be assigned to homogeneous ideal 2H diamond, and the local HRTEM observation remains an open puzzle.

cond-mat.mtrl-sci

Partition Scores Are Not System Scores: Deployment-Fidelity Gaps in Decomposed Algorithm Selection

Oracle-style quantities, including virtual best solvers, selected-portfolio VBS, virtual-best encodings, and best-in-family summaries, are widely reported as upper bounds on what a deployable selector could achieve. In decomposed algorithm selection, an analogous partition-level score grants an oracle choice of the best algorithm within the selected family; once the family selector is fixed, the deployable system must replace that within-family oracle with a learned within-family selector. We define the deployment-fidelity gap G(R) as the difference between partition-level and deployable end-to-end utility and derive two accounting consequences: a per-instance margin-regret stability condition that tells us when a partition-time family choice is deployment-optimal, and a sharp partition-only identification interval that, when it strictly crosses zero, prevents the partition-level report from certifying the deployable winner. Across five public algorithm-selection benchmarks spanning tabular AutoML and combinatorial CSP/SAT, every decomposed pipeline has positive G(R), ranging from 0.012 on TabZilla to 0.13 on PROTEUS-2014. Four of ten decomposed-versus-flat decisions have sign-changing point estimates; on PROTEUS-2014, a 33-point partition advantage shrinks to a 20-point end-to-end advantage. A training-side validation gap-correction diagnostic recovers the point-estimate deployable sign on all four sign-changing cells; it is a reporting aid, not a substitute for direct end-to-end evaluation. Partition and end-to-end scores should be reported side by side.

cs.AI

The Optimal Coefficients-Based Criterion for Primitive Quadratic Polynomials over Finite Fields

Let $\mathbb F_q$ be a finite field and consider quadratic polynomials $f(X)=X^2+bX+c\in\mathbb F_q[X]$ with primitive constant term $c$. We construct an optimal coefficients-based determining polynomial for primitive quadratic polynomials over every finite field. More precisely, for every primitive $c\in\mathbb F_q$, its specialization is the unique monic square-free polynomial whose roots are exactly the coefficients $b$ for which $X^2+bX+c$ is primitive. The construction is based on Lucas polynomials and Lucas atoms over finite fields. In odd characteristic, the optimal determining polynomial is the $(q+1)$-st Lucas atom. In characteristic $2$, this Lucas atom has multiplicity $2$ in the variable $B$, and the optimal polynomial is obtained by removing this multiplicity through the inverse Frobenius over $A=\mathbb F_q[C]/(Φ_{q-1}(C))$. We prove that the determining polynomial is unique for each fixed primitive constant term and, globally, unique as an element of $A[B]$. We also give equivalent criteria involving only recursively computable Lucas polynomials and, as an application, a first-zero coefficient description of the binomial order and order of an irreducible quadratic polynomial.

math.NT

Learning and retrieval for warm-starting charge-self-consistent DFT+DMFT

Charge-self-consistent (CSC) DFT+DMFT delivers quantitative correlated-electron physics one configuration at a time, making ensemble sampling dependent on reliable warm starts for its expensive fixed-point iteration. Here we compare retrieval from the most similar converged structure with an E(3)-equivariant model that predicts a physics-structured self-energy and Fermi level. Because the full CSC loop refines both initializations, every production result remains a converged DFT+DMFT solution. Across metallic Fe, correlated FeO, and Mott-insulating NiO, learning reduces the typical iterations to sustained convergence from 8 to 3, 8 to 3, and 6 to 1. Retrieval matches this median speed when a dense same-state archive is available, but several poor transplants reveal that structural similarity does not predict transplant quality. In a pre-registered volume window excluded from training and donor pools, learning retains its speed, whereas retrieval starts $3.4\times$ farther from the fixed point and approaches cold-start cost. Either initialization can select a distinct near-degenerate branch on a rugged CSC landscape. Applied end to end, the workflow generates over one thousand correlated energy and force labels for iron at Earth's-core conditions and trains an equivariant interatomic potential. Solid--liquid coexistence with 9216 atoms gives $T_m = 6225 \pm 42\,\mathrm{K}_{\mathrm{stat}}$ at 330~GPa, consistent with recent experiments. A 50-configuration DFT+DMFT audit resolves the potential's energy calibration but does not justify a corrected melting temperature, because four-atom cells cannot realize a liquid. The resulting regime map favors retrieval within dense coverage, amortized learning at and beyond its boundary, and solver refinement throughout.

cond-mat.mtrl-sci

Explicit Representatives and Sizes of Cyclotomic Cosets II: Cyclotomic Systems and Applications to Constacyclic Codes

Let $q=p^e$ be a prime power, let $\ell$ be a prime with $\gcd(\ell,qn)=1$, and let $n$ be a positive integer. While $q$-cyclotomic cosets are usually considered modulo a fixed integer, we consider their behavior along the sequence $n,\ell n,\ell^2n,\ldots$ of moduli. To describe the resulting compatibility among cyclotomic cosets, we introduce the $\ell$-adic $q$-cyclotomic system with base module $n$, defined as the projective limit of the spaces of $q$-cyclotomic cosets modulo $\ell^i n$. We associate to each $q$-cyclotomic coset modulo $n$ a cyclotomic $\ell$-adic integer and use its $\ell$-adic expansion to describe all compatible liftings through the successive levels. This leads to a classification of the corresponding sequences according to their splitting and stable behaviors, together with explicit formulas for the representatives and sizes of their components. We treat separately the cases of odd $\ell$ and $\ell=2$. These results are then used to determine representatives and sizes of $q$-cyclotomic cosets for arbitrary admissible parameters. As an application, we obtain explicit irreducible factorizations of binomials over finite fields and use them to describe the corresponding constacyclic codes.

math.NT

Can Video Diffusion Models Predict Past Frames? Bidirectional Cycle Consistency for Reversible Interpolation

Video frame interpolation aims to synthesize realistic intermediate frames between given endpoints while adhering to specific motion semantics. While recent generative models have improved visual fidelity, they predominantly operate in a unidirectional manner, lacking mechanisms to self-verify temporal consistency. This often leads to motion drift, directional ambiguity, and boundary misalignment, especially in long-range sequences. Inspired by the principle of temporal cycle-consistency in self-supervised learning, we propose a novel bidirectional framework that enforces symmetry between forward and backward generation trajectories. Our approach introduces learnable directional tokens to explicitly condition a shared backbone on temporal orientation, enabling the model to jointly optimize forward synthesis and backward reconstruction within a single unified architecture. This cycle-consistent supervision acts as a powerful regularizer, ensuring that generated motion paths are logically reversible. Furthermore, we employ a curriculum learning strategy that progressively trains the model from short to long sequences, stabilizing dynamics across varying durations. Crucially, our cyclic constraints are applied only during training; inference requires a single forward pass, maintaining the high efficiency of the base model. Extensive experiments show that our method achieves state-of-the-art performance in imaging quality, motion smoothness, and dynamic control on both 37-frame and 73-frame tasks, outperforming strong baselines while incurring no additional computational overhead.

cs.CV

Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoning

Large multimodal models have achieved strong reasoning on complex visual tasks, but their inference efficiency is often restricted by long chains of thought. A promising solution is to pair a small draft model with a large target model, enabling cooperative inference employing a routing signal that adaptively routes queries to either the draft or target model based on their difficulties for optimal efficiency and accuracy. Yet, the remaining bottleneck is to establish a reliable query difficulty signal under multimodal settings. Existing approaches designed for language models either rely on post-hoc token probabilities, which fall short in multimodal scenarios, or depend on supervised fine-tuning, which is a data-sensitive strategy. Both paradigms perform routing only after a complete output, and ignore whether the target model can actually solve the routed instances. To address this, we propose PRP, a Proactive Routing Paradigm that enables early decision-making by jointly evaluating the competence of both the draft and target models. Our Draft Rating Learning (DRL) equips the draft model with an internal confidence estimator, while Joint Rating Learning (JRL) predicts how well the target model can handle a given query, thereby prioritizing the allocation of samples it excels at rather than the hardest ones. These ratings enable fine-grained, instance-level \textbf{Proactive Routing} and substantially accelerate inference without compromising overall performance. Extensive experiments across multiple multimodal reasoning benchmarks validate our effectiveness and efficiency.

cs.CL

Benchmarking Multi-Modal Graph-based Social Media Popularity Prediction

Social media popularity prediction aims to forecast the future reach or influence of online content from early-stage observations. Accurate prediction enables key downstream applications, such as advertising optimization and strategic content planning by users, creators, and platforms. Despite substantial progress, existing popularity prediction works often fail to jointly consider multimodal content and temporal social interaction signals. Moreover, the literature remains highly fragmented across datasets, modalities, observation windows, prediction targets, and evaluation protocols. This fragmentation prevents fair comparison and obscures a systematic understanding of how textual, visual, temporal, and interaction-based signals jointly shape popularity dynamics. To address these challenges, we introduce MMG-Pop, a Multi-modal Graph-based Popularity Prediction benchmark, which unifies datasets, modalities, temporal interaction signals, and representative baselines under a standardized evaluation protocol. Furthermore, we propose MMG-PopNet, a unified multi-modal graph-based network that jointly models the aforementioned multi-modal signals and graph-structured social interactions. Extensive experiments on MMG-Pop, comprising four datasets across Bluesky and Reddit platforms, demonstrate the superior performance of MMG-PopNet and yield new insights into cross-platform training generalization, multi-task prediction benefits, multi-modality contributions, and LLM prediction limitation. These findings establish a unified foundation for future research on social dynamics modeling and intervention under heterogeneous modalities and socially-aware agentic ecosystem paradigms.

cs.SI

Pome: Parallelizing I/Os and Computations for Efficient LSM-tree-based Data Storage

CPU computations and I/O operations are fundamental to data storage systems. Storage systems conduct computations with their user threads, such as sorting data for orderliness. They handle I/Os mainly through system calls (syscalls) including file write, read, and fsync, which the OS's kernel threads perform with storage devices. Today, LSM-tree-based storage systems are widely deployed in production environments. Compaction is an essential operation that LSM-tree employs to maintain its tiered tree-like structure by re-sorting and re-storing data through computations and I/Os,respectively. In this paper, we first overhaul the procedure of a compaction. We find that computations and I/Os execute in sequential order. After re-sorting data, the user thread waits for a kernel thread to complete file write and fsync I/Os. These costly synchronous I/Os create a severely long critical path that affects the performance of LSM-tree. To address this issue, we propose parallelizing I/Os and computations for efficient LSM-tree-based data storage (Pome). Pome decouples computations from I/Os within each compaction by referring to its new protocol that moves I/O operations out of the critical path. To this end, it leverages the io_uring to perform asynchronous I/Os. Furthermore, regarding the potential I/O congestion caused by accelerated compactions, Pome incorporates an adaptive I/O rate limiter to achieve smooth execution. We prototype Pome on top of RocksDB. Experimental results demonstrate that Pome significantly improves the performance of RocksDB and outperforms several state-of-the-art LSM-tree variants.

cs.DB

StepGuard: Guarding Web Navigation via Single-Step Calibration

Web navigation requires agents to follow natural language goals, interact with web pages, and produce accurate answers. While recent advances leverage vision-language models and reinforcement learning, existing methods still suffer from single-step fragility due to reward misalignment and error propagation. To tackle the reward entanglement, we design Dynamic Dual-Policy Optimization (DDPO), which dynamically switches between a navigation-first mode for exploration and an answer-first mode for question-answering to mitigate reward conflict. To calibrate the single-step error, we propose Confidence-Guided Adaptive Navigation Reflection (CANR), a mechanism that estimates per-step confidence, triggers reflection only when necessary, and uses contrastive rewards to encourage self-correction to calibrate the single-step inaccuracy. With the above as the main components, we finally develop our StepGuard, a new framework of Guarding Web Navigation via Single-Step Calibration. Experiments demonstrate that our approach significantly improves navigation and answer accuracy, setting new state-of-the-art performance on standard web navigation benchmarks.

cs.AI

The Arithmetic Singleton Bound on the Hamming Distances of Simple-rooted Constacyclic Codes over Finite Fields

In this work, We introduce a new upper bound on the Hamming distance of simple-root constacyclic codes over finite fields, which we call the arithmetic Singleton bound. The main technical tool is the notion of a multiple equal-difference (MED) representation. Via the MED representations of the defining set of the generator polynomial of a simple-root constacyclic code, we obtain a family of upper bounds on its Hamming distance, among which the weakest one coincides with the Singleton bound, while the strongest one is defined to be the arithmetic Singleton bound for this code. Consequently, the arithmetic Singleton bound is always at least as strong as the classical Singleton bound, and is in fact strictly stronger in numerous nontrivial cases. The arithmetic Singleton bound partially measures the restriction on the Hamming distance of a simple-root constacyclic code imposed by its arithmetic structure. In particular, for an irreducible constacyclic code, the MED representations of the defining set of its generator polynomial are completely determined, via which the arithmetic Singleton bound is computed concretely. Finally for any simple-root cyclic code the arithmetic Singleton bound and the BCH bound are compared.

cs.IT

A Voxel-Based Quantum Computing Method (VBQC) for Solid Mechanics Problem

Quantum computing presents a promising method to overcome the efficiency and memory constraints in large-scale mechanical problems, with numerous successful applications demonstrated in fluid mechanics. However, solid mechanics problems usually require irregular grids for spatial discretization, due to the Lagrange formulations and complex boundaries, which makes the quantum simulation of the system matrix, e.g., the mass or stiffness matrix which is often referred to as the Hamiltonian in quantum computing, difficult to be effectively conducted. This study proposes a voxel-based quantum computing method (VBQC) for the quantum simulation of Hamiltonians in solid mechanics. VBQC applies voxel grids to discretize the spatial domain, thereby enabling the system matrix to exhibit the tridiagonal fractal property. Based on this property, the system matrix can be decomposed into three groups of fundamental matrices, $\mathbf{k}_{n}$, $\mathbf{c}_{n}$, and $\mathbf{q}_{n}$. This decomposition process is referred to as the KCQ decomposition. By integrating the KCQ decomposition with the quantum Fourier transform and the quantum multiplexer, VBQC enables efficient quantum simulation of Hamiltonians in solid mechanics. Three specific solid problems with different dimensions and numbers of variables are applied to preliminarily verify the correctness of the proposed VBQC for solid mechanics problems.

cs.CE

Dataset-aware entropy-maximized active learning for machine-learned interatomic potentials

We present an active learning framework for efficiently generating training data for machine-learned interatomic potentials (MLIPs). The method combines local entropy-driven molecular dynamics with global dataset-aware filtering: a per-configuration entropy term biases MD trajectories toward structurally diverse snapshots, while a global entropy measure, the log-determinant of the fingerprint covariance matrix of the entire dataset, selects only those configurations that provide genuinely new information. We employ dual covariance modes (per-atom for disordered structures and per-config for ordered phases) to achieve broad coverage of configuration space. Combined with a pre-trained foundation model (Allegro-OAM-L) and analytical fingerprint gradients from Gaussian overlap matrix eigenvalues, the framework produces high-quality domain-specific potentials with near- or sub-meV/atom accuracy on test data drawn from the same distribution at training-set sizes of order $10^{2}$ to $10^{3}$ entropy-selected DFT-labeled structures. We demonstrate the method on three systems spanning diverse bonding types and pressure-driven phase transitions: carbon (covalent), silicon (covalent/metallic), and NaCl (ionic). In learning curve comparisons against random molecular dynamics sampling at matched training set sizes ($N = 100$ to $800$), evaluated over three independent training-set draws per condition, entropy-driven sampling achieves a factor of approximately $3$ to $10$ lower energy MAE at $N = 800$ on in-distribution holdouts across the three systems, with the magnitude of the gain depending on the bonding type and the size at which the random-MD baseline saturates.

cond-mat.mtrl-sci

Cultivating Forensic Reasoning for Generalizable Multimodal Manipulation Detection

Recent advances in generative AI have significantly enhanced the realism of multimodal media manipulation, thereby posing substantial challenges to manipulation detection. Existing manipulation detection and grounding approaches predominantly focus on manipulation type classification under result-oriented supervision, which not only lacks interpretability but also tends to overfit superficial artifacts. In this paper, we argue that generalizable detection requires incorporating explicit forensic reasoning, rather than merely classifying a limited set of manipulation types, which fails to generalize to unseen manipulation patterns. To this end, we propose REFORM, a reasoning-driven framework that shifts learning from outcome fitting to process modeling. REFORM adopts a three-stage curriculum that first induces forensic rationales, then aligns reasoning with final judgments, and finally refines logical consistency via reinforcement learning. To support this paradigm, we introduce ROM, a large-scale dataset with rich reasoning annotations. Extensive experiments show that REFORM establishes new state-of-the-art performance with superior generalization, achieving 81.52% ACC on ROM, 76.65% ACC on DGM4, and 74.9 F1 on MMFakeBench.

cs.CV

SoK: Security of Autonomous LLM Agents in Agentic Commerce

Autonomous large language model (LLM) agents such as OpenClaw are pushing agentic commerce from human-supervised assistance toward machine actors that can negotiate, purchase services, manage digital assets, and execute transactions across on-chain and off-chain environments. Protocols such as the Trustless Agents standard (ERC-8004), Agent Payments Protocol (AP2), OKX Agent Payments Protocol (APP), the HTTP 402-based payment protocol (x402), Agent Commerce Protocol (ACP), the Agentic Commerce standard (ERC-8183), and Machine Payments Protocol (MPP) enable this transition, but they also create an attack surface that existing security frameworks do not capture well. This Systematization of Knowledge (SoK) develops a unified security framework for autonomous LLM agents in commerce and finance. We organize threats along five dimensions: agent integrity, transaction authorization, inter-agent trust, market manipulation, and regulatory compliance. From a systematically curated public corpus of academic papers, protocol documents, industry reports, and incident evidence, we derive 12 cross-layer attack vectors and show how failures propagate from reasoning and tooling layers into custody, settlement, market harm, and compliance exposure. We then propose a layered defense architecture addressing authorization gaps left by current agent-payment protocols. Overall, our analysis shows that securing agentic commerce is inherently a cross-layer problem that requires coordinated controls across LLM safety, protocol design, identity, market structure, and regulation. We conclude with a research roadmap and a benchmark agenda for secure autonomous commerce.

cs.CR

Physiology-Aware Masked Cross-Modal Reconstruction for Biosignal Representation Learning

Biosignals acquired from different locations on the body often provide temporally ordered views of the same underlying physiological process. However, most existing self supervised learning methods treat these signals as interchangeable views, overlooking the directional temporal dynamics that link them. A canonical example is the relationship between electrocardiography (ECG), which captures the electrical activation initiating each heartbeat, and photoplethysmography (PPG), which records the resulting peripheral pulse delayed by vascular dynamics. To capture this structured relationship, we introduce xMAE, a biosignal pretraining framework that leverages masked cross modal reconstruction across temporally ordered biosignals as a training time constraint to encourage physiologically meaningful timing structure in the learned representations. We show that pretraining with xMAE yields representations that outperform both unimodal and multimodal baselines on 15 of 19 downstream tasks, including cardiovascular outcome prediction, abnormal laboratory test detection, sleep staging, and demographic inference, while generalizing across devices, body locations, and acquisition settings. Further analysis suggests that the ECG PPG timing structure is reflected in the learned PPG representations. More broadly, xMAE demonstrates the effectiveness of incorporating temporal structure into multimodal pretraining when signals observe different stages of a shared underlying process. Code is available at https://github.com/hzhou3/xMAE.

cs.LG

AdapTime: Enabling Adaptive Temporal Reasoning in Large Language Models

Large language models have demonstrated strong reasoning capabilities in general knowledge question answering. However, their ability to handle temporal information remains limited. To address this limitation, existing approaches often involve external tools or manual verification and are tailored to specific scenarios, leading to poor generalizability. Moreover, these methods apply a fixed pipeline to all questions, overlooking the fact that different types of temporal questions require distinct reasoning strategies, which leads to unnecessary processing for simple cases and inadequate reasoning for complex ones. To this end, we propose AdapTime, an adaptive temporal reasoning method that dynamically executes reasoning steps based on the input context. Specifically, it involves three temporal reasoning actions: reformulate, rewrite and review, with an LLM planner guiding the reasoning process. AdapTime integrates seamlessly with state-of-the-art LLMs and significantly enhances their temporal reasoning capabilities without relying on external support. Extensive experiments demonstrate the effectiveness of our approach.

cs.CL

MultiDx: A Multi-Source Knowledge Integration Framework towards Diagnostic Reasoning

Diagnostic prediction and clinical reasoning are critical tasks in healthcare applications. While Large Language Models (LLMs) have shown strong capabilities in commonsense reasoning, they still struggle with diagnostic reasoning due to limited domain knowledge. Existing approaches often rely on internal model knowledge or static knowledge bases, resulting in knowledge insufficiency and limited adaptability, which hinder their capacity to perform diagnostic reasoning. Moreover, these methods focus solely on the accuracy of final predictions, overlooking alignment with standard clinical reasoning trajectories. To this end, we propose MultiDx, a two-stage diagnostic reasoning framework that performs differential diagnosis by analyzing evidence collected from multiple knowledge sources. Specifically, it first generates suspected diagnoses and reasoning paths by leveraging knowledge from web search, SOAP-formatted case, and clinical case database. Then it integrates multi-perspective evidence through matching, voting, and differential diagnosis to generate the final prediction.~Extensive experiments on two public benchmarks demonstrate the effectiveness of our approach.

cs.CL