arXiv ScienceSearch

arXiv subjects

Yu Liu

Publications and source records attributed to Yu Liu.

At least 19 recordsLinked to original sources

Hidden Magnetic Complexity Within a Simple van der Waals Ferromagnet Ce$_2$Te$_5$

Ce$_2$Te$_5$ is a layered $f$-electron van der Waals magnet in which reduced dimensionality and inequivalent Ce sites give rise to competing magnetic interactions. We investigate its magnetic ground state using muon spin relaxation ($μ$SR), neutron powder diffraction (NPD), and inelastic neutron scattering (INS). Zero-field $μ$SR reveals an onset of static magnetism below $T_{\mathrm{C}} = 5.0(1)$~K, followed by an additional anomaly in the internal field at $T_{\mathrm{2}} = 2.3(2)$~K, consistent with features observed in bulk thermodynamic and transport measurements. In contrast, NPD data collected between $0.05-8$~K reveal a single long-range ordered magnetic phase below $T_{\mathrm C}$, with the magnetic Bragg intensities vanishing at $T_{\mathrm C}$ with no evidence for additional structural or magnetic phase transitions down to base temperature. The ordered state is characterized by a commensurate propagation vector $\mathbf{k}=(0,0,0)$ and ferromagnetic alignment of Ce moments along the crystallographic $b$ axis. Remarkably, only one of the two crystallographically distinct Ce sites carries an ordered moment of $\sim 0.85(3)~μ_{\mathrm B}$ per Ce at 0.05~K. INS measurements establish the crystal electric field (CEF) energy scale of Ce$^{3+}$, revealing low-lying excitations at 7.94 and 21.46~meV and a Kramers doublet ground state with strong single-ion anisotropy. These results demonstrate that Ce$_2$Te$_5$ undergoes a single symmetry-breaking magnetic transition, while additional low-temperature anomalies reflect subtle modifications of the ordered state driven by competing interactions and CEF effects.

cond-mat.str-el

First Demonstration of Flip DRAM from Process, Architecture to System to Push DRAM Scaling beyond 4F2: 2F2 Self-aligned Flip Vertical Channel Transistor (FVCT) DRAM and Flip WL (FWL) 3D-DRAM

For the first time, we proposed a novel stacking technology for DRAM scaling by flipping and backside processes, making full use of DRAM wafer's backside and investigating it on both 4F2 and 3D-DRAM. For 4F2 VCT, 2F2 Flip VCT featuring self-aligned back-to-back stacked 1T1C bitcell, with various BL and WL configurations, were studied and key process modules such as self-aligned stacked vertical channel, BL and WL formations, wafer bonding and flipping, substrate thinning and low-R Co storage node (SN) were successfully developed, addressing the potential thermal, misalign and parasitic concerns in the Flip VCT process. A full DRAM DTCO framework was also established from device to mat and chip level. Compared to 4F2 VCT DRAM with the same mat size, 2F2 FVCT delivers 27.5% less parasitics, 11% better sense margin, 16.3% higher charge sharing (CS) speed and 50% less area. For 3D-DRAM, a brand-new flip WL staircase design with peripheral circuit innovations was studied and proved to have 25% density gain, 15.1% faster turn-on speed and 6.8% less CS time, proving further extendibility of flip technology on DRAM.

cond-mat.mes-hall

DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models

Reinforcement learning has emerged as a powerful tool for improving diffusion-based text-to-image models, but existing methods are largely limited to single-task optimization. Extending RL to multiple tasks is challenging: joint optimization suffers from cross-task interference and imbalance, while cascade RL is cumbersome and prone to catastrophic forgetting. We propose DiffusionOPD, a new multi-task training paradigm for diffusion models based on Online Policy Distillation (OPD). DiffusionOPD first trains task-specific teachers independently, then distills their capabilities into a unified student along the student own rollout trajectories. This decouples single-task exploration from multi-task integration and avoids the optimization burden of solving all tasks jointly from scratch. Theoretically, we lift the OPD framework from discrete tokens to continuous-state Markov processes, deriving a closed-form per-step KL objective that unifies both stochastic SDE and deterministic ODE refinement via mean-matching. We formally and empirically demonstrate that this analytic gradient provides lower variance and better generality compared to conventional PPO-style policy gradients. Extensive experiments show that DiffusionOPD consistently surpasses both multi-reward RL and cascade RL baselines in training efficiency and final performance, while achieving state-of-the-art results on all evaluated benchmarks.

cs.LG

Optimal Recursive Composition and Dyadic Phase Laws for Gradient Descent with Predetermined Stepsizes

Predetermined stepsize schedules featuring carefully chosen long steps have recently been shown to accelerate gradient descent (GD) on smooth convex functions. A prominent class of such schedules is built through recursive composition. In this paper, we characterize the convergence of these optimized recursive schedules, revealing a non-constant log-periodic modulation across prescribed horizons. Specifically, for symmetric recursive frameworks (primitive and OBS-S constructions), we prove that for every $N \geq 1$, the corresponding optimized schedules satisfy $f(x_{N-1})-f^\ast \le \frac{1}{2N^p Φ(\log_2N)-1} \frac{L}{2}\|x_0-x^\ast\|^2$, $ p=\log_2(1+\sqrt2)$, where $Φ$ is a positive, Lipschitz, nonconstant $1$-periodic function. We derive this by proving that balanced splitting is optimal at every horizon for these constructions, resolving a conjecture of Zhang and Jiang. Furthermore, for the asymmetric framework (the OBS-F construction), we show that although optimal splits are not necessarily balanced, the same Silver exponent asymptotically persists alongside a distinct log-periodic modulation.

math.OC

MADS: Multi-Agent Dialogue Simulation for Diverse Persuasion Data Generation

We propose MADS (Multi-Agent Dialogue Simulation), a scalable framework for generating persuasive multi-turn dialogues via agent self-play. MADS employs three coordinated agents: User Agents designed to simulate diverse persona-driven behaviors by leveraging personality signifiers such as Zodiac Signs and MBTI types, a Dialog Agent executing task-oriented persuasion strategies and an Optimization Agent evaluating and refining dialogue outcomes. We further validate its effectiveness through users' Chain-of-Attitude (CoA) modeling and dedicated LLMs' persuasion assessment. This approach enables low-cost generation of training data without human annotation, addressing key industry challenges such as lack of user data, cold-start evaluation difficulties, and prompt inefficiency. Applied to a real-world marketing scenario, MADS significantly improved the persuasion capacity of small LLMs, increasing the organic traffic conversion rate by 22.4% (from 1.83% to 2.24%) , demonstrating clear business value.

cs.CL

IDMate: Finite-temperature error bounds for window-resolved self-consistent-field screening

We formulate a finite-temperature residual test in IDMate that bounds window-resolved electronic errors without a spectral-gap assumption. For a fixed Hamiltonian and exact electron number, strong convexity of the matrix Fermi entropy bounds the density-matrix distance and free-energy error within a selected window. The bound remains finite at spectral crossings and extends to weighted k points with a shared chemical potential. We also derive a distance correction for particle-number mismatch. Across $3{,}586$ stress trials in $70$ seeded perturbation ladders, $1{,}942$ proposals satisfy the screen with no observed violation of the $0.05$ window-distance criterion plus its numerical allowance. This criterion differs from the uncorrected exact-trace bound, which six of ten historical in-loop candidates exceed at the numerical-error scale. An accept-or-recover loop replaces ten reference-map evaluations while meeting terminal comparison criteria in three configurations that include oracle-subspace controls. Additional candidates built only from preceding-iteration orbitals yield two acceptances and one abstention. The silicon candidate has a window distance of $1.891\times10^{-13}$ but a normalized real-space density error of $3.754\%$. Analytic examples separate errors from complement occupations and interblock coupling. Window-level accuracy therefore does not imply full-state accuracy; the screen tests compressed proposal quality, independently of nonlinear SCF convergence or net acceleration.

cond-mat.mtrl-sci

Magnetic phases of Kondo lattice materials Ce$_5$RhGe$_2$ and Ce$_5$IrGe$_2$

Single crystals of Ce$_5$RhGe$_2$ and Ce$_5$IrGe$_2$ have been systematically investigated by electrical resistivity, specific heat, and magnetization measurements. Together with Ce$_5$CoGe$_2$, all three compounds crystallize in the orthorhombic \emph{Pnma} structure, with the lattice parameters increasing monotonically from Co to Rh to Ir, consistent with the effect of negative chemical pressure. Magnetization measurements along the three principal crystallographic axes identify the \emph{a} axis as the easy magnetization direction throughout the series. Ce$_5$RhGe$_2$ exhibits ferromagnetic ordering with a Curie temperature of approximately 11.5 K and shows magnetic behavior closely resembling that of Ce$_5$CoGe$_2$. In contrast, Ce$_5$IrGe$_2$ undergoes two successive magnetic transitions at $T_{\rm M1}=12.7$ K and $T_{\rm M2}=11.8$ K, and there are multiple metamagnetic transitions under magnetic fields, giving rise to magnetization plateaus at fractions of the saturation magnetization $M_{\rm s}$ of approximately $M_{\rm s}/5$ and $M_{\rm s}/3$. The low-field metamagnetic transition along the easy axis shifts to lower field with decreasing temperature, and eventually a pronounced hysteresis loop is observed about zero-field, establishing that Ce$_5$IrGe$_2$ exhibits a ferrimagnetic ground state at the lowest measured temperatures.

cond-mat.str-el

The energetics of force errors in machine-learned molecular dynamics

The energetic effect of a force error depends on atomic motion. We establish a directional residual-work coefficient combining directional curvature mismatch with the spatial distribution of the residual response. For conservative potentials force-matched at an anchor, it determines the leading signed work at the first crossing of a small force-error budget. At a 474-atom lithium-electrolyte interface, predictions fixed before future reference evaluations differ from measurements by less than 4.6% of predicted work across 24 prescribed endpoints. Changing only the initial velocity direction at fixed structure and initial total kinetic energy reverses the force-work ranking. At the same admitted time of 0.25 fs, one direction gives an 11.6% larger maximum force residual but 36.1% less work. The reversal recurs at a second structure. The framework connects force tolerances to reference-energy transfer, providing a physical basis for potential assessment and adaptive reference allocation.

cond-mat.mtrl-sci

A PTAS for Non-Adaptive Stochastic Top-$k$ Sum under General Combinatorial Constraints

We study non-adaptive selection of a feasible set $S$ that maximizes the expected sum of the $k$ largest realized values among independent nonnegative discrete random variables. The same objective arises in team hiring and as VCG welfare in an $\ell$-unit auction. The main setting is a fixed-dimensional nonnegative packing family, whose natural LP has $d=O(1)$ packing inequalities with binary coefficients. We give a PTAS for every $k\ge 1$ on every such family, including binary one- and two-dimensional knapsack, by approximating the occupancy functional $p\mapsto\mathbb{E}[\min(k,N(p))]$ and realizing the resulting signatures in the packing LP. As a generic guarantee the scheme is essentially optimal: there is no FPTAS that works for every such $\mathcal{F}$ unless $P=NP$, and no EPTAS unless $W[1]=FPT$. An incomparable sufficient condition is a query-weight exact-sum oracle (DAG paths, matchings), which likewise yields a PTAS for every $k$. The same signatures give a PTAS for $\min_{S\in\mathcal{F}}\mathbb{E}[\mathrm{Top}_k(S)]$ on every fixed-$d$ covering family; two-dimensional covering knapsack rules out a generic FPTAS on that class.

cs.DS

Human-agent discovery of reconfigurable in-plane ferroelectric superdomain control

Automated experimentation is most effective when the observables, available actions, and objective are defined before the experiment starts, as is the case for Bayesian optimization. However, in many exploratory experiments, the variables that describe the sample must be extracted from the data, new operations emerge during the experiments, and the instrument budget is too small to learn the problem by trials. Here we introduce the Scanning Probe Agentic Research Cycle (SPARC) framework, in which a coding agent and a human operator share one microscope, one notebook, and two persistent memory files. FINDINGS.md stores graded conclusions about the experiment, whereas PITFALLS.md records learned failure modes of analysis and instrument. We apply SPARC to reconfigure the in-plane superdomain direction of a (111)-oriented PbZr0.2Ti0.8O3 film. In an operator-supervised campaign, the agent reanalyzed earlier manual measurements and developed an oriented lattice of stationary bias pulses with alternating polarity to reconfigure the superdomain direction. In a subsequent agent-controlled campaign, PITFALLS.md entries were compiled into checks that validate a design before any write. The experiments showed that spatial polarity alternation, instead of the exact matching between the lattice and lamellar periods, determines directional selection. Combining a raster scan with a masked pulse lattice printed the letters UTK into the superdomain orientation. The campaign also identified practical requirements for agentic experimentation where physical verification of instrument execution, the conditions under which stored findings remain valid, validation of new observables on instrument data, and robust control protocols.

cond-mat.mtrl-sci

Unveiling the Scaling Potential of Drain Merge through Active (DMtA) in CFETs: Breaking the Super-Via Bottlenecks and Unlocking New PPA Boosters

Drain merge (DM), a super via vertically connecting the common S/D terminals of stacked n/pFETs in Complementary FETs (CFETs), blocks further parasitic optimization and cell scaling. For the first time, this work systematically investigates the state-of-the-art Drain Merge through Active (DMtA), a revolutionary technology reported recently with the DM embedded in the active region, through a comprehensive DTCO framework spanning process integration, contact-configuration-dependent (CTCD) compact modeling, standardcell design, RO evaluation and block-level PPA benchmark on a 32-bit RISC-V Ibex core. By reducing DM parasitics and enabling DM-width optimization, DMtA improves RO frequency by 11.7% over its conventional Drain Merge through field (DMtF) counterpart. Active widening and Area Borrowing, the latter first reported in [8] and exploiting spatial slack in adjacent cells to further enlarge the nanosheet width (WNS), increase the maximum Ibex-core frequency by up to 34.8%. More importantly, DMtA also enables the once GAA-exclusive Hyper-cells on CFETs by merging the active regions across adjacent cell rows, providing a further 8.7% frequency gain. A post-routing floating-output-pin-aware optimization further removes redundant S/D contacts (CTs) and reduces power by 5.3%. Finally, DMtA facilitates more area-efficient 2.5T cell scaling by preserving single-row cell compatibility, reducing post-PR core area by 25.7%.

cond-mat.mes-hall

Inferring Urban Mobility Interactions from Aggregated Dynamics

Real-time urban governance depends not only on knowing where people are, but on how they move between places, directional flows that could be conventionally resolved by tracking individuals through space, i.e., expensive to sustain and built on traces that are highly unique and readily re-identifiable. Here we show that this directional structure need not be observed to be known: aggregated counts which cities already collect retain enough information to reconstruct the temporal evolution of origin-destination (OD) matrix. Using an uncertainty-aware physics-informed framework, we infer future OD flows from area-level counts alone across twelve mobility datasets from cities in the United States and China, reaching accuracy comparable to models that take historical OD matrices as input. Probabilistic modeling corrects the systematic underestimation of sparse, high-value corridors and yields calibrated predictions consistent with observed flows. Architectures that respect the generation-before-assignment logic of transport planning recover interactions more faithfully, indicating that location-level spatial heterogeneity should be preserved before pairwise interactions are reconstructed. Because inference requires only aggregated observations after training, recovering interactions this way reduces reliance on continuous individual-level tracking, pointing toward a more deployable and less exposure-heavy basis for real-time urban intelligence.

cs.LG

Back on Track: Aligning Rewards and States for Reasoning in Diffusion Large Language Models

Reinforcement learning (RL) holds immense promise for enhancing the reasoning capabilities of diffusion large language models (dLLMs). However, progress is fundamentally constrained by a dual misalignment between authentic generation trajectory and the gradient update process: (i) Process-reward misalignment. Sparse, terminal rewards are indiscriminately assigned to all intermediate steps of the generation process, failing to provide discriminative credit assignment. (ii) State-trajectory misalignment. Policy updates are often diverted toward artificial, out-of-trajectory states, squandering gradients on less informative samples. To address these limitations, we introduce Process Aligned Policy Optimization (PAPO), a novel framework that holistically aligns the RL update with the dLLM's generative trajectory via Step-Aware Process Rewards (SPR) that transform sparse terminal rewards into dense, step-wise credit, and Entropy-Guided Historical Re-enactment (EHR) that replays authentic trajectories at high-uncertainty steps. Extensive experiments on four benchmarks demonstrate that PAPO significantly outperforms baselines, achieving gains of 4.5% on GSM8K, 4.8% on MATH500, 42.2% on Countdown and 16.1% on Sudoku.

cs.CL

A Trustworthy Watermarking Framework for LLM-Generated Food Safety Content

Large language models are transforming many industries with their text generation abilities. However, their outputs can be easily tampered with, creating serious risks in critical areas such as food safety reporting. To protect the integrity and traceability of AI-generated content, this paper introduces ToSS (Token Oriented Repartitioning and Strategic Selection), a reliable authentication method using adaptive dual watermarking. The key innovation of ToSS is its dual watermark encoding approach that divides vocabulary tokens into black and white sublists, enabling precise bit-level embedding of traceability information. Additionally, an entropy adaptive mechanism dynamically selects text regions with high prediction uncertainty for watermark insertion, maintaining text fluency and factual accuracy while ensuring reliable traceability. Experiments on multiple datasets, including food domain texts, demonstrate that ToSS achieves leading performance in both watermark capacity and decoding accuracy.

cs.CR

DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents

As large language models evolve from question-answering systems into general-purpose agents, evaluation must move beyond static answer correctness to assess multimodal perception, multi-step execution, tool use, and artifact delivery. However, existing benchmarks are often tied to specific task types, execution environments, or scoring protocols, limiting their comparability, interpretability, and reliability for deployment decisions. We introduce DAREBench (Deployment-Aware and Reliable Evaluation of Models as Agents), a benchmark designed to capture workload variation and support reliable agent evaluation. Built on a shared OpenClaw execution environment, DAREBench organizes 233 tasks selected and adapted from 22 source benchmarks into a $2\times3$ workload matrix defined by input modality and execution form, and evaluates them under a unified contract-based protocol with evidence-based score auditing. We evaluate 23 commercial API models and 12 locally deployed open-weight models over 7,587 model--task runs, reporting accuracy and token consumption alongside reference costs for API models. Results show that no single model dominates all workload groups, text and multimodal tasks exhibit distinct accuracy--cost trade-offs, and local open-weight models are competitive in several groups but still trail frontier commercial models overall. These findings suggest that agent deployment and model selection should consider workload profiles, deployment mode, and accuracy--cost trade-offs rather than rely on a single aggregate score.

cs.AI

BF16 Component-Product Emulation of FP32 and FP64 GEMM on Intel AMX

Modern CPUs increasingly integrate high-throughput matrix engines optimized for low-precision AI workloads, while many scientific computing applications still rely on FP32 and FP64 GEMM to meet their numerical accuracy requirements. This mismatch motivates an algorithmic bridge that uses low-precision matrix products to emulate higher-precision GEMM. This paper presents a CPU-oriented method based on Intel Advanced Matrix Extensions (AMX) and BF16 matrix products. For FP32, each operand is decomposed into three BF16 components and six selected component products are evaluated, targeting FP32-level accuracy relative to oneMKL SGEMM without claiming elementwise or bitwise identity. For FP64 inputs within the supported BF16 exponent range, the method uses a simplified fixed six-slice Ozaki decomposition. Each retained BF16 product is first produced in FP32, then widened and accumulated in FP64. Four product-count settings retain 6, 10, 15, or 21 component products, exposing the accuracy--performance tradeoff relative to oneMKL DGEMM. The implementation combines precomputed packed component buffers, VNNI-packed $B$ panels, and an FP32 tile-resident operand-reuse schedule. On the tested square matrices, AMX-FP32 exceeds oneMKL SGEMM throughput. For AMX-FP64, low-product-count variants can exceed DGEMM at sufficiently large orders, while retaining more products improves accuracy at additional cost.

cs.MS

Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models

Modern vision-language models (VLMs) can directly answer many image-grounded questions, yet they often struggle with complex queries requiring fine-grained visual details or external knowledge. To acquire this missing evidence, agentic VLMs invoke tools such as image cropping, image search, and text search. However, existing training paradigms primarily evaluate tool-use based on final answer correctness, leaving evidence acquisition and utilization insufficiently supervised. This leads to two critical shortcomings: (i) models frequently issue redundant or off-target tool calls that fail to gather necessary evidence, and (ii) even when appropriate tools are called, models often fail to extract the necessary information from the resulting observations. To address these limitations, we introduce the NTEP (Necessary Tool-Evidence Path), a novel annotation scheme that explicitly specifies the essential external evidence and corresponding tool calls for each query. Building upon this, we propose NTEP-R (NTEP Reward), a supervision mechanism ensuring that each tool invocation strictly advances the reasoning process toward the final solution. Specifically, our approach rewards the agent for aligning its pre-call intent with a necessary evidence-seeking goal, and for ensuring the information summarized from the post-call observation aligns with the necessary evidence. Furthermore, we introduce a non-repeated-goal regularizer to penalize redundant calls that revisit satisfied NTEP goals. Extensive evaluations on seven image-grounded benchmarks demonstrate that our 8B-parameter instantiation, NTEP-8B, significantly improves both search-oriented accuracy and tool-use efficiency within a unified three-tool framework. These results highlight the critical value of fine-grained tool-evidence path supervision for training robust agentic VLMs.

cs.AI

Hierarchical automation of scanning probe microscopy through agentic orchestration and algorithmic control

Rapid advances in agentic artificial intelligence enable scientific systems to interpret open-ended objectives, combine heterogeneous information, invoke specialized tools, and revise experimental strategies as evidence accumulates. However, physical experimentation also contains many tasks for which agentic reasoning provides little advantage and can reduce reliability. Quantitative analysis, optimization, spatial targeting, validation, and instrument execution are often better posed as deterministic or algorithmic operations with explicit objectives and verifiable outputs. Here, we introduce a hierarchical architecture for autonomous experimentation that separates these roles. Agentic components interpret scientific intent, construct task-dependent experimental representations, evaluate accumulated evidence, and select high-level actions, whereas deterministic algorithms perform numerical analysis, coordinate selection, validation, and physical execution. We implement this architecture in piezoresponse force microscopy. Starting from a broad scientific question concerning the relation between local domain structure and polarization switching, the system constructs spatial descriptors from multichannel imaging, selects and analyzes local hysteresis measurements, adapts the spectroscopy waveform, and terminates the experiment when additional measurements cease to provide new evidence. The autonomous trajectory also identifies a confounding relationship between polarization state and domain-wall proximity and recognizes that the requested contrast is not independently represented within the available field of view. These results demonstrate a route toward scientific autonomy in which agents determine what evidence is required while algorithms determine how that evidence is acquired reproducibly and within validated physical constraints.

cond-mat.mtrl-sci