arXiv Science⌕ Search

arXiv subjects

Yu Liu

Publications and source records attributed to Yu Liu.

At least 37 records · Page 2Linked to original sources

Bare-Die Antiferromagnetic Computing

Semiconductor electronic devices are increasingly constrained by fundamental quantum tunneling effects and charge-based mechanisms, which severely limit further miniaturization, write-speed scaling, and environmental robustness of silicon-based technologies. These limitations are particularly prohibitive for deep-space exploration, where extreme temperatures, ultra-strong magnetic fields, and intense radiation rapidly incapacitate conventional electronics without massive shielding. Here, we present an intrinsically resilient, strain-mediated antiferromagnetic MnIr/PMN-PT edge processor that operates reliably as a bare die under temperatures up to 500 K, magnetic fields of 55 T, and radiation doses of 1.5 Mrad. By exploiting an input-modulated in situ self-refreshing encoding mechanism, the device performs nonlinear feature extraction and classification directly from raw analog signals, enabling an analog computing architecture that requires no time-frequency transformation. This architecture achieves 99.8% accuracy in speech recognition without digital preprocessing and 100% accuracy in astronaut visual object recognition. Furthermore, an all-hardware integrated drone vision system demonstrates real-time in situ command execution and autonomous navigation, delivering a terahertz-level response frequency and an ultra-low energy consumption of approximately 0.2 fJ per operation. This work expands the functional scope of antiferromagnetic devices beyond memory and logic, establishing them as a promising materials platform for energy-efficient physical computing and autonomous intelligence in extreme environments.

cond-mat.mes-hall↗

DiCoR: Decoupled Referent Disambiguation and Contour Recalibration for Efficient Referring Remote Sensing Image Segmentation

Referring remote sensing image segmentation (RRSIS) aims to delineate targets specified by natural language expressions in remote sensing imagery. Existing methods mainly follow joint fusion segmentation (JFS) or decoupled prompt segmentation (DPS). JFS is efficient but often suffers from limited accuracy because referent localization and mask delineation are optimized under a unified objective, whereas DPS separates localization from mask generation using spatial prompts and foundation segmenters at the cost of higher memory consumption and inference latency. To bridge this gap, we propose DiCoR, a decoupled referent disambiguation and contour recalibration framework built on an efficient JFS pipeline. DiCoR addresses two key challenges: distinguishing the correct referent from ambiguous candidates and refining coarse masks after localization. A disambiguation-aware localization guidance strategy ranks salient candidate regions with adaptive linguistic cues and injects the resulting localization prior into fused features. A lightweight contour recalibration module further predicts residual corrections to coarse logits under localized contour supervision, improving mask quality with limited computational overhead. Experiments on RefSegRS, RRSIS-D, and RISBench show that DiCoR achieves the best segmentation accuracy across all three benchmarks. On RefSegRS, it improves mIoU and gIoU by 4.90% and 2.16% over a competitive JFS method while running 2.4-6.5 times faster than the evaluated DPS method, demonstrating a favorable accuracy-efficiency trade-off. Code is available at https://github.com/zyGao1126/DiCoR.

cs.CV↗

Bond-selective modulation of scalar spin chirality in strained Mn4N

Scalar spin chirality (SSC) underlies a variety of Berry-phase-driven transport phenomena in noncoplanar magnets. Mn4N is a unique ferrimagnetic system in which both collinear and noncoplanar magnetic configurations have been reported at different lattice parameters, featuring vanishing and finite SSC, respectively. However, the microscopic mechanism governing the competition between these magnetic states and the associated modulation of SSC remains unclear. This issue is particularly important for Mn4N because its high magnetic ordering temperature (TN = 740 K) provides an attractive platform for exploring chiral magnetic states and their associated topological functionalities at elevated temperatures. Here, using first-principles calculations, we investigate the strain-driven evolution of the magnetic ground state and SSC in Mn4N and uncover the microscopic origin of the collinear-to-noncoplanar magnetic transition. We demonstrate that tensile strain continuously stabilizes the noncoplanar configuration and enhances SSC, as quantified by the magnitude of the chirality-order vector. Orbital-resolved crystal orbital Hamilton population and charge-density-difference analyses reveal that strain selectively weakens the Mn3c-N hybridization while preserving direct Mn3c-Mn3c interactions, leading simultaneously to the activation of the Mn3c in-plane magnetic moment and the suppression of N-mediated ferromagnetic superexchange between nearest-neighbor Mn3c atoms. This bond-selective electronic response governs the magnetic ground state and consequently controls the emergence and enhancement of SSC in Mn4N. Our work establishes bond selectivity as an effective strategy for engineering SSC in high-temperature ferrimagnetic materials.

cond-mat.mtrl-sci↗

Hidden magnetic order within the pressure induced superconducting dome of UTe2

Unconventional superconductivity typically occurs near magnetic instabilities, and the corresponding spin fluctuations are widely believed to play a crucial role in mediating electron pairing. UTe$_2$ is a promising candidate for exhibiting multiple spin-triplet superconducting phases when tuning with applied pressure and magnetic fields, but the nature of the magnetism driving these unconventional pairing states is undetermined. Our measurements of UTe$_2$ under applied pressures and magnetic fields reveal the presence of a magnetic order hidden within the pressure-induced superconducting dome, which vanishes together with the superconductivity once there is sufficiently high pressure to induce the three-dimensional antiferromagnetic phase. Extrapolation of the phase boundary of the hidden magnetic order, which is most likely antiferromagnetic in nature, points to a zero-temperature quantum critical point that coincides with the maximum transition temperature of the pressure-induced superconducting dome, suggesting that it could corresponds to the parent magnetic phase of the critical antiferromagnetic spin fluctuations driving the triplet superconductivity. These findings advance the understanding of the interplay of magnetism and superconductivity in an exemplar candidate triplet superconductor, which is necessary for revealing the microscopic origin of the different unconventional superconducting phases.

cond-mat.supr-con↗

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Video generation begins in text space by authoring a cinematic screenplay, then materializes into pixels. As contemporary video generators scale to 30 seconds and faithfully follow complex conditions, the textual prompt largely directs the production, planning how actions, camera trajectories, lighting, and sound unfold across multi-shot sequences. In this paper, we present WanPE, a 397B-parameter prompt enhancement model trained on 1.05M real-world videos to master director-level cinematic planning. WanPE formulates shot-level cinematic plans via video-grounded reverse construction and employs Semantic-Consistency GRPO (SC-GRPO) to faithfully preserve user requirements across shots and over time. To benchmark this capability, we curate WanPEval, a human-annotated testbed covering durations from 5 to 30 seconds across varying intent granularities, supported by approximately 11K blind pairwise assessments. When powering Wan3.0's video generator, WanPE-397B boosts human preference over raw user prompts by 10.66-18.84 points at 5-15 seconds and by a dramatic 50.86 points in the 30-second arena. Ablation studies show that reverse construction demonstrates clear superiority over forward rewriting, while SC-GRPO robustly preserves semantic fidelity across model scales. Ultimately, WanPE leads all evaluated commercial offerings at 5-15 seconds and remains competitive with Seedance 2.5 at 30 seconds.

cs.CV↗

Coupling Perception and Reasoning in Federated Multimodal Graph Foundation Models

Federated multimodal graph foundation models (GFMs) aim to adapt pretrained multimodal models to decentralized graph data, where each client owns a private multimodal graph and cannot share raw information. These models typically combine a multimodal Encoder that extracts semantic evidence from heterogeneous modalities and a graph neural network (GNN) that performs relational reasoning over graph structures. However, existing federated GFM adaptation methods mainly update graph-side modules while keeping the multimodal Encoder frozen, limiting adaptation to \emph{how information is propagated} while fixing \emph{what information is extracted}. Through empirical studies, we reveal that Encoder and GNN adaptations are not independent: Encoder adaptation is affected by graph relations, while cross-client module swapping reveals substantial pairing sensitivity between separately parameterized Encoder and GNN updates. Motivated by this observation, we propose \textbf{FedCORE}, a federated adaptation framework that represents Encoder and GNN updates through a shared low-dimensional latent state. FedCORE jointly optimizes this core from multimodal and structural signals and performs federated evolution directly in the shared state space, preserving compatibility between perception and reasoning adaptations. Extensive experiments demonstrate that FedCORE reduces the Encoder--GNN pairing gap from $30.6$ to $5.9$, corresponding to an $80.7\%$ reduction over independent joint adaptation.

cs.LG↗

MachEmbodied-U0: Unified Understanding and Generation Model for Embodied Intelligence

General-purpose robot control requires models to understand task intent, identify where to interact, capture how the scene evolves, and generate precise actions. Vision-language-action models provide strong semantic priors but typically do not explicitly model scene dynamics, while world-action models couple visual prediction with control without necessarily exposing the task-relevant semantic and spatial structure needed for fine-grained manipulation. We present MachEmbodied-U0 (ME-U0), a unified embodied foundation model connecting understanding and generation experts through a Mixture-of-Transformers architecture. Subtask prediction and affordance grounding guide joint visual-dynamics and action generation via flow matching. Visual dynamics encompass future RGB, depth, surface normals, and optical flow, providing complementary supervision for appearance, geometry, and motion. Multi-rate Rotary Position Encoding (MRPE) aligns visual dynamics with fine-grained control. We pretrain ME-U0 on approximately 4,200 hours of curated demonstrations from robotic datasets and egocentric datasets. Using only the supervision natively available in each downstream benchmark, ME-U0 achieves an average score of 17.66 on the RoboDojo simulation benchmark and average success rates of 99.0\% and 82.5\% on LIBERO and LIBERO-Plus, respectively. We additionally validate ME-U0 on real-world robotic manipulation tasks, demonstrating its effectiveness beyond simulation. Without corresponding downstream supervision, ME-U0 further demonstrates zero-shot subtask prediction, affordance grounding, and visual dynamics on simulated and real-world observations. Overall, ME-U0 combines competitive downstream control performance with transferable task-grounding and visual-dynamics capabilities across simulation and the real world.

cs.RO↗

Code Plans, Diffusion Renders: Open-Ended Generative World Modeling

We introduce \textbf{CoDeR}, a new paradigm for world modeling. Unlike existing video world models that implicitly represent world dynamics through visual observations, our system explicitly constructs an executable world with code and employs video generation models for visual realization. Specifically, we coordinate five complementary roles to translate high-level concepts into structured world rules, executable dynamics, and perceptual observations. This design enables \textit{long-term memory}, \textit{open-ended interactions}, \textit{autonomous world evolution}, and \textit{multi-agent scenarios}, where multiple entities can act, interact, and evolve persistently beyond the current observation. Extensive experiments demonstrate that our framework substantially extends the capabilities of existing world models, enabling long-term memory, open-ended interactions, autonomous evolution, and persistent multi-agent dynamics, while achieving state-of-the-art performance across multiple evaluation settings. Code and model weights will be made publicly available. Project Page: \href{https://becauseimbatman0.github.io/CoDeR}{CoDeR}.

cs.CV↗

ME-Dex 1.0: Bringing Heterogeneous Tactile Sensing into World Action Modeling

World Action Models bring the predictive capabilities of video models into robot action generation, providing a rich foundation for modeling future visual states. Tactile sensing complements this foundation with direct measurements of physical interaction. Some existing methods use tactile features as conditioning inputs without jointly predicting future tactile states, visual observations, and actions. Our key insight is that tactile signals, like video, provide observations of the evolving world state and should be modeled as future observations alongside video. We present ME-Dex-1.0 (MachEmbodied-Dex-1.0), a unified World Action Tactile Model for joint visual, tactile, and action learning. ME-Dex-1.0 adopts a Mixture-of-Transformers architecture comprising a Video Expert, a Tactile Expert, and an Action Expert, all trained with flow matching. We use shared attention connects the experts in intermediate layers, allowing action generation to draw on learned representations of visual and tactile dynamics during joint denoising. To support multi-source heterogeneous tactile inputs, a Canonical Hand Model and a Unified Tactile Autoencoder map tactile observations from different embodiments and sensing layouts into shared spatial and latent spaces. To address the limited availability of paired visual, tactile, and action data, we develop the Agentic Tactile Data Engine, an agent-based data production platform. It supplements RoboTwin and DexJoCo with tactile data recorded directly from force sensors during trajectory replay in simulation. Experiments on the RoboTwin, DexJoCo, and ManiFeel simulation platforms, together with real robot evaluations, demonstrate improved manipulation performance using both grippers and dexterous hands equipped with tactile sensing.

cs.CV↗

Optical Constants of Photochemical Haze Analogs in N2-CH4-CO Atmospheres from 0.4 to 28.6 μm

Photochemical hazes play an important role in shaping the spectra and radiative balance of N2-dominated planetary atmospheres. We present newly acquired FTIR and retrieved optical constants (N=n+ik) of laboratory-generated haze analogs from N2/CH4 and N2/CH4/CO gas mixtures under plasma discharge conditions. The retrievals use particle densities newly measured for the CH4-series and previously published particle densities for the CO-series. The experiments systematically explored CH4 concentrations from 0.5% to 10% and CO concentrations from 0% to 5% with fixed 5% CH4. Using measured particle densities together with the Beer-Lambert law and subtractive Kramers-Kronig (SKK) relation, we derived optical constants over the 350-25000 cm-1 (0.4-28.6 μm) spectral range, with the 0.4-25 μm results presented in the main text. The infrared spectra reveal prominent absorption features associated with hydrocarbon-, nitrogen-, and oxygen-bearing functional groups. Increasing CH4 abundance enhances aliphatic hydrocarbon features and corresponds to decreasing particle density, whereas increasing CO abundance promotes oxygen incorporation, broader mid-infrared absorptions, and higher particle density. The derived k spectra exhibit strong absorptions near ~3 μm, ~4.6 μm, and ~6-10 μm, while the real refractive index n generally ranges from ~1.2 to 1.7. The controlled CH4- and CO-series establish composition-dependent variations in haze optical properties. A benchmark comparison among Titan-, Pluto-, and Triton-like haze analogs then uses these experimentally identified trends to interpret the optical differences among N2-dominated planetary haze compositions. The density and optical constants provide laboratory constraints for atmospheric radiative transfer models and for interpreting planetary and exoplanetary spectra from spacecraft and telescopes.

astro-ph.EP↗

OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation

Reference-to-video (R2V) generation is evolving toward increasingly general and versatile reference control, giving rise to the emerging paradigm of omni R2V generation. However, existing benchmarks fall short of these emerging capabilities: their test cases cover limited reference types and compositions, and their evaluation protocols largely assess holistic reference consistency, overlooking whether reference factors are properly preserved, disentangled, and routed. Meanwhile, the high cost of constructing omni R2V training data makes suitable training resources scarce. To address these gaps, we introduce OmniVBench and the Omni-R2V Dataset for evaluating and training omni R2V models. OmniVBench expands R2V evaluation across broader reference types, fine-grained control tasks, and richer reference compositions, covering 7 task families and 18 fine-grained tasks spanning content, motion, style, structure, narrative, and multi-reference settings. We introduce factor-grounded evaluation with 12,172 case-specific checklist items, assessing whether intended reference factors are faithfully preserved, correctly disentangled and bound to their targets, and properly realized according to the instruction. We further introduce the Omni-R2V Dataset, bringing industrial-grade training resources for diverse R2V tasks to the broader research community. Drawing primarily on a large-scale corpus of professional video footage, it comprises 340K processed training samples spanning diverse reference types and multi-reference compositions. We develop task-specific pipelines for reference-target pair construction, offering a practical and scalable recipe for omni R2V data construction. Extensive evaluation of advanced open- and closed-source R2V models reveals clear performance gaps across task families and evaluation dimensions on OmniVBench, highlighting remaining limitations of current R2V models.

cs.CV↗

Tripartite Form Universality in Holographic Entropy Inequalities

To elucidate the meaning of holographic entropy inequalities (beyond subadditivity) which characterize the entanglement structure of geometric states in holography, arXiv:2309.06296 proposed the "tripartite form" for these inequalities, consisting of tripartite information and conditional tripartite information terms with unit coefficients. While this provides a compact and useful packaging of the inequalities, it is not a priori guaranteed that all inequalities can be recast in this form. Here we conjecture that they can, and present substantial evidence, by proving that the two known infinite families of holographic entropy inequalities found in arXiv:2309.15145 can indeed be written in the tripartite form. This is significant because such recasting is particularly nontrivial for these families. Apart from providing the explicit tripartite form expressions for every member of these two infinite families, we detail how we arrived at them, in the process deriving several useful identities which may serve as stepping stones to formulate further repackaging of the corresponding information quantities. We also illustrate the power of the tripartite form by proving a number of structural properties satisfied by any holographic entropy inequality.

hep-th↗

TorchCraft: Unified binder design by inverting an all-atom structure predictor

All-atom structure predictors model diverse molecular interactions, but using their learned structural priors for binder design remains challenging. Here we present TorchCraft, a unified binder-design framework that optimizes sequence logits through a frozen all-atom predictor. Implemented in TorchFold, TorchCraft combines confidence, contact, geometric, and sequence-prior objectives within a shared optimization procedure for minibinders, framework-conditioned VHHs, cyclic peptides, and ligand-binding proteins. Using pretrained AlphaFold 3 weights, TorchCraft generated representative minibinders and VHHs with experimentally measured binding across four targets in each format, without post hoc sequence redesign. Computational benchmarks further demonstrated the framework's applicability to cyclic peptides and ligand-conditioned pocket design. TorchCraft extends predictor inversion to multiple binder formats and molecular contexts, providing a common framework for reusing all-atom structural priors in design.

cs.AI↗

Identical-Particle Symmetry-Enabled Complete Coherent Control of Ultracold Atomic and Molecular Collisions

We show that exchange symmetry in collisions of identical particles enables symmetry-protected coherent control of the total scattering cross section. For identical fermions, antisymmetrization enforces a common control phase among contributing channels within a given parity sector, yielding maximal control visibility. For identical bosons, phase-locking persists but with reduced visibility due to additional exchange (satellite) contributions. Collisions of distinguishable particles lack this symmetry-imposed phase-locking, leading to lower controllability and visibility. We elucidate these principles through coupled-channel quantum-scattering calculations for lithium-lithium collisions, comparing the $^{6}\mathrm{Li}{-}^{6}\mathrm{Li}$ (identical fermions), $^{7}\mathrm{Li}{-}^{7}\mathrm{Li}$ (identical bosons), and $^{6}\mathrm{Li}{-}^{7}\mathrm{Li}$ (distinguishable) systems. Furthermore, in the identical-particle cases, symmetry-enforced phase-locking enables full control over the parity of the final state even beyond the ultracold regime. This mechanism is broadly applicable to identical-particle collisions, including homonuclear molecules for which established approaches--DC electric fields or microwave shielding--are ineffective or unavailable.

physics.atom-ph↗

Recursive Reasoning or Statistical Extrapolation? In-Context Learning in Multi-Agent Interdependent Decision-Making

In-context learning (ICL) enables large language model (LLM) agents to improve decisions using interaction history, yet it remains unclear whether such improvement reflects refined internal reasoning or mere extrapolation of statistical patterns. To disentangle these mechanisms, we study LLM agents in multi-agent incomplete-information games that require recursive belief reasoning. By constructing a public goods game and manipulating the statistical structure of historical feedback, we evaluate decision quality against a history-independent rational expectations equilibrium (REE) benchmark. Our experiments reveal that when historical statistical patterns are disrupted, the benefits of longer context largely vanish, degrading decision quality to the no-context baseline in a way sharply amplified by stronger strategic interdependence. These results suggest that, in such strategic environments, ICL behavior is more consistent with statistical extrapolation than with strategic reasoning. Our work extends the mechanistic study of ICL to strategic multi-agent settings, introduces REE as a diagnostic tool for distinguishing reasoning from extrapolation, and provides a reusable framework for probing the boundaries of LLM reasoning in recursive belief tasks.

cs.AI↗

PersonaPath: Towards Knowledge-Centric Personalized Learning Path Planning

Adaptive learning systems commonly formulate learning path planning as Exercise-Centric (EC) recommendation, where the next step is inferred from item-level interaction logs. Evaluating goal-oriented guidance additionally requires explicit learner goals and curriculum-scale prerequisites: learners with similar exercise records may need different paths toward their targets. We therefore study Knowledge-Centric (KC) personalized learning path planning, where a planner must reason over learner profiles, mastery states, and prerequisite knowledge structures to decide which textbook, unit, and concept should be studied next. To support this setting, we introduce PersonaPath, a benchmark that pairs 2,000 fine-grained learner personas with a hierarchical knowledge graph of 347 textbooks, 1,751 units, and 4,092 concepts across 77 subjects. We evaluate representative LLMs on PersonaPath. Results show that even the strongest LLM reaches only a 29.5% final pass rate in Basic Education, and that the main bottleneck lies in adaptivity, where no model exceeds 44.7% in tailoring paths to individual learners.

cs.CL↗

Hidden Magnetic Complexity Within a Simple van der Waals Ferromagnet Ce$_2$Te$_5$

Ce$_2$Te$_5$ is a layered $f$-electron van der Waals magnet in which reduced dimensionality and inequivalent Ce sites give rise to competing magnetic interactions. We investigate its magnetic ground state using muon spin relaxation ($μ$SR), neutron powder diffraction (NPD), and inelastic neutron scattering (INS). Zero-field $μ$SR reveals an onset of static magnetism below $T_{\mathrm{C}} = 5.0(1)$~K, followed by an additional anomaly in the internal field at $T_{\mathrm{2}} = 2.3(2)$~K, consistent with features observed in bulk thermodynamic and transport measurements. In contrast, NPD data collected between $0.05-8$~K reveal a single long-range ordered magnetic phase below $T_{\mathrm C}$, with the magnetic Bragg intensities vanishing at $T_{\mathrm C}$ with no evidence for additional structural or magnetic phase transitions down to base temperature. The ordered state is characterized by a commensurate propagation vector $\mathbf{k}=(0,0,0)$ and ferromagnetic alignment of Ce moments along the crystallographic $b$ axis. Remarkably, only one of the two crystallographically distinct Ce sites carries an ordered moment of $\sim 0.85(3)~μ_{\mathrm B}$ per Ce at 0.05~K. INS measurements establish the crystal electric field (CEF) energy scale of Ce$^{3+}$, revealing low-lying excitations at 7.94 and 21.46~meV and a Kramers doublet ground state with strong single-ion anisotropy. These results demonstrate that Ce$_2$Te$_5$ undergoes a single symmetry-breaking magnetic transition, while additional low-temperature anomalies reflect subtle modifications of the ordered state driven by competing interactions and CEF effects.

cond-mat.str-el↗

First Demonstration of Flip DRAM from Process, Architecture to System to Push DRAM Scaling beyond 4F2: 2F2 Self-aligned Flip Vertical Channel Transistor (FVCT) DRAM and Flip WL (FWL) 3D-DRAM

For the first time, we proposed a novel stacking technology for DRAM scaling by flipping and backside processes, making full use of DRAM wafer's backside and investigating it on both 4F2 and 3D-DRAM. For 4F2 VCT, 2F2 Flip VCT featuring self-aligned back-to-back stacked 1T1C bitcell, with various BL and WL configurations, were studied and key process modules such as self-aligned stacked vertical channel, BL and WL formations, wafer bonding and flipping, substrate thinning and low-R Co storage node (SN) were successfully developed, addressing the potential thermal, misalign and parasitic concerns in the Flip VCT process. A full DRAM DTCO framework was also established from device to mat and chip level. Compared to 4F2 VCT DRAM with the same mat size, 2F2 FVCT delivers 27.5% less parasitics, 11% better sense margin, 16.3% higher charge sharing (CS) speed and 50% less area. For 3D-DRAM, a brand-new flip WL staircase design with peripheral circuit innovations was studied and proved to have 25% density gain, 15.1% faster turn-on speed and 6.8% less CS time, proving further extendibility of flip technology on DRAM.

cond-mat.mes-hall↗