arXiv ScienceSearch

arXiv subjects

Kai Liu

Publications and source records attributed to Kai Liu.

At least 19 recordsLinked to original sources

Sample-Guided Exact Top-K Selection for Long-Context Sparse Attention

Sparse attention bounds downstream attention work by retaining a fixed-size subset of indexed tokens, but its standalone exact Top-$K$ stage must still process materialized score rows whose length grows with context. Production radix selectors discover their first actionable boundary only after a complete-row pass, forcing another row-scale traversal before exact refinement. We observe that locating a compact upper tail requires substantially less resolution than identifying the exact rank boundary, and that fixed-stride partial views of the current row remain calibrated to the corresponding complete-row rank across ragged lengths. We present HPC-Ops Top-K, a sample-guided exact selector for ragged sparse-attention score rows. A fixed-stride view proposes a row-local coarse boundary; the mandatory complete-row pass certifies its sufficiency, forms the admitted candidate set, and initializes exact FP32 refinement over the unresolved frontier. A nested secondary boundary and exact recovery handle underfilled proposals before any output is committed, so sampling controls common-path work but never correctness. The GPU implementation fuses complete-row certification and candidate formation, and combines persistent, KV-split, and direct-exact execution behind graph-capturable ragged-row dispatch. We evaluate HPC-Ops Top-K on indexer scores from Hy4-Preview. It outperforms the fastest verified external exact baseline by $1.29$--$1.75\times$ across 20 operator configurations, with a $1.55\times$ geometric-mean speedup. It further achieves $1.36\times$ and $1.48\times$ speedups on two framework-derived sparse-attention traces. The implementation is available in HPC-Ops, Tencent's open-source high-performance operator library for LLM inference, at https://github.com/Tencent/hpc-ops.

cs.DC

Reproducible capillary fluctuation analysis of solid-liquid interfaces for stiffness and anisotropy calculations

The capillary fluctuation method (CFM) is widely used to compute solid--liquid interfacial properties from atomistic simulations, but its accuracy depends on choices in interface construction, wave-vector selection, sampling, and simulation geometry. Here, we develop a diagnostics-driven workflow for reproducible CFM calculations using pure Al as a representative system. Employing both ribbon models and thick two-dimensional references, we show that apparent linearity of the fluctuation spectrum alone does not ensure reliable stiffness or anisotropy estimates. Instead, a reliable CFM analysis requires a fitting window consistent with both temporal sampling and continuum capillary-wave assumptions, systematic sensitivity tests of the interface identification procedure, explicit propagation of replica variability, and independent verification of model-thickness convergence. We further propose a practical thickness-selection rule based on coexistence-temperature consistency, which enables finite-size effects to be controlled while retaining the substantial computational efficiency of ribbon geometries. By making the main sources of uncertainty explicit and diagnosable, the proposed workflow improves the reliability of the CFM as a quantitative tool and provides a foundation for its broader application to complex solid--liquid interfaces.

cond-mat.mtrl-sci

Frame phase retrievability and state distinguishability of quantum channels

This survey introduces the role of frame phase retrieval in pure-state identification and information preservation by quantum channels. For unit vectors, the lift $x\mapsto xx^*$ identifies vectors that differ only by a global phase with the same rank-one quantum state, converts frame intensities into linear functionals of the lifted state, and makes the adjoint pullback of output observables an operator-valued measurement on the input system. From this viewpoint, a channel is phase retrievable exactly when it is injective on pure states. We develop this correspondence through Choi-rank-two linear combinations of Kraus operators, higher-rank relative spectra, structural obstructions, and constructions with prescribed Choi rank. We distinguish injective identification from full tomography, perfect one-shot discrimination, zero-error classical communication, and exact quantum correction. Twirling channels then provide a structured setting in which commutants, irreducible dimensions, multiplicities, orbit-frame orthogonality, coding indices, and phase-retrievable subspaces can be read from group representations.

quant-ph

Strain-driven orbital-selective reconstruction and bicollinear-to-stripe evolution in FeTe

FeTe, as a representative parent material among iron-based superconductors, provides an ideal platform for exploring the interplay among orbital-selective correlations, magnetism, and unconventional superconductivity. However, a unified picture of the correlated electronic structure and magnetism of FeTe under strain remains to be fully clarified. Here, combining density functional theory plus dynamical mean-field theory and Heisenberg model analysis, we uncover an orbital-selective reconstruction of the correlated electronic structure and reveal a strain-driven trajectory from bicollinear to stripe antiferromagnetism (AFM) via an intermediate competing staggered $n$-mer AFM regime in FeTe. Moderate strain gives rise to a regime where more coherent quasiparticles coexist with suppressed local moments. Further strain drives FeTe into an incoherent correlated regime with robust local moments and Fe-$3d_{z^2}$-dominated low-energy states. These results establish a strain-driven trajectory across distinct magnetic and correlated electronic states in FeTe.

cond-mat.supr-con

CrossMambaTuning: Synergistic Spatial and Cross-Layer Adaptation for Machine Vision Compression

To reduce deployment cost and retraining overhead, adapting pretrained learned image compression (LIC) models to downstream machine vision tasks has attracted growing attention. However, existing methods typically insert fine-tuning modules independently into frozen backbones, lacking explicit mechanisms for cross-layer coordination. To address this limitation, we propose a novel framework named CrossMambaTuning, which integrates State Space Models with cross-layer interaction mechanisms for parameter-efficient fine-tuning. Specifically, we design an efficient Mamba adapter equipped with task-specific prompts and multi-scale branching to precisely capture both local features and global dependencies. Furthermore, we introduce a Scale-Invariant Cross-Layer Adapter (SICA) utilizing a parameter-sharing strategy to fuse task information across different scales and reduce redundancy. Extensive experiments demonstrate that CrossMambaTuning achieves state-of-the-art (SOTA) performance on multiple machine vision tasks, reducing parameter overhead by 72\% compared to SOTA methods. Code is available at https://github.com/rsr1123/CrossMambaTuning.

cs.CV

GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture

Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question whether current VLA systems can benefit from more effective architectural design, scale to substantially larger and more heterogeneous data regimes, and achieve broader generalization across tasks and embodiments. To this end, we present GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments. Specifically, GigaBrain-0.7 unifies understanding, prediction, and action through a three-system architecture, scales pretraining to over 37,000 hours of heterogeneous embodied data, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation. Compared with the preceding GigaBrain-0 series and prior state-of-the-art models including $\pi_{0.5}$, GigaBrain-0.7 achieves substantial improvements in foundation zero-shot capabilities, language-conditioned instruction following, and post-training task success rates. In particular, on our in-house Maker H01 platform and mainstream robot embodiments, GigaBrain-0.7 demonstrates strong task adaptability and completion ability across both home and industrial scenarios. All training code and pretrained model weights will be released.

cs.RO

Layer-Number-Controlled Symmetry Breaking and Surface-State Transport in Rhombohedral Graphene Multilayers

Rhombohedral multilayer graphene hosts layer-polarized flat bands, providing an intriguing platform for correlated and topological electronic states; however, the role of layer number in governing symmetry breaking and surface screening remains elusive. Here we prepare rhombohedral graphene multilayers and systematically conduct electrical transport measurements. We uncover an unconventional layer dependence of phase transitions: the critical displacement field (D$_{c}$) for the layer-antiferromagnetic (LAF)-to-semimetal transitions remains constant across tetralayer to hexalayer graphene, whereas the D$_{c}$ for semimetal-to-layer-polarized-insulator (LPI) transition increases with layer number, defying unscreened Coulomb interaction models. In hexalayer graphene, surface-state-dominated transport emerges, with Landau levels (LLs) and resistive peaks selectively controlled by adjacent gates, a signature of strong interlayer screening absent in thinner stacks. High magnetic fields reveal valley-layer-locked LLs and dissipative states possibly from interlayer backscattering, highlighting the presence of decoupled surface states. Our findings establish layer number as a key tuning knob for engineering correlated and topological phases in rhombohedral graphene multilayers.

cond-mat.mes-hall

VOICE: A Vision-Omics Foundation Model Integrating Direct and Retrieval-Based Prediction of In-situ Single-Cell Gene Expression

Spatial transcriptomics can resolve gene expression at single-cell resolution, but it is costly, limited to targeted panels of a few hundred to a few thousand genes, and applicable to only a small number of samples. H&E imaging, by contrast, is cheap and collected routinely at scale. This makes predicting single-cell expression directly from morphology a practical way to bring molecular analysis to large tissue archives. We therefore present VOICE, a multimodal foundation model that predicts single-cell gene expression from H&E images using paired Xenium data. VOICE first aligns cell centered H&E morphology from a pathology foundation model with single-cell expression embeddings from a transcriptome foundation model, trained using contrastive learning over 23 million cells. Next it predicts expression through two branches. One branch directly regresses expression from morphology. The other branch retrieves measured expression from similar reference cells, recovering genes that do not have morphological signal. Because genes vary in morphological predictability, VOICE fuses the two branches with a per-gene weight. After training, VOICE generalizes to heldout patients, slides, and partially overlapping gene panels from Xenium, and it consistently outperforms prior single-cell expression prediction methods on seven metrics.

cs.CV

KV-Skill: Forging Expertise in the Model's Native Language

Task knowledge is commonly stored either as text in the prompt or as an update to model weights. Text is modular but must be interpreted on every use, while weight adaptation makes the resulting capability difficult to load, remove, or share independently. We introduce KV-Skill, a design space of external factorized operators that a frozen language model reads through a lightweight interface. KV-Skill supports two complementary paths. Registration converts an authored text skill into a text-derived operator and trains a shared per-backbone interface. Reward learning develops a compact latent operator directly from task outcomes, with or without an authored skill. Neither path adds positions to the prompt. Across ten benchmarks and four backbones from three model families, converting text to a KV-Skill consistently makes the same procedural knowledge more effective. On Qwen3.5-4B LiveMath, registration reaches 77.2 accuracy, compared with 23.4 for the source text skill, 52.0 for SkillOpt, and 64.5 for SoftSkill. Under matched reward training and parameter budgets, KV-Skill gives the best result in seven of eight matched settings against soft prefixes, prefix tuning, and LoRA. A post-hoc rank analysis further shows that text-derived operators retain nearly all of their benefit with one task-aligned direction per injection layer, while matched random directions fail. Finally, one shared interface retains three independently loadable KV-Skills without measurable forgetting. These results show that task knowledge can be acquired from text or experience, compressed into an external operator, and deployed separately from the backbone. Code is available at: https://github.com/shawnzhg/KV-Skill

cs.LG

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning

Human-in-the-loop reinforcement learning (HIL-RL) enables robots to learn contact-rich manipulation from limited real-world interaction, but deployment exposes three coupled limitations: static visual reward models fail under scene changes; independently sampled actions cause temporally inconsistent motion; and vision-based policies remain sensitive to appearance shifts. We present EvoHIL, a unified framework that adapts the reward model, action generator, and visual do main within a staged human-in-the-loop learning process. First, self-evolving reward (SER) adapts the success classifier from human-confirmed positives and provisional weak negatives. Second, Action Flow Stabilization (AFS) generates temporally coherent action chunks through flow matching, grounding policy updates in executed action prefixes and demonstrated behavior. Third, retention-aware offline fine-tuning replays relit interaction data while anchoring the AFS actor-critic to prior behavior, adapting the visual domain without additional robot interaction. Across six manipulation tasks on Franka FR3 and SO-101 arms under a controlled lighting shift, EvoHIL improves task success, agreement with human-confirmation labels, motion smoothness, and completion time relative to human-in-the-loop and imitation baselines.Project page: https://anonymous4366.github.io/EvoHIL/

cs.RO

ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs

Visual token pruning reduces the inference cost of multimodal large language models, but a fixed token ratio is poorly matched to text-rich inputs. In OCR-centric tasks, decisive evidence can be a small number, label, or field whose relevance is specified by the question; indiscriminate pruning can erase that evidence while retaining visually salient but irrelevant regions. We present ET-Prune, a training-free framework that casts pruning as evidence allocation. It derives question-conditioned evidence from a decoder-side partial query-key block, safeguards text-like spatial regions, and converts evidence uncertainty and density into a sample-specific token floor. Three progressive middle-layer events then move the sequence toward this budget, retaining more tokens for diffuse or text-dense evidence and pruning concentrated evidence more aggressively. At the observed point estimates from one deterministic pass per configuration, ET-Prune leads or ties among pruned methods in all six backbone-benchmark comparisons at roughly half tokens. On OCRBench-v2, it leads the strongest pruned baselines by 1.80 and 0.68 percentage points on Qwen3-VL-8B and InternVL3.5-8B, respectively, while retaining about half of the visual tokens; on MMBench v1.1, it reaches 0.8467 circular exact-matching accuracy versus 0.8437 for Vanilla at 54.45% average visual-token retention. These results show a favorable observed quality-cost trade-off for evidence-aware dynamic budgeting in text-rich multimodal inference.

cs.CV

Screening phonon-mediated superconductors from static orbital Hamiltonians

The first-principles search for superconductors is severely limited by the high cost of electron-phonon coupling (EPC) calculations. Here we develop a low-cost, physically transparent framework that identifies strong-EPC materials directly from static orbital-based Hamiltonians without explicit phonon perturbation calculations. Verification using density functional perturbation theory (DFPT) for representative superconductors shows that the framework captures semi-quantitatively the EPC scale at substantially lower computational cost. Applied to more than 36,000 compounds in the MattKeyBond database, it identifies 34 dynamically stable superconducting candidates with calculated $T_c > 10$ K after DFPT verification. These candidates reveal two distinct routes to relatively high-$T_c$ superconductivity: a metallized covalent $\sigma$-bond route that is more favorable for achieving high-$T_c$ superconductors, and a Fermi-level density-of-states accumulation route that can enhance $T_c$ but usually to a more limited extent.

cond-mat.supr-con

Decoding the Micromagnetic Hamiltonian from Magnetic Fingerprints

Extracting intrinsic magnetic Hamiltonians directly from magnetometry is challenging due to the high dimensionality of the parameter space and the degeneracy induced by ensemble averaging. Here, we introduce a collection of deep convolutional neural networks (CNNs) to extract the full phenomenological micromagnetic Hamiltonian directly from the magnetic fingerprints encoded within First-Order Reversal Curves (FORCs). We validate this approach via closed-loop verification, re-creating the input magnetometry for both simulated and experimental FORCs. To mitigate false positives, we deploy an `Alice--Bob' parallel network that quantifies prediction uncertainty based on solely the information in FORCs without any additional ground-truth knowledge. This framework provides a robust, machine-learning-assisted approach to unravel the underlying spin behaviors in complex magnetic systems

cond-mat.mtrl-sci

Emergent d-wave altermagnetism in chlorine-adsorbed FeSe monolayer

The recent emergence of altermagnetism has opened new frontiers in condensed matter physics, yet material platforms capable of hosting both intrinsic altermagnetic order and superconductivity remain exceedingly rare. Here, based on symmetry analysis and first-principles calculations, we propose a realistic route to engineer robust altermagnetism in monolayer FeSe, a prototypical iron-based superconductor. By designing a stoichiometric Fe2Se2Cl structure through single-side Cl adsorption and introducing gate-tunable hole doping, we achieve a highly stable altermagnetic ground state. Our calculations reveal a synergistic mechanism: hole doping firmly stabilizes the checkerboard magnetic order, while the asymmetric ligand environment intrinsically breaks the outof-plane spatial inversion symmetry. Consequently, this interplay induces a giant altermagnetic spin splitting of up to 620 meV. Crucially, we demonstrate that this altermagnetic state and its giant spin splitting are highly resilient, persisting even in a 10-layer slab model that accurately simulates the bulk limit. By introducing altermagnetism into the well-established FeSe-based superconducting family, our findings identify Fe2Se2Cl as a promising platform for spintronic applications and motivate future studies of the possible interplay between altermagnetism and superconductivity.

cond-mat.mtrl-sci

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time

Steering robot foundation models (RFMs) toward new task variants or user-preferred behaviors remains challenging, often requiring additional robot demonstrations, task-specific fine-tuning, or long-context conditioning. We present WAM-TTT, a test-time training framework for steering world action models from raw human videos. Rather than treating human videos as trajectories to imitate, WAM-TTT absorbs them into a lightweight adaptive memory inside a frozen WAM through self-supervised video prediction. To make this memory useful for control, we introduce a meta-training stage that aligns human demonstrations with robot behaviors using paired human-robot data and a key--value memory reconstruction objective. At test time, only unlabeled human videos are required to adapt the memory, while the pretrained WAM remains frozen. This enables efficient and reusable steering without robot actions, human-side annotations, or task-specific fine-tuning, while preserving the generalization ability of the foundation model. Extensive experiments show that WAM-TTT consistently outperforms in-context human-video conditioning baselines across diverse manipulation tasks and generalization settings.

cs.RO

A Conservative Time-Accurate Local Time-Stepping DG Scheme Based on a Weakly Compressible Model for Unsteady Low-Mach-Number Flows

This paper presents a conservative high-order discontinuous Galerkin (DG) method featuring time-accurate local time stepping for simulating low-Mach-number unsteady flows, based on a weakly compressible formulation. In this model, pressure is defined solely as a function of density, eliminating the need for a global pressure Poisson equation typical of incompressible solvers while preserving the locality and conservation of compressible schemes. This makes it suitable for low-speed unsteady flows and aeroacoustics. The spatial discretization uses a strong-form nodal DG spectral element method (DGSEM) on Gauss-Lobatto-Legendre points. Inviscid fluxes are handled by numerical fluxes tailored to the weakly compressible system; specifically, a two-rarefaction approximate Riemann solver is developed for the constant-sound-speed barotropic equation of state. Viscous terms employ the incomplete interior penalty Galerkin (IIPG) method. For time integration, a continuous extension Runge-Kutta (CERK) scheme constructs cell-local predictor polynomials for continuous-in-time volume reconstructions. Face fluxes are split into interior and common contributions: the former matches the volume quadrature, while the latter uses piecewise Gaussian quadrature from continuous predictors. This split preserves discrete summation-by-parts cancellation and ensures conservative inter-element flux exchange.

physics.flu-dyn

Rotatable Antenna-Enhanced Secure Integrated Sensing and Communications Under Imperfect CSI

A rotatable antenna (RA)-enhanced secure integrated sensing and communications system is investigated, where an RA-based transceiver simultaneously communicates with legitimate users and senses a target that is regarded as a potential eavesdropper. Under imperfect eavesdropping channel state information (CSI), a max-min data rate optimization problem is formulated by jointly optimizing the transmit beamforming, artificial noise (AN) covariance matrix, and transmit/receive boresights of RAs, subject to the maximum information leakage and minimum sensing power constraints. To address the highly non-convex problem, the information leakage and sensing power constraints are transformed into convex ones via S-Procedure method and Cauchy-Schwarz inequality, respectively. Subsequently, an alternating optimization algorithm is developed to decompose the reformulated problem into two subproblems. In particular, the transmit beamforming and AN covariance matrix are optimized by utilizing successive convex approximation and semi-definite relaxation methods, while the RA boresights are obtained by invoking the particle swarm optimization. Simulation results show that the RA-based scheme significantly outperforms the benchmarks, and offers enhanced robustness against imperfect CSI with the increase of the maximum rotation range.

eess.SP

Operator Learning for PDE Backstepping Control of Parabolic Equations on Time-Varying Domains

This paper develops a learning-based boundary control framework for stabilizing a parabolic equation defined on time-varying spatial domain. Although the partial differential equation (PDE) backstepping method provides a systematic theoretical framework for such moving-boundary systems, its real-time implementation is hindered by the need to repeatedly solve time-varying kernel PDEs on evolving domains. To overcome this limitation, we first formulate the time-varying backstepping design as an operator that maps the moving-boundary trajectory to the corresponding backstepping kernel. By mapping the time-varying domain of the backstepping kernel equation onto a fixed reference domain, we establish the continuous dependence of the kernel on the moving-boundary trajectory, which provides the theoretical basis for approximating the backstepping design operator by a neural operator. Based on the approximate kernel operator, we construct the corresponding boundary feedback controller to stabilize the system. It is shown that the closed-loop system admits an exponential decay estimate on any prescribed finite time interval. For numerical implementation, DeepONet is employed to learn the time-varying kernel operator from offline-generated numerical kernel solutions and is subsequently deployed online to generate the required time-varying kernels without repeatedly solving the kernel PDE. Numerical benchmarks demonstrate that the proposed neural-operator-based implementation bypasses repeated online solution of the time-varying kernel PDE, achieves a significant acceleration of close to three orders of magnitude compared with conventional numerical kernel solvers, and thus enables real-time stabilization of the system on time-varying spatial domain.

math.OC