arXiv ScienceSearch

arXiv subjects

Kai Xu

Publications and source records attributed to Kai Xu.

At least 19 recordsLinked to original sources

InterMASH: A Unified Geometric Representation for Grasp Synthesis

Grasp synthesis aims to generate stable and physically plausible hand--object interactions, and has become a fundamental problem in both human hand modeling and robotic manipulation. However, a unified representation across human and robotic hands is still lacking, mainly due to differences in hand morphology and surface modeling. Prior methods typically rely on either contact maps or dense implicit descriptors to represent interaction, but these representations are often incomplete or computationally expensive and redundant. We propose InterMASH, a unified geometric representation that establishes cross-embodiment correspondence using sphere-fixed anchors. At each anchor, low-degree spherical harmonics compactly encode local hand geometry, object geometry, and contact, forming an explicit and interpretable token sequence. Building on this natively tokenized structure, we introduce a conditional Diffusion Transformer that operates directly in the proposed InterMASH representation space and jointly generates hand geometry and contact, improving consistency and physical plausibility. Our method achieves competitive performance with state-of-the-art methods on key physical feasibility metrics in a large-scale ShadowHand benchmark, supports joint training across multiple hands, and shows that cross-embodiment fine-tuning with human grasp data can improve robotic grasp success and diversity. Project page is available at https://inter-mash.github.io/.

cs.RO

Pulse-Level Compilation of Measurement-Free Recovery in Transmon Circuits

Quantum error correction commonly relies on syndrome measurement, decoding, and conditional feedback. We numerically show that the conditional spectrum of an interacting transmon circuit can compile a local recovery rule into a fixed open-loop control cycle. In a four-bit repetition-code ring, the resulting input-independent control slows the decay of logical coherence under Pauli-\(X\) noise and stabilizes logical-one population under data relaxation relative to uncorrected references. Its local, bounded-degree architecture provides a hardware-native framework for extending compiled recovery control to larger quantum networks.

quant-ph

CGGT: Curve-Grounded Geometry Transformer for 3D Parametric Curve Reconstruction

Recovering editable 3D parametric curves from 2D images is a fundamental challenge in computer graphics, bridging pixel-based perception and vector-based CAD modeling. Existing NeRF- and 3DGS-based methods often rely on dense calibrated views, precomputed 2D edge maps, and costly per-scene optimization, limiting their applicability to casually captured real-world inputs. We propose CGGT, a Curve-Grounded Geometry Transformer that directly grounds 3D-consistent 2D curve instances in the image space from sparse, unposed multi-view images. CGGT combines a geometry-aware transformer encoder for multi-view feature learning with a curve-aware masked-attention decoder for cross-view instance association. In a single forward pass, it predicts camera parameters, dense depth maps, and instance-level 2D curve masks, which are then lifted into 3D and refined through a fast parametric optimization stage to recover compact, editable 3D curve primitives. To support structured curve learning, we introduce Wireframe-100K, a large-scale dataset comprising 100,000 CAD models with diverse topologies, realistic multi-view renderings, and accurate parametric curve annotations. Extensive experiments show that our framework achieves substantial improvements in both reconstruction accuracy and efficiency, particularly under challenging sparse-view settings and in separating persistent 3D structural edges from view-dependent image edges caused by silhouettes, textures, and appearance variations. Despite being trained solely on synthetic data, CGGT generalizes well to real-world images, demonstrating its potential for practical CAD-style wireframe reconstruction from unconstrained visual inputs.

cs.CV

Event-triggered Implicit Perturbation for Zeroth-Order Fine-Tuning of Spiking Transformers

Zeroth-order (ZO) optimization estimates gradients using only forward-pass evaluations, making it suitable for fine-tuning non-differentiable, event-driven spiking neural networks (SNNs). However, its deployment on in-memory computing (IMC) accelerators is constrained by the repeated read-modify-write (RMW) operations arising from explicit weight perturbation and the prohibitive hardware footprint of random number generators (RNGs) for statistically independent per-weight perturbations. To address these challenges, we propose an implicit-perturbation ZO (IPZO) architecture in which perturbation sums computed by an event-triggered perturbation generation unit (PGU) are combined with the weighted sums produced by the IMC array, eliminating perturbation-induced RMW operations while preserving weight-stationary execution of IMC. By exploiting spike sparsity, the PGU generates and accumulates perturbation contributions only for spike-activated weight rows, reducing the required row dimension of the RNG array. An address-driven XOR recombination scheme (PGU-XOR) is further introduced to mitigate the spatial correlations caused by direct RNG reuse (PGU-Reuse). The results show that (1) PGU-XOR matches software RNGs in accuracy on Spikingformer/CIFAR-10 (76.41% vs. 76.53%) and perplexity (PPL) on SpikeGPT/WikiText-2 (54.20 vs. 53.23), whereas PGU-Reuse degrades accuracy by 9.56 percentage points and increases PPL by 11.8; (2) implemented in a TSMC 16-nm CMOS technology, PGU-XOR incurs 40.3%-46.0% area and 15.2%-48.9% energy overhead per matrix-vector multiplication relative to PGU-Reuse, yet its faster convergence reduces the total perturbation energy to 0.51x that of PGU-Reuse at iso-accuracy; (3) IPZO reduces the perturbation energy to 0.46x-0.83x that of conventional explicit weight perturbation for a batch size of B=64 and T=4 time steps, with the advantage growing as BT decreases.

cs.AR

Inference-Time Scaling of Diffusion Language Models via Trajectory Refinement

Discrete diffusion models have recently emerged as strong alternatives to autoregressive language models, matching their performance through large-scale training. However, inference-time control remains relatively underexplored. In this work, we study how to steer generation toward desired rewards without retraining the models. Prior methods typically resample or filter within a single denoising trajectory, optimizing rewards step-by-step without trajectory-level refinement. We introduce particle Gibbs sampling for diffusion language models (PG-DLM), an inference-time algorithm enabling trajectory-level refinement. PG-DLM constructs a Markov chain over full denoising trajectories and applies a conditional sequential Monte Carlo kernel to resample them. By doing so, PG-DLM introduces a new scaling axis, the number of refinement iterations, which is unavailable to prior methods. Increasing iterations remains effective even as gains from adding more parallel samples saturate. Furthermore, PG-DLM enables adaptive compute allocation by performing additional iterations only when needed, leading to further efficiency gains. We derive theoretical guarantees for convergence and variance bounds, and analyze trade-offs across different scaling axes. Empirically, PG-DLM outperforms prior methods across compute budgets on reward-guided generation tasks. On GSM8K, it achieves 90.07% accuracy with 2.9 particles on average and 94.47% accuracy with 24.8 particles on average.

cs.LG

Intersection Bounds for BPS Strings in Six-Dimensional Supergravity

In six-dimensional $\mathcal{N}=(1,0)$ supergravity, the structure of tensor moduli space is governed by primitive BPS string charges known as BPS generators and their intersection pairing. We derive bounds on the intersection numbers of these generators from a purely effective field theory (EFT) perspective. Although gauge anomaly cancellation constrains intersections between generators supporting gauge algebras, bounds for E-strings intersecting generators with self-intersection numbers $-2$ and $-3$ have previously remained incomplete. We show that the Zariski decomposition, interpreted as the charge lattice counterpart of the attractor mechanism, together with current algebra embeddings on the E-string worldsheet theory, yields strong universal bounds on these intersection numbers. These results establish the finiteness of tensor charge intersection numbers up to duality. The underlying structure was identified through AI-guided investigation and is proven here analytically using EFT arguments.

hep-th

From Blind Search to Memory-Aware Evolution: Efficient DBMS Tuning via Collaborative Diagnosis and Utility-Aware Retrieval

Modern DBMSs expose multiple configurable components (e.g., knobs, query hints, and indexes) that jointly determine query performance. Multi-component tuning is challenging due to the large combinatorial search space and the difficulty of learning effective tuning policies under limited feedback. Existing approaches still rely on blind search over the configuration space and interaction-heavy policy learning, leading to high tuning overhead and limited performance gains. Recent advances in large language models (LLMs) enable knowledge-driven tuning, but existing LLM-based methods fail to effectively exploit online feedback and historical observations, often converging prematurely to suboptimal configurations. In this paper, we present EvoTune, a memory-aware evolution framework for multi-component DBMS tuning. EvoTune first localizes a query-specific high-impact subspace via collaborative diagnosis, which combines lightweight pattern learning with LLM-based reasoning. It further introduces a utility-aware retrieval policy that selects informative observations based on their resulting long-term performance improvement, instead of similarity-based retrieval. To support continual improvement, EvoTune organizes tuning feedback into a hierarchical memory and incrementally refines both subspace localization and tuning policies without requiring LLM fine-tuning. Extensive experiments show that EvoTune consistently outperforms state-of-the-art baselines, achieving up to 44.5% performance improvement under the same tuning budget and reaching the best competing baseline's final performance up to 3.9X faster.

cs.DB

Drawstrings and flexibility in the Geroch conjecture

In this paper, we observe new phenomena related to the structure of 3-manifolds satisfying lower scalar curvature bounds. We construct warped-product manifolds of almost nonnegative scalar curvature that converge to pulled string spaces in the Sormani-Wenger intrinsic flat topology. These examples extend the results of Lee-Naber-Neumayer \cite{LNN} to the case of dimension $3$. As a consequence, we produce the first counterexample to a conjecture of Sormani \cite{SormaniConj} on the stability of the Geroch Conjecture. Our example tests the appropriate hypothesis for a related conjecture of Gromov. On the other hand, we demonstrate a $W^{1,p}$-stability statement ($1\leq p<2$) for the Geroch Conjecture in the class of warped products.

math.DG

Layer-Parallel Inference Reduces Encrypted Nonlinear Depth in Transformers

Fully homomorphic encryption (FHE) enables computation on encrypted data, but practical encrypted Transformer inference is bottlenecked by the sequential composition of many nonlinear blocks. We study whether Structured Newton Layer Parallelism (SNLP) can make this inter-layer composition more FHE-friendly: each Transformer block still requires polynomial approximations for operations such as softmax and RMSNorm, but SNLP reduces the layerwise sequential nonlinear depth from L stages to a small number of solver iterations plus linear structured corrections. Using a simulation framework based on Chebyshev polynomial approximations, we measure error accumulation under sequential versus SNLP inference across 8 models and 4 architecture families. On a 0.5B IDN-trained model, SNLP reduces symbolic bootstraps from 53 to 20 (2.65x) with only +1.2% perplexity degradation, while lowering error amplification (1.36x vs. 1.42x). Across all tested models, SNLP has lower amplification than sequential inference. Ablations show that softmax approximation dominates the error budget and CKKS arithmetic noise is negligible in our setting, suggesting that SNLP is complementary to block-level FHE-friendly operator design rather than a replacement for it.

cs.LG

Connectivity-induced surface-loss penalty in superconducting qubit-coupler lattices

Recent advances in design and fabrication have increased the energy-relaxation times of isolated superconducting transmon qubits to the hundreds-of-microseconds regime, with reported values exceeding 500 $μ$s. However, the same progress has not automatically translated to multiqubit processors, where qubits are embedded in connected qubit-coupler lattices and often exhibit much shorter lifetimes than isolated qubits. To identify possible sources of this discrepancy, here we use finite-element simulation to investigate how surface participation ratios and the resulting surface dielectric loss change when a qubit is embedded in a flip-chip qubit-coupler lattice. Controlled comparisons show that higher connectivity can indeed lead to larger surface loss: in the simulated lattice, connecting a qubit to two and four couplers increases the surface loss by factors of 1.3 and 1.8, respectively. We attribute this change to the combined effects of added edge fields from coupling claws, field redistribution over the larger connected metal network, and hybridization with coupler modes. We further examine how this connectivity-induced surface-loss penalty depends on the geometric design parameters of both the qubit electrodes and the coupling claws, and derive guidelines for designing low-loss multiqubit processors.

quant-ph

New spectral Bishop-Gromov and Bonnet-Myers theorems and applications to isoperimetry

We show a sharp and rigid spectral generalization of the classical Bishop--Gromov volume comparison theorem: if a closed Riemannian manifold $(M,g)$ of dimension $n\geq3$ satisfies $$ λ_1\left(-\frac{n-1}{n-2}Δ+\mathrm{Ric}\right)\geq n-1, $$ then $\operatorname{vol}(M)\leq\operatorname{vol}(\mathbb S^{n})$, and $π_1(M)$ is finite. The constant $\frac{n-1}{n-2}$ cannot be improved, and if $\mathrm{vol}(M)=\mathrm{vol}(\mathbb S^n)$ holds, then $M\cong \mathbb S^{n}$. A sharp generalization of the Bonnet--Myers theorem is also shown under the same spectral condition. The proofs involve the use of a new unequally weighted isoperimetric problem, and unequally warped $μ$-bubbles. As an application, in dimensions $3\leq n\leq 5$, we infer sharp results on the isoperimetric structure at infinity of complete manifolds with nonnegative Ricci curvature and uniformly positive spectral biRicci curvature. Furthermore, the main result of this paper is applied in Mazet's recent solution of the stable Bernstein problem in $\mathbb R^6$.

math.DG

Connected sum of manifolds with spectral Ricci lower bounds

Let $n > 2$, $γ> \frac{n-1}{n-2}$, and $λ\in \mathbb{R}$. We prove that if $M$ and $N$ are two smooth $n$-manifolds that admit a complete Riemannian metric satisfying \[-γΔ+ \mathrm{Ric} > λ,\]then the connected sum $M \# N$ also admits such a metric. The construction geometrically resembles a Gromov-Lawson tunnel; the range $ γ> \frac{n-1}{n-2} $ is sharp for this to hold.

math.DG

HoloTetSphere: Unified TetSphere Mesh Reconstruction for Physical Simulations

Standard pipelines for physics-ready 3D reconstruction rely on a decoupled two-stage paradigm: extracting surface geometry followed by an error-prone tetrahedralization process. While recent Lagrangian methods like TetSphere Splatting attempt to bypass this by directly optimizing volumetric primitives, their homeomorphic constraints prevent topology-adaptive optimization. Consequently, they produce disjoint tetrahedra rather than a single connected mesh, rendering the structures unsuitable for further physical simulations. To address this, we propose a topology-adaptive framework for holistic tetrahedral mesh reconstruction through end-to-end topological and geometric optimization. First, by coupling Gaussian spheres to tetrahedral elements and leveraging edge connections, we estimate a continuous opacity field for differentiable element pruning. Next, jointly minimizing mesh smoothing energy and multi-view Gaussian rendering error drives alternating geometric refinement while preserving topological adaptivity. Consequently, our approach effectively constructs a unified and topologically coherent tetrahedral mesh. Extensive experiments demonstrate that our method outperforms state-of-the-art techniques by achieving superior geometric accuracy and producing coherent, single-connected tetrahedral meshes, thereby effectively bypassing the error-prone conventional tetrahedralization step for reconstructed surface meshes and streamlining downstream physical simulation.

cs.GR

SNLP: Layer-Parallel Inference via Structured Newton Corrections

Autoregressive language models execute Transformer layers sequentially, creating a latency bottleneck that is not removed by conventional tensor or pipeline parallelism. We study whether this layerwise dependency can be relaxed by treating the hidden-state trace across layers as the solution of a nonlinear residual equation and solving it with parallel Newton-style updates. While this view is principled, exact Newton corrections require expensive Jacobian-vector products and naive fixed-point iterations are unstable on trained Transformers. We introduce Structured Newton Layer Parallelism (SNLP), a training and inference framework that replaces exact layer Jacobians with cheap architecture-induced surrogate dynamics. In residual Transformers, this yields Identity Newton (IDN), where the correction reduces to a prefix-sum-like update; in mHC-style architectures, HC Newton (HCN) uses the model's residual mixing matrix. We also study SNLP-aware training, including pretraining regularization and direct SNLP-forward SFT. Experiments on Nanochat-scale Transformers show that SNLP exposes a practical speed-quality frontier: on 0.5B models, it reaches up to 2.58x wall-clock speedup, and a less aggressive configuration reaches 1.40x speedup without increasing PPL. The useful tradeoff comes from the biased finite-iteration computation induced by IDN/HCN rather than exact recovery of the sequential trace. We further show that SNLP-forward SFT can preserve downstream task accuracy, and that SNLP can serve as a drafter for self-speculative decoding while a sequential verifier preserves output correctness.

cs.LG

Infinity-harmonic functions and inverse mean curvature flow clusters

An $\infty$-harmonic function is a viscosity solution of $\nabla^2 u(\nabla u,\nabla u)=0$, or equivalently, an absolute minimizer of $\|\nabla u\|_{L^\infty}$. We prove a variety of new structural and regularity results in two dimensions, including: 1. $\infty$-harmonic functions in domains of $\mathbb{R}^2$ are $C^{1,1/3}$. 2. Critical points are isolated, and at each critical point, the solution has a unique quasiradial blow-up. 3. Entire solutions with polynomial growth have unique quasiradial blow-downs, and are determined by their Fourier modes at infinity. These results are consequences of a new theory relating $\infty$-harmonic functions to inverse mean curvature flow (IMCF) clusters -- which are piecewise weak solutions of IMCF with common obstacle-type boundary conditions on the interfaces (a simple example is an embedded family of cuspidal curves evolving by inverse curvature). This connection arises as the $p\to\infty$ limit of the classical duality between $p$-harmonic and $q$-harmonic functions in $\mathbb{R}^2$, where $\frac1p+\frac1q=1$.

math.AP

Inverse mean curvature flow with outer obstacle

We develop a new boundary condition for the weak inverse mean curvature flow, which gives canonical and non-trivial solutions in bounded domains. Roughly speaking, the boundary of the domain serves as an outer obstacle, and the evolving hypersurfaces are assumed to stick tangentially to the boundary upon contact. In smooth bounded domains, we prove an existence and uniqueness theorem for weak solutions, and establish $C^{1,α}$ regularity of the level sets up to the obstacle. The proof combines various techniques, including elliptic regularization, blow-up analysis, and certain parabolic estimates. As an analytic application, we address the well-posedness problem for the usual weak inverse mean curvature flow, showing that the initial value problem always admits a unique maximal (or innermost) weak solution.

math.DG

Few-Step Diffusion Language Models via Trajectory Self-Distillation

Diffusion large language models (DLLMs) have emerged as powerful generative models with the promise of fast text generation through parallel decoding. However, realizing this potential in practice remains challenging: reducing the number of decoding steps, typically causes a substantial degradation in output quality due to token factorization error. To alleviate this, we propose a self-distillation framework that trains a few-step student to match the generative trajectory of a full-step teacher. We theoretically and empirically show that trajectory-level supervision mitigates this factorization error, thereby enabling effective few-step decoding. We further incorporate Direct Discriminative Optimization (DDO), a reverse-KL objective that encourages mode-seeking toward the teacher's modes, yielding stronger performance on challenging reasoning tasks. Across reasoning and code-generation benchmarks, our method substantially narrows the gap between few-step and full-step decoding. The source code is available at https://github.com/Tyrion58/T3D.

cs.CL

A superconducting qutrit link beyond the qubit limit

Superconducting microwave links have enabled deterministic state transfer and remote entanglement between qubits, but deterministic links have so far operated with an effectively two-dimensional transmitted Hilbert space. Here we demonstrate a superconducting qutrit link between two independently packaged nodes connected by a microwave channel. Each node combines a transmon qutrit, a transmission resonator, and a tunable Purcell-filter interface, allowing the two remote microwave-photon interfaces to be matched in both frequency and bandwidth. We implement two transition-selective photon-mediated operations that transfer the $|e\rangle$ and $|f\rangle$ qutrit components in distinct temporal modes of the same channel. We tomographically characterize arbitrary qutrit-state transfer, obtaining a mean transferred-state fidelity of 83.68% and a qutrit process fidelity of 77.12%, exceeding both the classical qutrit-transfer benchmark and the best possible average fidelity of an effective qubit channel used to transmit an arbitrary qutrit. Using partial-transfer operations, we reconstruct a remote two-qutrit state with negativity 0.730, a tomography-inferred dense-coding capacity of 2.273 bits, and a tomography-inferred Collins-Gisin-Linden-Massar-Popescu (CGLMP) parameter $I_3=2.332$, all beyond the corresponding qubit or local bounds. These results demonstrate a superconducting microwave link that uses the native three-level structure of transmons as a genuine high-dimensional communication resource.

quant-ph