arXiv ScienceSearch

arXiv subjects

Xin Wan

Publications and source records attributed to Xin Wan.

At least 19 recordsLinked to original sources

ProactiveBench: Can Streaming Video Models Really Interact Like Humans?

Streaming video understanding requires models to process continuous multimodal input while maintaining temporal context. Existing evaluations are predominantly reactive: they query a model at a selected timestamp and therefore do not assess when it should respond. Proactive interaction instead requires monitoring a standing request, responding within an appropriate interval after the target event, and otherwise remaining silent. We introduce ProactiveBench, which evaluates models at one-second stream intervals without an explicit response cue. Its six subtasks vary trigger ambiguity and timing tolerance. Event Sensitivity geometrically combines response and silence rates on the same recording; four window-based subtasks distinguish early, in-window, and missed responses; and Duplicate Counting penalizes omissions and repetitions. Premature responses outnumber missed responses for four of the six evaluated systems, revealing a substantial gap in the temporal decision-making required for human-like interaction.

cs.LG

Kato's main conjecture for nonordinary modular forms

We prove Kato's main conjecture for modular forms at nonordinary primes for weights in the Fontaine-Laffaille range. The key ingredients in the proof are a reformulation of the conjecture in terms of signed Selmer groups due to Lei-Loeffler-Zerbes, certain $p$-adic families of Rankin-Eisenstein classes arising from the work of Lei-Loeffler-Zerbes and Kings-Loeffler-Zerbes, and the lower bound divisibility in an Iwasawa-Greenberg main conjecture for Rankin-Selberg $p$-adic $L$-functions obtained in our earlier work \cite{CLW}.

math.NT

ZoomV: Temporal Zoom-in for Efficient Long Video Understanding

Long video understanding poses a fundamental challenge for large video-language models (LVLMs) due to the overwhelming number of frames and the risk of losing essential context through naive downsampling. Inspired by the way humans watch videos on mobile phones, constantly zooming in on frames of interest, we propose ZoomV, a query-aware temporal zoom-in framework designed for efficient and accurate long video understanding. Specifically, ZoomV operates in three stages: (1) Temporal interests grounding: guided by the query, ZoomV retrieves relevant events and their associated temporal windows as candidates. (2) Event interests spotlighting: within pools of candidate windows, each window is scored through the model itself reflection and filtered accordingly, where higher-confidence windows are more representative. (3) Compact representation: the selected events are encoded and temporally downsampled to preserve critical semantics while significantly reducing redundancy. Extensive experiments demonstrate that ZoomV substantially outperforms prior video agent approaches. On temporal grounding, ZoomV unlocks the latent capability of LVLMs, achieving an 11.8% mIoU gain on Charades-STA. Remarkably, ZoomV further boosts accuracy on LVBench by 9.7%, underscoring its effectiveness on long-video benchmarks.

cs.CV

Zero-Clustering Geometry in Realistic Fractional Quantum Hall Wave Functions

The clustering pattern of zeros in the ground state of a fractional quantum Hall system is a defining feature of its topological properties. We analyze the geometrical fluctuations of the zeros around individual electrons and propose to use the displacement ratio of the zeros to visualize and measure the distance of a realistic state to a model wave function. The distribution of the zero displacement ratio behaves like an order parameter in the transition from a Laughlin phase to a topologically trivial one. The statistical comparison between quantum Hall states belonging to different Jain sequences leads to a composite fermion fluid description of the $ν= 1/5$ ground state with long-range Coulomb interaction that agrees almost perfectly for as few as $3$-$5$ electrons, overcoming the long-standing difficulties of accommodating the competing liquid and crystal orders at short distances.

cond-mat.str-el

Engineering two-body interaction for the Moore-Read State

Engineering interactions that stabilize non-Abelian fractional quantum Hall phases is a central challenge in strongly correlated topological matter and quantum simulation. We introduce a differentiable framework for inverse Hamiltonian design, in which Haldane pseudopotentials are optimized by gradient-based exact diagonalization to stabilize target fractional quantum Hall phases. In spherical geometry, the Haldane pseudopotentials are treated as variational parameters and optimized in a JAX-based exact-diagonalization framework. By directly maximizing the overlap between the many-body ground state and the Moore-Read state, we obtain a robust pseudopotential profile that has Pfaffian overlaps exceeding $99\%$ for systems up to $N_e=12$, substantially improving over conventional Coulomb interactions. Analyses of the neutral excitation spectrum and orbital entanglement spectrum further confirm that the optimized interaction stabilizes the Pfaffian topological phase. Our results demonstrate that essential features of the three-body Pfaffian parent Hamiltonian can be effectively encoded in a suitably designed two-body interaction. Furthermore, they identify a nearly universal exponentially decaying pseudopotential profile that stabilizes the Pfaffian phase and establishes a general framework toward engineering non-Abelian topological order in quantum simulation.

cond-mat.str-el

Energy Gap in Weakly Disordered Fractional Quantum Hall Liquids: Quantitative Comparison to GaAs Quantum Well Experiments at $ν= 1/3$

Based on a recent experiment in high-quality GaAs quantum wells [Phys. Rev. Lett. 127, 056801 (2021)], we present a microscopic study of the energy gap in two-dimensional electron gases at filling factor $ν=1/3$, explicitly incorporating both finite layer thickness and disorder effects. The finite layer thickness is modeled by solving the Poisson-Schrödinger equations for the experimental devices, yielding the electron wave functions in the perpendicular direction. Using these and the disorder energy extracted from the experiment, we estimate the charge gap and the mobility gap at $ν=1/3$ in the weakly disordered lowest Landau level. Remarkably, both gaps show good quantitative agreement with the activation gap measured from the experiment in narrow quantum wells. Our results also indicate the potential need of incorporating higher subbands to make accurate theoretical predictions of the energy gap in wide quantum wells.

cond-mat.str-el

SesQ: A Surface Electrostatic Simulator for Precise Energy Participation Ratio Simulation in Superconducting Qubits

An accurate and efficient numerical electromagnetic model for superconducting qubits is essential for characterizing and minimizing design-dependent dielectric losses. The energy participation ratio (EPR) is the commonly adopted metric used to evaluate these losses, but its calculation presents a severe multiscale computational challenge. Conventional finite element method (FEM) requires 3D volumetric meshing, leading to prohibitive computational costs and memory requirements when attempting to capture singular electric fields at nanometer-thin material interfaces. To address this bottleneck, we propose SesQ, a surface integral equation simulator tailored for the precise simulation of the EPR. By applying discretization on 2D surfaces, deriving a semi-analytical multilayer Green's function, and employing a dedicated non-conformal boundary mesh refinement scheme, SesQ accurately resolves singular edge fields without an explosive growth in the number of unknowns. Validations with analytically solvable models demonstrate that SesQ accelerates capacitance extraction by roughly two orders of magnitude compared to commercial FEM tools. While achieving comparable accuracy for capacitance extraction, SesQ delivers superior precision for EPR calculation. Simulations of practical transmon qubits further reveal that FEM approaches tend to significantly underestimate the EPR. Finally, the high efficiency of SesQ enables rapid iteration in the layout optimization, as demonstrated by minimizing the EPR of the qubit pattern, establishing the simulator as a powerful tool for the automated design of low-loss superconducting quantum circuits.

quant-ph

A refined non-vanishing of the $p$-adic logarithm of a rational point on an abelian variety

Inspired by a beautiful formula of Bertolini, Darmon, and Prasanna -- the oft-termed BDP formula -- we address questions about the non-vanishing of non-torsion points under $p$-adic logarithms of abelian varieties. We largely consider situations most applicable to ${\mathrm GL}_2$-type abelian varieties associated with Hilbert modular newforms and Heegner points. Not surprisingly, the main tool employed is the $p$-adic analytic subgroup theorem.

math.NT

$p$-adic Waldspurger Formula for Non-split Primes and Converse of Gross--Zagier and Kolyvagin Theorem

Let $p$ be a prime and $\mathcal{K}$ be an imaginary quadratic field. In this paper we generalize a recent construction of a new type of $p$-adic $L$-function and $p$-adic Waldspurger formula by Andreatta-Iovita for $p$ non-split in $\mathcal{K}$, as long as the $\mathrm{GL}_2$ automorphic representation is principal series at $p$. Then we develop a new kind of anticyclotomic local $\pm$-Iwasawa theory at $p$ for self-dual Hecke characters over imaginary quadratic fields (including elliptic curves $E/\mathbb{Q}$ with complex multiplication) which is valid for all ramification types of $p$ (split, inert and ramified, and allowing $p=2$). As the main consequence, we prove the converse of the Gross--Zagier--Kolyvagin theorem for self-dual CM characters: if the Selmer rank of $ψ$ is 1, then the analytic rank of $L(ψ,s)$ at $s=1$ is also 1. As corollaries, we prove Sylvester's conjecture (1879) on sums of two rational cubes, and Goldfeld's conjecture for CM elliptic curves over $\mathbb{Q}$ (conditional on work of A. Smith).

math.NT

MammothModa2: A Unified AR-Diffusion Framework for Multimodal Understanding and Generation

Unified multimodal models aim to integrate understanding and generation within a single framework, yet bridging the gap between discrete semantic reasoning and high-fidelity visual synthesis remains challenging. We present MammothModa2 (Mammoth2), a unified autoregressive-diffusion (AR-Diffusion) framework designed to effectively couple autoregressive semantic planning with diffusion-based generation. Mammoth2 adopts a serial design: an AR path equipped with generation experts performs global semantic modeling over discrete tokens, while a single-stream Diffusion Transformer (DiT) decoder handles high-fidelity image synthesis. A carefully designed AR-Diffusion feature alignment module combines multi-layer feature aggregation, unified condition encoding, and in-context conditioning to stably align AR's representations with the diffusion decoder's continuous latents. Mammoth2 is trained end-to-end with joint Next-Token Prediction and Flow Matching objectives, followed by supervised fine-tuning and reinforcement learning over both generation and editing. With roughly 60M supervised generation samples and no reliance on pre-trained generators, Mammoth2 delivers strong text-to-image and instruction-based editing performance on public benchmarks, achieving 0.87 on GenEval, 87.2 on DPGBench, and 4.06 on ImgEdit, while remaining competitive with understanding-only backbones (e.g., Qwen3-VL-8B) on multimodal understanding tasks. These results suggest that a carefully coupled AR-Diffusion architecture can provide high-fidelity generation and editing while maintaining strong multimodal comprehension within a single, parameter- and data-efficient model.

cs.CV

TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinforcement Learning

Temporal search aims to identify a minimal set of relevant frames from tens of thousands based on a given query, serving as a foundation for accurate long-form video understanding. Existing works attempt to progressively narrow the search space. However, these approaches typically rely on a hand-crafted search process, lacking end-to-end optimization for learning optimal search strategies. In this paper, we propose TimeSearch-R, which reformulates temporal search as interleaved text-video thinking, seamlessly integrating searching video clips into the reasoning process through reinforcement learning (RL). However, applying RL training methods, such as Group Relative Policy Optimization (GRPO), to video reasoning can result in unsupervised intermediate search decisions. This leads to insufficient exploration of the video content and inconsistent logical reasoning. To address these issues, we introduce GRPO with Completeness Self-Verification (GRPO-CSV), which gathers searched video frames from the interleaved reasoning process and utilizes the same policy model to verify the adequacy of searched frames, thereby improving the completeness of video reasoning. Additionally, we construct datasets specifically designed for the SFT cold-start and RL training of GRPO-CSV, filtering out samples with weak temporal dependencies to enhance task difficulty and improve temporal search capabilities. Extensive experiments demonstrate that TimeSearch-R achieves significant improvements on temporal search benchmarks such as Haystack-LVBench and Haystack-Ego4D, as well as long-form video understanding benchmarks like VideoMME and MLVU. Notably, TimeSearch-R establishes a new state-of-the-art on LongVideoBench with 4.1% improvement over the base model Qwen2.5-VL and 2.0% over the advanced video reasoning model Video-R1. Our code is available at https://github.com/Time-Search/TimeSearch-R.

cs.CV

SuperGrad: a differentiable simulator for superconducting processors

One significant advantage of superconducting processors is their extensive design flexibility, which encompasses various types of qubits and interactions. Given the large number of tunable parameters of a processor, the ability to perform gradient optimization would be highly beneficial. Efficient backpropagation for gradient computation requires a tightly integrated software library, for which no open-source implementation is currently available. In this work, we introduce SuperGrad, a simulator that accelerates the design of superconducting quantum processors by incorporating gradient computation capabilities. SuperGrad offers a user-friendly interface for constructing Hamiltonians and computing both static and dynamic properties of composite systems. This differentiable simulation is valuable for a range of applications, including optimal control, design optimization, and experimental data fitting. In this paper, we demonstrate these applications through examples and code snippets.

quant-ph

The effects of disorder in superconducting materials on qubit coherence

Introducing disorderness in the superconducting materials has been considered promising to enhance the electromagnetic impedance and realize noise-resilient superconducting qubits. Despite a number of pioneering implementations, the understanding of the correlation between the material disorderness and the qubit coherence is still developing. Here, we demonstrate a systematic characterization of fluxonium qubits with the superinductors made from titanium-aluminum-nitride with varied disorderness. From qubit noise spectroscopy, the flux noise and the dielectric loss are extracted as a measure of the coherence properties. Our results reveal that the $1/f$ flux noise dominates the qubit decoherence around the flux-frustration point, strongly correlated with the material disorderness; while the dielectric loss remains low under a wide range of material properties. From the flux-noise amplitudes, the areal density ($σ$) of the phenomenological spin defects and material disorderness are found to be approximately correlated by $σ\propto ρ_{xx}^3$, or effectively $(k_F l)^{-3}$. This work has provided new insights on the origin of decoherence channels within superconductors, and could serve as a useful guideline for material design and optimization.

quant-ph

Quantifying the Value of Revert Protection

Revert protection is a feature provided by some blockchain platforms that prevents users from incurring fees for failed transactions. We study the economic implications and benefits of revert protection in the context of priority gas auctions and maximal extractable value. We develop a model in which searchers bid for a top-of-block arbitrage opportunity under varying degrees of revert protection. This model applies to a broad range of settings, including bundle auctions on L1s and priority ordering sequencing rules on L2s. We quantify, in closed form, how revert protection improves equilibrium auction revenue, market efficiency, and blockspace efficiency.

cs.GT

What Drives Liquidity on Decentralized Exchanges? Evidence from the Uniswap Protocol

We study liquidity on decentralized exchanges (DEXs), identifying factors at the platform, blockchain, token pair, and liquidity pool levels with predictive power for market depth metrics. We introduce the v2 counterfactual spread metric, a novel criterion which assesses the degree of liquidity concentration in pools using the ``concentrated liquidity'' mechanism, allowing us to decompose the effect of a factor on market depth into two channels: total value locked (TVL) and concentration. We further explore how external liquidity from competing DEXs and private inventory on DEX aggregators influence market depth. We find that (i) gas prices, returns, and a DEX's share of trading volume affect liquidity through concentration, (ii) internalization of order flow by private market makers affects TVL but not the overall market depth, and (iii) volatility, fee revenue, and markout affect liquidity through both channels.

q-fin.TR

Dynamical Geometry of the Haldane Model under a Quantum Quench

We explore the time evolution of a topological system when the system undergoes a sudden quantum quench within the same nontrivial phase. Using Haldane's honeycomb model as an example, we show that equilibrium states in a topological phase can be distinguished by geometrical features, such as the characteristic momentum at which the half-occupied edge modes cross, the associated edge-mode velocity, and the winding vector about which the normalized pseudospin magnetic field winds along a great circle on the Bloch sphere. We generalize these geometrical quantities for non-equilibrium states and use them to visualize the quench dynamics of the topological system. In general, we find the pre-quench equilibrium state relaxes to the post-quench equilibrium state in an oscillatory fashion, whose amplitude decay as $t^{1/2}$. In the process, however, the characteristic winding vector of the non-equilibrium system can evolve to regimes that are not reachable with equilibrium states.

cond-mat.mes-hall

Zeta elements for elliptic curves and applications

Let $E$ be an elliptic curve defined over $\mathbb{Q}$ with conductor $N$ and $p\nmid 2N$ a prime. Let $L$ be an imaginary quadratic field with $p$ split. We prove the existence of $p$-adic zeta element for $E$ over $L$, encoding two different $p$-adic $L$-functions associated to $E$ over $L$ via explicit reciprocity laws at the primes above $p$. We formulate a main conjecture for $E$ over $L$ in terms of the zeta element, mediating different main conjectures in which the $p$-adic $L$-functions appear, and prove some results toward them. The zeta element has various applications to the arithmetic of elliptic curves. This includes a proof of main conjecture for semistable elliptic curves $E$ over $\mathbb{Q}$ at supersingular primes $p$, as conjectured by Kobayashi in 2002. It leads to the $p$-part of the conjectural Birch and Swinnerton-Dyer (BSD) formula for such curves of analytic rank zero or one, and enables us to present the first infinite families of non-CM elliptic curves for which the BSD conjecture is true. We provide further evidence towards the BSD conjecture: new cases of $p$-converse to the Gross--Zagier and Kolyvagin theorem, and $p$-part of the BSD formula for ordinary primes $p$. Along the way, we give a proof of a conjecture of Perrin-Riou connecting Beilinson--Kato elements with rational points.

math.NT

Iwasawa Main Conjecture for Supersingular Elliptic Curves and BSD conjecture

In this paper we prove the $\pm$-main conjecture of Iwasawa theory formulated by Kobayashi for elliptic curves with supersingular reduction at an odd prime $p$ such that $a_p=0$, using a key new observation that it can be reduced to another Iwasawa-Greenberg main conjecture, which is more accessible and proved here as a first step. Then we develop some generalized $\pm$ local theory and deduce the main conjecture. The argument uses in an essential way the recent study on explicit reciprocity law for Beilinson-Flach elements by Kings-Loeffler-Zerbes. We also prove as corollaries the $p$-part of the BSD formula at supersingular primes when the analytic rank is $0$ or $1$. The main result enables us to present in the Appendix a number of explicit infinite families of elliptic curves without complex multiplications for which we can now prove the full Birch-Swinnerton-Dyer conjecture. No such infinite families of curves without complex multiplication were known previously.

math.NT