arXiv ScienceSearch

arXiv subjects

Zipeng Wang

Publications and source records attributed to Zipeng Wang.

At least 19 recordsLinked to original sources

Video-HolmesV2: Can MLLMs Reason with Spatio-Temporal Audio-Visual Evidence in Long Videos?

Multimodal Large Language Models have demonstrated impressive video understanding, yet their ability to reason over long-form narratives is often masked by visual-centric evaluations and inefficient context processing. Existing benchmarks over-rely on visual heuristics while marginalizing auditory cues, effectively reducing models to "silent observers" that bypass genuine cross-modal reasoning. Moreover, standard dense sampling creates an evidence-context trade-off: increasing frames to capture evidence inevitably leads to attention distraction and token explosion. To bridge these gaps, we present Video-HolmesV2, a novel benchmark designed for Deep Audio-Visual Coupling. Unlike previous works, it enforces an Evidence-Based Evaluation, requiring models to justify answers with precise spatio-temporal audio-visual evidence, thereby reducing confounding effects of guessing and hallucinated evidence. To support this, we introduce: (1) a Multi-Model Cross-Verification pipeline to ensure task rigor; (2) a Spatio-temporal Evidence-Aware Metric for fine-grained calibration. Furthermore, we propose an Audio-Text Guided Token Compression framework. By fusing task intent with auditory anchors, our method distills high-value reasoning cues to mitigate long-context noise. In our evaluation, even strong proprietary models achieve below 60% accuracy, while our approach outperforms comparable open-source omni-models.

cs.CV

Phase transition in optimal hypercontractivity

We discover an exponent-dependent phase transition phenomenon for optimal hypercontractivity: for every prescribed $q_0>2$, there exists a reversible continuous-time Markov chain on three state space with normalized spectral gap whose $(2,q)$-optimal hypercontractivity time satisfies $$ \text{$t_{\mathrm{opt}}(2,q)=\frac12\log(q-1)$ if and only if $q\ge q_0$},$$ whereas the strict inequality $t_{\mathrm{opt}}(2,q)>\frac12\log(q-1)$ holds for $2<q<q_0$.

math.PR

Cancellation of complex kernels and sharp critical lines and endpoint theory for the Forelli--Rudin operators I: the purely hypersingular case

For $a,b,c\in\mathbb R$, we consider the Forelli--Rudin operators $$ T_{a,b,c}f(z):=(1-|z|^2)^a\int_{\mathbb D}\frac{(1-|w|^2)^b}{(1-z\overline w)^c}f(w)\,dA(w) $$ and their positive counterparts $$ S_{a,b,c}f(z):=(1-|z|^2)^a\int_{\mathbb D}\frac{(1-|w|^2)^b}{|1-z\overline w|^c}f(w)\,dA(w). $$ We obtain a complete and sharp classification of their weak- and restricted weak-type mapping properties in the hypersingular regime $$ Ω_{\mathcal H}:=\{(p,q):1\leq p,q\leq\infty,\ p>q\}, $$ thereby substantially extending the recent work of the first and fourth authors on hypersingular Bergman projections. One of the main discoveries of this work is an intrinsic cancellation phenomenon associated with the complex Forelli--Rudin kernel: at certain critical endpoints, cancellation creates a sharp separation between the two operators, with $T_{a,b,c}$ remaining bounded while its positive counterpart $S_{a,b,c}$ fails to be bounded. Perhaps surprisingly, this cancellation is invisible in the strong $L^p$--$L^q$ theory established by Zhao and Zhou in 2022, where the two operators have the same boundedness range, and emerges only at the weak- and restricted weak-type levels. Allowing the parameters $a,b,c$ to vary, we show that the collection of all weak-type Forelli--Rudin pairs in $Ω_{\mathcal{H}}$ is precisely $$ \mathcal{FR}_w=\{(p,q)\inΩ_{\mathcal H}:1<p\leq2\}, $$ whereas the collection of all restricted weak-type Forelli--Rudin pairs is $$ \mathcal{FR}_{rw}=\{(p,q)\inΩ_{\mathcal H}:p\neq\infty\}. $$ These ranges, together with all corresponding endpoint estimates and failures, are sharp. Our approach combines dyadic decompositions, probabilistic constructions, and weak-type Hardy estimates.

math.FA

Compact Toeplitz operators via the Berezin transform on radial weighted Bergman spaces

Let $ω$ be a radial $\widehat{\mathcal D}$-weight and $u$ be a bounded function on the unit disk $\mathbb D$. We prove that the Toeplitz operator \(T_{ω,u}\) is compact on \(A_ω^2\) if and only if its Berezin transform vanishes at the boundary. Our approach is based on a polynomial frame for $A_ω^2$ and a detailed localization analysis of the resulting infinite matrix representation of $T_{ω,u}$. Even in the unweighted Bergman space \(A^2\), our argument is new and does not rely on the classical translation operators. We further show that this Axler--Zheng compactness characterization does not extend, in general, to products of Toeplitz operators on \(\widehat{\mathcal D}\)-weighted Bergman spaces, and hence to the corresponding Toeplitz algebra generated by bounded symbols. More precisely, we construct a radial log-subharmonic \(\widehat{\mathcal D}\)-weight \(ω\) and bounded symbols \(u,v\) such that the product \(T_{ω,v}T_{ω,u}\) is noncompact, whereas its Berezin transform vanishes at the boundary.

math.CV

GeoBridge: Decoupled Semantic Conditioning for Generative Image Geolocalization

Multimodal large language models (MLLMs) have advanced image geolocalization mainly by improving how they reason about geographic cues. How that reasoning isdecoded into coordinates, however, has lagged behind. Predicting a place name for a geocoding API is discrete and lossy: it ignores image evidence and collapses multi-granular semantics into a coarse lookup. We argue that the bottleneck has shifted from what a model reasons to how that reasoning is represented for a continuous, geometry-aware decoder. We present GeoBridge, a role-decoupled conditioning mechanism that connects a frozen semantic MLLM to a frozen Riemannian flow-matching head that generates coordinates on the sphere. The central obstacle is arole conflict: supervising the condition with discrete semantic labels biases its representation toward class-discriminative geometry, at odds with the smooth manifold the generative head requires. GeoBridge keeps the semantic supervision decoupled from the condition interface: a separate projection forms the continuous condition the frozen head expects, injecting geographic priors without disturbing the spherical decoder. On IM2GPS3K, GeoBridge reaches 38.67/52.89/70.37 at the 25/200/750 km thresholds, improving over a place-name-to-API pipeline and reasoning-augmented direct prediction at these precision-relevant scales. GeoBridge is a decode-side algorithmic contribution, orthogonal and complementary to chain-of-thought reasoning. Code will be made publicly available.

cs.CV

Strong and weak-type estimates for radial weighted Bergman projections

We completely characterize the $L^p$-boundedness and the weak-type (1,1) estimate of radial weighted Bergman projections on the unit disk. Our result, in particular, confirms a conjecture proposed by Peláez and Rättyä in 2021 and thereby settles a longstanding problem in the area that was formally posed by Dostanić in 2004. Consequently, we establish the dichotomy that a radial weighted Bergman projection is bounded either only for $p=2$, or for all $p\in(1,\infty)$.

math.CV

Towards Interactive Global Geolocation Assistant

Global geolocation, which seeks to predict the geographical location of images captured anywhere in the world, is one of the most challenging tasks in the field of computer vision. In this paper, we introduce an innovative interactive global geolocation assistant named GaGA, built upon the flourishing large vision-language models (LVLMs). GaGA uncovers geographical clues within images and combines them with the extensive world knowledge embedded in LVLMs to determine the geolocations while also providing justifications and explanations for the prediction results. We further designed a novel interactive geolocation method that surpasses traditional static inference approaches. It allows users to intervene, correct, or provide clues for the predictions, making the model more flexible and practical. The development of GaGA relies on the newly proposed Multi-modal Global Geolocation (MG-Geo) dataset, a comprehensive collection of 5 million high-quality image-text pairs. GaGA achieves state-of-the-art performance on the GWS15k dataset, improving accuracy by 4.57% at the country level and 2.92% at the city level, setting a new benchmark. These advancements represent a significant leap forward in developing highly accurate, interactive geolocation systems with global applicability.

cs.CV

L^p-boundedness of the Bochner-Riesz operator

In this paper, we give a new approach to the Bochner-Riesz summability. As a result, we show that the Bochner-Riesz operator $\mathbf{S}^δ, 0<\Reδ<{1\over 2}$ is bounded on $\mathbf{L}^p(\mathbb{R}^n)$ for ${n-1\over 2n}\leq {1\over p}\leq{n+1\over 2n}$.

math.CA

GRTresna: An open-source code to solve the initial data constraints in numerical relativity

GRTresna is a multigrid solver designed to solve the constraint equations for the initial data required in numerical relativity simulations. In particular, it is focussed on scenarios with fundamental fields around black holes and inhomogeneous cosmological spacetimes. The code is based on the formalism in Aurrekoetxea, Clough \& Lim arXiv:2207.03125 and can be found at https://github.com/GRTLCollaboration/GRTresna

gr-qc

A new proof of maximal theorem on Heisenberg groups

Given $0\leqα<1$, we define \[\begin{array}{lr} \mathbf{M}_αf(u,v,t) = \sup_{ \mathbf{R} \ni (0,0,0)} {\rm vol} \{\mathbf{R}\}^{α-1} \iiint_\mathbf{R}\left|f [(u,v,t)\odot(ξ,η,τ)^{-1}]\right|dξdηdτ\end{array}\] where $\mathbf{R}\subset\mathbb{R}^{2n+1}$ is a rectangle parallel to the coordinates. Moreover, $\odot$ denotes the multiplication law on a real Heisenberg group. The $\mathbf{L}^p$-boundedness of $\mathbf{M}_0$ has been previously proved by M. Christ. We show $\mathbf{M}_α\colon\mathbf{L}^p(\mathbb{R}^{2n+1}) \to \mathbf{L}^q(\mathbb{R}^{2n+1})$ for $α={1\over p}-{1\over q},~ 1<p\leq q<\infty$ by applying a geometric covering lemma due to Córdoba and Fefferman.

math.CA

FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor Attention

3D reconstruction from multi-view images is a core challenge in computer vision. Recently, feed-forward methods have emerged as efficient and robust alternatives to traditional per-scene optimization techniques. Among them, state-of-the-art models like the Visual Geometry Grounding Transformer (VGGT) leverage full self-attention over all image tokens to capture global relationships. However, this approach suffers from poor scalability due to the quadratic complexity of self-attention and the large number of tokens generated in long image sequences. In this work, we introduce FlashVGGT, an efficient alternative that addresses this bottleneck through a descriptor-based attention mechanism. Instead of applying dense global attention across all tokens, FlashVGGT compresses spatial information from each frame into a compact set of descriptor tokens. Global attention is then computed as cross-attention between the full set of image tokens and this smaller descriptor set, significantly reducing computational overhead. Moreover, the compactness of the descriptors enables online inference over long sequences via a chunk-recursive mechanism that reuses cached descriptors from previous chunks. Experimental results show that FlashVGGT achieves reconstruction accuracy competitive with VGGT while reducing inference time to just 9.3% of VGGT for 1,000 images, and scaling efficiently to sequences exceeding 3,000 images. Our project page is available at https://wzpscott.github.io/flashvggt_page/.

cs.CV

Beyond Geometry: Artistic Disparity Synthesis for Immersive 2D-to-3D

Current 2D-to-3D conversion methods achieve geometric accuracy but are artistically deficient, failing to replicate the immersive and emotionally resonant experience of professional 3D cinema. This is because geometric reconstruction paradigms mistake deliberate artistic intent, such as strategic zero-plane shifts for pop-out effects and local depth sculpting, for data noise or ambiguity. This paper argues for a new paradigm: Artistic Disparity Synthesis, shifting the goal from physically accurate disparity estimation to artistically coherent disparity synthesis. We propose Art3D, a preliminary framework exploring this paradigm. Art3D uses a dual-path architecture to decouple global depth parameters (macro-intent) from local artistic effects (visual brushstrokes) and learns from professional 3D film data via indirect supervision. We also introduce a preliminary evaluation method to quantify cinematic alignment. Experiments show our approach demonstrates potential in replicating key local out-of-screen effects and aligning with the global depth styles of cinematic 3D content, laying the groundwork for a new class of artistically-driven conversion tools.

cs.CV

HeroGS: Hierarchical Guidance for Robust 3D Gaussian Splatting under Sparse Views

3D Gaussian Splatting (3DGS) has recently emerged as a promising approach in novel view synthesis, combining photorealistic rendering with real-time efficiency. However, its success heavily relies on dense camera coverage; under sparse-view conditions, insufficient supervision leads to irregular Gaussian distributions, characterized by globally sparse coverage, blurred background, and distorted high-frequency areas. To address this, we propose HeroGS, Hierarchical Guidance for Robust 3D Gaussian Splatting, a unified framework that establishes hierarchical guidance across the image, feature, and parameter levels. At the image level, sparse supervision is converted into pseudo-dense guidance, globally regularizing the Gaussian distributions and forming a consistent foundation for subsequent optimization. Building upon this, Feature-Adaptive Densification and Pruning (FADP) at the feature level leverages low-level features to refine high-frequency details and adaptively densifies Gaussians in background regions. The optimized distributions then support Co-Pruned Geometry Consistency (CPG) at parameter level, which guides geometric consistency through parameter freezing and co-pruning, effectively removing inconsistent splats. The hierarchical guidance strategy effectively constrains and optimizes the overall Gaussian distributions, thereby enhancing both structural fidelity and rendering quality. Extensive experiments demonstrate that HeroGS achieves high-fidelity reconstructions and consistently surpasses state-of-the-art baselines under sparse-view conditions.

cs.CV

The optimal hypercontractive constants for $\mathbb{Z}_3$ and biased Bernoulli random variables

We resolve a folklore problem of determining the optimal hypercontractive constants $r_{p,q}(\mathbb{Z}_3)$ for the cyclic group $\mathbb{Z}_3$ for all $1 < p < q < \infty$. More precisely, we have \[ r_{p,q}(\mathbb{Z}_3) = \frac{(1 + 2x)(1 - y)}{(1 + 2y)(1 - x)}, \] where $(x,y)$ is the unique solution in the open unit square $(0,1)\times (0,1)$ to the system of equations \begin{align*} \left\{ \begin{aligned} &\frac{1}{1+2x}\Big(\frac{1+2x^p}{3}\Big)^{\frac{1}{p}}=\frac{1}{1+2y}\Big(\frac{1+2y^q}{3}\Big)^{\frac{1}{q}},\\ &\frac{(1-x)(1-x^{p-1})}{1+2x^p}=\frac{(1-y)(1-y^{q-1})}{1+2y^q}. \end{aligned} \right. \end{align*} Consequently, for rational $p, q\in \mathbb{Q}$, the constants $r_{p,q}(\mathbb{Z}_3)$ are algebraic numbers which generally admit no radical expressions, since their often rather complicated minimal polynomials may have non-solvable Galois groups. Our formalism relies on a key observation: the existence of nontrivial critical extremizers. This approach can also be adapted to resolve a long-standing open problem -- determining all optimal $(p,q)$-hypercontractive constants for biased Bernoulli random variables, which are closely related to noise operators. Several noteworthy phenomena emerge from numerical simulations: the monotonicity of the hypercontractive constants in the parameters, and the appearance of intriguing limit shapes. These phenomena merit further investigation.

math.FA

$L^p$--$L^q$ estimates for Shimorin-type integral operators

Let $ν$ be a positive measure on $[0,1]$. A Shimorin-type operator $T_ν$ is an integral operator on the unit disk given by \[ T_νf(z) = \int_{\mathbb{D}} \frac{1}{1 - z\overlineλ} \left( \int_0^1 \frac{dν(r)}{1 - r z \overlineλ} \right) f(λ) \, dA(λ), \] which originates from Shimorin's work on Bergman-type kernel representations for logarithmically subharmonic weighted Bergman spaces. In this paper, we study $L^p$--$L^q$ estimates for $T_ν$. Unlike classical Bergman-type operators, the critical line on the $(1/p,1/q)$-plane that separates the boundedness and unboundedness regions of $T_ν$ is not immediately evident. Moreover, even along this line, new phenomena arise. In the present work, by introducing a quantity $c_ν$, \begin{itemize} \item we first determine the critical boundary in the $(1/p,1/q)$-plane for bounded $T_ν$; \item furthermore, on this critical line, we establish necessary and sufficient conditions for $T_ν$ which have standard Bergman-type $L^p$--$L^q$ estimates, meaning that it is bounded in the interior of the region and admits weak-type and BMO-type estimates at endpoints. \end{itemize}

math.CV