arXiv ScienceSearch

arXiv subjects

Haoming Wang

Publications and source records attributed to Haoming Wang.

At least 19 recordsLinked to original sources

Singular value decomposition of unbounded operators

The singular value decomposition has been established for matrices, Hilbert--Schmidt operators, trace-class operator, compact operators, and bounded operators, but surprisingly not for unbounded operators. Unfortunately, most interesting operators in applied math are unbounded, as any operators involving some form of derivatives --- gradient, exterior derivatives, Laplacians, Fourier and other transforms of derivatives, Hamiltonians, etc. --- are likely unbounded. In this article, we fill in this last missing piece by establishing the existence of singular value decompositions for unbounded operators in three natural forms: multiplication-operator, direct-integral, and operator-valued-measure. We show it inherits classical properties of finite-dimensional singular value decomposition including approximation results, relationships with fundamental subspaces, and the Moore--Penrose inverse. This discovery opens the door to the singular value decompositions of a myriad of well-known unbounded operators in mathematics, physics, statistics, and finnance --- gradients on Euclidean spaces and manifolds, Petrov--Galerkin method, finite-difference operators, Hilbert--Schmidt operators, Hilbert complexes, supersymmetric quantum mechanics, Sturm--Liouville theory, nonparametric density estimation, and the Black--Scholes equation. The resulting decompositions reveal a number of novel insights, among many others: bosonic and fermionic states in supersymmetric quantum mechanics arise as left and right singular vectors of generalized ladder operators; the Riesz transform appears as the left singular operator of the Euclidean gradient; and the Hodge decomposition follows directly from the singular value decompositions of the exterior derivatives.

math.SP

Identifiability of Nonnegative Tensor Decompositions via Positive Scattering

Identifiability of tensor decompositions is often established through linear-algebraic conditions on the factor families. For nonnegative decompositions, however, positivity provides additional information that is not captured by dimension and independence alone: nonnegative terms cannot cancel, and their supports constrain competing decompositions. We introduce a positive scattering term that quantifies this additional source of identifiability and combine it with the dimension budget underlying the Lovitz--Petrov generalization of Kruskal's theorem. For every subset of components, we obtain two sufficient conditions: a threshold of $2|S|-2$ guarantees minimality and nonnegative rank, while the stronger threshold $2|S|-1$ guarantees uniqueness among nonnegative decompositions of the same length. The key result is a positive splitting inequality for irreducible exchanges of nonnegative rank-one tensors, which combines the dimension constraint with support-induced geometric rigidity. Although the scattering term is defined through an optimization over intermediate factor spaces, we show that its mode costs are exactly $0$, $1$, or $+\infty$, yielding an exact activation characterization in terms of graph connectivity. The resulting criterion can strictly certify sparse nonnegative tensor decompositions beyond the reach of Kruskal and Lovitz--Petrov conditions, including examples for which those conditions fail even after reshaping. In the matrix case, the two criteria reduce respectively to full-rank factorization and two-sided separability.

stat.ML

The joint density of top k sample eigenvectors and principal subspace inference

In this paper, the exact joint density of the top $k$ sample eigenvectors is derived for any $p\times p$ population covariance matrix $\varSigma$, sample size $n > p -1$, and $1 \le k \le p$. Using symmetric functions such as zonal polynomials and a dual summation identity by I. G. Macdonald, the result is expressed as a series in terms of determinants of differential operators. A matrix Kummer transformation converts the resulting local alternating expansion into a globally absolutely convergent series over the entire positive definite matrix space. For $k=2$, the ordered eigenvalue integrals reduce to the Gauss hypergeometric function ${}_2F_1$ with a simple closed form when $n=p+1$. Previously, explicit formulas were only available for a scalar matrix $\varSigma = σ^2 I_p$ or a general matrix $\varSigma$ with $p=2$ or $k=1$, obtained by T. W. Anderson and T. Sugiyama, respectively. The frame law also induces an exact Grassmann density that provides a finite-sample benchmark for principal subspace inference.

math.ST

From Association to Causation: Improving Retrieval Precision of Retrieval-Augmented Generation via Causal Relations and an Attention Mechanism

Retrieval-Augmented Generation (RAG) grounds LLM generation on retrieved documents, but the standard terminal retrieval stage--dense-vector similarity, optionally followed by reranking--often returns documents that share keywords with the query without containing the needed information, a failure mode that grows with the knowledge base. We trace it to a conceptual gap: similarity captures only associational relations, whereas the documents that matter are linked to the query causally. We model the terminal retrieval stage with a causal graph grounded in Reichenbach's common cause principle: the keywords shared by the query and a retrieved document form a latent common cause A, and the document's residual keywords form a latent set B linking the document to the ideal output. Since a retrieved document is a collider (A -> d <- B), retrieval itself opens an associational path between the query and B, which licenses a training-free, attention-style re-scoring rule: the cosine similarity between the query embedding and the weighted centroid embedding of B. Unlike causality-enhanced RAG variants that model causal relations inside the knowledge content, our graph models the causal structure of the retrieval process itself. On a real 471-document enterprise knowledge base, the method promotes a relevant guideline from rank 6 to the top 3; on a controlled diagnostic corpus reproducing the keyword-stuffing regime, it improves the mean target rank from 2.88 to 1.25, while a trained cross-encoder reranker barely helps (2.63). Conversely, on three BEIR benchmarks the score underperforms the similarity baseline, delineating the applicability boundary: the method guards the keyword-stuffing regime of growing proprietary knowledge bases and complements neural rerankers; a corpus-level calibration gate selects the correct regime with >= 95% reliability. A fully local testbed demonstrates deployability.

cs.AI

Why does AI unlock new possibilities in STEM education? A Bibliometric Analysis of Trends and Future Agenda

STEM education faces challenges in personalization and interdisciplinary integration. AI technology has brought new possibilities, but the mechanisms by which AI reshapes the STEM education ecosystem require systematic investigation. This study employs bibliometric methods to analyze 242 publications from 2015-2025, constructing knowledge maps to reveal the evolutionary trajectory. The findings show that the field has transformed from intelligent tutoring systems to inquiry-based learning and computational thinking cultivation driven by LLMs. AI's key contribution lies in providing intelligent scaffolding that lowers the threshold for understanding knowledge. In this sense, AI is a core driving force promoting its shift from knowledge transmission to capability development.

cs.CY

BlockPython: A Process-Aware Agent-Supported Platform for the Transition from Block-Based to Python Programming

The transition from block-based to text-based programming requires learners to convert visible program structures into abstract textual expressions, which may create a cognitive gap between understanding computational concepts and expressing them in Python syntax. To support this transition, we designed and implemented BlockPython. The platform centers on bidirectional translation between blocks and Python and guides learners through four stages: Task Decomposition, Block-Based Practice, Code Challenge, and Extended Interaction. Across these stages, learners progressively establish connections among program structure, runtime behavior, and textual code. During learning, the platform continuously collects process evidence, including block artifacts, code versions, run outcomes, use of support, and dialogue. Deterministic diagnosis, program visualization, and the learning assistant use this evidence to identify different difficulties in computational understanding and Python expression. The rule-based system is responsible for program execution, objective evaluation, and stage control, while the learning assistant uses verified evidence to provide explanations, prompts, and guiding questions. This report describes the design rationale, learning workflow, and process-aware support mechanisms of BlockPython and provides a system-design reference for supporting the transition from block-based to text-based programming and for analyzing learning processes.

cs.AI

Continuous first-order logic of abstract harmonic spaces

This paper studies Brelot harmonic spaces within an unbounded many-sorted continuous first-order language and its expansions. 1. In the finite evaluation reduct of the harmonic language, the Brelot theory together with the finite interpolation scheme eliminates bounded harmonic quantifiers while leaving point quantifiers untouched. 2. The Martin compactification is homeomorphic to a compact subset of the local type space, with the Martin boundary corresponding to non-principal types. 3. If the compact convex space of normalized positive harmonic functions is a Choquet simplex, then each Martin representing measure on the minimal boundary corresponds to a unique Keisler measure supported on minimal types and admits a complete lift. 4. Harnack compactness also implies stability of normalized evaluation formulas, yielding canonical bases and a harmonic barycenter factoring through Keisler measures. These constructions are further applied to fine potential theory, where boundary poles of potentials are Keisler null. First-jet expansions encode the Radó--Kneser--Choquet and Lewy planar rigidity phenomena, while higher-dimensional failures are analyzed through thorn-forking chains in o-minimal expansions.

math.LO

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models

Decades of cognitive science establish that humans navigate environments by forming cognitive maps, defined as allocentric and topology-preserving representations of 3D space. While modern Vision-Language Models (VLMs) demonstrate emergent spatial reasoning from 2D egocentric inputs, it remains unclear whether they construct an analogous 3D internal representation. In this paper, we demonstrate that current VLMs do possess a latent topological map of 3D scenes, but it is heavily overshadowed by non-geometric visual semantics, such as color and shape. By isolating this spatial subspace through cross-scene linear feature extraction, we extract a clean spatial subspace that causally controls the model's spatial outputs. We mathematically shape this latent representation and prove its correspondence to the Laplacian eigenmaps of the scene's 3D Gaussian-kernel graph, converging to the physical 3D space in the continuous limit. Motivated by this geometric identification, we further introduce a mathematically principled latent regularization method for VLMs, based on Dirichlet energy. Applying this single-term regularizer to a minimal 500-step supervised VLM fine-tuning (SFT) on simple synthetic data yields significant improvements on real-world spatial benchmarks, outperforming standard SFT and competitive baselines by up to 12.1\% in spatial tasks involving scene topology understanding. Source code is available at https://github.com/pittisl/vlm-latent-shaping

cs.CV

On non-central distribution of the matrix ratio

We derive the distribution of the ratio of a non-central mean matrix and a sample covariance matrix. This aligns with the confluent term ${}_1F_1$ in the non-central uni-variate Student's $t$. Some extensions of matrix-variate distributions are considered.

math.ST

On incomplete Gamma and Beta integrals

This paper discusses the incomplete Gamma and Beta integrals involving the generalised hypergeometric function. The distribution of the largest and the smallest roots of a ratio arising in comparing the mean differences among groups is obtained as an application.

math.ST

Real gamma distribution on analytic bundles of flag varieties

This paper introduces four matrix normal distributions on analytic bundles of flag varieties, extending the separable covariance $\varPhi \otimes \varPsi$ with potentially variable-level ($\varPsi$) and/or sample-level ($\varPhi$) correlations. The joint distribution of sample variances and covariances, leading to the product moment distribution, is considered when precision matrices admit a specific form. Several well-known consequences, including the non-central Wishart distribution and normal quadratic forms, now appear as corollaries.

math.AG

MosaicThinker: On-Device Visual Spatial Reasoning for Embodied AI via Iterative Construction of Space Representation

When embodied AI is expanding from traditional object detection and recognition to more advanced tasks of robot manipulation and actuation planning, visual spatial reasoning from the video inputs is necessary to perceive the spatial relationships of objects and guide device actions. However, existing visual language models (VLMs) have very weak capabilities in spatial reasoning due to the lack of knowledge about 3D spatial information, especially when the reasoning task involve complex spatial relations across multiple video frames. In this paper, we present a new inference-time computing technique for on-device embodied AI, namely \emph{MosaicThinker}, which enhances the on-device small VLM's spatial reasoning capabilities on difficult cross-frame reasoning tasks. Our basic idea is to integrate fragmented spatial information from multiple frames into a unified space representation of global semantic map, and further guide the VLM's spatial reasoning over the semantic map via a visual prompt. Experiment results show that our technique can greatly enhance the accuracy of cross-frame spatial reasoning on resource-constrained embodied AI devices, over reasoning tasks with diverse types and complexities.

cs.CV

InfiniBench: Infinite Benchmarking for Visual Spatial Reasoning with Customizable Scene Complexity

Modern vision-language models (VLMs) are expected to have abilities of spatial reasoning with diverse scene complexities, but evaluating such abilities is difficult due to the lack of benchmarks that are not only diverse and scalable but also fully customizable. Existing benchmarks offer limited customizability over the scene complexity and are incapable of isolating and analyzing specific VLM failure modes under distinct spatial conditions. To address this gap, instead of individually presenting benchmarks for different scene complexities, in this paper we present InfiniBench, a fully automated, customizable and user-friendly benchmark generator that can synthesize a theoretically infinite variety of 3D scenes with parameterized control on scene complexity. InfiniBench uniquely translates scene descriptions in natural language into photo-realistic videos with complex and physically plausible 3D layouts. This is achieved through three key innovations: 1) a LLM-based agentic framework that iteratively refines procedural scene constraints from scene descriptions; 2) a flexible cluster-based layout optimizer that generates dense and cluttered scenes previously intractable for procedural methods; and 3) a task-aware camera trajectory optimization method that renders scenes into videos with full object coverage as VLM input. Experiments demonstrate that InfiniBench outperforms state-of-the-art procedural and LLM-based 3D generation methods in prompt fidelity and physical plausibility, especially in high-complexity scenarios. We further showcased the usefulness of InfiniBench, by generating benchmarks for representative spatial reasoning tasks including measurement, perspective-taking and spatiotemporal tracking.

cs.CV

Reasoning Path and Latent State Analysis for Multi-view Visual Spatial Reasoning: A Cognitive Science Perspective

Spatial reasoning is a core aspect of human intelligence that allows perception, inference and planning in 3D environments. However, current vision-language models (VLMs) struggle to maintain geometric coherence and cross-view consistency for spatial reasoning in multi-view settings. We attribute this gap to the lack of fine-grained benchmarks that isolate multi-view reasoning from single-view perception and temporal factors. To address this, we present ReMindView-Bench, a cognitively grounded benchmark for evaluating how VLMs construct, align and maintain spatial mental models across complementary viewpoints. ReMindView-Bench systematically varies viewpoint spatial pattern and query type to probe key factors of spatial cognition. Evaluations of 15 current VLMs reveals consistent failures in cross-view alignment and perspective-taking in multi-view spatial reasoning, motivating deeper analysis on the reasoning process. Explicit phase-wise analysis using LLM-as-a-judge and self-consistency prompting shows that VLMs perform well on in-frame perception but degrade sharply when integrating information across views. Implicit analysis, including linear probing and entropy dynamics, further show progressive loss of task-relevant information and uncertainty separation between correct and incorrect trajectories. These results provide a cognitively grounded diagnosis of VLM spatial reasoning and reveal how multi-view spatial mental models are formed, degraded and destabilized across reasoning phases. The ReMindView-Bench benchmark is available at https://huggingface.co/datasets/Xue0823/ReMindView-Bench, and the source codes of benchmark construction and VLM reasoning analysis are available at https://github.com/pittisl/ReMindView-Bench.

cs.AI

Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods

Spatial reasoning, which requires ability to perceive and manipulate spatial relationships in the 3D world, is a fundamental aspect of human intelligence, yet remains a persistent challenge for Multimodal large language models (MLLMs). While existing surveys often categorize recent progress based on input modality (e.g., text, image, video, or 3D), we argue that spatial ability is not solely determined by the input format. Instead, our survey introduces a taxonomy that organizes spatial intelligence from cognitive aspect and divides tasks in terms of reasoning complexity, linking them to several cognitive functions. We map existing benchmarks across text only, vision language, and embodied settings onto this taxonomy, and review evaluation metrics and methodologies for assessing spatial reasoning ability. This cognitive perspective enables more principled cross-task comparisons and reveals critical gaps between current model capabilities and human-like reasoning. In addition, we analyze methods for improving spatial ability, spanning both training-based and reasoning-based approaches. This dual perspective analysis clarifies their respective strengths, uncovers complementary mechanisms. By surveying tasks, benchmarks, and recent advances, we aim to provide new researchers with a comprehensive understanding of the field and actionable directions for future research.

cs.AI

Deciphering Personalization: Towards Fine-Grained Explainability in Natural Language for Personalized Image Generation Models

Image generation models are usually personalized in practical uses in order to better meet the individual users' heterogeneous needs, but most personalized models lack explainability about how they are being personalized. Such explainability can be provided via visual features in generated images, but is difficult for human users to understand. Explainability in natural language is a better choice, but the existing approaches to explainability in natural language are limited to be coarse-grained. They are unable to precisely identify the multiple aspects of personalization, as well as the varying levels of personalization in each aspect. To address such limitation, in this paper we present a new technique, namely \textbf{FineXL}, towards \textbf{Fine}-grained e\textbf{X}plainability in natural \textbf{L}anguage for personalized image generation models. FineXL can provide natural language descriptions about each distinct aspect of personalization, along with quantitative scores indicating the level of each aspect of personalization. Experiment results show that FineXL can improve the accuracy of explainability by 56\%, when different personalization scenarios are applied to multiple types of image generation models.

cs.LG

Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents

We present Game-TARS, a generalist game agent trained with a unified, scalable action space anchored to human-aligned native keyboard-mouse inputs. Unlike API- or GUI-based approaches, this paradigm enables large-scale continual pre-training across heterogeneous domains, including OS, web, and simulation games. Game-TARS is pre-trained on over 500B tokens with diverse trajectories and multimodal data. Key techniques include a decaying continual loss to reduce causal confusion and an efficient Sparse-Thinking strategy that balances reasoning depth and inference cost. Experiments show that Game-TARS achieves about 2 times the success rate over the previous sota model on open-world Minecraft tasks, is close to the generality of fresh humans in unseen web 3d games, and outperforms GPT-5, Gemini-2.5-Pro, and Claude-4-Sonnet in FPS benchmarks. Scaling results on training-time and test-time confirm that the unified action space sustains improvements when scaled to cross-game and multimodal data. Our results demonstrate that simple, scalable action representations combined with large-scale pre-training provide a promising path toward generalist agents with broad computer-use abilities.

cs.AI

On minimal predictable intensity of point processes

An adapted, right-continuous, non-decreasing, integer-valued process with unit jumps and starting at zero has a minimal predictable intensity if and only if it is a standard Poisson process under an absolutely continuous transformation of measures.

math.PR