arXiv Science⌕ Search

arXiv subjects

Zhiyuan Yang

Publications and source records attributed to Zhiyuan Yang.

At least 19 recordsLinked to original sources

The lower bound of shifted primes with large prime factors

In this paper, we consider the asymptotic density of $T_c(x):=\#\{p\leq x:P^+(p-1)\geq p^c\}$, where $P^+(n)$ denote the largest prime factor of $n$. We show that for $x\rightarrow\infty$, for any $0 0$ for $11/32 0.0436$. This improves a previous result by Liu-Wu-Xi (2020), who showed \begin{align*} \mathop{\lim\inf}_{x\rightarrow\infty}\frac{T_c(x)}{π(x)}&\geq\max\left(1-4ρ\left(\frac{1}{c}\right),1-c\right), \end{align*} for $0<c\leq 1/2$.

math.NT↗

Bi-exact Wreath-like Product Groups

In this short paper, we prove the bi-exactness of all wreath-like product groups $G\in\mathcal{WR}(A,B\curvearrowright I)$, whenever $A$ is amenable, $B$ is bi-exact, and the action $B\curvearrowright I$ has amenable stabilizers. As applications, we obtain solidity for the associated group von Neumann algebras and, by combining our result with existing rigidity theorems, new examples exhibiting the McDuff superrigidity.

math.OA↗

On the largest prime factors less than $y$ of consecutive shifted primes

For an integer $n > 1$, let $P^+(n)$ be the largest prime factor of $n$, and let $P_y^+(n)$ denote the largest prime factor of $n$ not exceeding $y$. One of Erdős and Turán's conjectures asserts that the asymptotic density of integers $n$ satisfying $P^+(n) 0$ such that \begin{align*} \#\{p\leq x:P_y^+(p-1)<P_y^+(p+1)\}\geq(h(α)+o(1))π(x). \end{align*} In particular, the function $h$ satisfies $\lim_{α\rightarrow 0^+}h(α)=1/2$. Similar result also holds for $\#\{n\leq x:P_y^+(n)<P_y^+(n+1)\}$. These improve Rivat's result (2001) and Wang's result (2019).

math.NT↗

SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space

World Action Models (WAMs) couple action generation with prediction of future states. Their effectiveness depends on whether future dynamics are modeled in a space that is both aligned with action generation and sufficiently geometry-aware to capture where and how actions change the scene. Existing WAMs typically satisfy only part of this requirement, relying on either perceptually heavy observation-space targets or auxiliary latent spaces that are not jointly structured for action relevance and geometry. We propose SG-WAM, a self-guided framework that learns geometry-aware action-conditioned dynamics directly in the policy-derived representation space. SG-WAM introduces learnable dynamics tokens and a Self-Guided World Predictor that forecasts their future latent states conditioned on intervening robot actions. Prediction targets are generated by an exponential moving average copy of the same policy backbone, providing stable supervision within the representation family used by the action expert. Geometric supervision further structures the policy image-token representations, providing spatially grounded context for the dynamics tokens and yielding a future-alignment space that is both action-relevant and geometry-aware. Latent future prediction, geometric grounding, and flow-matching action generation are jointly optimized end-to-end in a unified framework. Built on a 0.9B model without large-scale embodied pretraining, SG-WAM achieves 98.5% average success on LIBERO and 73% on LIBERO-Plus, while outperforming strong baselines in both in-distribution and out-of-distribution real-world evaluations.

cs.RO↗

Closure complexity of longest-edge bisection for triangular meshes

On triangular meshes, we analyze local mesh refinement based on longest-edge bisection equipped with the serial longest-edge propagation-path closure. Ties are resolved by terminal priority: if the incoming shared edge is a longest edge of the neighboring triangle, the pair is declared terminal and that edge is bisected. For every adaptive mesh sequence $\mathcal{T}_0, \mathcal{T}_1, \ldots, \mathcal{T}_L$ with a sequence of marked sets $\mathcal{M}_0, \mathcal{M}_1, \ldots, \mathcal{M}_{L-1}$, we prove the cumulative closure estimate $$\#\mathcal{T}_L-\#\mathcal{T}_0 \lesssim\sum_{\ell=0}^{L-1}\#\mathcal{M}_\ell.$$ The proof derives single-mark locality from the uniform multiplicative gap between possible descendant diameters implied by finite similarity classes, and converts this locality into the cumulative estimate through a Binev--Dahmen--DeVore-type charging argument.

math.NA↗

Convolution-type Bombieri-Vinogradov theorem with well-factorable weights, and its applications

In this paper, we consider the asymptotic density of $\#\{p\leq x:P^+(p-1)\geq p^c\}$ and $\#\{n\leq x:P^+(n) 0.299x. \end{align*} The first result constitutes an improvement upon that of Ding and Wang (2025), who obatined $\mathop{\lim\sup}_{x\rightarrow\infty} \frac{1}{π(x)}\#\{p\leq x:P^+(p-1)\geq p^c\}\leq \frac{7}{2}\log\frac{1}{c}$. The second result improves a previous result $0.280$ by the author (2026). The key to the proof is that for a special class of convolution forms equipped with well-factorable weights, we may use the level \(x^{5/8-o(1)}\) for Pascadi's prime-distribution result with triple-well-factorable weights. We also use Pascadi's estimation of incomplete Kloosterman sums.

math.NT↗

Closure complexity of Bänsch-type algorithms for tetrahedral mesh refinement

We prove, to our knowledge, the first unconditional cumulative closure estimates for the Arnold--Mukherjee--Pouly (AMP) refinement algorithm and the original face-marked tetrahedral algorithm of Bänsch on arbitrary conforming initial tetrahedral meshes. Let $T_0,\ldots,T_L$ be an adaptive mesh sequence generated by either algorithm, with $M_\ell\subseteq T_\ell$ denoting the marking set at step $\ell$. Then $$\#T_L-\#T_0\leq C_{\mathrm{clos}}(T_0)\sum_{\ell=0}^{L-1}\#M_\ell.$$ The proof is intrinsic to the physical three-dimensional mesh and requires neither an initial compatibility condition nor a higher-dimensional embedding. It organizes conformity refinements into a causal forest and combines a uniform horizontal-propagation estimate with a weighted packing argument to obtain an explicit closure constant. For the original Bänsch algorithm, every history-dependent resolution of the initial two-edge ambiguity is represented by one of finitely many AMP histories. The estimate therefore holds uniformly for arbitrary deterministic or nondeterministic choices. This resolves a long-standing complexity question for the Bänsch--AMP family.

math.NA↗

Relative Biexactness for Relative Hyperbolic Groups and Some Applications

In this paper, we confirm a conjecture of Ozawa and others asserting that every finitely generated, relatively hyperbolic, exact group is bi-exact (in the sense of Ozawa) relative to its natural peripheral structure. As a consequence, every such group gives rise to a prime group von Neumann algebra. As an application, we construct a continuum family of property (T), relatively hyperbolic groups $\{G_i\}_{i\in I}$ such that, for every fixed arbitrary free, ergodic, probability measure-preserving action $G_i \curvearrowright Z_i$, the collection of associated group measure space von Neumann algebras $\{L^\infty(Z_i)\rtimes G_i\}_{i\in I}$ are pairwise non-stably $\ast$-isomorphic.

math.OA↗

Gated Spatial Redundancy Projection for Pathology Transformer Attentions

Transformer models are increasingly used for whole-slide image analysis in computational pathology. Yet, WSIs differ fundamentally from natural images: neighbouring patches often contain highly similar tissue type, stain, texture, and cellular composition. We identify this local spatial redundancy as a pathology-specific failure mode of self-attention, where dominant neighbourhood features can be repeatedly mixed into patch-tokens and weaken subtle diagnostic or prognostic deviations. We propose Gated Spatial Redundancy Projection (Gated SRP), a lightweight drop-in correction module for self-attention layers. For each patch token and attention head, Gated SRP estimates a local redundancy axis from neighbouring value vectors, projects the attention output onto this axis, and applies a learned signed gate to correct the redundancy-aligned component geometrically. Across five TCGA survival cohorts, Gated SRP obtains the highest mean C-index among the compared attention variants in all cohorts, with an average improvement over the base attention, while adding only +0.02% parameters. Across five slide-level classification datasets, it improves the base attention on 12 of 16 reported metrics and achieves the best AUC on three datasets. Code is publicly available at https://github.com/AtlasAnalyticsLab/GatedSRP.

cs.CV↗

Simple Token-Efficient Vision-Language Model for Case-level Pathology Synoptic Report Generation

Generating clinically useful pathology reports for pathology cases from whole-slide images (WSIs) is challenging due to gigapixel resolution, long visual-token sequences, and the complexity of case-level reasoning, where a single case may contain multiple WSIs with heterogeneous tissues and ambiguous findings. We present a simple token-efficient vision--language model for case-level synoptic report generation that remains practical under constrained GPU memory. Our architecture follows a minimal three-component design: a frozen pathology patch encoder, a lightweight two-layer MLP vision-language aligner, and a large language model decoder, with an explicit WSI marker token to separate slides within a case. Training proceeds in two supervised stages: (1) aligner-only WSI captioning using heterogeneous WSI-text pairs, and (2) case-level supervised fine-tuning on case-report pairs for structured report generation. To reduce sequence length, we represent each slide using $512 \times 512$ patches at $5\times$ magnification, which reduces the average sequence length by up to $64\times$ times compared to the commonly used $20\times$ patches. Combined with efficient training techniques, we enable practical training with only half a NVIDIA H100 GPU. Across both training stages, our approach achieves high ROUGE-L/METEOR/BLEU-4 scores while being substantially more efficient in memory and runtime. In AI-based evaluations, our model is consistently preferred over strong baselines. Extensive ablations characterize performance-efficiency trade-offs and identify simple choices that improve robustness in multi-WSI settings. Overall, this work provides a strong, reproducible baseline for efficient pathology report generation, lowering the barrier to multi-WSI VLM research under limited compute. Code is available at https://github.com/AtlasAnalyticsLab/PathoSynVLM.

cs.CV↗

Relative biexactness and mixing in von Neumann algebras

We develop a new technique to upgrade relative biexactness in general von Neumann algebras: suppose that $\{N_i\}_{i\in I}\subset M$ are mixing and biexact subalgebras of a separable von Neumann algebra with expectation, and if $M$ is biexact relative to $\{N_i\}_{i\in I}$, then $M$ is biexact. This result yields several new examples of biexact von Neumann algebras, notably including amalgamated free products. By generalizing the relative biexactness results of Hoshino to the von Neumann algebra setting and applying our result above along with certain bimodule computations, we in fact obtain, as an application, a new classification result for biexactness for graph products of finite dimensional von Neumann algebras. This yields significant extensions of prior works of Caspers-Borst, and Blufstein-Goldman-Oyakawa.

math.OA↗

HoloMotion-1 Technical Report

In this report, we present HoloMotion-1, a humanoid motion foundation model for zero-shot whole-body motion tracking. A key innovation of HoloMotion-1 is to scale control-policy training with a large-scale hybrid motion corpus, where video-reconstructed motions from in-the-wild videos provide the dominant source of motion diversity, while curated motion-capture and in-house motion data provide higher-fidelity supervision and deployment-oriented coverage. This data regime enables HoloMotion-1 to move beyond conventional MoCap-only training and exposes the policy to substantially broader behaviors, capture conditions, and motion styles. Learning from such heterogeneous data introduces new challenges, including reconstruction noise, source-domain mismatch, uneven motion quality, and the need for temporal modeling under large behavioral variation. To address these challenges, HoloMotion-1 integrates large-capacity temporal modeling, a sparsely activated Mixture-of-Experts Transformer with KV-cache inference for real-time control, and a sequence-level training strategy that improves learning efficiency on extended motion sequences. Extensive experiments on multiple unseen motion benchmarks show that HoloMotion-1 generalizes robustly across diverse motion types and capture conditions, significantly improves tracking accuracy over prior methods, and transfers directly to a real humanoid robot without task-specific fine-tuning.

cs.RO↗

An Operator-Valued Haagerup Inequality for Hyperbolic Groups

We study an operator-valued generalization of the Haagerup inequality for Gromov hyperbolic groups. In 1978, U. Haagerup showed that if $f$ is a function on the free group $\mathbb{F}_r$ which is supported on the $k$-sphere $S_k=\{x\in \mathbb{F}_r:\ell(x)=k\}$, then the operator norm of its left regular representation is bounded by $(k+1)\|f\|_2$. An operator-valued generalization of it was started by U. Haagerup and G. Pisier. One of the most complete form was given by A. Buchholz, where the $\ell^2$-norm in the original inequality was replaced by $k+1$ different matrix norms associated to word decompositions (this type of inequality is also called Khintchine-type inequality). We provide a generalization of Buchholz's result for hyperbolic groups.

math.OA↗

Non-Hermitian Anomalous Scaling Engineering

Non-Hermitian systems exhibit anomalous scaling, a striking departure from conventional bulk laws, rooted in the non-Hermitian skin effect (NHSE). Here, we experimentally uncover this scaling and demonstrate its active control in a temporal photonic lattice. By tracking the real-time evolution of all eigenstates as system size varies, we directly observe scaling-driven spectral reshaping and eigenstate localization, revealing phenomena absent in Hermitian or NHSE-free lattices. In a Su-Schrieffer-Heeger lattice, scaling alone can trigger a non-Hermitian topological phase transition, with edge modes remaining protected. Crucially, Kerr interactions open the frontier of nonlinear non-Hermitian physics: weak nonlinearity accelerates or decelerates anomalous scaling, while strong nonlinearity suppresses it entirely. These results establish the first experimental platform for linear and nonlinear anomalous scaling engineering, paving the way for compact non-Hermitian devices and exploration of nonlinear and many-body non-Hermitian phenomena.

physics.optics↗

RoCo Challenge at AAAI 2026: Benchmarking Robotic Collaborative Manipulation for Assembly Towards Industrial Automation

Embodied Artificial Intelligence (EAI) is rapidly developing, gradually subverting previous autonomous systems' paradigms from isolated perception to integrated, continuous action. This transition is highly significant for industrial robotic manipulation, promising to free human workers from repetitive, dangerous daily labor. To benchmark and advance this capability, we introduce the Robotic Collaborative Assembly Assistance (RoCo) Challenge with a dataset towards simulation and real-world assembly manipulation. Set against the backdrop of human-centered manufacturing, this challenge focuses on a high-precision planetary gearbox assembly task, a demanding yet highly representative operation in modern industry. Built upon a self-developed data collection, training, and evaluation system in Isaac Sim, and utilizing a dual-arm robot for real-world deployment, the challenge operates in two phases. The Simulation Round defines fine-grained task phases for step-wise scoring to handle the long-horizon nature of the assembly. The Real-World Round mirrors this evaluation with physical gearbox components and high-quality teleoperated datasets. The core tasks require assembling an epicyclic gearbox from scratch, including mounting three planet gears, a sun gear, and a ring gear. Attracting over 60 teams and 170+ participants from more than 10 countries, the challenge yielded highly effective solutions, most notably ARC-VLA and RoboCola. Results demonstrate that a dual-model framework for long-horizon multi-task learning is highly effective, and the strategic utilization of recovery-from-failure curriculum data is a critical insight for successful deployment. This report outlines the competition setup, evaluation approach, key findings, and future directions for industrial EAI. Our dataset, CAD files, code, and evaluation results can be found at: https://rocochallenge.github.io/RoCo2026/.

cs.RO↗

FreeFly-Thinking : Aligning Chain-of-Thought Reasoning with Continuous UAV Navigation

Vision-Language Navigation aims to enable agents to understand natural language instructions and carry out appropriate navigation actions in real-world environments. Most work focuses on indoor settings, with little research in complex outdoor scenes. Current UAV Vision-and-Language Navigation models typically act as black boxes without explicit reasoning. We introduce FreeFly-thinking, an end-to-end VLN framework that converts the UAV agent's egocentric images and language instructions into a series of actions, inspired by environment of urban architecture proposed by OpenFly. We first construct a UAV dataset for navigation task, and then performing natural language chain of thought. We adopt a two-stage training strategy: Supervised fine-tuning and Reinforcement fine-tuning. Experiments on unseen test demonstrate a strong performance, presenting robustness and efficiency in UAV navigation issue.

cs.CV↗