arXiv ScienceSearch

arXiv subjects

Lin Niu

Publications and source records attributed to Lin Niu.

At least 19 recordsLinked to original sources

Birkhoff center and recurrent behavior of differentially positive systems on a homogeneous space

We study the structure of the Birkhoff center and the recurrent behavior of differentially positive systems whose linearizations along trajectories preserve a homogeneous cone field on a homogeneous space. Such cone fields arise naturally from general relativity and Lie theory. We establish an order-structural dichotomy for the Birkhoff center: every connected component of the Birkhoff center, as well as the support of any invariant measure, is either strongly ordered or unordered. This yields a comprehensive characterization of recurrent dynamics for differentially positive systems on homogeneous spaces, interpreted through the lens of the underlying order relation.

math.DS

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention

The quadratic cost of self-attention makes long-context inference prohibitively expensive, and proxy-based block-sparse attention has become a practical remedy. Existing methods typically rely on a proxy to predict a binary sparse mask and a kernel to consume this mask and perform sparse attention computation. Such an approach is effective under moderate budgets. However, as the budget tightens, the estimated proxy inevitably drops some salient blocks, while the kernel can only apply the sparse mask mechanically, leading to an evident drop in model accuracy. We propose CoSA, a two-stage training-free Sparse Attention under proxy-kernel CO-design, which couples a Kernel-Aware Proxy (KAP) with an Ordered-Skipping Kernel (OSK). In the first stage, the KAP selects blocks under a moderate budget and produces an ordered mask that prescribes the order in which KV pages are visited in the kernel inner loop. In the second stage, the OSK applies this mask and skips more blocks under a tightened budget given online-softmax statistics. Across mainstream LLM backbones and long-context benchmarks, CoSA attains higher accuracy at lower budgets. Impressively, CoSA achieves a 4.93$\times$ attention speedup and reduces end-to-end Time-to-First-Token by 2.53$\times$ under a context length of 128K with negligible performance degradation. Code is available at https://github.com/Tencent/AngelSlim.

cs.CL

PIVOT: Efficient Query-Group Indexing for Token-Level Sparse Attention

Token-level sparse attention, as implemented by DeepSeek Sparse Attention (DSA) in production systems, makes the downstream attention efficient but shifts the bottleneck to the indexer that feeds it. To select the top-k tokens for each query, the indexer must still score every preceding token, incurring a cost of O(L^2) per layer for a sequence of length L. We observe that this per-query scan is largely redundant: nearby queries select highly overlapping top-k tokens, and the indexer scores are long-tailed along the key axis. We exploit these properties in PIVOT, Proxy Indexing Via One full-prefix Traversal, a training-free, drop-in replacement for the DSA indexer that shares one prefix scan across a group of nearby queries. PIVOT aggregates a group into a single proxy query, performs one shared full-prefix scan to obtain a candidate set, and then selects a top-k for each query from that set. Two variants trade speed for fidelity: PIVOT-Reuse shares the proxy top-k across the group for maximum speed, whereas PIVOT-Refine re-scores the candidate set with the indexer of each query and then selects an individual top-k, matching the dense indexer at a small additional cost. A single algorithm covers both inference phases, differing only in how groups are formed: fixed-size groups of consecutive queries in prefill, and the queries decoded together in one multi-token prediction (MTP) step in decode. On DeepSeek-V3.2 and GLM-5.1 across LongBench and RULER, PIVOT matches the accuracy of the dense DSA indexer while accelerating it by up to 4x and reducing end-to-end latency by up to 1.6x at long context.

cs.CL

Hypergraph Erd\H{o}s--Rogers functions with consecutive clique sizes

For integers \(k\le s<t\), the hypergraph Erd\H{o}s--Rogers function \(f^{(k)}_{s,t}(n)\) is the largest integer \(m\) such that every \(n\)-vertex \(K_t^{(k)}\)-free \(k\)-graph contains a set of \(m\) vertices spanning no copy of \(K_s^{(k)}\). We prove that, for every fixed \(s\ge4\), \[ f^{(4)}_{s,s+1}(n)=(\log n)^{o(1)}, \] thereby resolving a problem posed by Conlon, Fox and Sudakov. The key input is a new \(3\)-uniform estimate: for every fixed \(s\ge3\), \(f^{(3)}_{s,s+1}(n)=O(\frac{\log n}{\log\log n})\), which improves the logarithmic upper bound of Dudek and Mubayi. The proof develops a probabilistic pair-coloring construction based on a robust auxiliary palette and hypergraph containers. As a further consequence, we obtain \(f^{(k)}_{k+1,k+2}(n)=(\log_{(k-3)} n)^{o(1)}\) for every fixed \(k\ge5\), making substantial progress towards a conjecture of Mubayi and Suk.

math.CO

Sharper Ramsey lower bounds from refined Gaussian estimates

Recently, Ma, Shen and Xie broke the Erd\H{o}s barrier for off-diagonal Ramsey numbers $R(\ell,C\ell)$, achieving the first exponential improvement over the classical lower bound for every $C>1$ and sufficiently large $\ell$. Hunter, Milojevi\'{c}, and Sudakov later gave a simplified proof using Gaussian random graphs and obtained better quantitative bounds. In this paper we prove a further improvement, and show that the exponent in the Ramsey lower bound can be increased by a strictly positive amount for every fixed $C>1$; as $C\to\infty$, the gain is asymptotically $\Theta(p_C^{-1/2}/\log C)$. The improvement is achieved by replacing the subgaussian estimate for truncated Gaussians with a sharp cumulant generating function bound.

math.CO

Effect of gap width on turbulent transition in Taylor-Couette flow

Simulations of the transitional flow in Taylor-Couette configuration are carried out to study the effect of the gap width on turbulent transition. The research results show that, under the same radius and the rotating speed of the inner cylinder, as the gap width increases, the flow becomes more stable. It is discovered that the average velocity distribution in the gap approaches the free vortex flow as the width increase and the stability of the flow is enhanced. It is found that, as the gap width increases, the maximum of the energy gradient function (from the energy gradient theory) in the gap decreases, which delays the turbulent transition. As such, the larger the gap width, the later the transition occurs. As the gap width increases, the Reynolds number based on the gap width alone is not able to characterize the flow behavior in Taylor-Couette flows, and the effect of the radius ratio should be taken into account.

physics.flu-dyn

Inverse Energy Cascade in Turbulent Taylor-Couette Flows

The inverse energy cascade in turbulent Taylor-Couette flow is studied in line with the results of the large eddy simulation. The simulation results show that the inverse energy cascade first occurs within the core region of the flow channel of the Taylor-Couette flow at higher Reynolds number. It is uncovered that this phenomenon is induced by the pulsed zero shear stress resulting from the singularities of the Navier-Stokes equation. In the core area between the two cylinders, the shear stress is nearly zero at higher Reynolds number. The turbulence generated there has high turbulent energy due to discontinuity of the tangential velocity. Since the energy transfer between the fluid layers is inhibited due to the low shear stress, the turbulent energy cannot be transferred along the radial direction, and small-scale vortices with high turbulent energy are produced. These small-scale vortices are located with the large-scale vortices and cannot be dissipated owing to low shear stress. A peak in the energy spectrum at middle frequency (or wave number) is formed due to the concentration of the small-scale vortices. As the number of the singular points of the Navier-Stokes equation increases with the increasing Reynolds number, the region with zero shear stress expands along the radial direction, intensifying nonlinear instability and energy accumulation. This, in turn, leads to more prominent peaks in the energy spectrum, resulting in a more pronounced inverse energy cascade.

physics.flu-dyn

Solitary wave structure of transitional flow in the wake of a sphere

The soliton-like coherent structure (SCS), which has been verified to exist in both transitional and turbulent boundary layers1-4, still poses a challenge in the understanding of its formation and behavior. In our previous study (Niu et al.5), the SCS was also found to exist in the transitional wake flow behind a sphere. In present study, the formation and evolution of the SCS is further investigated at four Reynolds numbers by numerical simulation. The results show that at the early stage of the turbulence transition, the SCS appears as a form of wave packet during the Tollmien-Schlichting (T-S) wave stage. With the increase of the Reynolds number, the SCS reaches its maximum amplitude downstream where the velocity discontinuity occurs. This position is located after the breakdown of the T-S wave and the three-dimensional structure is formed. Then, the SCS conserves its shape and amplitude over a long distance downstream. The relationships among the SCS, the spikes, the vortex structures, and the high-shear layers are further analyzed. It is found that the SCS in the wake flow has similarities to the phenomena observed in boundary layer flows during the turbulent transition. The vortex structures and high-shear layers mostly wrap around the border of the SCS. The vortex structure is considered to be as a consequence of the development of the SCS rather than its cause.

physics.flu-dyn

Stem: Rethinking Causal Information Flow in Sparse Attention

The quadratic computational complexity of self-attention remains a fundamental bottleneck for scaling Large Language Models (LLMs) to long contexts, particularly during the pre-filling phase. In this paper, we rethink the causal attention mechanism from the perspective of information flow. Due to causal constraints, tokens at initial positions participate in the aggregation of every subsequent token. However, existing sparse methods typically apply a uniform top-k selection across all token positions within a layer, ignoring the cumulative dependency of token information inherent in causal architectures. To address this, we propose Stem, a novel, plug-and-play sparsity module aligned with information flow. First, Stem employs the Token Position-Decay strategy, applying position-dependent top-k within each layer to retain initial tokens for recursive dependencies. Second, to preserve information-rich tokens, Stem utilizes the Output-Aware Metric. It prioritizes high-impact tokens based on approximate output magnitude. Extensive evaluations demonstrate that Stem achieves superior accuracy with reduced computation and pre-filling latency.

cs.LG

AngelSlim: A more accessible, comprehensive, and efficient toolkit for large model compression

This technical report introduces AngelSlim, a comprehensive and versatile toolkit for large model compression developed by the Tencent Hunyuan team. By consolidating cutting-edge algorithms, including quantization, speculative decoding, token pruning, and distillation. AngelSlim provides a unified pipeline that streamlines the transition from model compression to industrial-scale deployment. To facilitate efficient acceleration, we integrate state-of-the-art FP8 and INT8 Post-Training Quantization (PTQ) algorithms alongside pioneering research in ultra-low-bit regimes, featuring HY-1.8B-int2 as the first industrially viable 2-bit large model. Beyond quantization, we propose a training-aligned speculative decoding framework compatible with multimodal architectures and modern inference engines, achieving 1.8x to 2.0x throughput gains without compromising output correctness. Furthermore, we develop a training-free sparse attention framework that reduces Time-to-First-Token (TTFT) in long-context scenarios by decoupling sparse kernels from model architectures through a hybrid of static patterns and dynamic token selection. For multimodal models, AngelSlim incorporates specialized pruning strategies, namely IDPruner for optimizing vision tokens via Maximal Marginal Relevance and Samp for adaptive audio token merging and pruning. By integrating these compression strategies from low-level implementations, AngelSlim enables algorithm-focused research and tool-assisted deployment.

cs.LG

HunyuanVideo 1.5 Technical Report

We present HunyuanVideo 1.5, a lightweight yet powerful open-source video generation model that achieves state-of-the-art visual quality and motion coherence with only 8.3 billion parameters, enabling efficient inference on consumer-grade GPUs. This achievement is built upon several key components, including meticulous data curation, an advanced DiT architecture featuring selective and sliding tile attention (SSTA), enhanced bilingual understanding through glyph-aware text encoding, progressive pre-training and post-training, and an efficient video super-resolution network. Leveraging these designs, we developed a unified framework capable of high-quality text-to-video and image-to-video generation across multiple durations and resolutions. Extensive experiments demonstrate that this compact and proficient model establishes a new state-of-the-art among open-source video generation models. By releasing the code and model weights, we provide the community with a high-performance foundation that lowers the barrier to video creation and research, making advanced video generation accessible to a broader audience. All open-source assets are publicly available at https://github.com/Tencent-Hunyuan/HunyuanVideo-1.5.

cs.CV

Hunyuan3D Studio: End-to-End AI Pipeline for Game-Ready 3D Asset Generation

The creation of high-quality 3D assets, a cornerstone of modern game development, has long been characterized by labor-intensive and specialized workflows. This paper presents Hunyuan3D Studio, an end-to-end AI-powered content creation platform designed to revolutionize the game production pipeline by automating and streamlining the generation of game-ready 3D assets. At its core, Hunyuan3D Studio integrates a suite of advanced neural modules (such as Part-level 3D Generation, Polygon Generation, Semantic UV, etc.) into a cohesive and user-friendly system. This unified framework allows for the rapid transformation of a single concept image or textual description into a fully-realized, production-quality 3D model complete with optimized geometry and high-fidelity PBR textures. We demonstrate that assets generated by Hunyuan3D Studio are not only visually compelling but also adhere to the stringent technical requirements of contemporary game engines, significantly reducing iteration time and lowering the barrier to entry for 3D content creation. By providing a seamless bridge from creative intent to technical asset, Hunyuan3D Studio represents a significant leap forward for AI-assisted workflows in game development and interactive media.

cs.CV

HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels

Creating immersive and playable 3D worlds from texts or images remains a fundamental challenge in computer vision and graphics. Existing world generation approaches typically fall into two categories: video-based methods that offer rich diversity but lack 3D consistency and rendering efficiency, and 3D-based methods that provide geometric consistency but struggle with limited training data and memory-inefficient representations. To address these limitations, we present HunyuanWorld 1.0, a novel framework that combines the best of both worlds for generating immersive, explorable, and interactive 3D scenes from text and image conditions. Our approach features three key advantages: 1) 360{\deg} immersive experiences via panoramic world proxies; 2) mesh export capabilities for seamless compatibility with existing computer graphics pipelines; 3) disentangled object representations for augmented interactivity. The core of our framework is a semantically layered 3D mesh representation that leverages panoramic images as 360{\deg} world proxies for semantic-aware world decomposition and reconstruction, enabling the generation of diverse 3D worlds. Extensive experiments demonstrate that our method achieves state-of-the-art performance in generating coherent, explorable, and interactive 3D worlds while enabling versatile applications in virtual reality, physical simulation, game development, and interactive content creation.

cs.CV

Hunyuan3D 2.1: From Images to High-Fidelity 3D Assets with Production-Ready PBR Material

3D AI-generated content (AIGC) is a passionate field that has significantly accelerated the creation of 3D models in gaming, film, and design. Despite the development of several groundbreaking models that have revolutionized 3D generation, the field remains largely accessible only to researchers, developers, and designers due to the complexities involved in collecting, processing, and training 3D models. To address these challenges, we introduce Hunyuan3D 2.1 as a case study in this tutorial. This tutorial offers a comprehensive, step-by-step guide on processing 3D data, training a 3D generative model, and evaluating its performance using Hunyuan3D 2.1, an advanced system for producing high-resolution, textured 3D assets. The system comprises two core components: the Hunyuan3D-DiT for shape generation and the Hunyuan3D-Paint for texture synthesis. We will explore the entire workflow, including data preparation, model architecture, training strategies, evaluation metrics, and deployment. By the conclusion of this tutorial, you will have the knowledge to finetune or develop a robust 3D generative model suitable for applications in gaming, virtual reality, and industrial design.

cs.CV

Hunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets Generation

We present Hunyuan3D 2.0, an advanced large-scale 3D synthesis system for generating high-resolution textured 3D assets. This system includes two foundation components: a large-scale shape generation model -- Hunyuan3D-DiT, and a large-scale texture synthesis model -- Hunyuan3D-Paint. The shape generative model, built on a scalable flow-based diffusion transformer, aims to create geometry that properly aligns with a given condition image, laying a solid foundation for downstream applications. The texture synthesis model, benefiting from strong geometric and diffusion priors, produces high-resolution and vibrant texture maps for either generated or hand-crafted meshes. Furthermore, we build Hunyuan3D-Studio -- a versatile, user-friendly production platform that simplifies the re-creation process of 3D assets. It allows both professional and amateur users to manipulate or even animate their meshes efficiently. We systematically evaluate our models, showing that Hunyuan3D 2.0 outperforms previous state-of-the-art models, including the open-source models and closed-source models in geometry details, condition alignment, texture quality, and etc. Hunyuan3D 2.0 is publicly released in order to fill the gaps in the open-source 3D community for large-scale foundation generative models. The code and pre-trained weights of our models are available at: https://github.com/Tencent/Hunyuan3D-2

cs.CV

Almost sure convergence of differentially positive systems on a globally orderable manifold

Differentially positive systems are nonlinear systems whose linearization along trajectories preserves a cone field on a smooth Riemannian manifold. The structures of cone field come from general relativity and Lie theory. We prove that on a globally orderable manifold, the set of convergent points has full Riemann-Lebesgue measure, thus establishing almost sure convergence. This result thereby resolves a measure-theoretic form of the conjecture posed by Forni and Sepulchre in 2016 for such manifolds.

math.DS

Generic behavior of differentially positive systems on a globally orderable Riemannian manifold

Differentially positive systems are the nonlinear systems whose linearization along trajectories preserves a cone field on a smooth Riemannian manifold. One of the embryonic forms for cone fields in reality is originated from the general relativity. By utilizing the Perron-Frobenius vector fields and the $\Gamma$-invariance of cone fields, we show that generic (i.e.,``almost all" in the topological sense) orbits are convergent to certain single equilibrium. This solved a reduced version of Forni-Sepulchre's conjecture in 2016 for globally orderable manifolds.

math.DS

E-Sparse: Boosting the Large Language Model Inference through Entropy-based N:M Sparsity

Traditional pruning methods are known to be challenging to work in Large Language Models (LLMs) for Generative AI because of their unaffordable training process and large computational demands. For the first time, we introduce the information entropy of hidden state features into a pruning metric design, namely E-Sparse, to improve the accuracy of N:M sparsity on LLM. E-Sparse employs the information richness to leverage the channel importance, and further incorporates several novel techniques to put it into effect: (1) it introduces information entropy to enhance the significance of parameter weights and input feature norms as a novel pruning metric, and performs N:M sparsity without modifying the remaining weights. (2) it designs global naive shuffle and local block shuffle to quickly optimize the information distribution and adequately cope with the impact of N:M sparsity on LLMs' accuracy. E-Sparse is implemented as a Sparse-GEMM on FasterTransformer and runs on NVIDIA Ampere GPUs. Extensive experiments on the LLaMA family and OPT models show that E-Sparse can significantly speed up the model inference over the dense model (up to 1.53X) and obtain significant memory saving (up to 43.52%), with acceptable accuracy loss.

cs.LG