arXiv ScienceSearch

arXiv subjects

Jiaqi Ni

Publications and source records attributed to Jiaqi Ni.

6 recordsLinked to original sources

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

General reasoning represents a long-standing and formidable challenge in artificial intelligence. Recent breakthroughs, exemplified by large language models (LLMs) and chain-of-thought prompting, have achieved considerable success on foundational reasoning tasks. However, this success is heavily contingent upon extensive human-annotated demonstrations, and models' capabilities are still insufficient for more complex problems. Here we show that the reasoning abilities of LLMs can be incentivized through pure reinforcement learning (RL), obviating the need for human-labeled reasoning trajectories. The proposed RL framework facilitates the emergent development of advanced reasoning patterns, such as self-reflection, verification, and dynamic strategy adaptation. Consequently, the trained model achieves superior performance on verifiable tasks such as mathematics, coding competitions, and STEM fields, surpassing its counterparts trained via conventional supervised learning on human demonstrations. Moreover, the emergent reasoning patterns exhibited by these large-scale models can be systematically harnessed to guide and enhance the reasoning capabilities of smaller models.

cs.CL

DeepSeek-V3 Technical Report

We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, which were thoroughly validated in DeepSeek-V2. Furthermore, DeepSeek-V3 pioneers an auxiliary-loss-free strategy for load balancing and sets a multi-token prediction training objective for stronger performance. We pre-train DeepSeek-V3 on 14.8 trillion diverse and high-quality tokens, followed by Supervised Fine-Tuning and Reinforcement Learning stages to fully harness its capabilities. Comprehensive evaluations reveal that DeepSeek-V3 outperforms other open-source models and achieves performance comparable to leading closed-source models. Despite its excellent performance, DeepSeek-V3 requires only 2.788M H800 GPU hours for its full training. In addition, its training process is remarkably stable. Throughout the entire training process, we did not experience any irrecoverable loss spikes or perform any rollbacks. The model checkpoints are available at https://github.com/deepseek-ai/DeepSeek-V3.

cs.CL

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

We present DeepSeek-V2, a strong Mixture-of-Experts (MoE) language model characterized by economical training and efficient inference. It comprises 236B total parameters, of which 21B are activated for each token, and supports a context length of 128K tokens. DeepSeek-V2 adopts innovative architectures including Multi-head Latent Attention (MLA) and DeepSeekMoE. MLA guarantees efficient inference through significantly compressing the Key-Value (KV) cache into a latent vector, while DeepSeekMoE enables training strong models at an economical cost through sparse computation. Compared with DeepSeek 67B, DeepSeek-V2 achieves significantly stronger performance, and meanwhile saves 42.5% of training costs, reduces the KV cache by 93.3%, and boosts the maximum generation throughput to 5.76 times. We pretrain DeepSeek-V2 on a high-quality and multi-source corpus consisting of 8.1T tokens, and further perform Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) to fully unlock its potential. Evaluation results show that, even with only 21B activated parameters, DeepSeek-V2 and its chat versions still achieve top-tier performance among open-source models.

cs.CL

Littlewood-type theorems for Hardy spaces in infinitely many variables

Littlewood's theorem is one of the pioneering results in random analytic functions over the open unit disk. In this paper, we prove some analogues of this theorem for Hardy spaces in infinitely many variables. Our results not only cover finite-variable setting, but also apply in cases of Dirichlet series.

math.FA

Nevanlinna class, Dirichlet series and Szeg\"o's problem

This paper is associated with Nevanlinna class, Dirichlet series and Szeg\"o's problem in infinitely many variables. As we will see, there is a natural connection between these topics. The paper first introduces the Nevanlinna class and the Smirnov class in this context, and generalizes the classical theory in finitely many variables to the infinite-variable setting. These results applied to Szeg\"o's problem on Hardy spaces in infinitely many variables. Moreover, this paper is also devoted to the study of the correspondence between the Nevanlinna functions and Dirichlet series.

math.CV

Invariant subspaces of weighted Bergman spaces in infinitely many variables

This paper is concerned with polynomially generated multiplier invariant subspaces of the weighted Bergman space $A_{\boldsymbol{\beta}}^2$ in infinitely many variables. We completely classify these invariant subspaces under the unitary equivalence. Our results not only cover cases of both the Hardy space $H^{2}(\mathbb{D}_{2}^{\infty})$ and the Bergman space $A^{2}(\mathbb{D}_{2}^{\infty})$ in infinitely many variables, but also apply in finite-variable setting.

math.FA