arXiv ScienceSearch

arXiv subjects

Eric Harper

Publications and source records attributed to Eric Harper.

9 recordsLinked to original sources

Heterogeneous Parallelism for Multimodal Large Language Model Training

Foundation model training is becoming multimodal, from post-training pipelines to large-scale pretraining. As modality coverage broadens, context windows grow, and encoder LLM scales diverge, a single LLM-centric TP/CP/PP/DP/EP layout increasingly limits throughput. This coupling forces encoders to inherit LLM-driven sharding and placement choices that can add communication, limit encoder parallelism, or constrain the LLM schedule; the mismatch is most pronounced at long contexts, where LLM context parallelism is needed for the fused multimodal sequence but encoder inputs remain bounded. We present heterogeneous parallelism for multimodal large language model training, an abstraction that lets modules in one end-to-end graph use independent layouts and rank placements, supporting colocated execution on shared GPUs and non-colocated execution on disjoint rank sets. The key challenge is preserving boundary tensor semantics across independent layouts: forward activations must be materialized for the destination layout, while backward gradients must be routed back to the source layout. We address this with boundary communicators that implement forward and backward layout transforms, plus scheduling extensions for both placement modes. We evaluate optimized homogeneous, colocated heterogeneous, and non-colocated heterogeneous configurations across multimodal workloads and GPU scales to characterize when added layout and placement freedom exposes a better operating point. Across this sweep, colocated heterogeneity improves TFLOPS/GPU by up to 49.3%, while non-colocated heterogeneity improves aggregate token throughput by up to 13.0% and TFLOPS/GPU by up to 9.6%. We validate loss convergence parity against homogeneous baselines and release the system as an open-source Megatron-LM extension.

cs.LG

NVIDIA Nemotron 3: Efficient and Open Intelligence

We introduce the Nemotron 3 family of models - Nano, Super, and Ultra. These models deliver strong agentic, reasoning, and conversational capabilities. The Nemotron 3 family uses a Mixture-of-Experts hybrid Mamba-Transformer architecture to provide best-in-class throughput and context lengths of up to 1M tokens. Super and Ultra models are trained with NVFP4 and incorporate LatentMoE, a novel approach that improves model quality. The two larger models also include MTP layers for faster text generation. All Nemotron 3 models are post-trained using multi-environment reinforcement learning enabling reasoning, multi-step tool use, and support granular reasoning budget control. Nano, the smallest model, outperforms comparable models in accuracy while remaining extremely cost-efficient for inference. Super is optimized for collaborative agents and high-volume workloads such as IT ticket automation. Ultra, the largest model, provides state-of-the-art accuracy and reasoning performance. Nano is released together with its technical report and this white paper, while Super and Ultra will follow in the coming months. We will openly release the model weights, pre- and post-training software, recipes, and all data for which we hold redistribution rights.

cs.CL

Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

We present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model. Nemotron 3 Nano was pretrained on 25 trillion text tokens, including more than 3 trillion new unique tokens over Nemotron 2, followed by supervised fine tuning and large-scale RL on diverse environments. Nemotron 3 Nano achieves better accuracy than our previous generation Nemotron 2 Nano while activating less than half of the parameters per forward pass. It achieves up to 3.3x higher inference throughput than similarly-sized open models like GPT-OSS-20B and Qwen3-30B-A3B-Thinking-2507, while also being more accurate on popular benchmarks. Nemotron 3 Nano demonstrates enhanced agentic, reasoning, and chat abilities and supports context lengths up to 1M tokens. We release both our pretrained Nemotron 3 Nano 30B-A3B Base and post-trained Nemotron 3 Nano 30B-A3B checkpoints on Hugging Face.

cs.CL

The SU(N) Casson-Lin invariants for links

We introduce the $SU(N)$ Casson-Lin invariants for links $L$ in $S^3$ with more than one component. Writing $L = \ell_1 \cup \cdots \cup \ell_n$, we require as input an $n$-tuple $(a_1,\ldots, a_n) \in {\mathbb Z}^n$ of labels, where $a_j$ is associated with $\ell_j$. The $SU(N)$ Casson-Lin invariant, denoted $h_{N,a}(L)$, gives an algebraic count of certain projective $SU(N)$ representations of the link group $\pi_1(S^3 \smallsetminus L)$, and the family $h_{N,a}$ of link invariants gives a natural extension of the $SU(2)$ Casson-Lin invariant, which was defined for knots by X.-S. Lin and for 2-component links by Harper and Saveliev. We compute the invariants for the Hopf link and more generally for chain links, and we show that, under mild conditions on the labels $(a_1, \ldots, a_n)$, the invariants $h_{N,a}(L)$ vanish whenever $L$ is a split link.

math.GT

Virtual knot groups and almost classical knots

We define a group-valued invariant of virtual knots and relate it to various other group-valued invariants of virtual knots, including the extended group of Silver-Williams and the quandle group of Manturov and Bardakov-Bellingeri. A virtual knot is called almost classical if it admits a diagram with an Alexander numbering, and in that case we show that the group factors as a free product of the usual knot group and Z. We establish a similar formula for mod p almost classical knots, and we use these results to derive obstructions to a virtual knot K being mod p almost classical. Viewed as knots in thickened surfaces, almost classical knots correspond to those that are homologically trivial. We show they admit Seifert surfaces and relate their Alexander invariants to the homology of the associated infinite cyclic cover. We prove the first Alexander ideal is principal, recovering a result first proved by Nakamura et al. using different methods. The resulting Alexander polynomial is shown to satisfy a skein relation, and its degree gives a lower bound for the Seifert genus. We tabulate almost classical knots up to 6 crossings and determine their Alexander polynomials and virtual genus.

math.GT

Alexander invariants for virtual knots

Given a virtual knot $K$, we construct a group $VG_K$ called the virtual knot group, and we use the elementary ideals of $VG_K$ to define invariants of $K$ called the virtual Alexander invariants. For instance, associated to the $k=0$ ideal is a polynomial $H_K(s,t,q)$ in three variables which we call the virtual Alexander polynomial, and we show that it is closely related to the generalized Alexander polynomial $G_K(s,t)$ introduced by Sawollek, Kauffman-Radford, and Silver-Williams. We define a natural normalization of the virtual Alexander polynomial and show it satisfies a skein formula. We also introduce the twisted virtual Alexander polynomial associated to a virtual knot $K$ and a representation $\varrho \colon VG_K \to GL_n(R)$, and we define a normalization of the twisted virtual Alexander polynomial. As applications we derive bounds on the virtual crossing numbers of virtual knots from the virtual Alexander polynomial and twisted virtual Alexander polynomial.

math.GT

On instanton homology of corks W_n

We consider a family of corks, denoted $W_n$, constructed by Akbulut and Yasui. Each cork gives rise to an exotic structure on a smooth 4-manifold via a twist $τ$ on its boundary $Σ_n = \partial W_n$. We compute the instanton Floer homology of $Σ_n$ and show that the map induced on the instanton Floer homology by $τ: Σ_n \rightarrow Σ_n$ is non-trivial.

math.GT

Instanton Floer homology for two-component links

For any link of two components in an integral homology sphere, we define an instanton Floer homology whose Euler characteristic is the linking number between the components of the link. We relate this Floer homology to the Kronheimer-Mrowka instanton Floer homology of knots. We also show that, for two-component links in the 3-sphere, the Floer homology does not vanish unless the link is split.

math.GT

A Casson-Lin type invariant for links

We define an integer valued invariant for two-component links in S^3 by counting projective SU(2) representations of the link group having non-trivial second Stiefel-Whitney class. We show that our invariant is, up to sign, the linking number of the link. Our construction generalizes that of X.-S. Lin who defined a similar invariant for knots in S^3; his invariant equals half the knot signature.

math.GT