arXiv ScienceSearch

arXiv subjects

Daniel Persson

Publications and source records attributed to Daniel Persson.

At least 19 recordsLinked to original sources

Steerable Neural ODEs on Homogeneous Spaces

We introduce steerable neural ordinary differential equations on homogeneous spaces $M=G/H$. These models constitute a novel geometric extension of manifold neural ordinary differential equations (NODEs) that transport associated feature vectors transforming under the local symmetry group $H$. We interpret features as sections of associated vector bundles over $M$, and describe their evolution as parallel transport. This results in a coupled system of ODEs consisting of a flow equation on $M$ and a steering equation acting on features. We show that steerable NODEs are $G$-equivariant whenever the vector field generating the flow and the connection governing parallel transport are both $G$-invariant. Furthermore, we demonstrate how steerable NODEs incorporate existing NODE models and continuous normalizing flows on Lie groups. Our framework provides the geometric foundation for learning continuous-time equivariant dynamics of general vector-valued features on homogeneous spaces.

cs.LG

The Geometry of Polynomial Group Convolutional Neural Networks

We study polynomial group convolutional neural networks (PGCNNs) for an arbitrary finite group $G$. In particular, we introduce a new mathematical framework for PGCNNs using the language of graded group algebras. This framework yields two natural parameterizations of the architecture, based on Hadamard and Kronecker products, related by a linear map. We compute the dimension of the associated neuromanifold, verifying that it depends only on the number of layers and the size of the group. Furthermore, we show the general fiber of both parameterizations is trivial up to the regular group action and rescaling. Hence both parametrization maps are identifiable.

cs.LG

PEAR: Equal Area Weather Forecasting on the Sphere

Artificial intelligence is rapidly reshaping the natural sciences, with weather forecasting emerging as a flagship AI4Science application where machine learning models can now rival and even surpass traditional numerical simulations. Following the success of the landmark models Pangu Weather and Graphcast, outperforming traditional numerical methods for global medium-range forecasting, many novel data-driven methods have emerged. A common limitation shared by many of these models is their reliance on an equiangular discretization of the sphere which suffers from a much finer grid at the poles than around the equator. In contrast, in the Hierarchical Equal Area iso-Latitude Pixelization (HEALPix) of the sphere, each pixel covers the same surface area, removing unphysical biases. Motivated by a growing support for this grid in meteorology and climate sciences, we propose to perform weather forecasting with deep learning models which natively operate on the HEALPix grid. To this end, we introduce Pangu Equal ARea (PEAR), a transformer-based weather forecasting model which operates directly on HEALPix-features and outperforms the corresponding model on an equiangular grid, and other baselines, without any computational overhead. Furthermore, we perform numerical experiments on the equivariance properties of our setup and verify the performance of PEAR on climate model emulation.

cs.LG

Equivariant non-linear maps for neural networks on homogeneous spaces

This paper presents a novel framework for non-linear equivariant neural network layers on homogeneous spaces. The seminal work of Cohen et al. on equivariant $G$-CNNs on homogeneous spaces characterized the representation theory of such layers in the linear setting, finding that they are given by convolutions with kernels satisfying so-called steerability constraints. Motivated by the empirical success of non-linear layers, such as self-attention or input dependent kernels, we set out to generalize these insights to the non-linear setting. We derive generalized steerability constraints that any such layer needs to satisfy and prove the universality of our construction. The insights gained into the symmetry-constrained functional dependence of equivariant operators on feature maps and group elements informs the design of future equivariant neural network layers. We demonstrate how several common equivariant network architectures - $G$-CNNs, implicit steerable kernel networks, conventional and relative position embedded attention based transformers, and LieTransformers - may be derived from our framework.

cs.LG

Learning Chern Numbers of Topological Insulators with Gauge Equivariant Neural Networks

Equivariant network architectures are a well-established tool for predicting invariant or equivariant quantities. However, almost all learning problems considered in this context feature a global symmetry, i.e. each point of the underlying space is transformed with the same group element, as opposed to a local ``gauge'' symmetry, where each point is transformed with a different group element, exponentially enlarging the size of the symmetry group. Gauge equivariant networks have so far mainly been applied to problems in quantum chromodynamics. Here, we introduce a novel application domain for gauge-equivariant networks in the theory of topological condensed matter physics. We use gauge equivariant networks to predict topological invariants (Chern numbers) of multiband topological insulators. The gauge symmetry of the network guarantees that the predicted quantity is a topological invariant. We introduce a novel gauge equivariant normalization layer to stabilize the training and prove a universal approximation theorem for our setup. We train on samples with trivial Chern number only but show that our models generalize to samples with non-trivial Chern number. We provide various ablations of our setup. Our code is available at https://github.com/sitronsea/GENet/tree/main.

cs.LG

Equivariant Manifold Neural ODEs and Differential Invariants

In this paper, we develop a manifestly geometric framework for equivariant manifold neural ordinary differential equations (NODEs) and use it to analyse their modelling capabilities for symmetric data. First, we consider the action of a Lie group $G$ on a smooth manifold $M$ and establish the equivalence between equivariance of vector fields, symmetries of the corresponding Cauchy problems, and equivariance of the associated NODEs. We also propose a novel formulation, based on Lie theory for symmetries of differential equations, of the equivariant manifold NODEs in terms of the differential invariants of the action of $G$ on $M$, which provides an efficient parameterisation of the space of equivariant vector fields in a way that is agnostic to both the manifold $M$ and the symmetry group $G$. Second, we construct augmented manifold NODEs, through embeddings into flows on the tangent bundle $TM$, and show that they are universal approximators of diffeomorphisms on any connected $M$. Furthermore, we show that universality persists in the equivariant case and that the augmented equivariant manifold NODEs can be incorporated into the geometric framework using higher-order differential invariants. Finally, we consider the induced action of $G$ on different fields on $M$ and show how it can be used to generalise previous work, on, e.g., continuous normalizing flows, to equivariant models in any geometry.

cs.LG

HEAL-SWIN: A Vision Transformer On The Sphere

High-resolution wide-angle fisheye images are becoming more and more important for robotics applications such as autonomous driving. However, using ordinary convolutional neural networks or vision transformers on this data is problematic due to projection and distortion losses introduced when projecting to a rectangular grid on the plane. We introduce the HEAL-SWIN transformer, which combines the highly uniform Hierarchical Equal Area iso-Latitude Pixelation (HEALPix) grid used in astrophysics and cosmology with the Hierarchical Shifted-Window (SWIN) transformer to yield an efficient and flexible model capable of training on high-resolution, distortion-free spherical data. In HEAL-SWIN, the nested structure of the HEALPix grid is used to perform the patching and windowing operations of the SWIN transformer, enabling the network to process spherical representations with minimal computational overhead. We demonstrate the superior performance of our model on both synthetic and real automotive datasets, as well as a selection of other image datasets, for semantic segmentation, depth regression and classification tasks. Our code is publicly available at https://github.com/JanEGerken/HEAL-SWIN.

cs.CV

Massive Theta Lifts

We use Poincare series for massive Maass-Jacobi forms to define a "massive theta lift", and apply it to the examples of the constant function and the modular invariant j-function, with the Siegel-Narain theta function as integration kernel. These theta integrals are deformations of known one-loop string threshold corrections. Our massive theta lifts fall off exponentially, so some Rankin-Selberg integrals are finite without Zagier renormalization.

hep-th

Equivariance versus Augmentation for Spherical Images

We analyze the role of rotational equivariance in convolutional neural networks (CNNs) applied to spherical images. We compare the performance of the group equivariant networks known as S2CNNs and standard non-equivariant CNNs trained with an increasing amount of data augmentation. The chosen architectures can be considered baseline references for the respective design paradigms. Our models are trained and evaluated on single or multiple items from the MNIST or FashionMNIST dataset projected onto the sphere. For the task of image classification, which is inherently rotationally invariant, we find that by considerably increasing the amount of data augmentation and the size of the networks, it is possible for the standard CNNs to reach at least the same performance as the equivariant network. In contrast, for the inherently equivariant task of semantic segmentation, the non-equivariant networks are consistently outperformed by the equivariant networks with significantly fewer parameters. We also analyze and compare the inference latency and training times of the different networks, enabling detailed tradeoff considerations between equivariant architectures and data augmentation for practical problems. The equivariant spherical networks used in the experiments are available at https://github.com/JanEGerken/sem_seg_s2cnn .

cs.LG

BPS Algebras in 2D String Theory

We discuss a set of heterotic and type II string theory compactifications to 1+1 dimensions that are characterized by factorized internal worldsheet CFTs of the form $V_1\otimes \bar V_2$, where $V_1, V_2$ are self-dual (super) vertex operator algebras. In the cases with spacetime supersymmetry, we show that the BPS states form a module for a Borcherds-Kac-Moody (BKM) (super)algebra, and we prove that for each model the BKM (super)algebra is a symmetry of genus zero BPS string amplitudes. We compute the supersymmetric indices of these models using both Hamiltonian and path integral formalisms. The path integrals are manifestly automorphic forms closely related to the Borcherds-Weyl-Kac denominator. Along the way, we comment on various subtleties inherent to these low-dimensional string compactifications.

hep-th

Geometric Deep Learning and Equivariant Neural Networks

We survey the mathematical foundations of geometric deep learning, focusing on group equivariant and gauge equivariant neural networks. We develop gauge equivariant convolutional neural networks on arbitrary manifolds $\mathcal{M}$ using principal bundles with structure group $K$ and equivariant maps between sections of associated vector bundles. We also discuss group equivariant neural networks for homogeneous spaces $\mathcal{M}=G/K$, which are instead equivariant with respect to the global symmetry $G$ on $\mathcal{M}$. Group equivariant layers can be interpreted as intertwiners between induced representations of $G$, and we show their relation to gauge equivariant convolutional layers. We analyze several applications of this formalism, including semantic segmentation and object detection networks. We also discuss the case of spherical networks in great detail, corresponding to the case $\mathcal{M}=S^2=\mathrm{SO}(3)/\mathrm{SO}(2)$. Here we emphasize the use of Fourier analysis involving Wigner matrices, spherical harmonics and Clebsch-Gordan coefficients for $G=\mathrm{SO}(3)$, illustrating the power of representation theory for deep learning.

cs.LG

Fun with $F_{24}$

We study some special features of $F_{24}$, the holomorphic $c=12$ superconformal field theory (SCFT) given by 24 chiral free fermions. We construct eight different Lie superalgebras of "physical" states of a chiral superstring compactified on $F_{24}$, and we prove that they all have the structure of Borcherds-Kac-Moody superalgebras. This produces a family of new examples of such superalgebras. The models depend on the choice of an $\mathcal{N}=1$ supercurrent on $F_{24}$, with the admissible choices labeled by the semisimple Lie algebras of dimension 24. We also discuss how $F_{24}$, with any such choice of supercurrent, can be obtained via orbifolding from another distinguished $c=12$ holomorphic SCFT, the $\mathcal{N}=1$ supersymmetric version of the chiral CFT based on the $E_8$ lattice.

hep-th

Emergent Sasaki-Einstein geometry and AdS/CFT

We consider supergravity in five-dimensional Anti-De Sitter space $AdS_{5}$ with minimal supersymmetry, encoded by a Sasaki-Einstein metric on a five-dimensional compact manifold $M$. Our main result reveals how the Sasaki-Einstein metric emerges from a canonical state in the dual CFT, defined by a superconformal gauge theory in four dimensional Minkowski space $\mathbb{R}^{3,1}$in the t'Hooft limit where the rank $N$ tends to infinity. We obtain explicit finite $N-$approximations to the Sasaki-Einstein metric, expressed in terms of a canonical (i.e. background free) BPS-state on the gauge theory side. We also provide a string theory interpretation of the BPS-state in question, which sheds new light on the previously noted intriguing duality of giant gravitons.

hep-th

Eulerianity of Fourier coefficients of automorphic forms

We study the question of Eulerianity (factorizability) for Fourier coefficients of automorphic forms, and we prove a general transfer theorem that allows one to deduce the Eulerianity of certain coefficients from that of another coefficient. We also establish a `hidden' invariance property of Fourier coefficients. We apply these results to minimal and next-to-minimal automorphic representations, and deduce Eulerianity for a large class of Fourier and Fourier-Jacobi coefficients. In particular, we prove Eulerianity for parabolic Fourier coefficients with characters of maximal rank for a class of Eisenstein series in minimal and next-to-minimal representations of groups of ADE-type that are of interest in string theory.

math.NT

Fourier coefficients of minimal and next-to-minimal automorphic representations of simply-laced groups

In this paper we analyze Fourier coefficients of automorphic forms on a finite cover $G$ of an adelic split simply-laced group. Let $\pi$ be a minimal or next-to-minimal automorphic representation of $G$. We prove that any $\eta\in \pi$ is completely determined by its Whittaker coefficients with respect to (possibly degenerate) characters of the unipotent radical of a fixed Borel subgroup, analogously to the Piatetski-Shapiro--Shalika formula for cusp forms on $GL_n$. We also derive explicit formulas expressing the form, as well as all its maximal parabolic Fourier coefficient in terms of these Whittaker coefficients. A consequence of our results is the non-existence of cusp forms in the minimal and next-to-minimal automorphic spectrum. We provide detailed examples for $G$ of type $D_5$ and $E_8$ with a view towards applications to scattering amplitudes in string theory.

math.NT

A reduction principle for Fourier coefficients of automorphic forms

We consider a general class of Fourier coefficients for an automorphic form on a finite cover of a reductive adelic group ${\bf G}(\mathbb{A}_{\mathbb{K}})$, associated to the data of a `Whittaker pair'. We describe a quasi-order on Fourier coefficients, and an algorithm that gives an explicit formula for any coefficient in terms of integrals and sums involving higher coefficients. The maximal elements for the quasi-order are `Levi-distinguished' Fourier coefficients, which correspond to taking the constant term along the unipotent radical of a parabolic subgroup, and then further taking a Fourier coefficient with respect to a $\mathbb{K}$-distinguished nilpotent orbit in the Levi quotient. Thus one can express any Fourier coefficient, including the form itself, in terms of higher Levi-distinguished coefficients. In follow-up papers we use this result to determine explicit Fourier expansions of minimal and next-to-minimal automorphic forms on split simply-laced reductive groups, and to obtain Euler product decompositions of their top Fourier coefficients.

math.NT

Fourier coefficients attached to small automorphic representations of ${\mathrm{SL}}_n(\mathbb{A})$

We show that Fourier coefficients of automorphic forms attached to minimal or next-to-minimal automorphic representations of ${\mathrm{SL}}_n(\mathbb{A})$ are completely determined by certain highly degenerate Whittaker coefficients. We give an explicit formula for the Fourier expansion, analogously to the Piatetski-Shapiro-Shalika formula. In addition, we derive expressions for Fourier coefficients associated to all maximal parabolic subgroups. These results have potential applications for scattering amplitudes in string theory.

math.RT

Dualities in CHL-Models

We define a very general class of CHL-models associated with any string theory (bosonic or supersymmetric) compactified on an internal CFT C x T^d. We take the orbifold by a pair (g,\delta), where g is a (possibly non-geometric) symmetry of C and \delta is a translation along T^d. We analyze the T-dualities of these models and show that in general they contain Atkin-Lehner type symmetries. This generalizes our previous work on N=4 CHL-models based on heterotic string theory on T^6 or type II on K3 x T^2, as well as the `monstrous' CHL-models based on a compactification of heterotic string theory on the Frenkel-Lepowsky-Meurman CFT V^{\natural}.

hep-th