arXiv ScienceSearch

arXiv · 2106.07890

An enriched category theory of language: from syntax to semantics

Abstract

State of the art language models return a natural language text continuation from any piece of input text. This ability to generate coherent text extensions implies significant sophistication, including a knowledge of grammar and semantics. In this paper, we propose a mathematical framework for passing from probability distributions on extensions of given texts, such as the ones learned by today's large language models, to an enriched category containing semantic information. Roughly speaking, we model probability distributions on texts as a category enriched over the unit interval. Objects of this category are expressions in language, and hom objects are conditional probabilities that one expression is an extension of another. This category is syntactical -- it describes what goes with what. Then, via the Yoneda embedding, we pass to the enriched category of unit interval-valued copresheaves on this syntactical category. This category of enriched copresheaves is semantic -- it is where we find meaning, logical operations such as entailment, and the building blocks for more elaborate semantic concepts.

Explore related subjects

Keep this discovery

BibTeXRIS

Tai-Danae Bradley, John Terilla, Yiannis Vlassopoulos. 2021-06-15. An enriched category theory of language: from syntax to semantics. https://arxiv.org/abs/2106.07890

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A model structure for cartesian 2-fibrations

Cartesian 2-fibrations provide a way to understand indexed categories, but their classical ``straightening'' construction requires several layers of weak coherence data. This paper develops a homotopical framework that replaces much of this bookkeeping with a fully strict model. By using marked 2-categories to record the cartesian morphisms and 2-cells, we construct a model structure whose fibrant objects are precisely the cartesian 2-fibrations over a fixed 2-category $\mathcal{C}$. We then show that the marked Grothendieck construction identifies these 2-fibrations, up to weak equivalence, with strict 2-functors from $\mathcal{C}$ into $2\mathrm{Cat}$. As an additional contribution, we construct localizations of 2-categories that simultaneously invert selected morphisms and 2-cells.

math.CT

A Natural Fuzzy Order on Fuzzy Numbers

This paper introduces a natural fuzzy order on fuzzy numbers that extends the natural orders on real numbers and interval numbers. We investigate its completeness properties and show that the space of uniformly bounded fuzzy numbers is conically complete and conically cocomplete, and that it is complete if and only if the underlying continuous t-norm is the G\"odel t-norm. Moreover, it is proved that this space constitutes a \([0,1]\)-enriched domain if and only if the underlying continuous t-norm satisfies the (S) condition. These results provide a foundation for ordering fuzzy numbers.

math.CT

Noetherian forms of free non-symmetric operads

In this paper, we study certain categories of labeled finite rooted ordered trees over a fixed set of labels where each label is equipped with an arity: a fixed number of children that the vertex with the given label must have. Equivalently, these are expression trees for operations in a free non-symmetric operad. A morphism between these trees matches a pruning of one tree (a prefix) with an entire subtree of another (a suffix). We characterize such categories, up to isomorphism, in terms of suitable exactness properties. It turns out that these categories exhibit strong algebraic behavior, in the sense that every such category, when appended with a strict initial object, has a particularly nice noetherian form.

math.CT