arXiv ScienceSearch

arXiv subjects

Sam Adam-Day

Publications and source records attributed to Sam Adam-Day.

10 recordsLinked to original sources

Scaling Trends for Lie Detector Oversight in Preference Learning

Deceptive behavior in LLMs is costly to monitor and prevent, motivating approaches such as Scalable Oversight via Lie Detectors (SOLiD) (Cundy & Gleave, 2025), which uses lie detectors to identify responses for review by high-cost labelers. In this paper, we scale SOLiD to larger models and evaluate it in more diverse and realistic preference-learning settings. We find favorable scaling: undetected deception drops from 34% for 1B-parameter models to 14% for 405B-parameter models at a detector true positive rate of 99%, and expensive human labelers can be removed entirely from the fine-tuning phase without a statistically significant increase in deception. However, SOLiD is sensitive to distribution shift between detector training and preference-training data, which can drive detector false positive rates to impractical levels.

cs.AI

Convergence Laws for Extensions of First-Order Logic with Averaging

For many standard models of random structure, first-order logic sentences exhibit a convergence phenomenon on random inputs. The most well-known example is for random graphs with constant edge probability, where the probabilities of first-order sentences converge to 0 or 1. In other cases, such as certain ``sparse random graph'' models, the probabilities of sentences converge, although not necessarily to 0 or 1. In this work we deal with extensions of first-order logic with aggregate operators, variations of averaging. These logics will consist of real-valued terms, and we allow arbitrary Lipschitz functions to be used as ``connectives''. We show that some of the well-known convergence laws extend to this setting.

cs.LO

Neural Interactive Proofs

We consider the problem of how a trusted, but computationally bounded agent (a 'verifier') can learn to interact with one or more powerful but untrusted agents ('provers') in order to solve a given task. More specifically, we study the case in which agents are represented using neural networks and refer to solutions of this problem as neural interactive proofs. First we introduce a unifying framework based on prover-verifier games, which generalises previously proposed interaction protocols. We then describe several new protocols for generating neural interactive proofs, and provide a theoretical comparison of both new and existing approaches. Finally, we support this theory with experiments in two domains: a toy graph isomorphism problem that illustrates the key ideas, and a code validation task using large language models. In so doing, we aim to create a foundation for future work on neural interactive proofs and their application in building safer AI systems.

cs.AI

Almost Surely Asymptotically Constant Graph Neural Networks

We present a new angle on the expressive power of graph neural networks (GNNs) by studying how the predictions of real-valued GNN classifiers, such as those classifying graphs probabilistically, evolve as we apply them on larger graphs drawn from some random graph model. We show that the output converges to a constant function, which upper-bounds what these classifiers can uniformly express. This strong convergence phenomenon applies to a very wide class of GNNs, including state of the art models, with aggregates including mean and the attention-based mechanism of graph transformers. Our results apply to a broad class of random graph models, including sparse and dense variants of the Erd\H{o}s-R\'enyi model, the stochastic block model, and the Barab\'asi-Albert model. We empirically validate these findings, observing that the convergence phenomenon appears not only on random graphs but also on some real-world graphs.

cs.LG

The Intermediate Logic of Convex Polyhedra

We investigate a recent semantics for intermediate (and modal) logics in terms of polyhedra. The main result is a finite axiomatisation of the intermediate logic of the class of all polytopes -- i.e., compact convex polyhedra -- denoted PL. This logic is defined in terms of the Jankov-Fine formulas of two simple frames. Soundness of this axiomatisation requires extracting the geometric constraints imposed on polyhedra by the two formulas, and then using substantial classical results from polyhedral geometry to show that convex polyhedra satisfy those constraints. To establish completeness of the axiomatisation, we first define the notion of the geometric realisation of a frame into a polyhedron. We then show that any PL frame is a p-morphic image of one which has a special form: it is a 'sawed tree'. Any sawed tree has a geometric realisation into a convex polyhedron, which completes the proof.

math.LO

Zero-One Laws of Graph Neural Networks

Graph neural networks (GNNs) are the de facto standard deep learning architectures for machine learning on graphs. This has led to a large body of work analyzing the capabilities and limitations of these models, particularly pertaining to their representation and extrapolation capacity. We offer a novel theoretical perspective on the representation and extrapolation capacity of GNNs, by answering the question: how do GNNs behave as the number of graph nodes become very large? Under mild assumptions, we show that when we draw graphs of increasing size from the Erd\H{o}s-R\'enyi model, the probability that such graphs are mapped to a particular output by a class of GNN classifiers tends to either zero or to one. This class includes the popular graph convolutional network architecture. The result establishes 'zero-one laws' for these GNNs, and analogously to other convergence laws, entails theoretical limitations on their capacity. We empirically verify our results, observing that the theoretical asymptotic limits are evident already on relatively small graphs.

cs.LG

Uniform, rigid branchwise-real trees

A branchwise-real tree is a partial order which is a tree and in which every branch is isomorphic to a real interval. I give constructions of such trees which are both rigid (i.e. without non-trivial order-automorphisms) and uniform (in two different senses). Specifically, I show that there is a rigid branchwise-real tree in which every branching point has the same degree, one in which every point is branching and of the same degree, and finally one in which every point is branching of the same degree and which admits no order-preserving function into the reals. Trees are grown iteratively in stages, and a key technique is the construction (in ZFC) of a family of colourings of $(0,\infty)$ which is 'sufficiently generic', using these colourings to determine how to proceed with the construction.

math.LO

Bisimulations of potentialist systems

A potentialist system is a first-order Kripke model based on embeddings. I define the notion of bisimulation for these systems, and provide a number of examples. Given a first-order theory $T$, the system $\mathrm{Mod}(T)$ consists of all models of $T$. We can then take either all embeddings, or all substructure inclusions, between these models. I show that these two ways of defining $\mathrm{Mod}(T)$ are bitotally bisimilar. Next, I relate the notion of bisimulation to a generalisation of the Ehrenfeucht-Fra\"is\'e game, and use this to show the equivalence of the existence of a bisimulation with elementary equivalence with respect to an infinitary language. Finally, I consider the question of when a potentialist system is bitotally bisimilar to a system containing set-many models, providing too different sufficient conditions.

math.LO

Polyhedral completeness of intermediate logics: the Nerve Criterion

We investigate a recently-devised polyhedral semantics for intermediate logics, in which formulas are interpreted in n-dimensional polyhedra. An intermediate logic is polyhedrally complete if it is complete with respect to some class of polyhedra. The first main result of this paper is a necessary and sufficient condition for the polyhedral-completeness of a logic. This condition, which we call the Nerve Criterion, is expressed in terms of Alexandrov's notion of the nerve of a poset. It affords a purely combinatorial characterisation of polyhedrally-complete logics. Using the Nerve Criterion we show, easily, that there are continuum many intermediate logics that are not polyhedrally-complete but which have the finite model property. We also provide, at considerable combinatorial labour, a countably infinite class of logics axiomatised by the Jankov-Fine formulas of 'starlike trees' all of which are polyhedrally-complete. The polyhedral completeness theorem for these 'starlike logics' is the second main result of this paper.

math.LO

On the continuous gradability of the cut-point orders of $\mathbb R$-trees

An $\mathbb R$-tree is a certain kind of metric space tree in which every point can be branching. Favre and Jonsson posed the following problem in 2004: can the class of orders underlying $\mathbb R$-trees be characterised by the fact that every branch is order-isomorphic to a real interval? In the first part, I answer this question in the negative: there is a 'branchwise-real tree order' which is not 'continuously gradable'. In the second part, I show that a branchwise-real tree order is continuously gradable if and only if every well-stratified subtree is $\mathbb R$-gradable. This link with set theory is put to work in the third part answering refinements of the main question, yielding several independence results. For example, when $\kappa \geq \mathfrak c$, there is a branchwise-real tree order which is not continuously gradable, and which satisfies a property corresponding to $\kappa$-separability. Conversely, under Martin's Axiom at $\kappa$ such a tree does not exist.

math.LO