arXiv ScienceSearch

arXiv subjects

William Bialek

Publications and source records attributed to William Bialek.

At least 19 recordsLinked to original sources

Large language models and the entropy of English

We use large language models (LLMs) to uncover long-ranged structure in English texts from a variety of sources. The conditional entropy or code length in many cases continues to decrease with context length at least to $N\sim 10^4$ characters, implying that there are direct dependencies or interactions across these distances. A corollary is that there are small but significant correlations between characters at these separations, as we show from the data independent of models. The distribution of code lengths reveals an emergent certainty about an increasing fraction of characters at large $N$. Over the course of model training, we observe different dynamics at long and short context lengths, suggesting that long-ranged structure is learned only gradually. Our results constrain efforts to build statistical physics models of LLMs or language itself.

cond-mat.stat-mech

Ambiguous signals and efficient codes

In many biological networks the responses of individual elements are ambiguous. We consider a scenario in which many sensors respond to a shared signal, each with limited information capacity, and ask that the outputs together convey as much information as possible about an underlying relevant variable. In a low noise limit where we can make analytic progress, we show that individually ambiguous responses optimize overall information transmission.

physics.bio-ph

When many noisy genes optimize information flow

It often is emphasized that gene expression is noisy. A seemingly contradictory view is that control mechanisms have been optimized to squeeze as much information as possible out of a limited number of molecules. Here we revisit these issues in a simple model where a single transcription factor (TF) controls a large number of target genes. We include only the physically required noise sources: random arrival of TFs at their targets and counting noise in the synthesis and degradation of mRNA. If the cell has a limited total number of mRNA molecules, then the capacity to transmit information about TF concentration is maximized when these resources are distributed across the largest possible number of target genes. To realize this capacity the distribution of TF concentrations must be biased toward smaller values. Thus, in some limits, information transmission is optimized when individual expression levels are noisy. In addition, the dependence of information transmission on the parameters of this multi-gene system has a "sloppy" spectrum, so that optimal performance can co-exist with substantial variability.

physics.bio-ph

Context dependent adaptation in a neural computation

Brains adapt to the statistical structure of their input. In the visual system, local light intensities change rapidly, the variance of the intensity changes more slowly, and the dynamic range of contrast itself changes more slowly still. We use a motion-sensitive neuron in the fly visual system to probe this hierarchy of adaptation phenomena, delivering naturalistic stimuli that have been simplified to have a clear separation of time scales. We show that the neural response to visual motion depends on contrast, and this dependence itself varies with context. Using the spike-triggered average velocity trajectory as a response measure, we find that context dependence is confined to a low-dimensional space, with a single dominant dimension. Across a wide range of conditions this adaptation serves to match the integration time to the mean interval between spikes, reducing redundancy.

q-bio.NC

Neural subspaces, minimax entropy, and mean-field theory for networks of neurons

Recent advances in experimental techniques enable the simultaneous recording of activity from thousands of neurons in the brain, presenting both an opportunity and a challenge: to build meaningful, scalable models of large neural populations. Correlations in the brain are typically weak but widespread, suggesting that a mean-field approach might be effective in describing real neural populations, and we explore a hierarchy of maximum entropy models guided by this idea. We begin with models that match only the mean and variance of the total population activity, and extend to models that match the experimentally observed mean and variance of activity along multiple projections of the neural state. Confronted by data from several different brain regions, these models are driven toward a first-order phase transition, characterized by the presence of two nearly degenerate minima in the energy landscape, and this leads to predictions in qualitative disagreement with other features of the data. To resolve this problem we introduce a novel class of models that constrain the full probability distribution of activity along selected projections. We develop the mean-field theory for this class of models and apply it to recordings from 1000+ neurons in the mouse hippocampus. This 'distributional mean--field' model provides an accurate and consistent description of the data, offering a scalable and principled approach to modeling complex neural population dynamics.

physics.bio-ph

Optimization and variability can coexist

Many biological systems perform close to their physical limits, but promoting this optimality to a general principle seems to require implausibly fine tuning of parameters. Using examples from a wide range of systems, we show that this intuition is wrong. Near an optimum, functional performance depends on parameters in a "sloppy'' way, with some combinations of parameters being only weakly constrained. Absent any other constraints, this predicts that we should observe widely varying parameters, and we make this precise: the entropy in parameter space can be extensive even if performance on average is very close to optimal. This removes a major objection to optimization as a general principle, and rationalizes the observed variability.

q-bio.QM

How different are self and nonself?

Biological and artificial networks routinely make reliable distinctions between similar inputs, and the rules for making these distinctions are learned. In some ways, self/nonself discrimination in the immune system is similar, being both reliable and (partly) learned through thymic selection. In contrast to other examples, we show that the distributions of self and nonself peptides are nearly identical but strongly inhomogeneous. Reliable discrimination is possible only because self-peptides are a particular finite sample drawn out of this distribution, and T cells can target the spaces in between these samples. In conventional learning problems, this would constitute overfitting and lead to disaster. Here, the strong inhomogeneities imply instead that the immune system gains by targeting peptides which are similar to self, with maximum sensitivity for sequences just one or two substitutions away. This prediction from the structure of the underlying distribution in sequence space agrees, for example, with the observed responses to mutation derived cancer neoantigens.

q-bio.CB

Extended mean-field theories for networks of real neurons

If the behavior of a system with many degrees of freedom can be captured by a small number of collective variables, then plausibly there is an underlying mean-field theory. We show that simple versions of this idea fail to describe the patterns of activity in networks of real neurons. An extended mean-field theory that matches the distribution of collective variables is at least consistent, though shows signs that these networks are poised near a critical point, in agreement with other observations. These results suggest a path to analysis of emerging data on ever larger numbers of neurons.

physics.bio-ph

The FlEye camera: Sampling the joint distribution of natural scenes and motion

To make efficient use of limited physical resources, the brain must match its coding and computational strategies to the statistical structure of input signals. An attractive testing ground for these principles is the problem of motion estimation in the fly visual system: we understand the optics of the compound eye, have a quantitative description of input signals and noise from the retina, and can record from output neurons that encode estimates of different velocity components. Furthermore, recent work provides a nearly complete wiring diagram of the intervening circuitry. What is missing is a characterization of the visual signals and motions that flies encounter in a natural context. We attack this directly with the development of a specialized camera that matches the high temporal resolution, optical properties, and spectral sensitivity of the fly's eye; inertial motion sensors provide ground truth about rotations and translations through the world. We describe the design, construction, and performance characteristics of this FlEye camera. To illustrate the opportunities created by this instrument we use data on movies and motion to construct optimal local motion estimators that can be compared with the responses of the fly's motion sensitive neurons.

q-bio.NC

Direct estimates of irreversibility from time series

The arrow of time can be quantified through the Kullback-Leibler divergence ($D_{KL}$) between the distributions of forward and reverse trajectories in a system. Many approaches to estimate this rely on specific models, but the use of incorrect models can introduce uncontrolled errors. Here, we describe a model-free method that uses trajectory data directly to estimate the evidence for irreversibility over finite windows of time. To do this we build on previous work to identify and correct for errors that arise from limited sample size. Importantly, our approach accurately recovers $D_{KL} = 0$ in systems that adhere to detailed balance, and the correct nonzero $D_{KL}$ for data generated by well understood models of nonequilibrium systems. We apply our method to trajectories of neural activity in the retina as it responds to naturalistic inputs, and find evidence of irreversibility in single neurons, emphasizing the non-Markovian character of these data. These results open new avenues for investigating how the brain represents the arrow of time.

cond-mat.stat-mech

Moving boundaries: An appreciation of John Hopfield

The 2024 Nobel Prize in Physics was awarded to John Hopfield and Geoffrey Hinton, "for foundational discoveries and inventions that enable machine learning with artificial neural networks." As noted by the Nobel committee, their work moved the boundaries of physics. This is a brief reflection on Hopfield's work, its implications for the emergence of biological physics as a part of physics, the path from his early papers to the modern revolution in artificial intelligence, and prospects for the future.

physics.hist-ph

What makes it possible to learn probability distributions in the natural world?

Organisms and algorithms learn probability distributions from previous observations, either over evolutionary time or on the fly. In the absence of regularities, estimating the underlying distribution from data would require observing each possible outcome many times. Here we show that two conditions allow us to escape this infeasible requirement. First, the mutual information between two halves of the system should be consistently sub-extensive. Second, this shared information should be compressible, so that it can be represented by a number of bits proportional to the information rather than to the entropy. Under these conditions, a distribution can be described with a number of parameters that grows linearly with system size. These conditions are borne out in natural images and in models from statistical physics, respectively.

cond-mat.stat-mech

Probabilistic models, compressible interactions, and neural coding

In physics we often use very simple models to describe systems with many degrees of freedom, but it is not clear why or how this success can be transferred to the more complex biological context. We consider models for the joint distribution of many variables, as with the combinations of spiking and silence in large networks of neurons. In this probabilistic framework, we argue that simple models are possible if the mutual information between two halves of the system is consistently sub--extensive, and if this shared information is compressible. These conditions are not met generically, but they are met by real world data such as natural images and the activity in a population of retinal output neurons. We introduce compression strategies that combine the information bottleneck with an iteration scheme inspired by the renormalization group, and find that the number of parameters needed to describe the distribution of joint activity scales with the square of the number of neurons, even though the interactions are not well approximated as pairwise. Our results also show that this shared information is essentially equal to the information that individual neurons carry about natural visual inputs, which has surprising implications for the neural code.

q-bio.NC

Statistical mechanics for networks of real neurons

Perceptions and actions, thoughts and memories result from coordinated activity in hundreds or even thousands of neurons in the brain. It is an old dream of the physics community to provide a statistical mechanics description for these and other emergent phenomena of life. These aspirations appear in a new light because of developments in our ability to measure the electrical activity of the brain, sampling thousands of individual neurons simultaneously over hours or days. We review the progress that has been made in bringing theory and experiment together, focusing on maximum entropy methods and a phenomenological renormalization group. These approaches have uncovered new, quantitatively reproducible collective behaviors in networks of real neurons, and provide examples of rich parameter--free predictions that agree in detail with experiment.

cond-mat.dis-nn

Searching for long time scales without fine tuning

Most of animal and human behavior occurs on time scales much longer than the response times of individual neurons. In many cases, it is plausible that these long time scales emerge from the recurrent dynamics of electrical activity in networks of neurons. In linear models, time scales are set by the eigenvalues of a dynamical matrix whose elements measure the strengths of synaptic connections between neurons. It is not clear to what extent these matrix elements need to be tuned in order to generate long time scales; in some cases, one needs not just a single long time scale but a whole range. Starting from the simplest case of random symmetric connections, we combine maximum entropy and random matrix theory methods to construct ensembles of networks, exploring the constraints required for long time scales to become generic. We argue that a single long time scale can emerge generically from realistic constraints, but a full spectrum of slow modes requires more tuning. Langevin dynamics that generates patterns of synaptic connections drawn from these ensembles involves a combination of Hebbian learning and activity-dependent synaptic scaling.

physics.bio-ph

Maximum entropy models for patterns of gene expression

New experimental methods make it possible to measure the expression levels of many genes, simultaneously, in snapshots from thousands or even millions of individual cells. Current approaches to analyze these experiments involve clustering or low-dimensional projections. Here we use the principle of maximum entropy to obtain a probabilistic description that captures the observed presence or absence of mRNAs from hundreds of genes in cells from the mammalian brain. We construct the Ising model compatible with experimental means and pairwise correlations, and validate it by showing that it gives good predictions for higher-order statistics. We notice that the probability distribution of cell states has many local maxima. By labeling cell states according to the associated maximum, we obtain a cell classification that agrees well with previous results that use traditional clustering techniques. Our results provide quantitative descriptions of gene expression statistics and interpretable criteria for defining cell classes, supporting the hypothesis that cell classes emerge from the collective interaction of gene expression levels.

physics.bio-ph

A brief tutorial on information theory

At the 2023 Les Houches Summer School on Theoretical Biological Physics, several students asked for some background on information theory, and so we added a tutorial to the scheduled lectures. This is largely a transcript of that tutorial, lightly edited. It covers basic definitions and context rather than detailed calculations. We hope to have maintained the informality of the presentation, including exchanges with the students, while still being useful.

physics.bio-ph

Ambitions for theory in the physics of life

Theoretical physicists have been fascinated by the phenomena of life for more than a century. As we engage with more realistic descriptions of living systems, however, things get complicated. After reviewing different reactions to this complexity, I explore the optimization of information flow as a potentially general theoretical principle. The primary example is a genetic network guiding development of the fly embryo, but each idea also is illustrated by examples from neural systems. In each case, optimization makes detailed, largely parameter-free predictions that connect quantitatively with experiment

physics.bio-ph