arXiv ScienceSearch

arXiv · 2605.10721

Conformity Generates Collective Misalignment in AI Agents Societies

Abstract

Artificial intelligence safety research focuses on aligning individual language models with human values, yet deployed AI systems increasingly operate as interacting populations where social influence may override individual alignment. Here we show that populations of individually aligned AI agents can be driven into stable misaligned states through conformity dynamics. Simulating opinion dynamics across nine large language models and one hundred opinion pairs, we find that each agent's behavior is governed by two competing forces: a tendency to follow the majority and an intrinsic bias toward specific positions. Using tools from statistical physics, we derive a quantitative theory that predicts when populations become trapped in long-lived misaligned configurations, and identifies predictable tipping points where small numbers of adversarial agents can irreversibly shift population-level alignment even after manipulation ceases. These results demonstrate that individual-level alignment provides no guarantee of collective safety, calling for evaluation frameworks that account for emergent behavior in AI populations.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Giordano De Marzo, Alessandro Bellina, Claudio Castellano, Viola Priesemann, David Garcia. 2026-05-11. Conformity Generates Collective Misalignment in AI Agents Societies. https://arxiv.org/abs/2605.10721

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The Effective Gravitational Field of the Ball in Association Football

In association football, collective player motion is organized around the ball. We ask whether this many-agent motion can be described by an effective field analogous to gravitational attraction, emphasizing that the measured quantity is a radial drift velocity, as in overdamped dynamics, rather than a Newtonian acceleration. Using public tracking data from ten matches, we characterize this radial drift velocity through the distance-dependent mean D(r) and the ball-position-dependent field K(x,y). After excluding restarts and other constrained phases of play, both goalkeepers, the player nearest to the ball, and small player-ball separations, a persistent inward drift velocity remains, with a global mean of about 0.9 m/s. The distance dependence is not described by an inverse power of r: D(r) decreases at short range and then forms a broad plateau that persists to the largest measured separations. The mean K(x,y) is positive over most of the pitch, with a central depression about 15 percent below the non-goalmouth average and stronger suppression near the goalmouths, where the mean drift velocity becomes slightly negative, corresponding to weak effective repulsion. The role-resolved D(r) shows a pronounced defender minimum near 12 m and a pronounced attacker maximum near 60 m, whereas only defenders produce effective repulsion near the goals in K(x,y). These results show that football tracking data can reveal simple effective laws of active many-body motion while retaining clear signatures of player role and pitch geometry.

physics.soc-ph

Hyperscaling of spatial fluctuations constrains the development of urban populations

Urban populations exhibit fractal organization and systematic scaling regularities, yet the scaling exponents reported across cities vary substantially, challenging existing theory. Using 100~m gridded population maps for 109 urban regions in the Netherlands (2000--2023) and 368 major world cities (1975--2020), we recursively coarse-grain each city and quantify how the mean and variance of inhabitants in square grid cells of side length $\ell$ scale with $\ell$. This yields two exponents, $β$ from $\langle N_\ell\rangle\sim \ell^β$ and $γ$ from $\mathrm{Var}(N_\ell)\sim \ell^γ$, where in the small-$\ell$ limit $β$ equals the planar fractal dimension of populated space. Across cities within a given year, $γ$ depends linearly on $β$. Compiling $>$10,000 exponent estimates over five decades shows that this hyperscaling relation is robust yet non-universal: its slope and intercept vary across continents and drift systematically in time, trending toward the limiting form $γ\simeq 2+β$. A mean-field (independent-cell) argument predicts a quadratic mean--variance mapping and cannot reproduce the observed $β$--$γ$ dependence, implying strong spatial correlations. We derive a correlation-aware variance decomposition in which $γ$ is controlled by a correlation dimension $D_c$; in the correlation-dominated regime $γ=2+D_c$. If large maturing cities, as are the ones selected in our dataset, evolve to effective monofractal ($D_c\simeq β$) cities, the asymptotic prediction becomes $γ\simeq 2+β$, consistent with the observed temporal drift. This interdependence links urban geometry and fluctuations, provides an empirical constraint for mechanistic models of urban growth, and implies scaling predictions for spatial indicators built from local means and variances.

physics.soc-ph

Dynamic probabilistic decision networks

A new type of decision networks is suggested and its operation is analyzed. The network nodes are represented by intelligent agents who can denote either some biological beings, like humans, or neurons of the brain, or the nodes of artificial intelligence. The specifics of the network are in the following: It is probabilistic in the sense that the choice, accomplished by each agent, is characterized by the related probability. It is dynamic, with the probabilities varying in time due to the exchange of information between the agents. It is affective, because the agents choose between alternatives by taking account of utility as well as of biases and emotions. In general, it is heterogeneous, being composed of the groups of agents with different properties, for instance having long-term memory and short-term memory. The network dynamics, caused by the information exchange, results in decision error decrease. The network operation is illustrated by the example starting with the Allais paradox, its resolution, and the decision error diminution in the process of decision dynamics with information exchange. Resorting to machine-learning techniques it is possible to regulate the behavior of the network agents forcing them to choose particular alternatives.

physics.soc-ph