arXiv Science⌕ Search

arXiv · 2609.36739

Frontier Autolab: Organizational Memory, Adversarial Dissent and Temporal Leakage in Multi-Agent LLM Firms Across Fifty Years of Technological Change

Abstract

Multi-agent LLM systems are increasingly structured like organizations, with roles, critics and shared memory, yet they are evaluated on tasks that last minutes. We ask how such an organization behaves when the ground it stands on keeps moving. Frontier Autolab is a long-horizon testbed in which one simulated firm, voiced by sixteen role personas and a dedicated Red Team, must re-found itself in nine technology eras from 1990 to 2040. Each era is temporally gated: the firm decides from a dated briefing, a historian-judge then reveals what happened and scores the decision on a five-dimension rubric, and lessons enter a persistent Playbook. Six eras are scored against history, one against the live market and two are open forecasts. Across four trajectories (36 era decisions, 180 subscores) we find a consistent foresight-commitment gap: in all 24 historically scored eras the judge rated the firm's recognition of the coming shift above its choice of where to build (mean gap 1.9 points on a 10-point scale), because boards chose the layer their existing assets could reach. Organizational design shaped long-run character. A Red Team armed with numeric kill gates produced fifty years of gated pilots and no product, and the rubric rated this firm highest; firms whose memory stored market-structure lessons pivoted every era, while a firm whose memory stored only validation procedure kept one method throughout. We also show why such results are hard to trust. Scores rise across eras in every run while the judge's own hindsight subscore falls (within-run r = -0.58), so apparent learning is confounded with recall of history, and we trace further distortions to self-judging, briefing selection and score aggregation. We release all records and an API harness, and specify fictional and post-cutoff eras that would turn the testbed into a benchmark.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Bravish Ghosh. 2026-09-29. Frontier Autolab: Organizational Memory, Adversarial Dissent and Temporal Leakage in Multi-Agent LLM Firms Across Fifty Years of Technological Change. https://arxiv.org/abs/2609.36739

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Trusting the Inverse: Reliability-Aware Mapping for Simulation-Based Microstructure Estimation in Diffusion MRI

Diffusion-weighted MRI can probe tissue microstructure non-invasively, however interpreting microstructural parameter estimates remains challenging due to the intrinsic ambiguity of the inverse problem. While simulation-based approaches can incorporate increasingly realistic tissue models, they still lack voxel-wise characterisation of the reliability and degeneracy of the inferred parameters. Here, we introduce a reliability framework for simulation-based microstructure estimation based on three complementary scores that identify distinct sources of unreliability in the estimation process: out-of-distribution signals, local signal mismatch, and parameter degeneracy. The framework was implemented using a Monte Carlo dictionary of 1,050 synthetic voxels generated from geometrically realistic substrates, with parameter ranges grounded in electron microscopy measurements of rat corpus callosum. The dictionary spans biologically plausible axon radii (0.25-0.85 $μ$m), microscopic angular spread (0-10$^\circ$), packing densities (60-92%), and intrinsic diffusivities (1.75-3.0 $μ$m$^2$/ms). On synthetic data, the resulting Reliability Index correlated with actual estimation error (Spearman $ρ= -0.742$) and distinguished between extrapolation, poor local interpolation, and parameter degeneracy. Applied to in vivo corpus callosum DW-MRI in rat (256 voxels, four animals) and human (MGH-USC HCP, 18,765 voxels, nine subjects), 91% and 73% of voxels respectively exceeded $R > 0.5$. These results demonstrate how reliability-aware analysis can support the interpretation and future development of simulation-based diffusion MRI microstructure models.

cs.CE↗

ScentGen: Hierarchical Multimodal Olfactory Semantic Modeling for Molecular Odor Description Generation

In this paper, we introduce a molecular odor description generation task, which aims to generate natural language odor descriptions from molecular structures. Unlike conventional methods that describe molecular odor using discrete labels, this task generates expressive and human-interpretable sensory descriptions. To address this task, we propose a hierarchical multimodal olfactory semantic modeling framework, named ScentGen. ScentGen consists of three key components: an odor semantic planner, a semantic adapter, and a description generator. The odor semantic planner integrates complementary molecular information from 1D SMILES sequences, 2D molecular graphs, and 3D molecular conformations to learn discriminative and structured olfactory semantics. The semantic adapter further maps the learned olfactory representation into the hidden space of a large language model, transforming molecular odor semantics into language-compatible continuous prompts. Conditioned on these prompts, the description generator produces coherent odor descriptions that reflect plausible sensory characteristics of the input molecule. Considering the lack of molecular datasets with natural language odor descriptions, we further construct a molecular odor description dataset containing paired multimodal molecular representations and human-interpretable odor descriptions. Extensive experiments demonstrate that ScentGen generates coherent and expressive odor descriptions, providing a more flexible solution for molecular odor understanding beyond discrete odor label prediction.

cs.CE↗

SymbolicLM: Training Language Models as Symbolic Regressors

Large Language Models (LLMs) have shown promising capabilities in scientific reasoning, yet scientific discovery ultimately requires deriving precise laws directly from observational data, known as Symbolic Regression (SR). This poses a challenge for LLMs due to the gap between probabilistic text generation and the exact structural requirements of SR. Existing approaches rely on complex external scaffolds, which are computationally expensive and separate symbolic reasoning from the model itself. To address this limitation, we propose to directly equip LLMs with symbolic regression capabilities through dedicated numerical-symbolic and physical supervision. We introduce PhysSymbArena, a large-scale benchmark containing over 160,000 equations and 1.8B tokens of numerical-symbolic data with physical descriptions, enabling systematic training and evaluation. Based on PhysSymbArena, we develop SymbolicLM, which enhances the symbolic regression ability of LLMs through mathematical and physical supervision. During inference, we further introduce SymbolicSGA, a refinement framework that leverages quantitative feedback to iteratively improve generated equations. Experiments on multiple symbolic regression benchmarks show that SymbolicLM substantially improves structural recovery while maintaining competitive numerical fitting performance. These results demonstrate that symbolic regression can be explicitly learned as an intrinsic capability of LLMs.

cs.CE↗