arXiv Science⌕ Search

arXiv · 2610.09287

Scaling subjects in cross-modal alignment: video decoding with EEG foundation model

Abstract

Naturalistic visual decoding from EEG has long been constrained by cohort size: existing scaling literature caps out at a few dozen subjects, leading prior work to conclude that scaling subject cohorts yields minimal performance gains. Training a cross-modal EEG--video contrastive encoder on a cohort larger by more than an order of magnitude, we find the axis is productive but has an onset. Below $S\approx50$ --- the entirety of the range prior work occupies --- no model improves meaningfully over an untrained encoder; above it, decoding rises log-linearly in subject count with no saturation at the top of our ladder, on a fitted stimulus-feature probe and on fit-free movie-moment retrieval alike, and the gain transfers to a second, unseen film. How far the axis carries then depends on initialisation more than on capacity: an encoder initialised from an EEG foundation model scales ${\sim}1.6\times$ faster per doubling of the cohort than randomly initialised encoders at two depths, is the only one still converting subjects into retrieval accuracy at the top of the ladder, and converges on a fraction of the alignment compute. Subject scaling on shared naturalistic stimuli is thus an effective and currently unsaturated frontier for scaling EEG foundation models.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Dung Truong, Kuntal Kokate, Arnaud Delorme. 2026-10-07. Scaling subjects in cross-modal alignment: video decoding with EEG foundation model. https://arxiv.org/abs/2610.09287

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay

Understanding how humans and artificial intelligence systems predict and plan by interacting with their environment is a fundamental challenge at the intersection of neuroscience and machine learning. Most brain-encoding studies focus on aligning artificial models with brain activity during language comprehension or passive visual processing, while interactive brain alignment studies have to date been largely limited to reinforcement-learning (RL) agents and theory-based models. To address this gap, we study brain alignment of representative models from two foundation-model types, namely vision-language models (VLMs) and large-action models (LAMs), using fMRI recordings from participants playing naturalistic Atari-style video games. Specifically, we examine how action-focused and reasoning-focused prompts shape the models' internal representations and their alignment with fMRI brain activity. First, we find that both VLMs and LAMs achieve significantly higher voxel-wise encoding performance than RL baselines, with the advantage holding even under matched feature dimensionality. Second, compared to a no-prompt baseline, prompt-driven gains are larger in higher-order frontal-parietal and motor-planning regions than in early visual cortex, roughly 2-2.5$\times$ when averaged over region-of-interest (ROI) groups, although individual regions are heterogeneous. Third, variance partitioning reveals a qualitatively different representational organization. VLM representations are prompt-symmetric (12.4% unique action vs. 9.5% unique reasoning), whereas LAM representations are action-dominant (25.6% unique action vs.-8.2% unique reasoning), with the asymmetry strongest in frontal-motor cortex. Together, these results associate action specialization with distinct cortical alignment patterns in multimodal game-state representations, revealing differences hidden by similar prediction accuracy.

q-bio.NC↗

Connectome-Based Modeling of Mutation-Specific Amyloid-$β$ Aggregation in Familial Alzheimer's Disease

Familial amyloid-$β$ (A$β$) variants alter aggregation kinetics, but their interaction with structural brain connectivity remains incompletely understood. We developed a mutation-aware mechanistic model coupling a coarse-grained monomer--oligomer--fibril aggregation--fragmentation system to graph diffusion on the 540-node Budapest Reference Connectome component. Experimental A$β_{42}$ nucleation scores scaled primary nucleation rates for seven variants relative to wild type. Robustness was examined using mutation-score uncertainty, alternative kinetic mappings, global sensitivity analysis, seed and edge-weight perturbations, degree-preserving randomized connectomes, spatial propagation analysis, synthetic ABC-SMC parameter recovery, posterior prediction, and Chemical Langevin simulations. E22G showed the earliest threshold crossing and greatest cumulative oligomer burden, whereas A2V was delayed under the selected mapping. Mutation rankings persisted across tested network perturbations, although regional burden patterns depended on topology. Connectome distance from seed regions was associated with later oligomer arrival (Spearman $ρ\approx0.92$). Under inferred parameter uncertainty, the mean timing-rank correlation was 0.990 and cumulative-burden ordering was preserved in every posterior draw, while exact peak-amplitude ordering was less stable. Stochastic ensemble medians retained the deterministic ordering despite overlap among trajectories. Within this proof-of-concept framework, mutation-dependent kinetics primarily influence aggregation timing and cumulative burden, while connectivity shapes spatial propagation. Nominal background rates, model-time units, and synthetic parameter recovery limit interpretation to mechanistic comparisons rather than clinically calibrated prediction.

q-bio.NC↗

Feedback to the primary visual cortex is highly concentrated on the central visual field representation

Retrograde tracer injections in marmoset monkeys revealed a striking feature of neural circuitry in primate visual system: among neurons projecting to the primary visual cortex (striate cortex, V1), the proportion of extrastriate feedback neurons relative to thalamic feedforward neurons decreases markedly with visual field eccentricity. This supports the Central-Peripheral Dichotomy theory, which proposes fundamentally distinct computational specializations --- visual recognition and attentional selection, respectively --- in the central and peripheral visual fields.

q-bio.NC↗