arXiv Science⌕ Search

arXiv · 2609.36366

Cross-attention encoding models reveal dynamic spatiotemporal routing across human higher visual cortex

Abstract

Understanding how the brain parses actions and events from time-varying natural inputs is a central challenge in neuroscience. Recent work has used deep neural network (DNN) models to build stimulus-computable fMRI encoding models that predict single-voxel responses to complex natural videos. However, the majority of video-computable encoding models predict responses using simple linear mappings from model tokens, overlooking the spatiotemporal structure shared by video representations and neural responses. Recent cross-attention encoding models address this limitation for static images, enabling flexible stimulus-dependent weighting of image content across space. Here, we extend this framework to naturalistic video, using per-parcel cross-attention to dynamically route features from a self-supervised video model (V-JEPA-2) across both space and time, fitting this model to fMRI responses to short video clips. We compare joint spatiotemporal attention with factorized and selectively constrained alternatives, and find that joint routing improves predictions of brain responses to held-out videos across higher visual regions, most consistently in lateral and dorsal visual areas associated with dynamic motion perception. Moreover, our method provides interpretable, stimulus-specific attention maps that dynamically follow moving objects, revealing which locations and temporal moments contribute to each neural response. We further show that attention maps from parcels in different category-selective networks (face-, body-, scene-selective) differentially weight content in accordance with expected semantic selectivity. Together, this work provides a new computational framework for understanding how visual information is adaptively weighted by cortical populations during dynamic visual perception.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Iishaan Inabathini, Margaret M. Henderson. 2026-09-28. Cross-attention encoding models reveal dynamic spatiotemporal routing across human higher visual cortex. https://arxiv.org/abs/2609.36366

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Neural spikes as rare events

We consider the information transmission problem in neurons and its possible implications for learning in neural networks. Our approach is based on recent developments in statistical physics and complexity science. Combining sensory information from various modalities for perceptual decision-making offers several advantages and is essential for the survival of both humans and animals. Not much is known about which brain regions are involved in spatial localization using audiovisual integration. We explore this further by training mice in a task requiring audiovisual integration. We then record from the secondary motor cortex (M2) using high-density electrophysiology. Analyzing this data, we found neurons responsive to multimodal as well as unimodal auditory and visual stimuli. The neurons are generally more responsive to auditory, rather than visual, stimuli. There was low correlation between the auditory and visual responses. Some neurons were sensitive to the task mode, whether active or passive, with more neurons being responsive in the active mode. A relatively large percentage of neurons (10-11%) differed significantly in their response to left and right-sided auditory stimuli, but only in 1 of the 3 mice we recorded from. These findings suggest a role for M2 in multisensory decision making and should enable further research in this field. We then use branching process simulations to model neural activity. This would support temporal coding theory as a model for neural coding.

q-bio.NC↗

Association profile conditioning in a set-temporal transformer for cross-session intracortical motor decoding

Intracortical motor decoders degrade across sessions because the set of recorded units changes and persisting units can alter how their firing relates to behavior. Most existing methods update network weights on each new session or rely on unlabeled activity, which does not directly reveal such changes. We present APST, an Association Profile-conditioned Set-Temporal transformer that adapts to new sessions with all network weights frozen. From a few labeled calibration trials, APST summarizes how each unit's firing relates to behavior in a four-dimensional association profile computed in closed form. The profiles condition a set-attention encoder that accepts any number and order of units, followed by a causal transformer for streaming decoding. On held-out DANDI688 sessions from two monkeys, APST reaches velocity $R^2$ of $0.78$ and $0.81$, versus $0.40$ and $0.58$ for a variant that uses neural activity alone, and matches or exceeds an RNN fine-tuned on the same trials. On FALCON private held-out evaluation, it attains $R^2$ of $0.65$, $0.42$, and $0.44$ on M1, M2, and H1.

q-bio.NC↗

BrainWave: A Brain Signal Foundation Model for Clinical Applications

Neural electrical activity is fundamental to brain function, underlying a range of cognitive and behavioral processes, including movement, perception, decision-making, and consciousness. Abnormal patterns of neural signaling often indicate the presence of underlying brain diseases. The variability among individuals, the diverse array of clinical symptoms from various brain disorders, and the limited availability of diagnostic classifications, have posed significant barriers to formulating reliable model of neural signals for diverse application contexts. Here, we present BrainWave, the first foundation model for both invasive and non-invasive neural recordings, pretrained on more than 40,000 hours of electrical brain recordings (13.79 TB of data) from approximately 16,000 individuals. Our analysis show that BrainWave outperforms all other competing models and consistently achieves state-of-the-art performance in the diagnosis and identification of neurological disorders. We also demonstrate robust capabilities of BrainWave in enabling zero-shot transfer learning across varying recording conditions and brain diseases, as well as few-shot classification without fine-tuning, suggesting that BrainWave learns highly generalizable representations of neural signals. We hence believe that open-sourcing BrainWave will facilitate a wide range of clinical applications in medicine, paving the way for AI-driven approaches to investigate brain disorders and advance neuroscience research.

q-bio.NC↗