arXiv ScienceSearch

arXiv subjects

Carmen Amme

Publications and source records attributed to Carmen Amme.

2 recordsLinked to original sources

Active Visual Semantics: A large-scale MEG and eye-tracking dataset for understanding visual intelligence in action

Here we present the Active Visual Semantics (AVS) dataset, a large-scale collection of magnetoencephalography (MEG) and eye-tracking data recorded while five participants freely explored 4,080 natural scenes over 10 sessions each, yielding more than 200,000 fixation epochs in total. Unlike existing neuroimaging datasets that rely on passive viewing with enforced central fixation, AVS captures brain activity during active scene exploration, including self-generated saccades and fixations. A semantic captioning task on 25% of the trials provides behavioural measures linking gaze to scene understanding and memory. In addition to neural and behavioural data, AVS includes per-fixation object category labels, human-rated annotations of the appearance of fixation targets in the scene captioning task and pupil dynamics. Individual head stabilisation casts were used during MEG data collection, which alongside with structural MRI scans, enabled precise cross-session source reconstruction. Using artificial neural network (ANN) encoding models we demonstrate that individual fixation-aligned MEG epochs hold visual content-specific signal, despite the challenges that active scene viewing poses for MEG signal quality. Further, we use fixation-aligned representation similarity analysis (RSA) and demonstrate that we can derive fixation object category averages that yield representational geometries which are highly reliable across participants. Both in MEG sensor and source space this structure is validated by its robust alignment with ANN object-level representational geometry. Taken together, AVS provides a rich resource for investigating a large variety of questions regarding the neural mechanisms of active vision, object recognition and scene captioning during natural viewing, and the relationship between gaze behaviour and memory encoding.

q-bio.NC

Predicting upcoming visual features during eye movements yields scene representations aligned with human visual cortex

Natural scenes are complex arrangements of objects, surfaces, and backgrounds. For the brain's visual system to effectively operate, it needs to extract not only what objects are present, but also their spatial and semantic relations. We hypothesize that such structures may be learned, in a self-supervised fashion, by exploiting temporal regularities of natural active vision: each fixation reveals a glimpse that is related to the previous one via co-occurrence and saccade-conditioned spatial regularities. We instantiate this idea with Glimpse Prediction Networks (GPNs), recurrent models trained to predict the embedding of the next glimpse along human-like scanpaths. GPNs are shown to successfully extract complex scene information, including object co-occurrences and spatial object arrangements, and integrate information across glimpses. Importantly, GPN representations align strongly with human fMRI responses in mid and higher-level visual cortex and match, often outperform, alternative state-of-the-art ANN models, establishing next-glimpse-prediction as a biologically plausible route towards brain-aligned scene representations.

q-bio.NC