Uploaded October 2021 | Updated September 2026, 1 week ago
Will Xiao, Harvard University
Abstract: How does the brain support our ability to see? Studies of primate vision have typically focused on controlled viewing conditions exemplified by the rapid serial visual presentation (RSVP) task, where the subject must hold fixation while images are flashed briefly in randomized order. In contrast, during natural viewing, eyes move frequently, guided by subject-initiated saccades, resulting in a sequence of related sensory input. Thus, natural viewing departs from traditional assumptions of independent and unpredictable visual inputs, leaving it an open question how visual neurons respond in real life.
We recorded responses of interior temporal (IT) cortex neurons in macaque monkeys freely viewing natural images. We first examined responses of face-selective neurons and found that face neurons responded according to whether individual fixations were near a face, meticulously distinguishing single fixations. Second, we considered repeated fixations on very close-by locations, termed ‘return fixations.’ Responses were more similar during return fixations, and again distinguished individual fixations. Third, computation models could partially explain neuronal responses from an image crop centered on each fixation.
These results shed light on how the IT cortex does (and does not) contribute to our daily visual percept: a stable world despite frequent saccades.
Will Xiao, Harvard University
Abstract: How does the brain support our ability to see? Studies of primate vision have typically focused on controlled viewing conditions exemplified by the rapid serial visual presentation (RSVP) task, where the subject must hold fixation while images are flashed briefly in randomized order. In contrast, during natural viewing, eyes move frequently, guided by subject-initiated saccades, resulting in a sequence of related sensory input. Thus, natural viewing departs from traditional assumptions of independent and unpredictable visual inputs, leaving it an open question how visual neurons respond in real life.
We recorded responses of interior temporal (IT) cortex neurons in macaque monkeys freely viewing natural images. We first examined responses of face-selective neurons and found that face neurons responded according to whether individual fixations were near a face, meticulously distinguishing single fixations. Second, we considered repeated fixations on very close-by locations, termed ‘return fixations.’ Responses were more similar during return fixations, and again distinguished individual fixations. Third, computation models could partially explain neuronal responses from an image crop centered on each fixation.
These results shed light on how the IT cortex does (and does not) contribute to our daily visual percept: a stable world despite frequent saccades.


![Benchmarking Out-of-Distribution Generalization Capabilities of DNN-based Encoding Models for the...
[full title] Benchmarking Out-of-Distribution Generalization Capabilities of DNN-based Encoding Models for the Ventral Visual Cortex
Authors: Spandan Madan, Will Xiao, Mingran Cao, Hanspeter Pfister, Margaret Livingstone, Gabriel Kreiman
Link to Paper: https://arxiv.org/abs/2406.16935
Abstract: We characterized the generalization capabilities of DNN-based encoding models when predicting neuronal responses from the visual cortex. We collected textit{MacaqueITBench}, a large-scale dataset of neural population responses from the macaque inferior temporal (IT) cortex to over 300,000 images, comprising 8,233 unique natural images presented to seven monkeys over 109 sessions. Using textit{MacaqueITBench}, we investigated the impact of distribution shifts on models predicting neural activity by dividing the images into Out-Of-Distribution (OOD) train and test splits. The OOD splits included several different image-computable types including image contrast, hue, intensity, temperature, and saturation. Compared to the performance on in-distribution test images the conventional way these models have been evaluated models performed worse at predicting neuronal responses to out-of-distribution images, retaining as little as 20% of the performance on in-distribution test images. The generalization performance under OOD shifts can be well accounted by a simple image similarity metric the cosine distance between image representations extracted from a pre-trained object recognition model is a strong predictor of neural predictivity under different distribution shifts. The dataset of images, neuronal firing rate recordings, and computational benchmarks are hosted publicly at: https://bit.ly/3zeutVd Benchmarking Out-of-Distribution Generalization Capabilities of DNN-based Encoding Models for the...](https://i.ytimg.com/vi/bhXB-djJ9lc/mqdefault.jpg)







