Full text
77,191 characters
· extracted from
preprint-html
· click to expand
Higher-level spatial prediction in natural vision across mouse visual cortex | bioRxiv /* */ /* */ <!-- <!-- /*! * yepnope1.5.4 * (c) WTFPL, GPLv2 */ (function(a,b,c){function d(a){return"[object Function]"==o.call(a)}function e(a){return"string"==typeof a}function f(){}function g(a){return!a||"loaded"==a||"complete"==a||"uninitialized"==a}function h(){var a=p.shift();q=1,a?a.t?m(function(){("c"==a.t?B.injectCss:B.injectJs)(a.s,0,a.a,a.x,a.e,1)},0):(a(),h()):q=0}function i(a,c,d,e,f,i,j){function k(b){if(!o&&g(l.readyState)&&(u.r=o=1,!q&&h(),l.onload=l.onreadystatechange=null,b)){"img"!=a&&m(function(){t.removeChild(l)},50);for(var d in y[c])y[c].hasOwnProperty(d)&&y[c][d].onload()}}var j=j||B.errorTimeout,l=b.createElement(a),o=0,r=0,u={t:d,s:c,e:f,a:i,x:j};1===y[c]&&(r=1,y[c]=[]),"object"==a?l.data=c:(l.src=c,l.type=a),l.width=l.height="0",l.onerror=l.onload=l.onreadystatechange=function(){k.call(this,r)},p.splice(e,0,u),"img"!=a&&(r||2===y[c]?(t.insertBefore(l,s?null:n),m(k,j)):y[c].push(l))}function j(a,b,c,d,f){return q=0,b=b||"j",e(a)?i("c"==b?v:u,a,b,this.i++,c,d,f):(p.splice(this.i++,0,a),1==p.length&&h()),this}function k(){var a=B;return a.loader={load:j,i:0},a}var l=b.documentElement,m=a.setTimeout,n=b.getElementsByTagName("script")[0],o={}.toString,p=[],q=0,r="MozAppearance"in l.style,s=r&&!!b.createRange().compareNode,t=s?l:n.parentNode,l=a.opera&&"[object Opera]"==o.call(a.opera),l=!!b.attachEvent&&!l,u=r?"object":l?"script":"img",v=l?"script":u,w=Array.isArray||function(a){return"[object Array]"==o.call(a)},x=[],y={},z={timeout:function(a,b){return b.length&&(a.timeout=b[0]),a}},A,B;B=function(a){function b(a){var a=a.split("!"),b=x.length,c=a.pop(),d=a.length,c={url:c,origUrl:c,prefixes:a},e,f,g;for(f=0;f<d;f++)g=a[f].split("="),(e=z[g.shift()])&&(c=e(c,g));for(f=0;f<b;f++)c=x[f](c);return c}function g(a,e,f,g,h){var i=b(a),j=i.autoCallback;i.url.split(".").pop().split("?").shift(),i.bypass||(e&&(e=d(e)?e:e[a]||e[g]||e[a.split("/").pop().split("?")[0]]),i.instead?i.instead(a,e,f,g,h):(y[i.url]?i.noexec=!0:y[i.url]=1,f.load(i.url,i.forceCSS||!i.forceJS&&"css"==i.url.split(".").pop().split("?").shift()?"c":c,i.noexec,i.attrs,i.timeout),(d(e)||d(j))&&f.load(function(){k(),e&&e(i.origUrl,h,g),j&&j(i.origUrl,h,g),y[i.url]=2})))}function h(a,b){function c(a,c){if(a){if(e(a))c||(j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}),g(a,j,b,0,h);else if(Object(a)===a)for(n in m=function(){var b=0,c;for(c in a)a.hasOwnProperty(c)&&b++;return b}(),a)a.hasOwnProperty(n)&&(!c&&!--m&&(d(j)?j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}:j[n]=function(a){return function(){var b=[].slice.call(arguments);a&&a.apply(this,b),l()}}(k[n])),g(a[n],j,b,n,h))}else!c&&l()}var h=!!a.test,i=a.load||a.both,j=a.callback||f,k=j,l=a.complete||f,m,n;c(h?a.yep:a.nope,!!i),i&&c(i)}var i,j,l=this.yepnope.loader;if(e(a))g(a,0,l,0);else if(w(a))for(i=0;i (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0];var j=d.createElement(s);var dl=l!='dataLayer'?'&l='+l:'';j.src='//www.googletagmanager.com/gtm.js?id='+i+dl;j.type='text/javascript';j.async=true;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-M677548'); Skip to main content Home About Submit ALERTS / RSS Search for this keyword Advanced Search New Results Higher-level spatial prediction in natural vision across mouse visual cortex View ORCID Profile Micha Heilbron , View ORCID Profile Floris P de Lange doi: https://doi.org/10.1101/2025.05.15.654212 Micha Heilbron 1 Donders Institute for Brain, Cognition and Behaviour 2 Amsterdam Brain and Cognition, University of Amsterdam Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Micha Heilbron For correspondence: m.heilbron{at}uva.nl Floris P de Lange 1 Donders Institute for Brain, Cognition and Behaviour Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Floris P de Lange Abstract Full Text Info/History Metrics Preview PDF Abstract Theories of predictive processing propose that sensory systems constantly predict incoming signals, based on spatial and temporal context. However, evidence for prediction in sensory cortex largely comes from artificial experiments using simple, highly predictable stimuli, that arguably encourage prediction. Here, we test for sensory prediction during natural scene perception. Specifically, we use deep generative modelling to quantify the spatial predictability of receptive field (RF) patches in natural images, and compared those predictability estimates to brain responses in the mouse visual cortex – while rigorously accounting for established tuning to a rich set of low-level image features and their local statistical context — in a large scale survey of high-density recordings from the Allen Institute Brain Observatory. This revealed four insights. First, cortical responses across the mouse visual system are shaped by sensory predictability, with more predictable image patches evoking weaker responses. Secondly, visual cortical neurons are primarily sensitive to the predictability of higher-level image features, even in neurons in the primary visual areas that are preferentially tuned to low-level visual features. Third, unpredictability sensitivity is stronger in the superficial layers of primary visual cortex, in line with predictive coding models. Finally, these spatial prediction effects are independent of recent experience, suggesting that they rely on long-term priors about the structure of the visual world. Together, these results suggest visual cortex might predominantly predict sensory information at higher levels of abstraction – a pattern bearing striking similarities to recent, successful techniques from artificial intelligence for predictive self-supervised learning. 1 Introduction Theories of predictive processing propose that neural information processing critically relies on comparing (bottom-up) incoming signals to self-generated (top-down or locally recurrent) predictions, derived from a generative model ( Rao, 2024 ; Friston, 2005 ; Keller and Mrsic-Flogel, 2018 ), in order to drive both sensory inference and self-supervised learning. Over the past decade, predictive processing has become very influential, as a body of findings accumulated that support the framework ( Keller and Mrsic-Flogel, 2018 ; de Lange et al., 2018 ). Most of this work shows that stimuli evoke weaker responses when they are predictable and stronger responses when they are surprising; a phenomenon known as expectation suppression which is typically interpreted as reflecting sensory prediction errors. In perception, such effects are typically studied by experimentally manipulating stimulus probability – for instance by extensively exposing observers to one stimulus predictably following another stimulus (e.g. Meyer and Olson, 2011 ; Parras et al., 2017 ; Richter et al., 2022 ; see de Lange et al., 2018 ; Heilbron and Chait, 2018 for review). One limitation of this line of work is that it is unclear whether effects found under such artificial (and prediction-encouraging) conditions generalise to natural conditions. Typically, there are only a handful of stimuli that can occur in these experiments (e.g., a clockwise or counter-clockwise grating), whereas the realm of possibilities in natural conditions is nearly limitless. Moreover, because most studies use relatively coarse comparisons (e.g. expected vs unexpected stimulus), it typically remains unclear at what level of granularity sensory predictions are made: does sensory cortex make predictions about low-level, mid-level or high-level stimulus category features, or perhaps all of these at the same time? Advances in generative AI are allowing for a different way of studying predictive processing that can address both these issues. Instead of experimentally manipulating stimulus predictability, this involves estimating stimulus predictability using generative models, to test if predictable stimuli (along a certain dimension) evoke an attenuated neural response. Critically, this enables (i) studying predictions in natural conditions (i.e. without experimentally imposing the prediction on the observer), and (ii) assessing their granularity, as predictability can be computed at multiple levels of abstraction ( Heilbron et al., 2022 ; de Lange et al., 2022 ). Recently, Uran et al. (2022) introduced a framework to apply this approach to spatial predictability in visual scene perception. They recorded multi-unit activity from primary visual cortex in macaques viewing natural images. Then, using Deep Neural Networks (DNNs) trained for in-painting, they quantified the spatial predictability of image patches that were centered on the classical receptive field of the neurons, and found that less predictable patches evoked higher firing rates in macaque V1. Strikingly, in a later phase of processing (200-600 ms after stimulus onset) they found the unpredictability of high-level features of the image part drove activity most strongly. However, their analysis was based on a single area in macaque visual cortex, leaving it unclear how their findings would extend beyond primate vision and across the visual cortical hierarchy. Here, we build on their approach, and apply it to the Allen Institute Neuropixels dataset ( Siegle et al., 2021 ). This is a high-quality, comprehensive dataset of thousands of individual neurons across multiple cortical layers of the entire mouse visual cortical system of 32 mice. This dataset covers an extensive stimulus battery that includes controlled stimuli to characterise the tuning properties of individual neurons. This allowed us to quantify the spatial predictability for every stimulus, for every receptive field of every individual unit across the visual cortical hierarchy, while controlling for receptive field visual features and local statistical context (including contrast energy, orientation content, orientation consistency and coherence). This way, we ask four questions: (i) are cortical responses to natural images in visual cortex modulated by visual spatial predictability?; (ii) at what level of granularity do neurons in the visual cortex predict, and how does this differ across different areas of the visual system?; (iii) are prediction effects stronger in superficial compared to deep layers of cortex?; (iv) are these effects based on short-term online-learnt expectations, or on more long-term structural priors, independent of recent experience? To preview, we find that spatial predictability modulates cortical responses across the mouse visual cortex, with neurons most sensitive to high-level predictability even in primary visual areas. These spatial predictability effects were indepdent of recent experience, and in V1 they were strongest in superficial layers. Together, these results shed new light on sensory prediction, with implications for models of predictive processing and self-supervised learning in sensory cortex. 2 Results We analyzed the Allen Institute Neuropixels dataset ( Siegle et al., 2021 ), focusing on spiking responses to the “natural scenes” stimulus set, in which head-fixed, awake mice were mnocularly presented with full-field natural images, for 250 ms per presentation and 50 repetitions per stimulus ( Figure 1a ). Critically, stimuli were presented in a pseudo-random order without predictable temporal structure or prediction-encouraging task, allowing us to test for prediction in complex visual processing, without extraneous highly predictable structure. For our analyses, we selected neurons with well-defined receptive fields (see Methods for selection criteria) across six cortical visual areas that exhibit a hierarchical organization as established by Siegle et al. ( 2021 ). While our analyses encompass all six areas, our main focus was on primary visual cortex (V1), as it offered substantially greater coverage with the highest number of included neurons across the sample of mice. Download figure Open in new tab Figure 1. A) Dataset. We analysed cortical responses to natural images from the Allen Institute Visual Coding Neuropixels dataset, in which mice viewed full-field natural images for 250 ms per image, for a total of 5,900 trials ( Siegle et al., 2021 ). B) Prior to the natural scenes, flashing gratings were presented allowing to map the receptive field (RF) of every unit; we selected neurons based on their RF characteristics (see Methods ). C) Spatial predictability analysis: for every image and every neuron, the area spanning the receptive field was masked out and filled-in using a reconstructive autoencoder with skip connections (UNet; Liu et al. (2018) ). The predicted image patch was compared to the actual patch to compute image patch predictability ( Uran et al., 2022 ). D) To control for low-level image statistics, we generated multi-scale first-order (contrast energy) and second-order (spatial coherence) contrast images using Gabor filter banks ( Groen et al., 2013 ; Kay et al., 2013 ). From these, we derived a comprehensive set of baseline statistics (including measures of local contrast, orientation content and homogeneity, and coherence; see Methods ). 2.1 Firing rates throughout visual cortex are modulated by image patch predictability We first asked whether neural responses were modulated by overall image patch predictability. To compute image patch predictability for every patch in every image, we masked out the image patch that spans the receptive field of an image, then use a deep generative model to predict or “fill-in” the missing image patch, and compare the predicted image patch with the actual image patch to compute overall predictability. This was done for every classical receptive field of every unit of every image individually (see Figure 1c and Methods ). Cortical responses were assessed as firing rate activity time-courses (expressed as fold-change with respect to baseline) for every neuron, for every stimulus. To isolate the influence of image patch unpredictability from known drivers of visual cortical activity, we first fitted a robust baseline regression model to the firing rates at every timepoint. This baseline model incorporated a series of non-prediction-related variables, including running speed and a comprehensive set of low-level image statistics (e.g., local contrast energy across spatial frequencies, orientation content, and circular variance, indexing contour homogeneity; detailed in Methods 4.6). To evaluate whether unpredictability significantly modulated firing rate, we then evaluated whether adding our image patch unpredictability metric could explain unique, additional variance beyond this comprehensive baseline. To test this, we computed Δ R – the increase in cross-validated prediction performance when unpredictability was added to the baseline model ( Figure 2a-b ). Note that this cross-validated metric provides a conservative test for the presence of an effect, rather than a direct estimate of the effect size magnitude. Download figure Open in new tab Figure 2. A) Cross-validated prediction performance of the baseline model across the visual cortical areas, recapitulating the hierarchical organization established in Siegle et al. ( 2021 ). b) Average increase in cross-validated performance between 75-150 ms – see dotted lines in panel a) – when adding unpredictability to the baseline model. Small dots indicate individual animals; large dots with error bar indicate mean across animals and bootstrapped 95% CI. Note, Δ R provides a conservative test for the presence of an effect; for effect size, see β coefficients in panels c-d. c) Effect of unpredictability and contrast on average response in V1. Curves show the average response (blue curve), and the average response plus β coefficients of unpredictability (red) and contrast (gray). Note that firing rate is expressed as fold-change with respect to baseline (0-30 ms; see Methods ). d) Same as panel c) but showing just the modulation function ( β coefficients), rather than effect on average response. In both panels, curves show average coefficient across neurons for each animal, averaged across animals; shading shows bootstrapped 95% CI across animals. This revealed robust effects of unpredictability, both in V1 (mean Δ R = 0.0036, 95% CI [0.0026, 0.0046], bootstrap t -test with Bonferroni correction: p < 0.0001), and across downstream visual areas (VISl: mean Δ R = 0.0021, 95% CI [8.45 × 10 −4 , 0.0034], p = 0.006; VISrl: mean Δ R = 0.0022, 95% CI [0.0010, 0.0034], p = 0.0012; VISal: mean Δ R = 0.0014, 95% CI [4.47 × 10 −4 , 0.0026], p = 0.0144; VISam: mean Δ R = 0.0017, 95% CI [9.77 10 −4 , 0.0025], p < 0.0001), with the exception of VISpm where the effect was not significant (mean Δ R = 0.0012, 95% CI [−4.52 × 10 −4 , 0.0025], p = 0.282). Together, this shows that cortical responses are indeed sensitive to spatial unpredictability, across multiple cortical regions. Having established that spatial unpredictability impacts the neural response, we next examined the coefficients of the regression model to characterize the nature of this effect. The coefficients of this regression model can be interpreted as ‘modulation functions’ that quantify how a variable modulates the firing rate over time. As can be seen from Fig 2c-d , we observed a clear positive modulation by unpredictability (average coefficient between 75–150 ms: M = 0.077, 95% CI [0.059, 0.096], p < 0.0001, bootstrap t-test across animals). This indicates that more predictable image patches are associated with weaker neural responses (expectation suppression). For reference, the effect of unpredictability can be compared to the effect of local contrast energy in a unit’s preferred contrast channel. This reveals that the effect of contrast is stronger than the effect of unpredictability (paired difference between 75–150 ms: M = 0.132, 95% CI [0.108, 0.157], p < 0.0001). Critically, the contrast effect appears to modulate the neural response earlier ( Figure 2c-d ; average latency to 50% peak: M = 71.0 ms, SE = [68.8, 73.2] vs. unpredictability: M = 77.3 ms, SE = [75.1, 79.6]), in line with the notion that predictability effects rely more on recurrent processing. 2.2 Opposite tuning for image features and image predictability After confirming that overall image patch predictability modulated responses across mouse visual cortex, we next assessed the representational content of these predictions. This was done by estimating the neural effect not just of overall predictability, but of predictability at multiple levels of abstraction ( Uran et al., 2022 ; Heilbron et al., 2022 ). These multi-level predictability estimates were obtained by comparing the actual and predicted patch at increasingly abstract feature spaces. For this we used the convolutional layers of AlexNet ( Krizhevsky et al., 2012 ), a shallow CNN with a good fit to mouse visual cortex ( Nayebi et al., 2021 ). Because early CNN layers extract simple features (e.g. edges) and higher levels more complex ones (e.g. textures), performing the comparison at each layer allows to quantify predictability at multiple levels of abstraction ( Fig 3a-b ). Download figure Open in new tab Figure 3. A) Multi-level predictability analysis. To compute unpredictability at multiple levels, the predicted RF-patch and actual input RF-patch were fed into a CNN to extract a series of increasingly abstract features. The predicted and actual input patch were then compared for each level separately, to separately quantify the predictability of low- and high-level features. b) Dissociation between high-level and low-level predictability. Some illustrative example image patches at four quadrants of the high-level vs low-level unpredictabillity axes. One common example are textures (grass, moss, sandstone), for which the exact composition of low-level features (edges and orientations) are highly unpredictable, but the higher-level structure (the texture) is perfectly predictable; hence right top corner. See Methods for more details. c) Selectivity for image features (blue) and image feature predictability (red) in V1. Image feature encoding (blue) shows normalised encoding, i.e. correlation coefficient divided by the maximum encoding performance per unit, before averaging across units. Unpredictability sensitivity is defined as average coefficient in the 75-150 ms timewindow (see Figure 2 ). Dots show mean across animal-level averages; error bars show bootstrapped 95% confidence intervals. Focusing on V1 first, we computed the unpredictability sensitivity as the average coefficient between 75-150 ms, for each level of unpredictability. This revealed a clear positive relationship (see Figure 3c ): firing rates in V1 were least sensitive to low-level unpredictability, and most sensitive to high-level unpredictability; and this pattern was highly consistent across animals (mean slope = 0.149, 95% CI [0.096, 0.210], bootstrap: p < 0.0001), suggesting that early visual cortex predicts higher rather than lower-level visual features. While this is in line with recent observations ( Uran et al., 2022 ; Richter et al., 2024 ) the pattern is striking since the unpredictability-sensitivity to higher versus lower-level features is the opposite of the well-characterised feature-sensitivity of V1. Indeed, when we examined the tuning profile of the neurons, by performing a more standard encoding-modelling based feature-sensitivity analysis – i.e. predicting firing rates as a function of CNN image features, rather than the unpredictability – we found that V1 was tuned to low-level, instead of high-level, features. Namely, V1 firing rates were best predicted by low-level image features (from early CNN layers), and progressively worse by higher-level image features. This negative relationship recapitulates traditional bottom-up feature tuning, and was again highly consistent across animals (mean slope = −0.034, 95% CI [−0.039, −0.027], bootstrap test across animals: p < 0.0001). This dissociation between feature-sensitivity (preferring low-level features) and unpredictability-sensitivity (preferring high-level features) provides clear evidence that the high-level unpredictability effects are distinct from V1’s basic tuning to simple image features (see Discussion). Interestingly, when we repeated the same analysis for other downstream areas in the mouse visual cortex, we found largely the same pattern of effects: feature sensitivity was highest low-level image features, but unpredictability sensitivity was highest for high-level features – although we also found that, for encoding, later level visual areas are more sensitive to higher-level features than early visual areas (see Figure S1 , S2 ). The relatively modest differences in selectivity between V1 and higher visual areas is in line with earlier encoding analyses of mouse visual cortex ( Nayebi et al., 2021 ). As a final confirmatory control, we repeated the predictability-sensitivity analysis, but now using cross-validation: quantifying predictability effects as the increase in regression model fit when adding each unpredictabil-ity estimate to the baseline model. This reveals a very similar positive relationship for V1 ( Figure S3 a ), that is highly consistent across animals (mean slope across Δ R = 7.56 × 10 −4 , 95% CI [4.68 × 10 −4 , 1.09 × 10 −3 ], boot-strap t -test: p < 0.0001). Indeed, we find again a similar pattern across all cortical areas, where unpredictability effects are strongest for higher-level unpredictability and weakest for low-level unpredictability ( Figure S3 b ). Together, we observed that predictability sensitivity does not simply follow the neuron’s bottom-up feature sensitivity. Instead, we found a distinct preference for unpredictability of higher-level features. This results in a highly consistent pattern of opposite tuning: neurons are most sensitive to lower-level image features, but higher-level feature predictability. 2.3 Predictability effects are strongest in superficial layers of primary visual cortex A distinguishing feature of predictive coding models is that they postulate two distinct populations of neurons: error neurons that signal the mismatch between predictions and sensory input, and prediction or representation neurons that encode internal models of the environment ( Rao and Ballard, 1999 ; Friston, 2005 ; de Lange et al., 2018 ; Keller and Mrsic-Flogel, 2018 ). In hierarchical predictive coding models, predictions flow downward through the cortical hierarchy while prediction errors propagate forward. Because in cortex, feedforward signals originate from superficial layers (2/3) while feedback signals stem from deep layers (5/6), this anatomical asymmetry maps directly onto the functional asymmetry in predictive coding. Namely, superficial-layer neurons should primarily signal prediction errors, whereas deep-layer neurons should encode predictive representations of sensory input. The Visual Coding dataset allowed to test this hypothesized laminar organization, as it provides simultaneous recordings across all cortical layers of the visual system ( Figure 4a ). We took this opportunity by classifying neurons as either superficial (layer 2/3) or deep (layer 5/6) based on their precise recording depths (see Methods for detailed classification criteria). Then, we tested whether predictability effects (assuming they reflect sensory prediction errors) were indeed strongest in superficial layers of cortex. To this end, we quantified the modulation of firing rates (75-150 ms) by overall image patch predictability. Strikingly, in V1, we indeed identified this expected difference: the modulation by image patch predictability was consistently stronger for neurons in superficial layers compared to those in deeper layers (mean summed coefficient difference = 0.336 FC/σ , 95% CI [0.049, 0.605], p = 0.0173). Importantly, this laminar difference was specific to predictability and not observed for contrast energy (mean difference = −0.294, 95% CI [−0.949, 0.255], p = 0.2774; interaction, difference of difference: mean = 0.630, 95% CI [0.054, 1.347], p = 0.011). Download figure Open in new tab Figure 4. A) Neuropixels probe cortical activity across layers of the entire visual system, allowing to dissociate units in superficial (2/3) and deep (5-6) cortical layers. Figure adapted from Siegle et al. ( 2021 ). b) V1: effect of overall image patch predictability and contrast energy (average coefficient between 75-150ms), in neurons in superficial and deep cortical layers. Small dots with connecting lines indicate individual animals; big dots with error bars indicate mean across animal-averages, and bootstrapped 95% confidence interval. * indicates p<0.05. Cc Same as in b , but for the multi-level predictability estimates; dots indicate mean across animals, error bars indicate 95% confidence interval around the mean. d) V1: effect of overall image patch predictability over time (as in Fig. 2d ), split out for units in superficial versus deep layers. Solid line indicates mean modulation function across animals; shaded area indicates bootstrapped 95% confidence interval around the mean. For a similar plot for other cortical areas, see Figure S4 Although the observed laminar difference is subtle, it appears robust across methodological variations, as we found the same patterns not just for overall predictability, but for each of the multi-level predictability estimates ( Figure 4c ), and across time ( Figure 4d ). However, because the analysis involved subdividing the already sub-selected units into those classified as ‘superficial’ and ‘deep’ (and discarding others), the analysis could only be feasibly conducted in V1; we did not have enough sensitivity in other areas to meaningfully perform the comparison (see Figure S4 ). 2.4 No effect of short-term experience on predictability modulations Finally, we asked whether the predictability effects were dependent on short-term experience. One possibility would be that these effects reflect more long-term, ‘hard-wired’ priors about the structure of the natural visual world, largely independent of recent experience. However, the stimuli were repeated 50 times, so the animals built up extensive experience with the images. Since recent experience and familiarity are known to shape cortical processing in mouse visual cortex ( Homann et al., 2022 ; Garrett et al., 2020 ; Fritsche et al., 2022 ), such extensive repetition might enhance the predictability effect – or may even be required for it to emerge in the first place. To adjudicate between these possibilities, we estimated the effect separately for the first and second half of the dataset. Unsurprisingly, the average response was lower ( Figure 5 ) in the second half of trials ( ± 25 repetitions), compared to the first (difference between 75-150 ms: mean = 0.076, 95% CI [0.028, 0.119], p = 0.0124), potentially reflecting general adaptation. There was also a small but significant difference for contrast energy (mean = 0.024, 95% CI [0.002, 0.051], p = 0.0182), indicating that contrast energy modulated the neural response less strongly in the second half of trials. Strikingly, however, we observed no difference for the unpredictability effect (mean = 0.004, 95% CI [−0.006, 0.015], p = 0.4006; see Figure 5 ). Together, this suggests the spatial prediction effects we observed appear independent of recent experience. Download figure Open in new tab Figure 5. A) Primary visual cortex: average response (coefficient of time-resolved intercept) in first vs second half of the trials. b) Same but for β coefficient of contrast energy in the preferred contrast channel. c) Same but for image patch unpredictability. d) Difference waves (early-late) for all three effects. More extensive exposure leads to a reduced average response and reduced contrast effect, but no difference for unpredictability. In all panels, solid lines are mean coefficients across the animal-level means, shaded area indicates bootstrapped confidence interval around the across-animal mean. 3 Discussion We quantified the spatial predictability of local receptive field image patches for thousands of individual neurons using deep generative modelling, and related these predictability estimates to neural responses in a high-quality survey of visual cortex of mice viewing natural scenes. This revealed that neural responses throughout mouse visual cortex are shaped by sensory predictability, even after controlling for a range of low-level image statistics such as local contrast, orientation content, and contour-related properties. In line with predictive coding models, we found that this spatial prediction effect was stronger in superficial layers of primary visual cortex — but contrary to predictive coding models, this prediction error signal is in the wrong frame of reference: neurons are most sensitive to the predictability of higher-level features, even in regions tuned to low-level features. Finally, these effects occurred independent of recent experience with the stimuli, suggesting they reflect longer-term structural priors about the visual world rather than expectations based on the exposure to these specific stimuli. Together, these results provide a new window into how predictive processing operates under natural conditions. We observe strong modulations of cortical responses by spatial predictability, which suggests that visual cortex is constantly predicting incoming signals. The word “predicting” may seem odd here, as spatial predictions appear quite different from temporal predictions, i.e. predicting an upcoming stimulus, which are typically used to study predictive processing ( Richter et al., 2018 ; Parras et al., 2017 ; de Lange et al., 2018 ; Heilbron et al., 2022 ). Note however, that the canonical models that popularised predictive processing (e.g. Rao and Ballard, 1999 ; Lee and Mumford, 2003 ) deal exclusively with spatial (reconstructive) predictions, and do not include prediction over time. In that way, we study sensory prediction in the traditional predictive processing sense. What is shared between both types of prediction is the key assumption that the brain compares sensory input to predictions of what signals to expect based on a generative model and the spatial or temporal context. Plus, critically, the neural marker that responses are reduced when incoming signals are more in line with these (top-down or locally recurrent) contextual predictions. Spatial context has a long history in vision ( Biederman et al., 1982 ; Oliva and Torralba, 2007 ; Bar, 2004 ; Peelen et al., 2024 ). Perhaps the neurophysiologically best chacterised spatial context effects come from the center-surround literature. The canonical finding here is surround suppression , where responses to stimuli in the RF are reduced by the presence of particular surrounding stimuli ( Hubel and Wiesel, 1965 ; Knierim and Van Essen, 1992 ; Angelucci et al., 2017 ). This suppression is strongest when the content of the RF is coherent with and therefore predictable by the surround, yielding explanations based on redundancy reduction or predictive coding which cast surround suppression as a kind of expectation suppression, i.e. reduced prediction error ( Coen-Cagli et al., 2015 ; Self et al., 2014 ; Rao and Ballard, 1999 ). However, the opposite (surround excitation) is also observed ( Lamme and Roelfsema, 2000 ; Angelucci et al., 2017 ), and recent modelling-based analyses suggest it may be much more common than previously thought, in particular for natural stimuli ( Fu et al., 2024 ). Due to methodological differences, it is difficult to link our results directly to the center-surround effect literature, where the key analysis is based on comparing the content of the RF with the content of the surround, which is not what we do. Conceptually, though, the suppressive effect of predictability (or facilitating effect of unpredictability) we observe is in line with surround suppression, and in particular predictive or efficient coding based explanations thereof. However, they critically depart from typical surround suppression, where modulation often relies on similarly-tuned low-level features. Instead, we find a dissociation between bottom-up feature tuning and predictability tuning, with a dominant modulation by higher-level feature predictability , even in areas primarily sensitive to lower-level features. Predictability modulations were enhanced in superficial layers of cortex ( Figure 4 ). On face value, this may appear an important confirmation of predictive coding theories which critically and uniquely postulate such a laminar dissociation ( Rao and Ballard, 1999 ; Friston, 2005 ; Bastos et al., 2012 ), a distinction for which only recently supporting evidence is starting to accumulate Thomas et al. (2024) ; Furutachi et al. (2024) ; Jordan and Keller (2020) . However, we could only reliably observe the effect in V1, and critically, the magnitude of the effect was relatively small. Therefore, if one takes the predictability effect to reflect sensory prediction errors, this could even appear to speak against a strong laminar segregation, where prediction errors are only encoded in superficial layers. However, we caution against drawing strong conclusions in either directions, because the neural markers of prediction errors and the downstream effects of prediction errors are difficult to disentangle. Even in a classical predictive coding circuit — with error units in superficial and prediction units in deep layers — stronger prediction errors would result in stronger updates in prediction units, and therefore, in the aggregate, more spiking activity for more unpredictable stimuli in deep layers too. Therefore, while the laminar effect we observe is striking, strong conclusions about the existence of error/prediction units would arguably require analyses at a level of physiological granularity well beyond what is possible here. The predictability effects we observe appeared surprisingly independent of recent experience. This is notable because cortical processing in mouse visual cortex is typically highly malleable, with neural responses strongly shaped by recent experience and familiarity ( Homann et al., 2022 ; Garrett et al., 2020 ; Fritsche et al., 2022 ). This stability particularly contrasts with the findings of Seignette et al. (2024) , who, studying similar spatial context effects, found that responses to occluded scenes strengthen systematically as mice learn the complete versions of these images. These divergent findings reveal that even for spatial prediction, mouse visual cortex implements multiple prediction mechanisms in parallel: while some contextual effects rely on rapid learning of specific patterns, others appear to reflect more stable structural priors about the visual world. Perhaps most strikingly, we observed a divergence between feature tuning and predictability tuning: Neurons are most sensitive to the predictability of higher-level visual features, even in areas most sensitive to low-level features, like V1 ( Figure 3 ). This is a departure from classical predictive processing models, where the content of predictions aligns with the tuning for visual features, via hierarchical inference. For instance, early stages would reflect expected low-level features based on higher-order interpretations, such as when a faint edge is highly expected as a part of a key object boundary ( Rao and Ballard, 1999 ; Lee and Mumford, 2003 ). However, despite departing from theoretical models, our results align with emerging empirical evidence. The closest parallel is from Uran et al. (2022) who performed a similar analysis in macaque V1 and also found greatest sensitivity to higher-level feature predictability. This was further corroborated by Richter et al. (2024) who used representational similarity analysis (RSA) to examine the content of temporal predictions in a more classical, controlled experiment of visual stimuli predictably following other stimuli. Despite their distinct methodological approach, they also observed preferential encoding of higher-level predictability throughout the visual hierarchy. The convergence of these findings across multiple species (mouse, macaque, human) and experimental approaches (electrophysiology, fMRI) appears to point to a fundamental functional principle: visual cortex predicts sensory information primarily at higher levels of abstraction, even in early visual areas. Why would visual cortex predict sensory input at higher levels of abstraction? A tantalizing hint comes from recent advances in machine learning. There, researchers tried to develop general vision models by training networks to predict pixels in context ( Pathak et al., 2016 ; Vincent et al., 2008 ) – much in the same way as LLMs can learn general models of language by predicting words in context ( Radford et al., 2019 ; Devlin et al., 2019 ). However, initial attempts in vision were not nearly as successful as in language: While these models become effective at predicting pixels, they were not so effective for learning general representations that can be applied to other tasks (e.g. object recognition). Recently however, new techniques are starting to show success ( He et al., 2021 ; van den Oord et al., 2019 ; Shi et al., 2022 ; LeCun, 2022 ; Bizeul et al., 2025 ; Darcet et al., 2025 ). What these methods have in common is that they encourage predicting higher-level visual features instead of lower-level visual features. The intuition here is that – contrary to text – images contain very high levels of redundancy of low-level information, making low-level prediction tasks too simple to drive meaningful learning. By focusing on higher-level features instead, models are forced to develop more abstract and useful visual representations ( He et al., 2021 ; Bizeul et al., 2025 ). Our results suggest the brain may have found a similar strategy, predicting specifically at a higher level of abstraction, perhaps as a way for effective self-supervised learning of visual representations. Altogether, our results demonstrate that predictive processing is pervasive throughout visual cortex, operating predominantly at higher levels of abstraction - even in early sensory areas. This organization may reflect a fundamental neural strategy for learning robust visual representations from the statistical structure of the natural world. 4 Methods 4.1 Stimuli and data We analysed the ‘Brain Observatory 1.1’ subset of the Allen Institute Visual Coding Neuropixels dataset ( Siegle et al., 2021 ). This is an openly available dataset that surveys spiking activity across the mouse visual system, recorded using 6 simultaneously inserted Neuropixels probes. Mice viewed a range of full-field stimuli, presented monocularly to the right eye. We focus on two stimulus sets: Gabors – for Receptive Field (RF) mapping – and Natural Scenes. The latter comprises 118 scene images, presented for 250 ms, for 50 repetitions, resulting in 5900 trials ( Fig 1A ). For analysis, raw spike times were binned into 10 ms intervals to compute time-resolved firing rates. To enhance signal-to-noise ratio, firing rates were smoothed across time using a uniform moving average filter with a 30 ms window. Because different neurons exhibited widely varying firing rate ranges, we normalised the firing rates to decrease the variance between neurons and avoid that the population or animal-level averages would be dominated by high-firing-rate neurons, We normalized firing rates as fold change relative to baseline, defined as the mean activity during the initial 20 ms of stimulus presentation. This normalization was computed as (FR-baseline)/baseline, with a small pseudo-baseline value (of 0.5 Hz) added to avoid division by extremely small numbers. We further enhanced the signal-to-noise ratio by averaging responses across multiple presentations of the same stimulus. Specifically, the 50 repetitions of each natural scene were divided into 5 chunks, with each chunk representing the average response to approximately 10 presentations. This chunking approach preserved stimulus-specific response patterns while reducing noise, providing more reliable inputs for subsequent regression analyses. Finally, to minimize the influence of eye movements and blinks, we implemented quality control on each trial. Trials were excluded (i.e. removed before the averaging into chunks) if they contained blinks, or outlying horizontal or vertical eye positions, defined as those exceeding 2 standard deviations from the mean. 4.2 Receptive Field Mapping and unit selection RFs were modelled by fitting a 2d-Gaussian on the smoothed 2d spike histograms of the positions where gabors were presented. This resulted in RF for every identified unit. Subsequently, we selected all units for which this ideal RF that was on-screen and had a size between 150 and 750° 2 ; we did this because for much larger RFs the spatial predictability analysis did not work as the RFs became too large to inpaint. Moreover, we required that the “ideal” RF was a good fit, i.e. that the raw (smoothed) 2d histograms could be well approximated by the gaussian ( ρ > 0.7). Moreover, we required that the “ideal” RF was a good fit, i.e. that the raw (smoothed) 2d histograms could be well approximated by the gaussian ( ρ > 0.7). After applying these RF criteria and combining with the quality assurance criteria reported in Siegle et al. (2021) , this resulted in 3176 analysed cortical neurons in total that were located at the 6 pre-defined cortical areas of interest: 1173 neurons for V1, 354 neurons for VISl, 364 neurons for VISrl, 525 neurons for VISal, 253 neurons for VISpm, and 507 neurons for VISam. For each unit, we used its recorded CCF coordinates to putatively identify its corresponding cortical layer in the Allen Mouse Brain Atlas. Units were classified as “superficial” if they resided in cortical putative layers 1-3 and “deep” if they resided in putative layers 5-6. 4.3 Spatial predictability inpainting model Our inpainting model employed a PConvUNet architecture with partial convolutions, following Liu et al. Liu et al. (2018) . The encoder consisted of up to eight blocks, with the first using a 7×7 kernel followed by 5×5 kernels in the second and third blocks, and 3×3 kernels in subsequent blocks. Each encoder block combined a partial convolution with batch normalization (except the first layer) and ReLU activation. The partial convolution technique progressively filled missing regions by updating a binary mask after each operation, normalizing feature values based on the ratio of valid pixels. The decoder mirrored the encoder with eight blocks, each containing nearest-neighbor upsampling followed by concatenation with corresponding encoder features, partial convolution, batch normalization, and LeakyReLU activation. This skip-connection structure preserved spatial information lost during downsampling. During fine-tuning, encoder batch normalization layers were frozen to maintain pretrained feature extraction capabilities while allowing decoder parameters to adapt to the new dataset characteristics. The loss function combined reconstruction, content, and style components with weighted hyperparameters. Reconstruction loss incorporated Fourier domain metrics, while content and style losses compared VGG-16 acti-vations and their Gramian matrices, respectively, minimizing perceptual artifacts common in generative models. 4.4 Model training and fine-tuning We used a pretrained model initialized with VGG-16 weights and first trained on the Places2 dataset Zhou et al. (2017) for 60k iterations with a batch size of 64, focusing on large circular masks to improve performance on receptor-field-sized regions. We employed the Adam optimizer with a learning rate of 0.0005 and default momentum parameters ( β 1 = 0.9, β 2 = 0.999). For the final adaptation stage, we fine-tuned the model for 1500 steps on the Van Hateren Natural Images dataset ( Van Hateren and van der Schaaf, 1998 ), using a learning rate of 9.5 e − 5, excluding images used in the Allen Institute Dataset. This step specifically optimized the model for grayscale natural image statistics matching our experimental stimuli. We maintained the same hyperparameters during fine-tuning, with loss coefficients of 1.0 for valid regions, 12 for hole regions, 0.1 for total variation, 0.05 for perceptual loss, and 120.0 for style loss, balancing structural accuracy and perceptual quality. 4.5 Spatial predictability analysis For each neuron, we defined an RF mask as the area within the full-width at half-maximum of the gaussian receptive field. This mask was removed from the image and then in-painted by the model. For each image, the filled-in (predicted) RF patch was compared with the actual patch to compute predictability ( Fig 1A ), for this we took a rectangular mask spanning the RF circle with a dimension of times the diameter of the RF. Unpredictability was defined as the ℓ 2 distance between the observed and predicted patches at each convolutional layer of AlexNet, a shallow CNN with a good fit to mouse visual cortex ( Nayebi et al., 2021 ). We compute the discrepancy between predicted and actual patch in the feature space of a deep network because it captures perceptually relevant differences while remaining insensitive to pixel-level artifacts or distortions such as noise and intensity variations that don’t affect perceptual content, thus providing a more robust measure of perceptual distance and hence predictability ( Zhang et al., 2018 ). In our analyses, we use two types of spatial predictability estimates: the overall spatial predictability (e.g. in Figure 2 and elsewehere where indicated), and multi-level spatial predictability. The overall spatial predictability is meant to serve as an omnibus metric, for the research questions where we were not interested in feature-specific predictability. To reduce researcher degrees of freedom, we simply defined it as the average of the ℓ 2 distance for the individual layers. The multi-level predictability estimates, by contrast, where simply the layer-specific L2 distances. Because early CNN layers extract simple features (e.g. edges) and higher levels more complex ones (e.g. textures), performing the comparison at each layer allows to quantify predictability at multiple levels of abstraction ( Fig 3a ). While the predictability estimates at different levels are correlated, they also partially and systematically dissociate. Post-hoc inspection confirms they dissociate in predictable and interpretable ways. Some key illustrative examples are included in Figure 3 . Perhaps the most common stimulus patch that dissociates high and low-level predictability are textures, where the exact spatial configuration of low-level features (edges, orientations, and spatial frequencies) is highly unpredictable, but the higher-level statistical structure is perfectly predictable. This phenomenon aligns with the seminal work of Portilla and Simoncelli ( Portilla and Simoncelli, 2000 ), who demonstrated that textures can be mathematically defined and synthesized by matching higher-order correlations among low-level features (while having the low-level features themselves random and unpredictable). This principle explains why texture regions in our dataset occupy the upper-right quadrant in Figure 3B , exhibiting high low-level unpredictability but low high-level unpredictability. At other patches, both the high and low-level features are highly unpredictable, typically occurring at boundaries between distinct objects or scenes where contextual information from the surroundings provides insufficient constraints to predict either low-level features or higher-order structure within the receptive field. 4.6 Control variables To control for low-level statistics extracted in the feedforward sweep, we implemented a biologically-inspired contrast model that captures how early visual neurons respond to images, by computing both first-order and second-order contrast ( Groen et al., 2013 ; Ghebreab et al., 2009 ; Kay et al., 2013 ). The model computes local contrast using quadrature-phase Gabor filters ( Adelson and Bergen, 1985 ), which mimic the orientation and spatial frequency selectivity of V1 simple cells. Specifically, we constructed two filter banks: a first-order or contrast energy (CE) bank with spatial frequencies spanning 5 octaves (0.02-0.32 cycles per degree) and a second-order contrast or spatial coherence (SC) bank with slightly lower frequencies (0.015-0.24 cycles per degree). For each filter bank and spatial frequency, we computed the energy response E ( x, y ) by combining quadrature-phase filter responses F 0 and F π/2 according to . This energy-based formulation quantifies contrast along specific orientations (8 orientations, spanning 0- π radians). We then applied divisive normalization to these energy maps using local statistics: , where S ( x, y ) represents the local coefficient of variation (standard deviation divided by mean) computed over a Gaussian window. This step models the suppressive surround observed the visual system and enhances the representation of contrast boundaries. From these normalized contrast maps, we extracted several summary statistics for each neuron’s receptive field across all 118 natural scenes. Due to the limited number of unique images, we could not fit the contrast energy of an entire gabor pyramid to the neural response, and instead had to capture relevant contrast information with a minimal set of regressors. For each neuron, we first identified its preferred spatial frequency and orientation from the Gabor stimulus responses. We then computed the following sets of metrics within each receptive field: (1) five first-order contrast energies at five spatial frequency scales (0.02-0.32 cpd), representing the average normalized energy within the receptive field; (2) five second order contrast statistics (SC) at corresponding scales (0.015-0.24 cpd), computed as the inverse coefficient of variation ( µ/σ ) of the receptive field of the SC image; (3) preferred contrast energy, capturing energy specifically at each neuron’s preferred spatial frequency and orientation; (4) circular variance, which quantifies orientation homogeneity within the receptive field, allowing us to capture neural responses to extended contours ( Ringach et al., 2002 ); (5) relative orientation energies, indexing contrast at orientations ± 30 ° from the preferred orientation of the specific neuron, averaging across octaves; and (6) root-mean-square (RMS) contrast within the receptive field. Additionally, we incorporated running speed as a non-visual control variable, averaged across trial chunks, to account for locomotion-dependent modulation of visual responses. Altogether, this set comprehensive set of regressors enabled us to control for low-level feedforward processing while maintaining a statistically tractable model. 4.7 Regression analysis To assess the effect of spatial predictability on neural responses, we performed time-resolved regression analyses on the normalized firing rates. For each neuron and each timepoint, we fitted separate linear regression models, allowing us to track how the influence of predictability evolved over the visual response time course. We fitted two separate types of models, a baseline model containing just the control variables, plus a model additionally including spatial unpredictability. We constructed separate models for each predictability metric rather than including multiple predictability estimates simultaneously, since the correlations between different predictability metrics were often high, complicating the interpretation of the coefficients and Δ R . The regression analysis was performed separately for each neuron because each neuron has a unique receptive field, resulting in different feature matrices for each unit. For analyzing time-resolved data, we used a bin size of 10ms, which provided sufficient temporal resolution to observe the dynamics of predictability effects while maintaining adequate signal-to-noise ratio. We employed two complementary regression approaches. First, we estimated the full model coefficients ( β ) to perform inference on the relationship between predictability and neural responses. This allowed us to quantify the magnitude and timing of predictability effects across different visual areas and cortical layers. The standardized β coefficients from these models revealed how strongly neural activity was modulated by predictability compared to low-level features. Second, we used cross-validation to evaluate model performance. For each neuron, we implemented a shuffle-split cross-validation procedure (10 splits with 75% training, 25% testing) to compute correlation coefficients between predicted and actual neural responses. By comparing the performance of models with and without the predictability term, we calculated a difference in correlation coefficients (Δ R ), which quantified the unique contribution of predictability to explaining neural responses beyond what could be accounted for by low-level image statistics and behavioral variables alone. 4.8 Encoding-based feature sensitivity analysis To systematically quantify feature sensitivity across visual cortical areas, we implemented an encoding approach that assessed the predictive power of different feature levels on neural responses. Following Nayebi et al. (2021) , we employed a pretrained AlexNet CNN to extract hierarchical visual features from the stimulus set. For each unit, we constructed encoding models using feature maps from each AlexNet layer, trained to predict trial-averaged neural responses to natural scenes, with data partitioned into chunks to enhance signal-to-noise ratio while preserving stimulus-specific patterns. To accomodate the high dimensionality of CNN features relative to our stimuli (118 scenes), we employed Partial Least Squares regression with 25 components, capturing predictor space variance while avoiding overfitting. Model performance was evaluated using shuffle-split cross-validation (10 splits, 75% training, 25% testing) to compute correlation coefficients between predicted and actual neural responses, generating layer-specific encoding accuracies for each unit. Because our primary interest was in comparing relative preferences for stimulus features, we normalized cross-validated correlation scores by dividing by the maximum correlation achieved across layers for each neuron. This prevented high-SNR units from dominating population-level analyses while preserving relative layer preference patterns. The resulting normalized profiles enabled direct comparison of feature selectivity across different visual areas and cortical layers, providing complementary evidence to our spatial predictability analysis. 4.9 Statistical testing Statistical testing was performed using multi-level inference, first computing averages across units in one animal, and then performing statistics across animals. For simplicity, we ignored the fact that different animals had different amounts of (selected) neurons per region, and weighted the contribution of each animal equally. To compute p-values we used data-driven, non-analytical bootstrap t-tests, that involve resampling a null-distribution with zero mean (by removing the mean), counting across bootstraps how likely a t-value at least as extreme as the true t-value was to occur Each test used at least 10 4 bootstraps; p-values were computed without assuming symmetry (equal-tail bootstrap; Rousselet et al. 2019 ). Confidence intervals (in the figures and text) are also based on bootstrapping, with 10 4 resamples. 4.10 Data and code availability The Allen Institute Visual Coding Dataset is freely available online, via the Allen Institute. All code and files needed to reproduce final analyses will be published after revisions. Supplementary Figures Download figure Open in new tab Figure S1. Multi-level predictability analysis for all visual cortical areas. Same as red line in Figure 3 but for all cortical areas. Interestingly, all visual areas show a negative relationship, where sensitivity is highest to the lower-level features. However, in hierarchically higher cortical areas the relationship is more shallow, with less of a preference for lower level features, and lower encoding performance and more variance across units and animals. Dots show mean across animal-level averages; error bars show bootstrapped 95% confidence intervals around the mean. Download figure Open in new tab Figure S2. Multi-level feautre sensitivity analysis for all visual cortical areas. Same as blue line in Figure 3 but for all cortical areas. Interestingly, all visual areas show a similar positive relationship, where sensitivity is highest to the predictability of higher-level features, and lowest for to the predictability of lower-level features. However, in hierarchically higher cortical areas (such as VISam), the relationship appears steeper, so there is a stronger preference for higher-over-lower level features. Dots show mean across animal-level averages; error bars show bootstrapped 95% confidence intervals around the mean. Download figure Open in new tab Figure S3. Multi-level feature sensitivity analysis based on Δ R . a) V1. Same as red line in 3c, but using the cross-validated Δ R instead of the time-averaged coefficient as a metric of interest. b) Δ R -based unpredictability tuning analysis for all visual areas of interest. X-ticks indicate cortical area, colour indicates feature space or unpredictability level. Just like in the coefficient-based analysis ( Figures 3,S 1) we see the same positive relationship in all areas, where the unpredictability sensitivity highest for high-level unpredictability, and lowest for low-level unpredictability. Dots show mean across animal-level averages; error bars show bootstrapped 95% confidence intervals around the mean. Download figure Open in new tab Figure S4. Laminar comparison of overall image patch predictability across cortical areas. Same as Figure 4d , but for all cortical areas: lines with shaded error bars show the coefficient of overall image patch predictability over time, split out for units classified as ‘superficial’ or ‘deep’. Note that this sub-splitting of already sub-selected units reduces the number of neurons, and the number of animals with enough units in both categories. References ↵ Adelson , E. H. and Bergen , J. R. ( 1985 ). Spatiotemporal energy models for the perception of motion . Journal of the Optical Society of America. A, Optics and Image Science , 2 ( 2 ): 284 – 299 . OpenUrl CrossRef PubMed Web of Science ↵ Angelucci , A. , Bijanzadeh , M. , Nurminen , L. , Federer , F. , Merlin , S. , and Bressloff , P. C. ( 2017 ). Circuits and mechanisms for surround modulation in visual cortex . Annual review of neuroscience , 40 ( 1 ): 425 – 451 . OpenUrl CrossRef PubMed ↵ Bar , M. ( 2004 ). Visual objects in context . Nature Reviews Neuroscience , 5 ( 8 ): 617 – 629 . OpenUrl CrossRef PubMed Web of Science ↵ Bastos , A. M. , Usrey , W. M. , Adams , R. A. , Mangun , G. R. , Fries , P. , and Friston , K. J. ( 2012 ). Canonical microcircuits for predictive coding . Neuron , 76 ( 4 ): 695 – 711 . OpenUrl CrossRef PubMed Web of Science ↵ Biederman , I. , Mezzanotte , R. J. , and Rabinowitz , J. C. ( 1982 ). Scene perception: Detecting and judging objects undergoing relational violations . Cognitive psychology , 14 ( 2 ): 143 – 177 . OpenUrl CrossRef PubMed Web of Science ↵ Bizeul , A. , Sutter , T. , Ryser , A. , Schölkopf , B. , von Kügelgen , J. , and Vogt , J. E. ( 2025 ). From pixels to components: Eigenvector masking for visual representation learning . arXiv preprint arXiv:2502.06314. ↵ Coen-Cagli , R. , Kohn , A. , and Schwartz , O. ( 2015 ). Flexible gating of contextual influences in natural vision . Nature Neuroscience , 18 ( 11 ): 1648 – 1655 . OpenUrl CrossRef PubMed ↵ Darcet , T. , Baldassarre , F. , Oquab , M. , Mairal , J. , and Bojanowski , P. ( 2025 ). Cluster and predict latents patches for improved masked image modeling . arXiv preprint arXiv:2502.08769. ↵ de Lange , F. P. , Heilbron , M. , and Kok , P. ( 2018 ). How Do Expectations Shape Perception? Trends in Cognitive Sciences , 22 ( 9 ): 764 – 779 . OpenUrl CrossRef PubMed ↵ de Lange , F. P. , Schmitt , L.-M. , and Heilbron , M. ( 2022 ). Reconstructing the predictive architecture of the mind and brain . Trends in Cognitive Sciences . ↵ Devlin , J. , Chang , M.-W. , Lee , K. , and Toutanova , K. ( 2019 ). Bert: Pre-training of deep bidirectional transformers for language understanding . In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers), pages 4171 – 4186 . ↵ Friston , K. J. ( 2005 ). A theory of cortical responses . Philosophical Transactions of the Royal Society of London. Series B, Biological Sciences , 360 ( 1456 ): 815 – 836 . OpenUrl CrossRef PubMed ↵ Fritsche , M. , Solomon , S. G. , and de Lange , F. P. ( 2022 ). Brief stimuli cast a persistent long-term trace in visual cortex . Journal of Neuroscience , 42 ( 10 ): 1999 – 2010 . OpenUrl Abstract / FREE Full Text ↵ Fu , J. , Shrinivasan , S. , Baroni , L. , Ding , Z. , Fahey , P. G. , Pierzchlewicz , P. A. , Ponder , K. , Froebe , R. , Ntanavara , L. , Muhammad , T. , Willeke , K. F. , Wang , E. , Ding , Z. , Tran , D. T. , Papadopoulos , S. , Patel , S. , Reimer , J. , Ecker , A. S. , Pitkow , X. , Antolik , J. , Sinz , F. H. , Häfner , R. M. , Tolias , A. S. , and Franke , K. ( 2024 ). Pattern completion and disruption characterize contextual modulation in the visual cortex . ↵ Furutachi , S. , Franklin , A. D. , Aldea , A. M. , Mrsic-Flogel , T. D. , and Hofer , S. B. ( 2024 ). Cooperative thalamocortical circuit mechanism for sensory prediction errors . Nature , 633 ( 8029 ): 398 – 406 . OpenUrl CrossRef PubMed ↵ Garrett , M. , Manavi , S. , Roll , K. , Ollerenshaw , D. R. , Groblewski , P. A. , Ponvert , N. D. , Kiggins , J. T. , Casal , L. , Mace , K. , Williford , A. , et al. ( 2020 ). Experience shapes activity dynamics and stimulus coding of vip inhibitory cells . elife , 9 : e50340 . OpenUrl CrossRef PubMed ↵ Ghebreab , S. , Scholte , S. , Lamme , V. , and Smeulders , A. ( 2009 ). A Biologically Plausible Model for Rapid Natural Scene Identification . In Advances in Neural Information Processing Systems , volume 22 . Curran Associates, Inc . ↵ Groen , I. I. A. , Ghebreab , S. , Prins , H. , Lamme , V. A. F. , and Scholte , H. S. ( 2013 ). From Image Statistics to Scene Gist: Evoked Neural Activity Reveals Transition from Low-Level Natural Image Structure to Scene Category . Journal of Neuroscience , 33 ( 48 ): 18814 – 18824 . OpenUrl Abstract / FREE Full Text ↵ He , K. , Chen , X. , Xie , S. , Li , Y. , Dollár , P. , and Girshick , R. ( 2021 ). Masked Autoencoders Are Scalable Vision Learners . ↵ Heilbron , M. , Armeni , K. , Schoffelen , J.-M. , Hagoort , P. , and de Lange , F. P. ( 2022 ). A hierarchy of linguistic predictions during natural language comprehension . Proceedings of the National Academy of Sciences , 119 ( 32 ): e2201968119 . OpenUrl CrossRef PubMed ↵ Heilbron , M. and Chait , M. ( 2018 ). Great expectations: Is there evidence for predictive coding in auditory cortex? Neuroscience . ↵ Homann , J. , Koay , S. A. , Chen , K. S. , Tank , D. W. , and Berry , M. J. ( 2022 ). Novel stimuli evoke excess activity in the mouse primary visual cortex . Proceedings of the National Academy of Sciences , 119 ( 5 ): e2108882119 . OpenUrl Abstract / FREE Full Text ↵ Hubel , D. H. and Wiesel , T. N. ( 1965 ). Receptive fields and functional architecture in two nonstriate visual areas (18 and 19) of the cat . Journal of neurophysiology . ↵ Jordan , R. and Keller , G. B. ( 2020 ). Opposing influence of top-down and bottom-up input on excitatory layer 2/3 neurons in mouse primary visual cortex . Neuron , 108 ( 6 ): 1194 – 1206 . OpenUrl CrossRef PubMed ↵ Kay , K. N. , Winawer , J. , Mezer , A. , and Wandell , B. A. ( 2013 ). Compressive spatial summation in human visual cortex . Journal of Neurophysiology , 110 ( 2 ): 481 – 494 . OpenUrl CrossRef PubMed Web of Science ↵ Keller , G. B. and Mrsic-Flogel , T. D. ( 2018 ). Predictive Processing: A Canonical Cortical Computation . Neuron , 100 ( 2 ): 424 – 435 . OpenUrl CrossRef PubMed ↵ Knierim , J. J. and Van Essen , D. C. ( 1992 ). Neuronal responses to static texture patterns in area v1 of the alert macaque monkey . Journal of neurophysiology , 67 ( 4 ): 961 – 980 . OpenUrl CrossRef PubMed Web of Science ↵ Krizhevsky , A. , Sutskever , I. , and Hinton , G. E. ( 2012 ). ImageNet Classification with Deep Convolutional Neural Networks . In Advances in Neural Information Processing Systems , volume 25 . Curran Associates, Inc . ↵ Lamme , V. A. F. and Roelfsema , P. R. ( 2000 ). The distinct modes of vision offered by feedforward and recurrent processing . Trends in Neurosciences , 23 ( 11 ): 571 – 579 . OpenUrl CrossRef PubMed Web of Science ↵ LeCun , Y. ( 2022 ). A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27 . Open Review , 62 ( 1 ): 1 – 62 . OpenUrl ↵ Lee , T. S. and Mumford , D. ( 2003 ). Hierarchical bayesian inference in the visual cortex . Journal of the Optical Society of America A , 20 ( 7 ): 1434 – 1448 . OpenUrl CrossRef ↵ Liu , G. , Reda , F. A. , Shih , K. J. , Wang , T.-C. , Tao , A. , and Catanzaro , B. ( 2018 ). Image Inpainting for Irregular Holes Using Partial Convolutions . ↵ Meyer , T. and Olson , C. R. ( 2011 ). Statistical learning of visual transitions in monkey inferotemporal cortex . Proceedings of the National Academy of Sciences , 108 ( 48 ): 19401 – 19406 . OpenUrl Abstract / FREE Full Text ↵ Nayebi , A. , Kong , N. C. L. , Zhuang , C. , Gardner , J. L. , Norcia , A. M. , and Yamins , D. L. K. ( 2021 ). Shallow Unsupervised Models Best Predict Neural Responses in Mouse Visual Cortex . ↵ Oliva , A. and Torralba , A. ( 2007 ). The role of context in object recognition . Trends in cognitive sciences , 11 ( 12 ): 520 – 527 . OpenUrl CrossRef PubMed Web of Science ↵ Parras , G. G. , Nieto-Diego , J. , Carbajal , G. V. , Valdés-Baizabal , C. , Escera , C. , and Malmierca , M. S. ( 2017 ). Neurons along the auditory pathway exhibit a hierarchical organization of prediction error . Nature Communications , 8 ( 1 ): 2148 . OpenUrl CrossRef PubMed ↵ Pathak , D. , Krahenbuhl , P. , Donahue , J. , Darrell , T. , and Efros , A. A. ( 2016 ). Context encoders: Feature learning by inpainting . In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2536 – 2544 . ↵ Peelen , M. V. , Berlot , E. , and de Lange , F. P. ( 2024 ). Predictive processing of scenes and objects . Nature Reviews Psychology , 3 ( 1 ): 13 – 26 . OpenUrl PubMed ↵ Portilla , J. and Simoncelli , E. P. ( 2000 ). A parametric texture model based on joint statistics of complex wavelet coefficients . International journal of computer vision , 40 : 49 – 70 . OpenUrl CrossRef ↵ Radford , A. , Wu , J. , Child , R. , Luan , D. , Amodei , D. , Sutskever , I. , et al. ( 2019 ). Language models are unsupervised multitask learners . OpenAI blog , 1 ( 8 ): 9 . OpenUrl ↵ Rao , R. P. N. ( 2024 ). A sensory–motor theory of the neocortex . Nature Neuroscience , 27 ( 7 ): 1221 – 1235 . OpenUrl CrossRef PubMed ↵ Rao , R. P. N. and Ballard , D. H. ( 1999 ). Predictive coding in the visual cortex: A functional interpretation of some extra-classical receptive-field effects . Nature Neuroscience , 2 ( 1 ): 79 – 87 . OpenUrl CrossRef PubMed Web of Science ↵ Richter , D. , Ekman , M. , and de Lange , F. P. ( 2018 ). Suppressed Sensory Response to Predictable Object Stimuli throughout the Ventral Visual Stream . Journal of Neuroscience , 38 ( 34 ): 7452 – 7461 . OpenUrl Abstract / FREE Full Text ↵ Richter , D. , Heilbron , M. , and de Lange , F. P. ( 2022 ). Dampened sensory representations for expected input across the ventral visual stream . Oxford Open Neuroscience , 1 : kvac013 . OpenUrl ↵ Richter , D. , Kietzmann , T. C. , and de Lange , F. P. ( 2024 ). High-level visual prediction errors in early visual cortex . PLOS Biology , 22 ( 11 ): e3002829 . OpenUrl CrossRef PubMed ↵ Ringach , D. L. , Shapley , R. M. , and Hawken , M. J. ( 2002 ). Orientation selectivity in macaque v1: diversity and laminar dependence . Journal of neuroscience , 22 ( 13 ): 5639 – 5651 . OpenUrl Abstract / FREE Full Text ↵ Rousselet , G. , Pernet , D. C. , and Wilcox , R. R. ( 2019 ). An introduction to the bootstrap: A versatile method to make inferences by using data-driven simulations . ↵ Seignette , K. , de Kraker , L. , Papale , P. , Petro , L. S. , Hobo , B. , Montijn , J. S. , Self , M. W. , Larkum , M. E. , Roelfsema , P. R. , Muckli , L. , and Levelt , C. N. ( 2024 ). Experience-dependent predictions of feedforward and contextual information in mouse visual cortex . ↵ Self , M. W. , Lorteije , J. A. , Vangeneugden , J. , van Beest , E. H. , Grigore , M. E. , Levelt , C. N. , Heimel , J. A. , and Roelfsema , P. R. ( 2014 ). Orientation-tuned surround suppression in mouse visual cortex . Journal of Neuroscience , 34 ( 28 ): 9290 – 9304 . OpenUrl Abstract / FREE Full Text ↵ Shi , Y. , Siddharth , N. , Torr , P. , and Kosiorek , A. R. ( 2022 ). Adversarial masking for self-supervised learning . In International Conference on Machine Learning, pages 20026 – 20040 . PMLR . ↵ Siegle , J. H. , Jia , X. , Durand , S. , Gale , S. , Bennett , C. , Graddis , N. , Heller , G. , Ramirez , T. K. , Choi , H. , Luviano , J. A. , Groblewski , P. A. , Ahmed , R. , Arkhipov , A. , Bernard , A. , Billeh , Y. N. , Brown , D. , Buice , M. A. , Cain , N. , Caldejon , S. , Casal , L. , Cho , A. , Chvilicek , M. , Cox , T. C. , Dai , K. , Denman , D. J. , de Vries , S. E. J. , Dietzman , R. , Esposito , L. , Farrell , C. , Feng , D. , Galbraith , J. , Garrett , M. , Gelfand , E. C. , Hancock , N. , Harris , J. A. , Howard , R. , Hu , B. , Hytnen , R. , Iyer , R. , Jessett , E. , Johnson , K. , Kato , I. , Kiggins , J. , Lambert , S. , Lecoq , J. , Ledochowitsch , P. , Lee , J. H. , Leon , A. , Li , Y. , Liang , E. , Long , F. , Mace , K. , Melchior , J. , Millman , D. , Mollenkopf , T. , Nayan , C. , Ng , L. , Ngo , K. , Nguyen , T. , Nicovich , P. R. , North , K. , Ocker , G. K. , Ollerenshaw , D. , Oliver , M. , Pachitariu , M. , Perkins , J. , Reding , M. , Reid , D. , Robertson , M. , Ronellenfitch , K. , Seid , S. , Slaughterbeck , C. , Stoecklin , M. , Sullivan , D. , Sutton , B. , Swapp , J. , Thompson , C. , Turner , K. , Wakeman , W. , Whitesell , J. D. , Williams , D. , Williford , A. , Young , R. , Zeng , H. , Naylor , S. , Phillips , J. W. , Reid , R. C. , Mihalas , S. , Olsen , S. R. , and Koch , C. ( 2021 ). Survey of spiking in the mouse visual system reveals functional hierarchy . Nature . ↵ Thomas , E. R. , Haarsma , J. , Nicholson , J. , Yon , D. , Kok , P. , and Press , C. ( 2024 ). Predictions and errors are distinctly represented across v1 layers . Current Biology , 34 ( 10 ): 2265 – 2271 . OpenUrl CrossRef PubMed ↵ Uran , C. , Peter , A. , Lazar , A. , Barnes , W. , Klon-Lipok , J. , Shapcott , K. A. , Roese , R. , Fries , P. , Singer , W. , and Vinck , M. ( 2022 ). Predictive coding of natural images by V1 firing rates and rhythmic synchronization . Neuron , 110 ( 7 ): 1240 – 1257.e8 . OpenUrl CrossRef PubMed ↵ van den Oord , A. , Li , Y. , and Vinyals , O. ( 2019 ). Representation Learning with Contrastive Predictive Coding . arXiv :1807.03748 [cs, stat]. ↵ Van Hateren , J. H. and van der Schaaf , A. ( 1998 ). Independent component filters of natural images compared with simple cells in primary visual cortex . Proceedings of the Royal Society of London. Series B: Biological Sciences , 265 ( 1394 ): 359 – 366 . OpenUrl CrossRef PubMed Web of Science ↵ Vincent , P. , Larochelle , H. , Bengio , Y. , and Manzagol , P.-A. ( 2008 ). Extracting and composing robust features with denoising autoencoders . In Proceedings of the 25th international conference on Machine learning, pages 1096 – 1103 . ↵ Zhang , R. , Isola , P. , Efros , A. A. , Shechtman , E. , and Wang , O. ( 2018 ). The unreasonable effectiveness of deep features as a perceptual metric . In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 586 – 595 . ↵ Zhou , B. , Lapedriza , A. , Torralba , A. , and Oliva , A. ( 2017 ). Places: An image database for deep scene understanding . Journal of Vision , 17 ( 10 ): 296 – 296 . OpenUrl CrossRef View the discussion thread. Back to top Previous Next Posted May 19, 2025. Download PDF Email Thank you for your interest in spreading the word about bioRxiv. NOTE: Your email address is requested solely to identify you as the sender of this article. Your Email * Your Name * Send To * Enter multiple addresses on separate lines or separate them with commas. You are going to email the following Higher-level spatial prediction in natural vision across mouse visual cortex Message Subject (Your Name) has forwarded a page to you from bioRxiv Message Body (Your Name) thought you would like to see this page from the bioRxiv website. Your Personal Message CAPTCHA This question is for testing whether or not you are a human visitor and to prevent automated spam submissions. Share Higher-level spatial prediction in natural vision across mouse visual cortex Micha Heilbron , Floris P de Lange bioRxiv 2025.05.15.654212; doi: https://doi.org/10.1101/2025.05.15.654212 Share This Article: Copy Citation Tools Higher-level spatial prediction in natural vision across mouse visual cortex Micha Heilbron , Floris P de Lange bioRxiv 2025.05.15.654212; doi: https://doi.org/10.1101/2025.05.15.654212 Citation Manager Formats BibTeX Bookends EasyBib EndNote (tagged) EndNote 8 (xml) Medlars Mendeley Papers RefWorks Tagged Ref Manager RIS Zotero Tweet Widget Facebook Like Google Plus One Subject Area Neuroscience Subject Areas All Articles Animal Behavior and Cognition (7635) Biochemistry (17690) Bioengineering (13892) Bioinformatics (41936) Biophysics (21451) Cancer Biology (18588) Cell Biology (25499) Clinical Trials (138) Developmental Biology (13378) Ecology (19899) Epidemiology (2067) Evolutionary Biology (24320) Genetics (15609) Genomics (22506) Immunology (17736) Microbiology (40394) Molecular Biology (17181) Neuroscience (88603) Paleontology (666) Pathology (2832) Pharmacology and Toxicology (4824) Physiology (7641) Plant Biology (15152) Scientific Communication and Education (2045) Synthetic Biology (4294) Systems Biology (9825) Zoology (2271)
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.