(C) PLOS One This story was originally published by PLOS One and is unaltered. . . . . . . . . . . High-level visual prediction errors in early visual cortex [1] ['David Richter', 'Donders Institute For Brain', 'Cognition', 'Behaviour', 'Radboud University Nijmegen', 'Nijmegen', 'The Netherlands', 'Mind', 'Brain', 'Behavior Research Center'] Date: 2024-11 Perception is shaped by both incoming sensory input and expectations derived from our prior knowledge. Numerous studies have shown stronger neural activity for surprising inputs, suggestive of predictive processing. However, it is largely unclear what predictions are made across the cortical hierarchy, and therefore what kind of surprise drives this up-regulation of activity. Here, we leveraged fMRI in human volunteers and deep neural network (DNN) models to arbitrate between 2 hypotheses: prediction errors may signal a local mismatch between input and expectation at each level of the cortical hierarchy, or prediction errors may be computed at higher levels and the resulting surprise signal is broadcast to earlier areas in the cortical hierarchy. Our results align with the latter hypothesis. Prediction errors in both low- and high-level visual cortex responded to high-level, but not low-level, visual surprise. This scaling with high-level surprise in early visual cortex strongly diverged from feedforward tuning. Combined, our results suggest that high-level predictions constrain sensory processing in earlier areas, thereby aiding perceptual inference. Given that prediction error computation is a core mechanism of predictive processing, it is crucial to characterize what kind of visual surprise is tracked by the visual system. Here, we aimed to close this gap by exploring features reflected in the visual surprise response after statistical learning. Specifically, we asked (1) whether prediction errors come to reflect any visual feature tuning in predictive contexts, and if so, whether (2) this tuning is in line with the local visual features or inherited top-down. To do so, we exposed human volunteers to images that were either expected or unexpected in terms of their identity, given a preceding cue, while recording whole-brain fMRI. To quantify visual feature surprise across multiple levels of description, we used representational dissimilarity metrics derived from a visual deep neural network (DNN). Our results demonstrate that neural responses, and the fidelity of visual representations, across multiple visual cortical areas monotonically increased with how visually dissimilar a surprising object was compared to the expected object. Crucially, high-level visual dissimilarity accounted for the surprise induced increase of neural responses, including in the earliest visual cortical area, V1. Prediction errors thus appear to reflect surprise primarily in terms of high-level visual features, demonstrating that earlier visual areas inherit feature tuning usually associated with higher visual areas in predictive contexts, presumably due to feedback signals. In contrast to local feature tuning, prediction error tuning may be inherited top-down. Top-down inheritance is in line with hierarchical predictive processing, because predictions are proposed to be relayed top-down from higher to lower visual areas [ 2 ], and thus lower visual areas may come to reflect tuning properties of higher areas in predictive contexts due to the top-down prediction signals. Empirical support for this notion has been obtained in the macaque face processing system. For example, Schwiedrzik and Freiwald [ 25 ] showed increased neural activity in ML, a lower-level area in the macaque face processing hierarchy that is not tuned to identity, when monkeys viewed face stimuli that were surprising in terms of identity. These results could be explained by feedback signals from higher-order identity-tuned cells in the inferior temporal cortex. Whether a similar principle of top-down prediction error tuning inheritance applies across species, and to stimuli outside the narrow domain of face processing, and across the visual hierarchy remains unknown. If prediction and prediction error computations underlie perceptual inference, as suggested by predictive processing theories, we can stipulate that cortical predictions and the associated prediction error signatures must reflect stimulus features that are represented in the respective cortical area. For example, prediction errors in primary visual cortex (V1) may signal deviations from expectation in terms of simple features such as stimulus orientation, edges, and contrasts—i.e., visual features that V1 neurons are tuned to [ 19 ]. On the other hand, prediction errors in higher visual areas (HVC), for instance in fusiform gyrus, may reflect more complex high-level visual features, such as object identities, spatial relationships between object parts and more abstract concepts such as faces, commonly represented in those areas [ 20 – 22 ]. This account suggests that prediction errors mirror local feature tuning, unique to each visual cortical area. While some studies have provided indirect support for the feature specificity of sensory prediction errors by investigating tuning specific modulations [ 9 , 23 , 24 ], little evidence directly shows which visual feature surprise, if in fact any, is reflected in visual prediction errors. Predictive processing theories promise to provide a principled account of cortical computation [ 1 – 4 ]. One critical ingredient of predictive processing is the computation of prediction errors, i.e., the mismatch between prediction, usually thought of as a top-down signal, and bottom-up input. Such prediction errors then serve as input to the next level in the cortical hierarchy. The brain minimizes prediction errors by recurrently updating its predictions. This process enables the formation of a coherent, stable, and efficient representation of the world. Despite variations in specific predictive processing implementations [ 1 – 6 ], the core concept of prediction error computation is ubiquitous and supported by many empirical observations. For example, after visual statistical learning, visual cortex is sensitive to the likelihood of an object’s appearance. In particular, activity throughout the ventral visual stream has been shown to be attenuated to expected compared to unexpected appearances of the same stimuli [ 7 – 11 ]. This attenuation, also known as expectation suppression, has been observed across different species and modalities [ 12 – 14 ] and occurs also when predictions and stimuli are task-irrelevant [ 15 – 18 ]. Combined, expectation suppression has frequently been interpreted in the context of predictive processing as reflecting larger prediction errors for unexpected stimuli, and thus taken as crucial evidence that perception fundamentally relies on prediction [ 12 , 14 ]. An ROI analysis ( Fig 6B ) of the same 4 contrasts confirmed the whole-brain results. We observed reliable differences between models, which differed across ROIs (main effect of model: F (4,128) = 12.20, p < 0.001, = 0.28; interaction ROI by model: F (4.40,140.93) = 11.09, p < 0.001, = 0.26). Specifically, we found significantly stronger modulations of BOLD responses by high-level visual dissimilarity compared to all other 4 parametric modulators in V1 (paired t tests: all FDR corrected p < 0.001; all d > 0.99; see S3 and S4 Tables for details). Similar, albeit less pronounced results were found in LOC (Layer 8 versus Layer 2: p = 0.001, d z = 0.74; Layer 8 versus Animacy category: p = 0.010, d z = 0.61; Layer 8 versus Word2Vec: p = 0.060, d z = 0.47; Layer 8 versus Random layer 8: p = 0.078, d z = 0.44) and HVC (Layer 8 versus Layer 2: p = 0.028, d z = 0.50; Layer 8 versus Animacy category: p = 0.081, d z = 0.40; Layer 8 versus Word2Vec: p = 0.016, d z = 0.54; Layer 8 versus Random layer 8: p = 0.189, d z = 0.32). An additional negative modulation of prediction error magnitudes by word category surprise was observed in V1 (p = 0.020, d z = −0.60), and by animacy category surprise in LOC (p = 0.039, d z = −0.51), suggesting that prediction errors may be attenuated for more semantically dissimilar surprising images in V1 and for stimuli of a difference animacy category in LOC compared to the expected image. Finally, to ensure that our results were not dependent on the exact ROI mask definition and mask sizes, we repeated the analysis for an alternative ROI mask definition (using all stimulus driven voxels within the anatomically defined masks) and across multiple mask sizes. Results were qualitatively identical across ROI definitions ( S4 Fig ) and mask sizes ( S5 Fig ). (A) Whole-brain contrasts of the high-level visual feature model (layer 8) contrasted against 4 control variables. The top row shows that high-level visual models performed significantly better than low-level visual models (layer 2). Similarly, high-level visual surprise better accounted for prediction error magnitudes than the task-relevant animacy category of the unexpected stimuli (second row) and the semantic, word category surprise model (word2vec; third row). The bottom row shows that high-level visual dissimilarity significantly better explained prediction error magnitudes compared to an untrained but otherwise identical DNN layer 8. (B) ROI analysis including primary (V1), intermediate (LOC), and high-level visual cortex (HVC). Results confirm the whole-brain results, showing significant modulations of BOLD responses by high-level visual surprise (red) compared to low-level visual (blue), response category (green), and word category surprise (purple). Error bars indicate the 95% within-subject confidence intervals. Gray dots denote individual subjects. P values are FDR corrected. *** p < 0.001, ** p < 0.01, * p < 0.05. Data and code that support these findings are available at: https://doi.org/10.34973/8e49-2012 . DNN, deep neural network; HVC, higher visual cortex; LOC, lateral occipital complex; ROI, region of interest; V1, primary visual cortex. These controls revealed that high-level visual surprise (layer 8) best accounted for the data, significantly outperforming an untrained random layer 8 model, animacy category, word category (semantic), and the low-level visual surprise models in explaining prediction error magnitudes ( Fig 6A ). Statistically significant clusters were found in EVC as well as intermediate visual areas in LOC, and HVC in some contrasts. The exact extent of the modulation varies slightly between contrasts, but overall corroborate that prediction error magnitudes mainly result from high-level visual feature surprise and that none of the control variables likely account for the observed results. Corresponding whole-brain figures contrasting the control parametric modulators against baseline (no modulation) can be found in S3 Fig . Additionally, variance inflation factors (VIFs; S2 Table ) were smaller for layer 2 (VIF = 1.42) than layer 8 surprise (VIF = 1.88), suggesting that the absence of a significant modulation by low-level surprise was not due to problems with variance partitioning due to collinearity of the predictors in the GLM. In fact, all VIFs were significantly lower than the suggested threshold of VIF < 5 [ 35 ]. Multiple alternative explanations could account for a correlation of prediction error magnitudes with high-level visual features. To rule out alternative accounts for our observations, we included multiple control variables as parametric modulators in our GLM analysis. First, layer 8 representations from an untrained (i.e., random) but otherwise identical DNN were included. This contrast ruled out that the inherent structure of the DNN architecture or correlations in the input images caused the scaling of prediction errors with late layer representations. Second, animacy category was included in the GLM to assess whether high-level visual modulations result from a correlation of high-level visual features with the task-relevant dimension of animacy. Since participants had to distinguish between animate and inanimate entities in the images, prediction error modulations could potentially reflect task responses. Finally, because high-level visual features may significantly correlate with semantic, word category level surprise, we also contrasted the high-level visual model against a word2vec [ 33 , 34 ] derived model aimed at indexing nonvisual, semantic surprise. In addition to modulating BOLD responses, it is possible that high-level surprise may also result in sharper visual representations. To test this, we performed a decoding analysis, trained on the independent localizer data and tested on the main task data. Decoding accuracy of the object images (for details, see: Materials and methods : Statistical Analysis: Decoding as function of surprise) increased with layer 8 dissimilarity in V1 (t (32) = 5.50, p < 0.001, d z = 0.96, mean r = 0.08) and HVC (t (32) = 3.89, p < 0.001, d z = 0.68, mean r = 0.05), but not LOC (t (32) = 0.96, p = 0.342, d z = 0.17, mean r = 0.07). In other words, the more surprising an unexpected stimulus was in terms of high-level features, the better it could be decoded from V1 and HVC. To summarize, visual responses and the fidelity of the visual representations appeared to monotonically increase with high-level visual surprise across major parts of the ventral visual system. Next, we assessed the shape of the response modulation by high-level visual surprise by regressing BOLD responses for unexpected stimuli onto layer 8 dissimilarity using an ROI approach. Results, shown in Fig 5D , show that the increased BOLD response to surprising stimuli follows a positive monotonic association across the dissimilarity spectrum in all 3 ROIs (V1: t (32) = 9.35, p < 0.001, d z = 1.63, mean r = 0.09; LOC: t (32) = 3.27, p = 0.004, d z = 0.57, mean r = 0.03; HVC: t (32) = 2.05, p = 0.049, d z = 0.36, mean r = 0.02). In a subsequent analysis, we asked which DNN layer explained most neural variance of the prediction error response. We analyzed how neural responses to unexpected stimuli were scaled as a function of surprise indexed by each layer of the DNN. To this end, we regressed layers 1 to 8 dissimilarity onto single trial parameter estimates and determined for each voxel which layer had the largest explained variance. Results ( Fig 5C ) showed that prediction error magnitudes primarily scaled with high-level visual surprise (layer 8 and layer 7) across most parts of the ventral visual stream, including EVC, LOC, and HVC. We note additional minor clusters in HVC scaled by intermediate layer 4, as well as layers 1 and 3 in EVC and LOC, suggesting that some neural populations scaled with intermediate and low-level surprise as well. In sum, these results present a stark contrast to the modulation of responses during the prediction-free localizer ( Fig 3A ) where a clear gradient from early-to-late layers was observed, with early layers dominating EVC representations. In contrast, prediction errors appear to be scaled preferentially by high-level visual surprise across the visual system, including EVC. (A) Whole-brain results assessing the modulation of surprise responses as a function of high-level (top row) and low-level (bottom row) visual feature dissimilarity. The top row shows that surprise responses to unexpected images were increased if the image was more distant from the expected image in terms of high-level visual features. Color indicates the beta parameter estimate of the parametric modulation, with red and yellow representing increased responses. Black outlines denote statistically significant clusters (GRF cluster corrected). No significant modulation of sensory responses was observed by low-level visual surprise. (B) ROI analysis zooming in on ROIs in early visual (V1), intermediate (LOC), and HVC (encompassing occipito-temporal sulcus and fusiform cortex). Results mirror those of the whole-brain analysis, with significant modulations of the visual responses by high-level visual surprise (red), but not low-level visual surprise (blue). Error bars indicate the 95% within-subject confidence intervals. Gray dots denote individual subjects. P values are FDR corrected. *** p < 0.001, ** p < 0.01, * p < 0.05, = BF 10 < 1/3. (C) Prediction errors preferentially scale with high-level visual features (layer 8 and layer 7) throughout most of the visual system, including EVC, LOC and HVC. Color indicates the DNN layer with the largest effect (explained variance) on scaling the neural responses to surprising inputs. Cold colors (purple–blue) represent early layers (i.e., low-level visual features), while warm colors (yellow–red) indicate late layers (i.e., high-level visual features). Analysis was masked to visual cortex and thresholded at a liberal z ≥ 1.96 (i.e., p < 0.05, two-sided) to explore the landscape of prediction error modulations across DNN layers. Results strongly contrast with those observed for prediction-free visual responses during the localizer ( Fig 3A ). (D, E) ROI analysis regressing BOLD responses ( D ) or decoded true class probability ( E ) onto high-level visual dissimilarity. Results indicate a monotonic relationship between high-level surprise and BOLD responses across all 3 ROIs, as well as decoding performance in V1 and HVC. The chance level for decoding the true class probability is 0.125. For display purposes dissimilarities were ranked and averaged across participants, while regression models were fit per participant on the correlation distances. Data and code that support these findings are available at: https://doi.org/10.34973/8e49-2012 . DNN, deep neural network; EVC, early visual cortex; HVC, higher visual cortex; LOC, lateral occipital complex; ROI, region of interest; V1, primary visual cortex. Our whole-brain results were corroborated by a region of interest (ROI) analysis, depicted in Fig 5B (ROI masks are illustrated in Fig 3B ). Results showed a strong difference between high- and low-level visual feature models in modulating BOLD surprise responses (main effect of model: F (1,32) = 42.78, p < 0.001, = 0.57). We found reliable modulations of surprise responses by high-level visual, but no significant modulation by low-level visual features, in primary visual cortex (V1: Layer 8: t [ 32 ] = 6.79, p < 0.001, d z = 1.18; Layer 2: W = 190, p = 0.159, d z = −0.32, BF 10 = 0.58), intermediate visual areas in the lateral occipital complex (LOC: Layer 8: W = 126, p = 0.017, d z = 0.55; Layer 2: t (32) = −0.68, p = 0.504, d z = −0.12, BF 10 = 0.23) and high-level visual cortex (HVC: Layer 8: t (32) = 2.59, p = 0.029, d z = 0.45; Layer 2: t (32) = −0.70, p = 0.586, d z = −0.12, BF 10 = 0.23). Contrasting the modulation by layer 8 against layer 2 dissimilarity confirmed that layer 8 modulated visual surprise responses significantly more than layer 2 in V1 (t (32) = 8.05, p < 0.001, d z = 1.40), LOC (t (32) = 4.24, p < 0.001, d z = 0.74), and HVC (t (32) = 2.85, p = 0.008, d z = 0.50). Hence, the larger the visual dissimilarity of a surprising stimulus in terms of high-level visual features, the more vigorous the visual response across the ventral visual stream, including V1. No corresponding modulation by low-level visual features was observed, suggesting that high-level features are predominantly reflected in visual prediction error signals. A complementary control analysis ( S2 Fig ), using stimulus uninformative voxels (i.e., chance or near chance-level decoding accuracy) showed no modulation by either layer 2 or layer 8 dissimilarity, suggesting that the modulation of visual responses by layer 8 dissimilarity is specific to stimulus selective voxels. Results, depicted in Fig 5A , demonstrated that surprise responses scaled significantly with high-level visual dissimilarity (layer 8) in visual cortex, encompassing early and intermediate visual areas (cluster size 779 voxel, 6,232 mm 3 ; S1 Table contains additional details). That is, the more an unexpected stimulus diverged from the expected image in terms of high-level visual features, the more the sensory response increased in magnitude. Surprisingly, we did not find any modulation of neural responses by low-level visual dissimilarity (layer 2) anywhere in visual cortex. In other words, even in EVC prediction error magnitudes were modulated by high-level but not low-level visual surprise, as indexed by layer 2. In the example in Fig 4A , this corresponds to the unexpected image of the woman eliciting a larger prediction error in V1 compared to the unexpected guitar. On the other hand, the low-level surprise elicited by the unexpected guitar would not result in an additional up-regulation of prediction errors in visual cortex beyond the associated high-level visual surprise. (A) If you expect to see the first guitar on the left, what kind of visual features does visual cortex predict? Low-level visual features, illustrated next to the expected image, concern local oriented edges, spatial frequency, and similar properties. High-level visual features entail more complex visual representations, such as texture-like features [ 32 ], core object parts and their relationships, features commonly shared between instances of an object, irrespective of the specific depiction. Depending on which features are predicted, the 2 “seen” images will result in different prediction error magnitudes. The image of the woman is very different in high-level visual features, but shares some local orientation with the expected guitar, hence resulting primarily in high-level visual surprise. On the other hand, the image of the other guitar is very different in terms of low-level visual features, as it is differently rotated compared to the expected guitar, but it is still a guitar and thus shares high-level visual features. The key question of the analysis is whether and where in the visual system low-level or high-level surprise results in larger prediction errors. (B) Analysis procedure. The left side shows a single trial with a letter cue predicting a specific image. Below multiple unexpected images are illustrated, which were presented on other trials. The right side depicts the extraction of low-level and high-level visual feature dissimilarity from layer 2 (low-level) and layer 8 (high-level) of the visual DNN. Dissimilarity of the unexpected seen image relative to the expected stimulus was then added as a parametric modulator in the first level GLMs of the fMRI analysis. For illustration purposes, the graph at the bottom uses data from one example participant to depict how BOLD responses (example data from V1; ordinate) are modulated as a function of low-level (blue) and high-level (red) visual dissimilarity (abscissa). A positive slope thus indicates that the more dissimilar a seen image was relative to the expected stimulus the more vigorous the neural response. Dots indicating individual dissimilarities (i.e., distances of seen unexpected images compared to the expected image) are arranged in rows, reflecting the granularity of the available surprise distances for this example participant. In the example data, the image of the unexpected woman would hence result in larger prediction errors compared to the image of the unexpected guitar, because of the larger high-level surprise. Critically, this analysis only uses BOLD data from unexpected image trials. Additional control variables for task relevance (animacy) and word meaning, discussed in more detail later, were also included. We performed this parametric modulation analysis in a voxel-wise fashion across the whole brain. DNN, deep neural network; GLM, general linear model. Next, we turned our attention to the nature of prediction errors, asking which visual features, if any, they reflect across multiple regions along the ventral visual hierarchy. Fig 4 illustrates the analysis rationale. If low-level visual features, such as local oriented edges and spatial frequency, are predicted in a specific cortical area (e.g., V1) then prediction error magnitudes should scale with low-level surprise in that area. As an example, expecting a specific image of a guitar but seeing an image of another guitar from a different angle should yield a large prediction error as the 2 images are different from a low-level visual feature standpoint. On the other hand, if high-level visual features, such as more abstract and general guitar features (e.g., a neck and guitar body) are predicted, invariant to local orientation, then the unexpected guitar should not yield a strong prediction error, whereas seeing an image of an unexpected category should result in a large prediction error (even if low-level features are similar). Based on our analyses of DNN alignment with fMRI localizer data and prior work [ 27 ], we decided to use layer 2 of the visual DNN as low-level feature model and layer 8 (before softmax of the output layer) as high-level visual feature models. This allowed for maximal differentiation from the low-level model. To model how the BOLD response changed as a function of how visually surprising an unexpected stimulus was in terms of low-level (layer 2) and high-level (layer 8) visual features respectively, we used the dissimilarity of each surprising image, compared to the stimulus expected on that trial, as parametric modulators in the fMRI GLM analysis (for details, see Materials and methods : Data analysis). Using this approach, variance not uniquely associated with either regressor will not contribute to the observed results. Thus, in addition, we included multiple control variables, such as task-relevant stimulus animacy and word-level (semantic) dissimilarity. (A) RSA of visual responses shows that (feedforward) visual responses during the prediction-free localizer were best explained by a gradient of low-level to high-level visual features going up the ventral visual hierarchy. EVC responses aligned more closely with early DNN layers, indicative of low-level visual feature processing (depicted in cold colors: purple to blue). HVC areas, like the fusiform gyrus, show a greater correlation with late DNN layers, representing high-level visual feature processing (illustrated in warm colors: yellow to red). Analysis was masked to visual cortex and thresholded at z > 3.1 (p < 0.001, uncorrected) of the RSA. (B) ROI masks, depicting voxels included in the anatomically and functionally defined masks. Color indicates the ROI: Blue = V1, Purple = LOC, Orange = HVC. Opacity indicates the proportion of participants whose individual masks included the voxel. For visualization, full opacity corresponds to a proportion of 0.1, with voxel inclusion proportions reaching up to ~0.7. Data and code that support these findings are available at: https://doi.org/10.34973/8e49-2012 . DNN, deep neural network; EVC, early visual cortex; HVC, higher visual cortex; ROI, region of interest; RSA, representational similarity analysis; V1, primary visual cortex. Results ( Fig 3A ), showed a gradient from early to high visual cortex with the corresponding early to late DNN layers best explaining neural variance. Specifically, early visual cortex (EVC) responses were best explained by early layers of the DNN, intermediate visual areas in LOC by early to intermediate layers, and HVC responses, for example, in the fusiform gyrus, were best explained by intermediate to late layers of the DNN. These outcomes affirm previous reports [ 27 – 31 ] and validate that our visual feature models, including the low-level (layer 2) and high-level visual model (layer 8), accounted for cortical variance elicited by visual stimulation in an expected pattern of a low-to-high level gradient. In addition, we explored representations of layer 2 and 8 using feature visualization and decoding techniques to further support our assumption that these layer representations serve as low-level and high-level visual feature spaces respectively. For details, see S1 Fig . In brief, results showed that layer 2 represents low-level visual information, such as Gabor-like orientation features, while layer 8 has more complex texture-like [ 32 ] and categorical tuning. While (visual) DNNs have been used to successfully explain a variety of neural data [ 26 ], we first ensured that the specific feature models utilized here were able to explain visual responses in our data in a prediction-free context. Specifically, we extracted the correlation distance between images from an implementation of AlexNet trained on ecoset [ 27 ]. Then, we performed representational similarity analysis (RSA) using the representational distances derived from the DNN layers and from the fMRI data obtained during localizer runs. During these runs, each image was presented in isolation, without preceding cue and without predictive associations between the images. To assess possible contributions of all DNN layers, we repeated this analysis using representational dissimilarity matrices (RDMs) from each of the DNN layers separately. Finally, for each voxel (sphere searchlight) we determined the best DNN layer for explaining neural variance based on the RSA results (correlation coefficient). (A) RTs to expected objects were faster compared to unexpected images requiring the same or different response as the expected stimulus. (B) Accuracy of responses was high overall, but responses were more accurate to expected compared to unexpected images requiring a different response than the expected stimulus. Error bars indicate the 95% within-subject confidence intervals. *** p < 0.001, ** p < 0.01, * p < 0.05. Data and code that support these findings are available at: https://doi.org/10.34973/8e49-2012 . RT, reaction time. Participants learned and used the underlying statistical regularities to predict inputs thereby facilitating performance ( Fig 2 ), in terms of response times (RT; F (1.4,43.3) = 23.5, p < 0.001, = 0.42) and response accuracy (F (1.4,46.3) = 7.8, p = 0.003, = 0.20). Specifically, RTs to expected images (501 ms) were faster compared to unexpected images requiring the same button press (509 ms; t (32) = 2.41, p = 0.019, d z = 0.14) and different button press (524 ms; t (32) = 6.77, p < 0.001, d z = 0.39). The former contrast demonstrates that surprising stimuli resulted in slower responses even when the stimulus required the same response (i.e., was of the same animacy category) as the expected input. Additionally, adjusting responses resulted in even slower RTs (unexpected same versus unexpected different: t (32) = 4.36, p < 0.001, d z = 0.25). Response accuracy only showed a reliable decrements when unexpected stimuli required a different response, thus likely reflecting response errors due to invalid prediction (expected versus unexpected different: t (32) = 3.75, p = 0.001, d z = 0.61; expected versus unexpected same: t (32) = 0.79, p = 0.43, d z = 0.13; unexpected same versus unexpected different: t (32) = 2.97, p = 0.008, d z = 0.48). Overall behavioral facilitation thus demonstrated that participants benefitted from prediction. Moreover, accuracy in general was very high (>95%), indicating effective task compliance. (A) A single trial, showing a letter cue (500 ms) followed by an image (500 ms) and a variable ITI (approximately 5,000 ms). The image was expected or unexpected given the preceding letter cue. Participants responded by button press to the images, indicating whether the entity in the image was animate or inanimate. (B) TPM determining the associations between cues and images. Each of the 8 images was associated with one of the 8 letter cues. The expected image was 7 times more likely to appear than any other image given its cue. Numbers in each cell indicate the number of trials per run. The specific cue-image associations were randomized and differed between participants. Moreover, the set of 8 images also varied for different participants. (C) Two cycles of a localizer trial. During the localizer, one image was presented repeatedly (500 ms on, 300 ms off) for 12,000 ms. The identity of the images was not predictable. Participants responded to a high brightness version of the images, which was shown once during each trial for one cycle. ISI, interstimulus interval; ITI, intertrial interval; TPM, transitional probability matrix. Human volunteers (n = 33) viewed images that could contain either animate or inanimate entities. Each image was preceded by a letter cue that probabilistically predicted the identity of the image. A trial is depicted in Fig 1A . The expected image was seven times more likely to follow its associated letter cue compared to each of the 7 unexpected images (the transitional probability matrix (TPM) is depicted in Fig 1B ). The same images appeared both as expected and unexpected stimuli, with the expectation status only contingent on the cue after which the image appeared. Participants were tasked to classify the content of the image as animate or inanimate. Further details are in the Materials and methods: Stimuli and experimental paradigm section. Discussion Hierarchical predictive processing theories [1–4,36] have received significant attention as they propose a fundamental framework for cortical computation. Numerous studies have corroborated the main tenets of predictive processing, such as demonstrating that sensory responses to surprising inputs are enhanced compared to expected ones, likely reflecting larger sensory prediction errors [9,15–17,37]. However, while evidence for this core mechanism of predictive processing has been shown across modalities, paradigms, and species [12,14,38], it remains unknown what kind of surprise is reflected in these putative prediction errors. Here, we set out to elucidate the nature of the neural modulation by visual prediction errors and thus what information is predicted across the visual hierarchy. Visual prediction errors scale with high-level visual surprise Using fMRI and representational distance measures derived from a visual DNN, our data showed that throughout multiple visual cortical areas sensory responses to unexpected images scale with the representational distance of a seen unexpected stimulus relative to the expected input. Specifically, responses monotonically scaled with high-level visual feature surprise: the larger the high-level deviation from expectation, the larger the prediction error response. Interestingly, and in contrast to feedforward processing, even early visual areas, such as V1, predominantly responded to high-level visual over low-level surprise, while being best explained in terms of low-level features in cases where no prediction was possible (localizer). These results are in line with a recent study demonstrating that firing rates in macaque V1 correlate with spatial predictability of high-level and not low-level visual features [39]. The increased neural activity for high-level visual surprise in EVC, a region that is not tuned to high-level visual features during feedforward processing, may suggest that predictions are relayed top-down and hence result in the observed inheritance of feature surprise from higher areas in earlier processing stages. Thus, our results support and extend previous studies by demonstrating that (1) top-down inheritance during predictive vision generalizes across species and recording modalities; and (2) crucially appears to be a general principle of visual sensory processing evident across multiple cortical areas that together encompass the ventral visual stream [25,40]. Additionally, the more surprising objects were in terms of high-level features, the better they could be decoded by a linear classifier. Therefore, our data suggested that both the magnitude but also the stimulus information contained in the visual response scaled with high-level surprise. This may imply a possible functional role in enhancing visual representations of inputs that do not match our high-level predictions of the world. Moreover, our results further demonstrate that predictive signatures and top-down inheritance of high-level feature surprise can arise, at least in humans, with little exposure to the predictive regularities, requiring only several dozen exposures rather than extensive exposure as in the case of studies in nonhuman primates [25]. This flexibility of the visual system to learn and rapidly utilize novel sensory priors to generate high-level predictions to inform sensory processing in earlier stages of the hierarchy further supports the hypothesis that predictive top-down signaling is a ubiquitous and general principle of visual processing. What kind of mechanism may underlie the here observed modulations of sensory responses? Earlier neurophysiological work in macaques have found that the visual prediction error response in IT has roughly the same latency as the response to the stimulus itself [15], whereas high-level surprise affects a late stage of neural responses [25,39]. We speculate that our observed fMRI BOLD correlates reflect a similar late-stage top-down modulation of sensory processing. This interpretation aligns well with our observation that the feedforward response, mostly reflected during the prediction-free localizer, is dominated by local tuning properties (e.g., low-level features in V1), while recurrent processing due to prediction during the main task, relying on feedback and reflecting high-level visual surprise, takes time to arise given the necessary computations and signal relaying across multiple cortical areas. Taken together, our results suggest that following the activation of high-level areas, during the initial feedforward sweep, a prediction is subsequently relayed down the processing hierarchy modulating the sustained phase of neural responses across earlier areas. Elaborating further on this account, one possible explanation for the observed results is that predictions allow the visual system to settle on a valid perceptual interpretation more quickly, as predictions match the bottom-up inputs. In contrast, unexpected input may require stronger [15] and more prolonged processing within the visual system to arrive at the best interpretation of the current inputs, potentially leading to the larger BOLD signal and superior decoding of stimulus representations observed here. Put differently, our results may suggest that the larger the high-level difference between what is expected and what is observed the more cycles of recurrent processing along the ventral visual stream may be necessary to solve object recognition. To clarify, here we are not suggesting that visual processing in V1 primarily reflects high-level visual information, but rather that its modulation as a function of expectation reflects high-level surprise, potentially implemented by cortico-cortical feedback. Our analyses also demonstrated that the modulation of prediction error responses by high-level visual surprise was not explained by the task-relevant animacy category dimension or by word-level (semantic) surprise. These results thus suggest that visual prediction errors are predominantly modulated by high-level visual and not abstract linguistic or response-related surprise. This preference aligns with our earlier proposition of facilitated perceptual inference, with visual prediction errors scaling primarily with visual surprise due to the role of top-down feedback in constraining visual interpretations in lower areas. Consequently, because high-level visual areas in the ventral stream encode primarily high-level visual features, this surprise signal is expressed in terms of high-level visual features instead of abstract non-visual representations. In the whole-brain analyses (Figs 5A and 6A), we found reliable up-regulations of prediction error amplitudes by high-level visual surprise in the visual system, but not in the motor system or other areas outside the visual system, such as inferior frontal gyrus or anterior insula, known to generate prediction errors especially when predictions are task-relevant [10,41–43]. This suggests that the modulation by high-level visual surprise primarily concerns facilitated perceptual inference rather than facilitated decision-making or response initiation. Yet, this does not mean that predictions do not facilitate these processes as well. The focus of our investigation was the modulation of neural responses for different unexpected inputs. A direct comparison of unexpected compared to expected inputs (i.e., expectation suppression) can be found in S6 Fig and matches previous studies showing additional prediction error signatures in decision, attention, and motor processing-related areas [10,41–43], suggesting that prediction facilitates processing across multiple cortical systems. A supplemental analysis decoding stimulus identity is reported in S5 Table. Flexible prediction (error) tuning While we found no reliable up-regulation of prediction error magnitudes by any of the control models (also see: S3 Fig), this does not imply that the visual system exclusively encodes high-level visual surprise. Although we dismissed that animacy category explained our results, this does not rule out that task requirements shaped the acquisition and generation of predictions, and consequently the scaling of prediction errors. Indeed, it is possible that the focus on high-level features, due to the task requirements, shaped what kind of features were predicted, given the substantial effect that tasks can have on visual processing [44]. Thus, our results are also consistent with recent models of adaptive efficient coding [45]. Specifically, because our task required a high-level decision, such task demands might be reflected in the adaptive compression (silencing) of non-task relevant sensory representations due to top-down feedback, thereby resulting in the observed scaling of early visual responses with high-level visual surprise. Moreover, evolutionary it is advantageous to develop flexible predictions, allowing the visual system to adapt to environmental requirements. In line with this hypothesis, recent evidence suggests that semantic (word level) priors can be used to generate category specific sensory predictions, even in early visual cortex [46]. Therefore, more abstract, semantic representations could modulate visual prediction errors under certain conditions. Additionally, considering the nature of the utilized DNN, it is probable that our high-level visual feature model captured a wide range of category variability, with animacy representing just one such (binary) dimension. On the other hand, it is also plausible that visual prediction errors can reflect low-level features, particularly if required by the environment. Indeed, in the present data small additional clusters across multiple areas scaled with intermediate-level surprise, suggesting variety in the surprise reflected in visual prediction errors. Furthermore, employing a different task specifically targeting low-level features, or using stimuli that are predictable predominantly in terms of these features [9] might indeed lead to strong representations of low-level surprise. However, here we observe for naturalistic stimuli, which inherently support a hierarchy of predictions, that lower-level predictions may be subsumed by high-level ones, mirroring recent results in language processing [47]. In addition, it is likely that there are limits to the flexibility imposed by the architecture and representational constraints of the visual system. As argued above, high-level visual predictions may be relayed top-down and constrain visual interpretations in lower areas, because neurons in higher visual areas are tuned to high-level features. In contrast, it is less likely that high-level visual areas have the necessary representational structure and acuity to predict detailed low-level features, and hence these areas may be unable to constrain visual interpretations in EVC for low-level visual features to the same degree as for high-level predictions. Nonetheless, we believe that it is essential for future work to chart the extent to which prediction error tuning is flexibly adjusted to reflect environmental and task demands. Limitations There are some limitations to the present research. A superior low-level visual feature model may have yielded prediction error modulations by such features. In addition, the specific stimulus set could have discouraged or obscured low-level surprise modulations. Nevertheless, we demonstrated that early DNN layers best explain feedforward visual activity in EVC during the prediction-free localizer runs. This observation underscores that the low-level feature model employed in our study reliably accounted for the visual responses elicited by our stimuli in the EVC, outperforming intermediate or high-level feature models. Consequently, the absence of low-level surprise modulations in our results cannot be attributed to a general inadequacy of the low-level model. Moreover, we chose a commonly used visual DNN. This and similar models have repeatedly been shown to share representational geometry with visual cortex [27–31] despite important differences between how DNNs and cortex represent stimuli (e.g., receptive fields). Finally, there is no evidence for a strong positive, but sub-threshold, modulation of visual responses by low-level visual surprise evident in either the whole-brain or ROI results. This suggests that it is unlikely that a quantitative issue, such as low statistical power due to a subpar low-level feature model or non-ideal stimulus set for evoking low-level visual predictions, explains the current results. A limitation of the present approach is that neighboring layers of visual DNNs tend to correlate and hence explain significant shared variance of visual representations (S8 Fig), thus contraindicating the inclusion of all layers into the same GLM. For this reason, our key analyses focused on 2 layers, layer 2 as low-level and layer 8 as high-level visual feature model. Future work could improve on the generalization of the present results by utilizing different, ideally more complex stimulus sets and improved feature models, thereby further investigating whether (other aspects of) low-level features may modulate visual prediction errors. The word category (semantic) model has similar limitations. While word2vec and related models have been widely used to index semantic dissimilarity [48–50], they may no longer constitute state-of-the-art models. Thus, given a superior semantic feature model [51] or different task, visual prediction errors could be modulated by semantic surprise. However, while improvements in the semantic model are possible, it seems unlikely that incremental improvements in the model can fully account for the observed results, especially in EVC, given that semantic modulations tend to arise later in the processing hierarchy. In sum, while we cannot rule out that in some circumstances other models, stimuli, or tasks could result in low-level visual or linguistic semantic features modulating sensory prediction errors, our data indicates a propensity of the visual system to scale prediction errors in terms of high-level visual surprise. The precise features represented by visual DNNs are poorly understood. Visualizing features (S1 Fig) can help to gain an intuition for the feature tuning, but ultimately hand-crafted and better controlled feature models might be necessary to further elucidate the precise features that are reflected in visual prediction errors. That said, our conclusions are not dependent on layers 2 and 8 specifically; see S7 Fig for comparable results using layers 1 or 3 instead of layer 2 as low-level, and layer 7 instead of layer 8 as high-level visual surprise model. In other words, we do not propose that layer 8 of the utilized DNN is particularly important, rather it should be seen as one instance of a high-level feature space correlating with visual representations and hence scaling with visual cortical surprise. Interestingly, surprise modulations appeared to be more pronounced in V1 compared to LOC or HVC. This raises the question of whether such a distinction reflects a functionally significant difference between these ROIs, such as V1 acting as a high-resolution blackboard [52] particularly sensitive to predictive modulations, or if this result reflects a more mundane consequence of differences in neurovascular coupling that leads to superior signal-to-noise ratios in V1. The exact nature of this observation remains to be elucidated. [END] --- [1] Url: https://journals.plos.org/plosbiology/article?id=10.1371/journal.pbio.3002829 Published and (C) by PLOS One Content appears here under this condition or license: Creative Commons - Attribution BY 4.0. via Magical.Fish Gopher News Feeds: gopher://magical.fish/1/feeds/news/plosone/