(C) PLOS One This story was originally published by PLOS One and is unaltered. . . . . . . . . . . Language-specific neural dynamics extend syntax into the time domain [1] ['Cas W. Coopmans', 'Max Planck Institute For Psycholinguistics', 'Nijmegen', 'The Netherlands', 'Donders Institute For Brain', 'Cognition', 'Behaviour', 'Radboud University', 'Centre For Language Studies', 'Helen De Hoop'] Date: 2025-01 In this study, we investigated when the brain projects its knowledge of syntax onto speech during natural story listening. Atemporal syntactic structures were extended into the time domain via the use of incremental node count, whose temporal distribution showed that the delta band is a syntactically relevant timescale in our stimuli. Using a forward modeling approach to map node counts onto source-reconstructed delta-band activity, we then compared three parsing models that differ in the dynamics of structure building. A key finding of these analyses is that neural source dynamics most strongly reflect node counts derived from a top-down parsing model. This model postulates syntactic structure early in time, suggesting that predictive structure building is an important component of Dutch sentence comprehension. The additional (though weaker) effects of the bottom-up and left-corner predictors indicate that integratory and mildly predictive parsing also play a role, and suggest that people’s processing strategy might be flexibly adapted to the specific properties of the linguistic input. 4.1. Predictive structure building in the brain Node counts derived from all 3 parsing models explained unique variance in delta-band MEG activity, consistent with recent studies showing a relationship between delta-band activity and syntactic processing [32,33,35,50–55]. The results also align with prior studies that use neuro-computational language models to study syntactic processing [16,19,21–23,27,28,34], but they nevertheless advance our understanding of the neural encoding of syntax due to the inclusion of low-level (linguistic and non-linguistic) predictors in our TRF models. In addition to a number of acoustic predictors, the null models also contained information-theoretic predictors (e.g., surprisal, entropy) that partially reflect syntactic information and therefore explain some of the variance in brain activity whose origin is syntactic. By including these semi-syntactic control predictors, we stacked the cards against us and likely underestimated the “true” neural response to syntactic structure. To determine to what extent our syntactic effects are influenced by the presence of these control predictors, we ran an additional analysis with TRF models from which surprisal, entropy, and word frequency were omitted. Removing these statistical predictors increases the variance explained by each syntactic predictor, but the pattern of results remains unchanged: the top-down model still outperforms the 2 other parsing models (see S5 Fig and S1 Text, section 2). The fact that the effects of syntactic processing are robust and stable, both in the presence and in the absence of statistical control predictors, underscores the relevance of syntax for the neural mechanisms underlying language comprehension [35]. Of all syntactic predictors, node counts derived from a top-down parsing method were the strongest syntactic predictor of brain activity in language-relevant areas. These effects peaked twice within the first 500 ms after word onset and encompassed mostly superior and middle temporal, and inferior and middle frontal areas in the left hemisphere. The predictiveness of top-down node counts is somewhat at odds with previous studies that have looked at different parsing models in naturalistic comprehension, which either find that top-down methods are less predictive of brain activity [23,34,85] or that they do not differ from other parsing methods [21]. What could account for the strong top-down effects? One explanation is that top-down node counts capture the predictive nature of language processing well. There is substantial evidence from psycholinguistics that people generate structural predictions across a variety of syntactic constructions [39,86–89], and they do so in naturalistic contexts as well [15,18,90]. These predictive structure-building processes are mostly associated with activity in the left PTL [21,23,91–93] and left IFG [91], both of which were responsive to the top-down predictor in the current study. Matchin and colleagues [92] suggest that the PTL is involved in predictive activation of sentence-level syntactic representations and/or increased maintenance of the syntactic representations associated with lexical items when they are presented in a sentential context (see also [93]). On both interpretations, the PTL encodes structural representations that can be activated in a predictive fashion and are later to be integrated with the sentence-level syntactic representation in IFG [94–96]. Importantly, this process does not proceed in a purely feedforward manner, but rather relies on recurrent connections between temporal and frontal regions [97]. That the TRF of the top-down predictor is bimodal is consistent with this idea (Figs 4B and 6C). The first peak around 100 ms was strongest in posterior regions and likely reflects predictive structure building, which can occur in the PTL within the first 150 ms after word onset [92]. Indeed, the early timing of this effect is consistent with the predictive nature and temporal spell-out of top-down parsing, which is maximally eager and postulates syntax early in time. Previous electrophysiological studies have shown that phrase-structure violations elicit similarly early responses, including an early left anterior negativity [98–100] and a syntactic Mismatch Negativity [101,102], which have been linked to left superior temporal cortex in lesion and localization studies [98–100,102]. The fact that phrase-structure violations are detected so early strongly suggests that the brain generates predictions about the syntactic properties of upcoming material, perhaps in the form of precompiled phrase structure representations [94,98,101]. The second, later peak in the top-down TRF was strongest in more anterior regions and is largely consistent with syntactic surprisal effects, both in terms of timing [15,18,90] and spatial distribution [103]. It might therefore reflect responses related to the disconfirmation of predicted structures based on incoming information, which triggers a revision of the initial parse of the sentence. An additional observation is that the effects of the top-down predictor, and to a lesser extent also those of the left-corner predictor, are somewhat bilateral. A tentative explanation for this finding is that syntactic prediction, as quantified by top-down and left-corner node counts, is demanding and therefore requires support from right-hemispheric regions. These areas are presumably not the locus of the syntactic representations and computations themselves but might be activated when processing demands are increased [104]. Understanding how these MEG time courses are related to the structural neuroanatomy of syntax is an important topic for future research, which would benefit from a multi-methodological approach. Compared to the top-down TRF, bottom-up effects were relatively weak in amplitude. This is quite striking because, of all syntactic features, the bottom-up predictor is most strongly correlated with properties of the acoustic signal. That is, bottom-up node count quantifies syntactic phrase closure, which is related to the closure of prosodic phrases [79–81] and moderately correlated with the strength of prosodic boundaries across all story segments (r = 0.34 in our stimuli). To determine whether any of the syntactic effects might be explained in terms of brain responses to prosody, which also modulates delta-band activity [76–78], we included prosodic boundary strength as predictor in our TRF models. The results show that while the top-down effects are quite stable, the unique variance explained by the bottom-up and left-corner predictors is significantly reduced by the addition of prosodic boundary strength as regressor (see S6 Fig and S1 Text, section 3). It thus seems that the integratory effects of bottom-up and left-corner node count are modulated by prosody, pointing to a potential role for prosody in supporting the inference of syntactic structure [82,105]. The top-down effect, however, is not affected by prosodic boundary strength, providing additional support that it reflects abstract structure building. More generally, this result reinforces the idea, alluded to the introduction, that perception requires the brain to add information to its analysis of the physical environment. Spoken language comprehension is more than passive extraction of information from speech—it is an active inference process, during which the brain adds its knowledge of syntactic structure. Because syntactic effects cannot be driven entirely by what is in the signal, they must at least partially reflect what the brain perceptually infers from that signal. It seems that the bottom-up and top-down effects are inversely related. Strong top-down effects, which indicate predictive processing, are accompanied by weak bottom-up effects, and vice versa. One of the advantages of prediction in language comprehension is that it can reduce the burden on future integration processes [106]. If a structural representation has already been pre-built or pre-activated, correctly predicted incoming words only have to be inserted into the existing structure, so integration costs for these words are low. Because integration costs are approximated via bottom-up node count, any effects of bottom-up node count should be reduced if people successfully engage in predictive processing. This account would thus predict that when top-down metrics modulate brain activity for a given sentence, bottom-up metrics will not provide a good fit for that same sentence, and conversely, when top-down metrics do not provide a good fit, bottom-up metrics should be highly predictive. While it should be investigated in future work whether the effects of top-down and bottom-up metrics indeed go hand in hand in this anti-correlated way, the results of 2 naturalistic studies with spontaneous speech are consistent with this possibility. Contrasting with coherent audiobook narratives, spontaneously produced speech contains dysfluencies and corrections, which might make participants less inclined to rely on predictive processing. In an fMRI study by Giglio and colleagues [34], English speakers had to listen to other people’s verbal summaries of a TV episode. The authors found that bottom-up node counts modulated activity in language-relevant brain areas (LIFG and LPTL) more strongly than top-down node counts. Second, an EEG study by Agmon and colleagues [107] investigated the neural encoding of different linguistic features when Hebrew speakers listened to a spontaneously generated narrative. Using a TRF analysis on low-frequency activity, they found that the neural response for words closing a clause (bottom-up) was stronger than the neural response for words opening a clause (top-down). Both studies indicate that in the comprehension of spontaneous speech the brain relies relatively strongly on integratory processing. A possible implication of these findings is that parsing strategies can be flexibly adapted to the specific properties of the current linguistic input (e.g., grammatical properties, sentence complexity, reliability of predictive cues), such that people are less likely to engage in predictive structure building if the input contains ungrammatical sentences that make predicting ineffective [106,108]. As an exploratory analysis of the potential trade-off between predictive and integratory structure building, we investigated whether predictability (quantified through surprisal) modulates the demands on bottom-up structure building (see S7 Fig and S1 Text, section 4; see also [85]). How exactly the demands on predictive and integratory structure building dynamically change over the course of sentence processing should be investigated in future research. Effects of the left-corner predictor were also relatively weak compared to effects of the top-down predictor, especially after the addition of prosodic boundary strength as control regressor (S6 Fig). The comparatively weak effects of left-corner node counts might indicate that left-corner metrics are insufficiently predictive to account for the comprehension of head-final constructions of Dutch. The left corner of head-initial structures is very informative, which could explain why left-corner parsing metrics successfully predict brain activity of participants comprehending English [23,28]. We suggested in the introduction that these effects might be weaker in languages with head-final constructions, in particular if speakers of these languages adopt predictive parsing strategies. However, in seeming conflict with this possibility, a recent study in Japanese, a strictly head-final language, showed that a left-corner parsing model outperformed a top-down parsing model in left inferior frontal and temporal-parietal regions [109]. One relevant difference with our study is that they used a complexity metric that considers the number of possible syntactic analyses at each word (i.e., modeling ambiguity resolution), rather than directly quantifying the number of operations that are required to build the correct structure (i.e., node count for a one-path syntactic parse tree). These metrics do not reflect the same process. It is therefore possible that brain activity corresponding to ambiguity resolution is best modeled by considering the number of syntactic analyses following a left-corner strategy, and that, when the most likely analysis is chosen, the structure-building process itself is best modeled via top-down metrics. To better understand such diverging results, it is important that future studies take into account how syntactic properties (of different languages) affect the suitability of different parsing strategies. As a case in point, it is commonly mentioned that the left-corner strategy predicts processing breakdown for exactly those constructions that are difficult to process. Left-corner parsers have the property that their memory demands increase in proportion to the number of embeddings in center-embedded constructions, while they remain constant for both right- and left-branching structures [36–38]. Sentences with multiple levels of center-embedding indeed quickly over-tax working memory resources [31], supporting the cognitive plausibility of the left-corner method. Intriguingly, however, processing difficulty for center-embedded constructions is not consistent across languages. For example, it has been found that speakers of German (a language with head-final VPs, like Dutch) are hindered less than English speakers during the comprehension of multiply center-embedded sentences [43]. Language-dependent findings such as these reinforce the idea that people’s ability to generate syntactic predictions might be dependent on the specific grammatical properties of the language. Because the average dependency lengths in head-final languages are longer than those in head-initial languages [110], speakers of head-final languages have more experience with processing longer dependencies and might therefore rely more strongly on a processing strategy that facilitates the comprehension of these structures. In all, the fact that node counts derived from a top-down parser best explain brain activity of people listening to Dutch stories might have to do with certain grammatical properties of Dutch, including its head-final VPs, which make left-corner prediction inadequate. That this conclusion can be reached merely by extending the approach to a closely related Germanic language underscores the need for more work on typologically diverse languages, whose structural properties invite different parsing strategies that might rely on different brain regions to varying degrees [44]. Overall, the fronto-temporal language network is remarkably consistent across speakers of different languages [111], but structural differences within this network can be induced by experience with sentence structures that elicit different processing behavior [112]. [END] --- [1] Url: https://journals.plos.org/plosbiology/article?id=10.1371/journal.pbio.3002968 Published and (C) by PLOS One Content appears here under this condition or license: Creative Commons - Attribution BY 4.0. via Magical.Fish Gopher News Feeds: gopher://magical.fish/1/feeds/news/plosone/