Neural Tracking of Linguistic Predictors in Spontaneous Conversational Speech

preprint OA: closed CC-BY-NC-ND-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

This study investigates whether neural tracking of linguistic information extends from read speech to spontaneous conversation. Using the temporal response function (TRF) framework, we validate our approach on a read-speech EEG dataset and then apply it to EEG recordings from natural conversations. We observe reliable neural tracking of key linguistic predictors, including word onset, part-of-speech surprisal, and lexical surprisal, in spontaneous speech, with effects around 200, 400, and 600 ms. These results provide new evidence that linguistic neural tracking operates in natural conversational settings and confirm the feasibility of EEG studies in ecologically valid contexts.
Full text 38,460 characters · extracted from oa-pdf · 7 sections · click to expand

Abstract

This study investigates whether neural tracking of lin- guistic information extends from read speech to sponta- neous conversation. Using the temporal response func- tion (TRF) framework, we validate our approach on a read-speech EEG dataset and then apply it to EEG recordings from natural conversations. We observe reli- able neural tracking of key linguistic predictors, includ- ing word onset, part-of-speech surprisal, and lexical sur- prisal, in spontaneous speech, with effects around 200, 400, and 600 ms. These results provide new evidence that linguistic neural tracking operates in natural con- versational settings and confirm the feasibility of EEG studies in ecologically valid contexts. Keywords:TRF; Conversation; EEG; Surprisal

Introduction

The study of the neural bases of language comprehen- sion relies on a range of techniques aimed at identify- ing the impact of specific linguistic phenomena in the brain signal. A substantial body of work has investi- gated different levels of linguistic processing by linking them to event-related potentials (ERPs) emerging at dis- tinct time windows in the brain signal after word onset. More recently, a complementary approach has been pro- posed, known as the temporal response function (TRF), which consists in predicting the brain signal from lin- guistic predictors (Brodbeck et al., 2023; Crosse et al., 2016; Ding & Simon, 2012). The temporal response function framework enables the fine-grained investigation of linguistic predictors, in- dividually or jointly, by modeling the correspondence between continuous linguistic signals and brain activity over extended time scales. Using this approach, pre- vious studies have demonstrated neural tracking of the speech envelope (Brodbeck & Simon, 2020; Ding et al., 2014) as well as of multiple linguistic predictors, includ- ing phoneme and word onsets, lexical surprisal, and se- mantic dissimilarity, highlighting the flexibility of TRFs for probing different levels of language processing (Brod- erick et al., 2018; Chalehchaleh et al., 2025; Gillis et al., 2021; Heilbron et al., 2022; Weissbart et al., 2020). However, most TRF studies rely on passive listening to read speech under controlled EEG conditions, typically using audiobook stimuli. In contrast, only a few studies have examined natural conversational speech (Goldstein et al., 2025; Silem et al., 2025; Zada et al., 2024), largely due to the challenges of collecting and analyzing EEG data in ecological settings where speech production and movement introduce substantial noise. In this study, we address the challenge of analyzing the neural correlates of speech in natural settings by in- vestigating whether findings from read speech generalize to spontaneous speech. In the naturalistic setting, the brain not only processes incoming speech but also plans upcoming responses, engaging additional neural systems. It is also harder to predict due to disfluencies and greater variability in pacing/repairs. We hypothesize that prin- cipal linguistic predictors contributing to neural tracking are also active during spontaneous speech. First, we introduce a processing pipeline that trains models on single linguistic predictors and validate it on an existing read-speech dataset (Bhattasali et al., 2020), successfully replicating previously reported results. Sec- ond, we apply the same pipeline to a corpus of sponta- neous conversational speech (Boudin et al., 2023). Our results confirm neural tracking of several lin- guistic predictors, including word onset, part-of-speech surprisal, and lexical surprisal in spontaneous speech with robust effects in canonical linguistic time windows around 200, 400, and 600 ms. To our knowledge, this study provides the first evidence of linguistic neural tracking in spontaneous speech and demonstrates the feasibility of using EEG data collected in naturalistic conversational settings Related works One major advantage of TRFs is their ability to capture brain activity over extended time periods, representing a substantial methodological advance for the analysis of natural language. However, TRF studies have so far re- lied almost exclusively on read speech (Broderick et al., 2018; Chalehchaleh et al., 2025; Dou et al., 2025; Gillis et al., 2021; Heilbron et al., 2022; Weissbart et al., 2020). In the present study, we seek to advance our understand- ing of the neural bases of language in ecological contexts by focusing on spontaneous rather than read speech. Predictor selection is a central issue in TRF studies. Extensive prior work has highlighted the crucial role of speech envelope tracking in neural responses (Brodbeck .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted May 14, 2026. ; https://doi.org/10.64898/2026.05.13.722865doi: bioRxiv preprint & Simon, 2020; Ding et al., 2014; Lalor & Foxe, 2010), demonstrating effects of rhythmic structure and acoustic onsets (Ding & Simon, 2014), as well as higher-level pre- dictors such as phoneme onsets (Brodbeck et al., 2023; Donhauser & Baillet, 2020). These predictors typically elicit a negative deflection around 100 ms, sometimes fol- lowed by a second negativity around 250 ms, resembling the N250 or phonological mismatch negativity (Dou et al., 2025; Gillis et al., 2021). At higher linguistic levels, word-based predictors such as word onset, surprisal, and semantic dissimilarity evoke later and distinct neural responses depending on their representational level. For instance, part-of-speech (POS) surprisal has been linked to effects in the 200 – 500 ms time window (Heilbron et al., 2022). At the lexical level, predictors including word onset, lexical sur- prisal, and word frequency are typically associated with responses around 400 ms, often interpreted as reflecting an N400 component (Broderick et al., 2018; Dou et al., 2025; Weissbart et al., 2020). Semantic-level predictors have also been investigated, most notably semantic dis- similarity (Broderick et al., 2018). These studies likewise report N400-like effects, although such findings are not always consistently replicated (Gillis et al., 2021). Recent work has examined brain responses in terms of both latency and duration, highlighting the distinct temporal profiles of individual predictors — not only in their predictive power (Dou et al., 2025), but also in their position within the hierarchy of linguistic process- ing (Gwilliams et al., 2025).

Methods

Data We used two datasets, to contrast a more controlled, passive listening scenario with a richer, more naturalistic and dynamically interactive conversational setting. The first corpus is the Alice dataset (Bhattasali et al., 2020; J. R. Brennan, 2023), which includes EEG record- ings from 49 participants listening to the opening chapter of Alice in Wonderland (12.4 min; segmented into 12 tri- als). The data was recorded using 61 electrodes at 500 Hz. Several participants had already been excluded by the original authors due to experimental errors, unmet behavioral criteria, or excessive noise, leaving 33 par- ticipants for the analysis. We excluded four additional participants because their first trial was missing. For the second dataset, we used the SMYLE corpus (Boudin et al., 2023), a French multimodal dataset com- bining audio, video, and EEG recordings. It includes 30 dyads (16h) who first engage in storytelling and then in free conversation. Storytelling consists of three tasks: recounting a short video (the Pear Story (Chafe, 1981)), pitching a movie, book, or video game, and describ- ing a memorable vacation. Two listener conditions were employed: attentive, in which listeners followed and re- sponded naturally, and distracted, in which listeners se- cretly counted words starting with /t /. For this study, we selected 19 dyads after excluding dyads with exces- sive noise and focused on the storytelling task, in which one participant acts solely as listener to reduce EEG noise. The corpus provides enriched orthographic tran- scriptions (Blache et al., 2017), segmented into Inter- Pausal Units (IPUs) and annotated for laughter, disflu- encies, repetitions, truncated words, and elisions. Tran- scriptions were normalized, tokenized, and time-aligned with the speech signal using the SPPAS toolkit (Bigi, 2012). EEG data were recorded with two 64-channel BioSemi systems (10 – 20 layout) at 2048 Hz. EEG Pre-processing EEG preprocessing was conducted using MNE-Python 1.10 (Gramfort, 2013). Alice dataset bad channels were pre-marked by the authors (Bhattasali et al., 2020; J. R. Brennan, 2023). For SMYLE participants, noisy or artifact-ridden channels were marked as bad via visual inspection of raw signals and power spectra. Participants with>20% bad channels were excluded, resulting in two Alice and four SMYLE participants removed. Signals were referenced to the common average, band-pass fil- tered (0.5 – 30 Hz, FIR), and bad channels interpolated using spherical splines. To limit interpolation to <15% (Crosse et al., 2021), three electrodes were excluded per dataset, leaving 58 for Alice and 61 for SMYLE. EEG signals were downsampled to 256 Hz. We investigated broad band frequencies as well as frequencies from the delta band. We included the delta band, since word- related speech features naturally occur at 1–4 Hz, match- ing the temporal dynamics of delta oscillations For the SMYLE dataset, recorded during natural con- versations, a semi-automatic artifact removal procedure using ICA was applied. Signals were scaled to unit vari- ance and whitened via PCA, then FastICA (Hyvarinen, 1999) extracted ICs according to the data rank. The ICLabel method was used to inspect and classify ICs as eye-blink, muscle, or cardiac artifacts (Li et al., 2022; Pion-Tonachini et al., 2019). In parallel, a human inspec- tor classified ICs by visually examining IC time-series, topography and power spectrum. After excluding the ICLabel- and human-identified noise ICs, the EEG sig- nals were reconstructed using all of the remaining ICs. Linguistic Features We aim to investigate the neural correlates of a set of linguistic features, including higher-level features such as word and part-of-speech (POS) surprisal as well as low- level features like word onset and the speech envelope. Surprisal Estimation using LLMs. The word sur- prisal and the POS surprisal were estimated using LLMs: GPT-2 (Radford et al., 2019) for Alice and GPT-fr (Simoulin & Crabb´ e, 2021) fine-tuned on French con- .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted May 14, 2026. ; https://doi.org/10.64898/2026.05.13.722865doi: bioRxiv preprint versation for SMYLE. We argue that English LLMs are better equipped to model spoken English due to the sub- stantially larger volume of available data and greater rep- resentation of spoken language compared with French. To address this gap, we fine-tuned the GPT-fr base model on a SMYLE-derived conversational dataset with LoRA applied to all layers, using transcriptions from both the storytelling and free-conversation tasks with all disfluencies preserved. The model was trained on samples of 10 consecutive turns separated by the marker. Training ran for five epochs with AdamW, us- ing the following parameters: learning rate = 0.002, 500- step warmup, batch size = 8, LoRA rank = 32, α= 32, dropout ratio = 0.05, and gradient clipping at 1. We passed the transcriptions (of the first Alice chapter or the concatenated turns of the speaker for SMYLE) through the LLM and obtained the logits for each word. These were transformed into conditional probabilities for each word by applying softmax and choosing the most probable word. We obtain the word surprisal as follows: surprisal(wi) =−log(P (wi|w1...wi−1)) For the POS surprisal, we followed to approach of Heil- bron et al. (2022), i.e. top-k nucleus sampling with k = 40 and p = 0.9. Specifically, the candidate set con- sisted of the most probable tokens whose cumulative probability reached 90%, with a minimum of the top 40 tokens always included. We restricted the maximum k to be 300 to reduce computational costs. We then cal- culated the POS tag using the Spacy library1 for each of the top-k nucleus sampled tokens as well as the actual target word. The POS surprisal is given by: surprisalPOS(wi) =−log (∑ t∈TP (wt|context) ∑ a∈AP(w a|context) ) withTas the group of top-k nucleus sampled words hav- ing the same POS tag as the target wordw i andAas the group ofalltop-k nucleus sampled word. All words were derived given the same context sequence as wi. Constructing Continuous Feature Signals. With our approach described above, we get discrete scores at each word onset. Since TRFs work on signals, we need to construct a continuous signal from these discrete values. For this, we initiate a continuous time series, or rather an array with the sampling rate of 256 Hz matching the du- ration of the conversation, set to zero throughout. Spikes scaled by the previously estimated surprisal value of the corresponding word were inserted at word onset times provided by the corpora. This yields a continuous sur- prisal representation. This procedure was repeated for the POS surprisal. 1https://spacy.io/ withfr core news lgfor French and en core web lgfor English. Because surprisal impulses occur at word boundaries, we additionally modeled a word onset feature to disso- ciate neural responses to linguistic information from re- sponses driven purely by boundary timing. This control regressor was generated using the same procedure, but with impulses of amplitude 1 placed at each onset. For the envelope, the amplitude envelope was ex- tracted using the Hilbert transform implemented in the Eelbrain toolbox (Brodbeck et al., 2023), and then re- sampled to the target sampling rate of 256 Hz. No normalization was applied to word onset since this feature was binary encoded. As suggested in (Crosse et al., 2021), the envelope was normalized by its standard deviation to maintain positive values. Given that sur- prisal and POS surprisal have identical timings as word onset, the non-zero values of the two features were z- score normalized. TRF Modeling Temporal response functions are widely used to study the relationship between linguistic predictors and EEG signal by modeling how an input predictor, when con- volved with a response function, predicts brain activity. In practice, models are trained separately for each par- ticipant/dyad and feature, estimating time-lagged pa- rameters that capture how predictors contribute to the EEG signal at different latencies. Both, feature and neu- ral signal were normalized before training (except for word onset). Model performance is then evaluated by comparing the predicted EEG signal with the observed data across electrodes. The estimated parameters, or weights, indicate how variations in the stimulus relate to EEG signals, reflect- ing the strength of coupling between the linguistic input and the neural response. When this coupling is strong, it is informative to examine whether its temporal profile corresponds to known ERP components. For instance, a strong coupling between lexical surprisal and EEG ac- tivity around 400 ms after stimulus onset may be inter- preted as reflecting an N400-like effect. In our study, the temporal lag range was defined from -200 ms to 800 ms, yielding 256 discrete time points given a sampling rate of 256 Hz. TRF weights were com- puted by minimizing the mean squared error between the recorded EEG signal and its model-based prediction. To reduce the risk of overfitting, for each dyad/participant, the EEG signals and the corresponding features were split into training and testing sets. For each partici- pant in the Alice dataset, the first 10 trials (83% of the data) were used for training, the last two trials were reserved for testing. For each dyad in the SMYLE cor- pus, the initial 90% of the data was used for training and the remaining 10% was reserved for testing. TRF models were fitted separately on the training set of each dyad/participant using ridge regression, with parame- terλbeing optimized via 5-fold cross-validation, dur- .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted May 14, 2026. ; https://doi.org/10.64898/2026.05.13.722865doi: bioRxiv preprint (a) Broadband Alice SMYLE (b) Deltaband Alice SMYLE Figure 1: Spatiotemporal clusters were identified using a permutation-based approach for the envelope. Left: Scalp topog- raphy of the T-map averaged across significant post-stimulus time window. Electrodes belonging to significant clusters are marked by white circles, with the color scale indicating the magnitude of the statistics. Right: Averaged time-course of the envelope responses. The blue shaded region indicates the significant time window corresponding to the topographic map. Table 1: Overall prediction accuracy (mean Pearson’s r) of TRF models for each feature for either broad (0.5 − 30 Hz) or delta (0.5 −4 Hz) bands. EEG Band Feature Pearson’s r (Alice) Pearson’s r (SMYLE) Broad Envelope 0.0395 0.0198 Word Onset 0.0486 0.0164 POS Surprisal 0.0131 -0.0032 Surprisal 0.0017 -0.0018 Delta Envelope 0.0539 0.0252 Word Onset 0.0663 0.0235 POS Surprisal 0.0223 -0.0050 Surprisal 0.0072 0.0011 ing which 100 candidate values logarithmically spaced between 10−4and 10 12 were evaluated. Model perfor- mance was quantified on the test set using the Pearson correlation coefficient of the predicted and observed EEG signals.

Results

For both datasets, we trained TRF models for each par- ticipant using individual linguistic predictors. Over- all model performance is summarized in Table 1. TRF significance was assessed with a mass uni- variate spatiotemporal cluster-based permutation test (Maris & Oostenveld, 2007) using MNE-Python’s spatio temporal cluster 1samp test. The method computes univariate t-statistics across electrodes and time points, clusters adjacent points exceeding p = 0.01, and compares each cluster’s mass (sum of t-values) against a permutation-based null distribution of maxi- mum cluster mass (10, 000 iterations) to determine sig- nificance (p< 0.05). Envelope Many studies have described acoustic neural tracking, in particular showing the predictive power of the speech envelope (Crosse et al., 2016; Ding & Simon, 2014; Dren- nan & Lalor, 2019; Yasmin et al., 2023). These works report several early effects starting as early as 50 ms and extending up to 250 ms. These effects are often associated with the N1–P2 complex, known to encode speech information. With regard to neural tracking, dif- ferent studies have shown amplitude peaks at 50 – 80 ms (Ding et al., 2014) and at 130 – 200 ms (Drennan & Lalor, 2019; Yasmin et al., 2023). Our results (see Figure 1) on the Alice corpus show a similar double peak: the first one very early, around 30 ms (p< 0.019), and the second at about 170 ms (78 - 164ms, p< 0.048). This double peak therefore confirms a well-known effect in the neural tracking of read speech. Our results in natural conversation show a pattern that could correspond to an amplitude modulation in the same temporal window, between 30 and 200 ms, showing the N1–P2 complex. A significant cluster only appears around 400 to 700 ms. By comparing the results obtained for acoustic track- ing in the delta band (0.5 - 4 Hz), which is frequently used in this field (Weissbart et al., 2020), we instead ob- serve a greater similarity between the conditions (Figure 1b; p< 2×10−4for Alice; p< 3×10−3for SMYLE). In both cases, we indeed observe an overall comparable pattern across the entire time window studied, with a marked amplitude between 30 and 200 ms, this ampli- tude being slightly earlier and sharper in the case of read speech compared to spontaneous speech. Word Onset Word onset is an important predictor for neural track- ing. The literature reports an amplitude increase from 50 to 300 ms after word onset, and a second peak around 300 to 400 ms, usually interpreted as an N400 (Gillis et .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted May 14, 2026. ; https://doi.org/10.64898/2026.05.13.722865doi: bioRxiv preprint (a) Broadband Alice SMYLE (b) Deltaband Alice SMYLE Figure 2: Spatiotemporal clusters of the weights trained on word onset responses were identified using a permutation-based approach. Significant clusters are highlighted by the topography and the blue shadow in the time courses. al., 2021; Karunathilake et al., 2023). In our results (see Figure 2), for read speech we observe a double-peak phe- nomenon around 150 ms and 230 ms, which then gradu- ally decreases up to 500 ms (p< 0.0014). This observa- tion doesn’t exactly match the findings in the literature regarding the second peak, which appears earlier here. This early effect, more pronounced around 200 ms, can possibly be associated with attentional phenomena typ- ically observed in this time window (usually the P200).

Results

for spontaneous speech, on the other hand, show a peak around 300 – 500 ms (p< 0.05). This second ob- servation corresponds more clearly to an interpretation in terms of the N400. In both cases, even though this phenomenon is less marked for read speech, it is interest- ing to note the strong predictive power of word onset in the 300 – 500 ms time window in both conditions (even in the absence of a clear peak for read speech). By exam- ining the neural tracking of the word onset in the delta band, the results show a clearer peak between 150 and 350 ms for read speech (p< 7×10−4) and between 150 and 450 ms for spontaneous speech (p< 0.002). In this case, we observe more similar patterns across both con- ditions over a relatively broad shared time interval. In both cases, a peak around 300 ms is observed, which has been reported in studies describing a P300–N400 com- plex related to categorization and surprise. POS Surprisal POS surprisal as a neural predictor is expected to show effects between 200 and 500 ms (J. Brennan & Hale, 2019; Heilbron et al., 2022). The effects associated with this predictor are more specifically syntactic, with a prominent impact around 400 ms, corresponding to ef- fects observed in the P300–N400 complex related to pre- diction and memory access (Bornkessel-Schlesewsky & Schlesewsky, 2019). In addition, a later effect around 600 – 800 ms has also been reported (Heilbron et al., 2022). We observe similar effects starting around 100 ms (82 - 129 ms,p< 0.045), with a significant peak at 400 ms (see Figure 3; p< 0.02) and extending into a later effect in the Alice dataset (473 - 800ms; p< 9×10−4). Interest- ingly, the same pattern is also observed for spontaneous speech. Focusing more specifically on the delta band reveals the same POS surprisal impact pattern within the 200 – 700 ms time window (p< 0.0025 for Alice), with a comparable pattern observed for both read and (a) Broadband Alice SMYLE (b) Deltaband Alice SMYLE Figure 3: Spatiotemporal clusters of the weights trained on POS surprisal responses were identified using a permutation- based approach. Significant clusters are highlighted by the topography and the blue shadow in the time courses. .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted May 14, 2026. ; https://doi.org/10.64898/2026.05.13.722865doi: bioRxiv preprint (a) Broadband Alice SMYLE (b) Deltaband Alice SMYLE Figure 4: Spatiotemporal clusters of the weights trained on surprisal responses were identified using a permutation-based approach. Significant clusters are highlighted by the topography and the blue shadow in the time courses. conversational speech. Surprisal The question of neural tracking of surprise has been in- vestigated in several studies using read speech. These studies have generally shown an effect of lexical sur- prisal around 400 ms (Chalehchaleh et al., 2025; Gau- thier & Levy, 2023; Gillis et al., 2021; Heilbron et al., 2022; Weissbart et al., 2020). This is an expected effect, which can be associated with the N400 component in ERP terms. It should be noted, however, that in these studies the results were characterized by small effect sizes and were often limited to a small number of electrodes. In addition, one study showed that multiple frequency bands can be shaped by surprisal, with in particular late effects observed between 200 and 1000 ms in the beta and gamma frequency bands (Weissbart et al., 2020). Our results are more contrasted (see Figure 4). For the Alice dataset, as in the literature, we observe a slight impact of surprisal between 250 and 450 ms, although this effect is not very pronounced, thus confirming the limited effects reported in previous work. This effect ap- pears in a more diffuse manner for spontaneous speech, with a possible impact starting around 350 ms. How- ever, our results also show a significant late effect around 700 ms (p< 0.02). Interestingly, these findings are very different from those obtained with the word onset pre- dictor, which may seem surprising. The analysis of neu- ral tracking in the delta band does not allow us to draw comparable and significant conclusions between read and spontaneous speech either. Nevertheless, a pattern can be observed for read speech, with a first effect around 150 ms and a second around 300 ms, which is not visible for spontaneous speech.

Discussion

& Conclusion As expected, TRF models trained on read speech show higher predictive performance than those trained on nat- uralistic spontaneous speech (Table 1). In both con- texts, linguistic features exhibit a similar performance pattern: low-level features outperform high-level fea- tures, and restricting neural tracking to the delta band substantially improves model performance. Unlike the narrative speech in Alice (audio), natural conversation in SMYLE provides intensive multimodal input, requir- ing the brain to integrate visual, auditory, linguistic, and social information. This environmental complexity may increase EEG variability, including greater spatial vari- ation across electrodes and temporal asynchrony, result- ing in fewer significant clusters in SMYLE than in Alice. However, when analysis is restricted to the delta band, typically associated with word-level speech frequency, the results show remarkable consistency between Alice and SMYLE (see Figure 1 and 2). Our objective was to assess the feasibility of study- ing neural tracking in spontaneous speech. We used temporal response functions, which rely on extracting linguistic predictors from speech signals. Here, we em- ployed commonly used predictors spanning low-level fea- tures (speech envelope, word onset) and higher-level fea- tures (part-of-speech and lexical surprisal). To this end, we fine-tuned a large language model on spontaneous speech to extract predictors tailored to this setting and used them to predict EEG signals. Our results first allowed us to replicate findings from the literature for these predictors using a read-speech dataset. Importantly, results from spontaneous speech revealed comparable neural tracking effects. Together, these findings pave the way for new investigations of spontaneous speech, complementing the controlled speech paradigms traditionally used in neurolinguistics. Further studies using multiple frequency bands are ex- pected to discriminate different cognitive processes in natural conversations.Additionally, it would be valuable to examine predictors using multivariate models to assess the unique variance contributed by each feature. Future research should also explore the use of models trained on larger corpora or multilingual models to assess whether these factors can improve performance and generalisabil- ity. .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted May 14, 2026. ; https://doi.org/10.64898/2026.05.13.722865doi: bioRxiv preprint Funding This work, carried out within the Institute of Conver- gence ILCB, was supported by grants from France 2030 (ANR-16-CONV-0002) and the Initiative d’Excellence d’Aix-Marseille Universit´ e – AMIDEX, grant no. AMX- 22-CEX-057.

References

Bhattasali, S., Brennan, J., Luh, W.-M., Franzluebbers, B., & Hale, J. (2020). The alice datasets: FMRI & EEG observations of natural language comprehension. In N. Calzolari, F. B´ echet, P. Blache, K. Choukri, C. Cieri, T. Declerck, S. Goggi, H. Isahara, B. Mae- gaard, J. Mariani, H. Mazo, A. Moreno, J. Odijk, & S. Piperidis (Eds.),Proceedings of the twelfth language resources and evaluation conference(pp. 120–125). Eu- ropean Language Resources Association. Bigi, B. (2012). Sppas: A tool for the phonetic segmen- tations of speech. Proceedings of the Eighth Interna- tional Conference on Language Resources and Evalu- ation, 1748–1755. Blache, P., Bertrand, R., Ferr´ e, G., Pallaud, B., Pr´ evot, L., & Rauzy, S. (2017). The corpus of interactional data: A large multimodal annotated resource. Hand- book of linguistic annotation, 1323–1356. Bornkessel-Schlesewsky, I., & Schlesewsky, M. (2019). Toward a neurobiologically plausible model of language-related, negative event-related potentials. Frontiers in psychology, 10, 298. Boudin, A., Bertrand, R., Rauzy, S., Houl` es, M., Legou, T., Ochs, M., & Blache, P. (2023). Smyle: A new multimodal resource of talk-in-interaction including neuro-physiological signal. Companion Publication of the 25th International Conference on Multimodal In- teraction, 344–352. Brennan, J., & Hale, J. (2019). Hierarchical structure guides rapid linguistic predictions during naturalistic listening. PLoS ONE, 14. Brennan, J. R. (2023). EEG datasets for naturalistic lis- tening to “alice in wonderland” (version 2). Brodbeck, C., Das, P., Gillis, M., Kulasingham, J. P., Bhattasali, S., Gaston, P., Resnik, P., & Simon, J. (2023). Eelbrain, a python toolkit for time-continuous analysis with temporal response functions. eLife, 12. Brodbeck, C., & Simon, J. Z. (2020). Continuous speech processing. Current Opinion in Physiology, 18, 25–31. Broderick, M. P., Anderson, A. J., Di Liberto, G. M., Crosse, M. J., & Lalor, E. C. (2018). Electrophysio- logical correlates of semantic dissimilarity reflect the comprehension of natural, narrative speech. Current Biology, 28(5), 803–809. Chafe, W. (1981). The pear stories: Cognitive, cultural, and linguistic aspects of narrative production. Lan- guage in Society, 10(3). Chalehchaleh, A., Winchester, M., & Di Liberto, G. M. (2025). Robust assessment of the cortical encoding of word-level expectations using the temporal response function. Journal of Neural Engineering, 22(1). Crosse, M. J., Di Liberto, G. M., Bednar, A., & Lalor, E. C. (2016). The multivariate temporal response func- tion (mtrf) toolbox: A matlab toolbox for relating neu- ral signals to continuous stimuli. Frontiers in human neuroscience, 10, 604. Crosse, M. J., Zuk, N. J., Di Liberto, G. M., Nidiffer, A. R., Molholm, S., & Lalor, E. C. (2021). Linear Mod- eling of Neurophysiological Responses to Speech and Other Continuous Stimuli: Methodological Considera- tions for Applied Research. Frontiers in Neuroscience, 15, 705621. Ding, N., Chatterjee, M., & Simon, J. Z. (2014). Robust cortical entrainment to the speech envelope relies on the spectro-temporal fine structure. NeuroImage, 88, 41–46. Ding, N., & Simon, J. (2012). Neural coding of contin- uous speech in auditory cortex during monaural and dichotic listening. Journal of Neurophysiology, 107(1), 78–89. Ding, N., & Simon, J. Z. (2014). Cortical entrainment to continuous speech: Functional roles and interpreta- tions. Frontiers in Human Neuroscience, Volume 8 - 2014. Donhauser, P., & Baillet, S. (2020). Two distinct neu- ral timescales for predictive speech processing.Neuron, 105(2). Dou, J., Anderson, A. J., White, A. S., Norman- Haignere, S. V., & Lalor, E. C. (2025). Dynamic mod- eling of eeg responses to natural speech reveals earlier processing of predictable words. PLOS Computational Biology, 21(4), 1–28. Drennan, D. P., & Lalor, E. C. (2019). Cortical tracking of complex sound envelopes: Modeling the changes in response with intensity. eneuro, 6(3). Gauthier, J., & Levy, R. (2023). The neural dynam- ics of word recognition and integration. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 980–995. Gillis, M., Vanthornhout, J., Simon, J. Z., Francart, T., & Brodbeck, C. (2021). Neural markers of speech com- prehension: Measuring eeg tracking of linguistic speech representations, controlling the speech acoustics.Jour- nal of Neuroscience, 41(50), 10316–10329. Goldstein, A., Wang, H., Niekerken, L., Schain, M., Zada, Z., Aubrey, B., Sheffer, T., Nastase, S. A., Gazula, H., Singh, A., Rao, A., Choe, G., Kim, C., Doyle, W., Friedman, D., Devore, S., Dugan, P., Has- sidim, A., Brenner, M., . . . Hasson, U. (2025). A uni- fied acoustic-to-speech-to-language embedding space captures the neural basis of natural language pro- .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted May 14, 2026. ; https://doi.org/10.64898/2026.05.13.722865doi: bioRxiv preprint cessing in everyday conversations. Nature Human Be- haviour, 9(5), 1041–1055. Gramfort, A. (2013). MEG and EEG data analysis with MNE-Python. Frontiers in Neuroscience, 7. https : //doi.org/10.3389/fnins.2013.00267 Gwilliams, L., Marantz, A., Poeppel, D., & King, J.-R. (2025). Hierarchical dynamic coding coordinates speech comprehension in the human brain. Proceed- ings of the National Academy of Sciences, 122(42), e2422097122. Heilbron, M., Armeni, K., Schoffelen, J.-M., Hagoort, P., & de Lange, F. P. (2022). A hierarchy of linguistic pre- dictions during natural language comprehension. Pro- ceedings of the National Academy of Sciences, 119(32). Hyvarinen, A. (1999). Fast and robust fixed-point al- gorithms for independent component analysis. IEEE Transactions on Neural Networks, 10(3), 626–634. Karunathilake, I. D., Dunlap, J. L., Perera, J., Presacco, A., Decruy, L., Anderson, S., Kuchinsky, S. E., & Si- mon, J. Z. (2023). Effects of aging on cortical repre- sentations of continuous speech. Journal of neurophys- iology, 129(6), 1359–1377. Lalor, E. C., & Foxe, J. J. (2010). Neural responses to uninterrupted natural speech can be extracted with precise temporal resolution. European Journal of Neu- roscience, 31(1), 189–193. Li, A., Feitelberg, J., Saini, A. P., H¨ ochenberger, R., & Scheltienne, M. (2022). MNE-ICALabel: Automat- ically annotating ICA componentswith ICLabel in Python. Journal of Open Source Software, 7(76), 4484. Maris, E., & Oostenveld, R. (2007). Nonparametric sta- tistical testing of EEG- and MEG-data. Journal of Neuroscience Methods, 164(1), 177–190. Pion-Tonachini, L., Kreutz-Delgado, K., & Makeig, S. (2019). ICLabel: An automated electroencephalo- graphic independent component classifier, dataset, and website. NeuroImage, 198, 181–197. Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. (2019). Language models are un- supervised multitask learners. OpenAI blog, 1(8), 9. Silem, O., Fleig, M., Blache, P., Boudin, A., Oufaida, H., & Becerra-Bonache, L. (2025). Exploring neural correlates of predictability in natural face-to-face con- versation. Proceedings of the Annual Meeting of the Cognitive Science Society, 47. Simoulin, A., & Crabb´ e, B. (2021). Un mod` ele trans- former g´ en´ eratif pr´ e-entrain´ e pour le fran¸ cais. In P. De- nis, N. Grabar, A. Fraisse, R. Cardon, B. Jacquemin, E. Kergosien, & A. Balvet (Eds.), Taln-2021. ATALA. Weissbart, H., Kandylaki, K., & Reichenbach, T. (2020). Cortical tracking of surprisal during continuous speech comprehension. pg - 155-166. Journal of cognitive neu- roscience, 32(1), 155–166. Yasmin, S., Irsik, V. C., Johnsrude, I. S., & Herrmann, B. (2023). The effects of speech masking on neural track- ing of acoustic and semantic features of natural speech. Neuropsychologia, 186, 108584. Zada, Z., Goldstein, A., Michelmann, S., Simony, E., Price, A., Hasenfratz, L., Barham, E., Zadbood, A., Doyle, W., Friedman, D., Dugan, P., Melloni, L., De- vore, S., Flinker, A., Devinsky, O., Nastase, S. A., & Hasson, U. (2024). A shared model-based linguis- tic space for transmitting our thoughts from brain to brain in natural conversations.Neuron, 112(18), 3211– 3222.e5. .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted May 14, 2026. ; https://doi.org/10.64898/2026.05.13.722865doi: bioRxiv preprint

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-pdf

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-24T02:00:01.246996+00:00
License: CC-BY-NC-ND-4.0