Keywords
speech production; electroencephalography; hierarchical structure 37
38
39
40
41
preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for thisthis version posted March 25, 2026. ; https://doi.org/10.64898/2026.03.23.713128doi: bioRxiv preprint
1. Introduction 42
43
Language is an organised communication system in which smaller linguistic units are embedded 44
within larger structures: phonemes form syllables, which combine to form words, and words can 45
be used to build phrases and sentences (Rauschecker and Scott, 2009; Peelle and Davis, 2012). 46
This hierarchical structure is reflected in the neurobiology of both speech perception (Ding et al., 47
2015) and production (Sengupta and Nasir, 2016; Tremblay, Deschamps and Gracco, 2016; de 48
Heer et al., 2017) and supported by temporally distinct patterns of neural activity (Giraud and 49
Poeppel, 2012; Peelle and Davis, 2012) . A speech perception study (Ding et al., 2015) used 50
isochronous streams of syllables that could be grouped into words , phrases or sentences at 51
different fixed rates, and showed that cortical electrophysiological activity synchronised (i.e., 52
entrained) not only to the underlying syllabic rhythm but also to the word , phrase and sentence 53
rhythms. This frequency -tagging paradigm demonstrates that neural oscillations within posterior 54
temporal and inferior frontal cortices track the hierarchical linguistic structure of speech perception. 55
In other words, the brain may have special neural signatures at different levels of speech hierarchy. 56
Similarly, continuous speech production involves transforming abstract linguistic representations 57
into precisely timed articulatory actions, engaging cortical and subcortical regions . Foundational 58
neurocognitive models, such as the ‘Blueprint of the Speaker’ (Indefrey and Levelt, 2004) or the 59
dual-stream framework (Hickok and Poeppel, 2007; Hickok, 2012) also describe speech production 60
using a similar hierarchy. Overall, brain regions including motor and premotor cortices are 61
associated with motor planning and execution, auditory cortices with sensory-feedback prediction 62
and monitoring of speech outcomes (Houde and Jordan, 1998; Tian and Poeppel, 2010; 63
Christoffels et al., 2011), and basal ganglia and cerebellar structures with timing and sequencing 64
of motor programmes (Knolle et al., 2019; Archila-Meléndez et al., 2020; Jorge et al., 2022). These 65
regions may play different roles at different hierarchical levels of speech. 66
67
In motor-speech control, an efference copy (i.e., a forward model) is an internal representation of 68
the sensory consequences of a planned motor command (Crapse and Sommer, 2008) , which 69
enables the rapid evaluation of incoming sensory feedback for error detection and correction (Ford, 70
Roach and Mathalon, 2010; Niziolek, Nagarajan and Houde, 2013; Simmonds et al., 2014) . At 71
present, a key challenge of speech production is disentangling different levels of planning from 72
production proper and monitoring of predicted sensory consequences (Tremblay, Deschamps and 73
Gracco, 2016) . Producing any utterance involves a continuously updated planning scope that 74
spans upcoming segments and often the next word. Consequently, neural activity around a 75
speech-production boundary (e.g., a syllable or word) may reflect (i) syllabic sequencing and 76
articulatory chunking, (ii) lexical selection and word-level planning, and (iii) sensory predictions and 77
error monitoring, all partially overlapping in time. Hence, to study the full-scale neural processes 78
involved in speech production requires continuous speaking paradigms, while experimentally 79
preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for thisthis version posted March 25, 2026. ; https://doi.org/10.64898/2026.03.23.713128doi: bioRxiv preprint
isolating the different neural processes involved. Interestingly, debate remains on what are the 80
basic units of articulatory planning. Some studies suggest that syllables are the units of articulatory 81
planning (Laganaro, Valente and Perret, 2012; Bürki et al., 2016), whereas others postulate that 82
speech planning can proceed at the word level (Piai et al., 2015; Assaneo and Poeppel, 2018) . 83
84
Distinguishing between speech planning, monitoring and execution processes may also shed light 85
on the neural disruptions present in motor-speech disorders, including stuttering, where fluent 86
transitions between speech units are compromised (Indefrey, 2011; Chang and Guenther, 2020). 87
Interestingly, the relative position of speech gestures within words affects stuttering behaviour. 88
Studies have shown considerably higher dysfluency rates at word, phrase and sentence boundary 89
transitions, compared to within-word syllabic transitions (Bloodstein and Ratner, 2008; Buhr and 90
Zebrowski, 2009) . This dysfluency disproportion could arise from either encoding, prediction or 91
monitoring (Loucks et al., 2007; Ghazanfar and Eliades, 2014). Furthermore, the basal ganglia play 92
a critical role in motor control by modulating the cortex for motor initiation/execution, based on the 93
direct and indirect striatum -based pathways (Schulz et al., 2005) . These cortico-basal ganglia-94
thalamo-cortical circuits (aka, the cortico-basal ganglia loop) are essential for the execution of 95
motor commands, including motor speech gestures. Their role in stuttering has been suggested, 96
as well as, possible parallels between stuttering and Parkinson’s disease (PD) (Alm, 2004; Giraud 97
et al., 2008). Beyond the possible role of the direct and indirect pathways in stuttering, a more 98
recent link between stuttering and the basal ganglia has been proposed based on failures within a 99
faster cortico-basal ganglia loop that short -cuts the striatum, and instead stimulates the sub -100
thalamic nucleus directly - the hyperdirect pathway (HDP) (Neef, Anwander and Friederici, 2015). 101
The HDP leads to hyper-fast inhibitory signals to the cortex (Nambu, Tokuno and Takada, 2002). 102
Precision in inhibitory motor control via the HDP is suggested as necessary to stop or correct 103
ongoing motor actions, thereby contributing to sensorimotor control in speech production (Whillier 104
et al., 2018; Usler, 2022) . Importantly, the HDP may also be critically involved in motor/speech 105
initiation (Usler, 2022; Nambu et al., 2023) by resetting the primary motor cortex (M1) prior to the 106
arrival of the excitatory direct pathway that ultimately leads to muscle contraction (Nambu, 2011). 107
In turn, cortical inhibition by the HDP may be particularly important at higher hierarchical levels of 108
motor planning requiring M1 reset, including transitions between words, phrases and sentences. 109
Importantly, impairments along the HDP has been implicated in stuttering (Usler, 2022) . 110
111
Transitions between linguistic units, whether within a word ( e.g., syllable-to-syllable) or between 112
words (e.g., word-to-word), are hypothesized to engage distinct neural oscillatory mechanisms 113
(Cao, Thut and Gross, 2017; Cao et al. , 2024; Orpella et al. , 2024) . Prior studies have used 114
delayed-production paradigms (Kittilstved et al., 2018; Archila-Meléndez et al., 2020) , in which 115
speech planning and production proper (i.e., speech execution) are separated by an experimental 116
cue. While this approach helps to dissociate planning- from execution-related activity, it limits our 117
preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for thisthis version posted March 25, 2026. ; https://doi.org/10.64898/2026.03.23.713128doi: bioRxiv preprint
understanding of how neural dynamics support online transitions during continuous speech 118
production. In addition, while it is possible to instruct participants to produce isolated speech units 119
of different hierarchical levels (e.g., isolated syllables versus words) (Silbert et al., 2014), probing 120
these mechanisms in sequential/continuous speaking tasks is crucial. For instance, producing the 121
syllable ba in different contexts: i) in isolation; ii) as the first syllable in the word banana; or iii) as 122
the final syllable in the word scuba, may involve distinct neural processes (from planning to 123
monitoring) for the ba component, despite requiring the same broad articulatory gesture s. 124
However, these conditions hold co-articulatory differences that confound how the syllable ba, in 125
this example, is produced in the sequence context. If, instead, target syllables are embedded in a 126
syllabic sequence with matching co-articulatory conditions, but under different word-boundaries, it 127
becomes possible to study hierarchical speech production processes in a controlled manner. We 128
attempt to provide an analogy of how we to test this question using English language: if the word 129
‘scu’ and ‘nana’ would be legal english words, the sequences ‘scuba nana’ and ‘scu banana’ would 130
be valid phrases allowing to study how the target syllable ba plays different hierarchical roles in the 131
sequences, while preserving its phonological/phonetic form and co-articulatory effects. In practice, 132
working with pseudoword pairs enable s the construction of controlled stimulus sequences (that 133
occur across different syllable positions) across experimental conditions, which we use in this 134
study. 135
EEG time-frequency analyses, such as event-related spectral perturbations (ERSP) allow tracking 136
transient power changes across frequency bands in response to time-locked experimental events 137
(Delorme and Makeig, 2004), which can be used to separate hierarchical transitions and explore 138
the neural dynamics of speech production (Vos et al., 2010; Jenson et al., 2014). Beta oscillatory 139
activity (~15-30 Hz) over frontal cortices - often prominent over right inferior frontal and premotor 140
regions - has been associated with preparatory set/maintenance and sequential control during 141
speech and orofacial actions (Pfurtscheller and Lopes da Silva, 1999; Weiss and Mueller, 2012). 142
Converging evidence indicates a functional subdivision whereby low beta (~1 5-20 Hz) is more 143
closely associated with the indirect basal ganglia pathway whereas high beta (~20-30 Hz) reflects 144
hyperdirect cortico-subthalamic signals, coupling the premotor cortex (e.g., inferior frontal regions 145
- IFG, specially the right IFG) and the supplementary motor areas (SMA) with the STN (Oswal et 146
al., 2021; Herz et al., 2023; Cao et al., 2024). Changes in time-frequency provide neural signatures 147
for different aspects of speech production, which may help unveil the neurobiological failures that 148
contribute to stuttering, mainly around word-boundaries. Clarifying how oscillatory signals encode 149
word-to-word transitions in continuous speech may therefore open new venues to understand 150
different neural dysfunctions that can contribute to future research on motor-speech disorders 151
(Neef et al., 2018). 152
In the present study we investigate EEG time-frequency changes during paced overt speech 153
production to compare within-word and between -word syllabic transitions , in the context of 154
preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for thisthis version posted March 25, 2026. ; https://doi.org/10.64898/2026.03.23.713128doi: bioRxiv preprint
pseudowords. Fluent p articipants produced pseudo-word pairs forming a sequence of six 155
consonant-vowel (CV) syllables, at the pace of a visual metronome. Importantly, pseudo-word pairs 156
consisted of either a two-syllable word followed by a four-syllable word (i.e., condition 2+4) or two 157
three-syllable words (i.e., condition 3+3). This experimental design enabled the direct comparison 158
of neural activity at syllable-to-syllable and word-to-word transitions in a controlled and time-locked 159
manner. In turn, p articipants produced the same exact syllables, in the same exact sequence 160
position, but under different relative positions within a word ( i.e., first versus last syllable). 161
Accordingly, we focus on EEG time-frequency changes time-locked to the production of the third 162
syllable, as this syllable is an event that corresponds to a word -to-word transition in the 2+4 163
condition and a within-word syllable-to-syllable transition in 3+3 condition. 164
165
2. Methodology 166
167
2.1 Participants 168
Twenty-one fluent Portuguese-speaking adults (14 females, 5 males), aged between 18 and 53 169
years (Mean ± Standard Deviation (SD) = 29.4 ± 9.6), participated in the study. Two participants 170
were excluded due to technical errors during the EEG-audio recordings. All participants were right-171
handed, had completed secondary education or higher education (bachelor’s, master’s, or doctoral 172
level) and reported normal hearing and no history of psychiatric, neurological, or language-related 173
disorders. Participation was voluntary, and no compensation was given for taking part in the study. 174
The study was approved by the ethical committee of University of Algarve, Portugal , and 175
performed in accordance with the Declaration of Helsinki and Oviedo Convention. All participants 176
provided written informed consent prior to testing. 177
178
2.2 Stimuli 179
Stimuli consisted of sequences of two pseudo-words designed to isolate neural activity related to 180
syllable- and word-level transitions (see Figure 1B). Each sequence contained two pseudo-words 181
and conformed to one of two structures: condition 2+4 (e.g., náfa dacalána), in which a disyllabic 182
pseudo-word was followed by a tetrasyllabic pseudo-word, and condition 3+3 (e.g., náfada calána), 183
in which two trisyllabic pseudo-words were presented. Both structures were matched for length (six 184
syllables per trial). All syllables followed a CV structure ( ‘ba’, ‘da’, ‘fa’, ‘ma’, ‘na’, ‘sa’, ‘ta’, ‘ca’, ‘la’, 185
and ‘ga’) corresponding to the International Phonetic Alphabet (IPA) consonant sounds /b/, /d/, /f/, 186
/m/, /n/, /s/, /t/, /k/,/l/, and /g/ followed by the vowel /a/. We minimised bilabial closures at and before 187
the target syllable 3 to reduce perioral EMG while maintaining clear tongue -tip versus tongue-188
dorsum articulations. Labial consonants were restricted to speech positions after syllable 3; i.e., no 189
labial articulations occurred at syllable position 2 or 3 to reduce EMG artifacts on the EEG signal 190
preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for thisthis version posted March 25, 2026. ; https://doi.org/10.64898/2026.03.23.713128doi: bioRxiv preprint
(Whitham et al., 2007; Stepp, 2012). The third syllable was always tongue-driven (using /d/ or /k/ 191
consonants). The pseudo-words were phonotactically legal in Portuguese and carried stress on 192
the first syllable of the first pseudo-word and the penultimate syllable of the second pseudo-word, 193
mirroring natural stress patterns while avoiding lexical or semantic familiarity effects (Arruda et al., 194
2023). Hence, the stress pattern of the individual syllables was the same for all stimuli (and across 195
conditions). Finally, written stimuli were presented in lowercase text using a uniform sans-serif font 196
(size 24), centred on the screen with a grey background using Presentation® software 197
(Neurobehavioral Systems, Inc., Berkeley, CA, USA; www.neurobs.com). 198
2.3 Experimental Procedure 199
The experiment took place in an electrically shielded EEG laboratory at the Cognitive Neuroscience 200
Group of University of Algarve. EEG data were recorded using a 128-channel BioSemi ActiveTwo 201
system (BioSemi B.V., Amsterdam, the Netherlands) with electrodes positioned according to the 202
international 10/20 system. The vertex sensor (i.e., Cz) was centred between the inion and the 203
nasion anatomical fiducials on the sagittal plane, and the midpoint between the ears on the coronal 204
plane. Conductive gel was applied to each electrode site using a blunt-tipped plastic syringe, and 205
electrode impedance was verified to minimize offsets, in line with BioSemi specifications. Once 206
EEG preparation was completed, participants seated comfortably in front of a computer screen at 207
approximately 70 cm viewing distance. 208
209
The experiment consisted of 160 trials in total, divided into 4 blocks of 40 trials each. Each block 210
contained five repetitions of 4 sequences from condition 2+4 and 4 of condition 3+3, fully 211
randomised. Within a block, the same sequence was never presented under both conditions. 212
Hence, blocks differed in how syllables were grouped, providing variation while preserving the 213
same underlying syllabic material across blocks. The order of blocks was counterbalanced across 214
participants, ensuring that all participants completed both conditions under comparable exposure 215
and repetition constraints . Following EEG cap placement and impedance checks, participants 216
began each trial with a practice phase followed by a paced production using a visual metronome 217
(see Figure 1A). The practice phase had the target pseudo -word pair written on the screen, and 218
participants repeatedly uttered the target sequence three times. Upon completion of the trial 219
practice phase, participants pressed the ‘space bar’ button to signal their readiness to proceed to 220
the paced production phase. In the paced production phase, the written sequence disappeared 221
from the screen, and a green fixation cross appeared, flashing every 700 ms (1.43 Hz). Participants 222
were instructed to produce one syllable per flash, aligning their speech with the visual rhythm. 223
Because the lexical boundary shifted across sequence structures, the transitions flanking syllable 224
3 differed by condition. In condition 2+4, the transitions between syllables 2 and 3 were between-225
word transitions and in condition 3+3, were within-word transitions. Conversely, the transitions 226
between syllables 3 and 4 were within -word transitions in condition 2+4 and between -word 227
preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for thisthis version posted March 25, 2026. ; https://doi.org/10.64898/2026.03.23.713128doi: bioRxiv preprint
transitions in condition 3+3. This counter balancing enabled within-item contrast of between-word 228
vs within-word syllabic transitions while controlling segmental content and pacing . 229
230
231
Figure 1 232
A) Experimental design. ITI – inter-trial-interval. B) Stimuli. Two sequence types were used: Condition 2+4 233
and Condition 3+3. C) Behavioural results corresponding to speech onsets flanking syllable 3. Red and green 234
dots depict individual trial syllable speech timings, relative to syllable 3, per subject (red = syllable 2; green 235
= syllable 4); vertical lines mark the metronome timings. Mean (SD) speech timings are shown at the top. D) 236
Averaged speech-envelope responses. Group speech envelope traces time-locked to syllable 3 (grey = 237
individual participants; coloured = condition averages). E) Regions of interest. Bilateral inferior frontal gyrus 238
(IFG), superior temporal gyrus (STG), and supplementary motor areas (SMA) used for hypothesis-driven 239
visualisation in ICA cluster selection. 240
241
242
243
244
preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for thisthis version posted March 25, 2026. ; https://doi.org/10.64898/2026.03.23.713128doi: bioRxiv preprint
2.4 Data Recording and Pre -processing 245
Speech responses were simultaneously recorded with EEG at 16384 Hz (16-bit resolution) using 246
a Shure SM58 microphone placed approximately 15 cm from the participants’ mouth. Continuous 247
EEG was recorded from the 128-channel BioSemi ActiveTwo system (BioSemi B.V., Amsterdam, 248
the Netherlands), with a sampling rate of 512 Hz. Electrodes were arranged according to the 249
international 10/20 system. The BioSemi CMS/DRL system served as the reference and ground 250
during acquisition. Audio data were processed to identify speech response onsets. First, speech 251
recordings were subjected to a Hilbert transformation, and the resulting amplitude envelope was 252
low-passed filtered at 8 Hz and resampled to 256 Hz. Within each trial, six local envelopes (one 253
per syllable) were detected to provide syllable-level temporal markers for identifying the production 254
of the 6 syllables composing our stimuli. Custom Matlab routines were used for audio processing. 255
EEG preprocessing was performed in EEGLAB (Delorme and Makeig, 2004), including custom-256
made MATLAB scripts (version R2021b; MathWorks, Natick, MA). The continuous EEG recordings 257
were i) resampled to 256 Hz and ii) high-pass filtered at 1 Hz to remove slow signal drifts. iii) Noisy 258
and flat-line channels were identified and removed using the ‘clean_rawdata’ EEGLAB function, 259
after which data were iv) re-referenced to the average of all remaining channels. Across datasets, 260
a mean of 97.9 channels ( SD = 9.6) remained following bad -channel rejection. Next, v) 261
independent component analysis (ICA; INFOMAX) was performed, and components reflecting 262
ocular, muscular, or other stereotypical artefacts were automatically identified and removed using 263
‘ICLabel’ and ‘ICFlag’ EEGLAB functions (Pion-Tonachini, Kreutz-Delgado and Makeig, 2019). vi) 264
Channels previously excluded were subsequently interpolated back into the dataset using the 265
’interp’ method. vii) Following ICA cleaning, the data were low-pass filtered at 70 Hz and a band-266
stop (48-52 Hz) filter was applied to further suppress line-power noise (i.e., 50 Hz energy). Next, 267
we segmented the continuous EEG into epochs from -1 to +1 s ecs centred on the third -syllable 268
production event. Time-locking our EEG analysis to the third syllable minimises contamination from 269
sequence-onset preparatory motor activity, while still providing symmetric context for preceding 270
and following syllable transitions. Although the speech envelope peak can lag the true acoustic 271
onset by a small amount, the same alignment procedure and fixed pacing were used in both 272
conditions, and behavioural timing (Figure 1D) confirmed consistent alignment across participants. 273
Event-related spectral perturbations (ERSPs) were computed for each trial and independent 274
component (IC), providing a time-frequency representation of oscillatory time-frequency changes 275
across conditions. 276
277
278
2.5 Data Analysis and Statistics 279
280
Independent Component Analysis and Source Clustering 281
282
Equivalent current dipoles were estimated for each independent component using a standard 283
preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for thisthis version posted March 25, 2026. ; https://doi.org/10.64898/2026.03.23.713128doi: bioRxiv preprint
boundary element head model. Components with residual variance greater than 15% were 284
excluded. Next, dipoles were computed for each IC component and participant. Finally, 20 clusters 285
were made by matching dipole locations and spectral activity between 3 and 40 Hz (see Figure 2). 286
Cluster sizes ranged from 19 to 197 ICs (M = 93.1, SD = 40.0). With this approach, we took 287
advantage of the high-density EEG montage (i.e., 128 channels), by performing group statistics at 288
the EEG source level instead of relying on EEG -channel signals, which are known to have low 289
spatial specificity (Delorme & Makeig, 2004 ). 290
291
292
Figure 2 293
Independent Component (IC) clusters, and centroid location (red sphere) in Talairach coordinates. Blue 294
spheres are individual ICs forming each cluster. 295
preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for thisthis version posted March 25, 2026. ; https://doi.org/10.64898/2026.03.23.713128doi: bioRxiv preprint
Time-Frequency Analysis 296
297
EEG was segmented into epochs from -1 to +1 secs aligned to the speech response time of the 298
third syllable (target syllable) of the sequence. Time-locking our EEG analysis to the third syllable 299
minimises contamination from sequence -onset preparatory motor activity, while still providing 300
symmetric context for preceding and following syllable transitions. Although the speech-envelope 301
peaks may lag true acoustic onset by a small amount, the same alignment procedure and fixed 302
pacing were used in both conditions, and behavioural timing (Figure 1D) confirmed consistent 303
alignment across participants. Event-related spectral perturbations (ERSPs) were computed for 304
each trial and independent component (IC), providing a time -frequency representation of time-305
frequency power changes across conditions. 306
307
ERSPs were averaged within each cluster and assessed statistically using a permutation test 308
(based on 2,000 condition -label permutations). For hypothesis -driven visualisation, we report 309
clusters corresponding to bilateral IFG, bilateral STG, and bilateral SMA (Figure 1E); example 310
ERSPs are shown for L-IFG (Cluster 1), R-IFG (Cluster 3), and L-STG (Cluster 8) (see figure 3). 311
ERSPs for all 20 clusters are provided in figure 5. Multiple comparisons over the time-frequency 312
bins were corrected using FDR (q = 0.05) (Benjamini and Hochberg, 1995) (i.e., 200 time-points 313
by 68 frequencies). Behavioural contrasts were evaluated with two-tailed, paired t-tests. We refer 314
to decreases in band-limited power as event-related desynchronisation (ERD) and increases as 315
event-related synchronisation (ERS), reporting frequency bands and time windows relative to the 316
epoch average (Delorme and Makeig, 2004). 317
preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for thisthis version posted March 25, 2026. ; https://doi.org/10.64898/2026.03.23.713128doi: bioRxiv preprint
318
319
Figure 3 320
ERSPs in the three clusters matching the selected ROIs (Clusters 1, 3, and 8). First and second left time-321
frequency plots depict the grand -averaged ERSPs per condition. Third and fourth plots depict ERSP 322
statistics: uncorrected effects (p<0.05) and FDR-corrected (q = .05). A) left-IFG cluster. B) right-IFG cluster. 323
C) left-STG cluster. 324
325
326
3 Results 327
328
3.1 Behavioural Results 329
330
Participants successfully aligned their speech production with the visual metronome (see Figure 331
1C-D). Timing between successive syllables around syllable 3 differed slightly, but significantly by 332
condition. The interval between syllable 2 and 3 was significantly shorter in the 3+3 condition in 333
comparison to the 2+4 condition, mean (SD) = 17 ms (34.49), t(18) = 2.15, p = .045. The interval 334
between syllable 3 and 4 was longer in 3+3 than in 2+4, mean (SD) = 29 ms (43.75), t(18) = -2.90, 335
p = 0.010. Group speech-envelope traces time-locked to syllable 3 confirmed consistent alignment 336
of the third -syllable envelope peak across participants and conditions (Figure 1D). 337
preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for thisthis version posted March 25, 2026. ; https://doi.org/10.64898/2026.03.23.713128doi: bioRxiv preprint
338
339
3.2 Time-frequency results (Event -Related Spectral Perturbations) 340
341
Time-frequency differences were analysed within the three ROI clusters (Clusters 1, 3, and 8; see 342
Figure 2). EEG statistics are based on permutation -based condition tests (2,000 label 343
permutations) with FDR for multiple comparisons correction, across time-frequency bins (q = .05). 344
Consistent with our analysis plan, bilateral IFG, bilateral STG, and bilateral SMA were treated as 345
confirmatory ROIs (FDR correction within time-frequency bins). No other ROI clusters yielded FDR-346
corrected effects. For completeness, ERSPs for all 20 clusters are shown in figure 5. 347
348
The left IFG (Cluster 1; tal: -41, 46, 6) showed largely similar ERSPs across conditions, with only 349
a small, isolated FDR -corrected significant alpha patch (~ 10 -14 Hz) about 450 ms prior to the 350
target syllable 3 (Figure 3A). Broader uncorrected differences (p<0.05) did not survive FDR 351
correction. In the right IFG (Cluster 3; tal: 49, 39, 9), a robust pre -pk3 modulation survived FDR 352
correction, spanning ~ 4 -30 Hz (theta/alpha into low beta; Figure 3B, FDR panel). Power 353
modulations were greater for the 2+4 condition in comparison to the 3+3 condition, with additional 354
weaker patches evident at the uncorrected level in the early post-syllable-3 window. In the left STG 355
(Cluster 8; tal: -59, -39, 3), we observed two brief FDR-significant time-frequency patches (figure 356
3C): a pre-syllable 3 alpha patch (~ 9-11 Hz, [-400 to -300] msecs) and a short post-syllable 3 high-357
beta patch (~ 26-30 Hz, [0 to +100] msecs). Outside these windows, no sustained corrected effects 358
were present. The right STG showed no FDR -corrected differences (see figure 3 and 4 ). 359
Considering all 20 clusters, the clearest FDR-corrected modulation emerged in right IFG, indicating 360
a right-lateralised inferior frontal preparatory signature preceding the target syllable 3. Left IFG and 361
left STG contributed only limited, and no other clusters yielded FDR-corrected effects. See figure 362
4 for a summary of the significant ERSP contrasts in the confirmatory ROIs. 363
364
365
366
preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for thisthis version posted March 25, 2026. ; https://doi.org/10.64898/2026.03.23.713128doi: bioRxiv preprint
367
368
Figure 4 369
ERSP summary per frequency band (delta, theta, alpha, low beta and high beta) per cluster (cluster 1 – left 370
IFG, cluster 2 – right IFG, cluster 8 – left STG). Blue line is condition 2+4 and red line is condition 3+3. 371
Shaded areas around the blue and red lines represent standard errors from the group mean. Black horizontal 372
bars on top of each cluster-band plot depict uncorrected p values (p<0.05), and the yellow line depict FDR 373
corrected p values (q<0.05). 374
375
376
377
4 Discussion 378
379
Although language is a hierarchically organized communication system, segmental boundaries in 380
natural speech are often concealed due to coarticulation. The perceptual challenge for a listener is to 381
recover speech boundaries from continuously presented speech , which involves specific neural 382
oscillations (Ding et al., 2015). In speech production , the same hierarchical levels become relevant 383
for understanding motor speech disorders such stuttering, but their neural signatures remain 384
underspecified (Neef et al., 2015). Here, we investigated differences in cortical oscillations across 385
different levels of hierarchy -syllable forming word boundary vs. within -word syllable boundary . 386
Clustering of ICA components , obtained from high-density EEG, identified three clusters within our 387
preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for thisthis version posted March 25, 2026. ; https://doi.org/10.64898/2026.03.23.713128doi: bioRxiv preprint
regions of interest , which showed significantly different time-frequency signatures when comparing 388
within-word and between-word syllabic transitions: the left and right IFG and the left STG (Figure 2). 389
The importance of these clusters has been previously identified in the literature, as a distinguishing 390
part of the language and speech network that supports speech planning versus monitoring, 391
respectively (Hickok, 2012; Tremblay, Deschamps and Gracco, 2016). Across these clusters, the most 392
robust differences in modulation were in the right IFG spanning both theta and beta oscillatory bands 393
(Figure 3). 394
We observed a robust effect spanning theta and beta oscillations around 500 ms prior to the target 395
syllable, which was a between-word transition in the 2+4 condition and a within-word syllabic transition 396
in the 3+3 condition. This effect indicates that frontal control networks adjust preparatory states 397
according to the hierarchical role of the upcoming syllable, maintaining within-word chunking versus 398
reconfiguration at a boundary. This pattern aligns with reports that right-frontal beta oscillations index 399
preparatory and sequential control during speech (Pfurtscheller and Lopes da Silva, 1999; Weiss and 400
Mueller, 2012; Piai et al., 2015). Furthermore, theta oscillations may contribute to joint operations with 401
beta rhythms, namely in sensory prediction (Arnal and Giraud, 2012) or cortico-basal ganglia circuits 402
(Oswal et al., 2021) . Importantly, the right inferior frontal cortex is central in inhibitory control via 403
hyperdirect projections to the subthalamic nucleus, which is expressed in beta neural rhythms 404
(Nambu, 2011; Jorge et al., 2022). 405
406
These findings have important implications for future research on motor speech disorders, particularly 407
stuttering (Rocha, Carmona and Correia, 2025) . In stuttering, sensorimotor processes that support 408
fluent transitions between speech units are often disrupted, leading to higher dysfluency rates at word, 409
phrase, and sentence boundaries compared to within-word syllabic transitions (Bloodstein and Ratner, 410
2008; Buhr and Zebrowski, 2009). Neuroimaging studies indicate that atypical function in the cortico-411
basal ganglia circuits, including both the direct and indirect pathways, may contribute to speech 412
disruptions (Alm, 2004; Giraud et al., 2008), and recent work highlights the HDP as crucial for inhibitory 413
control and preparatory motor regulation in speech (Nambu, 2011; Neef et al., 2015) . The HDP 414
conveys rapid signals from frontal inferior cortical regions to the STN, enabling stopping and 415
adjustment of ongoing motor commands, which supports sensorimotor integration and accurate 416
coordination of speech movements (Whillier et al., 2018; Usler, 2022) . Dysfunction in these circuits 417
may impair the ability to reset or prepare the motor cortex at motor planning boundaries (e.g., word 418
boundaries), providing a possible explanation for why stuttering often emerges at transitions between 419
hierarchical speech units. 420
421
Although the focus of this study was on the segmental level of speech production, prosodic information 422
can also define word boundaries. In this study, the syllables marking a word boundary also mark a 423
difference in prosody. Other studies implicate the right IFG in the perceptual categorization of 424
emotionally-valenced prosodic information (Sammler et al., 2015; Zhang et al., 2026). The right STG 425
preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for thisthis version posted March 25, 2026. ; https://doi.org/10.64898/2026.03.23.713128doi: bioRxiv preprint
is also implicated in prosody and modulated by its emotional salience (Sammler et al., 2015), however, 426
right STG condition-differences did not emerge in this study. One possibility is that the right IFG is part 427
of a dorsal stream (Hickok and Poeppel, 2007; Ausilio, Craighero and Fadiga, 2012; Correia, Jansma 428
and Bonte, 2015) associated with sensorimotor transformations, which can include information about 429
prosody and may therefore inherently contribute to decisions about word boundaries. Nevertheless, 430
the right IFG emerged in the comparison between the 2+4 and 3+3 conditions, where stress pattern 431
was held constant. Hence, prosody differences cannot fully account for the right IFG activation 432
differences found between conditions, here. The results are more likely consistent with a role for the 433
right IFG in a basal-ganglia cortical loop that supports segmental sequences forming larger units, 434
namely word units. 435
Several possible limitations should be considered. First, the sample size (N=19), although typical for 436
EEG studies may limit sensitivity to smaller effects between condition 2+4 and 3+3, especially given 437
the need to correct for multiple time -frequency bins tested per IC cluster. Second, within -word 438
transitions were 17 milliseconds shorter between syllable s 2 and 3, and 29 m illiseconds shorter 439
between syllables 3 and 4. These timing differences may also contribute to the ERSP estimates. Third, 440
paced speech constrains generalization of the findings to more natural and prosodically richer speech 441
production; but strict pacing and pseudo-words enhance control and timing, allowing a more controlled 442
analysis of EEG data in epochs less likely contaminated by speech responses (i.e., speech -related 443
EMG).. Finally, while the most parsimonious interpretation of the results in the right IFG cluster is 444
related to the cortico-basal ganglia circuit underlying the HDP, we cannot fully rule out other factors 445
confirming this as its sole role. Scalp EEG has limited coverage for subcortical regions (i.e., reduced 446
signal-to-noise-ratio from deep brain sources) and thus does not provide a direct measure . In the 447
future, simultaneous extracranial and intracranial EEG recordings that more directly target subcortical 448
regions (e.g., from the STN in PD patients undergoing deep -brain stimulation - DBS) can provide 449
additional insights into the interpretation of the results. Altogether, the behavioural and neural data 450
indicate that hierarchical structure shapes both the timing and pre-speech cortical dynamics of speech 451
production. The right -lateralised IFG oscillatory signature preceding the target syllable suggests 452
predictive control mechanisms tuned to whether the upcoming syllable completes or starts a word, 453
complementing accounts of hierarchical temporal scaffolding in speech perception and production 454
(Ghitza, Giraud and Poeppel, 2013; Ding et al., 2015; Piai et al., 2015). 455
456
457
458
459
preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for thisthis version posted March 25, 2026. ; https://doi.org/10.64898/2026.03.23.713128doi: bioRxiv preprint
460
461
462
Figure 5 463
ERSPs in all the 20 clusters (see Figure 2 for cluster locations). First and second left time-frequency plots 464
depict the grand -averaged ERSPs per condition. Third and fourth plots depict ERSP statistics based on 465
permutation testing (2,000 label permutations): third plot is uncorrected effects (p<0.05), and fourth plot is 466
FDR-corrected ( q = .05). 467
468
469
470
preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for thisthis version posted March 25, 2026. ; https://doi.org/10.64898/2026.03.23.713128doi: bioRxiv preprint
References
471
Alm, P. A. (2004) ‘Stuttering and the basal ganglia circuits : a critical review of possible relations’, 472
37, pp. 325 –369. doi: 10.1016/j.jcomdis.2004.03.001. 473
Archila-Meléndez, M. E. et al. (2020) ‘Combining Gamma With Alpha and Beta Power Modulation 474
for Enhanced Cortical Mapping in Patients With Focal Epilepsy’, Frontiers in Human Neuroscience, 475
14, p. 555054. doi: 10.3389/fnhum.2020.555054. 476
Arnal, L. H. and Giraud, A. L. (2012) ‘Cortical oscillations and sensory predictions’, Trends in 477
Cognitive Sciences. Elsevier Ltd, 16(7), pp. 390 –398. doi: 10.1016/j.tics.2012.05.003. 478
Assaneo, M. F. and Poeppel, D. (2018) ‘The coupling between auditory and motor cortices is rate-479
restricted: Evidence for an intrinsic speech -motor rhythm’, Science Advances, 4(2). doi: 480
10.1126/sciadv.aao3842. 481
Ausilio, A. D., Craighero, L. and Fadiga, L. (2012) ‘The contribution of the frontal lobe to the 482
perception of speech’, Journal of Neurolinguistics . Elsevier Ltd, 25(5), pp. 328 –335. doi: 483
10.1016/j.jneuroling.2010.02.003. 484
Benjamini, Y. and Hochberg, Y. (1995) ‘Controlling the False Discovery Rate: A Practical and 485
Powerful Approach to Multiple Testing’, Journal of the Royal Statistical Society Series B: Statistical 486
Methodology , 57(1), pp. 289 –300. doi: 10.1111/j.2517 -6161.1995.tb02031.x. 487
Bloodstein, O. and Ratner, N. B. (2008) A Handbook on Stuttering . New York: Delmar. 488
Buhr, A. and Zebrowski, P. (2009) ‘Sentence position and syntactic complexity of stuttering in early 489
childhood : A longitudinal study’, 34, pp. 155 –172. doi: 10.1016/j.jfludis.2009.08.001. 490
Bürki, A. et al. (2016) ‘Sequential processing during noun phrase production’, Cognition, 146, pp. 491
90–99. doi: 10.1016/j.cognition.2015.09.002. 492
Cao, C. et al. (2024) ‘Low-beta versus high-beta band cortico-subcortical coherence in movement 493
inhibition and expectation’, Neurobiology of Disease , 201, p. 106689. doi: 494
10.1016/j.nbd.2024.106689. 495
Cao, L., Thut, G. and Gross, J. (2017) ‘The role of brain oscillations in predicting self -generated 496
sounds’, NeuroImage , 147. doi: 10.1016/j.neuroimage.2016.11.001. 497
Chang, S. E. and Guenther, F. H. (2020) ‘Involvement of the Cortico -Basal Ganglia-498
Thalamocortical Loop in Developmental Stuttering’, Frontiers in Psychology , 10(January). doi: 499
10.3389/fpsyg.2019.03088. 500
Christoffels, I. K. et al. (2011) ‘The sensory consequences of speaking: Parametric neural 501
cancellation during speech in auditory cortex’, PLoS ONE, 6(5). doi: 502
10.1371/journal.pone.0018307. 503
Correia, J. M., Jansma, B. M. B. and Bonte, M. (2015) ‘Decoding articulatory features from fmri 504
responses in dorsal speech regions’, Journal of Neuroscience , 35(45), pp. 15015 –15025. doi: 505
10.1523/JNEUROSCI.0977-15.2015. 506
Crapse, T. B. and Sommer, M. A. (2008) ‘Corollary discharge across the animal kingdom’, Nature 507
Reviews Neuroscience, 9(8), pp. 587 –600. doi: 10.1038/nrn2457. 508
Delorme, A. and Makeig, S. (2004) ‘EEGLAB: an open source toolbox for analysis of single -trial 509
EEG dynamics’, Journal of Neuroscience Methods , 13, pp. 9–21. 510
preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for thisthis version posted March 25, 2026. ; https://doi.org/10.64898/2026.03.23.713128doi: bioRxiv preprint
Ding, N. et al. (2015) ‘Cortical tracking of hierarchical linguistic structures in connected speech’, 511
Nature Neuroscience , 19(1), pp. 158 –164. doi: 10.1038/nn.4186. 512
Ford, J. M., Roach, B. J. and Mathalon, D. H. (2010) ‘Assessing corollary discharge in humans 513
using noninvasive neurophysiological methods’, Nature Protocols, 5(6), pp. 1160 –1168. doi: 514
10.1038/nprot.2010.67. 515
Ghazanfar, A. A. and Eliades, S. J. (2014) ‘The neurobiology of primate vocal communication’, 516
Current Opinion in Neurobiology. Elsevier Ltd, 28, pp. 128–135. doi: 10.1016/j.conb.2014.06.015. 517
Ghitza, O., Giraud, A. L. and Poeppel, D. (2013) ‘Neuronal oscillations and speech perception: 518
Critical-band temporal envelopes are the essence’, Frontiers in Human Neuroscience, 6(JAN), pp. 519
1–4. doi: 10.3389/fnhum.2012.00340. 520
Giraud, A. L. et al. (2008) ‘Severity of dysfluency correlates with basal ganglia activity in persistent 521
developmental stuttering’, Brain and Language , 104(2), pp. 190 –199. doi: 522
10.1016/j.bandl.2007.04.005. 523
Giraud, A. L. and Poeppel, D. (2012) ‘Cortical oscillations and speech processing: Emerging 524
computational principles and operations’, Nature Neuroscience . Nature Publishing Group, 15(4), 525
pp. 511–517. doi: 10.1038/nn.3063. 526
de Heer, W. A. et al. (2017) ‘The hierarchical cortical organization of human speech processing’, 527
Journal of Neuroscience , 37(27), pp. 6539 –6557. doi: 10.1523/JNEUROSCI.3267 -16.2017. 528
Herz, D. M. et al. (2023) ‘Dynamic modulation of subthalamic nucleus activity facilitates adaptive 529
behavior’, PLOS Biology. Edited by A. Gail, 21(6), p. e3002140. doi: 10.1371/journal.pbio.3002140. 530
Hickok, G. (2012) ‘Computational neuroanatomy of speech production’, Nature Reviews 531
Neuroscience, 13(2), pp. 135 –145. doi: 10.1038/nrn3158. 532
Hickok, G. and Poeppel, D. (2007) ‘The cortical organization of speech processing’, Nature 533
Reviews Neuroscience, 8(5), pp. 393 –402. doi: 10.1038/nrn2113. 534
Houde, J. F. and Jordan, M. I. (1998) ‘Sensorimotor adaptation in speech production’, Science, 535
279(5354), pp. 1213 –1216. doi: 10.1126/science.279.5354.1213. 536
Indefrey, P. (2011) ‘The spatial and temporal signatures of word production components: A critical 537
update’, Frontiers in Psychology , 2(OCT), pp. 1 –16. doi: 10.3389/fpsyg.2011.00255. 538
Indefrey, P. and Levelt, W. J. M. (2004) ‘The spatial and temporal signatures of word production 539
components’, Cognition, 92(1–2), pp. 101 –144. doi: 10.1016/j.cognition.2002.06.001. 540
Jenson, D. et al. (2014) ‘Temporal dynamics of sensorimotor integration in speech perception and 541
production: Independent component analysis of EEG data’, Frontiers in Psychology, 5(JUL), pp. 1–542
17. doi: 10.3389/fpsyg.2014.00656. 543
Jorge, A. et al. (2022) ‘Hyperdirect connectivity of opercular speech network to the subthalamic 544
nucleus’, Cell Reports. The Author(s), 38(10), p. 110477. doi: 10.1016/j.celrep.2022.110477. 545
Kittilstved, T. et al. (2018) ‘The Effects of Fluency Enhancing Conditions on Sensorimotor Control 546
of Speech in Typically Fluent Speakers: An EEG Mu Rhythm Study’, Frontiers in Human 547
Neuroscience, 12(April), pp. 1 –15. doi: 10.3389/fnhum.2018.00126. 548
Knolle, F. et al. (2019) ‘Auditory Predictions and Prediction Errors in Response to Self -Initiated 549
Vowels’, Frontiers in Neuroscience , 13(October), pp. 1 –11. doi: 10.3389/fnins.2019.01146. 550
Laganaro, M., Valente, A. and Perret, C. (2012) ‘Time course of word production in fast and slow 551
preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for thisthis version posted March 25, 2026. ; https://doi.org/10.64898/2026.03.23.713128doi: bioRxiv preprint
speakers: A high density ERP topographic study’, NeuroImage . Elsevier Inc., 59(4), pp. 3881 –552
3888. doi: 10.1016/j.neuroimage.2011.10.082. 553
Loucks, T. M. J. et al. (2007) ‘Human brain activation during phonation and exhalation: Common 554
volitional control for two upper airway functions’, NeuroImage , 36(1), pp. 131 –143. doi: 555
10.1016/j.neuroimage.2007.01.049. 556
Nambu, A. (2011) ‘Somatotopic organization of the primate basal ganglia’, Frontiers in 557
Neuroanatomy , 5(APRIL), pp. 1–9. doi: 10.3389/fnana.2011.00026. 558
Nambu, A. et al. (2023) ‘Dynamic Activity Model of Movement Disorders: The Fundamental Role of 559
the Hyperdirect Pathway’, Movement Disorders, 38(12), pp. 2145–2150. doi: 10.1002/mds.29646. 560
Nambu, A., Tokuno, H. and Takada, M. (2002) ‘Functional significance of the cortico -subthalamo-561
pallidal “hyperdirect” pathway’, Neuroscience Research, 43(2), pp. 111–117. doi: 10.1016/S0168 -562
0102(02)00027 -5. 563
Neef, N. E. et al. (2015) ‘Speech dynamics are coded in the left motor cortex in fluent speakers but 564
not in adults who stutter’, Brain, 138(3), pp. 712 –725. doi: 10.1093/brain/awu390. 565
Neef, N. E. et al. (2018) ‘Structural connectivity of right frontal hyperactive areas scales with 566
stuttering severity’, Brain, 141(1), pp. 191 –204. doi: 10.1093/brain/awx316. 567
Neef, N. E., Anwander, A. and Friederici, A. D. (2015) ‘The Neurobiological Grounding of 568
Persistent Stuttering: from Structure to Function’, Current Neurology and Neuroscience Reports , 569
15(9). doi: 10.1007/s11910 -015-0579-4. 570
Niziolek, C. A., Nagarajan, S. S. and Houde, J. F. (2013) ‘What does motor efference copy 571
represent? evidence from speech production’, Journal of Neuroscience, 33(41), pp. 16110–16116. 572
doi: 10.1523/JNEUROSCI.2137 -13.2013. 573
Orpella, J. et al. (2024) ‘Reactive Inhibitory Control Precedes Overt Stuttering Events’, 574
Neurobiology of Language , 5(2), pp. 432 –453. doi: 10.1162/nol_a_00138. 575
Oswal, A. et al. (2021) ‘Neural signatures of hyperdirect pathway activity in Parkinson’s disease’, 576
Nature Communications . Springer US, 12(1), pp. 1 –14. doi: 10.1038/s41467 -021-25366-0. 577
Peelle, J. E. and Davis, M. H. (2012) ‘Neural oscillations carry speech rhythm through to 578
comprehension’, Frontiers in Psychology , 3(SEP), pp. 1–17. doi: 10.3389/fpsyg.2012.00320. 579
Pfurtscheller, G. and Lopes da Silva, F. H. (1999) ‘Event -related EEG/MEG synchronization and 580
desynchronization: basic principles’, Clinical Neurophysiology, 110(11), pp. 1842 –1857. doi: 581
10.1016/S1388 -2457(99)00141 -8. 582
Piai, V. et al. (2015) ‘Beta oscillations reflect memory and motor aspects of spoken word 583
production’, Human Brain Mapping , 36(7), pp. 2767 –2780. doi: 10.1002/hbm.22806. 584
Pion-Tonachini, L., Kreutz-Delgado, K. and Makeig, S. (2019) ‘ICLabel: An automated 585
electroencephalographic independent component classifier, dataset, and website’, NeuroImage , 586
198, pp. 181 –197. doi: 10.1016/j.neuroimage.2019.05.026. 587
Rauschecker, J. P. and Scott, S. K. (2009) ‘Maps and streams in the auditory cortex: Nonhuman 588
primates illuminate human speech processing’, Nature Neuroscience , 12(6), pp. 718 –724. doi: 589
10.1038/nn.2331. 590
Rocha, M. F., Carmona, J. and Correia, J. M. (2025) ‘EEG responses to auditory cues predict 591
fluency variability and stuttering intervention outcome’. doi: 10.1101/2025.02.21.635719. 592
preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for thisthis version posted March 25, 2026. ; https://doi.org/10.64898/2026.03.23.713128doi: bioRxiv preprint
Sammler, D. et al. (2015) ‘Report Dorsal and Ventral Pathways for Prosody’, pp. 3079 –3085. doi: 593
10.1016/j.cub.2015.10.009. 594
Schulz, G. M. et al. (2005) ‘Functional neuroanatomy of human vocalization: An H215O PET 595
study’, Cerebral Cortex , 15(12), pp. 1835 –1847. doi: 10.1093/cercor/bhi061. 596
Sengupta, R. and Nasir, S. M. (2016) ‘The predictive roles of neural oscillations in speech motor 597
adaptability’, Journal of Neurophysiology , 115(5), pp. 2519 –2528. doi: 10.1152/jn.00043.2016. 598
Silbert, L. J. et al. (2014) ‘Coupled neural systems underlie the production and comprehension of 599
naturalistic narrative speech’, Proceedings of the National Academy of Sciences of the United 600
States of America, 111(43), pp. E4687 –E4696. doi: 10.1073/pnas.1323812111. 601
Simmonds, A. J. et al. (2014) ‘Sensory -motor integration during speech production localizes to 602
both left and right plana temporale’, Journal of Neuroscience , 34(39), pp. 12963 –12972. doi: 603
10.1523/JNEUROSCI.0336-14.2014. 604
Tian, X. and Poeppel, D. (2010) ‘Mental imagery of speech and movement implicates the dynamics 605
of internal forward models’, Frontiers in Psychology , 1(OCT), pp. 1 –23. doi: 606
10.3389/fpsyg.2010.00166. 607
Tremblay, P., Deschamps, I. and Gracco, V. L. (2016) ‘Neurobiology of Speech Production’, in 608
Hickok, G. and Small, S. L. (eds) Neurobiology of Language . Academic Press, pp. 741–750. doi: 609
10.1016/B978 -0-12-407794-2.00059-6. 610
Usler, E. R. (2022) ‘Why Stuttering Occurs: The Role of Cognitive Conflict and Control’, Topics in 611
Language Disorders , 42(1), pp. 24 –40. doi: 10.1097/TLD.0000000000000275. 612
Vos, D. M. et al. (2010) ‘Removal of muscle artifacts from EEG recordings of spoken language 613
production’, Neuroinformatics , 8(2), pp. 135 –150. doi: 10.1007/s12021 -010-9071-0. 614
Weiss, S. and Mueller, H. M. (2012) ‘“Too many betas do not spoil the broth”: The role of beta 615
brain oscillations in language processing’, Frontiers in Psychology , 3(JUN), pp. 1 –15. doi: 616
10.3389/fpsyg.2012.00201. 617
Whillier, A. et al. (2018) ‘Adults who stutter lack the specialised prespeech facilitation found in non-618
stutterers’, PLoS ONE, 13(10), pp. 1 –26. doi: 10.1371/journal.pone.0202634. 619
Zhang, Y. et al. (2026) ‘More than words : word predictability , prosody , gesture and mouth 620
movements in natural language comprehension’, (January). doi: 621
10.1098/rspb.2021.0500/879281/rspb.2021.0500.pdf. 622
623
624
625
626
627
628
629
630
preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for thisthis version posted March 25, 2026. ; https://doi.org/10.64898/2026.03.23.713128doi: bioRxiv preprint
preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for thisthis version posted March 25, 2026. ; https://doi.org/10.64898/2026.03.23.713128doi: bioRxiv preprint
preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for thisthis version posted March 25, 2026. ; https://doi.org/10.64898/2026.03.23.713128doi: bioRxiv preprint
preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for thisthis version posted March 25, 2026. ; https://doi.org/10.64898/2026.03.23.713128doi: bioRxiv preprint
preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for thisthis version posted March 25, 2026. ; https://doi.org/10.64898/2026.03.23.713128doi: bioRxiv preprint
preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission.
The copyright holder for thisthis version posted March 25, 2026. ; https://doi.org/10.64898/2026.03.23.713128doi: bioRxiv preprint