Methods
90
Participants 91
We recruited thirty-three volunteers (22 women, Mage = 23.8 years, SDage = 2.6 years, range = 18 to 92
31) with normal or corrected -to-normal vision, and with no history of epileptic attacks o r 93
neuropsychological conditions that could interfere with the examined study effects. The sample size 94
required to derive a reliable effect was estimated based on (Wolff et al., 2017), though our estimation 95
was limited by the fact that all previous work was in a WM context. One participant did not finish the 96
experiment because they were unwell, and following data inspection, two participants were removed 97
because of poor data quality due to a large number of high impedance channels, and one because of 98
stimulus trigger issues. Thus, EEG-based analyses were conducted based on 29 participants. For 99
behavioural analyses, the first four participants were excluded because of missing button press 100
triggers, which, with the further exclusion of the participant who did not complete the experiment, 101
resulted in an analysis of 28 participants (participants with noisy EEG data were included in the 102
behavioural analysis). 103
Participants were informed about the details of the experiment in advance —including its 104
duration, protocol, and methods —but were left naïve with respect to the purpose and hypotheses 105
associated with the presentation of visual pings. Participants provided their written consent, and after 106
the experiment, they were debriefed and given information about the central manipulation and 107
hypothesis upon request, and they were compensated for their time with £9 per volunteered hour. The 108
study was approved by the Ethical committee of the College of Science and Engineering of the 109
University of Glasgow (Application number: 300210113). 110
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted April 23, 2024. ; https://doi.org/10.1101/2024.04.19.590215doi: bioRxiv preprint
4
111
Stimulus and apparatus 112
The presentation of stimuli was controlled using PsychoPy (version 2021.2.3; Peirce et al., 2019 ) 113
running on Windows 10. Stimuli were presented on a CRT monitor (53.3 cm; 1024 by 768 pixels) 114
operating at a refresh rate of 60 Hz. Participants were seated in a magnetically shielded room in a 115
chinrest 65 cm from the screen , or at an approximately similar distance from the screen outside the 116
chinrest if they experienced discomfort. Throughout the experiment, a fixation cross (with a visual 117
angle of 0.44°) was presented in the centre of a constantly presented grey background (RGB = 128 118
128 128 ; Psycho Py default ). All centrally presented stimuli overrode the fixation dot. The visual 119
impulse (i.e., ping) was a single full-contrast bullseye stimulus presented at the centre of the screen 120
for 200 m illiseconds (ms; w ith a diameter of 13° and 0.31° cycles per degree ). The ping was 121
generated using MATLAB and edited using GIMP (GNU Image Manipulation Program version 122
2.10.32). 123
In the main memory task, participants learned associations between action verbs and images, 124
and were later prompted with the action verb to retrieve the associated image. The action verbs were 125
selected based on usage frequency (largely based on Linde-Domingo et al., 2019 ) and the image 126
stimulus set was a combination of 192 colour images collated across various royalty free databases, 127
including the Bank of Standardized Stimuli (BOSS, Brodeur et al., 2010), and the SUN database (Xiao 128
et al., 2010). The selected 192 images were constructed to follow a nested category structure of three 129
embedded hierarchical levels. At the top level, the set consisted of 96 objects and 96 scenes, which 130
were in turn composed at the middle level of 48 animate and 48 inanimate objects and 48 indoor and 131
48 outdoor scenes. Moving down to the bottom level, each of the middle level categories branched 132
out into 4 categories (e.g., for animate objects: birds, insects, mammals, and marine animals), each of 133
which contained 12 specific instances (e.g., twelve specific birds). We chose this nested hierarchy of 134
stimulus categories because we did not know a priori what dimension of retrieved memories would be 135
effectively decodable, so we included multiple levels of abstraction and chose one level based on pre-136
defined criteria (See Level Selection). The objects were presented on a white square matching in size 137
to scene images ( i.e., the visual degrees of all stimulus categories were 13°). Key presses were 138
registered using a standard QWERTY keyboard. 139
140
Procedure 141
The main experiment consisted of 8 blocks, each with an encoding, distractor, recall, and recognition 142
phase (Fig. 1A). In total, the main experiment lasted between approximately 45 and 6 5 minutes 143
depending on the duration of self-paced breaks and electrode impedance maintenance. Before the 144
main experiment, participants were provided with a practice run that covered each phase using 145
example verbs and images that were not used in the main experiment. A standardized set of verbal 146
instructions were provided to guide participants through the practice run. If the participant reported not 147
understanding the task or if they did not give accurate responses , the practice run and instructions 148
were repeated. Then, the main experiment commenced, throughout which EEG was acquired. At the 149
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted April 23, 2024. ; https://doi.org/10.1101/2024.04.19.590215doi: bioRxiv preprint
5
start of each experimental phase, a screen was presented with a reminder of the task instructions and 150
required response keys. 151
152
153
Figure 1 . Paradigm and behavioural results . (A) Experimental paradigm. The encoding phase 154
consisted of a word-image pair learning task. This was followed by a distractor task intended to wash 155
out working memory effects. Then, during the critical recall phase, participants were cued with words 156
to retrieve the paired image while visual perturbations (pings) were presented in 75% of trials. In a 157
fourth phase, recognition performance was tested (not displayed). (B) Average performance during 158
the recognition task for trials with and without pings , collapsing across blocks for each participant . 159
Datapoints are individual participants. (C) Average recognition performance per participant (i.e., 160
collapsing blocks). (D) Average recognition performance per block (i.e., collapsing participants) . (E) 161
Average reaction time during encoding for subsequently recognized and forgotten trials, collapsing 162
across blocks. Note: in B, C, and D, the y-axis is truncated due to high recognition performance. 163
164
165
Ping No ping
0.8
0.9
1
recognition [%]
2468
Block number
Average
per blockAverage
per participant
Average
per condition
51 0 1 52 0 2 5
Participant
Encoding EEG
Recall
Word-image pairs
+
jump
+
+
jump
+
Distractor
task
Reaction time during encoding
Recognized
Reaction time [ms]
Forgotten
n=2194
n=46
01 0 0 0 2 0 0 0 3 0 0 0 4 0 0 0 5000
E
A
DCB
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted April 23, 2024. ; https://doi.org/10.1101/2024.04.19.590215doi: bioRxiv preprint
6
In the encoding phase, participants learned to build a mental association between action 166
verbs and paired images. First, a verb was presented for 1500 ms (white, OpenSans font). Then, after 167
1000 ms, the associated image was presented until the spacebar was pressed to indicate the 168
association was encoded (with a 6000 ms limit). Then, after a 1000 ms delay, the next verb was 169
presented. During each block’s encoding phase, 10 unique verb -image pairs were learned in one 170
shot. This resulted in 80 encoded pairs across the full experiment , with the images pseudo-randomly 171
selected from the full stimulus set such as to maintain an equal distribution of top -level stimulus 172
categories (40 objects and 40 scenes) and fully random selection over nested middle and bottom 173
levels for each ping and no-ping condition. 174
The distractor phase that followed was included to flush out WM effects. Here, participants 175
performed an odd-even task lasting 20 seconds. A number between 1 and 99 was presented in the 176
centre of the screen (white, OpenSans font), and participants were instructed to press left key for odd 177
numbers, and right key for even numbers. Following a left or right key press, the next number was 178
presented immediately. Participants’ average performance was displayed at the end of the distractor 179
phase, marked as the proportion of correct responses. This data was not further analysed. 180
Next in each block , the recall phase tested our central manipulation of a ping-based visual 181
perturbation. In this phase, participants recalled the learned verb-image associations of the encoding 182
phase. First, one of the ten encoded verbs was presented for 2000 ms , serving as the retrieval cue 183
that prompted recall of the associated image. In 75% of trials, a visual impulse was presented in 184
either of three time bin s: between 500 to 833.33 ms (“early ping”), 833.34 to 1116.67 ms (“middle 185
ping”), or 1116.68 to 1500 ms (“late ping”) after the onset of the retrieval cue , with a uniform 186
distribution of possible ping times within each bin. This window was chosen on the basis that previous 187
research on cued recall paradigms suggests this is the moment of maximum memory reinstatement 188
(Staresina & Wimber, 2019) . In 25% of trials, no visual impulse was presented in order to derive a 189
baseline for statistical hypothesis testing. Participants pressed the left key to indicate that they had 190
forgotten the image associated with the verb cue, or right key to indicate they remembered it. Key 191
presses only resulted in a new trial after 1700 ms following retrieval cue onset (i.e., 200 ms after the 192
latest possible ping). With presses earlier than that , nothing happened . Participants were given a 193
visual indication that key presses were available via disappearance of the retrieval cue (at its offset; 194
2000 ms). During the recall phase, each of the 10 encoded verb -image pairs were tested four times , 195
resulting in 40 recall trials per block, and 320 trials in total , comprising 160 objects and 160 scenes . 196
Within participants, each of the four conditions —early, middle, late, and no ping —were configured to 197
present object and scene images equally often (i.e., the top-level stimulus category), with the nested 198
mid and bottom -level categories randomized. The sequence of presented s timulus level catego ries, 199
pinging conditions, and verb -image pairs was fully randomized within and across blocks to mitigate 200
order effects. For the within block randomization, while the 40 recall trials were fully randomized, we 201
ensured the same pair was never recalled twice in direct succession. 202
Finally, since the cued recall phase only included subjective memory judgments, a recognition 203
phase was included to obtain an objective measure of memory performance for the verb-image pairs. 204
During this two-alternative forced choice task, o ne of the 10 encoded verbs was presented in the 205
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted April 23, 2024. ; https://doi.org/10.1101/2024.04.19.590215doi: bioRxiv preprint
7
centre of the screen, with two images (visual angle of 7.8 °) presented underneath, one on the left-206
hand, and one on the right-hand side of the screen. Participants chose which of the two images was 207
paired with the central action verb using a left or right key press (with a 5000 ms time limit). The 208
location of the correctly paired image was randomized between the left and right location. The lure 209
image was always another old image from the immediately preceding encoding phase. Each of the 10 210
encoded verb-image pairs was tested once in a random sequence. Note that we designed this study 211
to expend most of the available study time on the recall phase to maximize the statistical power of our 212
main analysis, with the recognition phase serving chiefly as a basic check to ensure participants were 213
not skipping through the experiment without memorizing verb-image pairs. 214
215
EEG acquisition and preprocessing 216
The data was recorded using a 64-channel passive EEG BrainVision system ( BrainAmp MR; Brain 217
Products) with a sampling rate of 1000 Hz. For our recording software we used BrainVision Recorder 218
(Brain Products) . The 64 Ag/AgCl electrodes were positioned in accordance with the extended 219
international 10-20 system. Due to a necessary change in the recording system, a different EEG cap 220
type (EasyCap) was used for participants 1 to 14 (subset 1) and 15 to 33 (subset 2) . In the first 221
subset, the ground electrode was located on the back of the head, below occipital electrode Oz, and 222
two EOG channels were used to monitor eye movements (placed below and next to the eye ; VEOG 223
and HEOG). In the second subset, the ground electrode was on the midline frontal location AFz, and 224
one EOG channel was used to measure eye movements (placed below the eye ; VEOG ). 225
Furthermore, the cap used in the second subset included channels FT9 and FT10. For event related 226
potential analyses, we included only electrodes common to both caps to enable a universal 227
visualization of brain activity . Most electrode impedances were kept below 25 kiloΩ, and electrodes 228
with outlier impedances were removed during preprocessing, with their associated data interpolated 229
(see below). 230
Preprocessing was performed using FieldTrip (Oostenveld et al., 2011) in MATLAB (the 231
MathWorks). First, the continuous EEG data was split up into two datasets: one with all trials epoched 232
relative to retrieval cues, and one with trials epoched relative to pings and no -ping (defined by 233
randomly sampling ping times of the pinged trials, yielding so -called “pseudo-pings”). Put differently, 234
the data was locked once to 𝑡 = 0 defined as the retrieval cue, and once to 𝑡 = 0 defined as the 235
manipulation of interest or a baseline alternative. In both cases, the epoched trials were 4 seconds in 236
duration (-1 to 3 seconds relative to the event of interest). 237
Each dataset was filtered between 0.05 and 80 Hz and downsampled to 250 Hz . Next, bad 238
trials and channels with outlier impedance levels were manually removed via visual inspection . 239
Subsequently, eye movement and muscle artefacts were identified and removed using ICA 240
decomposition, and removed channels were interpolated using spline interpolation (with the FieldTrip 241
function ft_scalpcurrentdensity). Finally, the data was re -referenced using a common average and a 242
Laplacian method (current source density), deriving separate data structures for cue-locked and ping-243
locked analyses. 244
245
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted April 23, 2024. ; https://doi.org/10.1101/2024.04.19.590215doi: bioRxiv preprint
8
Behavioural analysis 246
The experiment was designed to result in high or even ceiling memory performance in order to obtain 247
a maximal number of successfully remembered trials, and to optimally evaluate the central hypothesis 248
of a ping -induced decodability enhancement . We report objective performance for the memory test 249
conducted in the recognition phase, both across pinging conditions (Fig. 1B), participants (Fig. 1C), 250
and across blocks (Fig. 1D). We also report subjective judgments during the recall phase, quantifying 251
how often participants report remembering versus forgetting the word -image pair. Reaction time (RT) 252
during the recall phase is uninformative, because as described in the Procedure section, the response 253
key was locked until 1700 ms after cue onset, at which point participants likely had already retrieved 254
the associated image (Staresina & Wimber, 2019) . Indeed, participants reported actively waiting for 255
response buttons to become available. Thus, we instead analysed RT during the encoding phase as a 256
function of whether the word -image pair was subsequently recognized or not. These RT data were 257
collapsed across participants and blocks (Fig. 1E). For the proceeding analyses, both subsequently 258
recognized and forgotten trials were included. 259
260
ERP Analysis 261
For the ERP analyses, only channels common to both electrode cap subsets were used . We applied 262
two types of ERP analyses, one locked to (pseudo -)pings and one to retrieval cues. FieldTrip was 263
used to downsample the data to 250 Hz and a band-pass filter between 0.2 and 40 Hz was used. The 264
data was baseline -corrected from -200 ms to 0 ms from events of interest. For ERP traces, we 265
calculated the average activity across posterior channels (C3, C4, P3, P4, O1, O2, Cz, Pz, Oz, CP1, 266
CP2, C1, C2, P1, P2, CP3, CP4, PO3, PO4, PO7, PO8, CPz, POz ). For ERP topographies, we used 267
the 61 channels common to both ERP cap types. We statistically evaluated whether pings resulted in 268
higher amplitude ERPs compared to no -ping trials using non-parametric Monte Carlo permutation 269
tests applied to each channel, correcting for multiple comparisons using Bonferroni correction as 270
implemented in FieldTrip, averaging activity from 200 to 400 ms after pseudo-pings (alpha = 0.05; 105 271
randomizations). 272
273
MVPA analysis 274
For MVPA, all EEG channels available per electrode cap type were used except EOG channels. 275
Depending on the analysis, w e trained and tested either a multi -class LDA using FieldTrip 276
(ft_timelockstatistics), or a binary -class LDA using the MVPA Light toolbox (Treder, 2020) . We 277
classified EEG data re-referenced using a Laplacian transform on the basis that it accentuates local 278
patterns (Kayser & Tenke, 2015) . All classifier analyses were performed on the recall phase, where 279
our main hypothesis could be evaluated. Unless specified otherwise, analyses were carried out on the 280
retrieval cue-locked dataset. We downsampled the data from 250 Hz to 50 Hz by applying a moving 281
average with a window length of 140 ms, moving in steps of 20 ms . During each step, a Gaussian-282
weighted mean was applied in which the centre data sample of the window was multiplied by 1, and 283
the tail samples by 0.15 (FWHM = ~81 ms). In a subsequent step, sample by sample, the data was z-284
scored across channels (i.e., setting every channel to mean = 0 and standard deviation = 1), followed 285
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted April 23, 2024. ; https://doi.org/10.1101/2024.04.19.590215doi: bioRxiv preprint
9
by training and testing using LDA . To evaluate decoder performance, w e applied k -fold cross 286
validation (5 folds, with 25 repetitions). For binary class decoding, w e used area under the receiver 287
operating characteristics curve (AUC) as a performance metric because it adjusts for class 288
imbalances (Grootswagers et al., 2017; Xie & Qiu, 2007) . For multi-class decoding, where standard 289
AUC is unavailable, we used accuracy and factored in level-specific differences in chance levels. To 290
infer decoding performance values under the null hypothesis , depending on the analysis, we either 291
used no-ping trials or ping trials with shuffled class labels (100 1st-level permutations, each with 3 292
repetitions). All analyses were restricted to the period before button presses were made (i.e., < 2000 293
ms). 294
295
Level selection 296
We used a multi -class LDA on no-ping trials to determine which retrieved stimulus category (top, 297
middle, or bottom level) is most robustly detectable in the data when our main experimental 298
manipulation was not applied. This level was then locked in for subsequent analyses that relate to our 299
key hypothesis of ping -induced decoder enhancement. We selected the level with a high baseline 300
performance to offer a conservative starting point from which we could establish whether pings are a 301
powerful tool to further enhance decodability. However, as we will see in the results, stimulus 302
selection rationales matter minimally because we found no reliable level differences in the no -ping 303
decoder across levels to begin with. For statistics, we performed a Wilcoxon rank sum test comparing 304
the empirical and shuffled decoding performance for each level, in the way described in the next 305
section. 306
307
Main analysis 308
For the statistical analysis of the main hypothesis, we used two-level permutation testing for the ping 309
versus shuffle decodability comparison, and a Wilcoxon ranked sum test for the ping versus no -ping 310
comparison. The former approach, which is based on van Bree et al., 2022 , implemented the 311
following algorithm in pseudo-code—applied window-by-window: 312
1) For each 2 nd-level permutation (105 times): Grab one random window -specific decodability 313
value from the 1st-level distribution of the 25 permutations of each participant and average the 314
result. This yields 105 permuted averages. 315
2) Generate one empirical p-value by calculating the percentile of the average empirical 316
decoding value within the distribution of permuted averages. 317
The latter approach involved taking the Wilcoxon signed-rank test between the distribution of 318
empirical decoder results and 1st-level permutation results across participants . We opted for a 319
Wilcoxon test over cluster -based methods because it makes minimal assumptions about the 320
distribution of decoding results (Wilcoxon, 1945; Grootswagers et al., 2017). For both approaches, we 321
adjusted the resulting p-values across windows for their false discover y rate (FDR). Since the p -322
values are not independent across time, we applied the approach by Benjamini & Yekutieli (2001). 323
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted April 23, 2024. ; https://doi.org/10.1101/2024.04.19.590215doi: bioRxiv preprint
10
Finally, for ping-locked analyses we restricted statistical analyses between 0 and 500 ms from 324
ping onset. For analyses locked to retrieval cue, we analysed 500 to 2000 ms from cue, which is the 325
approximate range where memory reactivation is maximal (Staresina & Wimber, 2019). 326
327
Condition-relative decoding peaks 328
In addition to our main analysis, we carried out a presumably more sensitive analysis to evaluate the 329
possibility of ping-induced decoding enhancements. We reasoned that even if visual pings do not 330
offer an enhancement of LTM decoding performance that is strong enough to emerge in a direct ping-331
to-no ping or ping -to-shuffle comparison, there could still be a weaker effect that is detectable by 332
factoring in the relative order of decoding peaks across pinging conditions. Specifically, we tested 333
whether trials with an early, middle, and late ping tended to have, respectively, earlier, later, and even 334
later decoding performance peaks. In other words, we tested to what extent decoding peaks captured 335
ping presentation order s (see Linde-Domingo et al., 2019; Mirjalili et al., 2021 for similar peak 336
selection approaches). 337
First, we took every participant’s SOA -specific decod ing time series —early, middle, and 338
late—and extracted one peak (specified below). Then, we calculated a peak order distance (POD) per 339
participant, defined as the absolute serial distance between the order of extracted peaks and true ping 340
presentation order, given by the formula: 341
342
𝑃𝑂𝐷!"!#!"$%&'()*+ = ' 𝑎𝑏𝑠(𝑝𝑒𝑎𝑘 − 𝑡𝑟𝑢𝑒) 343
344
For example, if the decoder peak came first for early ping trials (1 − 1), third for middle pings trials 345
(3 − 2), and second for late ping trials (2 − 3), this would amount to a POD of two. We divided PODs 346
by the maximum distance (4), normalizing the score between zero and one: 347
348
𝑃𝑂𝐷 = ∑ 𝑎𝑏𝑠(𝑝𝑒𝑎𝑘 − 𝑡𝑟𝑢𝑒)
𝑚𝑎𝑥𝑖𝑚𝑢𝑚 𝑑𝑖𝑠𝑡𝑎𝑛𝑐𝑒 349
350
On this distance metric, lower values indicate a closer correspondence between ping -induced peaks 351
and condition presentation order, which in turn confers stronger evidence for ping -based decoding 352
enhancement. For our statistical evaluation, we used a two-level permutation approach (similar to van 353
Bree et al., 2022). Specifically, we compared the distribution of empirical PODs with PODs calculated 354
across 106 second-level permutations, randomly grabbing from the pool of first-level shuffled decoder 355
time courses. The p-values were defined by the resulting percentile of the empirical POD within the 356
distribution of second-level shuffled PODs (one-sided test, empirical < permuted). 357
For the detection of decoder peaks in this analysis , we detected the maximum peak in the 358
derivative of the cumulative sum of decoding time series. We chose this peak detection method over 359
more standard approaches—such as simply extracting the largest peak from raw decoding series —360
because independent simulations revealed that this algorithm is most powerful at detecting true POD 361
effects, outperforming a range of competing approaches (Supplementary Materials; Section 2). 362
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted April 23, 2024. ; https://doi.org/10.1101/2024.04.19.590215doi: bioRxiv preprint
11
363
References
556
Barbosa, J., Lozano -Soldevilla, D., & Compte, A. (2021). Pinging the brain with visual impulses 557
reveals electrically active, not activity -silent, working memories. PLoS Biology , 19(10), 558
e3001436. https://doi.org/10.1371/journal.pbio.3001436 559
Benjamini, Y., & Yekutieli, D. (2001). The control of the false discovery rate in multiple testing under 560
dependency. The Annals of Statistics , 29(4), 1165 –1188. 561
https://doi.org/10.1214/aos/1013699998 562
Beukers, A. O., Buschman, T. J., Cohen, J. D., & Norman, K. A. (2021). Is Activity Silent Working 563
Memory Simply Episodic Memory? Trends in Cognitive Sciences , 25(4), 284 –293. 564
https://doi.org/10.1016/j.tics.2021.01.003 565
Brodeur, M. B., Dionne -Dostie, E., Montreuil, T., & Lepage, M. (2010). The Bank of Standardized 566
Stimuli (BOSS), a New Set of 480 Normative Photos of Objects to Be Used as Visual Stimuli 567
in Cognitive Research. PLOS ONE , 5(5), e10773. 568
https://doi.org/10.1371/journal.pone.0010773 569
Cruzat, J., Torralba, M., Ruzzoli, M., Fernández, A., Deco, G., & Soto-Faraco, S. (2021). The phase of 570
Theta oscillations modulates successful memory formation at encoding. Neuropsychologia, 571
154, 107775. https://doi.org/10.1016/j.neuropsychologia.2021.107775 572
Danker, J. F., & Anderson, J. R. (2010). The ghosts of brain states past: Remembering reactivates 573
the brain regions engaged during encoding. Psychological Bulletin , 136(1), 87 –102. 574
https://doi.org/10.1037/a0017937 575
Deuker, L., Olligs, J., Fell, J., Kranz, T. A., Mormann, F., Montag, C., Reuter, M., Elger, C. E., & 576
Axmacher, N. (2013). Memory Consolidation by Replay of Stimulus -Specific Neural Activity. 577
Journal of Neuroscience , 33(49), 19373 –19383. https://doi.org/10.1523/JNEUROSCI.0414-578
13.2013 579
Dijkstra, N., Mostert, P., Lange, F. P. de, Bosch, S., & van Gerven, M. A. (2018). Differential temporal 580
dynamics during visual imagery and perception. eLife, 7, e33904. 581
https://doi.org/10.7554/eLife.33904 582
Duncan, D. H., van Moorselaar, D., & Theeuwes, J. (2023). Pinging the brain to reveal the hidden 583
attentional priority map using encephalography. Nature Communications , 14(1), Article 1. 584
https://doi.org/10.1038/s41467-023-40405-8 585
Favila, S. E., Kuhl, B. A., & Winawer, J. (2022). Perception and memory have distinct spatial tuning 586
properties in human visual cortex. Nature Communications , 13(1), 5864. 587
https://doi.org/10.1038/s41467-022-33161-8 588
Favila, S. E., Lee, H., & Kuhl, B. A. (2020). Transforming the Concept of Memory Reactivation. 589
Trends in Neurosciences, 43(12), 939–950. https://doi.org/10.1016/j.tins.2020.09.006 590
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted April 23, 2024. ; https://doi.org/10.1101/2024.04.19.590215doi: bioRxiv preprint
18
Fritch, H. A., MacEvoy, S. P., Thakral, P. P., Jeye, B. M., Ross, R. S., & Slotnick, S. D. (2020). The 591
anterior hippocampus is associated with spatial memory encoding. Brain Research, 1732, 592
146696. https://doi.org/10.1016/j.brainres.2020.146696 593
Fuster, J. M., & Alexander, G. E. (1971). Neuron Activity Related to Short -Term Memory. Science, 594
173(3997), 652–654. https://doi.org/10.1126/science.173.3997.652 595
Goldman-Rakic, P. S. (1995). Cellular basis of working memory. Neuron, 14(3), 477 –485. 596
https://doi.org/10.1016/0896-6273(95)90304-6 597
Grootswagers, T., Wardle, S. G., & Carlson, T. A. (2017). Decoding Dynamic Brain Patterns from 598
Evoked Responses: A Tutorial on Multivariate Pattern Analysis Applied to Time Series 599
Neuroimaging Data. Journal of Cognitive Neuroscience , 29(4), 677 –697. 600
https://doi.org/10.1162/jocn_a_01068 601
Haque, R. U., Wittig, J. H., Damera, S. R., Inati, S. K., & Zaghloul, K. A. (2015). Cortical Low -602
Frequency Power and Progressive Phase Synchrony Precede Successful Memory Encoding. 603
Journal of Neuroscience , 35(40), 13577 –13586. https://doi.org/10.1523/JNEUROSCI.0687-604
15.2015 605
Haxby, J. V., Connolly, A. C., & Guntupalli, J. S. (2014). Decoding Neural Representational Spaces 606
Using Multivariate Pattern Analysis. Annual Review of Neuroscience , 37(1), 435 –456. 607
https://doi.org/10.1146/annurev-neuro-062012-170325 608
Josselyn, S. A., & Tonegawa, S. (2020). Memory engrams: Recalling the past and imagining the 609
future. Science, 367(6473), eaaw4325. https://doi.org/10.1126/science.aaw4325 610
Kamiński, J., & Rutishauser, U. (2020). Between persistently active and activity -silent frameworks: 611
Novel vistas on the cellular basis of working memory. Annals of the New York Academy of 612
Sciences, 1464(1), 64–75. https://doi.org/10.1111/nyas.14213 613
Kandemir, G., & Akyürek, E. G. (2023). Impulse perturbation reveals cross -modal access to sensory 614
working memory through learned associations. NeuroImage, 274, 120156. 615
https://doi.org/10.1016/j.neuroimage.2023.120156 616
Kandemir, G., Wilhelm, S. A., Axmacher, N., & Akyürek, E. G. (2023). Maintenance of colour 617
memoranda in activity-quiescent working memory states: Evidence from impulse perturbation 618
(p. 2023.07.03.547526). bioRxiv. https://doi.org/10.1101/2023.07.03.547526 619
Kayser, J., & Tenke, C. E. (2015). On the benefits of using surface Laplacian (Current Source 620
Density) methodology in electrophysiology. International Journal of Psychophysiology : 621
Official Journal of the International Organization of Psychophysiology , 97(3), 171 –173. 622
https://doi.org/10.1016/j.ijpsycho.2015.06.001 623
Kerrén, C., van Bree, S., Griffiths, B. J., & Wimber, M. (2022). Phase separation of competing 624
memories along the human hippocampal theta rhythm. eLife, 11, e80633. 625
https://doi.org/10.7554/eLife.80633 626
Kragel, J. E., Ezzyat, Y., Sperling, M. R., Gorniak, R., Worrell, G. A., Berry, B. M., Inman, C., Lin, J. -627
J., Davis, K. A., Das, S. R., Stein, J. M., Jobst, B. C., Zaghloul, K. A., Sheth, S. A., Rizzuto, D. 628
S., & Kahana, M. J. (2017). Similar patterns of neural activity predict memory function during 629
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted April 23, 2024. ; https://doi.org/10.1101/2024.04.19.590215doi: bioRxiv preprint
19
encoding and retrieval. NeuroImage, 155, 60 –71. 630
https://doi.org/10.1016/j.neuroimage.2017.03.042 631
Kuhl, B. A., Rissman, J., & Wagner, A. D. (2012). Multi-voxel patterns of visual category 632
representation during episodic encoding are predictive of subsequent memory. 633
Neuropsychologia, 50(4), 458–469. https://doi.org/10.1016/j.neuropsychologia.2011.09.002 634
Linde-Domingo, J., Treder, M. S., Kerrén, C., & Wimber, M. (2019). Evidence that neural information 635
flow is reversed between object perception and object reconstruction from memory. Nature 636
Communications, 10(1), 179. https://doi.org/10.1038/s41467-018-08080-2 637
Madore, K. P., & Wagner, A. D. (2022). Readiness to remember: Predicting variability in episodic 638
memory. Trends in Cognitive Sciences , 26(8), 707 –723. 639
https://doi.org/10.1016/j.tics.2022.05.006 640
Maguire, E. A. (2014). Memory consolidation in humans: New evidence and opportunities. 641
Experimental Physiology, 99(3), 471–486. https://doi.org/10.1113/expphysiol.2013.072157 642
Martín-Buro, M. C., Wimber, M., Henson, R. N., & Staresina, B. P. (2020). Alpha Rhythms Reveal 643
When and Where Item and Associative Memories Are Retrieved. The Journal of 644
Neuroscience: The Official Journal of the Society for Neuroscience , 40(12), 2510 –2518. 645
https://doi.org/10.1523/JNEUROSCI.1982-19.2020 646
Masse, N. Y., Rosen, M. C., & Freedman, D. J. (2020). Reevaluating the Role of Persistent Neural 647
Activity in Short -Term Memory. Trends in Cognitive Sciences , 24(3), 242 –258. 648
https://doi.org/10.1016/j.tics.2019.12.014 649
Merkow, M. B., Burke, J. F., & Kahana, M. J. (2015). The human hippocampus contributes to both the 650
recollection and familiarity components of recognition memory. Proceedings of the National 651
Academy of Sciences, 112(46), 14378–14383. https://doi.org/10.1073/pnas.1513145112 652
Mirjalili, S., Powell, P., Strunk, J., James, T., & Duarte, A. (2021). Context Memory Encoding and 653
Retrieval Temporal Dynamics are Modulated by Attention across the Adult Lifespan. eNeuro, 654
8(1), ENEURO.0387-20.2020. https://doi.org/10.1523/ENEURO.0387-20.2020 655
Moliadze, V., Zhao, Y., Eysel, U., & Funke, K. (2003). Effect of transcranial magnetic stimulation on 656
single-unit activity in the cat primary visual cortex. The Journal of Physiology, 553(Pt 2), 665–657
679. https://doi.org/10.1113/jphysiol.2003.050153 658
Mormann, F., Fell, J., Axmacher, N., Weber, B., Lehnertz, K., Elger, C. E., & Fernández, G. (2005). 659
Phase/amplitude reset and theta –gamma interaction in the human medial temporal lobe 660
during a continuous word recognition memory task. Hippocampus, 15(7), 890 –900. 661
https://doi.org/10.1002/hipo.20117 662
Mueller, J., Legon, W., Opitz, A., Sato, T. F., & Tyler, W. J. (2014). Transcranial Focused Ultrasound 663
Modulates Intrinsic and Evoked EEG Dynamics. Brain Stimulation , 7(6), 900 –908. 664
https://doi.org/10.1016/j.brs.2014.08.008 665
Oostenveld, R., Fries, P., Maris, E., & Schoffelen, J. -M. (2011). FieldTrip: Open source software for 666
advanced analysis of MEG, EEG, and invasive electrophysiological data. Computational 667
Intelligence and Neuroscience, 2011, 156869. https://doi.org/10.1155/2011/156869 668
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted April 23, 2024. ; https://doi.org/10.1101/2024.04.19.590215doi: bioRxiv preprint
20
Pearson, J., Naselaris, T., Holmes, E. A., & Kosslyn, S. M. (2015). Mental Imagery: Functional 669
Mechanisms and Clinical Applications. Trends in Cognitive Sciences , 19(10), 590 –602. 670
https://doi.org/10.1016/j.tics.2015.08.003 671
Peirce, J., Gray, J. R., Simpson, S., MacAskill, M., Höchenberger, R., Sogo, H., Kastman, E., & 672
Lindeløv, J. K. (2019). PsychoPy2: Experiments in behavior made easy. Behavior Research 673
Methods, 51(1), 195–203. https://doi.org/10.3758/s13428-018-01193-y 674
Rissman, J., & Wagner, A. D. (2012). Distributed representations in memory: Insights from functional 675
brain imaging. Annual Review of Psychology , 63, 101–128. https://doi.org/10.1146/annurev-676
psych-120710-100344 677
Rizzuto, D. S., Madsen, J. R., Bromfield, E. B., Schulze -Bonhage, A., Seelig, D., Aschenbrenner -678
Scheibe, R., & Kahana, M. J. (2003). Reset of human neocortical oscillations during a working 679
memory task. Proceedings of the National Academy of Sciences , 100(13), 7931 –7936. 680
https://doi.org/10.1073/pnas.0732061100 681
Schreiner, T., Petzka, M., Staudigl, T., & Staresina, B. P. (2021). Endogenous memory reactivation 682
during sleep in humans is clocked by slow oscillation -spindle complexes. Nature 683
Communications, 12(1), Article 1. https://doi.org/10.1038/s41467-021-23520-2 684
Staresina, B. P., Reber, T. P., Niediek, J., Boström, J., Elger, C. E., & Mormann, F. (2019). 685
Recollection in the human hippocampal -entorhinal cell circuitry. Nature Communications , 686
10(1), 1503. https://doi.org/10.1038/s41467-019-09558-3 687
Staresina, B. P., & Wimber, M. (2019). A Neural Chronometry of Memory Recall. Trends in Cognitive 688
Sciences, 23(12), 1071–1085. https://doi.org/10.1016/j.tics.2019.09.011 689
Stokes, M. G. (2015). “Activity -silent” working memory in prefrontal cortex: A dynamic coding 690
framework. Trends in Cognitive Sciences , 19(7), 394 –405. 691
https://doi.org/10.1016/j.tics.2015.05.004 692
Ten Oever, S., De Weerd, P., & Sack, A. T. (2020). Phase -dependent amplification of working 693
memory content and performance. Nature Communications , 11(1), 1832. 694
https://doi.org/10.1038/s41467-020-15629-7 695
Ter Wal, M., Linde-Domingo, J., Lifanov, J., Roux, F., Kolibius, L. D., Gollwitzer, S., Lang, J., Hamer, 696
H., Rollings, D., Sawlani, V., Chelvarajah, R., Staresina, B., Hanslmayr, S., & Wimber, M. 697
(2021). Theta rhythmicity governs human behavior and hippocampal signals during memory -698
dependent tasks. Nature Communications, 12(1), 7048. https://doi.org/10.1038/s41467-021-699
27323-3 700
Treder, M. S. (2020). MVPA -Light: A Classification and Regression Toolbox for Multi -Dimensional 701
Data. In Frontiers in Neuroscience (Vol. 14). 702
https://www.frontiersin.org/article/10.3389/fnins.2020.00289 703
van Bree, S., Melcón, M., Kolibius, L. D., Kerrén, C., Wimber, M., & Hanslmayr, S. (2022). The brain 704
time toolbox, a software library to retune electrophysiology data to brain dynamics. Nature 705
Human Behaviour, 6(10), 1430–1439. https://doi.org/10.1038/s41562-022-01386-8 706
Wilcoxon, F. (1945). Individual Comparisons by Ranking Methods. Biometrics Bulletin, 1(6), 80–83. 707
https://doi.org/10.2307/3001968 708
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted April 23, 2024. ; https://doi.org/10.1101/2024.04.19.590215doi: bioRxiv preprint
21
Wolff, M. J., Ding, J., Myers, N. E., & Stokes, M. G. (2015). Revealing hidden states in visual working 709
memory using electroencephalography. Frontiers in Systems Neuroscience , 9. 710
https://doi.org/10.3389/fnsys.2015.00123 711
Wolff, M. J., Jochim, J., Akyürek, E. G., Buschman, T. J., & Stokes, M. G. (2020). Drifting codes within 712
a stable coding scheme for working memory. PLOS Biology , 18(3), e3000625. 713
https://doi.org/10.1371/journal.pbio.3000625 714
Wolff, M. J., Jochim, J., Akyürek, E. G., & Stokes, M. G. (2017). Dynamic hidden states underlying 715
working-memory-guided behavior. Nature Neuroscience , 20(6), 864 –871. 716
https://doi.org/10.1038/nn.4546 717
Xiao, J., Hays, J., Ehinger, K. A., Oliva, A., & Torralba, A. (2010). SUN database: Large -scale scene 718
recognition from abbey to zoo. 2010 IEEE Computer Society Conference on Computer Vision 719
and Pattern Recognition, 3485–3492. https://doi.org/10.1109/CVPR.2010.5539970 720
Xie, J., & Qiu, Z. (2007). The effect of imbalanced data sets on LDA: A theoretical and empirical 721
analysis. Pattern Recognition, 40(2), 557–562. https://doi.org/10.1016/j.patcog.2006.01.009 722
Xue, G. (2018). The Neural Representations Underlying Human Episodic Memory. Trends in 723
Cognitive Sciences, 22(6), 544–561. https://doi.org/10.1016/j.tics.2018.03.004 724
Yang, C., He, X., & Cai, Y. (2023). Reactivating and reorganizing activity-silent working memory: Two 725
distinct mechanisms underlying pinging the brain (p. 2023.07.16.549254). bioRxiv. 726
https://doi.org/10.1101/2023.07.16.549254 727
728
729
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted April 23, 2024. ; https://doi.org/10.1101/2024.04.19.590215doi: bioRxiv preprint
22
Supplementary Materials 730
1. Event-related potential 731
732
733
Supplementary Figure 1. Retrieval cue-locked ERP . The purple trace reflects the average cue -734
locked response for each participant across posterior EEG channels. The grey horizontal line 735
represents cue onset. For more details, see the Methods section in the main text. The amplitude on 736
the y-axis is in arbitrary units. 737
738
739
740
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted April 23, 2024. ; https://doi.org/10.1101/2024.04.19.590215doi: bioRxiv preprint
23
Supplementary Figure 2. Retrieval ping-locked ERP. The purple trace reflects the average ping-741
locked response for each participant across posterior EEG channels. The grey horizontal line 742
represents ping onset. For more details, see the Methods section in the main text. The amplitude on 743
the y-axis is in arbitrary units. 744
745
746
747
748
Supplementary Figure 3. Retrieval cue-locked topographies. These topographical plots represent 749
the average cue -locked activity across participants. The colours represent the difference in EEG 750
activity before and after cue onset in arbitrary units (red colours represent activity post > activitypre and 751
vice versa for blue colours). No statistical analysis was carried out for these topographical contrasts. 752
For more details, see the Methods section in the main text. 753
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted April 23, 2024. ; https://doi.org/10.1101/2024.04.19.590215doi: bioRxiv preprint
24
754
755
Supplementary Figure 4. Retrieval ping -locked topographies (ping vs. no ping trials) . These 756
topographical plots represent the average cue -locked activity across participants. The colours 757
represent the difference in EEG activity between ping and no-ping (red colours represent activityping > 758
activityno ping and vice versa for blue colours). No statistical analysis was carried out for these 759
topographical contrasts. For more details, see the Methods section in the main text. 760
761
762
Channel Early ping (p-val) Middle ping (p-val) Late ping (p-val)
Fp1 0.032 0.616 0.246
Fpz 0.089 0.079 0.011
Fp2 0.042 0.537 0.115
AF8 0.119 0.422 0.318
AF7 0.014 0.272 0.954
AF3 0.439 0.23 0.123
AF4 0.712 0.541 0.014
F7 0.002 0.346 0.358
F5 0.068 0.439 0.33
F3 0.119 0.477 0.693
F1 0.597 0.662 0.119
Fz 0.053 0.551 0.003
F2 0.013 0.473 0
F4 0.341 0.939 0.049
F6 0.427 0.559 0.707
F8 0.131 0.826 0.825
FT8 0.001 0.097 0.049
FC6 0.177 0.142 0.78
FC4 0.962 0.176 0.881
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted April 23, 2024. ; https://doi.org/10.1101/2024.04.19.590215doi: bioRxiv preprint
25
FC2 0.245 0.276 0.09
FC1 0.969 0.503 0.881
FC3 0.176 0.279 0.28
FC5 0.011 0.083 0.112
FT7 0.002 0.298 0.127
T7 0.004 0.047 0.043
C5 0.013 0.027 0.051
C3 0.002 0.022 0.01
C1 0.148 0.049 0.04
Cz 0.144 0.155 0.114
C2 0.305 0.182 0.104
C4 0.02 0.004 0.014
C6 0.003 0 0.003
T8 0.002 0.002 0.003
TP10 0.002 0 0
TP8 0 0 0
CP6 0 0 0
CP4 0.001 0 0.001
CP2 0.006 0.001 0.001
CPz 0.019 0.006 0.01
CP1 0.272 0.002 0.003
CP3 0 0 0.001
CP5 0.002 0.001 0.001
TP7 0.002 0.008 0.001
TP9 0 0 0
P7 0 0 0
P5 0 0 0
P3 0 0 0
P1 0 0 0
Pz 0.001 0 0
P2 0.001 0 0
P4 0 0 0
P6 0 0 0
P8 0 0 0
PO8 0 0 0
PO4 0 0 0
POz 0 0 0
PO3 0 0 0
PO7 0 0 0
O1 0 0 0
Oz 0.001 0 0
O2 0 0 0
763
Supplementary Table 1. P-values associated with inset topographies in main text Fig. 2; rounded to 764
three decimal points. 765
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted April 23, 2024. ; https://doi.org/10.1101/2024.04.19.590215doi: bioRxiv preprint
26
2. Peak order analysis simulation 766
2.1 Time series simulation 767
We used MATLAB (the MathWorks) to generate time series with two components: (1) a peak at a 768
fixed time point (1000 ms), and (2) autocorrelated noise generated using a random walk procedure. 769
We matched several characteristics of the simulated time series to our empirical decoding data, 770
including the analysis period (500 to 2000 ms) , sampling rate (50 Hz) , and the number of (virtual) 771
participants (N = 29) . The signal -to-noise (SNR) ratio of the simulation was set to 1.15 , qualitatively 772
matching peaks observed in the empirical data. We found that varying the SNR does not significantly 773
alter the results. We generated 1000 trials per participant, resulting in 29000 trials in total. 774
775
2.2 Analysis 776
We included a smoothing parameter that implemented one of four smoothing methods : no filter, a 777
Gaussian filter, a Savitzky-Golay filter, and a median filter. We also included a window size for 778
smoothing, set to 10 samples for our main analysis. We compared the performance of eight peak 779
detection methods, evaluating each of them based on the absolute distance between estimated peaks 780
and true peaks—amounting to a simplified version of the peak order distance score described under 781
condition-relative decoding peaks in the main text . The winning method was locked in for our 782
empirical analysis. We tested eight peak detection methods: 783
(1) Low-pass approach, where the maximum peak was computed after a low -pass filter was 784
applied to the time series. 785
(2) Maximum value approach, which simply computed the maximum value per time series 786
regardless of whether the surrounding data was peak-like. 787
(3) Cumulative sum approach, which computed the maximum peak in the derivative of the 788
cumulative sum of the data. 789
(4) Cumulative integral approach, which computed the maximum peak in the cumulative integral 790
of the data via the trapezoidal method. 791
(5) Integral cumulative sum approach, which worked as the previous method but which operates 792
over the cumulative sum rather than raw time series. 793
(6) Wavelet transform-based method, which finds the maximum peak in a wavelet decomposed 794
version of the data. 795
(7) Hilbert transform-based method, which find the maximum peak in the amplitude fluctuations in 796
the envelope of the time series. 797
(8) Cross-correlation method, which finds the time lag with a maximal correlation between the 798
signal and iteratively shifted versions of itself. 799
800
2.3 Results 801
We found that approach 5—the i ntegral cumulative sum approach —reliably achieves low absolute 802
distance errors across parameters (Supplementary Figure 5). These results were generally 803
unchanged across adjustments of the parameters (to evaluate this, we refer to the code published 804
with this manuscript). Thus, we used approach 5 in our main peak order detection analysis. 805
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted April 23, 2024. ; https://doi.org/10.1101/2024.04.19.590215doi: bioRxiv preprint
27
806
Supplementary Figure 6. In simulated time series, the integral cumulative sum approach works best 807
for detecting a peak in noisy time series. The red circle indicates the best -performing method, and 808
yellow the second best-performing method. Errors were computed based on the absolute distance in 809
milliseconds (ms) between estimated and true peak location. 810
811
3. Class and trial number decoding simulation 812
We speculated based on a qualitative inspection of the empirical decoding results that the number of 813
trials (Ntrials) and classes (Nclasses) reduces the statistical significance of decoding results. We 814
evaluated this intuition by demonstrating using simulations that these two parameters do indeed 815
influence the variance of shuffled and empirical results, which in turn affects p-values but only if there 816
is a true effect in the data. 817
818
3.1 Time series simulation 819
Using MATLAB, w e generated one ground truth vector of class labels which represented the true 820
class structure in the simulated data. This vector contained a random sequence of integers randomly 821
grabbed between the interval 1 and Nclasses. For example, with 16 classes, the ground truth pattern 822
might have contained a sequence of [2,7,15,4,13,17] and with 2 classes a sequence of [2,2,1,2,1,2]. 823
Then, to simulate shuffled decoding results, we generated a distribution of random sequences 824
of integers identical to the ground truth procedure, but with newly generated random integers. These 825
random sequences represented shuffled decoding results and were scored based on their average 826
element-wise correspondence to the ground truth pattern —which is how decoding accuracy is 827
normally computed. For example, if the permuted vector is [2,1,2,2,1,1] and the true sequence is 828
[2,2,1,2,1,2], the accuracy would be 50% because half of the class labels correspond to the true 829
structure. Trivially, with increasing repetitions the shuffled distribution will approach chance level 830
predictions of the ground truth pattern (i.e., the expected value is exactly at 1/Nclasses). 831
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted April 23, 2024. ; https://doi.org/10.1101/2024.04.19.590215doi: bioRxiv preprint
28
Finally, to simulate empirical decoding results, we again generated a distribution of random 832
integers identical to the procedure for shuffled and ground truth decoding results. However, for these 833
data we manually injected between 0% and 60% of the ground truth pattern into the otherwise 834
random vector, effectively modulating decoding accuracy. With 0% of the ground truth injected, there 835
is no statistically detectable difference in accuracy between empirical and shuffled decoding results, 836
because the vectors are equally random. With 60%, the encoding results are substantially more 837
accurate than shuffled results, yielding above chance decoding accuracy. 838
We simplified our simulation by operationalizing the variable Ntrials as the number of elements 839
in the vector , allowing us to efficiently investigate how the number of observations influences 840
statistical tests. We also compared Nclasses = 2 and Nclasses = 16, which respectively match the number 841
of classes for top- and bottom-level category decoding in our main experiment. Both Ntrials and Nclasses 842
were independently manipulated in a 2 ∗ 2 factorial design, allowing us to evaluate the contribution of 843
each variable toward statistical outcomes (as a function of effect size). 844
845
3.2 Results 846
First, with respect to N classes, we found that increasing the number of classes reduces the spread of 847
both shuffled and empirical decoding results (Supplementary Figure 7; columns). This happens both if 848
there is no true effect in the empirical data, and when a significant proportion of the ground truth is 849
inserted into the empirical data . Second, we found that Ntrials similarly reduces the variance of both 850
shuffled and decoding results, both across low and high N classes (Supplementary Figure 7; top and 851
bottom half). Thus, we conclude that both factors modulate the likelihood of finding a significant 852
difference between empirical and shuffled results, but only if there is a true effect in the data. Indeed, 853
as we can glean from the results based on non -existent effects, the distributions of empirical and 854
shuffled will overlap regardless of N trials or Nclasses (Supplementary Figure 7; left half). In contrast, if 855
there is an effect (60% injected ground truth), both N trials and N classes independently increase the 856
distributional distance between empirical and shuffled accuracy values. 857
858
3.3 Discussion 859
We found that N trials and N classes independently reduce the variance of accuracy results, which will 860
affect statistical tests between empirical and shuffled distributions but only if there is an effect in the 861
data. As suggested in the main text, these findings suggest that statistical analyses that depend on 862
variance comparisons between empirical and shuffled distributions should be interpreted with care if it 863
is done across conditions with varying Ntrials and Nclasses. With regard to our main analysis for example, 864
the fact that the decoder based on pinged trials yields more significant decodability compared to the 865
decoder based on no-pinged trials should be interpreted with caution because there are differences in 866
Ntrials between the two conditions that could partially or fully explain this effect. More generally, we 867
found that the condition with more trials or more classes is by default more likely to yield significant p-868
values—but only if a true effect exist. 869
870
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted April 23, 2024. ; https://doi.org/10.1101/2024.04.19.590215doi: bioRxiv preprint
29
871
Supplementary Figure 7. The effects of class and trial number on decoding accuracy. Both the 872
number of classes (columns) and trials (top vs. bottom half) influences the distance between shuffled 873
and empirical distributions—but only if there is an effect in the data (left vs. right half). 874
875
These findings may be a manifestation of the classical notion of statistical power in statistical 876
analysis but within the less intuitive context of decoding accuracy . Our interpretation then is not that 877
Ntrials and N classes must necessarily be equal between conditions for a statistical comparison to be 878
meaningful. Rather, we wanted to err on the side of caution and ensure that analyses where power 879
differences could possibly explain condition differences (e.g., Fig. 3 and Fig. 4A and 4B in the main 880
text) do not inform subsequent analyses and scientific interpretations by themselves . Instead, we 881
supplemented each of the implicated analyses with additional rationale (in the case of Fig. 3) or 882
analyses that do not involve empirical -to-shuffle decoding comparisons. Indeed, Fig. 4C and Fig. 4D 883
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted April 23, 2024. ; https://doi.org/10.1101/2024.04.19.590215doi: bioRxiv preprint
30
involve direct comparisons between empirical and shuffled distributions , sidestepping the issue 884
altogether. 885
.CC-BY 4.0 International licensemade available under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is
The copyright holder for this preprintthis version posted April 23, 2024. ; https://doi.org/10.1101/2024.04.19.590215doi: bioRxiv preprint