Explainable Eeg for Auditory Attention Decoding | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Explainable Eeg for Auditory Attention Decoding Fatma ÖZCAN This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8503109/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Human communication involves the simultaneous processing of sounds. In order to function in various complex auditory conditions, the auditory system distinguishes sound sources by reducing the impact of ambient noise. Auditory attention decoding (AAD) uses brain signals to identify conversations being listened to in an environment where several people are speaking simultaneously. Our goal was to use electroencephalogram (EEG) signals for AAD in order to predict the sound being listened to based on neural activity in the auditory cortex with deep learning models. The EEG signals are converted into graphs and then presented at the input of the ResNet-101 architecture. We obtained an AUC value of 94.6% with data from 4 participants and 87.7% with data from 18 participants. We then evaluated the model using the GradCAM technique to understand which characteristics and types of EEG components are used for AAD. Finally, auditory attention is elucidated through explainability. This study shows that alpha, beta and sensorimotor rhythms are more or less intense depending on the type of speech. Although preliminary, this study shows the potential of combining an explainability to increase the robustness of EEG source analysis for AAD. Auditory attention decoding Electroencephalography Scalogram Auditory Cortex Explainable Artificial Intelligence Figures Figure 1 Figure 2 Figure 3 INTRODUCTION The human brain is capable of exceptional speed and precision in processing sensory data. In neuroscience, understanding the complexity of these processes is one of the main objectives [ 1 ]. Human communication involves the simultaneous processing of sounds [ 2 ]. In order to perform in a variety of complex listening conditions, the auditory system distinguishes sound sources by reducing the impact of ambient noise. Auditory attention decoding (AAD) uses brain signals to identify conversations being listened to in an environment where several people are talking simultaneously [ 3 ]. The ability to focus one's attention on a single speaker in an environment is an essential skill of the human auditory system [ 4 ]. Electroencephalography provides valuable information about the brain's information processing in a practical and inexpensive way. The neural correlates of selective attention can be decoded from the EEG [ 5 ]. EEG, which can reproduce the cocktail party effect, shows promise for the study of AAD. At a cocktail party, there are multiple sound sources, inducing energetic masking in the spectro-temporal space, and the content conveyed can form informational masking, which creates an additional cognitive load for the system. The auditory system must therefore separate the sources using contextual information and spatial cues at its disposal [ 6 ]. In an environment where several speakers are present, the speaker who attracts attention is selected as the one who shows the strongest correlation between neural data and the speech envelope [ 1 , 7 , 8 ]. Cortical tracking of the acoustic envelope is a mechanism whereby the electrical activity of the brain varies in response to the acoustic envelope of the stimulus. The synchronisation between the vocal envelope and neural activity in the auditory regions plays a key role in speech processing [ 9 , 10 ] Understanding speech in a noisy environment is a major challenge for people with hearing loss. Speech stimuli are therefore more valuable than clicks and tones for assessing hearing [ 11 ]. Understanding the process of solving complex perception problems will enable the development of artificial systems that can help people with sensory impairments [ 12 , 13 ]. The decoding of AAD via a brain-computer interface has seen considerable growth since its introduction by Mesgarani and Chang using electrocorticographic recordings [ 14 ]. Deep learning models have shown promising results in various tasks related to brain signal processing for AAD [ 1 , 12 , 15 ]. Recent advances in eXplainable Artificial Intelligence (XAI) offer new possibilities for AAD by enabling the learning of informative features [ 16 – 20 ] Although progress has been made in detecting auditory attention based on EEG, it is necessary to conduct trials with new methods in this field. It is still unclear whether EEG responses to speech offer an advantage in predicting speech intelligibility [ 11 ] and EEG signals generally suffer from a low signal-to-noise ratio, making it difficult to decode complex cognitive functions [ 1 ]. Furthermore, the neural mechanisms that enable selective attention dependent on sound processing in complex auditory scenes are not well understood [ 6 ]. Our objective is to predict, using current and explainable deep learning methods, the sound heard from neural activity in the auditory cortex. In our study, EEG signals recorded when two different voices are heard, on 4 and 18 subjects, will be used for AAD. This work consists of verifying, using a current deep learning method, the task of discriminating between speakers who are taken into account and those who are not, while highlighting the characteristics that are important for classification. Here, the EEG signals are converted into images. We then extract attention-related information from these images. Next, the deep learning model is fine-tuned using this data. Finally, auditory attention is elucidated through explainability. The ultimate goal of this study is to enable the detection of hearing-impaired individuals using a non-invasive method and to develop artificial systems that can assist them. RELATED WORK Several studies have been conducted in the field of AAD. Tanveer et al., in their study, analysed 64 EEG channels extracted from subjects, with fallowing scenarios: target speech/distracting music, target speech/distracting speech, target music/distracting speech, and target music/distracting music. The performance of the models on time windows of 3, 5, 10, and 20 seconds was 54.1, 57.0, 66.5, and 80.1 (%) respectively [ 10 ]. The work of Bollens et al. describes the auditory EEG challenge, organised as part of the major signal processing challenges 2024. The challenge provides EEG recordings from 105 subjects who listened to continuous speech, in the form of audiobooks or podcasts, while their brain activity was recorded [ 21 ]. Several studies have been conducted to address this challenge. The EEG signals of 34 subjects listening to violin sounds were collected using a Go/No-Go paradigm in the study by Meng et al. The modified, lightweight EEGNet model was proposed for EEG-based pitch classification. The average classification accuracy was 77% for intra-subject modelling [ 22 ]. Ciccarelli et al. used both wet and dry EEG systems on 11 people per class. The results indicate that decoding accuracy reached 81% with wet EEG and 87% with dry EEG [ 23 ]. In our study, we will use the Technical University of Denmark (DTU) dataset. Using the entire dataset and applying convolutional neural networks (CNN), results of approximately 80% were obtained. With the help of a generative network, the data was multiplied and reached approximately 90%. As part of their study, the researchers who collected the data achieved 78% accuracy with 10-second windows [ 24 ]. Sridhar et al., using the same data, achieved an accuracy of 81% with 10-second windows [ 1 ]. Geravanchizadeh et al. achieved an accuracy of 80.12% [ 4 ]. In the review by Sun et al., in the comparison table, on 1-second windows, 71.9% was obtained on the DTU dataset [ 12 ]. In the work of Yuan et al., using the same dataset with one-second windows, the accuracy was 78.3% [ 8 ]. Kausar et al., using the AD-GAN network and nearly 85,000 generated data points in addition to 12,800, achieved an accuracy of 90.1% for 10-second sequences on the same dataset [ 7 ]. EskandariNasab et al. obtained 89.4% accuracy in 10-second windows when comparing healthy and sick individuals [ 25 ]. Huang et al. achieved 89.4% accuracy on 3-second windows with the DBGAN generative network [ 26 ]. In this study, using a novel approach, we will apply transfer learning using two-dimensional data and a CNN, utilising only a portion of the data. The ability to understand and evaluate how AI algorithms reach their conclusions is achieved through explainability. Explainable AI approaches aim to make deep learning models, which are ‘black boxes’, more transparent and reliable. Techniques can be used to determine which parts of the brain or which signal properties have the most influence on hearing and AAD. In the field of AAD, Fiorini's work involves using interpretable AI to examine the overall importance of features using Shapley additive explanations (SHAP) [ 27 ]. The work of Almadhor et al. uses deep learning models to detect schizophrenia from EEG data. Important EEG features that influence the model's decisions are highlighted by the SHAP and LIME explainability tools [ 28 ]. Mahjoory et al. sought to use CNN as an interpretable model to discover task-specific interactions between brain regions, rather than simply using it as a black-box decoder. To this end, the CNN model was applied for 10 cortical regions from five-second inputs. Using only these features for decoding, the model achieved a median accuracy of 77.56% for intra-participant classification and 65.14% for inter-participant classification (alpha, delta theta, beta) [ 5 ]. Magosso et al. collected EEG signals during a hidden attention task, which required directing attention to the left or right visual field. They focused their work on explainable artificial intelligence, combining a convolutional neural network and an explanation technique to reveal the most relevant regions of the entire cortex [ 16 ]. In our work, we will apply the GradCAM tool to explain the model used. For the first time, XAI work will be carried out on DUT data. MATERIAL AND METHODS The auditory cortex's response will be extracted to analyse the neural coding of the target sound being listened to when two sounds are heard simultaneously. Deep learning techniques will provide new contributions to the results of existing studies. Neural data will be presented in the form of a Scalogram to the input of the pre-trained Resnet101 architecture. A graphical summary of the work is presented in Fig. 1 . EEG signals from 14 channels obtained from 18 individuals listening to sounds involving the reading of different texts coming from the right and left for 50 seconds by both women and men were converted into scalogram images over 30 trials. These images, obtained from 10-second windows, were processed using the resnet101 architecture, classified, and the results were compared. As a result of the classification, significant features were evaluated using the XAI approach (Fig. 3 a). The Dataset This study uses the standardised, publicly available dataset from the Technical University of Denmark (DTU), created as part of the COCOHA (Cognitive Control of a Hearing Aid) project, based on the task of discriminating between attentive and inattentive speakers for hearing aids [ 29 ]. It includes 64-channel EEG recordings from 18 participants with normal hearing who listened attentively to one of two competing speech streams (A and B) from Danish audiobooks in simulated acoustic environments. The study protocol was approved by the Scientific Ethics Committee of the Capital Region of Denmark. All subjects gave written informed consent in accordance with the Declaration of Helsinki. Subjects were instructed to pay attention to one of two professional speakers (male and female), with the direction to which they were to pay attention alternating between left and right. In each experimental trial, a two-speaker scenario was created. The order of acoustic conditions, gender and target flow position, as well as the order of story presentation, were randomised from trial to trial. EEG acquisition was configured according to the international 10–20 system. Each subject performed 60 trials (30 trials were used in this study), each lasting 50 seconds (Fig. 2 ). The following EEG analyses were performed on the envelope representations of the vocal stimuli. The raw vocal waveforms were decomposed by a gammatone filter bank into 128 sub-bands with centre frequencies between 100 Hz and 8,000 Hz. The individually extracted narrowband Hilbert envelopes were transformed according to a power law to simulate the perception of sound volume in the auditory system, then averaged to obtain the broadband envelope [ 24 ]. Electroencephalogram There are 66 EEG data channels. Electrodes placed on the right and left auditory areas. The auditory cortex area, for the left side, FT7, T7, FP7, C5, CP5, C3, CP3, for the right side, FT8, T8, FP8, C6, CP6, C4, CP4 electrodes were used. Scalogram 2D scalogram images are created by applying Continuous Wavelet Transform (CWT) to the raw EEG data. Time frequency features are obtained from the scalogram images [ 31 ]. A scalogram is the absolute value of the coefficients of the continuous wavelet transform (CWT) of a signal represented as a function of time and frequency [ 32 ]. This study adopts the 2D scalogram approach, which uses CWT with pre-processed signals. CWT uses inner products to calculate the similarity between a waveform and an examination function such as the Fourier transform [ 33 ]. The CWT is obtained by windowing the signal with a scaled and time-shifted wavelet. The CWT is obtained using the Morse analytical wavelet. The Fourier transform of the generalised Morse waveform is as follows (1) : $$\:{\psi\:}_{P,\gamma\:}\left(\omega\:\right)\:=U\left(\omega\:\right){a}_{P,\gamma\:}{\omega\:}^{\frac{{P}^{2}}{\gamma\:}}{e}^{{-\omega\:}^{\gamma\:}}$$ 1 where U(ω) is the unit step, a P , γ are normalisation constants, P 2 is the time-bandwidth product, and γ characterises the symmetry of the Morse waveform [ 34 – 36 ]. Classification model: ResNet-101 convolutional neural network Deep neural networks trained on large image collections can be used for small, labelled datasets. The pre-trained network ResNet101 will be used in this study. Deep learning-based image classification models learn to recognise features in an image by learning to recognise features from Scalogram representations. ResNet-101 is a convolutional neural network with a depth of 101 layers. The network was trained on over a million images in the ImageNet database and can classify images into 1,000 object categories. The network has an input image size of 224 by 224 [ 35 , 37 , 38 ]. We can say that the pre-trained ResNet-101 network enables us to achieve excellent transfer learning results. Transfer learning was performed by modifying the final layers of ResNet101 using weight transfer. In other words, fine-tuning was performed using new data. To obtain the probability distribution, the softmax output layer has been removed and a fully connected layer with two outputs for the classes and a new softmax output layer have been added. The criteria used to evaluate the performance of the classification are important for comparison with other results obtained. Accuracy, recall/sensitivity and precision and the AUC (Area Under the Curve) value were considered to evaluate the results. Explainable Artificial Intelligence Explainability is achieved by estimating the marginal contribution of each feature to a prediction using GradCAM. Gradient-weighted class activation mapping (Grad-CAM) is an explainability technique that can be used to help understand the predictions made by a deep neural network. Grad-CAM determines the importance of each neuron in a network prediction by considering the gradients of the target flowing through the deep network [ 39 ]. The Grad-CAM interpretability technique uses the gradients of the classification score relative to the final convolutional feature map. The parts of an observation with a high value for the Grad-CAM map are those that have the greatest impact on the network's score for that class. The gradients are grouped in spatial and temporal dimensions to determine the importance weights of the neurons. These weights are then used to linearly combine the activation maps and determine which features are most important for prediction [ 35 ]. In the test dataset, all correctly classified scalograms can each generate a GradCAM map. The averaged image is obtained by calculating the numerical average pixel by pixel. The study was conducted on an Apple MacBook M2 Pro with 16 GB of memory, a total of 12 cores, and a 19-core graphics processor. Matlab 2025b was used. RESULTS We will carry out our work using part of the data, without using all of the tests, with transfer learning, and with a less time-consuming process. In this study, our objective will be to perform AAD through transfer learning using a limited number of 10-second windows. The data comes from the sounds of women or men coming from the right or left, recorded under any acoustic conditions (with or without echo). Therefore, this study did not take into account different configurations, different trials, whether the sound was heard from the right or left, whether the speaker was female or male, or whether the EEG signals were taken from the right or left auditory region. Table 1 Tets classification results for 4 participants and 18 participants. Transfer learning with ResNet-101. Accuracy, sensitivity, precision, AUC values. Accuracy Sensitivity/Recall Precision AUC 4 Participants Speech A Speech B 87.2 86.0 88.5 88.2 86.3 94.6 18 Participants Speech A Speech B 79.7 79.5 80.0 79.9 79.6 87.7 The results of a test. The values are given as percentages(%). The signals from the 14 channels located on the auditory areas during 30 trials are transformed into scalogram images over 10 seconds. 90% of the data was used for training and 10% for testing. During training, care was taken to avoid overfitting or underfitting. In Table 1 , the results for 5 participants and 18 participants are presented separately. The dataset is balanced with 15,330 images per class for 18 participants and approximately 4,000 images for 4 participants. Training for 4 participants lasted approximately 2 hours and 8 hours for 18 participants. High performance was achieved with 87.2% accuracy and 94.6% AUC value for 4 participants and 79.7% accuracy and 87.7% AUC value for 18 participants. In this study, with the use of deep learning techniques, the encoding of the sound envelope considered in a noisy environment in the brain was solved with high performance. The results of the study on explainability are shown in Fig. 3 . Sample GradCAM map images are shown in Fig. 3 b. The averaged image of the GradCAM maps is shown in Fig. 3 c. DISCUSSION Deep learning models have yielded promising results in various tasks related to brain signal processing. This study has provided new insights using an explainability approach while attempting to decode the auditory attention code using EEG data in competing speech scenarios. In this study, the auditory attention code was decoded in short-term conversations using transfer learning with 10-second windows. This study was performed using two-dimensional convolutional neural networks with approximately half the data and lower computational requirements. Different configurations, different trials, whether the sound was heard from the right or left, whether the speaker was female or male, and whether the EEG signals were taken from the right or left auditory region were not considered. If they had been, the results would have been much higher. The reconstruction accuracy for attended and ignored speech, was calculated for ResNet-101. Table 1 shows that for the model trained using data classified from the EEG signals of four individuals, we achieved an accuracy of 87.2% and an AUC of 94.6%. Eger, 0.5 < AUC < 1 indicates that the model has some ability to distinguish between categories. The closer the value is to 1, the better the model's performance. This remains relatively high compared to the literature studied. In the context of classifying data from 18 individuals, we achieved an accuracy of 79.7% and an AUC of 87.7%. The more people with different EEG waves, the lower the model's performance results. In this case, inter-subject modelling far exceeds the random accuracy of 50%. Cortical activity phase-locks to the envelope of an attended speech source, and the encoding of the considered sound envelope in the brain has been solved with high performance. In Table 1 , we can also see that the sensitivity and precision results for both speech A (79.5% and 79.9% respectively) and speech B (80% and 79.6% respectively) are very similar. This shows us that the two classes are equally distinguishable. Using 10-second windows allows us to evaluate speech and/or language. Shorter durations are used more for evaluating voice and/or sound. In our case, our work allows us to evaluate not only hearing but also the integration, identification, and comprehension of what is heard. Furthermore, through testing, we found that using data from only 14 electrodes near the auditory lobes maintains better performance compared to using 64 electrodes. Existing AAD algorithms often exploit the powerful modelling capabilities of deep learning, but few of them take XAI into account. The results of the study on explainability are shown in Fig. 3 . The averaged image of the GradCAM maps is shown in Fig. 3 c. The model exhibits strong interpretability, revealing that the left and right auditory lobes are more active during speaker scenarios. We can see in Fig. 3 c that the identification of the two speeches listened to is based on the increase or decrease in alpha and beta waves, and more specifically in the Sensorimotor Rhythm (SMR) [ 40 ]. This could provide to neuroscientists a new and useful data-driven analysis tool, aimed not only at decoding but also at analyzing functional connectivity estimates. EEG rhythms allow us to observe different functional states of the brain. Generally, high-frequency, low-amplitude rhythms are associated with alertness and wakefulness, or with the dreaming phases of sleep. When the cortex is most engaged in analysing information from sensory input or internal processes (wakefulness), cortical neuron activity is relatively high [ 41 ]. Alpha waves (8 to 12 Hz) are associated with states of calmness during wakefulness [ 41 ]. Attentional processing or cognitive tasks attenuate the alpha waves. alpha waves may be used to predict mistakes, with open eyes can be a predictor of visual information processing in working memory and perceptual visual learning [ 42 ]. Beta waves (15 to 30 Hz) correspond to periods of normal or intense activity during the day [ 40 ]. Beta rhythm brain activity is characteristic of normal wakefulness when the subject has their eyes open and is performing a perceptual (vision, hearing, touch) or mental (arithmetic, cognitive, complex) task. Low-amplitude beta waves with multiple rapid frequency changes are often associated with an active, busy, or even anxious state of mind, with high concentration. Sudden peaks in beta activity are associated with enhanced sensory feedback during control. [ 42 ]. The sensorimotor rhythm (SMR) (12 to 15 Hz) corresponds to the state of concentration just before performing an action [ 40 ]. From a phenomenological point of view, its amplitude tends to increase when the sensorimotor regions of the brain are at rest, such as in a state of immobility. The functional significance of SMR is linked, among other things, to a state of relaxed alertness without movement and reduced sensory processing during periods of motor rest [ 42 ]. This information is consistent with what we obtain on XAI maps in terms of frequency and activation. Certain areas of the GradCAM map are activated or not. This shows that alpha, beta and SMR waves are more or less intense depending on the type of speech. Although preliminary, this study shows the potential of combining an XAI approach with more traditional methods to increase the robustness of EEG source analysis in hearing and AAD. The same combined approach can be easily transposed to study other cognitive tasks. The results of this study have important implications for understanding and assessing auditory attention, which is essential for applications such as brain-computer interface (BCI) systems and the development of hearing aids. Sound can be easily detected even in very noisy environments, so our results demonstrate the existence of the neural foundations necessary for human communication in realistic situations. This is also demonstrated by the brain's electrical activity. Deep learning models have achieved promising results in various tasks related to brain signal processing. Neuroscience and Brain-Computer Interface Research Identifying Hearing-Impaired Individuals Using a Non-Invasive Method Developing Artificial Systems CONCLUSIONS AND LIMITATIONS In this study, we uses EEG signals for AAD in order to predict the sound being listened to based on neural activity in the auditory cortex with deep learning models. The EEG signals are converted into graphs and then presented at the input of the ResNet-101 architecture. We obtained an AUC value of 87.7% with data from 18 participants. We then used XAI the GradCAM technique to understand which characteristics and types of EEG components are used for AAD. Then, auditory attention is elucidated through explainability. This study shows that alpha, beta and sensorimotor rhythms are more or less intense depending on the type of speech. This Study can be used as a reference for artificial intelligence algorithms that process sound in real environments. In the future, the same study could be conducted to investigate the decoding of the neural code of sounds in noise within the subcortical region of the auditory system, for example in the inferior colliculus, independently of the presence of the auditory cortex. The effect of factors such as the gender of the person, the gender of the listener on AAD performance and subjective experience has not been studied. The AAD algorithm must demonstrate accurate performance in a variety of real-world listening situations to be implemented in future hearing aids. Abbreviations AAD : Auditory attention decoding EEG : ElectroEncephaloGram GradCAM : Gradient-weighted class activation mapping XAI : eXplainable Artificial Intelligence DTU : Technical University of Denmark CNN : convolutional neural networks DBGAN : Dual Branch Generative Adversarial Network SHAP : Shapley additive explanations LIME: Local interpretable model-agnostic explanations COCOHA : Cognitive Control of a Hearing Aid CWT : Continuous Wavelet Transform AUC : Area Under the Curve SMR : SensoriMotor Rhythm BCI : Brain-Computer Interface Declarations Ethics approval and consent to participate This research did not require ethical approval as the data is publicly available and ethical approval has been obtained by the data publishers. This statement is not applicable to the present study. Consent to participate statement is not applicable to the present study Consent for publication Not applicable Availability of data and materials The data that support the endings of this study are available publicly from the the Technical University of Denmark, created as part of the COCOHA. The dataset is publicly accessible via https://zenodo.org/records/1199011 Competing Interests The author has no relevant financial or non-financial interests to disclose. Funding The author declare that no funds, grants, or other support were received during the preparation of this manuscript. Author Contributions The author, Fatma ÖZCAN, contributed to the design and implementation of the study, performed the data analysis, and wrote the manuscript. Acknowledgements Not applicable References Sridhar, G. (2025). Improving auditory attention decoding in noisy environments for listeners with hearing impairment through contrastive learning . Journal Of Neural Engineering , 22(3). Ceravolo, L. (2024). Functional and causal neural mechanisms of human voice perception in noisy situations BioRxiv. Mai, A., Hillyard, S. A., & Strauss, D. J. (2025). Linear modeling of brain activity during selective attention to continuous speech: the critical role of the N1 effect in event-related potentials to acoustic edges. Cognitive Neurodynamics , 19 (1), 110. Geravanchizadeh, M., Shaygan Asl, A., & Danishvar, S. (2024). Selective Auditory Attention Detection Using Combined Transformer and Convolutional Graph Neural Networks . Bioengineering (Basel) , 11(12). Mahjoory, K., Bahmer, A., & Henry, M. J. (2024). Convolutional neural networks can identify brain interactions involved in decoding spatial auditory attention. Plos Computational Biology , 20 (8), e1012376. Alishbayli, A. (2024). Processing of Statistically Defined Sounds in the Auditory Cortex . Kausar, T., et al. (2024). Auditory-GAN: deep learning framework for improved auditory spatial attention detection. PeerJ Comput Sci , 10 , e2394. Yuan, L., et al. (2025). Frequency-Based Alignment of EEG and Audio Signals Using Contrastive Learning and SincNet for Auditory Attention Detection . arxiv. Issa, M. F., et al. (2024). On the speech envelope in the cortical tracking of speech. Neuroimage , 297 , 120675. Tanveer, M. A. (2024). Envelope Based Deep Source Separation and EEG Auditory Attention Decoding for Speech and Music , in EUSIPCO 2024 . Deoisres, S., et al. (2025). Comparing approaches for predicting behavioural speech-in-noise performance using cortical responses to unattended stimuli. Hearing Research , 457 , 109197. Sun, Q., et al. (2025). Attention Detection Using EEG Signals and Machine Learning: A Review (p. 22). Machine Intelligence Research. Thornton, M., Mandic, D., & Reichenbach, T. (2024). Comparison of linear and nonlinear methods for decoding selective attention to speech from ear-EEG recordings . arXiv . Mesgarani, N., & Chang, E. F. (2012). Selective cortical representation of attended speaker in multi-talker speech perception. Nature , 485 (7397), 233–236. Thornton, M., Mandic, D., & Reichenbach, T. (2022). Robust decoding of the speech envelope from EEG recordings through deep neural networks . Journal Of Neural Engineering , 19(4). Magosso, E., Bruno, P., & Borra, D. (2025). Combining EEG Oscillation Analysis and Explainable Artificial Intelligence for Characterizing Visuospatial Attention , in Springer Nature Switzerland AG 2025 . Di Martino, F., & Delmastro, F. (2023). Explainable AI for clinical and remote health applications: a survey on tabular and time series data. Artificial Intelligence Review , 56 (6), 5261–5315. Dissanayake, T., et al. (2021). A Robust Interpretable Deep Learning Classifier for Heart Anomaly Detection Without Segmentation (p. 25). IEEE JOURNAL OF BIOMEDICAL AND HEALTH INFORMATICS. Tjoa, E., & Guan, C. (2021). A Survey on Explainable Artificial Intelligence (XAI): Toward Medical XAI (p. 32). IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS. Vilone, G., & Longo, L. (2021). Notions of explainability and evaluation approaches for explainable artificial intelligence . Information Fusion , 76. Bollens, L. (2025). Auditory EEG Decoding Challenge for ICASSP 2024 . IEEE Signal Processing , 6. Meng, Q., et al. (2025). EEG-based cross-subject passive music pitch perception using deep learning models. Cognitive Neurodynamics , 19 (1), 6. Ciccarelli, G., et al. (2019). Comparison of Two-Talker Attention Decoding from EEG with Nonlinear Neural Networks and Linear Methods. Scientific Reports , 9 (1), 11538. Wong, D. D. E., et al. (2018). A Comparison of Regularization Methods in Forward and Backward Models for Auditory Attention Decoding. Front Neurosci , 12 , 531. EskandariNasab, M., et al. (2024). A GRU-CNN model for auditory attention detection using microstate and recurrence quantification analysis. Scientific Reports , 14 (1), 8861. Huang, S., Wang, Y., & Luo, H. (2025). A dual-branch generative adversarial network with self-supervised enhancement for robust auditory attention decoding . Engineering Applications of Artificial Intelligence. Fiorini, L., Lucia, S., & Pietrogiacomi, F. (2024). Insights_Into_Pre-Stimulus_Activity_for_EEG-Based_Machine_Learning_Sensory_Classification , in IEEE . Almadhor, A., et al. (2025). An interpretable XAI deep EEG model for schizophrenia diagnosis using feature selection and attention mechanisms. Frontiers In Oncology , 15 , 1630291. Fuglsang, S. A., Wong, D. E., & Hjortkjaer, J. (2018). EEG and audio dataset for auditory attention decoding . Zenodo. Jianbiao, M., et al. (2023). EEG signal classification of tinnitus based on SVM and sample entropy. Comput Methods Biomech Biomed Engin , 26 (5), 580–594. Aslan, Z., & Akin, M. (2022). A deep learning approach in automated detection of schizophrenia using scalogram images of EEG signals. Phys Eng Sci Med , 45 (1), 83–96. Byeon, Y. H., Pan, S. B., & Kwak, K. C. (2019). Intelligent Deep Models Based on Scalograms of Electrocardiogram Signals for Biometrics . Sensors (Basel) , 19(4). Loey, M., & Mirjalili, S. (2021). COVID-19 cough sound symptoms classification from scalogram image representation using deep learning models. Computers In Biology And Medicine , 139 , 105020. Lilly, J. M., & Olhede, S. C. (2012). Generalized Morse wavelets as a superfamily of analytic wavelets. IEEE Transactions on Signal Processing , 60 (11), 6036–6041. Mathworks (2025). Matlab R2025a . Olhede, S. C., & Walden, A. T. (2002). Generalized morse wavelets. IEEE Transactions on Signal Processing , 50 (11), 2661–2670. He, K. (2016). Deep Residual Networks . He, K., et al. (2015). Deep Residual Learning for Image Recognition . arxiv. Ramprasaath, R., et al. (2017). Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization . arxiv. Pouessel, C. (2019). Tout savoir exactement de manière simple sur les ondes cérébrales . ; Available from: https://www.cabinet-neurofeedback.fr/2019/06/23/les-ondes-cerebrales/ Bear, M. F., Connors, B. W., & Paradiso, M. A. (2015). Neuroscience – Exploring the Brain. Éditions Pradel. Wikipedia (2025). Rythme sensorimoteur . ; Available from: https://fr.wikipedia.org/wiki/Rythme_sensorimoteur Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8503109","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":573570334,"identity":"623b44d3-d8ee-4970-a7a8-796c7604ccbd","order_by":0,"name":"Fatma ÖZCAN","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA/UlEQVRIiWNgGAWjYBAC9gYkzsEPFUCSmbkBu1oo4DnADGczHpY4A9LCSLwW5gO8bWCtBLSwnz+64QeDjT1/+9kDByTn1UbztwO1/KjYhlsLTzLbzR6GNGaJM3kJBwq3Hc+dcZixgbHnzG2cWuwZktlu8DAcZjNgyDE4ILntWG4DUAszYxtuLTz8j9lu/mE4zGPA/8bgAO+cY7nzCWqRSGa7DbRFwkACaAtvQ03uBsJaHpvdljFIM5C48cbgsMSxA7kbgVoO4vMLD3/is5tvKoAh1p9j/PFDTV3uvPOHDz74UYFbCwQYwFmHweQBAupRQB0pikfBKBgFo2CEAAAVs1nbLsGEBwAAAABJRU5ErkJggg==","orcid":"","institution":"Kahramanmaras Sutcu Imam University","correspondingAuthor":true,"prefix":"","firstName":"Fatma","middleName":"","lastName":"ÖZCAN","suffix":""}],"badges":[],"createdAt":"2026-01-02 20:08:16","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8503109/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8503109/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":100366977,"identity":"0b5919a8-5ecf-4c56-9ee2-8a29d97cb650","added_by":"auto","created_at":"2026-01-16 07:56:42","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":5680323,"visible":true,"origin":"","legend":"","description":"","filename":"AADandEEGandXAI2.docx","url":"https://assets-eu.researchsquare.com/files/rs-8503109/v1/6bbcb730c819f7aa30367517.docx"},{"id":100136176,"identity":"008f480c-5a7b-48f5-bf4a-3b0e576aedf7","added_by":"auto","created_at":"2026-01-13 10:44:46","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":3051,"visible":true,"origin":"","legend":"","description":"","filename":"29881e3636cb446980f31197adfc27c8.json","url":"https://assets-eu.researchsquare.com/files/rs-8503109/v1/f32d6f2f4a822fabc87c5782.json"},{"id":100136170,"identity":"0a738694-bd89-4556-8b43-4b0be0b850c2","added_by":"auto","created_at":"2026-01-13 10:44:45","extension":"xml","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":84197,"visible":true,"origin":"","legend":"","description":"","filename":"29881e3636cb446980f31197adfc27c81enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-8503109/v1/4cfebcbf678f95ee2711fe19.xml"},{"id":100136169,"identity":"ee60a5e4-dafa-4afe-bae4-a14cf3fe25f0","added_by":"auto","created_at":"2026-01-13 10:44:45","extension":"png","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":51223,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8503109/v1/a33f8b9f7b37da26e77d8860.png"},{"id":100136174,"identity":"c8c219ea-f9f0-4d2b-a387-8cde84f7f5d4","added_by":"auto","created_at":"2026-01-13 10:44:46","extension":"png","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":47820,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8503109/v1/1c7278a64b0f3d220ec800cc.png"},{"id":100368233,"identity":"47e0ef57-8225-4af0-a0dc-8411baac2948","added_by":"auto","created_at":"2026-01-16 07:57:44","extension":"png","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":182818,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8503109/v1/99e8d2ade5e66bac8ae4ca24.png"},{"id":100136179,"identity":"960817b9-b175-45f2-bc3f-7dc816695989","added_by":"auto","created_at":"2026-01-13 10:44:46","extension":"xml","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":83115,"visible":true,"origin":"","legend":"","description":"","filename":"29881e3636cb446980f31197adfc27c81structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-8503109/v1/9439b50eb87f0e69b97f5002.xml"},{"id":100136177,"identity":"358cfdee-5429-4fe4-8bdd-16231c0100e9","added_by":"auto","created_at":"2026-01-13 10:44:46","extension":"html","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":92880,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-8503109/v1/f01d0d9d6378ef70cdf2b9e5.html"},{"id":100136173,"identity":"2211eaf0-61c5-404b-9eed-fb16303c56df","added_by":"auto","created_at":"2026-01-13 10:44:46","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":341669,"visible":true,"origin":"","legend":"\u003cp\u003eFunctional diagram\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8503109/v1/6468e94b2ed3d58c6e208b70.png"},{"id":100136172,"identity":"b95d42e1-b5bf-4fdb-a31b-08d0e1a60ca8","added_by":"auto","created_at":"2026-01-13 10:44:45","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":148398,"visible":true,"origin":"","legend":"\u003cp\u003eSummary of the experimental design for data collection [3]. Participants took part in a selective auditory attention task with two speakers. Continuous vocal stimuli were presented using in-ear headphones placed at a 60-degree angle to the midline, and the attention task was performed randomly for 30 trials (each lasting 50 seconds).\u003c/p\u003e","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8503109/v1/da9ab55ed96efbaf4c6ced66.jpeg"},{"id":100367286,"identity":"a331c5e8-4ddf-45f2-94bd-85adb1dc6d79","added_by":"auto","created_at":"2026-01-16 07:56:56","extension":"jpeg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":889524,"visible":true,"origin":"","legend":"\u003cp\u003e(\u003cstrong\u003ea\u003c/strong\u003e) Examples of scalograms for the attended speech, A or B. The graphical representation has the time (10 secondes) as the abcissa and the frequency (in Hz) as the ordinate. On the right-hand side of the image (from blue to yellow for the lowest to highest intensity) the intensity is shown. (\u003cstrong\u003eb\u003c/strong\u003e) XAI GradCAM map from correctly classified image for two classes. The red areas mark the zones used by the model to make the related classification. (\u003cstrong\u003ec\u003c/strong\u003e) Average images of XAI maps in (b) for two classes.\u003c/p\u003e","description":"","filename":"floatimage3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8503109/v1/2e7dfb8ab00c7c32f15fa916.jpeg"},{"id":103392859,"identity":"2c393ac4-0d9e-4b3d-8464-b8916d545604","added_by":"auto","created_at":"2026-02-25 08:12:59","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1934730,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8503109/v1/77a44fb8-4794-4028-ba4d-6419964e5a9a.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"\u003cp\u003eExplainable Eeg for Auditory Attention Decoding\u003c/p\u003e","fulltext":[{"header":"INTRODUCTION","content":"\u003cp\u003eThe human brain is capable of exceptional speed and precision in processing sensory data. In neuroscience, understanding the complexity of these processes is one of the main objectives [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. Human communication involves the simultaneous processing of sounds [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. In order to perform in a variety of complex listening conditions, the auditory system distinguishes sound sources by reducing the impact of ambient noise. Auditory attention decoding (AAD) uses brain signals to identify conversations being listened to in an environment where several people are talking simultaneously [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. The ability to focus one's attention on a single speaker in an environment is an essential skill of the human auditory system [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eElectroencephalography provides valuable information about the brain's information processing in a practical and inexpensive way. The neural correlates of selective attention can be decoded from the EEG [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. EEG, which can reproduce the cocktail party effect, shows promise for the study of AAD. At a cocktail party, there are multiple sound sources, inducing energetic masking in the spectro-temporal space, and the content conveyed can form informational masking, which creates an additional cognitive load for the system. The auditory system must therefore separate the sources using contextual information and spatial cues at its disposal [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIn an environment where several speakers are present, the speaker who attracts attention is selected as the one who shows the strongest correlation between neural data and the speech envelope [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. Cortical tracking of the acoustic envelope is a mechanism whereby the electrical activity of the brain varies in response to the acoustic envelope of the stimulus. The synchronisation between the vocal envelope and neural activity in the auditory regions plays a key role in speech processing [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]\u003c/p\u003e \u003cp\u003eUnderstanding speech in a noisy environment is a major challenge for people with hearing loss. Speech stimuli are therefore more valuable than clicks and tones for assessing hearing [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. Understanding the process of solving complex perception problems will enable the development of artificial systems that can help people with sensory impairments [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. The decoding of AAD via a brain-computer interface has seen considerable growth since its introduction by Mesgarani and Chang using electrocorticographic recordings [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eDeep learning models have shown promising results in various tasks related to brain signal processing for AAD [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. Recent advances in eXplainable Artificial Intelligence (XAI) offer new possibilities for AAD by enabling the learning of informative features [\u003cspan additionalcitationids=\"CR17 CR18 CR19\" citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]\u003c/p\u003e \u003cp\u003eAlthough progress has been made in detecting auditory attention based on EEG, it is necessary to conduct trials with new methods in this field. It is still unclear whether EEG responses to speech offer an advantage in predicting speech intelligibility [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e] and EEG signals generally suffer from a low signal-to-noise ratio, making it difficult to decode complex cognitive functions [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. Furthermore, the neural mechanisms that enable selective attention dependent on sound processing in complex auditory scenes are not well understood [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eOur objective is to predict, using current and explainable deep learning methods, the sound heard from neural activity in the auditory cortex. In our study, EEG signals recorded when two different voices are heard, on 4 and 18 subjects, will be used for AAD. This work consists of verifying, using a current deep learning method, the task of discriminating between speakers who are taken into account and those who are not, while highlighting the characteristics that are important for classification. Here, the EEG signals are converted into images. We then extract attention-related information from these images. Next, the deep learning model is fine-tuned using this data. Finally, auditory attention is elucidated through explainability. The ultimate goal of this study is to enable the detection of hearing-impaired individuals using a non-invasive method and to develop artificial systems that can assist them.\u003c/p\u003e\n\u003ch3\u003eRELATED WORK\u003c/h3\u003e\n\u003cp\u003eSeveral studies have been conducted in the field of AAD. Tanveer et al., in their study, analysed 64 EEG channels extracted from subjects, with fallowing scenarios: target speech/distracting music, target speech/distracting speech, target music/distracting speech, and target music/distracting music. The performance of the models on time windows of 3, 5, 10, and 20 seconds was 54.1, 57.0, 66.5, and 80.1 (%) respectively [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. The work of Bollens et al. describes the auditory EEG challenge, organised as part of the major signal processing challenges 2024. The challenge provides EEG recordings from 105 subjects who listened to continuous speech, in the form of audiobooks or podcasts, while their brain activity was recorded [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. Several studies have been conducted to address this challenge. The EEG signals of 34 subjects listening to violin sounds were collected using a Go/No-Go paradigm in the study by Meng et al. The modified, lightweight EEGNet model was proposed for EEG-based pitch classification. The average classification accuracy was 77% for intra-subject modelling [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]. Ciccarelli et al. used both wet and dry EEG systems on 11 people per class. The results indicate that decoding accuracy reached 81% with wet EEG and 87% with dry EEG [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIn our study, we will use the Technical University of Denmark (DTU) dataset. Using the entire dataset and applying convolutional neural networks (CNN), results of approximately 80% were obtained. With the help of a generative network, the data was multiplied and reached approximately 90%. As part of their study, the researchers who collected the data achieved 78% accuracy with 10-second windows [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]. Sridhar et al., using the same data, achieved an accuracy of 81% with 10-second windows [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. Geravanchizadeh et al. achieved an accuracy of 80.12% [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. In the review by Sun et al., in the comparison table, on 1-second windows, 71.9% was obtained on the DTU dataset [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. In the work of Yuan et al., using the same dataset with one-second windows, the accuracy was 78.3% [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eKausar et al., using the AD-GAN network and nearly 85,000 generated data points in addition to 12,800, achieved an accuracy of 90.1% for 10-second sequences on the same dataset [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. EskandariNasab et al. obtained 89.4% accuracy in 10-second windows when comparing healthy and sick individuals [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. Huang et al. achieved 89.4% accuracy on 3-second windows with the DBGAN generative network [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIn this study, using a novel approach, we will apply transfer learning using two-dimensional data and a CNN, utilising only a portion of the data.\u003c/p\u003e \u003cp\u003eThe ability to understand and evaluate how AI algorithms reach their conclusions is achieved through explainability. Explainable AI approaches aim to make deep learning models, which are \u0026lsquo;black boxes\u0026rsquo;, more transparent and reliable. Techniques can be used to determine which parts of the brain or which signal properties have the most influence on hearing and AAD. In the field of AAD, Fiorini's work involves using interpretable AI to examine the overall importance of features using Shapley additive explanations (SHAP) [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. The work of Almadhor et al. uses deep learning models to detect schizophrenia from EEG data. Important EEG features that influence the model's decisions are highlighted by the SHAP and LIME explainability tools [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]. Mahjoory et al. sought to use CNN as an interpretable model to discover task-specific interactions between brain regions, rather than simply using it as a black-box decoder. To this end, the CNN model was applied for 10 cortical regions from five-second inputs. Using only these features for decoding, the model achieved a median accuracy of 77.56% for intra-participant classification and 65.14% for inter-participant classification (alpha, delta theta, beta) [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. Magosso et al. collected EEG signals during a hidden attention task, which required directing attention to the left or right visual field. They focused their work on explainable artificial intelligence, combining a convolutional neural network and an explanation technique to reveal the most relevant regions of the entire cortex [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIn our work, we will apply the GradCAM tool to explain the model used. For the first time, XAI work will be carried out on DUT data.\u003c/p\u003e"},{"header":"MATERIAL AND METHODS","content":"\u003cp\u003eThe auditory cortex's response will be extracted to analyse the neural coding of the target sound being listened to when two sounds are heard simultaneously. Deep learning techniques will provide new contributions to the results of existing studies. Neural data will be presented in the form of a Scalogram to the input of the pre-trained Resnet101 architecture. A graphical summary of the work is presented in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eEEG signals from 14 channels obtained from 18 individuals listening to sounds involving the reading of different texts coming from the right and left for 50 seconds by both women and men were converted into scalogram images over 30 trials. These images, obtained from 10-second windows, were processed using the resnet101 architecture, classified, and the results were compared. As a result of the classification, significant features were evaluated using the XAI approach (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ea).\u003c/p\u003e\n\u003ch3\u003eThe Dataset\u003c/h3\u003e\n\u003cp\u003eThis study uses the standardised, publicly available dataset from the Technical University of Denmark (DTU), created as part of the COCOHA (Cognitive Control of a Hearing Aid) project, based on the task of discriminating between attentive and inattentive speakers for hearing aids [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]. It includes 64-channel EEG recordings from 18 participants with normal hearing who listened attentively to one of two competing speech streams (A and B) from Danish audiobooks in simulated acoustic environments. The study protocol was approved by the Scientific Ethics Committee of the Capital Region of Denmark. All subjects gave written informed consent in accordance with the Declaration of Helsinki.\u003c/p\u003e \u003cp\u003eSubjects were instructed to pay attention to one of two professional speakers (male and female), with the direction to which they were to pay attention alternating between left and right. In each experimental trial, a two-speaker scenario was created. The order of acoustic conditions, gender and target flow position, as well as the order of story presentation, were randomised from trial to trial. EEG acquisition was configured according to the international 10\u0026ndash;20 system. Each subject performed 60 trials (30 trials were used in this study), each lasting 50 seconds (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). The following EEG analyses were performed on the envelope representations of the vocal stimuli. The raw vocal waveforms were decomposed by a gammatone filter bank into 128 sub-bands with centre frequencies between 100 Hz and 8,000 Hz. The individually extracted narrowband Hilbert envelopes were transformed according to a power law to simulate the perception of sound volume in the auditory system, then averaged to obtain the broadband envelope [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e].\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\n\u003ch3\u003eElectroencephalogram\u003c/h3\u003e\n\u003cp\u003eThere are 66 EEG data channels. Electrodes placed on the right and left auditory areas. The auditory cortex area, for the left side, FT7, T7, FP7, C5, CP5, C3, CP3, for the right side, FT8, T8, FP8, C6, CP6, C4, CP4 electrodes were used.\u003c/p\u003e\n\u003ch3\u003eScalogram\u003c/h3\u003e\n\u003cp\u003e2D scalogram images are created by applying Continuous Wavelet Transform (CWT) to the raw EEG data.\u003c/p\u003e \u003cp\u003eTime frequency features are obtained from the scalogram images [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. A scalogram is the absolute value of the coefficients of the continuous wavelet transform (CWT) of a signal represented as a function of time and frequency [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e]. This study adopts the 2D scalogram approach, which uses CWT with pre-processed signals. CWT uses inner products to calculate the similarity between a waveform and an examination function such as the Fourier transform [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. The CWT is obtained by windowing the signal with a scaled and time-shifted wavelet. The CWT is obtained using the Morse analytical wavelet. The Fourier transform of the generalised Morse waveform is as follows (1) :\u003cdiv id=\"Equ1\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ1\" name=\"EquationSource\"\u003e\n$$\\:{\\psi\\:}_{P,\\gamma\\:}\\left(\\omega\\:\\right)\\:=U\\left(\\omega\\:\\right){a}_{P,\\gamma\\:}{\\omega\\:}^{\\frac{{P}^{2}}{\\gamma\\:}}{e}^{{-\\omega\\:}^{\\gamma\\:}}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e1\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003ewhere \u003cem\u003eU(ω)\u003c/em\u003e is the unit step, \u003cem\u003ea\u003c/em\u003e\u003csub\u003e\u003cem\u003eP\u003c/em\u003e,\u003cem\u003eγ\u003c/em\u003e\u003c/sub\u003e are normalisation constants, \u003cem\u003eP\u003c/em\u003e\u003csup\u003e2\u003c/sup\u003e is the time-bandwidth product, and \u003cem\u003eγ\u003c/em\u003e characterises the symmetry of the Morse waveform [\u003cspan additionalcitationids=\"CR35\" citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e].\u003c/p\u003e\n\u003ch3\u003eClassification model: ResNet-101 convolutional neural network\u003c/h3\u003e\n\u003cp\u003eDeep neural networks trained on large image collections can be used for small, labelled datasets. The pre-trained network ResNet101 will be used in this study. Deep learning-based image classification models learn to recognise features in an image by learning to recognise features from Scalogram representations.\u003c/p\u003e \u003cp\u003eResNet-101 is a convolutional neural network with a depth of 101 layers. The network was trained on over a million images in the ImageNet database and can classify images into 1,000 object categories. The network has an input image size of 224 by 224 [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e, \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e, \u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]. We can say that the pre-trained ResNet-101 network enables us to achieve excellent transfer learning results.\u003c/p\u003e \u003cp\u003eTransfer learning was performed by modifying the final layers of ResNet101 using weight transfer. In other words, fine-tuning was performed using new data.\u003c/p\u003e \u003cp\u003eTo obtain the probability distribution, the softmax output layer has been removed and a fully connected layer with two outputs for the classes and a new softmax output layer have been added.\u003c/p\u003e \u003cp\u003eThe criteria used to evaluate the performance of the classification are important for comparison with other results obtained. Accuracy, recall/sensitivity and precision and the AUC (Area Under the Curve) value were considered to evaluate the results.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eExplainable Artificial Intelligence\u003c/h2\u003e \u003cp\u003eExplainability is achieved by estimating the marginal contribution of each feature to a prediction using GradCAM.\u003c/p\u003e \u003cp\u003eGradient-weighted class activation mapping (Grad-CAM) is an explainability technique that can be used to help understand the predictions made by a deep neural network. Grad-CAM determines the importance of each neuron in a network prediction by considering the gradients of the target flowing through the deep network [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]. The Grad-CAM interpretability technique uses the gradients of the classification score relative to the final convolutional feature map. The parts of an observation with a high value for the Grad-CAM map are those that have the greatest impact on the network's score for that class. The gradients are grouped in spatial and temporal dimensions to determine the importance weights of the neurons. These weights are then used to linearly combine the activation maps and determine which features are most important for prediction [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIn the test dataset, all correctly classified scalograms can each generate a GradCAM map. The averaged image is obtained by calculating the numerical average pixel by pixel. The study was conducted on an Apple MacBook M2 Pro with 16 GB of memory, a total of 12 cores, and a 19-core graphics processor. Matlab 2025b was used.\u003c/p\u003e \u003c/div\u003e"},{"header":"RESULTS","content":"\u003cp\u003eWe will carry out our work using part of the data, without using all of the tests, with transfer learning, and with a less time-consuming process. In this study, our objective will be to perform AAD through transfer learning using a limited number of 10-second windows. The data comes from the sounds of women or men coming from the right or left, recorded under any acoustic conditions (with or without echo). Therefore, this study did not take into account different configurations, different trials, whether the sound was heard from the right or left, whether the speaker was female or male, or whether the EEG signals were taken from the right or left auditory region.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eTets classification results for 4 participants and 18 participants. Transfer learning with ResNet-101. Accuracy, sensitivity, precision, AUC values.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSensitivity/Recall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003ePrecision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eAUC\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e4 Participants\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSpeech A\u003c/p\u003e \u003cp\u003eSpeech B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e87.2\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e86.0\u003c/p\u003e \u003cp\u003e88.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e88.2\u003c/p\u003e \u003cp\u003e86.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e94.6\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e18 Participants\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSpeech A\u003c/p\u003e \u003cp\u003eSpeech B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e79.7\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e79.5\u003c/p\u003e \u003cp\u003e80.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e79.9\u003c/p\u003e \u003cp\u003e79.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e87.7\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe results of a test. The values are given as percentages(%).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe signals from the 14 channels located on the auditory areas during 30 trials are transformed into scalogram images over 10 seconds. 90% of the data was used for training and 10% for testing. During training, care was taken to avoid overfitting or underfitting.\u003c/p\u003e \u003cp\u003eIn Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, the results for 5 participants and 18 participants are presented separately. The dataset is balanced with 15,330 images per class for 18 participants and approximately 4,000 images for 4 participants. Training for 4 participants lasted approximately 2 hours and 8 hours for 18 participants.\u003c/p\u003e \u003cp\u003eHigh performance was achieved with 87.2% accuracy and 94.6% AUC value for 4 participants and 79.7% accuracy and 87.7% AUC value for 18 participants. In this study, with the use of deep learning techniques, the encoding of the sound envelope considered in a noisy environment in the brain was solved with high performance.\u003c/p\u003e \u003cp\u003eThe results of the study on explainability are shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e. Sample GradCAM map images are shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eb. The averaged image of the GradCAM maps is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ec.\u003c/p\u003e"},{"header":"DISCUSSION","content":"\u003cp\u003eDeep learning models have yielded promising results in various tasks related to brain signal processing. This study has provided new insights using an explainability approach while attempting to decode the auditory attention code using EEG data in competing speech scenarios.\u003c/p\u003e \u003cp\u003eIn this study, the auditory attention code was decoded in short-term conversations using transfer learning with 10-second windows. This study was performed using two-dimensional convolutional neural networks with approximately half the data and lower computational requirements. Different configurations, different trials, whether the sound was heard from the right or left, whether the speaker was female or male, and whether the EEG signals were taken from the right or left auditory region were not considered. If they had been, the results would have been much higher.\u003c/p\u003e \u003cp\u003eThe reconstruction accuracy for attended and ignored speech, was calculated for ResNet-101. Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e shows that for the model trained using data classified from the EEG signals of four individuals, we achieved an accuracy of 87.2% and an AUC of 94.6%. Eger, 0.5 \u0026lt; AUC \u0026lt; 1 indicates that the model has some ability to distinguish between categories. The closer the value is to 1, the better the model's performance. This remains relatively high compared to the literature studied. In the context of classifying data from 18 individuals, we achieved an accuracy of 79.7% and an AUC of 87.7%. The more people with different EEG waves, the lower the model's performance results. In this case, inter-subject modelling far exceeds the random accuracy of 50%. Cortical activity phase-locks to the envelope of an attended speech source, and the encoding of the considered sound envelope in the brain has been solved with high performance.\u003c/p\u003e \u003cp\u003eIn Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, we can also see that the sensitivity and precision results for both speech A (79.5% and 79.9% respectively) and speech B (80% and 79.6% respectively) are very similar. This shows us that the two classes are equally distinguishable.\u003c/p\u003e \u003cp\u003eUsing 10-second windows allows us to evaluate speech and/or language. Shorter durations are used more for evaluating voice and/or sound. In our case, our work allows us to evaluate not only hearing but also the integration, identification, and comprehension of what is heard.\u003c/p\u003e \u003cp\u003eFurthermore, through testing, we found that using data from only 14 electrodes near the auditory lobes maintains better performance compared to using 64 electrodes.\u003c/p\u003e \u003cp\u003eExisting AAD algorithms often exploit the powerful modelling capabilities of deep learning, but few of them take XAI into account. The results of the study on explainability are shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e. The averaged image of the GradCAM maps is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ec. The model exhibits strong interpretability, revealing that the left and right auditory lobes are more active during speaker scenarios. We can see in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ec that the identification of the two speeches listened to is based on the increase or decrease in alpha and beta waves, and more specifically in the Sensorimotor Rhythm (SMR) [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e]. This could provide to neuroscientists a new and useful data-driven analysis tool, aimed not only at decoding but also at analyzing functional connectivity estimates.\u003c/p\u003e \u003cp\u003eEEG rhythms allow us to observe different functional states of the brain. Generally, high-frequency, low-amplitude rhythms are associated with alertness and wakefulness, or with the dreaming phases of sleep. When the cortex is most engaged in analysing information from sensory input or internal processes (wakefulness), cortical neuron activity is relatively high [\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e]. Alpha waves (8 to 12 Hz) are associated with states of calmness during wakefulness [\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e]. Attentional processing or cognitive tasks attenuate the alpha waves. alpha waves may be used to predict mistakes, with open eyes can be a predictor of visual information processing in working memory and perceptual visual learning [\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e]. Beta waves (15 to 30 Hz) correspond to periods of normal or intense activity during the day [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e]. Beta rhythm brain activity is characteristic of normal wakefulness when the subject has their eyes open and is performing a perceptual (vision, hearing, touch) or mental (arithmetic, cognitive, complex) task. Low-amplitude beta waves with multiple rapid frequency changes are often associated with an active, busy, or even anxious state of mind, with high concentration. Sudden peaks in beta activity are associated with enhanced sensory feedback during control. [\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e]. The sensorimotor rhythm (SMR) (12 to 15 Hz) corresponds to the state of concentration just before performing an action [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e]. From a phenomenological point of view, its amplitude tends to increase when the sensorimotor regions of the brain are at rest, such as in a state of immobility. The functional significance of SMR is linked, among other things, to a state of relaxed alertness without movement and reduced sensory processing during periods of motor rest [\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e]. This information is consistent with what we obtain on XAI maps in terms of frequency and activation. Certain areas of the GradCAM map are activated or not. This shows that alpha, beta and SMR waves are more or less intense depending on the type of speech. Although preliminary, this study shows the potential of combining an XAI approach with more traditional methods to increase the robustness of EEG source analysis in hearing and AAD. The same combined approach can be easily transposed to study other cognitive tasks.\u003c/p\u003e \u003cp\u003eThe results of this study have important implications for understanding and assessing auditory attention, which is essential for applications such as brain-computer interface (BCI) systems and the development of hearing aids. Sound can be easily detected even in very noisy environments, so our results demonstrate the existence of the neural foundations necessary for human communication in realistic situations. This is also demonstrated by the brain's electrical activity. Deep learning models have achieved promising results in various tasks related to brain signal processing.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\u003cul\u003e \u003cli\u003e \u003cp\u003eNeuroscience and Brain-Computer Interface Research\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eIdentifying Hearing-Impaired Individuals Using a Non-Invasive Method\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eDeveloping Artificial Systems\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003cp\u003e\u003c/p\u003e "},{"header":"CONCLUSIONS AND LIMITATIONS","content":"\u003cp\u003eIn this study, we uses EEG signals for AAD in order to predict the sound being listened to based on neural activity in the auditory cortex with deep learning models. The EEG signals are converted into graphs and then presented at the input of the ResNet-101 architecture. We obtained an AUC value of 87.7% with data from 18 participants. We then used XAI the GradCAM technique to understand which characteristics and types of EEG components are used for AAD. Then, auditory attention is elucidated through explainability. This study shows that alpha, beta and sensorimotor rhythms are more or less intense depending on the type of speech.\u003c/p\u003e\u003cp\u003eThis Study can be used as a reference for artificial intelligence algorithms that process sound in real environments. In the future, the same study could be conducted to investigate the decoding of the neural code of sounds in noise within the subcortical region of the auditory system, for example in the inferior colliculus, independently of the presence of the auditory cortex.\u003c/p\u003e\u003cp\u003eThe effect of factors such as the gender of the person, the gender of the listener on AAD performance and subjective experience has not been studied. The AAD algorithm must demonstrate accurate performance in a variety of real-world listening situations to be implemented in future hearing aids.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cp\u003eAAD : Auditory attention decoding\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eEEG : ElectroEncephaloGram\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eGradCAM :\u0026nbsp;Gradient-weighted class activation mapping\u003c/p\u003e\n\u003cp\u003eXAI : eXplainable Artificial Intelligence\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eDTU : Technical University of Denmark\u003c/p\u003e\n\u003cp\u003eCNN : convolutional neural networks\u003c/p\u003e\n\u003cp\u003eDBGAN : Dual Branch Generative Adversarial Network\u003c/p\u003e\n\u003cp\u003eSHAP : Shapley additive explanations\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eLIME:\u0026nbsp;Local interpretable model-agnostic explanations\u003c/p\u003e\n\u003cp\u003eCOCOHA : Cognitive Control of a Hearing Aid\u003c/p\u003e\n\u003cp\u003eCWT : Continuous Wavelet Transform\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eAUC : Area Under the Curve\u003c/p\u003e\n\u003cp\u003eSMR : SensoriMotor Rhythm\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eBCI : Brain-Computer Interface\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis research did not require ethical approval as the data is publicly available and ethical approval has been obtained by the data publishers.\u003c/p\u003e\n\u003cp\u003eThis statement is not applicable to the present study.\u003c/p\u003e\n\u003cp\u003eConsent to participate statement is not applicable to the present study\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe data that support the endings of this study are available publicly from the the Technical University of Denmark, created as part of the COCOHA. The dataset is publicly accessible via https://zenodo.org/records/1199011\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting Interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe author has no relevant financial or non-financial interests to disclose.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe author declare that no funds, grants, or other support were received during the preparation of this manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe author, Fatma \u0026Ouml;ZCAN, contributed to the design and implementation of the study, performed the data analysis, and wrote the manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eSridhar, G. (2025). \u003cem\u003eImproving auditory attention decoding in noisy environments for listeners with hearing impairment through contrastive learning\u003c/em\u003e. \u003cem\u003eJournal Of Neural Engineering\u003c/em\u003e, 22(3).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCeravolo, L. (2024). \u003cem\u003eFunctional and causal neural mechanisms of human voice perception in noisy situations\u003c/em\u003e BioRxiv.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMai, A., Hillyard, S. A., \u0026amp; Strauss, D. J. (2025). Linear modeling of brain activity during selective attention to continuous speech: the critical role of the N1 effect in event-related potentials to acoustic edges. \u003cem\u003eCognitive Neurodynamics\u003c/em\u003e, \u003cem\u003e19\u003c/em\u003e(1), 110.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGeravanchizadeh, M., Shaygan Asl, A., \u0026amp; Danishvar, S. (2024). \u003cem\u003eSelective Auditory Attention Detection Using Combined Transformer and Convolutional Graph Neural Networks\u003c/em\u003e. \u003cem\u003eBioengineering (Basel)\u003c/em\u003e, 11(12).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMahjoory, K., Bahmer, A., \u0026amp; Henry, M. J. (2024). Convolutional neural networks can identify brain interactions involved in decoding spatial auditory attention. \u003cem\u003ePlos Computational Biology\u003c/em\u003e, \u003cem\u003e20\u003c/em\u003e(8), e1012376.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAlishbayli, A. (2024). \u003cem\u003eProcessing of Statistically Defined Sounds in the Auditory Cortex\u003c/em\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKausar, T., et al. (2024). Auditory-GAN: deep learning framework for improved auditory spatial attention detection. \u003cem\u003ePeerJ Comput Sci\u003c/em\u003e, \u003cem\u003e10\u003c/em\u003e, e2394.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYuan, L., et al. (2025). \u003cem\u003eFrequency-Based Alignment of EEG and Audio Signals Using Contrastive Learning and SincNet for Auditory Attention Detection\u003c/em\u003e. arxiv.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIssa, M. F., et al. (2024). On the speech envelope in the cortical tracking of speech. \u003cem\u003eNeuroimage\u003c/em\u003e, \u003cem\u003e297\u003c/em\u003e, 120675.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTanveer, M. A. (2024). \u003cem\u003eEnvelope Based Deep Source Separation and EEG Auditory Attention Decoding for Speech and Music\u003c/em\u003e, in \u003cem\u003eEUSIPCO 2024\u003c/em\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDeoisres, S., et al. (2025). Comparing approaches for predicting behavioural speech-in-noise performance using cortical responses to unattended stimuli. \u003cem\u003eHearing Research\u003c/em\u003e, \u003cem\u003e457\u003c/em\u003e, 109197.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSun, Q., et al. (2025). \u003cem\u003eAttention Detection Using EEG Signals and Machine Learning: A Review\u003c/em\u003e (p. 22). Machine Intelligence Research.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThornton, M., Mandic, D., \u0026amp; Reichenbach, T. (2024). \u003cem\u003eComparison of linear and nonlinear methods for decoding selective attention to speech from ear-EEG recordings\u003c/em\u003e. \u003cem\u003earXiv\u003c/em\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMesgarani, N., \u0026amp; Chang, E. F. (2012). Selective cortical representation of attended speaker in multi-talker speech perception. \u003cem\u003eNature\u003c/em\u003e, \u003cem\u003e485\u003c/em\u003e(7397), 233\u0026ndash;236.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThornton, M., Mandic, D., \u0026amp; Reichenbach, T. (2022). \u003cem\u003eRobust decoding of the speech envelope from EEG recordings through deep neural networks\u003c/em\u003e. \u003cem\u003eJournal Of Neural Engineering\u003c/em\u003e, 19(4).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMagosso, E., Bruno, P., \u0026amp; Borra, D. (2025). \u003cem\u003eCombining EEG Oscillation Analysis and Explainable Artificial Intelligence for Characterizing Visuospatial Attention\u003c/em\u003e, in \u003cem\u003eSpringer Nature Switzerland AG 2025\u003c/em\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDi Martino, F., \u0026amp; Delmastro, F. (2023). Explainable AI for clinical and remote health applications: a survey on tabular and time series data. \u003cem\u003eArtificial Intelligence Review\u003c/em\u003e, \u003cem\u003e56\u003c/em\u003e(6), 5261\u0026ndash;5315.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDissanayake, T., et al. (2021). \u003cem\u003eA Robust Interpretable Deep Learning Classifier for Heart Anomaly Detection Without Segmentation\u003c/em\u003e (p. 25). IEEE JOURNAL OF BIOMEDICAL AND HEALTH INFORMATICS.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTjoa, E., \u0026amp; Guan, C. (2021). \u003cem\u003eA Survey on Explainable Artificial Intelligence (XAI): Toward Medical XAI\u003c/em\u003e (p. 32). IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVilone, G., \u0026amp; Longo, L. (2021). \u003cem\u003eNotions of explainability and evaluation approaches for explainable artificial intelligence\u003c/em\u003e. \u003cem\u003eInformation Fusion\u003c/em\u003e, 76.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBollens, L. (2025). \u003cem\u003eAuditory EEG Decoding Challenge for ICASSP 2024\u003c/em\u003e. \u003cem\u003eIEEE Signal Processing\u003c/em\u003e, 6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMeng, Q., et al. (2025). EEG-based cross-subject passive music pitch perception using deep learning models. \u003cem\u003eCognitive Neurodynamics\u003c/em\u003e, \u003cem\u003e19\u003c/em\u003e(1), 6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCiccarelli, G., et al. (2019). Comparison of Two-Talker Attention Decoding from EEG with Nonlinear Neural Networks and Linear Methods. \u003cem\u003eScientific Reports\u003c/em\u003e, \u003cem\u003e9\u003c/em\u003e(1), 11538.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWong, D. D. E., et al. (2018). A Comparison of Regularization Methods in Forward and Backward Models for Auditory Attention Decoding. \u003cem\u003eFront Neurosci\u003c/em\u003e, \u003cem\u003e12\u003c/em\u003e, 531.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEskandariNasab, M., et al. (2024). A GRU-CNN model for auditory attention detection using microstate and recurrence quantification analysis. \u003cem\u003eScientific Reports\u003c/em\u003e, \u003cem\u003e14\u003c/em\u003e(1), 8861.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuang, S., Wang, Y., \u0026amp; Luo, H. (2025). \u003cem\u003eA dual-branch generative adversarial network with self-supervised enhancement for robust auditory attention decoding\u003c/em\u003e. Engineering Applications of Artificial Intelligence.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFiorini, L., Lucia, S., \u0026amp; Pietrogiacomi, F. (2024). \u003cem\u003eInsights_Into_Pre-Stimulus_Activity_for_EEG-Based_Machine_Learning_Sensory_Classification\u003c/em\u003e, in \u003cem\u003eIEEE\u003c/em\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAlmadhor, A., et al. (2025). An interpretable XAI deep EEG model for schizophrenia diagnosis using feature selection and attention mechanisms. \u003cem\u003eFrontiers In Oncology\u003c/em\u003e, \u003cem\u003e15\u003c/em\u003e, 1630291.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFuglsang, S. A., Wong, D. E., \u0026amp; Hjortkjaer, J. (2018). \u003cem\u003eEEG and audio dataset for auditory attention decoding\u003c/em\u003e. Zenodo.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJianbiao, M., et al. (2023). EEG signal classification of tinnitus based on SVM and sample entropy. \u003cem\u003eComput Methods Biomech Biomed Engin\u003c/em\u003e, \u003cem\u003e26\u003c/em\u003e(5), 580\u0026ndash;594.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAslan, Z., \u0026amp; Akin, M. (2022). A deep learning approach in automated detection of schizophrenia using scalogram images of EEG signals. \u003cem\u003ePhys Eng Sci Med\u003c/em\u003e, \u003cem\u003e45\u003c/em\u003e(1), 83\u0026ndash;96.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eByeon, Y. H., Pan, S. B., \u0026amp; Kwak, K. C. (2019). \u003cem\u003eIntelligent Deep Models Based on Scalograms of Electrocardiogram Signals for Biometrics\u003c/em\u003e. \u003cem\u003eSensors (Basel)\u003c/em\u003e, 19(4).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLoey, M., \u0026amp; Mirjalili, S. (2021). COVID-19 cough sound symptoms classification from scalogram image representation using deep learning models. \u003cem\u003eComputers In Biology And Medicine\u003c/em\u003e, \u003cem\u003e139\u003c/em\u003e, 105020.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLilly, J. M., \u0026amp; Olhede, S. C. (2012). Generalized Morse wavelets as a superfamily of analytic wavelets. \u003cem\u003eIEEE Transactions on Signal Processing\u003c/em\u003e, \u003cem\u003e60\u003c/em\u003e(11), 6036\u0026ndash;6041.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMathworks (2025). \u003cem\u003eMatlab R2025a\u003c/em\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOlhede, S. C., \u0026amp; Walden, A. T. (2002). Generalized morse wavelets. \u003cem\u003eIEEE Transactions on Signal Processing\u003c/em\u003e, \u003cem\u003e50\u003c/em\u003e(11), 2661\u0026ndash;2670.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHe, K. (2016). \u003cem\u003eDeep Residual Networks\u003c/em\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHe, K., et al. (2015). \u003cem\u003eDeep Residual Learning for Image Recognition\u003c/em\u003e. arxiv.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRamprasaath, R., et al. (2017). \u003cem\u003eGrad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization\u003c/em\u003e. arxiv.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePouessel, C. (2019). \u003cem\u003eTout savoir exactement de mani\u0026egrave;re simple sur les ondes c\u0026eacute;r\u0026eacute;brales\u003c/em\u003e. ; Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.cabinet-neurofeedback.fr/2019/06/23/les-ondes-cerebrales/\u003c/span\u003e\u003cspan address=\"https://www.cabinet-neurofeedback.fr/2019/06/23/les-ondes-cerebrales/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBear, M. F., Connors, B. W., \u0026amp; Paradiso, M. A. (2015). \u003cem\u003eNeuroscience \u0026ndash; Exploring the Brain.\u003c/em\u003e \u0026Eacute;ditions Pradel.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWikipedia (2025). \u003cem\u003eRythme sensorimoteur\u003c/em\u003e. ; Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://fr.wikipedia.org/wiki/Rythme_sensorimoteur\u003c/span\u003e\u003cspan address=\"https://fr.wikipedia.org/wiki/Rythme_sensorimoteur\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Auditory attention decoding, Electroencephalography, Scalogram, Auditory Cortex, Explainable Artificial Intelligence","lastPublishedDoi":"10.21203/rs.3.rs-8503109/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8503109/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eHuman communication involves the simultaneous processing of sounds. In order to function in various complex auditory conditions, the auditory system distinguishes sound sources by reducing the impact of ambient noise. Auditory attention decoding (AAD) uses brain signals to identify conversations being listened to in an environment where several people are speaking simultaneously. Our goal was to use electroencephalogram (EEG) signals for AAD in order to predict the sound being listened to based on neural activity in the auditory cortex with deep learning models. The EEG signals are converted into graphs and then presented at the input of the ResNet-101 architecture. We obtained an AUC value of 94.6% with data from 4 participants and 87.7% with data from 18 participants. We then evaluated the model using the GradCAM technique to understand which characteristics and types of EEG components are used for AAD. Finally, auditory attention is elucidated through explainability. This study shows that alpha, beta and sensorimotor rhythms are more or less intense depending on the type of speech. Although preliminary, this study shows the potential of combining an explainability to increase the robustness of EEG source analysis for AAD.\u003c/p\u003e","manuscriptTitle":"Explainable Eeg for Auditory Attention Decoding","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-01-13 10:44:41","doi":"10.21203/rs.3.rs-8503109/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"68349ccb-01a0-413f-b10c-a2f2643db09e","owner":[],"postedDate":"January 13th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-02-25T08:12:07+00:00","versionOfRecord":[],"versionCreatedAt":"2026-01-13 10:44:41","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8503109","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8503109","identity":"rs-8503109","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.