{"paper_id":"33fc3ac8-7e36-421c-a6ce-75011d494f31","body_text":"Detection of COVID-19 in smartphone-based breathing\nrecordings using CNN-BiLSTM: a pre-screening deep\nlearning tool\nMohanad Alkhodari 1*¤, Ahsan H. Khandoker 1,\n1 Healthcare Engineering Innovation Center (HEIC), Department of Biomedical\nEngineering, Khalifa University, Abu Dhabi, UAE\n¤Current Address: Healthcare Engineering Innovation Center (HEIC), Department of\nBiomedical Engineering, Khalifa University, Abu Dhabi, UAE\n* mohanad.alkhodari@ku.ac.ae\nAbstract\nThis study was sought to investigate the feasibility of using smartphone-based breathing\nsounds within a deep learning framework to discriminate between COVID-19, including\nasymptomatic, and healthy subjects. A total of 480 breathing sounds (240 shallow and\n240 deep) were obtained from a publicly available database named Coswara. These sounds\nwere recorded by 120 COVID-19 and 120 healthy subjects via a smartphone microphone\nthrough a website application. A deep learning framework was proposed herein the\nrelies on hand-crafted features extracted from the original recordings and from the\nmel-frequency cepstral coefficients (MFCC) as well as deep-activated features learned by\na combination of convolutional neural network and bi-directional long short-term memory\nunits (CNN-BiLSTM). Analysis of the normal distribution of the combined MFCC values\nshowed that COVID-19 subjects tended to have a distribution that is skewed more\ntowards the right side of the zero mean (shallow: 0.59\n±1.74, deep: 0.65 ±4.35). In\naddition, the proposed deep learning approach had an overall discrimination accuracy\nof 94.58% and 92.08% using shallow and deep recordings, respectively. Furthermore,\nit detected COVID-19 subjects successfully with a maximum sensitivity of 94.21%,\nspecificity of 94.96%, and area under the receiver operating characteristic (AUROC)\ncurves of 0.90. Among the 120 COVID-19 participants, asymptomatic subjects (18\nsubjects) were successfully detected with 100.00% accuracy using shallow recordings and\n88.89% using deep recordings. This study paves the way towards utilizing smartphone-\nbased breathing sounds for the purpose of COVID-19 detection. The observations\nfound in this study were promising to suggest deep learning and smartphone-based\nbreathing sounds as an effective pre-screening tool for COVID-19 alongside the current\nreverse-transcription polymerase chain reaction (RT-PCR) assay. It can be considered as\nan early, rapid, easily distributed, time-efficient, and almost no-cost diagnosis technique\ncomplying with social distancing restrictions during COVID-19 pandemic.\nIntroduction 1\nCorona virus 2019 (COVID-19), which is a novel pathogen of the severe acute respiratory 2\nsyndrome coronavirus 2 (SARS-Cov-2), appeared first in late November 2019 and ever 3\nsince, it has caused a global epidemic problem by spreading all over the world [1]. 4\nAccording to the world heath organization (WHO) April 2021 report [2], there have been 5\nSeptember 18, 2021 1/20\n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted September 22, 2021. ; https://doi.org/10.1101/2021.09.18.21263775doi: medRxiv preprint \nNOTE: This preprint reports new research that has not been certified by peer review and should not be used to guide clinical practice.\n\nnearly 150 million confirmed cases and over 3 million deaths since the pandemic broke 6\nout in 2019. Additionally, the United States (US) have reported the highest number 7\nof cumulative cases and deaths with over 32.5 million and 500,000, respectively. These 8\nhuge numbers have caused many healthcare services to be severely burdened especially 9\nwith the ability of the virus to develop more genomic variants and spread more readily 10\namong people. India, which is one of the world’s biggest suppliers of vaccines, is now 11\nseverely suffering from the pandemic after the explosion of cases due to a new variant of 12\nCOVID-19. It has reached more than 17.5 million confirmed cases, setting it behind the 13\nUS as the second worst hit country [2,3]. 14\nCOVID-19 patients usually range from being asymptomatic to developing pneumonia 15\nand in severe cases, death. In most reported cases, the virus remains incubation 16\nfor a period of 1 to 14 days before the symptoms of an infection start arising [4]. 17\nPatients carrying COVID-19 have exhibited common signs and symptoms including 18\ncough, shortness of breath, fever, fatigue, and other acute respiratory distress syndromes 19\n(ARDS) [5, 6]. Most infected people suffer from mild to moderate viral symptoms, 20\nhowever, they end up by being recovered. On the other hand, patients who develop 21\nsevere symptoms such as severe pneumonia are mostly people over 60 years of age 22\nwith conditions such as diabetes, cardiovascular diseases (CVD), hypertension, and 23\ncancer [4, 5]. On most cases, the early diagnosis of COVID-19 helps in preventing its 24\nspreading and development to severe infection stages. This is usually done by following 25\nsteps of early patient isolation and contact tracing. Furthermore, timely medication and 26\nefficient treatment reduces symptoms and results in lowering the mortality rate of this 27\npandemic [7]. 28\nThe current gold standard in diagnosing COVID-19 is the reverse-transcription 29\npolymerase chain reaction (RT-PCR) assay [8,9]. It is the most commonly used technique 30\nworldwide to successfully confirm the existence of this viral infection. Additionally, 31\nexaminations of the ribonucleic acid (RNA) in patients carrying the virus provide further 32\ninformation about the infection, however, it requires longer time for diagnosis and is not 33\nconsidered as accurate as other diagnostic techniques [10]. The integration of computed 34\ntomography (CT) screening is another effective diagnostic tool (sensitivity ≥ 90%) that 35\noften provides supplemental information about the severity and progression of COVID-19 36\nin lungs [11, 12]. CT imaging is not recommended for patients at the early stages of 37\nthe infection, i.e., showing asymptomatic to mild symptoms. It provides useful details 38\nabout the lungs in patients with moderate to severe stages due to the disturbance in the 39\npulmonary tissues and its corresponding functions [13]. However, CT imaging may not 40\nbe available in all public healthcare services, especially for countries who are swamped 41\nwith the pandemic, due to its costs and additional maintenance requirements. Therefore, 42\nbiological signals, such as coughing and breathing sounds, could be another promising 43\ntool to indicate the existence of the viral infection [14]. In addition, due to the simplicity 44\nin recording respiratory signals, lung sounds could carry useful information about the 45\nviral infection, and thus, could set an early alert to the patient before moving on with 46\nfurther medication procedures. In addition, the new emerging algorithms in artificial 47\nintelligence (AI) could be a key to enhance the sensitivity of detection for positive cases 48\ndue to its ability to generalize over a wide set of data [15]. 49\nMany studies have investigated the information carried by respiratory sounds in 50\npatients tested positive for COVID-19 [16 –18]. Furthermore, it has been found that 51\nvocal patterns extracted from COVID-19 patients’ speech recordings carry indicative 52\nbiomarkers for the existence of the viral infection [19]. In addition, a telemedicine 53\napproach was also explored to observe evidences on the sequential changes in respiratory 54\nsounds as a result of COVID-19 infection [20]. Most recently, AI was utilized in one 55\nstudy to recognize COVID-19 in cough signals [21] and in another to evaluate the severity 56\nof patients’ illness, sleep quality, fatigue, and anxiety through speech recordings [22]. 57\nSeptember 18, 2021 2/20\n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted September 22, 2021. ; https://doi.org/10.1101/2021.09.18.21263775doi: medRxiv preprint \n\nFig 1. A graphical abstract of the complete procedure followed in this study . The input data includes breathing\nsounds collected from an open-access database for respiratory sounds (Coswara [23]) recorded via smartphone microphone. The\ndata includes a total of 240 participants, out of which 120 subjects were suffering from COVID-19, while the remaining 120 were\nhealthy (control group). A deep learning framework was then utilized based on hand-crafted features extracted by feature\nengineering techniques, as well as deep-activated features extracted by a combination of convolutional and recurrent neural\nnetwork. The performance was then evaluated and further discussed on the use of artificial intelligence (AI) as a successful\npre-screening tool for COVID-19.\nDespite of the high levels of performance achieved in the aforementioned AI-based 58\nstudies, further investigations on the capability of respiratory sounds in carrying useful 59\ninformation about COVID-19 are still required, especially when embedded within the 60\nframework of sophisticated AI-based algorithms. Furthermore, due to the explosion in 61\nthe number of confirmed positive COVID-19 cases all over the world, it is essential to 62\nensure providing a system capable of recognizing the disease in signals recording through 63\nportable devices, such as computers or smartphones, instead of regular clinic-based 64\nelectronic stethoscopes. 65\nMotivated by the aforementioned, a complete deep learning approach is proposed in 66\nthis paper for a successful detection of COVID-19 using only breathing sounds recorded 67\nthrough a microphone of a smartphone device (Fig. 1). The proposed approach serves 68\nas a rapid, no-cost, and easily distributed pre-screening tool for COVID-19, especially 69\nfor countries who are in a complete lockdown due to the wide spread of the pandemic. 70\nAlthough the current gold standard, RT-PCR, provides high success rates in detecting 71\nthe viral infection, it has various limitations including the high expenses involved with 72\nequipment and chemical agents, requirement of expert nurses and doctors for diagnosis, 73\nviolation of social distancing, and the long testing time required to obtain results (2- 74\n3 days). Thus, the development of a deep learning model overcomes most of these 75\nlimitations and allows for a better revival in the healthcare and economic sectors in 76\nseveral countries. 77\nFurthermore, the novelty of this work lies in utilizing smartphone-based breathing 78\nrecordings within this deep learning model, which, when compared to conventional 79\nrespiratory auscultation devices, i.e., electronic stethoscopes, are more preferable due 80\nto their higher accessibility by wider population. This plays an important factor in 81\nobtaining medical information about COVID-19 patients in a timely manner while at 82\nthe same time maintaining an isolated behaviour between people. Additionally, this 83\nstudy covers patients who are mostly from India, which is severely suffering from a new 84\ngenomic variant (first reported in December 2020) of COVID-19 capable of escaping the 85\nimmune system and most of the available vaccines [2,24]. Thus, it gives an insight on 86\nthe ability of AI algorithms in detecting this viral infection in patients carrying this new 87\nSeptember 18, 2021 3/20\n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted September 22, 2021. ; https://doi.org/10.1101/2021.09.18.21263775doi: medRxiv preprint \n\nvariant, including asymptomatic. Lastly, the study presented herein investigates signal 88\ncharacteristics contaminated within shallow and deep breathing sounds of COVID-19 89\nand healthy subjects through deep-activated attributes (neural network activations) of 90\nthe original signals as well as wide attributes (hand-crafted features) of the signals and 91\ntheir corresponding mel-frequency cepstrum (MFC). The utilization of one-dimensional 92\n(1D) signals within a successful deep learning framework allows for a simple, yet effective, 93\nAI design that does not require heavy memory requirements. This serves as a suitable 94\nsolution for further development of telemedicine and smartphone applications for COVID- 95\n19 (or other pandemics) that can provide real-time results and communications between 96\npatients and clinicians in an efficient and timely manner. Therefore, as a pre-screening 97\ntool for COVID-19, this allows for a better and faster isolation and contact tracing than 98\ncurrently available techniques. 99\nMaterials and Methods 100\nDataset collection and subjects information 101\nThe dataset used in this study was obtained from Coswara [23], which is a project 102\naiming towards providing an open-access database for respiratory sounds of healthy 103\nand unhealthy individuals, including those suffering from COVID-19. The project is a 104\nworldwide respiratory data collection effort that was first initiated in August, 7th 2020. 105\nEver since, it has collected data from more than 1,600 participants (Male: 1185, Female: 106\n415) from allover the world (mostly Indian population). The database was approved by 107\nthe Indian institute of science (IISc), human ethics committee, Bangalore, India, and 108\nconforms to the ethical principles outlined in the declaration of Helsinki. No personally 109\nidentifiable information about participants was collected and the participants’ data was 110\nfully anonymized during storage in the database. 111\nThe database includes breath, cough, and voice sounds acquired via crowdsourcing 112\nusing an interactive website application that was built for smartphone devices [25]. The 113\naverage interaction time with the application was 5-7 minutes. All sounds were recorded 114\nusing the microphone of a smartphone and sampled with a sampling frequency of 48 kHz. 115\nThe participants had the freedom to select any device for recording their respiratory 116\nsounds, which reduces device-specific bias in the data. The audio samples (stored in 117\n.WAV format) for all participants were manually curated through a web interface that 118\nallows multiple annotators to go through each audio file and verify the quality as well as 119\nthe correctness of labeling. All participants were requested to keep a 10 cm distance 120\nbetween the face and the device before starting the recording. 121\nSo far, the database had a COVID-19 participants’ count of 120, which is almost 122\n1-10 ratio to healthy (control) participants. In this study, all COVID-19 participants’ 123\ndata was used, and the same number of samples from the control participants’ data 124\nwas randomly selected to ensure a balanced dataset. Therefore, the dataset used in this 125\nstudy had a total of 240 subjects (COVID-19: 120, Control: 120). The demographic 126\nand clinical information of the selected subjects is provided in Table 1. Furthermore, 127\nonly breathing sounds of two types, namely shallow and deep, were obtained from every 128\nsubject and used for further analysis (examples from the shallow breathing dataset are 129\nshown in Fig. 2). To ensure the inclusion of maximum information from each breathing 130\nrecording as well as to cover at least 2-4 breathing cycles (inhale and exhale), a total of 131\n16 seconds were considered, as the normal breathing pattern in adults ranges between 132\n12 to 18 breaths per minute [26]. All recordings with less than 16 seconds were padded 133\nwith zeros. Furthermore, the final signals were resampled with a sampling frequency of 134\n4 kHz. 135\nSeptember 18, 2021 4/20\n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted September 22, 2021. ; https://doi.org/10.1101/2021.09.18.21263775doi: medRxiv preprint \n\n(a) COVID-19 (asymptomatic)\n (b) COVID-19 (mild)\n (c) COVID-19 (moderate)\n(d) Healthy\n (e) Healthy\n (f) Healthy\nFig 2. Examples from the shallow breathing sounds recorded via smartphone\nmicrophone along with the corresponding spectrogram for: (a-c) COVID-19 subjects\n(asymptomatic, mild, moderate), (d-f) healthy subjects.\nDeep learning framework 136\nThe deep learning framework proposed in this study (Fig. 3) includes a combination 137\nof hand-crafted features as well as deep-activated features learned through model’s 138\ntraining and reflected as time-activations of the input. To extract hand-crafted features, 139\nvarious algorithm and functions were used to obtain signal attributes from the original 140\nbreathing recording and from its corresponding mel-frequency cepstral coefficients 141\n(MFCC). In addition, deep-activated learned features were obtained from the original 142\nbreathing recording through a combined neural network that consists of convolutional 143\nand recurrent neural networks. Each part of this framework is briefly described in the 144\nfollowing subsections. 145\nHand-crafted features 146\nThese features refer to signal attributes that are extracted manually through various 147\nalgorithms and functions in a process called feature engineering. The advantage of 148\nfollowing such process is that it can extract internal and hidden information within 149\ninput data, i.e., sounds, and represent it as single or multiple values. Thus, additional 150\nknowledge about the input data can be obtained and used for further analysis and 151\nevaluation. Hand-crafted features were initially extracted from the original breathing 152\nrecordings, then, they were also extracted from the MFCC transformation of the signals. 153\nThe features included in this study are, 154\nKurtosis and Skewness: In statistics, kurtosis is a quantification measure for 155\nthe degree of extremity included within the tails of a distribution relative to the tails 156\nof a normal distribution. The more the distribution is outlier-prone, the higher the 157\nkurtosis values, and vice-versa. A kurtosis of 3 indicates that the values follow a normal 158\ndistribution. On the other hand, skewness is a measure for the asymmetry of the data 159\nthat deviates it from the mean of the normal distribution. If the skewness is negative, 160\nthen the data are more spread towards the left side of the mean, while a positive skewness 161\nSeptember 18, 2021 5/20\n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted September 22, 2021. ; https://doi.org/10.1101/2021.09.18.21263775doi: medRxiv preprint \n\nT able 1. The demographic and clinical information of COVID-19 and\nhealthy (control) subjects included in the study .\nCategory CO\nVID-19 Healthy\n(Control)Asymptomatic Mild Moderate Overall\nDemographic\ninformation\nNum\nber of subjects 18 91 11 120 120\nAge\n(Mean±Std)\n20-77\n(32.65±13.69)\n15-70\n(33.43±12.99)\n23-65\n(43.33±15.46)\n15-77\n(34.04±13.45)\n15-70\n(36.02±13.06)\nSex (Male / Female) 10 / 8 65 / 25 7 / 3 82 / 36 85/35\nComorbidities\nDiab\netes 1 7 1 9 11\nHypertension 0 6 1 7 6\nChronic lung disease 1 1 0 2 0\nIschemic heart disease 1 3 0 4 0\nPneumonia 0 3 0 3 0\nHealth\nconditions\nF\never 0 39 6 45 1\nCold 0 37 4 41 6\nCough 0 44 4 48 13\nMuscle pain 2 15 5 22 1\nLoss of smell 0 15 3 18 0\nSore throat 0 26 3 29 2\nFatigue 1 18 3 22 1\nBreathing Difficulties 1 7 6 14 0\nDiarrhoea 0 1 0 1 0\nindicates data spreading towards the right side of the mean [27]. A skewness of zero 162\nindicates that the values follow a normal distribution. Kurtosis ( k) and skewness (s) 163\ncan be calculated as, 164\nk =E\n[(X −µ)4\nσ4\n]\n(1)\ns =E\n[(X −µ)3\nσ3\n]\n(2)\nwhere X included input values, µ and σ are the mean and standard deviation values 165\nof the input, respectively, and E is an expectation operator. 166\nSample entropy: In physiological signals, the sample entropy (SampEn) provides 167\na measure for complexity contaminated within time sequences. It can be calculated 168\nthough the negative natural logarithm of a probability that segments of length m match 169\ntheir consecutive segments under a value of tolerance (r ) [28] as follows, 170\nSampEn = −log\n( segmentA\nsegmentA+1\n)\n(3)\nwhere segmentA is the first segment in the time sequence and segmentA+1 is the 171\nconsecutive segment. 172\nSpectral entropy: To measure time series irregularity, spectral entropy (SE) pro- 173\nvides a frequency domain entropy measure as a sum of the normalize signal spectral 174\npower [29]. Based on Shannon’s entropy, the SE can be calculated as, 175\nSeptember 18, 2021 6/20\n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted September 22, 2021. ; https://doi.org/10.1101/2021.09.18.21263775doi: medRxiv preprint \n\nFig 3. The framework of deep learning followed in this study. The framework includes a combination of hand-crafted\nfeatures and deep-activated features. Deep features were obtained through a combined convolutional and recurrent neural\nnetwork (CNN-BiLSTM), and the final classification layer uses both features sets to discriminate between COVID-19 and\nhealthy subjects.\nSE = −\nN∑\nn=1\nP (n) ×log(P (n)) (4)\nwhere N is the total number of frequency points and P (n) is the probability distri- 176\nbution of the power spectrum. 177\nF ractal dimension: Higuchi and Katz [30, 31] provided two methods to measure 178\nstatistically the complexity in a time series. More specifically, fractal dimension measures 179\nprovide an index for characterizing how much a time series is self-similar over some 180\nregion of space. Higuchi ( HFD ) and Katz (KFD ) fractal dimensions can be calculated 181\nas, 182\nHFD = log(L(r))\nlog(1/r) (5)\nKFD = log(N)\nlog(N) +log(d/L(r)) (6)\nwhere L(k) is the length of the fractal curve, r is the selected time interval, N is the 183\nlength of the signal, and d is the maximum distance between an initial point to other 184\npoints. 185\nZero-crossing rate: To measure the number of times a signal has passed through 186\nthe zero point, a zero-crossing rate (ZCR) measure is provided. In other words, ZCR 187\nSeptember 18, 2021 7/20\n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted September 22, 2021. ; https://doi.org/10.1101/2021.09.18.21263775doi: medRxiv preprint \n\nrefers to the rate of sign-changes in the signals’ data points. It can be calculated as 188\nfollows, 189\nZCR = 1\nT\nT∑\nt=1\n(|xt −xt+1|) (7)\nwherext = 1 if the signal has a positive value at time stept and a value of 0 otherwise. 190\nMel-frequency cepstral coefficients (MFCC): To better represent speech and 191\nvoice signals, MFCC provides a set of coefficients of the discrete cosine transformed 192\n(DCT) logarithm of a signal’s spectrum (mel-frequency cepstrum (MFC)). It is considered 193\nas an overall representation of the information contaminated within signals regarding 194\nthe changes in its different spectrum bands [32,33]. Briefly, to obtain the coefficients, 195\nthe signals goes through several steps, namely windowing the signal, applying discrete 196\nFourier transform (DFT), calculating the log energy of the magnitude, transforming the 197\nfrequencies to the Mel-scale, and applying inverse DCT. 198\nIn this work, 13 coefficients (MFCC-1 to MFCC-13) were obtained from each breathing 199\nsound signal. For every coefficient, the aforementioned features were extracted and 200\nstored as an additional MFCC hand-crafted features alongside the original breathing 201\nsignals features. 202\n0.0.1 Deep-activated features 203\nThese features refer to attributes extracted from signals through a deep learning process 204\nand not by manual feature engineering techniques. The utilization of deep learning 205\nallows for the acquisition of optimized features extracted through deep convolutional 206\nlayers about the structural information contaminated within signals. Furthermore, it 207\nhas the ability to acquire the temporal (time changes) information carried through time 208\nsequences [34,35]. Such optimized features can be considered as a complete representation 209\nof the input data generated iteratively through an automated learning process. To achieve 210\nthis, we used an advanced neural network based on a combination of convolutional neural 211\nnetwork and bi-directional long short-term memory (CNN-BiLSTM). 212\nNeural network architecture: The structure of the network starts by 1D convo- 213\nlutional layers. In deep learning, convolutions refer to a multiple number of dot products 214\napplied to 1D signals on pre-defined segments. By applying consecutive convolutions, 215\nthe network extracts deep attributes (activations) to form an overall feature map for the 216\ninput data [35]. A single convolution on an input x0\ni = [x1,x 2,...,x n], where n is the 217\ntotal number of points, is usually calculated as, 218\nclj\ni =h(bj +\nM∑\nm=1\nwj\nmxj\ni+m−1) (8)\nwhere l is the layer index, h is the activation function, b is the bias of the jth feature 219\nmap, M is the kernel size, wj\nm is the weight of the jth feature map and mth filter index. 220\nIn this work, three convolutional layers were used to form the first stage of the deep 221\nneural network. The kernel sizes of each layer are [9, 1], [5, 1], and [3, 1], respectively. 222\nFurthermore, the number of filters increases as the network becomes deeper, that is 16, 223\n32, and 64, respectively. Each convolutional layer was followed by a max-pooling layer 224\nto reduce the dimensionality as well as the complexity in the model. The max-pooling 225\nkernel size decreases as the network gets deeper with a [8, 1], [4, 1], and [2, 1] kernels 226\nfor the three max-pooling layers, respectively. It is worth noting that each max-pooling 227\nlayer was followed by a batch normalization (BN) layer to normalize all filters as well as 228\nby a rectified linear unit (ReLU) layer to set all values less than zero in the feature map 229\nto zero. The complete structure is illustrated in Fig. 3. 230\nSeptember 18, 2021 8/20\n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted September 22, 2021. ; https://doi.org/10.1101/2021.09.18.21263775doi: medRxiv preprint \n\nThe network continues with additional extraction of temporal features through bi- 231\ndirectional LSTM units. In recurrent neural networks, LSTM units allows for the 232\ndetection of long short-term dependencies between time sequence data points. Thus, it 233\novercomes the issues of exploding and vanishing gradients in chain-like structures during 234\ntraining [34,36]. An LSTM block includes a collection of gates, namely input ( i), output 235\n(o), and forget (f) gates. These gates handle the flow of data as well as the processing 236\nof the input and output activations within the network’s memory. The information of 237\nthe main cell (C t) at any instance (t) within the block can be calculated as, 238\nCt =ftCt−1 +itct (9)\nwhere ct is the input to the main cell and Ct−1 includes the information at the 239\nprevious time instance. 240\nIn addition, the network performs hidden-units (ht) activations on the output and 241\nmain cell input using a sigmoid function as follows, 242\nht =otσ(ct) (10)\nFurthermore, a bi-drectional functionality (BiLSTM) allows the network to process 243\ndata in both the forward and backward direction as follows, 244\nyt =W− →h y\n− →\nhN +W← −h y\n← −\nhN +by (11)\nwhere\n− →\nhN and\n← −\nhN\nare the outputs of the hidden layers in the forward and backward 245\ndirections, respectively, for all N levels of stack and by is a bias vector. 246\nIn this work, a BiLSTM hidden units functionality was selected with a total number 247\nof hidden units of 256. Thus, the resulting output is a 512 vector (both directions) of 248\nthe extracted hidden-units of every input. 249\nBiLSTM activations: To be able to utilize the parameters that the BiLSTM units 250\nhave learned, the activations that correspond to each hidden-unit were extracted from 251\nthe network for each input signal. Recurrent neural network activations of a pre-trained 252\nnetwork are vectors that carry the final learned attributes about different time steps 253\nwithin the input [37]. In this work, these activations were the final signal attributes 254\nextracted from each input signal. Such attributes are referred to as deep-activated 255\nfeatures in this work (Fig. 3). Furthermore, they were concatenated with the hand- 256\ncrafted features alongside age and sex information and used for the final predictions by 257\nthe network. 258\n0.0.2 Network configuration and training scheme 259\nPrior to deep learning model training, several data preparation and network fine-tuning 260\nsteps were followed including data augmentation, best features selection, deciding the 261\ntraining and testing scheme, and network parameters configuration. 262\nData augmentation: Due to the small sample size available, it is critical for deep 263\nlearning applications to include augmented data. Instead of training the model on the 264\nexisting dataset only, data augmentation allows for the generation of new modified copies 265\nof the original samples. These new copies have similar characteristics of the original data, 266\nhowever, they are slightly adjusted as if they are coming from a new source (subject). 267\nSuch procedure is essential to expose the deep learning model to more variations in the 268\ntraining data. Thus, making it robust and less biased when attempting to generalize the 269\nparameters on new data [38]. Furthermore, it was essential to prevent the model from 270\nover-fitting, where the model learns exactly the input data only with a very minimal 271\ngeneralization capabilities for unseen data [39]. 272\nSeptember 18, 2021 9/20\n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted September 22, 2021. ; https://doi.org/10.1101/2021.09.18.21263775doi: medRxiv preprint \n\nIn this study, 3,000 samples per class were generated using two 1D data augmentation 273\ntechniques as follows, 274\n• Volume control: Adjusts the strength of signals in decibels (dB) for the generated 275\ndata [40] with a probability of 0.8 and gain ranging between -5 and 5 dB. 276\n• Time shift: Modifies time steps of the signals to illustrate shifting in time for the 277\ngenerated data [41] with a shifting range of [-0.005 to 0.005] seconds. 278\nBest features selection: To ensure the inclusion of the most important hand- 279\ncrafted features within the trained model, a statistical univariate chi-square test (χ2-test) 280\nwas applied. In this test, a feature is decided to be important if the observed statistical 281\nanalysis using this feature matches with the expected one, i.e., label [42]. Furthermore, 282\nan important feature indicates that it is considered significant in discriminating between 283\ntwo categories with a p-value< 0.05. The lower the p-value, the more the feature is 284\ndependent on the category label. The importance score can then be calculated as, 285\nscore = −log(p) (12)\nIn this work, hand-crafted features extracted from the original breathing signals 286\nand from the MFCC alongside the age and sex information were selected for this test. 287\nThe best 20 features were included in the final best features vector within the final 288\nfully-connected layer (along with the deep-activated features) for predictions. 289\nT raining configuration: To ensure the inclusion of the whole available data, a 290\nleave-one-out training and testing scheme was followed. In this scheme, a total of 240 291\niterations (number of input samples) were applied, where in each iteration, an i th subject 292\nwas used as the testing subject, and the remaining subjects were used for model’s training. 293\nThis scheme was essential to be followed to provide a prediction for each subject in the 294\ndataset. 295\nFurthermore, the network was optimized using adaptive moment estimation (ADAM) 296\nsolver [43] and with a learning rate of 0.001. The L2-regularization was set to 10 6 and 297\nthe mini-batch size to 32. 298\nPerformance evaluation 299\nThe performance of the proposed deep learning model in discriminating COVID-19 from 300\nhealthy subjects was evaluated using traditional evaluation metrics including accuracy, 301\nsensitivity, specificity, precision, and F1-score. These metrics can be calculated as, 302\nAccuracy = TP +TN\nTP +TN +FP +FN (13)\nSensitivity = TP\nTP +FN (14)\nSpecificity = TN\nTN +FP (15)\nPrecision = TP\nTP +FP (16)\nF 1 −score = 2TP\n2TP +FP +FN (17)\nwhere TP is the true positive, TN is the true negative, FP is the false positive, and 303\nFN is the false negative numbers in the confusion matrix. 304\nSeptember 18, 2021 10/20\n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted September 22, 2021. ; https://doi.org/10.1101/2021.09.18.21263775doi: medRxiv preprint \n\n(a) COVID-19 (asymptomatic)\n (b) COVID-19 (mild)\n (c) COVID-19 (moderate)\n(d) Healthy\n (e) Healthy\n (f) Healthy\nFig 4. Examples of the mel-frequency cepstral coefficients (MFCC)\nextracted from the shallow breathing dataset and illustrated as a normal\ndistribution of summed coefficients. (a-c) COVID-19 subjects (asymptomatic,\nmild, moderate), (d-f) healthy subjects.\nAdditionally, the area under the receiver operating characteristic (AUROC) curves 305\nwas analysed for each category to show the true positive rate (TPR) versus the false 306\npositive rate (FPR). 307\nResults 308\nAnalysis of MFCC 309\nExamples of the 13 MFCC extracted from the original shallow breathing signals are 310\nillustrated in Fig. 4 for COVID-19 and healthy subjects. Furthermore, the figure shows 311\nMFCC values (after summing all coefficients) distributed as a normal distribution. From 312\nthe figure, the normal distribution of COVID-19 subjects was slightly skewed to the right 313\nside of the mean, while the normal distribution of the healthy subjects was more towards 314\nthe zero mean, indicating that it better in representing a normal distribution. Tables 2 315\nand 3 show the values of the combined MFCC values, kurtosis, and skewness among all 316\nCOVID-19 and healthy subjects (mean±std) for the shallow and deep breathing datasets, 317\nrespectively. In both datasets, the kurtosis and skewness values for COVID-19 subjects 318\nwere slightly higher than healthy subjects. Furthermore, the average combined MFCC 319\nvalues for COVID-19 were less than those for the healthy subjects. More specifically, 320\nin the shallow breathing dataset, a kurtosis and skewness of 4.65 ±15.97 and 0.59 ±1.74 321\nwas observed for COVID-19 subjects relative to 4.47 ±20.66 and 00.19 ±1.75 for healthy 322\nsubjects. On the other hand, using the deep breathing dataset, COVID-19 subjects had 323\nSeptember 18, 2021 11/20\n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted September 22, 2021. ; https://doi.org/10.1101/2021.09.18.21263775doi: medRxiv preprint \n\nT able 2. Normal distribution analysis (mean ±std) of the combined\nmel-frequency cepstral coefficients (MFCCs) using the shallow breathing\ndataset.\nCategory Normal\ndistribution analysis\nCombined MFCC values Kurtosis Skewness\nCOVID-19\nAsymp. -0.29±0.78 1.92 ±0.94 0.45±0.78\nMild -0.24±0.70 5.46 ±18.22 0.61±1.95\nModerate -0.26±0.70 2.26 ±1.85 0.60±0.83\nOverall -0.25 ±0.72 4.65 ±15.97 0.59±1.74\nHealthy -0.11±0.75 4.47 ±20.66 0.19±1.75\nT able 3. Normal distribution analysis (mean ±std) of the combined\nmel-frequency cepstral coefficients (MFCCs) using the deep breathing\ndataset.\nCategory Normal\ndistribution analysis\nCombined MFCC values Kurtosis Skewness\nCOVID-19\nAsymp. -0.05±0.69 2.61 ±2.12 -0.17±0.97\nMild -0.15±0.63 26.63 ±15.51 0.91 ±4.95\nModerate 0.01±0.58 2.54±1.12 -0.15±0.96\nOverall -0.12 ±0.64 20.82 ±12.99 0.65 ±4.35\nHealthy 0.12±0.60 3.23±6.06 -0.36±1.08\na kurtosis and skewness of 20.82 ±152.99 and 0.65 ±4.35 compared to lower values of 324\n3.23±6.06 and -0.36±1.08 for healthy subjects. 325\nDeep learning performance 326\nThe overall performance of the proposed deep learning model is shown in Fig. 5. From 327\nthe figure, the model correctly predicted 113 and 114 COVID-19 and healthy subjects, 328\nrespectively, using the shallow breathing dataset out of the 120 total subjects (Fig. 5(a)). 329\nIn addition, only 7 COVID-19 subjects were miss-classified as healthy, whereas only 6 330\nsubjects were wrongly classified as carrying COVID-19. The correct predictions number 331\nwas slightly lower using the deep breathing dataset with a 109 and 112 for COVID-19 332\nand healthy subjects, respectively. In addition, wrong predictions were also slightly 333\nhigher with 11 COVID-19 and 8 healthy subjects. Therefore, the confusion matrices 334\nshow percentages of proportion of 94.20% and 90.80% for COVID-19 subjects using 335\nthe shallow and deep datasets, respectively. On the other hand, healthy subjects had 336\npercentages of 95.00% and 93.30% for both datasets, respectively. 337\nThe evaluation metrics (Fig. 5(b)) calculated from these confusion matrices returned 338\nan accuracy measure of 94.58% and 92.08% for the shallow and deep datasets, respectively. 339\nFurthermore, the model had a sensitivity and specificity measures of 94.21%/94.96% 340\nfor the shallow dataset and 93.16%/91.06% for the deep dataset. The precision was the 341\nhighest measure obtained for the shallow dataset (95.00%), where as the deep dataset 342\nhad the lowest value in the precision with a 90.83%. Lastly, the F1-score measures 343\nSeptember 18, 2021 12/20\n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted September 22, 2021. ; https://doi.org/10.1101/2021.09.18.21263775doi: medRxiv preprint \n\n(a) Model predictions and confusion matrices\n(b) Evaluation metrics\n (c) Receiver operating characteristic (ROC) curves\nFig 5. The performance of the deep learning model in predicting\nCOVID-19 and healthy subjects using shallow and deep breathing datasets.\n(a) model’s predictions for both datasets and he corresponding confusion matrices, (b)\nevaluation metrics including accuracy, sensitivity, specificity, precision, and F1-score, (d)\nreceiver operating characteristic (ROC) curves and corresponding area under the curve\n(AUROC) for COVID-19 and healthy subjects using both datasets.\nreturned 94.61% and 91.98% for both datasets, respectively. 344\nTo analyze the AUROC, Fig. 5(c) shows the ROC curves of predictions using both 345\nthe shallow and deep datasets. The shallow breathing dataset had an overall AUROC of 346\n0.90 in predicting COVID-19 and healthy subjects, whereas the deep breathing dataset 347\nhad a 0.86 AUROC, which is slightly lower performance in the prediction process. 348\nAdditionally, the model had high accuracy measures in predicting asymptomatic 349\nCOVID-19 subjects (Fig 6). Using the shallow breathing dataset, the model had 350\na 100.00% accuracy by predicting all subjects correctly. On the other hand, using 351\nthe deep breathing dataset, the model achieved an accuracy of 88.89% by missing 352\ntwo asymptomatic subjects. It is worth noting that few subjects had close scores 353\n(probabilities) to 0.5 using both datasets, however, the model correctly discriminated 354\nthem from healthy subjects. 355\nSeptember 18, 2021 13/20\n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted September 22, 2021. ; https://doi.org/10.1101/2021.09.18.21263775doi: medRxiv preprint \n\nFig 6. Asymptomatic COVID-19 subjects’ predictions based on the\nproposed deep learning model. The model had a decision boundary of 0.5 to\ndiscriminate between COVID-19 and healthy subjects. The values represent a\nnormalized probability regrading the confidence in predicting these subjects as carrying\nCOVID-19.\nDiscussion 356\nThis study demonstrated the importance of using deep learning for the detection of 357\nCOVID-19 subjects, especially those who are asymptomatic. Furthermore, it elaborated 358\non the significance of biological signals, such as breathing sounds, in acquiring useful 359\ninformation about the viral infection. Unlike the conventional lung auscultation tech- 360\nniques, i.e., electronic stethoscopes, to record breathing sounds, the study proposed 361\nherein utilized breathing sounds recorded via a smartphone microphone. The observa- 362\ntions found in this study (highest accuracy: 94.58%) strongly suggest deep learning as 363\na pre-screening tool for COVID-19 as well as an early detection technique prior to the 364\ngold standard RT-PCR assay. 365\nSmartphone-based breathing recordings 366\nAlthough current lung auscultation techniques provide high accuracy measures in de- 367\ntecting respiratory diseases [44 –46], it requires subjects to be present at hospitals for 368\nequipment setup and testing preparation prior to data acquisition. Furthermore, it 369\nrequires the availability of an experienced person, i.e., clinician or nurse, to take data 370\nfrom patients and store it in a database. Therefore, utilizing a smartphone device 371\nto acquire such data allows for a faster data acquisition process from subjects or pa- 372\ntients while at the same time, provides highly comparable and acceptable diagnostic 373\nperformance. In addition, smartphone-based lung auscultation ensures a better social 374\ndistancing behaviour during lock downs due to pandemics such as COVID-19, thus, it 375\nallows for a rapid and time-efficient detection of diseases despite of strong restrictions. 376\nSeptember 18, 2021 14/20\n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted September 22, 2021. ; https://doi.org/10.1101/2021.09.18.21263775doi: medRxiv preprint \n\nBy visually inspecting COVID-19 and healthy subjects’ breathing recordings (Fig. 2), 377\nan abnormal nature was usually observed by COVID-19 subjects, while healthy subjects 378\nhad a more regular pattern during breathing. This could be related to the hidden 379\ncharacteristics of COVID-19 contaminated within lungs and exhibited during lung 380\ninhale and exhale [47 –49]. Additionally, the MFCC transformation of these recordings 381\n(Fig. 4(a-c)) returned similar observations. By quantitatively evaluating these coefficients 382\nwhen combined, COVID-19 subjects had a unique distribution (positively skewed) that 383\ncan be easily distinguished from the one of healthy subjects. This gives an indication 384\nabout the importance of further extracting the internal attributes carried not only by 385\nthe recordings themselves, but rather by the additional MFC transformation of such 386\nrecordings. Additionally, the asymptomatic subjects had a distribution of values that 387\nwas close in shape to the distribution of healthy subjects (Fig. 4(a)), however, it was 388\nskewed towards the right side of the zero mean. This may be considered as a strong 389\nattribute when analyzing COVID-19 patients who do not exhibit any symptoms and 390\nthus, discriminating them easily from healthy subjects. 391\nDiagnosis of COVID-19 using deep learning 392\nIt is essential to be able to gain the benefit of the recent advances in AI and computerized 393\nalgorithms, especially during these hard times of COVID-19 spread worldwide. Deep 394\nlearning not only provides high levels of performance, it also reduces the dependency 395\non experts, i.e., clinicians and nurses, who are now suffering in handling the pandemic 396\ndue to the huge and rapidly increasing number of infected patients [50 –52]. Recently, 397\nthe detection of COVID-19 using deep learning has reached high levels of accuracy 398\nthrough two-dimensional (2D) lung CT images [53 –55]. Despite of such performance 399\nin discriminating and detecting COVID-19 subjects, CT imaging is considered high 400\nin cost and requires extra time to acquire testing data and results. Furthermore, it 401\nutilizes excessive amount of ionizing radiations (X-ray) that are usually harmful to 402\nthe human body, especially for severely affected lungs. Therefore, the integration of 403\nbiological sounds, as in breathing recordings, within a deep learning framework overcomes 404\nthe aforementioned limitations, while at the same time provides acceptable levels of 405\nperformance. 406\nThe proposed deep learning framework had high levels of accuracy (94.58%) in 407\ndiscriminating between COVID-19 and healthy subjects. The structure of the framework 408\nwas built to ensure a simple architecture, while at the same time to provide advanced 409\nfeatures extraction and learning mechanisms. The combination between hand-crafted 410\nfeatures and deep-activated features allowed for maximized performance capabilities 411\nwithin the model, as it learns through hidden and internal attributes as well as deep 412\nstructural and temporal characteristics of recordings. The high sensitivity and specificity 413\nmeasures (94.21% and 94.96%, respectively) obtained in this study prove the efficiency 414\nof deep learning in distinguishing COVID-19 subjects (AUROC: 0.90). Additionally, it 415\nsupports the field of deep learning research on the use of respiratory signals for COVID-19 416\ndiagnostics [21, 56]. Alongside the high performance levels, it was interesting to observe 417\na 100.00% accuracy in predicting asymptomatic COVID-19 subjects. This could enhance 418\nthe detection of this viral infection at a very early stage and thus, preventing it from 419\ndeveloping to mild and moderate conditions or spreading to other people. 420\nFurthermore, this high performance levels were achieved through 1D signals instead 421\nof 2D images, which allowed the model to be simple and not memory exhausting. In 422\naddition, due to its simplicity and effective performance, it can be easily embedded 423\nwithin smartphone applications and internet-of-things tools to allow real-time and direct 424\nconnectivity between the subject and family for care or healthcare authorities for services. 425\nSeptember 18, 2021 15/20\n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted September 22, 2021. ; https://doi.org/10.1101/2021.09.18.21263775doi: medRxiv preprint \n\nClinical relevance 426\nThe utilization of smartphone-based breathing recordings within a deep learning frame- 427\nwork may have the potential to provide a non-invasive, zero-cost, rapid pre-screening tool 428\nfor COVID-19 in low-infected as well as servery-infected countries. Furthermore, it may 429\nbe useful for countries who are not able of providing the RT-PCR test to everyone due 430\nto healthcare, economic, and political difficulties. Furthermore, instead of performing 431\nRT-PCR tests on daily or weekly basis, the proposed framework allows for easier, cost 432\neffective, and faster large-scale detection, especially for counties/areas who are putting 433\nhigh expenses on such tests due to logistical complications. Alongside the rapid nature of 434\nthis approach, many healthcare service could be revived significantly by decreasing the 435\ndemand on clinicians or nurses. In addition, due to the ability of successfully detecting 436\nasymptomatic subjects, it can decrease the need for extra equipment and costs associated 437\nwith further medication after the development of the viral infection in patients. 438\nClinically, it is better to have a faster connection between COVID-19 subjects and 439\nmedical practitioners or health authorities to ensure continues monitoring for such cases 440\nand at the same time maintain successful contact tracing and social distancing. By 441\nembedding such approach within a smartphone applications or cloud-based networks, 442\nmonitoring subjects, including those who are healthy or suspected to be carrying the 443\nvirus, does not require the presence at clinics or testing points. Instead, it can be 444\nperformed real-time through a direct connectivity with a medical practitioners. In 445\naddition, it can be completely done by the subject himself to self-test his condition prior 446\nto taking further steps towards the RT-PCR assay. Therefore, such approach could set 447\nan early alert to people, especially those who interacted with COVID-19 subjects or are 448\nasymptomatic, to go and further diagnose their case. Considering such mechanism in 449\ndetecting COVID-19 could provide a better and well-organized approach that results in 450\nless demand for clinics and medical tests, and thus, enhances back the healthcare and 451\neconomic sectors in various countries worldwide. 452\nConclusion 453\nThis study suggests smartphone-based breathing sounds as a promising indicator for 454\nCOVID-19 cases. It further recommends the utilization of deep learning as a pre- 455\nscreening tool for such cases prior to the gold standard RT-PCR tests. The overall 456\nperformance found in this study (accuracy 94.58%) in discriminating between COVID-19 457\nand healthy subjects shows the potential of such approach. This study paves the way 458\ntowards implementing deep learning in COVID-19 diagnostics by suggesting it as a rapid, 459\ntime-efficient, and no-cost technique that does not violate social distancing restrictions 460\nduring pandemics such as COVID-19. 461\nAcknowledgement 462\nThis work was supported by a grant (award number: 8474000132) from the Healthcare 463\nEngineering Innovation Center (HEIC) at Khalifa University, Abu Dhabi, UAE, and 464\nby grant (award number: 29934) from the Department of Education and Knowledge 465\n(ADEK), Abu Dhabi, UAE. 466\nReferences\n1. Dong E, Du H, Gardner L. An interactive web-based dashboard to track COVID-19\nin real time. The Lancet infectious diseases. 2020;20(5):533–534.\nSeptember 18, 2021 16/20\n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted September 22, 2021. ; https://doi.org/10.1101/2021.09.18.21263775doi: medRxiv preprint \n\n2. World health organization (WHO). COVID-19 Weekly epidemi-\nological update;. https://www.who.int/publications/m/item/\nweekly-epidemiological-update-on-covid-19---20-april-2021.\n3. Padma T. India’s COVID-vaccine woes-by the numbers. Nature. 2021;.\n4. World health organization (WHO). Report of the WHO-\nChina Joint Mission on Coronavirus Disease 2019 (COVID-\n19);. https://www.who.int/docs/default-source/coronaviruse/\nwho-china-joint-mission-on-covid-19-final-report.pdf.\n5. Paules, Catharine I and Marston, Hilary D and Fauci, Anthony S. Coronavirus\ninfections more than just the common cold. Jama. 2020;323(8):707–708.\n6. Menni C, Valdes AM, Freidin MB, Sudre CH, Nguyen LH, Drew DA, et al. Real-\ntime tracking of self-reported symptoms to predict potential COVID-19. Nature\nmedicine. 2020;26(7):1037–1040.\n7. Liang T, et al. Handbook of COVID-19 prevention and treatment. The First\nAffiliated Hospital, Zhejiang University School of Medicine Compiled According\nto Clinical Experience. 2020;68.\n8. Lim J, Lee J. Current laboratory diagnosis of coronavirus disease 2019. The\nKorean Journal of Internal Medicine. 2020;35(4):741.\n9. Wu D, Gong K, Arru CD, Homayounieh F, Bizzo B, Buch V, et al. Severity\nand Consolidation Quantification of COVID-19 From CT Images Using Deep\nLearning Based on Hybrid Weak Labels. IEEE Journal of Biomedical and Health\nInformatics. 2020;24(12):3529–3538.\n10. Zhang N, Wang L, Deng X, Liang R, Su M, He C, et al. Recent advances in the\ndetection of respiratory virus infection in humans. Journal of medical virology.\n2020;92(4):408–417.\n11. Fang Y, Zhang H, Xie J, Lin M, Ying L, Pang P, et al. Sensitivity of chest CT for\nCOVID-19: comparison to RT-PCR. Radiology. 2020; p. 200432.\n12. Ai T, Yang Z, Hou H, Zhan C, Chen C, Lv W, et al. Correlation of chest CT and\nRT-PCR testing in coronavirus disease 2019 (COVID-19) in China: a report of\n1014 cases. Radiology. 2020; p. 200642.\n13. Rubin GD, Ryerson CJ, Haramati LB, Sverzellati N, Kanne JP, Raoof S, et al.\nThe role of chest imaging in patient management during the COVID-19 pandemic:\na multinational consensus statement from the Fleischner Society. Chest. 2020;.\n14. Brown C, Chauhan J, Grammenos A, Han J, Hasthanasombat A, Spathis D, et al.\nExploring Automatic Diagnosis of COVID-19 from Crowdsourced Respiratory\nSound Data. arXiv preprint arXiv:200605919. 2020;.\n15. Faezipour M, Abuzneid A. Smartphone-Based Self-Testing of COVID-19 Using\nBreathing Sounds. Telemedicine and e-Health. 2020;.\n16. hui Huang Y, jun Meng S, Zhang Y, sheng Wu S, Zhang Y, wei Zhang Y, et al.\nThe respiratory sound features of COVID-19 patients fill gaps between clinical\ndata and screening methods. medRxiv. 2020;.\n17. Wang B, Liu Y, Wang Y, Yin W, Liu T, Liu D, et al. Characteristics of Pulmonary\nauscultation in patients with 2019 novel coronavirus in china. 2020;.\nSeptember 18, 2021 17/20\n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted September 22, 2021. ; https://doi.org/10.1101/2021.09.18.21263775doi: medRxiv preprint \n\n18. Deshpande G, Schuller B. An Overview on Audio, Signal, Speech, & Language\nProcessing for COVID-19. arXiv preprint arXiv:200508579. 2020;.\n19. Quatieri TF, Talkar T, Palmer JS. A framework for biomarkers of covid-19\nbased on coordination of speech-production subsystems. IEEE Open Journal of\nEngineering in Medicine and Biology. 2020;1:203–206.\n20. Noda A, Saraya T, Morita K, Saito M, Shimasaki T, Kurai D, et al. Evidence\nof the Sequential Changes of Lung Sounds in COVID-19 Pneumonia Using a\nNovel Wireless Stethoscope with the Telemedicine System. Internal Medicine.\n2020;59(24):3213–3216.\n21. Laguarta J, Hueto F, Subirana B. COVID-19 Artificial Intelligence Diagnosis\nusing only Cough Recordings. IEEE Open Journal of Engineering in Medicine\nand Biology. 2020;1:275–281.\n22. Han J, Qian K, Song M, Yang Z, Ren Z, Liu S, et al. An Early Study on Intelligent\nAnalysis of Speech under COVID-19: Severity, Sleep Quality, Fatigue, and Anxiety.\narXiv preprint arXiv:200500096. 2020;.\n23. Sharma N, Krishnan P, Kumar R, Ramoji S, Chetupalli SR, Ghosh PK, et al.\nCoswara–A Database of Breathing, Cough, and Voice Sounds for COVID-19\nDiagnosis. arXiv preprint arXiv:200510548. 2020;.\n24. Organization WH, et al. COVID-19 Weekly Epidemiological Update, 25 April\n2021. 2021;.\n25. Indian institute of science. Project Coswara — IISc;. https://coswara.iisc.ac.\nin/team.\n26. Barrett KE, Barman SM, Boitano S, Brooks HL, et al.. Ganong’s review of medical\nphysiology; 2016.\n27. Groeneveld RA, Meeden G. Measuring skewness and kurtosis. Journal of the\nRoyal Statistical Society: Series D (The Statistician). 1984;33(4):391–399.\n28. Richman JS, Moorman JR. Physiological time-series analysis using approximate\nentropy and sample entropy. American Journal of Physiology-Heart and Circulatory\nPhysiology. 2000;.\n29. Shannon CE. A mathematical theory of communication. The Bell system technical\njournal. 1948;27(3):379–423.\n30. Higuchi T. Approach to an irregular time series on the basis of the fractal theory.\nPhysica D: Nonlinear Phenomena. 1988;31(2):277–283.\n31. Katz MJ. Fractals and the analysis of waveforms. Computers in biology and\nmedicine. 1988;18(3):145–156.\n32. Zheng F, Zhang G, Song Z. Comparison of different implementations of MFCC.\nJournal of Computer science and Technology. 2001;16(6):582–589.\n33. Rabiner L, Schafer R. Theory and applications of digital speech processing.\nPrentice Hall Press; 2010.\n34. M Schuster, K Paliwal. Bidirectional recurrent neural networks. IEEE Transactions\non Signal Processing. 1997;45(11):2673–2681. doi:10.1109/78.650093.\nSeptember 18, 2021 18/20\n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted September 22, 2021. ; https://doi.org/10.1101/2021.09.18.21263775doi: medRxiv preprint \n\n35. Schmidhuber, J¨ urgen. Deep learning in neural networks: An overview. Neural\nnetworks. 2015;61:85–117.\n36. S Hochreiter and J Schmidhuber. Long short-term memory. Neural computation.\n1997;9(8):1735–1780.\n37. Campos-Taberner M, Garc´ ıa-Haro FJ, Mart´ ınez B, Izquierdo-Verdiguier E,\nAtzberger C, Camps-Valls G, et al. Understanding deep learning in land use\nclassification based on Sentinel-2 time series. Scientific reports. 2020;10(1):1–12.\n38. Shorten C, Khoshgoftaar TM. A survey on image data augmentation for deep\nlearning. Journal of Big Data. 2019;6(1):60.\n39. Christian B, Griffiths T. Algorithms to live by: The computer science of human\ndecisions. Macmillan; 2016.\n40. Nanni L, Maguolo G, Paci M. Data augmentation approaches for improving\nanimal audio classification. Ecological Informatics. 2020; p. 101084.\n41. Salamon J, Bello JP. Deep convolutional neural networks and data augmen-\ntation for environmental sound classification. IEEE Signal Processing Letters.\n2017;24(3):279–283.\n42. Greenwood PE, Nikulin MS. A guide to chi-squared testing. vol. 280. John Wiley\n& Sons; 1996.\n43. Qian N. On the momentum term in gradient descent learning algorithms. Neural\nnetworks. 1999;12(1):145–151.\n44. Gurung A, Scrafford CG, Tielsch JM, Levine OS, Checkley W. Computerized\nlung sound analysis as diagnostic aid for the detection of abnormal lung sounds:\na systematic review and meta-analysis. Respiratory medicine. 2011;105(9):1396–\n1403.\n45. Shi L, Du K, Zhang C, Ma H, Yan W. Lung Sound Recognition Algorithm Based\non VGGish-BiGRU. IEEE Access. 2019;7:139438–139449.\n46. Shuvo SB, Ali SN, Swapnil SI, Hasan T, Bhuiyan MIH. A lightweight cnn model for\ndetecting respiratory diseases from lung auscultation sounds using emd-cwt-based\nhybrid scalogram. IEEE Journal of Biomedical and Health Informatics. 2020;.\n47. Wang B, Liu Y, Wang Y, Yin W, Liu T, Liu D, et al. Characteristics of Pul-\nmonary auscultation in patients with 2019 novel coronavirus in china. Respiration.\n2020;99(9):755–763.\n48. Huang Y, Meng S, Zhang Y, Wu S, Zhang Y, Zhang Y, et al. The respiratory\nsound features of COVID-19 patients fill gaps between clinical data and screening\nmethods. medRxiv. 2020;.\n49. Noda A, Saraya T, Morita K, Saito M, Shimasaki T, Kurai D, et al. Evidence\nof the Sequential Changes of Lung Sounds in COVID-19 Pneumonia Using a\nNovel Wireless Stethoscope with the Telemedicine System. Internal Medicine.\n2020;59(24):3213–3216.\n50. Mehta S, Machado F, Kwizera A, Papazian L, Moss M, Azoulay ´E, et al. COVID-\n19: a heavy toll on health-care workers. The Lancet Respiratory Medicine.\n2021;9(3):226–228.\nSeptember 18, 2021 19/20\n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted September 22, 2021. ; https://doi.org/10.1101/2021.09.18.21263775doi: medRxiv preprint \n\n51. O’Flynn-Magee K, Hall W, Segaric C, Peart J. GUEST EDITORIAL: The impact\nof Covid-19 on clinical practice hours in pre-licensure registered nurse programs.\nTeaching and Learning in Nursing. 2021;16(1):3.\n52. Anders RL. Engaging nurses in health policy in the era of COVID-19. In: Nursing\nforum. vol. 56. Wiley Online Library; 2021. p. 89–94.\n53. Li C, Dong D, Li L, Gong W, Li X, Bai Y, et al. Classification of Severe and Critical\nCovid-19 Using Deep Learning and Radiomics. IEEE Journal of Biomedical and\nHealth Informatics. 2020;24(12):3585–3594.\n54. Meng L, Dong D, Li L, Niu M, Bai Y, Wang M, et al. A Deep Learning Prognosis\nModel Help Alert for COVID-19 Patients at High-Risk of Death: A Multi-Center\nStudy. IEEE Journal of Biomedical and Health Informatics. 2020;24(12):3576–\n3584.\n55. Jiang Y, Chen H, Loew M, Ko H. COVID-19 CT Image Synthesis with a\nConditional Generative Adversarial Network. IEEE Journal of Biomedical and\nHealth Informatics. 2020;.\n56. Mouawad P, Dubnov T, Dubnov S. Robust Detection of COVID-19 in Cough\nSounds. SN Computer Science. 2021;2(1):1–13.\nSeptember 18, 2021 20/20\n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted September 22, 2021. ; https://doi.org/10.1101/2021.09.18.21263775doi: medRxiv preprint","source_license":"CC-BY-4.0","license_restricted":false}