Development of a Deep Learning Technique to Analyze Air Quality from Social Media Response

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract One of the prominent new-age methods used by today’s population in spreading awareness or drawing attention on an issue or concern is through social media platforms. In this study, the responses of general public to air quality that they share on a popular social media platform – Twitter, was taken as a virtual quantity that helps in measuring and analyzing the prevalent air quality. Machine learning technique based on self attention network was used to sort, clean and classify large amount of air pollution related Twitter responses extracted during 2019-2020 at Delhi in India. The temporal correlation of tweet response volumes with the ground monitored concentrations of air pollution and word cloud analysis were used to analyze the attitude of the public towards urban air quality. These analyses lead to the development of an air quality prediction model using ‘spaCy’ – similarity analysis method which attempts to predict the ground concentration range of pollutant PM2.5 from tweets received on a particular day.
Full text 118,775 characters · extracted from preprint-html · click to expand
Development of a Deep Learning Technique to Analyze Air Quality from Social Media Response | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Development of a Deep Learning Technique to Analyze Air Quality from Social Media Response Thushara Sudheish Kumbalaparambil, Ratish Menon, Vishnu P Radhakrishnan, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-1174813/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 6 You are reading this latest preprint version Abstract One of the prominent new-age methods used by today’s population in spreading awareness or drawing attention on an issue or concern is through social media platforms. In this study, the responses of general public to air quality that they share on a popular social media platform – Twitter, was taken as a virtual quantity that helps in measuring and analyzing the prevalent air quality. Machine learning technique based on self attention network was used to sort, clean and classify large amount of air pollution related Twitter responses extracted during 2019-2020 at Delhi in India. The temporal correlation of tweet response volumes with the ground monitored concentrations of air pollution and word cloud analysis were used to analyze the attitude of the public towards urban air quality. These analyses lead to the development of an air quality prediction model using ‘spaCy’ – similarity analysis method which attempts to predict the ground concentration range of pollutant PM 2.5 from tweets received on a particular day. Twitter air pollution self attention network Delhi PM2.5 Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 1. Introduction Increasing air pollution has become a major concern for the environmental quality of life in the urban areas. Recent studies visualize the grim reality of air pollution and its health risks across the cities of the world (Dandona, 2018; Pant et al., 2019; IQAir, 2019; WHO, 2020). For strategizing any control or eradication program, first step would be to monitor air quality at possibly high temporal and spatial resolutions to identify the source, location, quality and quantity of pollutants. The existing regulatory monitoring networks, especially in developing countries such as India, however make sparse spatial and temporal measurements (CPCB, 2020). Any additional information in real time, even if qualitative, would be an advantage for air quality management initiatives in countries like India which has 6 out of the 10 most polluted cities in the world (IQAir, 2019). With increased access to internet and emergence of social media platforms, people are now expressing their views and relation to the events around them like never before. Social media have grown tremendously in number and popularity ever since its emergence and now include multitude of platforms such as Twitter, Facebook, LinkedIn, YouTube channels, blogs, chat rooms, discussion forums etc. India with its population of ~1.3 billion people has over 560 million internet users and is the second largest online market in the world. The internet accessibility and use in the country largely varied based on factors like gender and socio-economic divide (Statista, 2020b). Analyzing public behavior on such platforms are not new to researchers and have been part of routine business analytics (Fan and Gordon, 2014). Many studies have also explored their potential for monitoring environmental events such as natural disasters in real time (Sakaki et al., 2010; Lindsay, 2011; Earle et al., 2011; Li et al., 2012; Kent and Capello Jr., 2013; Middleton et al., 2013, Singh et al., 2019). Social media responses also have the potential to extract useful information on air pollution (Robinson and Fialkowski, 2010; Jiang et al., 2015; Gurajala and Matthews, 2018; Jiang et al. 2019). It has been seen that the exposed people tend to share pictures, videos, blogs and tweets about air quality events almost immediately on social media within minutes of occurrence of an event (Robinson and Fialkowski, 2010). These user generated contents (UGC) enable us to derive real time information on air pollution. The response of users or ‘human sensors’ may be direct in the form of immediate posts, tweets, blogs etc or indirect like re-sharing an already existing content or supporting it. This rise in popularity of social media has given people access to a humongous volume of information which can be used in a variety of valuable areas. Extracting and analyzing UGC to derive actionable knowledge is challenging task. However, recent approaches using machine learning (ML) techniques have been found promising for the purpose (Jackoway et al., 2011; Gurajala and Matthews, 2018). Such ML techniques are exploited curiously by new researchers are air quality prediction and forecasting as Xu et al. (2020), Chang et al. (2021) and Zhang et al. (2021). Use of attention network in deep machine learning algorithms which is now widely recognised in sequence modeling as shown by Bahdanau et al. (2014) and Vaswani et al. (2017) to maximize performance in models using attention mechanism to connect between the input and output. Long - Short Term Memory (LSTM) are known to be effective for time series prediction tasks and multiple LSTM units are stacked to form the more effective Bi-directional Long - Short Term Memory (Bi-LSTM) models(Zhang et al., 2021). Wiedemann et al. (2018) and Wu et al. (2019) has also presented inspiring Bi-LSTM model architectures. Table 1 shows a tabulation of some of the previous works that attempted to predict air quality using different ML techniques. It is seen that the influence of social media for air quality prediction is lesser explored than using air quality parameters and other polluting factors. Also self attention networks have not been prominently used for air quality prediction using social sensors. Table1 Comparison between various reference papers that performed air quality prediction Reference papers Monitored air quality parameter Source/ Social media platform used for prediction Country or city under study Prediction Models/Techniques used, efficiency obtained Jiang et al. (2015) AQI Sina Weibo(Chinese Twitter) Beijing Gradient tree boosting (GTB) - 59% Gurajala and Matthews (2018) PM 2.5 Twitter Paris, Delhi, London Air quality not predicted Jiang et al. (2019) AQI Twitter California, Idaho, Illinois, Indiana, Ohio Natural Language Processing (NLP) - 6.9–17.7% improvement with social media intervention over base line method. Xu et al. (2020) PM 2.5 Historical meteorological data, road network data, administrative boundary vector data, POI data. Beijing, Tianjin, Hebei Temporal-spatial-regression-tree model, Grid prediction model - 90% Chang et al. (2021) PM 2.5 Local and neighboring station data, chimney and abroad pollution data Taiwan Aggregate LSTM - better than GTB, LSTM, SVR* Zhang et al. (2021). PM 2.5 PM 2.5 - Hourly, daily, restructured multi hour data Beijing LSTM, Bi-LSTM, EMD**- BiLSTM, - More than 95% *SVR - Support vector regression **EMD - Empirical mode decomposition In the present work, Twitter responses related to air quality in Delhi were analyzed during 2019-2020 to develop a deep learning model that can predict the concentration of PM 2.5 (mass concentration of particulates having an aerodynamic diameter less than 2.5µm) from the tweets. A self attention network based classifier was used to characterize the tweets. The PM 2.5 data from continuous ambient air quality monitoring stations (CAAQMS) within the city were used to train the deep learning algorithm. The aim of the study is to: analyze the fit and relevance of the physically monitored air quality parameter with the tweeting volume and behavior of people in an urban city of a developing country like Delhi, India and thus; evaluate the ability to predict pollution concentrations from this virtual medium; develop a new self attention deep learning classification model with high accuracy to classify tweets to those indicating poor air quality (Class I), good air quality (Class II) and neutral or noise tweets (Class 0) with minimal human intervention. develop a machine learning model to predict the PM 2.5 concentration by analysing a tweet. The further organization of the paper is as follows: methodologies and machine learning techniques followed for the accomplishment of the objectives are elaborated in Section 2 . The results and discussions of the temporal and behavioral analysis of the study are discussed in Section 3 along with the efficiency on the classification and prediction models. Finally Section 4 wraps up the paper with the conclusions drawn, challenges faced by the study and the future scope of work. 2. Methodology This section discusses in detail the extraction of the social media responses from the chosen area under study during 2019 - 2020 and the various analyses performed on them along with the descriptions of the machine learning models used for the classification of the extracted data and the prediction of PM 2.5 concentration using the classified data. 2.1 Selection of study area and the social media platform Twitter was chosen as the social media platform for the present study. Tweet response volumes, defined as the sum of direct and indirect response of people to air quality, for the period of 12 months ranging from March 2019 to February 2020 was used as the data for the analysis. Indirect responses included likes and re-tweets for a particular tweet. There have been works like that of Jiang et al. (2015) who have considered the effects of direct and indirect responses separately but we chose to consider them as a collective response of the people of Delhi to their air quality issues. Twitter is peculiar for its short and crisp contents limited over 280 characters and its prominence at the time of an event. Twitter has also become an increasingly important tool in politics, mass media communications, and many more such that it touches almost every sector of life. Many world leaders, governments, ministries, influencers, institutions, news channels have their official Twitter accounts to make announcements, influence and engage with the general public. In India there were 13.15 million Twitter users as of April 2020 (Statista, 2020a). Delhi, capital city of India, was selected as a representative urban area to demonstrate the methodology presented in this paper. A preliminary analysis of tweets related to air pollution from the 100 most populated cities in India as part of this study identified Delhi as the city with maximum number of air quality episodes (IQAir, 2019) and maximum number of tweets related to air quality. Also, a public health emergency was declared at Delhi in November 2019 as the air quality index exceeded more than 3 times the ‘hazardous’ level (CNN, 2020). 2.2 Data extraction Tweets were extracted from Twitter by attaining access to Twitter streaming API. The tweets were fetched with the help of keywords pertaining to air pollution along with the name of the city. About 15 different keyword combinations with Delhi (as shown in Table 2) were used for tweet extraction for the study period. The keywords thus chosen were different combinations of ‘air quality’, ‘air pollution’ and ‘smog’. The data set collected included attributes such as the day, date and time of the tweet, the tweet text and the number of likes and retweets received by that particular tweet. In this manner, a data set of around 82,000 tweets were created. Tweets only in the English language were considered in this study as Twitter users in India are pre-dominantly English speakers (Poell and Rajagopalan, 2015). In order to reduce the noise in the data set, few words were used to filter out unnecessary tweets like those generated by automated websites on a periodic basis regardless of the change in pollution intensity. Inorder to relate Twitter activity with ambient air quality information, PM 2.5 data from 38 CAAQMS stations within Delhi during the period of March 2019 to February 2020 were collected from the CPCB data repository (CCR-CPCB, 2020). The datasets generated and analysed during the current study are available from the corresponding author on reasonable request. Table 2 Keyword combinations used to extract tweets during study period Keyword Combinations Air Quality + Delhi Air Pollution + Delhi DelhiChokes AirQuality + Delhi AirPollution + Delhi Choke + Delhi AirQualityDelhi AirPollutionDelhi Clean air + Delhi DelhiAirQuality DelhiAirPollution DelhiEmergency Delhi + Smog DelhiSmog Delhi + RightToBreathe 2.3 Data classification Each tweet, within the allowed 280 characters contained words indicating good or bad air quality. Extracted dataset also contained some tweets that are either irrelevant or neutral with respect to the air quality. Segregating these tweets into various classes was an important step in the present study. The tweet data set was manually read and sorted by a team of 15 environmental engineers and the data set was sorted into 3 classes to segregate them into: poor air quality tweets (Class I), good air quality tweets (Class II) and noise or neutral tweets (Class 0). The sorted data set was then checked for variations in sorting and resorted (if needed) by an expert team of 3 environmental engineers. The sorted set consisted of tweets in Class I and II which contributed as indicators of the ambient air quality whereas Class 0 consisted of tweets generated by bots, irrelevant tweets, etc. Examples of tweets in their respective classes is shown in Table 3. The classified set was then pre processed and used to train a self attention model that would automatically classify a sample tweet into its suitable classes. Table 3 Examples of classified tweets S. No. Tweets Classification 1. New Delhi: Over 1.2 million people died in India due to air pollution in 2017, said a global report on air pollution Class 0 2. #DelhiAirEmergency #DelhiPollution #DelhiBachao #DelhiAirQuality #DelhiNCRPollution Class 0 3. Kab tak zindagi katoga bd or cigar mein, kuch din to gujaro delhi, in ncr Class 0 4. Just landed in #Delhi, the air here is just unbreathable Class 1 5. Amazing Air Quality today in Delhi! Enjoy the blue sky and clean air while it lasts. Class 2 2.4 Data analysis The following analyses were performed on the extracted and classified data set. 2.4.1 Temporal analysis of tweet response volume in relation to PM 2.5 concentrations The volume of tweet responses related to air quality was compared with the ambient concentrations of PM 2.5 so as to analyze similarity in their trend. The comparison was done by graphical and statistical methods. 2.4.1.1 Graphical analysis Semi log time series graphs for the tweet response volumes for each of Class I and Class II tweets were plotted against the ground average PM 2.5 concentrations obtained from CPCB for the duration of the study. Twitter activity was plotted on a logarithmic scale considering their vast data range. The coherence of peaks and falls in the graph were analyzed. The graphical representation and the inferences drawn from it are discussed in Section 3.1. 2.4.1.2 Statistical analysis In-order to understand the statistical relationship between the two data sets, box plots were made with PM 2.5 on x-axis and tweet response volume on y-axis. The data was sorted on the basis of the classification of PM 2.5 values into 6 categories as defined by the CPCB namely Good( 0 - 30µg/m 3 ), Satisfactory( 31 - 60µg/m 3 ), Moderate( 61 - 90µg/m 3 ), Poor( 91 - 120µg/m 3 ), Very Poor( 121 - 250µg/m 3 ) and Severe ( above 250µg/m 3 ) for plotting the box plots. The relationship of tweet response volumes of each class (Class I and II) in these categories were analyzed and the results are given in section 3.2. 2.4.2 Behavioral analysis using word clouds Word clouds were generated corresponding to each season to understand how people express the changes in the air quality through words using a python program utilizing necessary supporting libraries. India has 4 meteorological seasons namely Winter (January, February), Pre Monsoon (March – May), Southwest Monsoon (June – September) and Post Monsoon (October – December) as per Indian meteorological Department (IMD). A prominent transition in the air quality levels in Delhi was observed during post monsoon and winter seasons having the worst air quality which improved over pre monsoon and southwest monsoon seasons. 2.5 Machine Learning Techniques In this study, machine learning techniques were used for (i) automated classification of the tweet data set and (ii) for the prediction of PM 2.5 concentration range from tweet response. The program codes constructed for the classification, training and prediction of the data in this work is available from the corresponding author on reasonable request. 2.5.1 Automated data classification model The tweets extracted over a year along with the likes and retweets appended to them comprised the dataset for the study. The dataset was a highly imbalanced multiclass data set with around 62,000 tweets in Class 0, 18000 tweets in Class I and 1700 tweets in Class II. As initial step, the manually classified dataset was tokenized into words. Preprocessing was performed in order to remove special symbols, stop words, punctuations, Twitter handles, etc. Natural Language Toolkit(NLTK) library was used for the same. The data set was further groomed by performing processes like word indexing, integer encoding and assigning them with pad descriptions before using the data for supervised learning. The respective classes of the tweets were label encoded. The processed data set was split into a training set and test set as 90% and 10% respectively for a 10 fold cross validation analysis to evaluate the overall performance of the model and a test-train split of 80%-20% for result analysis. The present study uses a multilayer classification model. The first layer is an embedding layer in which an embedding matrix is constructed from the data set using pertained global vector (GloVe) embedding for creating weights for the embedding layer (Pennington et al., 2014). The second layer is a BiLSTM layer (Units: 100) where the input sequences are analyzed in the forward and backward direction; that is, if a tweet has 'n' number of words (w 1 , w 2 ,...w n ) then the tweet is first analyzed to from w 1 to w n by forward LSTM and w n to w 1 by backward LSTM and forms a word feature (Wf) for each word on the basis of both the forward analysis (f F ) and backward analysis (f B ) such that Wf = f F ‖ f B where ‖ stands for concatenation function. Similarly a word feature is formed for every word of the data file. After this the data flows into a self attention layer (SeqSelfAttention) layer where words are assigned weights and added to the resultant word features of the previous layer based on the relative importance of that word to the entire tweet. This layer improves the efficiency of the classification as 'attention' is given to words based on their importance. For further detailed reference on Bi-LSTM and attention neural networks Bahdanau et al, 2014; Vaswani et al., 2017; Wiedemann et al, 2018; Zhang et al., 2018; Wu et al, 2019; Xu et al., 2020; Chang et al., 2021 and Zhang et al., 2021 may be used. Python libraries like Keras (https://keras.io/) and Keras_self_attention (https://pypi.org/project/keras-self-attention/) was used to support this layer. The context vectors or sentence feature vector that are obtained as the output of the attention layer using weighted sum function is then given to two dense layers or the fully connected layers where all the inputs and outputs are connected to all the neurons in each layer. The both dense layer contains 50 units and is activated with Rectified Linear Unit (ReLu) layer activation function. It is then flattened in the next layer, the flatten layer. The final output is obtained in a dense layer which has units equal to the number of output classes and is then given output activation by softmax function which provides the probabilities of the potential outcomes. Other parameters used in this modelling architecture includes: Epochs=5, Batch size:128, Optimizer: Adam. The model architecture is shown in Fig. 1. 2.5.2 PM 2.5 prediction model The prediction model used similarity analysis to find the most similar tweets to a test tweet to predict the probable ground PM 2.5 concentration that would have been recorded on that day. The flow chart of the model is shown in Fig. 2. This model used ‘spaCy’ library available in python in order to find the top 10% of the most similar tweets of the test tweet from the season wise training set. A test tweet first undergoes model classification as mentioned in the previous section and then from the date of the tweet identify respective season’s training set. The test tweet was then analyzed for similarity with each tweet of its corresponding training set and a similarity index was assigned to each training tweet. The training set tweets are then arranged in descending of the cosine similarity index and the first 10% of the training set in this sorted set was taken as the resultant ‘most similar tweets set’. The arithmetic mean of PM 2.5 concentrations corresponding to that similarity set was then reported as the output result of the model. The usefulness of the model was analyzed by measuring the correctly predicted concentrations as the percentage of the total predictions under each category specified by CPCB (CPCB, 2014). This was also represented as a histogram plot to understand the social media behavior with respect to change in air quality. 3. Results And Discussions The data classification model sorted the tweets into their corresponding Class (Class 0, I or II) with an overall accuracy of 96.7% from the 10 fold cross validation testing and an accuracy of 87.4% for the model using test-train split of 80% - 20%. On the classified data further analyses were done, the results of which are discussed in the following sections. 3.1 Temporal variation of tweet responses with PM 2.5 concentrations: Graphical analysis A time series analysis for tweet responses of Class I and II along with corresponding PM 2.5 concentration was conducted for the duration of study (Fig. 3). It was observed that there is rise in peak of the tweet response volume along with rise in pollution in the air and also a corresponding dip in tweet response volume with a fall in air pollution. It was noticed that there is more response in Class I than Class II. The larger response volume in Class I might be due to the alertness of the public to poor quality of air that brings distress to their daily lives and their eagerness to spread the information to others, draw attention of government or other related or powerful individuals in the society, apart from the fact that the quality of air remained poor for a major part of the study period. The low volume in Class II is primarily due to lesser number of good air quality days in Delhi and also might be because good air quality days are appreciated mostly after a spike in pollution and not much while quality of air has been staying good for a period of time. During pre-monsoon season (March – May 2019), Delhi had its worst air quality towards the middle of May and that was the time when the highest tweet response (around 200) in Class I was recorded. The day with best air quality of the season reported no Class I tweets. High volume in Class II tweet response volume was found during the early days of March when the air quality used to improve to moderate category after a poor air quality period. Most of the times when the PM 2.5 concentrations fell within the desirable limit there was some Class II tweet response activity. A sudden fall in pollutant level resulted in more tweet responses than a gradual fall. During the period June – September 2019 (southwest monsoon season), the average PM 2.5 concentration levels in Delhi were mostly within the desirable limit. Therefore, Class I tweet response activity was comparatively lesser (below 100) compared to other seasons. There were more number of days with minimal or no Class I tweet response volume. Whereas in the case of Class II, the maximum tweet response volumes were recorded on the day with best air quality of the year. Even though the quality of air in this season was mostly good, there wasn’t a tweet response for each of those days. Class II responses were mostly observed on good air quality days after an increase in pollution or if the pollutant concentrations were well below the desirable limits. In the post monsoon season the average PM 2.5 concentration in Delhi was seldom within the desirable limits that is below 60μg/m 3 as specified in National Ambient Air Quality Standards by CPCB. This was the infamous times when the stubble burning in northern parts of India highly polluted the air in the northern states and public health emergency was declared in Delhi (CNN, 2020). This season also coincides with the Indian festival Diwali which is mostly celebrated with fireworks, crackers, etc. Hence, for almost every day of this season, there was Class I tweet response. Class I tweet response volume was varied from 100 units and ranges as high as above 50000 on the most polluted day of the year when average PM 2.5 concentration over Delhi was around 550μg/m 3 . The time series of the tweet response volume of this season resembles much with its PM 2.5 concentrations. There was a high response in Class II as well. This was due to the fact that even a small fall in pollutant concentration was a relief to the citizen that they readily responded, increase the volume in Class II. In the winter season, Class I response was observed almost every day of the season owing to the fact that the average PM 2.5 concentration in Delhi during this season was seldom within the desirable limits. As the air pollution was not bad as in the previous season, the tweet response volume range has reduced. The low volume of response on 1st January 2020 in contrary to the high ground PM 2.5 concentration on the same day might have been a result of deviation of attention of public from air pollution to New Year celebrations and activities. As the air quality in this season was generally poor, the volume of tweet responses in Class II was less. However, Class II tweets were prominent with sharp fall in pollution as seen in the previous cases. 3.2 Statistical relationship of tweet response volume with PM 2.5 To analyze the statistical relationship between the tweet response volumes and the average ground PM 2.5 concentrations, box plots were made for each class as shown in Fig. 4. For this analysis the data was classified according to the different PM 2.5 concentration categories as specified by CPCB. It was seen from Fig. 4(a) that the tweet response volumes for Class I were the largest when the PM 2.5 concentrations were in the poor - severe category and the lowest when it was in the good – moderate category. This proves that the people were more responsive on Twitter and posted tweets indicating poor air quality during the times when the quality of air was in the poor – severe conditions and the Class I tweet response volumes were proportional to the severity of air pollution. The small volume of Class I tweets even on days with good – moderate quality of air could be because the air quality of a day was poorer than the previous day though the concentrations were not above the desired limits. It could also be a response of people visiting the city for the first time and finding the air quality to be bad compared to the place from which they arrived. From the present dataset, it was difficult to verify the behavior of the floating population. Class II tweets were made not only based on present day air quality, but also when there was a decrease in pollution levels from previous days. Therefore, for the statistical analysis, the change in PM 2.5 concentration every 3 days (∆ 3 ) was calculated and the cumulative Class II tweet response volume during that period was considered and plotted to a box plot. From Fig. 4b, it was observed that days with good - satisfactory PM 2.5 concentrations had smaller ∆ 3 and therefore had lesser tweet response volumes. During the very poor – severe air pollution days, the tweets response volumes were more in Class I category than Class II. The pollution was too high to have anything positive to tweet about. It was only while the air quality improved from severe to moderate the people were more responsive to tweet in a positive manner which explains the larger tweet response volume observed under moderate-poor category. 3.3 Behavioral analysis from word clouds Word clouds were generated from the model classified & pre processed data segregated season wise in order to study the responsive nature of public about air quality over various situations and seasons. Word clouds helps in identifying the most frequent words in a season, the nature of emotions expressed through these words & the focus of people. The word cloud for each season of the study period is shown in Fig. 5. The word clouds for the different season in the study period collectively suggest that the citizens of Delhi were much interested in following the air quality index(AQI) of the city regularly. People seems to pay continuous attention to the news updates which are a major source of the AQI indices apart from dedicated apps, Twitter handles, etc and like and retweet tweets put up by news channel in regard with the varying air quality conditions. During poor air quality seasons, the tweeting pattern seem to be dependent on how distressed they feel about air pollution along with concern for health and calling for help and attention from the government and related departments. During good air quality season, people were found to tweet about the content and happiness in having a clear sky and asking the government to maintain the good conditions. The reason for an improvement or deterioration of the air quality is also of interest to the people. From the word cloud for pre monsoon season, it was evident that the air quality was mostly in the ‘moderate’ to ‘severe’ zone. The social media responses have reported when the air quality turned ‘poor’ or when it is ‘deteriorate’-ing or when it has shown an ‘improve’-ment. This was also the summer season in Delhi and there was tweets correlating air quality with ‘temperature’ and ‘dust’ and how it was an ‘unhealthy’ condition and difficult to ‘breathe’. In the monsoon season, tweets were mostly about ‘improve’-d air quality and tweets about ‘rain’, ‘wind’, and ‘storm’ with an ‘AQI’ mostly in the ‘good’, ‘satisfactory’ or ‘moderate’ ranges and asking the ‘government’ to maintain the ‘clean’ ‘sky’. The post monsoon season had large number of tweets in comparison to the other seasons and the wordcloud had the maximum words. The public mostly spoke about how the condition was ‘severe’, ‘poor’, ‘toxic’, ‘hazardous’, and an ‘emergency’ situation which ‘choked’ and affected the daily life of the ‘people’ and led to ‘school’ to shut and ‘flights’ delayed or redirected. They reported about the effect of ‘diwali’ too in turning Delhi into a ‘gas chamber’. Even the slightest ‘improve’-ment brought a ‘relief’ causing people to tweet about it. In the winter season, tweets were manly about ‘poor’ quality of air still ‘remain’-ing in the city with its share of ‘improve’-ments. But the condition was mostly ‘toxic’ and ‘bad’ as ‘hell’. The ‘fog’ contributed much to the ‘smog’. 3.4 Prediction model results Machine learning based prediction model that can estimate PM 2.5 concentrations from the content of tweets was developed in the present study. Fig. 6 shows the percentage accuracy of these predictions for each PM 2.5 concentration category (CPCB, 2014). Prediction accuracy was found to be high for extreme conditions of air quality such as good air quality (80%) and severe air quality (99%) categories. This was because the public could clearly experience these conditions without ambiguity, who then reflected these experiences with appropriate words in their tweet and thus making it easily predictable with our prediction model. As the PM 2.5 concentration moved into other categories, mixed tweeting behavior was observed thereby reducing the prediction accuracy for the model. Also, for a city like Delhi, the citizens were used to moderate to poor air quality conditions for the major part of the year and thus their reactivity to moderate air pollution situations were not so prominent. It was the deviations from the moderate conditions that made the people more alert and produced larger specific tweets suitable for predicting accurately. 4. Conclusion Present study explored the social media behavior of urban dwellers towards air pollution and developed a deep learning based model to predict the concentration of PM 2.5 in the city based on the air pollution related social media responses. The population of Delhi expressed their agony and disappointment on urban air quality through various modes and present study analyzed their response to air quality through popular social media forum - Twitter. The Twitter generated data was sorted before analysis using a well efficient self attention network. The relationship of the Twitter responses to actual pollution level was inferred from temporal and statistical analysis of average PM 2.5 concentrations with tweet response volumes received per day over various season of a year and categories of PM 2.5 concentration ranges. The emotions and attitudes that the Delhi public hold towards air quality are depicted through the words they use in a tweet was analyzed using word clouds which showed that the Delhi public closely follows the AQI in Delhi through news streams or dedicated apps or through social media. It is evident from these analyses that social media is a powerful tool to monitor air quality variations in an urban city like Delhi. The PM concentration prediction results showed that there was high prediction efficiency (above 80%) for extreme air quality conditions and lower prediction efficiency for moderate air quality conditions. It was observed that limited people in India geo tag their tweets and therefore it was difficult to find location of specific tweets unless they mention the location in the tweet. The present study has only considered the text in tweets for analysis here. There were Twitter responses in other forms such as pictures, re-directing internet links, etc. which were not considered in this study. This may be overcome by developing a system to identify multiple languages, image interpretations, etc. Also tweets in regional languages were also overlooked in this study and could further develop the models to include more vernacular languages of the area and mixed scripts using cross lingual knowledge transfer and analysis. The nature of response of people on Twitter in terms of likes or retweets may be influenced by tweets of influential persons, popular organisations, bots, etc. This influence maybe taken into attention and the current study could be further advanced with the analysis of account origins of the tweets. It was seen during the pilot study of this study that many India cities face air quality issues but most of them do not have an active Twitter response culture. Metro cities showed relatively better responses and Delhi with its severe air quality issues had high response volume. Though the study focused on the social media behavior of people in Delhi and depends on several factors such as access to social media, environmental awareness among public, range of changes in air quality experienced by public etc., the methodology is replicable to any urban area in the world. Twitter was also not very popular in India as compared to other social media platforms like Facebook, YouTube, Instagram, etc. Such other social media platforms may also be explored in future to analyze if they hold higher relevance to real time air pollution related responses monitoring in India than Twitter for similar studies. In case of prediction of pollutant concentrations, similarity analysis might be insufficient for peculiar case of air quality conditions unfamiliar to the current training data set such as the unduly fall in air pollution levels in India during Covid-19 pandemic lockdown. Declarations Data Availability PM 2.5 data is available from CPCB at https://app.cpcbccr.com/ccr/#/caaqm-dashboard-all/caaqm-landing. Twitter data can be extracted using Twitter streaming API [https://developer.twitter.com/en/docs/twitter-api]. Data are also available from the corresponding author on reasonable request. Funding No funding was received to assist with the preparation of this manuscript. Ethical Approval Not applicable. Consent to Participate Not applicable Consent to Publish Not applicable Decleration of Competing Interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. CRediT authorship contribution statement Thushara Sudheish Kumbalaparambil : Formal analysis, Writing -original draft. Ratish Menon : Supervision, Writing- original draft, Writing- review & editing. Vishnu P Radhakrishnan :Programming, Data curation. Vinod P Nair : Formal analysis, Writing-review & editing References Bahdanau D, Cho K, Bengio Y (2014). Neural machine translations by jointly learning to align and translate. ICLR 2015, arXiv:1409.0473 CCR-CPCB (2020). https://app.cpcbccr.com/ccr/#/caaqm-dashboard-all/caaqm-landing , last accessed on 01-12-2021 Chang Y, Chiao H, Abimannan S, Huang Y, Tsai Y, and Lin K (2020). An LSTM-based aggregated model for air pollution forecasting. Atmospheric Pollution Research, Volume 11, Issue 8, August 2020, Pages 1451-1463 CNN (2020). https://edition.cnn.com/2020/02/25/health/most-polluted-cities-india-pakistan-intl-hnk/index.html , last accessed on 01-12-2021 CPCB (Central Pollution Control Board), 2014. National Air Quality Index. Control of Urban Pollution Series, CUPS/82/2014-15 CPCB (2020). https://cpcb.nic.in/monitoring-network-3/ , last accessed on 01-12-2021 Dandona L (2018). The impact of air pollution on deaths, disease burden, and life expectancy across the states of India: The Global Burden of Disease Study 2017. Lancet Planet Health; 3: e26–39 Earle P S, Bowden D C, and Guy M (2011). Twitter earthquake detection: Earthquake monitoring in a social world. Annals of Geophysics, 54(6), 708–715 Fan W and Gordon M (2014). The Power of Social Media Analytics. Communications of the ACM. 57. 74-81. 10.1145/2602574 Gurajala S and Matthews J N (2018). Twitter data analysis to understand societal response to air quality. Proceedings of the 9th International Conference on Social Media and Society ,July 2018 ,Pages 82–90, https://doi.org/10.1145/3217804.3217900 IQAir (2019). World Air Quality Report - Region & City PM2.5 Ranking. https://www.iqair.com/world-most-polluted-cities/world-air-quality-report-2019-en.pdf , last accessed on 01-12-2021 Jackoway A, Samet H and Sankaranarayanan J (2011). Identification of Live News Events using Twitter. Proceedings of the 3rd ACM SIGSPATIAL International Workshop on Location-Based Social Networks, Pages 25-32 Jiang J, Sun X, Wang W and Young S (2019). Enhancing Air Quality Prediction with Social Media and Natural Language Processing. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy, July 28 - August 2, 2019, Pages 2627–2632 Jiang W, Wang Y, Tsou M-H, and Fu X (2015). Using Social Media to Detect Outdoor Air Pollution and Monitor Air Quality Index (AQI): A Geo-Targeted Spatiotemporal Analysis Framework with Sina Weibo (Chinese Twitter). PLoS ONE 10(10): e0141185. Kent J D and Capello Jr. H T (2013). Spatial patterns and demographic indicators of effective social media content during the Horsethief Canyon fire of 2012. Cartography and Geographic Information Science, 40:2,78-89 Li R, Lei K H, Khadiwala R and Chang K C (2012). TEDAS: A Twitter-based Event Detection and Analysis System. IEEE 28th International Conference on Data Engineering, 1084-4627/12 Lindsay B R (2011). Social Media and Disasters: Current Uses, Future Options, and Policy Considerations. Congressional Research Service, Washington, DC, 7-5700 Middleton S E, Middleton L and Modafferi S (2013). Real-Time Crisis Mapping of Natural Disasters Using Social Media. IEEE Intelligent Systems 29 (2), 9-17 Pant P, Lal R M, Guttikunda, S K, Russell AG, Nagpure A S, Ramaswami A and Peltier R E (2019). Monitoring particulate matter in India: recent trends and future outlook. Air Quality, Atmosphere & Health.12, 45-58 Pennington J, Socher R and Manning C D (2014). GloVe: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), 1532–1543 Poell, T and Rajagopalan S (2015). Connecting Activists and Journalists: Twitter communication in the aftermath of the 2012 Delhi rape. Journalism Studies, 16:5, 719-733 Robinson E M, and Fialkowski W E (2010). Air Twitter: Using Social Media And Scientific Data To Sense Air Quality Events. 2010 IEEE International Geoscience and Remote Sensing Symposium, July 25-30, 2010, Honolulu, Hawaii, USA Sakaki T, Okazaki M and Matsuo Y (2010). Earthquake shakes Twitter users: real-time event detection by social sensors. Proceedings of the 19th international conference on World Wide Web, Pages 851–860 Singh J P, Dwivedi Y K, Rana N P, Kumar A and Kapoor K K (2019). Event classification and location prediction from tweets during disasters. Annals of Operations Research, 283,737-757 Statista (2020a). https://www.statista.com/statistics/242606/number-of-active-twitter-users-in-selected-countries/ , last accessed on 01-12-2021 Statista (2020b). https://www.statista.com/statistics/255146/number-of-internet-users-in-india/ , last accessed on 01-12-2021 Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez A N, Kaiser L and Polosukhin I (2017). Attention is all you need. 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA. WHO (2020). https://www.who.int/health-topics/air-pollution#tab=tab_1 , last accessed on 01-12-2021. Wiedemann G, Ruppert E, Jindal R, and Biemann C (2018). Transfer Learning from LDA to BiLSTM-CNN for Offensive Language Detection in Twitter. Proceedings of GermEval 2018, 14th Conference on Natural Language Processing (KONVENS 2018). arXiv:1811.02906 Wu G, Tang G, Wang Z, Zhang Z, and Wang Z (2019). An Attention-Based BiLSTM-CRF Model for Chinese Clinic Named Entity Recognition. Special Section On Data-Enabled Intelligence For Digital Health, Vol. 7, pages 113942-113949 Xu C, Tong T, Zhang W and Meng M (2020). Fine-grained prediction of PM2.5 concentration based on multisource data and deep learning. Atmospheric Pollution Research, Volume 11, Issue 10, October 2020, Pages 1728-1737 Zhang L, Liu P, Zhao L, Wang G, Zhang W and Liu J (2021). Air quality predictions with a semi-supervised bidirectional LSTM neural network. Atmospheric Pollution Research, Volume 12, Issue 1, January 2021, Pages 328-339 Zhang Y, Wang J and Zhang X (2018). YNU-HPCC at SemEval-2018 Task 1: BiLSTM with Attention Based Sentiment Analysis for Affect in Tweets. Proceedings of the 12th International Workshop on Semantic Evaluation (SemEval-2018), Pages 273–278 Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Major Revision 22 Mar, 2022 Reviews received at journal 07 Feb, 2022 Reviewers invited by journal 07 Feb, 2022 Editor invited by journal 20 Jan, 2022 Editor assigned by journal 12 Jan, 2022 First submitted to journal 15 Dec, 2021 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-1174813","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":81986204,"identity":"1b802d6d-e84b-429c-8e25-19b81ba1e0ab","order_by":0,"name":"Thushara Sudheish Kumbalaparambil","email":"","orcid":"","institution":"SCMS School of Engineering and Technology","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Thushara","middleName":"Sudheish","lastName":"Kumbalaparambil","suffix":""},{"id":81986205,"identity":"1ce12f55-a1c7-40b2-af62-c02940d0f88b","order_by":1,"name":"Ratish Menon","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABBUlEQVRIiWNgGAWjYJCCA4wNYJrxAZEamOFamA2I1sIA1cImQZQG8/bzBw/+3LFNzpy991jlj5pt9vLRh49JFzDY5Ms7YNcicyaZ4TDvmdvGlj3n0m7zHLuduPFcWpr0DIY0y40HsGuRYABqYWy7nbjhRo7ZbcaG2wmGPTxm0jwMhw0MG3Bo4X/McPBn2+16kJbCnw237QlrkUhmOMDbdjvBAKiFgbfhNuN8HqgWeRzel5B4bHAYqMVww5kzxtIgv2zgYUu2nmGQZoAryCX4Ex9/BDpM3uB4j+HHHzW37eV7mA/eLqiwMZDH4TBMYHAAFFcGEAZxAGQ4M4wxCkbBKBgFowAIABUuXBI5eaRPAAAAAElFTkSuQmCC","orcid":"https://orcid.org/0000-0002-8391-1672","institution":"SCMS School of Engineering and Technology","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Ratish","middleName":"","lastName":"Menon","suffix":""},{"id":81986206,"identity":"20cb5d57-b954-4f84-b5ae-855bf4352290","order_by":2,"name":"Vishnu P Radhakrishnan","email":"","orcid":"","institution":"SCMS School of Engineering and Technology","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Vishnu","middleName":"P","lastName":"Radhakrishnan","suffix":""},{"id":81986207,"identity":"1db9f6d2-1407-4b1f-bae0-8ac2d4691569","order_by":3,"name":"Vinod P Nair","email":"","orcid":"","institution":"Cochin University of Science and Technology","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Vinod","middleName":"P","lastName":"Nair","suffix":""}],"badges":[],"createdAt":"2021-12-15 17:12:23","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-1174813/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-1174813/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":18044722,"identity":"ef274747-8899-4bbd-bec9-770011f7ac7f","added_by":"auto","created_at":"2022-02-08 20:57:31","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":82252,"visible":true,"origin":"","legend":"\u003cp\u003eClassification model architecture\u003c/p\u003e\u003cp\u003e\u003cbr\u003e\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-1174813/v1/09a0c48eef195aa9f6cda475.jpeg"},{"id":18044915,"identity":"b195bfe2-b957-4652-9e43-5515da176334","added_by":"auto","created_at":"2022-02-08 21:00:32","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":10398,"visible":true,"origin":"","legend":"\u003cp\u003eFlowchart of the PM\u003csub\u003e2.5 \u003c/sub\u003econcentration prediction model\u003c/p\u003e","description":"","filename":"Fig02.png","url":"https://assets-eu.researchsquare.com/files/rs-1174813/v1/60eb995f074b281cd41b1d19.png"},{"id":18044480,"identity":"9a30a1fd-045a-4b9e-ae6d-8e7b01ff55fa","added_by":"auto","created_at":"2022-02-08 20:54:31","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":83066,"visible":true,"origin":"","legend":"\u003cp\u003e(a) Time series of Class I tweet response and (b) Time series of Class II tweet response\u0026nbsp;\u003c/p\u003e","description":"","filename":"Fig03.png","url":"https://assets-eu.researchsquare.com/files/rs-1174813/v1/c1993cac0f68eef7c9fa408c.png"},{"id":18044477,"identity":"30165b89-bcce-44aa-b385-94ace92aa665","added_by":"auto","created_at":"2022-02-08 20:54:31","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":40373,"visible":true,"origin":"","legend":"\u003cp\u003e\tStatistical variations of the tweet volume responses with PM\u003csub\u003e2.5\u003c/sub\u003e concentration ranges for (a) Class I \u0026amp; (b) Class II.\u003c/p\u003e","description":"","filename":"Fig04.png","url":"https://assets-eu.researchsquare.com/files/rs-1174813/v1/e2437ce99929bd98f101222e.png"},{"id":18044482,"identity":"dfa6a40b-d7ca-4a88-9a72-1f488ddd62b4","added_by":"auto","created_at":"2022-02-08 20:54:32","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":615642,"visible":true,"origin":"","legend":"\u003cp\u003eWord clouds for (a) pre monsoon, (b) southwest monsoon, (c) post monsoon and (d) winter season.\u003c/p\u003e","description":"","filename":"Fig05.png","url":"https://assets-eu.researchsquare.com/files/rs-1174813/v1/bf0ad4377673319f19ff989b.png"},{"id":18044723,"identity":"d4246e07-3748-45b2-95eb-224a9279b390","added_by":"auto","created_at":"2022-02-08 20:57:32","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":27166,"visible":true,"origin":"","legend":"\u003cp\u003ePrediction accuracy of the model within different PM\u003csub\u003e2.5 \u003c/sub\u003econcentration categories\u003c/p\u003e","description":"","filename":"Fig06.png","url":"https://assets-eu.researchsquare.com/files/rs-1174813/v1/2a978c4851767495e74af36a.png"},{"id":18044916,"identity":"3c955f4f-b5c4-42a0-a155-c4ed0a9a6263","added_by":"auto","created_at":"2022-02-08 21:00:34","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":582499,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-1174813/v1/35eb500f-6bb7-4f3e-8600-eadd606eed85.pdf"}],"financialInterests":"","formattedTitle":"Development of a Deep Learning Technique to Analyze Air Quality from Social Media Response","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eIncreasing air pollution has become a major concern for the environmental quality of life in the urban areas. Recent studies visualize the grim reality of air pollution and its health risks across the cities of the world (Dandona, 2018; Pant et al., 2019; IQAir, 2019; WHO, 2020). For strategizing any control or eradication program, first step would be to monitor air quality at possibly high temporal and spatial resolutions to identify the source, location, quality and quantity of pollutants. The existing regulatory monitoring networks, especially in developing countries such as India, however make sparse spatial and temporal measurements (CPCB, 2020). Any additional information in real time, even if qualitative, would be an advantage for air quality management initiatives in countries like India which has 6 out of the 10 most polluted cities in the world (IQAir, 2019).\u003c/p\u003e \u003cp\u003eWith increased access to internet and emergence of social media platforms, people are now expressing their views and relation to the events around them like never before. Social media have grown tremendously in number and popularity ever since its emergence and now include multitude of platforms such as Twitter, Facebook, LinkedIn, YouTube channels, blogs, chat rooms, discussion forums etc. India with its population of ~1.3 billion people has over 560 million internet users and is the second largest online market in the world. The internet accessibility and use in the country largely varied based on factors like gender and socio-economic divide (Statista, 2020b). Analyzing public behavior on such platforms are not new to researchers and have been part of routine business analytics (Fan and Gordon, 2014). Many studies have also explored their potential for monitoring environmental events such as natural disasters in real time (Sakaki et al., 2010; Lindsay, 2011; Earle et al., 2011; Li et al., 2012; Kent and Capello Jr., 2013; Middleton et al., 2013, Singh et al., 2019).\u003c/p\u003e \u003cp\u003eSocial media responses also have the potential to extract useful information on air pollution (Robinson and Fialkowski, 2010; Jiang et al., 2015; Gurajala and Matthews, 2018; Jiang et al. 2019). It has been seen that the exposed people tend to share pictures, videos, blogs and tweets about air quality events almost immediately on social media within minutes of occurrence of an event (Robinson and Fialkowski, 2010). These user generated contents (UGC) enable us to derive real time information on air pollution. The response of users or \u0026lsquo;human sensors\u0026rsquo; may be direct in the form of immediate posts, tweets, blogs etc or indirect like re-sharing an already existing content or supporting it. This rise in popularity of social media has given people access to a humongous volume of information which can be used in a variety of valuable areas. Extracting and analyzing UGC to derive actionable knowledge is challenging task. However, recent approaches using machine learning (ML) techniques have been found promising for the purpose (Jackoway et al., 2011; Gurajala and Matthews, 2018). Such ML techniques are exploited curiously by new researchers are air quality prediction and forecasting as Xu et al. (2020), Chang et al. (2021) and Zhang et al. (2021).\u003c/p\u003e \u003cp\u003eUse of attention network in deep machine learning algorithms which is now widely recognised in sequence modeling as shown by Bahdanau et al. (2014) and Vaswani et al. (2017) to maximize performance in models using attention mechanism to connect between the input and output. Long - Short Term Memory (LSTM) are known to be effective for time series prediction tasks and multiple LSTM units are stacked to form the more effective Bi-directional Long - Short Term Memory (Bi-LSTM) models(Zhang et al., 2021). Wiedemann et al. (2018) and Wu et al. (2019) has also presented inspiring Bi-LSTM model architectures.\u003c/p\u003e \u003cp\u003eTable 1 shows a tabulation of some of the previous works that attempted to predict air quality using different ML techniques. It is seen that the influence of social media for air quality prediction is lesser explored than using air quality parameters and other polluting factors. Also self attention networks have not been prominently used for air quality prediction using social sensors.\u003c/p\u003e \u003cp\u003e \u003cb\u003eTable1\u003c/b\u003e \u003c/p\u003e \u003cp\u003eComparison between various reference papers that performed air quality prediction\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Taba\" border=\"1\"\u003e \u003ccolgroup cols=\"6\"\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eReference papers\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eMonitored air quality parameter\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSource/ Social media platform used for prediction\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eCountry or city under study\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003ePrediction Models/Techniques used, efficiency obtained\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eJiang et al. (2015)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAQI\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c4\" namest=\"c3\"\u003e \u003cp\u003eSina Weibo(Chinese Twitter)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eBeijing\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eGradient tree boosting (GTB) - 59%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGurajala and Matthews (2018)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePM\u003csub\u003e2.5\u003c/sub\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c4\" namest=\"c3\"\u003e \u003cp\u003eTwitter\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eParis, Delhi, London\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eAir quality not predicted\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eJiang et al. (2019)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAQI\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c4\" namest=\"c3\"\u003e \u003cp\u003eTwitter\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eCalifornia, Idaho, Illinois, Indiana, Ohio\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eNatural Language Processing (NLP) - 6.9\u0026ndash;17.7% improvement with social media intervention over base line method.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eXu et al. (2020)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePM\u003csub\u003e2.5\u003c/sub\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c4\" namest=\"c3\"\u003e \u003cp\u003eHistorical meteorological data, road network data, administrative boundary vector data, POI data.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eBeijing, Tianjin, Hebei\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eTemporal-spatial-regression-tree model, Grid prediction model - 90%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eChang et al. (2021)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePM\u003csub\u003e2.5\u003c/sub\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c4\" namest=\"c3\"\u003e \u003cp\u003eLocal and neighboring station data, chimney and abroad pollution data\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eTaiwan\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eAggregate LSTM - better than GTB, LSTM, SVR*\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eZhang et al. (2021).\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePM\u003csub\u003e2.5\u003c/sub\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c4\" namest=\"c3\"\u003e \u003cp\u003ePM\u003csub\u003e2.5\u003c/sub\u003e - Hourly, daily, restructured multi hour data\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eBeijing\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eLSTM, Bi-LSTM, EMD**- BiLSTM, - More than 95%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"6\"\u003e*SVR - Support vector regression\u003c/td\u003e\u003c/tr\u003e \u003ctr\u003e\u003ctd colspan=\"6\"\u003e**EMD - Empirical mode decomposition\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eIn the present work, Twitter responses related to air quality in Delhi were analyzed during 2019-2020 to develop a deep learning model that can predict the concentration of PM\u003csub\u003e2.5\u003c/sub\u003e (mass concentration of particulates having an aerodynamic diameter less than 2.5\u0026micro;m) from the tweets. A self attention network based classifier was used to characterize the tweets. The PM\u003csub\u003e2.5\u003c/sub\u003e data from continuous ambient air quality monitoring stations (CAAQMS) within the city were used to train the deep learning algorithm. The aim of the study is to:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eanalyze the fit and relevance of the physically monitored air quality parameter with the tweeting volume and behavior of people in an urban city of a developing country like Delhi, India and thus;\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eevaluate the ability to predict pollution concentrations from this virtual medium;\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003edevelop a new self attention deep learning classification model with high accuracy to classify tweets to those indicating poor air quality (Class I), good air quality (Class II) and neutral or noise tweets (Class 0) with minimal human intervention.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003edevelop a machine learning model to predict the PM\u003csub\u003e2.5\u003c/sub\u003e concentration by analysing a tweet.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eThe further organization of the paper is as follows: methodologies and machine learning techniques followed for the accomplishment of the objectives are elaborated in Section \u003cspan refid=\"Sec2\" class=\"InternalRef\"\u003e2\u003c/span\u003e. The results and discussions of the temporal and behavioral analysis of the study are discussed in Section \u003cspan refid=\"Sec14\" class=\"InternalRef\"\u003e3\u003c/span\u003e along with the efficiency on the classification and prediction models. Finally Section \u003cspan refid=\"Sec18\" class=\"InternalRef\"\u003e4\u003c/span\u003e wraps up the paper with the conclusions drawn, challenges faced by the study and the future scope of work.\u003c/p\u003e"},{"header":"2. Methodology","content":"\u003cp\u003eThis section discusses in detail the extraction of the social media responses from the chosen area under study during \u0026nbsp;2019 - 2020 and the various analyses performed on them along with the descriptions of the machine learning models used for the classification of the extracted data and the prediction of PM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003econcentration using the classified data.\u003c/p\u003e\n\u003cp\u003e2.1 Selection of study area and the social media platform\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTwitter was chosen as the social media platform\u0026nbsp;for the present study. Tweet response volumes, defined as the sum of direct and indirect response of people to air quality, for the period of 12 months ranging from March 2019 to February 2020 was used as the data for the analysis. Indirect responses included likes and re-tweets for a particular tweet.\u0026nbsp;There have been works like that of Jiang et al. (2015) who have considered the effects of direct and indirect responses separately but we chose to consider them as a collective response of the people of Delhi to their air quality issues.\u0026nbsp;Twitter is peculiar for its short and crisp contents limited over\u0026nbsp;280\u0026nbsp;characters and its prominence at the time of an event. Twitter has also become an increasingly important tool in politics, mass media communications, and many more such that it touches almost every sector of life. Many world leaders, governments, ministries, influencers, institutions, news channels have their official Twitter accounts to make announcements, influence and engage with the general public. In India there\u0026nbsp;were 13.15 million\u0026nbsp;Twitter users as of April 2020 (Statista, 2020a).\u003c/p\u003e\n\u003cp\u003eDelhi, capital city of India, was selected as a representative urban area to demonstrate the methodology presented in this paper. A preliminary analysis of tweets related to air pollution from the 100 most populated cities in India as part of this study identified Delhi as the city with maximum number of air quality episodes (IQAir, 2019) and maximum number of tweets related to air quality. Also, a public health emergency was declared at Delhi in November 2019 as the air quality index exceeded more than 3 times the \u0026lsquo;hazardous\u0026rsquo; level (CNN, 2020).\u003c/p\u003e\n\u003cp\u003e2.2 Data extraction\u003c/p\u003e\n\u003cp\u003eTweets were extracted from Twitter by attaining access to Twitter streaming API. The tweets were fetched with the help of keywords pertaining to air pollution along with the name of the city. About 15 different keyword combinations with Delhi (as shown in Table 2) were used for tweet extraction for the study period. The keywords thus chosen were different combinations of \u0026lsquo;air quality\u0026rsquo;, \u0026lsquo;air pollution\u0026rsquo; and \u0026lsquo;smog\u0026rsquo;. The data set collected included attributes such as the day, date and time of the tweet, the tweet text and the number of likes and retweets received by that particular tweet. In this manner, a data set of around 82,000 tweets were created. Tweets only in the English language were considered in this study as Twitter users in India are pre-dominantly English speakers (Poell and Rajagopalan, 2015). \u0026nbsp; In order to reduce the noise in the data set, few words were used to filter out unnecessary tweets like those generated by automated websites on a periodic basis regardless of the change in pollution intensity. Inorder to relate Twitter activity with ambient air quality information, PM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003e\u003csub\u003e\u0026nbsp;\u003c/sub\u003edata\u0026nbsp;from 38 CAAQMS stations within Delhi\u0026nbsp;during the period of March 2019 to February 2020 were\u0026nbsp;collected\u0026nbsp;from the\u0026nbsp;CPCB data repository\u0026nbsp;(CCR-CPCB, 2020).\u0026nbsp;The datasets generated and analysed during the current study are available from the corresponding author on reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u0026nbsp;Table 2\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;Keyword combinations used to extract tweets during study period\u003c/p\u003e\n\u003ctable border=\"1\" cellpadding=\"0\" cellspacing=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"33.333333333333336%\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"33.333333333333336%\"\u003e\n \u003cp\u003e\u003cstrong\u003eKeyword Combinations\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"33.333333333333336%\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"33.333333333333336%\"\u003e\n \u003cp\u003eAir Quality + Delhi\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"33.333333333333336%\"\u003e\n \u003cp\u003eAir Pollution + Delhi\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"33.333333333333336%\"\u003e\n \u003cp\u003eDelhiChokes\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"33.333333333333336%\"\u003e\n \u003cp\u003eAirQuality + Delhi\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"33.333333333333336%\"\u003e\n \u003cp\u003eAirPollution + Delhi\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"33.333333333333336%\"\u003e\n \u003cp\u003eChoke + Delhi\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"33.333333333333336%\"\u003e\n \u003cp\u003eAirQualityDelhi\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"33.333333333333336%\"\u003e\n \u003cp\u003eAirPollutionDelhi\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"33.333333333333336%\"\u003e\n \u003cp\u003eClean air + Delhi\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"33.333333333333336%\"\u003e\n \u003cp\u003eDelhiAirQuality\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"33.333333333333336%\"\u003e\n \u003cp\u003eDelhiAirPollution\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"33.333333333333336%\"\u003e\n \u003cp\u003eDelhiEmergency\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"33.333333333333336%\"\u003e\n \u003cp\u003eDelhi + Smog\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"33.333333333333336%\"\u003e\n \u003cp\u003eDelhiSmog\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"33.333333333333336%\"\u003e\n \u003cp\u003eDelhi + RightToBreathe\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e2.3 Data classification\u003c/p\u003e\n\u003cp\u003eEach tweet, within the allowed 280 characters contained words indicating good or bad air quality. Extracted dataset also contained some tweets that are either irrelevant or neutral with respect to the air quality. Segregating these tweets into various classes was an important step in the present study. The tweet data set was manually read and sorted by a team of 15 environmental engineers and the data set was sorted into 3 classes to segregate them into: poor air quality tweets (Class I), good air quality tweets (Class II) and noise or neutral tweets (Class 0). The sorted data set was then checked for variations in sorting and resorted (if needed) by an expert team of 3 environmental engineers.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThe sorted set consisted of tweets in Class I and II which contributed as indicators of the ambient air quality whereas Class 0 consisted of tweets generated by bots, irrelevant tweets, etc. Examples of tweets in their respective classes is shown in Table 3. The classified set was then pre processed and used to train a self attention model that would automatically classify a sample tweet into its suitable classes.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 3\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eExamples of classified tweets\u003c/p\u003e\n\u003ctable border=\"1\" cellpadding=\"0\" cellspacing=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"8.955223880597014%\"\u003e\n \u003cp\u003e\u003cstrong\u003eS. No.\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"70.64676616915423%\"\u003e\n \u003cp\u003e\u003cstrong\u003eTweets\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.398009950248756%\"\u003e\n \u003cp\u003e\u003cstrong\u003eClassification\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"8.955223880597014%\"\u003e\n \u003cul\u003e\n \u003cli\u003e1. \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u003c/li\u003e\n \u003c/ul\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"70.64676616915423%\"\u003e\n \u003cp\u003eNew Delhi: Over 1.2 million people died in India due to air pollution in 2017, said a global report on air pollution\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.398009950248756%\"\u003e\n \u003cp\u003eClass 0\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"8.955223880597014%\"\u003e\n \u003cul\u003e\n \u003cli\u003e2. \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u003c/li\u003e\n \u003c/ul\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"70.64676616915423%\"\u003e\n \u003cp\u003e#DelhiAirEmergency #DelhiPollution #DelhiBachao #DelhiAirQuality #DelhiNCRPollution\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.398009950248756%\"\u003e\n \u003cp\u003eClass 0\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"8.955223880597014%\"\u003e\n \u003cul\u003e\n \u003cli\u003e3. \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u003c/li\u003e\n \u003c/ul\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"70.64676616915423%\"\u003e\n \u003cp\u003eKab tak zindagi katoga bd or cigar mein, kuch din to gujaro delhi, in ncr\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.398009950248756%\"\u003e\n \u003cp\u003eClass 0\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"8.955223880597014%\"\u003e\n \u003cul\u003e\n \u003cli\u003e4. \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u003c/li\u003e\n \u003c/ul\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"70.64676616915423%\"\u003e\n \u003cp\u003eJust landed in #Delhi, the air here is just unbreathable\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.398009950248756%\"\u003e\n \u003cp\u003eClass 1\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"8.955223880597014%\"\u003e\n \u003cul\u003e\n \u003cli\u003e5. \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u003c/li\u003e\n \u003c/ul\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"70.64676616915423%\"\u003e\n \u003cp\u003eAmazing Air Quality today in Delhi! Enjoy the blue sky and clean air while it lasts.\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.398009950248756%\"\u003e\n \u003cp\u003eClass 2\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e2.4 Data analysis\u003c/p\u003e\n\u003cp\u003eThe following analyses were performed on the extracted and classified data set.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e2.4.1 Temporal analysis of tweet response volume in relation to PM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003econcentrations\u003c/p\u003e\n\u003cp\u003eThe volume of tweet responses related to air quality was compared with the ambient concentrations of PM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003e so as to analyze similarity in their trend. The comparison was done by graphical and statistical methods.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e2.4.1.1 Graphical analysis\u003c/p\u003e\n\u003cp\u003eSemi log time series graphs for the tweet response volumes for each of Class I and Class II tweets were plotted against the ground average PM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003econcentrations obtained from CPCB for the duration of the study. Twitter activity was plotted on a logarithmic scale considering their vast data range. The coherence of peaks and falls in the graph were analyzed. The graphical representation and the inferences drawn from it are discussed in Section 3.1.\u003c/p\u003e\n\u003cp\u003e2.4.1.2 Statistical analysis\u003c/p\u003e\n\u003cp\u003eIn-order to understand the statistical relationship between the two data sets, box plots were made with PM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003eon x-axis and tweet response volume on y-axis. The data was sorted on the basis of the classification of PM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003evalues into 6 categories as defined by the CPCB namely Good( 0 - 30\u0026micro;g/m\u003csup\u003e3\u003c/sup\u003e), Satisfactory( 31 - 60\u0026micro;g/m\u003csup\u003e3\u003c/sup\u003e), Moderate( 61 - 90\u0026micro;g/m\u003csup\u003e3\u003c/sup\u003e), Poor( 91 - 120\u0026micro;g/m\u003csup\u003e3\u003c/sup\u003e), Very Poor( 121 - 250\u0026micro;g/m\u003csup\u003e3\u003c/sup\u003e) and Severe ( above 250\u0026micro;g/m\u003csup\u003e3\u003c/sup\u003e) for plotting the box plots. The relationship of tweet response volumes of each class (Class I and II) in these categories were analyzed and the results are given in section 3.2.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e2.4.2 \u003cem\u003eBehavioral analysis using word clouds\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eWord clouds were generated \u0026nbsp;corresponding to each season to understand how people express the changes in the air quality through words using a python program utilizing necessary supporting libraries. India has 4 meteorological seasons namely Winter (January, February), Pre Monsoon (March \u0026ndash; May), Southwest Monsoon (June \u0026ndash; September) and Post Monsoon (October \u0026ndash; December) as per Indian meteorological Department (IMD). A prominent transition in the air quality levels in Delhi was observed during \u0026nbsp;post monsoon and winter seasons having the worst air quality which improved over pre monsoon and southwest monsoon seasons.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e2.5 Machine Learning Techniques\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eIn this study, machine learning techniques were used for (i) automated classification of the tweet data set and (ii) for the prediction of PM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003econcentration range from tweet response. The program codes constructed for the classification, training and prediction of the data in this work is available from the corresponding author on reasonable request.\u003c/p\u003e\n\u003cp\u003e2.5.1 Automated data classification model\u003c/p\u003e\n\u003cp\u003eThe tweets extracted over a year along with the likes and retweets appended to them comprised the dataset for the study. The dataset was a highly imbalanced multiclass data set with around 62,000 tweets in Class 0, 18000 tweets in Class I and 1700 tweets in Class II. As initial step, the manually classified dataset was tokenized into words. Preprocessing was performed in order to remove special symbols, stop words, punctuations, Twitter handles, etc. Natural Language Toolkit(NLTK) library was used for the same. The data set was further groomed by performing processes like word indexing, integer encoding and assigning them with pad descriptions before using the data for supervised learning. The respective classes of the tweets were label encoded. The processed data set was split into a training set and test set as 90% and 10% respectively for a 10 fold cross validation analysis to evaluate the overall performance of the model and a test-train split of 80%-20% for result analysis.\u003c/p\u003e\n\u003cp\u003eThe present study uses a multilayer classification model. The first layer is an embedding layer in which an embedding matrix is constructed from the data set using pertained global vector (GloVe) embedding for creating weights for the embedding layer (Pennington et al., 2014). The second layer is a BiLSTM layer (Units: 100) where the input sequences are analyzed in the forward and backward direction; that is, if a tweet has \u0026apos;n\u0026apos; number of words (w\u003csub\u003e1\u003c/sub\u003e, w\u003csub\u003e2\u003c/sub\u003e,...w\u003csub\u003en\u003c/sub\u003e) then the tweet is first analyzed to from w\u003csub\u003e1\u003c/sub\u003e to w\u003csub\u003en\u003c/sub\u003e by forward LSTM and w\u003csub\u003en\u003c/sub\u003e to w\u003csub\u003e1\u0026nbsp;\u003c/sub\u003eby backward LSTM and forms a word feature (Wf) for each word on the basis of both the forward analysis (f\u003csub\u003eF\u003c/sub\u003e) and backward analysis (f\u003csub\u003eB\u003c/sub\u003e) such that\u003c/p\u003e\n\u003cp\u003eWf = f\u003csub\u003eF\u0026nbsp;\u003c/sub\u003e‖\u003csub\u003e\u0026nbsp;\u003c/sub\u003ef\u003csub\u003eB\u003c/sub\u003e\u003c/p\u003e\n\u003cp\u003ewhere ‖ stands for concatenation function.\u003c/p\u003e\n\u003cp\u003eSimilarly a word feature is formed for every word of the data file. After this the data flows into a self attention layer (SeqSelfAttention) layer\u0026nbsp;where words are assigned weights and added to the resultant word features of the previous layer based on the relative importance of that word to the entire tweet. This layer improves the efficiency of the classification as \u0026apos;attention\u0026apos; is given to words based on their importance. For further detailed reference on Bi-LSTM and attention neural networks Bahdanau et al, 2014; Vaswani et al., 2017; Wiedemann et al, 2018; Zhang et al., 2018; Wu et al, 2019; Xu et al., 2020; Chang et al., 2021 and Zhang et al., 2021 may be used. Python libraries like Keras (https://keras.io/) and Keras_self_attention (https://pypi.org/project/keras-self-attention/) was used to support this layer.\u003c/p\u003e\n\u003cp\u003eThe context vectors or sentence feature vector that are obtained as the output of the attention layer using weighted sum function is then given to two dense layers or the fully connected layers where all the inputs and outputs are connected to all the neurons in each layer. The both dense layer contains 50 units and is activated with Rectified Linear Unit (ReLu) layer activation function. It is then flattened in the next layer, the flatten layer. The final output is obtained in a dense layer which has units equal to the number of output classes and is then given output activation by softmax function which provides the probabilities of the potential outcomes. Other parameters used in this modelling architecture includes: Epochs=5, Batch size:128, Optimizer: Adam. The model architecture is shown in Fig. 1.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e2.5.2 PM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003eprediction model\u003c/p\u003e\n\u003cp\u003eThe prediction model used similarity analysis to find the most similar tweets to a test tweet to predict the probable ground PM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003econcentration that would have been recorded on that day. The flow chart of the model is shown in Fig. 2. This model used\u0026nbsp;\u0026lsquo;spaCy\u0026rsquo; library available in python in order to find the top 10% of the most similar tweets of the test tweet from the season wise training set. A test tweet first undergoes model classification as mentioned in the previous section and then\u0026nbsp;from\u0026nbsp;the date of the tweet\u0026nbsp;identify\u0026nbsp;respective season\u0026rsquo;s training set. The test tweet\u0026nbsp;was\u0026nbsp;then analyzed\u0026nbsp;for similarity with each tweet of its corresponding training set and a similarity index\u0026nbsp;was\u0026nbsp;assigned to each training tweet. The training set tweets\u0026nbsp;are then arranged in descending of the cosine similarity index and the first 10% of the training set in this sorted set\u0026nbsp;was\u0026nbsp;taken as the resultant\u0026nbsp;\u0026lsquo;most\u0026nbsp;similar tweets set\u0026rsquo;. The\u0026nbsp;arithmetic mean of PM\u003csub\u003e2.5\u003c/sub\u003e concentrations corresponding to that similarity set was then reported as the output result of the model.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThe usefulness of the model was analyzed by measuring the correctly predicted concentrations as the percentage of the total predictions under each category specified by CPCB (CPCB, 2014). This was also represented as a histogram plot to understand the social media behavior with respect to change in air quality.\u003c/p\u003e"},{"header":"3. Results And Discussions","content":"\u003cp\u003eThe data classification model sorted the tweets into their corresponding Class (Class 0, I or II) with an overall accuracy of 96.7% from the 10 fold cross validation testing and an accuracy of 87.4% for the model using test-train split of 80% - 20%. On the classified data further analyses were done, the results of which are discussed in the following sections.\u003c/p\u003e\n\u003ch2\u003e3.1 Temporal variation of tweet responses with PM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003econcentrations: Graphical analysis\u003c/h2\u003e\n\u003cp\u003eA time series analysis for tweet responses of Class I and II along with corresponding PM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003econcentration was conducted for the duration of study (Fig. 3). \u0026nbsp;It was observed that there is rise in peak of the tweet response volume along with rise in pollution in the air and also a corresponding dip in tweet response volume with a fall in air pollution. \u0026nbsp;It was noticed that there is more response in Class I than Class II. The larger response volume in Class I might be due to the alertness of the public to poor quality of air that brings distress to their daily lives and their eagerness to spread the information to others, draw attention of government or other related or powerful individuals in the society, apart from the fact that the quality of air remained poor for a major part of the study period. The low volume in Class II is primarily due to lesser number of good air quality days in Delhi and also might be because good air quality days are appreciated mostly after a spike in pollution and not much while quality of air has been staying good for a period of time.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eDuring pre-monsoon season (March \u0026ndash; May 2019), Delhi had its worst air quality towards the middle of May and that was the time when the highest tweet response (around 200) in Class I was recorded. \u0026nbsp;The day with best air quality of the season reported no Class I tweets. High volume in Class II tweet response volume was found during the early days of March when the air quality used to improve to moderate category after a poor air quality period. Most of the times when the PM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003econcentrations fell within the desirable limit there was some Class II tweet response activity. A sudden fall in pollutant level resulted in more tweet responses than a gradual fall.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eDuring the period June \u0026ndash; September 2019 (southwest monsoon season), the average PM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003econcentration levels in Delhi were mostly within the desirable limit. Therefore, Class I tweet response activity was comparatively lesser (below 100) compared to other seasons. There were more number of days with minimal or no Class I tweet response volume. Whereas in the case of Class II, the maximum tweet response volumes were recorded on the day with best air quality of the year. Even though the quality of air in this season was mostly good, there wasn\u0026rsquo;t a tweet response for each of those days. Class II responses were mostly observed on good air quality days after an increase in pollution or if the pollutant concentrations were well below the desirable limits.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eIn the post monsoon season the average PM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003econcentration in Delhi was seldom within the desirable limits that is below 60\u0026mu;g/m\u003csup\u003e3\u003c/sup\u003e as specified in National Ambient Air Quality Standards by CPCB. This was the infamous times when the stubble burning in northern parts of India highly polluted the air in the northern states and public health emergency was declared in Delhi (CNN, 2020). This season also coincides with the Indian festival Diwali which is mostly celebrated with fireworks, crackers, etc. \u0026nbsp;Hence, for almost every day of this season, there was Class I tweet response. Class I tweet response volume was varied from 100 units and ranges as high as above 50000 on the most polluted day of the year when average PM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003econcentration over Delhi was around 550\u0026mu;g/m\u003csup\u003e3\u003c/sup\u003e. The time series of the tweet response volume of this season resembles much with its PM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003econcentrations. There was a high response in Class II as well. This was due to the fact that even a small fall in pollutant concentration was a relief to the citizen that they readily responded, increase the volume in Class II.\u003c/p\u003e\n\u003cp\u003eIn the winter season, Class I response was observed almost every day of the season owing to the fact that the average PM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003econcentration in Delhi during this season was seldom within the desirable limits. As the air pollution was not bad as in the previous season, the tweet response volume range has reduced. The low volume of response on 1st January 2020 in contrary to the high ground PM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003econcentration on the same day might have been a result of deviation of attention of public from air pollution to New Year celebrations and activities. As the air quality in this season was generally poor, the volume of tweet responses in Class II was less. However, Class II tweets were prominent with sharp fall in pollution as seen in the previous cases.\u003c/p\u003e\n\u003ch2\u003e3.2 Statistical relationship of tweet response volume \u0026nbsp;with PM\u003csub\u003e2.5\u003c/sub\u003e\u003c/h2\u003e\n\u003cp\u003eTo analyze the statistical relationship between the tweet response volumes and the average ground PM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003econcentrations, box plots were made for each class as shown in Fig. 4. For this analysis the data was classified according to the different PM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003econcentration categories as specified by CPCB.\u003c/p\u003e\n\u003cp\u003eIt was seen from Fig. 4(a) that the tweet response volumes for Class I were the largest when the PM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003econcentrations were in the poor - severe category and the lowest when it was in the good \u0026ndash; moderate category. This proves that the people were more responsive on Twitter and posted tweets indicating poor air quality during the times when the quality of air was in the poor \u0026ndash; severe conditions and the Class I tweet response volumes were proportional to the severity of air pollution. The small volume of Class I tweets even on days with good \u0026ndash; moderate quality of air could be because the air quality of a day was poorer than the previous day though the concentrations were not above the desired limits. It could also be a response of people visiting the city for the first time and finding the air quality to be bad compared to the place from which they arrived. From the present dataset, it was difficult to verify the behavior of the floating population.\u003c/p\u003e\n\u003cp\u003eClass II tweets were made not only based on present day air quality, but also when there was a decrease in pollution levels from previous days. Therefore, for the statistical analysis, the change in PM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003econcentration every 3 days (∆\u003csub\u003e3\u003c/sub\u003e) was calculated and the cumulative Class II tweet response volume during that period was considered and plotted to a box plot. From Fig. 4b, it was observed that days with good - satisfactory PM\u003csub\u003e2.5\u003c/sub\u003e concentrations had smaller ∆\u003csub\u003e3\u0026nbsp;\u003c/sub\u003eand therefore had lesser tweet response volumes. During the very poor \u0026ndash; severe air pollution days, the tweets response volumes were more in Class I category than Class II. The pollution was too high to have anything positive to tweet about. It was only while the air quality improved from severe to moderate the people were more responsive to tweet in a positive manner which explains the larger tweet response volume observed under moderate-poor \u0026nbsp;category.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003e3.3 Behavioral analysis from word clouds\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eWord clouds were\u0026nbsp;generated from the model classified \u0026amp; pre processed data segregated season wise in order to study the responsive nature of public about air quality over various situations and seasons. Word clouds helps in identifying the most frequent words in a season, the nature of emotions expressed through these words \u0026amp; the focus of people. The word cloud for each season of the study period is shown in Fig. 5.\u003c/p\u003e\n\u003cp\u003eThe word clouds for the different season in the study period collectively suggest that the citizens of Delhi were much interested in following the air quality index(AQI) of the city regularly. People seems to pay continuous attention to the news updates which are a major source of the AQI indices apart from dedicated apps, Twitter handles, etc and like and retweet tweets put up by news channel in regard with the varying air quality conditions. During poor air quality seasons, the tweeting pattern seem to be dependent on how distressed they feel about air pollution along with concern for health and calling for help and attention from the government and related departments. During good air quality season, people were found to tweet about the content and happiness in having a clear sky and asking the government to maintain the good conditions. The reason for an improvement or deterioration of the air quality is also of interest to the people.\u003c/p\u003e\n\u003cp\u003eFrom the word cloud for pre monsoon season, it was evident that the air quality was mostly in the \u0026lsquo;moderate\u0026rsquo; to \u0026lsquo;severe\u0026rsquo; zone. The social media responses have reported when the air quality turned \u0026lsquo;poor\u0026rsquo; or when it is \u0026lsquo;deteriorate\u0026rsquo;-ing or when it has shown an \u0026lsquo;improve\u0026rsquo;-ment. This was also the summer season in Delhi and there was tweets correlating air quality with \u0026lsquo;temperature\u0026rsquo; and \u0026lsquo;dust\u0026rsquo; and how it was an \u0026lsquo;unhealthy\u0026rsquo; condition and difficult to \u0026lsquo;breathe\u0026rsquo;. \u0026nbsp;In the monsoon season, tweets were mostly about \u0026lsquo;improve\u0026rsquo;-d air quality and tweets about \u0026lsquo;rain\u0026rsquo;, \u0026lsquo;wind\u0026rsquo;, and \u0026lsquo;storm\u0026rsquo; with an \u0026lsquo;AQI\u0026rsquo; mostly in the \u0026lsquo;good\u0026rsquo;, \u0026lsquo;satisfactory\u0026rsquo; or \u0026lsquo;moderate\u0026rsquo; ranges and asking the \u0026lsquo;government\u0026rsquo; to maintain the \u0026lsquo;clean\u0026rsquo; \u0026lsquo;sky\u0026rsquo;. The post monsoon season had large number of tweets in comparison to the other seasons and the wordcloud had the maximum words. The public mostly spoke about how the condition was \u0026lsquo;severe\u0026rsquo;, \u0026lsquo;poor\u0026rsquo;, \u0026lsquo;toxic\u0026rsquo;, \u0026lsquo;hazardous\u0026rsquo;, and an \u0026lsquo;emergency\u0026rsquo; situation which \u0026lsquo;choked\u0026rsquo; and affected \u0026nbsp;the daily life of the \u0026lsquo;people\u0026rsquo; and led to \u0026lsquo;school\u0026rsquo; to shut and \u0026lsquo;flights\u0026rsquo; delayed or redirected. They reported about the effect of \u0026lsquo;diwali\u0026rsquo; too in turning Delhi into a \u0026lsquo;gas chamber\u0026rsquo;. Even the slightest \u0026lsquo;improve\u0026rsquo;-ment brought a \u0026lsquo;relief\u0026rsquo; causing people to tweet about it. In the winter season, tweets were manly about \u0026lsquo;poor\u0026rsquo; quality of air still \u0026lsquo;remain\u0026rsquo;-ing in the city with its share of \u0026lsquo;improve\u0026rsquo;-ments. But the condition was mostly \u0026lsquo;toxic\u0026rsquo; and \u0026lsquo;bad\u0026rsquo; as \u0026lsquo;hell\u0026rsquo;. The \u0026lsquo;fog\u0026rsquo; contributed much to the \u0026lsquo;smog\u0026rsquo;.\u003c/p\u003e\n\u003ch2\u003e3.4 Prediction model results\u003c/h2\u003e\n\u003cp\u003eMachine learning based prediction model that can estimate PM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003econcentrations from the content of tweets was developed in the present study. Fig. 6 shows the percentage accuracy of these predictions for each PM\u003csub\u003e2.5\u003c/sub\u003e concentration category (CPCB, 2014).\u003c/p\u003e\n\u003cp\u003ePrediction accuracy was found to be high for extreme conditions of air quality such as good air quality (80%) and severe air quality (99%) categories. This was because the public could clearly experience these conditions without ambiguity, who then reflected these experiences with appropriate words in their tweet and thus making it easily predictable with our prediction model. As the PM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003econcentration moved into other categories, mixed tweeting behavior was observed thereby reducing the prediction accuracy for the model. Also, for a city like Delhi, the citizens were used to moderate to poor air quality conditions for the major part of the year and thus their reactivity to moderate air pollution situations were not so prominent. It was the deviations from the moderate conditions that made the people more alert and produced larger specific tweets suitable for predicting accurately. \u0026nbsp;\u003c/p\u003e"},{"header":"4. Conclusion","content":"\u003cp\u003ePresent study explored the social media behavior of urban dwellers towards air pollution and developed a deep learning based model to predict the concentration of PM\u003csub\u003e2.5\u003c/sub\u003e in the city based on the air pollution related social media responses. The population of Delhi expressed their agony and disappointment on urban air quality through various modes and present study analyzed their response to air quality through popular social media forum - Twitter. The Twitter generated data was sorted before analysis using a well efficient self attention network. The relationship of the Twitter responses to actual pollution level was inferred from temporal and statistical analysis of average PM\u003csub\u003e2.5\u003c/sub\u003e concentrations with tweet response volumes received per day over various season of a year and categories of PM\u003csub\u003e2.5\u003c/sub\u003e concentration ranges. The emotions and attitudes that the Delhi public hold towards air quality are depicted through the words they use in a tweet was analyzed using word clouds which showed that the Delhi public closely follows the AQI in Delhi through news streams or dedicated apps or through social media. It is evident from these analyses that social media is a powerful tool to monitor air quality variations in an urban city like Delhi. The PM concentration prediction results showed that there was high prediction efficiency (above 80%) for extreme air quality conditions and lower prediction efficiency for moderate air quality conditions.\u003c/p\u003e \u003cp\u003eIt was observed that limited people in India geo tag their tweets and therefore it was difficult to find location of specific tweets unless they mention the location in the tweet. The present study has only considered the text in tweets for analysis here. There were Twitter responses in other forms such as pictures, re-directing internet links, etc. which were not considered in this study. This may be overcome by developing a system to identify multiple languages, image interpretations, etc. Also tweets in regional languages were also overlooked in this study and could further develop the models to include more vernacular languages of the area and mixed scripts using cross lingual knowledge transfer and analysis. The nature of response of people on Twitter in terms of likes or retweets may be influenced by tweets of influential persons, popular organisations, bots, etc. This influence maybe taken into attention and the current study could be further advanced with the analysis of account origins of the tweets. It was seen during the pilot study of this study that many India cities face air quality issues but most of them do not have an active Twitter response culture. Metro cities showed relatively better responses and Delhi with its severe air quality issues had high response volume. Though the study focused on the social media behavior of people in Delhi and depends on several factors such as access to social media, environmental awareness among public, range of changes in air quality experienced by public etc., the methodology is replicable to any urban area in the world. Twitter was also not very popular in India as compared to other social media platforms like Facebook, YouTube, Instagram, etc. Such other social media platforms may also be explored in future to analyze if they hold higher relevance to real time air pollution related responses monitoring in India than Twitter for similar studies. In case of prediction of pollutant concentrations, similarity analysis might be insufficient for peculiar case of air quality conditions unfamiliar to the current training data set such as the unduly fall in air pollution levels in India during Covid-19 pandemic lockdown.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eData Availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003ePM\u003csub\u003e2.5\u0026nbsp;\u003c/sub\u003edata is available from CPCB at https://app.cpcbccr.com/ccr/#/caaqm-dashboard-all/caaqm-landing. Twitter data can be extracted using Twitter streaming API [https://developer.twitter.com/en/docs/twitter-api]. Data are also available from the corresponding author on reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNo funding was received to assist with the preparation of this manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthical Approval\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent to Participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent to Publish\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDecleration of Competing Interest\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCRediT authorship contribution statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eThushara Sudheish Kumbalaparambil\u003c/strong\u003e: Formal analysis, Writing -original draft. \u003cstrong\u003eRatish Menon\u003c/strong\u003e: Supervision, Writing- original draft, Writing- review \u0026amp; editing. \u003cstrong\u003eVishnu P Radhakrishnan\u003c/strong\u003e:Programming, Data curation. \u003cstrong\u003eVinod P Nair\u003c/strong\u003e: Formal analysis, Writing-review \u0026amp; editing\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eBahdanau D, Cho K, Bengio Y (2014). Neural machine translations by jointly learning to align and translate. ICLR 2015, arXiv:1409.0473\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCCR-CPCB (2020). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://app.cpcbccr.com/ccr/#/caaqm-dashboard-all/caaqm-landing\u003c/span\u003e\u003c/span\u003e, last accessed on 01-12-2021\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChang Y, Chiao H, Abimannan S, Huang Y, Tsai Y, and Lin K (2020). An LSTM-based aggregated model for air pollution forecasting. Atmospheric Pollution Research, Volume 11, Issue 8, August 2020, Pages 1451-1463\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCNN (2020). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://edition.cnn.com/2020/02/25/health/most-polluted-cities-india-pakistan-intl-hnk/index.html\u003c/span\u003e\u003c/span\u003e, last accessed on 01-12-2021\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCPCB (Central Pollution Control Board), 2014. National Air Quality Index. Control of Urban Pollution Series, CUPS/82/2014-15\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCPCB (2020). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://cpcb.nic.in/monitoring-network-3/\u003c/span\u003e\u003c/span\u003e, last accessed on 01-12-2021\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDandona L (2018). The impact of air pollution on deaths, disease burden, and life expectancy across the states of India: The Global Burden of Disease Study 2017. Lancet Planet Health; 3: e26\u0026ndash;39\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEarle P S, Bowden D C, and Guy M (2011). Twitter earthquake detection: Earthquake monitoring in a social world. Annals of Geophysics, 54(6), 708\u0026ndash;715\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFan W and Gordon M (2014). The Power of Social Media Analytics. Communications of the ACM. 57. 74-81. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1145/2602574\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGurajala S and Matthews J N (2018). Twitter data analysis to understand societal response to air quality. Proceedings of the 9th International Conference on Social Media and Society ,July 2018 ,Pages 82\u0026ndash;90, \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1145/3217804.3217900\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIQAir (2019). World Air Quality Report - Region \u0026amp; City PM2.5 Ranking. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.iqair.com/world-most-polluted-cities/world-air-quality-report-2019-en.pdf\u003c/span\u003e\u003c/span\u003e, last accessed on 01-12-2021\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJackoway A, Samet H and Sankaranarayanan J (2011). Identification of Live News Events using Twitter. Proceedings of the 3rd ACM SIGSPATIAL International Workshop on Location-Based Social Networks, Pages 25-32\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJiang J, Sun X, Wang W and Young S (2019). Enhancing Air Quality Prediction with Social Media and Natural Language Processing. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy, July 28 - August 2, 2019, Pages 2627\u0026ndash;2632\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJiang W, Wang Y, Tsou M-H, and Fu X (2015). Using Social Media to Detect Outdoor Air Pollution and Monitor Air Quality Index (AQI): A Geo-Targeted Spatiotemporal Analysis Framework with Sina Weibo (Chinese Twitter). PLoS ONE 10(10): e0141185.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKent J D and Capello Jr. H T (2013). Spatial patterns and demographic indicators of effective social media content during the Horsethief Canyon fire of 2012. Cartography and Geographic Information Science, 40:2,78-89\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi R, Lei K H, Khadiwala R and Chang K C (2012). TEDAS: A Twitter-based Event Detection and Analysis System. IEEE 28th International Conference on Data Engineering, 1084-4627/12\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLindsay B R (2011). Social Media and Disasters: Current Uses, Future Options, and Policy Considerations. Congressional Research Service, Washington, DC, 7-5700\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMiddleton S E, Middleton L and Modafferi S (2013). Real-Time Crisis Mapping of Natural Disasters Using Social Media. IEEE Intelligent Systems 29 (2), 9-17\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePant P, Lal R M, Guttikunda, S K, Russell AG, Nagpure A S, Ramaswami A and Peltier R E (2019). Monitoring particulate matter in India: recent trends and future outlook. Air Quality, Atmosphere \u0026amp; Health.12, 45-58\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePennington J, Socher R and Manning C D (2014). GloVe: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), 1532\u0026ndash;1543\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePoell, T and Rajagopalan S (2015). Connecting Activists and Journalists: Twitter communication in the aftermath of the 2012 Delhi rape. Journalism Studies, 16:5, 719-733\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRobinson E M, and Fialkowski W E (2010). Air Twitter: Using Social Media And Scientific Data To Sense Air Quality Events. 2010 IEEE International Geoscience and Remote Sensing Symposium, July 25-30, 2010, Honolulu, Hawaii, USA\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSakaki T, Okazaki M and Matsuo Y (2010). Earthquake shakes Twitter users: real-time event detection by social sensors. Proceedings of the 19th international conference on World Wide Web, Pages 851\u0026ndash;860\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSingh J P, Dwivedi Y K, Rana N P, Kumar A and Kapoor K K (2019). Event classification and location prediction from tweets during disasters. Annals of Operations Research, 283,737-757\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStatista (2020a). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.statista.com/statistics/242606/number-of-active-twitter-users-in-selected-countries/\u003c/span\u003e\u003c/span\u003e, last accessed on 01-12-2021\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStatista (2020b). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.statista.com/statistics/255146/number-of-internet-users-in-india/\u003c/span\u003e\u003c/span\u003e, last accessed on 01-12-2021\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez A N, Kaiser L and Polosukhin I (2017). Attention is all you need. 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWHO (2020). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.who.int/health-topics/air-pollution#tab=tab_1\u003c/span\u003e\u003c/span\u003e, last accessed on 01-12-2021.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWiedemann G, Ruppert E, Jindal R, and Biemann C (2018). Transfer Learning from LDA to BiLSTM-CNN for Offensive Language Detection in Twitter. Proceedings of GermEval 2018, 14th Conference on Natural Language Processing (KONVENS 2018). arXiv:1811.02906\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWu G, Tang G, Wang Z, Zhang Z, and Wang Z (2019). An Attention-Based BiLSTM-CRF Model for Chinese Clinic Named Entity Recognition. Special Section On Data-Enabled Intelligence For Digital Health, Vol. 7, pages 113942-113949\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXu C, Tong T, Zhang W and Meng M (2020). Fine-grained prediction of PM2.5 concentration based on multisource data and deep learning. Atmospheric Pollution Research, Volume 11, Issue 10, October 2020, Pages 1728-1737\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang L, Liu P, Zhao L, Wang G, Zhang W and Liu J (2021). Air quality predictions with a semi-supervised bidirectional LSTM neural network. Atmospheric Pollution Research, Volume 12, Issue 1, January 2021, Pages 328-339\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang Y, Wang J and Zhang X (2018). YNU-HPCC at SemEval-2018 Task 1: BiLSTM with Attention Based Sentiment Analysis for Affect in Tweets. Proceedings of the 12th International Workshop on Semantic Evaluation (SemEval-2018), Pages 273\u0026ndash;278\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":true,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"environmental-science-and-pollution-research","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"espr","sideBox":"Learn more about [Environmental Science and Pollution Research](https://www.springer.com/journal/11356)","snPcode":"11356","submissionUrl":"https://submission.nature.com/new-submission/11356/3","title":"Environmental Science and Pollution Research","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Twitter, air pollution, self attention network, Delhi, PM2.5 ","lastPublishedDoi":"10.21203/rs.3.rs-1174813/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-1174813/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eOne of the prominent new-age methods used by today\u0026rsquo;s population in spreading awareness or drawing attention on an issue or concern is through social media platforms. In this study, the responses of general public to air quality that they share on a popular social media platform \u0026ndash; Twitter, was taken as a virtual quantity that helps in measuring and analyzing the prevalent air quality. Machine learning technique based on self attention network was used to sort, clean and classify large amount of air pollution related Twitter responses extracted during 2019-2020 at Delhi in India. The temporal correlation of tweet response volumes with the ground monitored concentrations of air pollution and word cloud analysis were used to analyze the attitude of the public towards urban air quality. These analyses lead to the development of an air quality prediction model using \u0026lsquo;spaCy\u0026rsquo; \u0026ndash; similarity analysis method which attempts to predict the ground concentration range of pollutant PM\u003csub\u003e2.5\u003c/sub\u003e from tweets received on a particular day.\u003c/p\u003e","manuscriptTitle":"Development of a Deep Learning Technique to Analyze Air Quality from Social Media Response","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2022-02-08 20:54:30","doi":"10.21203/rs.3.rs-1174813/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major Revision","date":"2022-03-22T06:21:30+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2022-02-07T10:25:21+00:00","index":0,"fulltext":""},{"type":"reviewersInvited","content":"","date":"2022-02-07T10:04:26+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"Environmental Science and Pollution Research","date":"2022-01-20T22:48:26+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2022-01-12T05:04:31+00:00","index":"","fulltext":""},{"type":"submitted","content":"Environmental Science and Pollution Research","date":"2021-12-15T12:12:10+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"environmental-science-and-pollution-research","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"espr","sideBox":"Learn more about [Environmental Science and Pollution Research](https://www.springer.com/journal/11356)","snPcode":"11356","submissionUrl":"https://submission.nature.com/new-submission/11356/3","title":"Environmental Science and Pollution Research","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"67b91850-69d6-4c12-b88a-b547797db255","owner":[],"postedDate":"February 8th, 2022","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2022-08-29T09:24:14+00:00","versionOfRecord":[],"versionCreatedAt":"2022-02-08 20:54:30","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-1174813","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-1174813","identity":"rs-1174813","version":["v1"]},"buildId":"7rjqhiLT3MXkJMwkYKINL","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0