Machine Learning Algorithms for Forecasting Air Quality Index: A Predictive Analysis in theTaj Trapezium Zone (TTZ) of Agra

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Clean air is vital for sustaining life, and its quality directly impacts health. As industrialization progresses and populations grow, air pollution has become increasingly prevalent, emerging as a significant societal challenge. Air pollution has wide-ranging negative impacts on human health, including increased risk of early death and various ailments such as skin irritation., pulmonary infections, respiratory ailments, pneumonia, lung cancer, and cardiac complications. The purpose of this study is to use machine learning algorithms to predict the Taj Trapezium Zone’s air quality index. Such predictions can inform preventive action to mitigate air pollution. The study compares four algorithmic approaches: Light Gradient Boosting Machine (LightGBM), categorical boosting (Catboost), adaptive boosting (AdaBoost), and extreme gradient boosting (XGBoost). This algorithm’s performance is evaluated based on several parameters, including R-SQUARE, Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and ((MAE).
Full text 92,343 characters · extracted from preprint-html · click to expand
Machine Learning Algorithms for Forecasting Air Quality Index: A Predictive Analysis in theTaj Trapezium Zone (TTZ) of Agra | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Machine Learning Algorithms for Forecasting Air Quality Index: A Predictive Analysis in theTaj Trapezium Zone (TTZ) of Agra Swati Varshney, Jitendra Nath Shrivastava, Neha Gupta This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6358438/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Clean air is vital for sustaining life, and its quality directly impacts health. As industrialization progresses and populations grow, air pollution has become increasingly prevalent, emerging as a significant societal challenge. Air pollution has wide-ranging negative impacts on human health, including increased risk of early death and various ailments such as skin irritation., pulmonary infections, respiratory ailments, pneumonia, lung cancer, and cardiac complications. The purpose of this study is to use machine learning algorithms to predict the Taj Trapezium Zone’s air quality index. Such predictions can inform preventive action to mitigate air pollution. The study compares four algorithmic approaches: Light Gradient Boosting Machine (LightGBM), categorical boosting (Catboost), adaptive boosting (AdaBoost), and extreme gradient boosting (XGBoost). This algorithm’s performance is evaluated based on several parameters, including R-SQUARE, Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and ((MAE). Air Quality Index Particulate Matter Gaseous Pollutant Graded Response Action Plan(GRAP) Commission for Air Quality Management (CAQM) Figures Figure 1 Figure 2 Figure 3 1. Introduction One of the most important environmental issues facing contemporary society is air pollution, sometimes known as the invisible danger. Low-quality air, particularly in metropolitan regions, affects millions of people globally, leading to serious health consequences, including cardiovascular and respiratory conditions and mental health problems. As industries grow, transportation increases, and the effects of natural sources such as wildfires and volcanic activity combined with human-made emissions, the quality of the air we breathe continues to degrade. This research focuses on the comprehensive impact of air quality on both individual health and societal structures, with particular emphasis on how air pollution contributes to social inequities, economic losses, and healthcare burdens. Air plays a vital role in our life. The good quality of air that we breathe from a fresh environment improves mental health and reduces stress and all health issues. But with the advent of industrial development the curse of poor air quality is poisoning modern society. Currently, all eyes are on improving this negative factor as much as possible. The sources of polluted air are in two categories: -1. Natural Source, 2. Manmade source. Natural sources like natural phenomena emit some poisonous gases for example SO2, NO2, CO2, Sulphate etc. Manmade sources like the burning of fossil fuels, the greenhouse effect, and transportation emissions emit polluted oxygen, nitrogen, sulfur, CH4, N2O, etc. The current tool that gives the current status of the air quality is the Air Quality Index (AQI). Government organizations issue AQIs to measure air pollution levels and alert the public to potential hazards. It shows the daily status of the contaminants. More serious health issues are indicated by higher AQI levels. Based on the concentrations of air pollutants, the AQI is determined namely PM2.5, PM 10, NOx, CO, O3, and SO2 over a specific period, and health advisories are associated with the ranges into which the results are categorized. The current SMOG (Smoke + Fog) phenomena of winter is creating havoc in NCR and adjoining TTZ (TAJ TRAPESIUM ZONE) [1] area every winter from October to February. To Combat this air pollution menace, CAQM is the apex body of the Central Government. It has devised GRAP (Graded Response Action Plan) as AQI deteriorates. The Graded Response Action Plan (GRAP) is invoked at least three days before the air quality index (AQI) reaches the projected levels for a particular stage. The notable feature of this plan is the projected level of AQI and 3 3-day time frame for remedial GRAP action. It means there is a need for accurate prediction of AQI levels in short and long Term. The large difference in projected and actual value will trigger a false GRAP action. Which will cause a substantial loss of economic resources. Against this backdrop, there are increasing need to use newer prediction methods in the areas of Deep Learning, Machine Learning, Neural networks, AI, etc. The first-generation traditional forecasting tools are from the statistical and numerical domain. In the AQI complex domain, they lack high forecasting accuracy. This intricacy and amount of data cannot be handled by traditional methods. To handle the complexity of large data sets use of neural networks and or fuzzy logic has gained importance. However, in current research machine learning in AQI forecasting is showing promising results. The only rider is that with these researches ONE SIZE FIT All ML models are not possible. Each Geographic area or Urban center Has a peculiar pollution array. Such as NCR has PARALI-born pollution while TTZ has glass foundry/furnace-born pollution as a special factor. In this scenario, there is a need to study area-specific AQI forecasting techniques. 2. Objectives The following are the primary goals of this study paper: Assessment of major pollutants with specific reference to TTZ. Analysis of meteorological data, weather patterns, traffic flow, and industrial activity in the TTZ area. Identifying unusual pollutant sources, urban hotspots, or extreme weather conditions affecting air quality. Examining machine learning and deep learning algorithms for AQI prediction that is appropriate for the TTZ region. To prepare suitable DATA sets for these algorithms. To test best fitting models with evaluation help of Statistical parameters. To scale up results in real-world applications for AQI prediction. 3. Literature Survey A Final Report is created to create a ground framework for AQI pollution in Agra(TTZ) [1]. The present deep learning technique for air pollution concentration prediction was examined from the viewpoints of temporal, spatial, and spatiotemporal models. [2]. To estimate the air quality index for Ahmedabad, Gujarat, this paper compares many machine learning techniques, including SARIMA, SVM, and LSTM. [3] Frequent air pollution forecasting and monitoring are necessary to preserve ambient air quality. A study of metrological factors, particulate matter, and gaseous pollutants is important. This paper further deals with deep learning models for the city of Visakhapatnam AP [4].The study uses cutting-edge methods like SMOTE to identify the ideal dataset for AQI forecasts. The paper also works with SVR, RFR, and CatBoost Regression for four metro cities in India, mainly Delhi, Calcutta, Bangalore, and Hyderabad [5].The study develops a novel technique and technology for automatically identifying fog, air pollution, and air quality from a car. It uses IOT-based sensors for AQ prediction. It uses SVR, Machine Learning models to activate alarms in vehicles [6]. The study suggested an Internet of Things-enabled method for tracking and managing air pollution in a cloud computing setting. The paper uses the LR-MSV model for AQI forecasting [7]. The paper extends the comprehensive review of deep learning methods for AQI prediction. It gives an analysis of both deep learning method/deep learning methods and their applications for the AQI field in terms of expertise, application, and deficiency. It compares experiments on both types of methods. It gives a clear advantage for using deep learning methods for AQI prediction [8]. This paper deals with IOT-based real-time AQI calculation. It uses an IOT-enabled air pollution monitoring system based on current data only without using historical data only. This helps in validating various prediction models Vis avis actual AQI parameter in real time [9]. The paper deals with deep learning models using pollution data as well as metrological data in dependent and independent formats. It proposed four modules in the framework. The first module (AQIfp) calculates the pollutant concentration of harmful gases. In the Second module (AQIp) using historical air quality for calculating air quality. The third module (AQIm) uses meteorological features for finding AQI. In the fourth module (AQIc) combine the two modules AQIp and AQIm. [10]. The comparative examination of AQI measurement using different deep learning algorithms is the main emphasis of this paper. Like a Decision tree, Boosting with classical machine learning models such as ARIMA. It indicates that deep learning models have better AQI prediction performance [11]. The development of coupled CNN-long short-term memory (DCNN-LSTM) and deep convolutional neural networks (DCNN) is the subject of the study. It integrates AQI time-series data and meteorological elements. Long short-term memory (LSTM) networks are skilled at identifying both short-term and long-term correlations, whereas deep convolutional layers are excellent at extracting useful information and identifying the internal structure of input time series. [12].The paper presented a comparison of different machine-learning models. The authors determined that Extreme Gradient Boosting (XGB) is the most effective algorithm for the datasets examined. This research focused on making predictions one hour in advance. The findings suggest that this short-term prediction is inadequate for managing air quality [13]. This study employs a combination of Support Vector Regression and Long Short-Term Memory to categorize the AQI Index. The researchers have conducted a comparative analysis between this approach and existing methodologies. [14]. The researcher employed Support Vector Regression (SVR) and Random Forest Regression (RFR) to construct a predictive model for the Air Quality Index (AQI) in Beijing. To evaluate the performance of these regression models, the researcher utilized various accuracy metrics. [15]. In this Sang-hai case study, the researcher employed five models to forecast the Air Quality Index (AQI) using six air quality parameters. The findings suggest that tree-based models outperform neural networks in terms of regression metrics. [16].In this study, the researcher utilized data from 2014 to 2019 as the training set and data from 2020 to 2021 as the testing set. The author employed a CEEMDAN-ARMA-LSTM model for conducting predictions and analysis [17]. Two hybrid models for AQI prediction have been proposed by the author of this work. According to the authors, hybrid models outperform other machine learning models [18]. This study suggests using a hybrid model that combines regression, machine learning, and input variable selection (IVS) to anticipate and predict the daily amounts of particulate matter [19]. According to the author, ground stations are impractical for AqI prediction. Satellite data is more useful for AQI assessment and prediction [20] 4. Materials and procedures 4.1 Study Area Uttar Pradesh is home to many historical monuments. One of the most famous monuments in Uttar Pradesh is the Taj Mahal. This research article focuses on the Taj Mahal as the study area, which is illustrated in Fig. 1 , and TAJ TRAPEZIUM ZONE Fig. 2 was formed to control the adverse impact of AQI deterioration mainly on TAJ MAHAL. The 10,400 square kilometer Taj Trapezium Zone (TTZ) surrounds the Taj Mahal and serves to shield it from pollution. Agra Fort, Fatehpur Sikri, and the Taj Mahal are three World Heritage Sites that are part of the TTZ. Because TTZ is formed like a trapezoid and is situated near the Taj Mahal, it gets its name. The main pollutant concerns in TTZ are Foundry, Furnace, Leather /Shoe Industry, Petha industry, and Glass/Bengals making. Here also CAQM is invoking GRAP actions on a strict basis. Thus there is a need for better AQI level prediction in the TTZ area. However, very few studies on AQI prediction in the TTZ area have commenced. The current study aims to provide one step ahead in this direction. Six air quality monitoring stations exist in Agra: Manoharpur, Rohta, Sanjay, Sector 3 B Awas Vikas Colony, Shahjahan Garden, and Shastripuram. For this research, real-time data was collected from the Shastripuram station. The study employed four algorithms to analyze this real-time data, aiming to identify the most effective machine-learning model for accurate AQI prediction. By examining historical AQI trends, we can forecast future patterns, which will assist in implementing GRAP measures. 4.2 What is AQI? The AQI is a tool that gives the current status of air quality. Government organizations issue the AQI to measure air pollution levels and alert the public to hazards. The AQI is calculated based on air pollutant concentrations, namely PM2.5, PM 10, NOx, CO, O3, and SO2 over a specific period. CAQM: The Commission for Air Quality Management is the apex body of the central Government. It has devised GRAP (Graded Response Action Plan) as AQI deteriorates. Table 1 presents a comprehensive overview of the levels of concern associated with deteriorating Air Quality Index (AQI) levels, utilizing a spectrum of distinct colors to effectively convey the varying ranges of AQI. Each color is meticulously chosen to represent specific thresholds of air quality, clearly indicating the corresponding level of health risk for the general population. Table 2 delineates the various stages of the Graded Response Action Plan (GRAP) that are activated as AQI levels escalate, providing a structured framework for response measures based on the increasing severity of air pollution. Table 1 : Level of Concern as AQI deteriorates AQI Level of Concern 0-50 Good 50-100 Satisfactory 100-200 Moderate 200-300 Poor 300-400 Very Poor 400-450 Severe >450+ Severe + Table 2 : CAQM action plan as the level of Concern increases AQI Level of Concern CAQM Action 0-50 Good 50-100 Satisfactory 100-200 Moderate 200-300 Poor GRAP-1 300-400 Very Poor GRAP-2 400-450 Severe GRAP-3 >450+ Severe + GRAP-4 4.3 Machine Learning Algorithms Predicting outcomes three days in advance of reaching a specific threshold value within a given range presents significant difficulties. Our research employed four distinct algorithms to examine meteorological information using various parameters. The algorithms utilized in this investigation are LightGBM, CatBoost, XGBoost, and AdaBoost. 1) Light Gradient Boosting Machine ( LightGBM): It is a swift, high-performance framework built on decision trees. This algorithm employs a leaf-wise growth strategy and utilizes maximum delta values for expansion. It implements a histogram-based approach and offers a distributed and efficient architecture. LightGBM accelerates training and enhances accuracy through variable bucketing while consuming less memory. Although it excels with complex datasets, it may occasionally encounter overfitting issues. The algorithm’s design allows it to handle intricate data structures effectively. 2) Categorical Boosting (CatBoost): It is a variation of the boosting algorithm that can process both categorical and numerical data. This method eliminates the need for preprocessing categorical information, which significantly reduces time and effort for users. CatBoost incorporates techniques to mitigate overfitting issues and is specifically developed to handle regression and classification tasks involving large datasets. The algorithm expands in a depth-wise manner until it reaches its maximum level. 3) Extreme Gradient Boosting (XGBoost): It is a strong algorithm that is well-known for its remarkable performance, speed, accuracy, and efficiency. This method operates by Combining multiple weak learners and focuses on rectifying and enhancing the shortcomings of existing models. XGBoost constructs a robust learner model through an iterative process of merging numerous weak learners. It effectively manages intricate relationships and possesses the capability to mitigate overfitting. This algorithm employs parallel processing for computational efficiency. 4) Adaptive Boosting (AdaBoost): It is an ensemble technique that sequentially adds weak learner models. It operates similarly to XGBoost but with a key difference in the alpha factor, which is indirectly linked to the weak learners’ errors. This approach works well for jobs involving both regression and classification. Once the alpha parameter is determined, greater importance is assigned to weak learners. A lower weightage indicates that the weak learner is correcting errors, thus becoming a more effective model. AdaBoost is considered a valuable approach in machine learning for improving predictive accuracy. 4.4 Evaluation of Performance and Selection Model This research paper employs four criteria to determine the optimal algorithm for analyzing Taj Trapezium Zone datasets. The evaluation of performance will be based on these criteria, leading to the selection of the most suitable model for our datasets. 1) R Square : The coefficient of determination is another name for R-squared. It displays the way the data fits into a regression curve or line. R Squared Formula The following formulas are used to calculate the coefficient of determination or R2: R 2 = 1 – \(\:\frac{Rs}{Ts}\) (1) Where, where RS, the needed R Squared value by R2, and the total sum denote the residual sum of squares by TS. 2) Mean Squared Error : It calculates the average squared difference between the dataset’s actual values and its anticipated values. Mean Squared Error Formula $$\:\frac{1}{n}\sum\:_{i=0}^{n}{\left({y}_{i}-{\widehat{y}}_{i}\right)}^{2}$$ 2 Where n is the dataset’s number of observations. The observation’s true value is denoted by y i . The expected value of the i th observation is represented by ŷi. 3) Root Mean Square Error : One statistical metric that illustrates the magnitude of a fluctuating quantity is the Root Mean Square (RMS). Root Mean Square formula For a data set of n values, that is, x1, x2, x3,.... xn, the root mean square value is given as, X(rms)= \(\:\sqrt{\frac{{x}_{1}^{2}+{x}_{2}^{2}+{x}_{3}^{2}\cdot\:\cdots\:{x}_{n}^{2}}{n}}\) (3) Here, X(rms) is the data set’s assigned n observations’ root mean square value. 4) Mean Absolute Error : A straightforward yet effective metric for assessing the precision of regression models is mean absolute error, or MAE. Mean Absolute Error Formula X(mae) = \(\:\frac{1}{n}\sum\:_{i=1}^{n}\left({y}_{i}-{\widehat{y}}_{i}\right)\) (4) Where n is the number of observations in the dataset. y i is the actual value of the observation. ŷ i is the predicted value of the i th observation. 4.5 Methodology The proposed methodology to achieve the above objectives is discussed and illustrated in Fig. 3 4.5.1 Data Collection : Gather datasets on air quality from PCBs. Combine demographic data to understand variations in pollution components and geographic variations. Collect AQI data and historical records of trends and studies from secondary sources such as websites etc. 4.5.2 Pre-processing: Clean and process raw data to remove noise and handle missing values. Feature selection and modification to create relevant input variables (e.g., pollutant levels over time, population density). 4.5.3 Exploratory Data Analysis: Use of SMOTE Algorithm for dataset balancing. To highlight any performance variations that can result from balancing, both balanced and unbalanced datasets will be kept and used. 4.5.4 Dataset Split : In a typical machine learning process, the dataset is then divided into test and training data. It facilitates comparing the models’ accuracy to actual data. The Pareto principle dictates that the split is in an 80:20 ratio. 4.5.5 ML Algorithms : Currently, each selected regression model is utilized to make predictions, and as previously stated, its accuracy is measured for each balanced and imbalanced dataset. The research proposes to evaluate the following Algorithm models Light GBM (Light Gradient Boosting Machine), Catboost ( Categorical boosting) AdaBoost (Adaptive boosting) XGBoost (Extreme gradient boosting) 4.5.6 Performance Evaluation : Metrics including accuracy, mean absolute error (MAE), mean squared error (MSE), root mean squared error (RMSE), and R-SQUARE are used to test and compare each method’s performance 5. Result To scale up results in real-world applications for AQI prediction. Improved decision-making for Pollution Control Boards for GRAP planning. An App-based aid tool for local governments and individuals to reduce pollution exposure and forecast dangerous levels. Early detection of harmful pollution events, improving public health response times. 6. Conclusion This research paper introduces a reliable prediction model utilizing air quality datasets from the Taj Trapezium Zone. Various algorithms are evaluated to determine the most effective predictive model for these datasets. The findings of this study have potential applications for future forecasting. For the next research, estimates for specific city districts will also be provided using satellite images and more comprehensive data. Artificial intelligence (AI) would be an additional field to investigate to improve the models’ efficacy and inventiveness. This would assist in identifying the industrial regions that produce the greatest pollution. Our work would also get more detailed if we tried other algorithms and extended the study. The goal is to identify trends and offer options for raising a city’s air quality index. Investigating the most important contributing variables and effective strategies to reduce them is worthwhile. Additionally, by examining our dataset in greater detail to look for any interesting trends, The result will be THE BEAUTIFUL TAJ, FOREVER TAJ. Declarations Author Contribution Swati Varshney, Jitendra Nath Shrivastava, Neha Gupta reviewed the manuscript References M. Sharma, “Air quality assessment, trend analysis, emission inventory and source apportionment study in the city of Agra, 2021,” 2021. M. Zareba, H. Dlugosz, T. Danek, and E. Weglinska, “Big-data-driven machine learning for enhancing spatiotemporal air pollution pattern analysis,” Atmosphere, vol. 14, p. 760, Apr. 2023. N. N. Maltare and S. Vahora, “Air quality index prediction using machine learning for Ahmedabad city,” Digital Chemical Engineering, vol. 7, p. 100093, Mar. 2023. G. Ravindran, G. Hayder, K. Kanagarathinam, A. Alagumalai, and C. Sonne, “Air quality prediction by machine learning models: A predictive study on the Indian coastal city of Visakhapatnam,” Chemosphere, vol. 338, p. 139518, 2023. N. S. Gupta, Y. Mohta, K. Heda, R. Armaan, B. Valarmathi, and G. Arulkumaran, “Prediction of air quality index using machine learning techniques: A comparative analysis,” Journal of Environmental and Public Health, 2023. M. Dhanalakshmi, “A survey paper on vehicles emitting air quality and prevention of air pollution by using IoT along with machine learning approaches,” Turkish Journal of Computer and Mathematics Education (TURCOMAT), vol. 12, pp. 5950–5962, 2021. Dhanalakshmi and Radha, “Discretized linear regression and multiclass support vector-based air pollution forecasting technique,” Int. J. Eng. Trends Technol., vol. 70, pp. 315–323, Oct. 2022. B. Zhang, Y. Rong, R. Yong, D. Qin, M. Li, G. Zou, and J. Pan, “Deep learning for air pollutant concentration prediction: A review,” Atmos.Environ. (1994), vol. 290, p. 119347, Dec. 2022. B. Ghose and Z. Rehena, “Real-time air pollution monitoring system employing IoT,” in 2024 2nd International Conference on Intelligent Data Communication Technologies and Internet of Things (IDCIoT), pp. 53–58, IEEE, Jan. 2024. S. Sachdeva, H. Singh, S. Bhatia, and P. Goswami, “An integrated framework for predicting air quality index using pollutant concentration and meteorological data,” Multimed.Tools Appl., vol. 83, pp. 46967– 46996, Oct. 2023. A. Mishra and Y. Gupta, “Comparative analysis of air quality index prediction using deep learning algorithms,” Spat. Inf. Res., vol. 32, pp. 63–72, Feb. 2024. A. Barthwal and A. K. Goel, “Advancing air quality prediction models in urban India: a deep learning approach integrating DCNN and LSTM architectures for AQI time-series classification,” Model. Earth Syst. Environ., Feb. 2024. Y. Lee, I. Na, and Y. Son, “Evaluation of machine learning application on the prediction of particulate matter concentrations in small/medium-sized city,” Journal of Korean Society of Environmental Engineers, 2024. R. Janarthanan, P. Partheeban, K. Somasundaram, and P. Navin Elam-parity, “A deep learning approach for prediction of air quality index in a metropolitan city,” Sustainable Cities and Society, vol. 67, p. 102720, 2021. H. Liu, Q. Li, D. Yu, and Y. Gu, “Air quality index and air pollutant concentration prediction based on machine learning algorithms,” Applied Sciences, vol. 9, no. 19, 2019. X. Liu and H. Guo, “Air quality indicators and aqi prediction coupling long-short term memory (lstm) and sparrow search algorithm (SSA): A case study of Shanghai,” Atmospheric Pollution Research, vol. 13, no. 10, p. 101551, 2022 Y. Sun and J. Liu, “Aqi prediction based on ceemdan-arma-lstm,” Sustainability, vol. 14, no. 19, 2022. S. Zhu, X. Lian, H. Liu, J. Hu, Y. Wang, and J. Che, “Daily air quality index forecasting with hybrid models: A China case,” Environmental Pollution, vol. 231, pp. 1232–1244, 2017. M. T. Udristioiu, Y. EL Mghouchi, and H. Yildizhan, “Prediction, modeling, and forecasting of pm and aqi using hybrid machine learning,” Journal of Cleaner Production, vol. 421, p. 138496, 2023. S. Tinku Singh, Nikhil Sharma, and M. Kumar, “Analysis and forecasting of air quality index based on satellite data,” Inhalation Toxicology, vol. 35, no. 1–2, pp. 24–39, 2023. PMID: 36602767. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6358438","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":439991724,"identity":"9bf60315-5e23-4a92-be97-6ae11fdecb3a","order_by":0,"name":"Swati Varshney","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABLUlEQVRIiWNgGAWjYLCCBBAhASIq2Hjk2RuADAMLfBoYGxBazvDJGfYcAGmRwK+FAaaFsU3OmOEGwgSsQH5G8vMHD9vs8uVn9x58zMNmltg48/nVDT8KJBj427sTsGkxuJFm2JDYlmy54c65ZGMenrTEdumcsps9QIdJnDm7AasW6QTDhoQzzAYGEjlm0jwSxxIbZ+ek3eABajGQyMWqRX52+keglnoD+Rk55r95DP4nNtw8k3bzDx4tDLdzgLZUHDZguJFjxsyTwAb0Pvux2/hsMbj/pnBGQsVxA4MbOcaScw6wAQM5h+22jIEEDy6/yPcc3/Dxh0E1yGGGH97+A0Xl8Wc33/yxkeNv78XuMGTAxAOmeAzAJEHlIMD4A0yxPyBK9SgYBaNgFIwYAABmW2W6qhfR+AAAAABJRU5ErkJggg==","orcid":"","institution":"Invertis University Bareilly (U.P)","correspondingAuthor":true,"prefix":"","firstName":"Swati","middleName":"","lastName":"Varshney","suffix":""},{"id":439991725,"identity":"23db41e1-e1a1-4658-b0df-f7ae4129d440","order_by":1,"name":"Jitendra Nath Shrivastava","email":"","orcid":"","institution":"Invertis University Bareilly (U.P)","correspondingAuthor":false,"prefix":"","firstName":"Jitendra","middleName":"Nath","lastName":"Shrivastava","suffix":""},{"id":439991726,"identity":"44017881-e300-44c5-b2dd-f5931b050574","order_by":2,"name":"Neha Gupta","email":"","orcid":"","institution":"Symbiosis University of Applied Science Indore (M.P)","correspondingAuthor":false,"prefix":"","firstName":"Neha","middleName":"","lastName":"Gupta","suffix":""}],"badges":[],"createdAt":"2025-04-02 07:08:19","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6358438/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6358438/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":80201132,"identity":"87ba4f5f-fb1e-4a59-8a06-b535955e7419","added_by":"auto","created_at":"2025-04-09 06:47:04","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":302444,"visible":true,"origin":"","legend":"\u003cp\u003eStudy Area in India\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-6358438/v1/dea4933d7f1318da1d4b0224.png"},{"id":80201131,"identity":"f68f774f-c1ff-4e79-b86d-c6eb1e1a7898","added_by":"auto","created_at":"2025-04-09 06:47:04","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":568116,"visible":true,"origin":"","legend":"\u003cp\u003eTaj trapezium zone\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-6358438/v1/73a7e96a5af789a03a174f42.png"},{"id":80201130,"identity":"61b7a470-f9c5-4924-ac59-2b1f8ea4b32c","added_by":"auto","created_at":"2025-04-09 06:47:04","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":47472,"visible":true,"origin":"","legend":"\u003cp\u003eProposed Methodology Flow Chart\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-6358438/v1/73f359e381b973325abd9c74.png"},{"id":80202032,"identity":"c26d19b7-97a1-46d7-9b3b-d80feff2c4f4","added_by":"auto","created_at":"2025-04-09 06:55:05","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1538123,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6358438/v1/8199acd4-68b7-4f63-9ddc-a9398d535caa.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Machine Learning Algorithms for Forecasting Air Quality Index: A Predictive Analysis in theTaj Trapezium Zone (TTZ) of Agra","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eOne of the most important environmental issues facing contemporary society is air pollution, sometimes known as the invisible danger. Low-quality air, particularly in metropolitan regions, affects millions of people globally, leading to serious health consequences, including cardiovascular and respiratory conditions and mental health problems. As industries grow, transportation increases, and the effects of natural sources such as wildfires and volcanic activity combined with human-made emissions, the quality of the air we breathe continues to degrade. This research focuses on the comprehensive impact of air quality on both individual health and societal structures, with particular emphasis on how air pollution contributes to social inequities, economic losses, and healthcare burdens. Air plays a vital role in our life. The good quality of air that we breathe from a fresh environment improves mental health and reduces stress and all health issues. But with the advent of industrial development the curse of poor air quality is poisoning modern society. Currently, all eyes are on improving this negative factor as much as possible. The sources of polluted air are in two categories: -1. Natural Source, 2. Manmade source. Natural sources like natural phenomena emit some poisonous gases for example SO2, NO2, CO2, Sulphate etc. Manmade sources like the burning of fossil fuels, the greenhouse effect, and transportation emissions emit polluted oxygen, nitrogen, sulfur, CH4, N2O, etc. The current tool that gives the current status of the air quality is the Air Quality Index (AQI). Government organizations issue AQIs to measure air pollution levels and alert the public to potential hazards. It shows the daily status of the contaminants. More serious health issues are indicated by higher AQI levels. Based on the concentrations of air pollutants, the AQI is determined namely PM2.5, PM 10, NOx, CO, O3, and SO2 over a specific period, and health advisories are associated with the ranges into which the results are categorized. The current SMOG (Smoke\u0026thinsp;+\u0026thinsp;Fog) phenomena of winter is creating havoc in NCR and adjoining TTZ (TAJ TRAPESIUM ZONE) [1] area every winter from October to February. To Combat this air pollution menace, CAQM is the apex body of the Central Government. It has devised GRAP (Graded Response Action Plan) as AQI deteriorates. The Graded Response Action Plan (GRAP) is invoked at least three days before the air quality index (AQI) reaches the projected levels for a particular stage. The notable feature of this plan is the projected level of AQI and 3 3-day time frame for remedial GRAP action. It means there is a need for accurate prediction of AQI levels in short and long Term. The large difference in projected and actual value will trigger a false GRAP action. Which will cause a substantial loss of economic resources. Against this backdrop, there are increasing need to use newer prediction methods in the areas of Deep Learning, Machine Learning, Neural networks, AI, etc. The first-generation traditional forecasting tools are from the statistical and numerical domain. In the AQI complex domain, they lack high forecasting accuracy. This intricacy and amount of data cannot be handled by traditional methods. To handle the complexity of large data sets use of neural networks and or fuzzy logic has gained importance. However, in current research machine learning in AQI forecasting is showing promising results. The only rider is that with these researches ONE SIZE FIT All ML models are not possible. Each Geographic area or Urban center Has a peculiar pollution array. Such as NCR has PARALI-born pollution while TTZ has glass foundry/furnace-born pollution as a special factor. In this scenario, there is a need to study area-specific AQI forecasting techniques.\u003c/p\u003e"},{"header":"2. Objectives","content":"\u003cp\u003eThe following are the primary goals of this study paper:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003eAssessment of major pollutants with specific reference to TTZ.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAnalysis of meteorological data, weather patterns, traffic flow, and industrial activity in the TTZ area. Identifying unusual pollutant sources, urban hotspots, or extreme weather conditions affecting air quality.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eExamining machine learning and deep learning algorithms for AQI prediction that is appropriate for the TTZ region.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eTo prepare suitable DATA sets for these algorithms.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eTo test best fitting models with evaluation help of Statistical parameters.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eTo scale up results in real-world applications for AQI prediction.\u003c/p\u003e\n\u003c/li\u003e\n\u003c/ul\u003e"},{"header":"3. Literature Survey","content":"\u003cp\u003eA Final Report is created to create a ground framework for AQI pollution in Agra(TTZ) [1]. The present deep learning technique for air pollution concentration prediction was examined from the viewpoints of temporal, spatial, and spatiotemporal models. [2]. To estimate the air quality index for Ahmedabad, Gujarat, this paper compares many machine learning techniques, including SARIMA, SVM, and LSTM. [3] Frequent air pollution forecasting and monitoring are necessary to preserve ambient air quality. A study of metrological factors, particulate matter, and gaseous pollutants is important. This paper further deals with deep learning models for the city of Visakhapatnam AP [4].The study uses cutting-edge methods like SMOTE to identify the ideal dataset for AQI forecasts. The paper also works with SVR, RFR, and CatBoost Regression for four metro cities in India, mainly Delhi, Calcutta, Bangalore, and Hyderabad [5].The study develops a novel technique and technology for automatically identifying fog, air pollution, and air quality from a car. It uses IOT-based sensors for AQ prediction. It uses SVR, Machine Learning models to activate alarms in vehicles [6]. The study suggested an Internet of Things-enabled method for tracking and managing air pollution in a cloud computing setting. The paper uses the LR-MSV model for AQI forecasting [7]. The paper extends the comprehensive review of deep learning methods for AQI prediction. It gives an analysis of both deep learning method/deep learning methods and their applications for the AQI field in terms of expertise, application, and deficiency. It compares experiments on both types of methods. It gives a clear advantage for using deep learning methods for AQI prediction [8]. This paper deals with IOT-based real-time AQI calculation. It uses an IOT-enabled air pollution monitoring system based on current data only without using historical data only. This helps in validating various prediction models Vis avis actual AQI parameter in real time [9]. The paper deals with deep learning models using pollution data as well as metrological data in dependent and independent formats. It proposed four modules in the framework. The first module (AQIfp) calculates the pollutant concentration of harmful gases. In the Second module (AQIp) using historical air quality for calculating air quality. The third module (AQIm) uses meteorological features for finding AQI. In the fourth module (AQIc) combine the two modules AQIp and AQIm. [10]. The comparative examination of AQI measurement using different deep learning algorithms is the main emphasis of this paper. Like a Decision tree, Boosting with classical machine learning models such as ARIMA. It indicates that deep learning models have better AQI prediction performance [11]. The development of coupled CNN-long short-term memory (DCNN-LSTM) and deep convolutional neural networks (DCNN) is the subject of the study. It integrates AQI time-series data and meteorological elements. Long short-term memory (LSTM) networks are skilled at identifying both short-term and long-term correlations, whereas deep convolutional layers are excellent at extracting useful information and identifying the internal structure of input time series. [12].The paper presented a comparison of different machine-learning models. The authors determined that Extreme Gradient Boosting (XGB) is the most effective algorithm for the datasets examined. This research focused on making predictions one hour in advance. The findings suggest that this short-term prediction is inadequate for managing air quality [13]. This study employs a combination of Support Vector Regression and Long Short-Term Memory to categorize the AQI Index. The researchers have conducted a comparative analysis between this approach and existing methodologies. [14]. The researcher employed Support Vector Regression (SVR) and Random Forest Regression (RFR) to construct a predictive model for the Air Quality Index (AQI) in Beijing. To evaluate the performance of these regression models, the researcher utilized various accuracy metrics. [15]. In this Sang-hai case study, the researcher employed five models to forecast the Air Quality Index (AQI) using six air quality parameters. The findings suggest that tree-based models outperform neural networks in terms of regression metrics. [16].In this study, the researcher utilized data from 2014 to 2019 as the training set and data from 2020 to 2021 as the testing set. The author employed a CEEMDAN-ARMA-LSTM model for conducting predictions and analysis [17]. Two hybrid models for AQI prediction have been proposed by the author of this work. According to the authors, hybrid models outperform other machine learning models [18]. This study suggests using a hybrid model that combines regression, machine learning, and input variable selection (IVS) to anticipate and predict the daily amounts of particulate matter [19]. According to the author, ground stations are impractical for AqI prediction. Satellite data is more useful for AQI assessment and prediction [20]\u003c/p\u003e"},{"header":"4. Materials and procedures","content":"\u003cdiv id=\"Sec5\" class=\"Section2\"\u003e\n \u003ch2\u003e4.1 Study Area\u003c/h2\u003e\n \u003cp\u003eUttar Pradesh is home to many historical monuments. One of the most famous monuments in Uttar Pradesh is the Taj Mahal. This research article focuses on the Taj Mahal as the study area, which is illustrated in Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e, and TAJ TRAPEZIUM ZONE Fig.\u0026nbsp;2 was formed to control the adverse impact of AQI deterioration mainly on TAJ MAHAL. The 10,400 square kilometer Taj Trapezium Zone (TTZ) surrounds the Taj Mahal and serves to shield it from pollution. Agra Fort, Fatehpur Sikri, and the Taj Mahal are three World Heritage Sites that are part of the TTZ. Because TTZ is formed like a trapezoid and is situated near the Taj Mahal, it gets its name. The main pollutant concerns in TTZ are Foundry, Furnace, Leather /Shoe Industry, Petha industry, and Glass/Bengals making. Here also CAQM is invoking GRAP actions on a strict basis. Thus there is a need for better AQI level prediction in the TTZ area. However, very few studies on AQI prediction in the TTZ area have commenced. The current study aims to provide one step ahead in this direction. Six air quality monitoring stations exist in Agra: Manoharpur, Rohta, Sanjay, Sector 3 B Awas Vikas Colony, Shahjahan Garden, and Shastripuram. For this research, real-time data was collected from the Shastripuram station. The study employed four algorithms to analyze this real-time data, aiming to identify the most effective machine-learning model for accurate AQI prediction. By examining historical AQI trends, we can forecast future patterns, which will assist in implementing GRAP measures.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\n \u003ch2\u003e4.2 What is AQI?\u003c/h2\u003e\n \u003cp\u003eThe AQI is a tool that gives the current status of air quality. Government organizations issue the AQI to measure air pollution levels and alert the public to hazards. The AQI is calculated based on air pollutant concentrations, namely PM2.5, PM 10, NOx, CO, O3, and SO2 over a specific period.\u003c/p\u003e\n \u003cp\u003eCAQM: The Commission for Air Quality Management is the apex body of the central Government. It has devised GRAP (Graded Response Action Plan) as AQI deteriorates. Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e presents a comprehensive overview of the levels of concern associated with deteriorating Air Quality Index (AQI) levels, utilizing a spectrum of distinct colors to effectively convey the varying ranges of AQI. Each color is meticulously chosen to represent specific thresholds of air quality, clearly indicating the corresponding level of health risk for the general population. Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e delineates the various stages of the Graded Response Action Plan (GRAP) that are activated as AQI levels escalate, providing a structured framework for response measures based on the increasing severity of air pollution.\u003c/p\u003e\n \u003cdiv class=\"gridtable\"\u003e\n \u003cdiv class=\"colspec\" align=\"left\"\u003e\u003cstrong\u003eTable 1\u003c/strong\u003e: Level of Concern as AQI deteriorates\u0026nbsp;\u003c/div\u003e\n \u003cdiv class=\"colspec\" align=\"left\"\u003e\u0026nbsp;\u003c/div\u003e\n \u003c/div\u003e\n \u003cdiv class=\"gridtable\"\u003e\n \u003cdiv class=\"colspec\" align=\"left\"\u003e\u0026nbsp;\u003c/div\u003e\n \u003cdiv align=\"center\" style='margin:0in;text-align:justify;line-height:14.0pt;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\n \u003ctable style=\"border: none;width:344.6pt;border-collapse:collapse;\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 172.3pt;border-width: 1pt 1pt 3pt;border-style: solid;border-color: white;border-image: initial;background: rgb(214, 220, 229);padding: 0.75pt 5.4pt 0in;height: 15.35pt;vertical-align: top;\"\u003e\n \u003cp style='margin-top:0in;margin-right:0in;margin-bottom:0in;margin-left:1.0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cstrong\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003eAQI\u003c/span\u003e\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 172.3pt;border-top: 1pt solid white;border-left: none;border-bottom: 3pt solid white;border-right: 1pt solid white;background: rgb(214, 220, 229);padding: 0.75pt 5.4pt 0in;height: 15.35pt;vertical-align: top;\"\u003e\n \u003cp style='margin:0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cstrong\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003e\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; Level of Concern\u0026nbsp;\u003c/span\u003e\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 172.3pt;border-right: 1pt solid white;border-bottom: 1pt solid white;border-left: 1pt solid white;border-image: initial;border-top: none;background: rgb(112, 173, 71);padding: 0.75pt 5.4pt 0in;height: 16.3pt;vertical-align: top;\"\u003e\n \u003cp style='margin-top:0in;margin-right:0in;margin-bottom:0in;margin-left:1.0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cstrong\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003e0-50\u003c/span\u003e\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 172.3pt;border-top: none;border-left: none;border-bottom: 1pt solid white;border-right: 1pt solid white;background: rgb(112, 173, 71);padding: 0.75pt 5.4pt 0in;height: 16.3pt;vertical-align: top;\"\u003e\n \u003cp style='margin-top:0in;margin-right:0in;margin-bottom:0in;margin-left:1.0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003eGood\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 172.3pt;border-right: 1pt solid white;border-bottom: 1pt solid white;border-left: 1pt solid white;border-image: initial;border-top: none;background: yellow;padding: 0.75pt 5.4pt 0in;height: 3.9pt;vertical-align: top;\"\u003e\n \u003cp style='margin-top:0in;margin-right:0in;margin-bottom:0in;margin-left:1.0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cstrong\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003e50-100\u003c/span\u003e\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 172.3pt;border-top: none;border-left: none;border-bottom: 1pt solid white;border-right: 1pt solid white;background: yellow;padding: 0.75pt 5.4pt 0in;height: 3.9pt;vertical-align: top;\"\u003e\n \u003cp style='margin-top:0in;margin-right:0in;margin-bottom:0in;margin-left:1.0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003eSatisfactory\u0026nbsp;\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 172.3pt;border-right: 1pt solid white;border-bottom: 1pt solid white;border-left: 1pt solid white;border-image: initial;border-top: none;background: rgb(255, 192, 0);padding: 0.75pt 5.4pt 0in;height: 3.9pt;vertical-align: top;\"\u003e\n \u003cp style='margin-top:0in;margin-right:0in;margin-bottom:0in;margin-left:1.0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cstrong\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003e100-200\u003c/span\u003e\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 172.3pt;border-top: none;border-left: none;border-bottom: 1pt solid white;border-right: 1pt solid white;background: rgb(255, 192, 0);padding: 0.75pt 5.4pt 0in;height: 3.9pt;vertical-align: top;\"\u003e\n \u003cp style='margin-top:0in;margin-right:0in;margin-bottom:0in;margin-left:1.0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003eModerate\u0026nbsp;\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 172.3pt;border-right: 1pt solid white;border-bottom: 1pt solid white;border-left: 1pt solid white;border-image: initial;border-top: none;background: red;padding: 0.75pt 5.4pt 0in;height: 3.9pt;vertical-align: top;\"\u003e\n \u003cp style='margin-top:0in;margin-right:0in;margin-bottom:0in;margin-left:1.0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cstrong\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003e200-300\u003c/span\u003e\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 172.3pt;border-top: none;border-left: none;border-bottom: 1pt solid white;border-right: 1pt solid white;background: red;padding: 0.75pt 5.4pt 0in;height: 3.9pt;vertical-align: top;\"\u003e\n \u003cp style='margin-top:0in;margin-right:0in;margin-bottom:0in;margin-left:1.0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003ePoor\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 172.3pt;border-right: 1pt solid white;border-bottom: 1pt solid white;border-left: 1pt solid white;border-image: initial;border-top: none;background: rgb(112, 48, 160);padding: 0.75pt 5.4pt 0in;height: 3.9pt;vertical-align: top;\"\u003e\n \u003cp style='margin-top:0in;margin-right:0in;margin-bottom:0in;margin-left:1.0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cstrong\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003e300-400\u003c/span\u003e\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 172.3pt;border-top: none;border-left: none;border-bottom: 1pt solid white;border-right: 1pt solid white;background: rgb(112, 48, 160);padding: 0.75pt 5.4pt 0in;height: 3.9pt;vertical-align: top;\"\u003e\n \u003cp style='margin-top:0in;margin-right:0in;margin-bottom:0in;margin-left:1.0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003eVery Poor\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 172.3pt;border-right: 1pt solid white;border-bottom: 1pt solid white;border-left: 1pt solid white;border-image: initial;border-top: none;background: rgb(192, 0, 0);padding: 0.75pt 5.4pt 0in;height: 3.9pt;vertical-align: top;\"\u003e\n \u003cp style='margin-top:0in;margin-right:0in;margin-bottom:0in;margin-left:1.0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cstrong\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003e400-450\u003c/span\u003e\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 172.3pt;border-top: none;border-left: none;border-bottom: 1pt solid white;border-right: 1pt solid white;background: rgb(192, 0, 0);padding: 0.75pt 5.4pt 0in;height: 3.9pt;vertical-align: top;\"\u003e\n \u003cp style='margin-top:0in;margin-right:0in;margin-bottom:0in;margin-left:1.0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003eSevere\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 172.3pt;border-right: 1pt solid white;border-bottom: 1pt solid white;border-left: 1pt solid white;border-image: initial;border-top: none;background: rgb(196, 89, 17);padding: 0.75pt 5.4pt 0in;height: 20.8pt;vertical-align: top;\"\u003e\n \u003cp style='margin-top:0in;margin-right:0in;margin-bottom:0in;margin-left:1.0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cstrong\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003e\u0026gt;450+\u003c/span\u003e\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 172.3pt;border-top: none;border-left: none;border-bottom: 1pt solid white;border-right: 1pt solid white;background: rgb(196, 89, 17);padding: 0.75pt 5.4pt 0in;height: 20.8pt;vertical-align: top;\"\u003e\n \u003cp style='margin-top:0in;margin-right:0in;margin-bottom:0in;margin-left:1.0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003eSevere +\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n \u003cdiv class=\"colspec\" align=\"left\"\u003e\u0026nbsp;\u003c/div\u003e\n \u003cdiv class=\"colspec\" align=\"left\"\u003e\u0026nbsp;\u003c/div\u003e\n \u003cp\u003e\u003cstrong\u003eTable 2\u003c/strong\u003e: CAQM action plan as the level of Concern increases\u003c/p\u003e\n \u003cdiv align=\"center\" style='margin:0in;text-align:justify;line-height:14.0pt;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\n \u003ctable style=\"border: none;border-collapse: collapse;width: 498px;\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 105.85pt;border-width: 1pt 1pt 3pt;border-style: solid;border-color: white;border-image: initial;background: rgb(214, 220, 229);padding: 0.75pt 5.4pt 0in;height: 20.55pt;vertical-align: top;\"\u003e\n \u003cp style='margin:0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cstrong\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003eAQI\u003c/span\u003e\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 127.55pt;border-top: 1pt solid white;border-left: none;border-bottom: 3pt solid white;border-right: 1pt solid white;background: rgb(214, 220, 229);padding: 0.75pt 5.4pt 0in;height: 20.55pt;vertical-align: top;\"\u003e\n \u003cp style='margin:0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cstrong\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003eLevel of Concern\u0026nbsp;\u003c/span\u003e\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 140.1pt;border-top: 1pt solid white;border-left: none;border-bottom: 3pt solid white;border-right: 1pt solid white;background: rgb(214, 220, 229);padding: 0.75pt 5.4pt 0in;height: 20.55pt;vertical-align: top;\"\u003e\n \u003cp style='margin:0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cstrong\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003eCAQM Action\u003c/span\u003e\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 105.85pt;border-right: 1pt solid white;border-bottom: 1pt solid white;border-left: 1pt solid white;border-image: initial;border-top: none;background: rgb(112, 173, 71);padding: 0.75pt 5.4pt 0in;height: 4.85pt;vertical-align: top;\"\u003e\n \u003cp style='margin:0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cstrong\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003e0-50\u003c/span\u003e\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 127.55pt;border-top: none;border-left: none;border-bottom: 1pt solid white;border-right: 1pt solid white;background: rgb(112, 173, 71);padding: 0.75pt 5.4pt 0in;height: 4.85pt;vertical-align: top;\"\u003e\n \u003cp style='margin:0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003eGood\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 140.1pt;border-top: none;border-left: none;border-bottom: 1pt solid white;border-right: 1pt solid white;background: rgb(112, 173, 71);padding: 0.75pt 5.4pt 0in;height: 4.85pt;vertical-align: top;\"\u003e\n \u003cp style='margin-top:0in;margin-right:0in;margin-bottom:0in;margin-left:1.0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cspan style='font-size:16px;line-height:150%;font-family: \"Times New Roman\",serif;'\u003e\u0026nbsp;\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 105.85pt;border-right: 1pt solid white;border-bottom: 1pt solid white;border-left: 1pt solid white;border-image: initial;border-top: none;background: yellow;padding: 0.75pt 5.4pt 0in;height: 4.85pt;vertical-align: top;\"\u003e\n \u003cp style='margin:0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cstrong\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003e50-100\u003c/span\u003e\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 127.55pt;border-top: none;border-left: none;border-bottom: 1pt solid white;border-right: 1pt solid white;background: yellow;padding: 0.75pt 5.4pt 0in;height: 4.85pt;vertical-align: top;\"\u003e\n \u003cp style='margin:0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003eSatisfactory\u0026nbsp;\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 140.1pt;border-top: none;border-left: none;border-bottom: 1pt solid white;border-right: 1pt solid white;background: yellow;padding: 0.75pt 5.4pt 0in;height: 4.85pt;vertical-align: top;\"\u003e\n \u003cp style='margin-top:0in;margin-right:0in;margin-bottom:0in;margin-left:1.0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cspan style='font-size:16px;line-height:150%;font-family: \"Times New Roman\",serif;'\u003e\u0026nbsp;\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 105.85pt;border-right: 1pt solid white;border-bottom: 1pt solid white;border-left: 1pt solid white;border-image: initial;border-top: none;background: rgb(255, 192, 0);padding: 0.75pt 5.4pt 0in;height: 4.85pt;vertical-align: top;\"\u003e\n \u003cp style='margin:0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cstrong\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003e100-200\u003c/span\u003e\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 127.55pt;border-top: none;border-left: none;border-bottom: 1pt solid white;border-right: 1pt solid white;background: rgb(255, 192, 0);padding: 0.75pt 5.4pt 0in;height: 4.85pt;vertical-align: top;\"\u003e\n \u003cp style='margin:0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003eModerate\u0026nbsp;\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 140.1pt;border-top: none;border-left: none;border-bottom: 1pt solid white;border-right: 1pt solid white;background: rgb(255, 192, 0);padding: 0.75pt 5.4pt 0in;height: 4.85pt;vertical-align: top;\"\u003e\n \u003cp style='margin-top:0in;margin-right:0in;margin-bottom:0in;margin-left:1.0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cspan style='font-size:16px;line-height:150%;font-family: \"Times New Roman\",serif;'\u003e\u0026nbsp;\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 105.85pt;border-right: 1pt solid white;border-bottom: 1pt solid white;border-left: 1pt solid white;border-image: initial;border-top: none;background: red;padding: 0.75pt 5.4pt 0in;height: 4.85pt;vertical-align: top;\"\u003e\n \u003cp style='margin:0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cstrong\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003e200-300\u003c/span\u003e\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 127.55pt;border-top: none;border-left: none;border-bottom: 1pt solid white;border-right: 1pt solid white;background: red;padding: 0.75pt 5.4pt 0in;height: 4.85pt;vertical-align: top;\"\u003e\n \u003cp style='margin:0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003ePoor\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 140.1pt;border-top: none;border-left: none;border-bottom: 1pt solid white;border-right: 1pt solid white;background: red;padding: 0.75pt 5.4pt 0in;height: 4.85pt;vertical-align: top;\"\u003e\n \u003cp style='margin-top:0in;margin-right:0in;margin-bottom:0in;margin-left:1.0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cspan style='font-size:16px;line-height:150%;font-family: \"Times New Roman\",serif;'\u003eGRAP-1\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 105.85pt;border-right: 1pt solid white;border-bottom: 1pt solid white;border-left: 1pt solid white;border-image: initial;border-top: none;background: rgb(112, 48, 160);padding: 0.75pt 5.4pt 0in;height: 4.85pt;vertical-align: top;\"\u003e\n \u003cp style='margin:0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cstrong\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003e300-400\u003c/span\u003e\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 127.55pt;border-top: none;border-left: none;border-bottom: 1pt solid white;border-right: 1pt solid white;background: rgb(112, 48, 160);padding: 0.75pt 5.4pt 0in;height: 4.85pt;vertical-align: top;\"\u003e\n \u003cp style='margin:0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003eVery Poor\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 140.1pt;border-top: none;border-left: none;border-bottom: 1pt solid white;border-right: 1pt solid white;background: rgb(112, 48, 160);padding: 0.75pt 5.4pt 0in;height: 4.85pt;vertical-align: top;\"\u003e\n \u003cp style='margin-top:0in;margin-right:0in;margin-bottom:0in;margin-left:1.0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cspan style='font-size:16px;line-height:150%;font-family: \"Times New Roman\",serif;'\u003eGRAP-2\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 105.85pt;border-right: 1pt solid white;border-bottom: 1pt solid white;border-left: 1pt solid white;border-image: initial;border-top: none;background: rgb(192, 0, 0);padding: 0.75pt 5.4pt 0in;height: 4.85pt;vertical-align: top;\"\u003e\n \u003cp style='margin:0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003e400-450\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 127.55pt;border-top: none;border-left: none;border-bottom: 1pt solid white;border-right: 1pt solid white;background: rgb(192, 0, 0);padding: 0.75pt 5.4pt 0in;height: 4.85pt;vertical-align: top;\"\u003e\n \u003cp style='margin:0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003eSevere\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 140.1pt;border-top: none;border-left: none;border-bottom: 1pt solid white;border-right: 1pt solid white;background: rgb(192, 0, 0);padding: 0.75pt 5.4pt 0in;height: 4.85pt;vertical-align: top;\"\u003e\n \u003cp style='margin-top:0in;margin-right:0in;margin-bottom:0in;margin-left:1.0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cspan style='font-size:16px;line-height:150%;font-family: \"Times New Roman\",serif;'\u003eGRAP-3\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 105.85pt;border-right: 1pt solid white;border-bottom: 1pt solid white;border-left: 1pt solid white;border-image: initial;border-top: none;background: rgb(132, 60, 12);padding: 0.75pt 5.4pt 0in;height: 4.85pt;vertical-align: top;\"\u003e\n \u003cp style='margin:0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cstrong\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003e\u0026gt;450+\u003c/span\u003e\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 127.55pt;border-top: none;border-left: none;border-bottom: 1pt solid white;border-right: 1pt solid white;background: rgb(132, 60, 12);padding: 0.75pt 5.4pt 0in;height: 4.85pt;vertical-align: top;\"\u003e\n \u003cp style='margin:0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cspan style='font-size:16px;line-height:150%;font-family:\"Times New Roman\",serif;'\u003eSevere +\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 140.1pt;border-top: none;border-left: none;border-bottom: 1pt solid white;border-right: 1pt solid white;background: rgb(132, 60, 12);padding: 0.75pt 5.4pt 0in;height: 4.85pt;vertical-align: top;\"\u003e\n \u003cp style='margin-top:0in;margin-right:0in;margin-bottom:0in;margin-left:1.0in;text-align:justify;line-height:150%;font-size:13px;font-family:\"Palatino Linotype\",serif;color:black;'\u003e\u003cspan style='font-size:16px;line-height:150%;font-family: \"Times New Roman\",serif;'\u003eGRAP-4\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n \u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec7\" class=\"Section2\"\u003e\n \u003ch2\u003e4.3 Machine Learning Algorithms\u003c/h2\u003e\n \u003cp\u003ePredicting outcomes three days in advance of reaching a specific threshold value within a given range presents significant difficulties. Our research employed four distinct algorithms to examine meteorological information using various parameters. The algorithms utilized in this investigation are LightGBM, CatBoost, XGBoost, and AdaBoost.\u003c/p\u003e\n \u003cp\u003e1) Light Gradient Boosting Machine ( LightGBM): It is a swift, high-performance framework built on decision trees. This algorithm employs a leaf-wise growth strategy and utilizes maximum delta values for expansion. It implements a histogram-based approach and offers a distributed and efficient architecture. LightGBM accelerates training and enhances accuracy through variable bucketing while consuming less memory. Although it excels with complex datasets, it may occasionally encounter overfitting issues. The algorithm\u0026rsquo;s design allows it to handle intricate data structures effectively.\u003c/p\u003e\n \u003cp\u003e2) Categorical Boosting (CatBoost): It is a variation of the boosting algorithm that can process both categorical and numerical data. This method eliminates the need for preprocessing categorical information, which significantly reduces time and effort for users. CatBoost incorporates techniques to mitigate overfitting issues and is specifically developed to handle regression and classification tasks involving large datasets. The algorithm expands in a depth-wise manner until it reaches its maximum level.\u003c/p\u003e\n \u003cp\u003e3) Extreme Gradient Boosting (XGBoost): It is a strong algorithm that is well-known for its remarkable performance, speed, accuracy, and efficiency. This method operates by Combining multiple weak learners and focuses on rectifying and enhancing the shortcomings of existing models. XGBoost constructs a robust learner model through an iterative process of merging numerous weak learners. It effectively manages intricate relationships and possesses the capability to mitigate overfitting. This algorithm employs parallel processing for computational efficiency.\u003c/p\u003e\n \u003cp\u003e4) Adaptive Boosting (AdaBoost): It is an ensemble technique that sequentially adds weak learner models. It operates similarly to XGBoost but with a key difference in the alpha factor, which is indirectly linked to the weak learners\u0026rsquo; errors. This approach works well for jobs involving both regression and classification. Once the alpha parameter is determined, greater importance is assigned to weak learners. A lower weightage indicates that the weak learner is correcting errors, thus becoming a more effective model. AdaBoost is considered a valuable approach in machine learning for improving predictive accuracy.\u003c/p\u003e\n \u003cdiv class=\"gridtable\"\u003e\n \u003cdiv class=\"colspec\" align=\"left\"\u003e\u0026nbsp;\u003c/div\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e\n \u003ch2\u003e4.4 Evaluation of Performance and Selection Model\u003c/h2\u003e\n \u003cp\u003eThis research paper employs four criteria to determine the optimal algorithm for analyzing Taj Trapezium Zone datasets. The evaluation of performance will be based on these criteria, leading to the selection of the most suitable model for our datasets.\u003c/p\u003e\n \u003cp\u003e1) \u003cstrong\u003eR Square\u003c/strong\u003e: The coefficient of determination is another name for R-squared. It displays the way the data fits into a regression curve or line.\u003c/p\u003e\n \u003cp\u003eR Squared Formula\u003c/p\u003e\n \u003cp\u003eThe following formulas are used to calculate the coefficient of determination or R2:\u003c/p\u003e\n \u003cp\u003eR\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;1 \u0026ndash; \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\frac{Rs}{Ts}\\)\u003c/span\u003e\u003c/span\u003e (1)\u003c/p\u003e\n \u003cp\u003eWhere, where RS, the needed R Squared value by R2, and the total sum denote the residual sum of\u003c/p\u003e\n \u003cp\u003esquares by TS.\u003c/p\u003e\n \u003cp\u003e2) \u003cstrong\u003eMean Squared Error\u003c/strong\u003e: It calculates the average squared difference between the dataset\u0026rsquo;s actual values and its anticipated values.\u003c/p\u003e\n \u003cp\u003eMean Squared Error Formula\u003c/p\u003e\n \u003cdiv id=\"Equ1\" class=\"Equation\"\u003e\n \u003cdiv id=\"FileID_Equ1\" class=\"mathdisplay\"\u003e$$\\:\\frac{1}{n}\\sum\\:_{i=0}^{n}{\\left({y}_{i}-{\\widehat{y}}_{i}\\right)}^{2}$$\u003c/div\u003e\n \u003cdiv class=\"EquationNumber\"\u003e2\u003c/div\u003e\n \u003c/div\u003e\n \u003cp\u003eWhere n is the dataset\u0026rsquo;s number of observations. The observation\u0026rsquo;s true value is denoted by y\u003csub\u003ei\u003c/sub\u003e. The expected value of the i\u003csup\u003eth\u003c/sup\u003e observation is represented by ŷi.\u003c/p\u003e\n \u003cp\u003e3) \u003cstrong\u003eRoot Mean Square Error\u003c/strong\u003e: One statistical metric that illustrates the magnitude of a fluctuating quantity is the Root Mean Square (RMS).\u003c/p\u003e\n \u003cdiv class=\"BlockQuote\"\u003e\n \u003cp\u003eRoot Mean Square formula For a data set of n values, that is, x1, x2, x3,.... xn, the root mean square value is given as,\u003c/p\u003e\n \u003cp\u003eX(rms)= \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\sqrt{\\frac{{x}_{1}^{2}+{x}_{2}^{2}+{x}_{3}^{2}\\cdot\\:\\cdots\\:{x}_{n}^{2}}{n}}\\)\u003c/span\u003e\u003c/span\u003e (3)\u003c/p\u003e\n \u003c/div\u003e\n \u003cp\u003eHere, X(rms) is the data set\u0026rsquo;s assigned n observations\u0026rsquo; root mean square value.\u003c/p\u003e\n \u003cp\u003e4) \u003cstrong\u003eMean Absolute Error\u003c/strong\u003e: A straightforward yet effective metric for assessing the precision of regression models is mean absolute error, or MAE.\u003c/p\u003e\n \u003cp\u003eMean Absolute Error Formula\u003c/p\u003e\n \u003cdiv class=\"BlockQuote\"\u003e\n \u003cp\u003eX(mae) = \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\frac{1}{n}\\sum\\:_{i=1}^{n}\\left({y}_{i}-{\\widehat{y}}_{i}\\right)\\)\u003c/span\u003e\u003c/span\u003e (4)\u003c/p\u003e\n \u003cp\u003eWhere\u003c/p\u003e\n \u003c/div\u003e\n \u003cul\u003e\n \u003cli\u003e\n \u003cp\u003en is the number of observations in the dataset.\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003ey\u003csub\u003ei\u003c/sub\u003e is the actual value of the observation.\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eŷ\u003csub\u003ei\u003c/sub\u003e is the predicted value of the i\u003csup\u003eth\u003c/sup\u003e observation.\u003c/p\u003e\n \u003c/li\u003e\n \u003c/ul\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e\n \u003ch2\u003e4.5 Methodology\u003c/h2\u003e\n \u003cp\u003eThe proposed methodology to achieve the above objectives is discussed and illustrated in Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e\u003c/p\u003e\n \u003cdiv id=\"Sec10\" class=\"Section3\"\u003e\n \u003ch2\u003e\u003cstrong\u003e4.5.1 Data Collection\u003c/strong\u003e:\u003c/h2\u003e\n \u003cul\u003e\n \u003cli\u003e\n \u003cp\u003eGather datasets on air quality from PCBs.\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eCombine demographic data to understand variations in pollution components and geographic variations.\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eCollect AQI data and historical records of trends and studies from secondary sources such as websites etc.\u003c/p\u003e\n \u003c/li\u003e\n \u003c/ul\u003e\n \u003c/div\u003e\n \u003cdiv id=\"Sec11\" class=\"Section3\"\u003e\n \u003ch2\u003e4.5.2 Pre-processing:\u003c/h2\u003e\n \u003cul\u003e\n \u003cli\u003e\n \u003cp\u003eClean and process raw data to remove noise and handle missing values.\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eFeature selection and modification to create relevant input variables (e.g., pollutant levels over time, population density).\u003c/p\u003e\n \u003c/li\u003e\n \u003c/ul\u003e\n \u003c/div\u003e\n \u003cdiv id=\"Sec12\" class=\"Section3\"\u003e\n \u003ch2\u003e4.5.3 Exploratory Data Analysis:\u003c/h2\u003e\n \u003cul\u003e\n \u003cli\u003eUse of SMOTE Algorithm for dataset balancing.\u003c/li\u003e\n \u003cli\u003eTo highlight any performance variations that can result from balancing, both balanced and unbalanced datasets will be kept and used.\u003c/li\u003e\n \u003c/ul\u003e\u003cstrong\u003e4.5.4 Dataset Split\u003c/strong\u003e: In a typical machine learning process, the dataset is then divided into test and training data. It facilitates comparing the models\u0026rsquo; accuracy to actual data. The Pareto principle dictates that the split is in an 80:20 ratio.\u003cp\u003e\u003cstrong\u003e4.5.5 ML Algorithms\u003c/strong\u003e: Currently, each selected regression model is utilized to make predictions, and as previously stated, its accuracy is measured for each balanced and imbalanced dataset. The research proposes to evaluate the following Algorithm models\u003c/p\u003e\n \u003cul\u003e\n \u003cli\u003e\n \u003cp\u003eLight GBM (Light Gradient Boosting Machine),\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eCatboost ( Categorical boosting)\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eAdaBoost (Adaptive boosting)\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eXGBoost (Extreme gradient boosting)\u003c/p\u003e\n \u003c/li\u003e\n \u003c/ul\u003e\u003cstrong\u003e4.5.6 Performance Evaluation\u003c/strong\u003e: Metrics including accuracy, mean absolute error (MAE), mean squared error (MSE), root mean squared error (RMSE), and R-SQUARE are used to test and compare each method\u0026rsquo;s performance\n \u003c/div\u003e\n\u003c/div\u003e"},{"header":"5. Result","content":"\u003cul\u003e\n \u003cli\u003e\n \u003cp\u003eTo scale up results in real-world applications for AQI prediction.\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eImproved decision-making for Pollution Control Boards for GRAP planning.\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eAn App-based aid tool for local governments and individuals to reduce pollution exposure and forecast dangerous levels.\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eEarly detection of harmful pollution events, improving public health response times.\u003c/p\u003e\n \u003c/li\u003e\n\u003c/ul\u003e"},{"header":"6. Conclusion","content":"\u003cp\u003eThis research paper introduces a reliable prediction model utilizing air quality datasets from the Taj Trapezium Zone. Various algorithms are evaluated to determine the most effective predictive model for these datasets. The findings of this study have potential applications for future forecasting. For the next research, estimates for specific city districts will also be provided using satellite images and more comprehensive data. Artificial intelligence (AI) would be an additional field to investigate to improve the models\u0026rsquo; efficacy and inventiveness. This would assist in identifying the industrial regions that produce the greatest pollution. Our work would also get more detailed if we tried other algorithms and extended the study. The goal is to identify trends and offer options for raising a city\u0026rsquo;s air quality index. Investigating the most important contributing variables and effective strategies to reduce them is worthwhile. Additionally, by examining our dataset in greater detail to look for any interesting trends, The result will be THE BEAUTIFUL TAJ, FOREVER TAJ.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eSwati Varshney, Jitendra Nath Shrivastava, Neha Gupta reviewed the manuscript\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003e M. Sharma, \u0026ldquo;Air quality assessment, trend analysis, emission inventory and source apportionment study in the city of Agra, 2021,\u0026rdquo; 2021.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e M. Zareba, H. Dlugosz, T. Danek, and E. Weglinska, \u0026ldquo;Big-data-driven machine learning for enhancing spatiotemporal air pollution pattern analysis,\u0026rdquo; Atmosphere, vol. 14, p. 760, Apr. 2023.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e N. N. Maltare and S. Vahora, \u0026ldquo;Air quality index prediction using machine learning for Ahmedabad city,\u0026rdquo; Digital Chemical Engineering, vol. 7, p. 100093, Mar. 2023.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e G. Ravindran, G. Hayder, K. Kanagarathinam, A. Alagumalai, and C. Sonne, \u0026ldquo;Air quality prediction by machine learning models: A predictive study on the Indian coastal city of Visakhapatnam,\u0026rdquo; Chemosphere, vol. 338, p. 139518, 2023.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e N. S. Gupta, Y. Mohta, K. Heda, R. Armaan, B. Valarmathi, and G. Arulkumaran, \u0026ldquo;Prediction of air quality index using machine learning techniques: A comparative analysis,\u0026rdquo; Journal of Environmental and Public Health, 2023.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e M. Dhanalakshmi, \u0026ldquo;A survey paper on vehicles emitting air quality and prevention of air pollution by using IoT along with machine learning approaches,\u0026rdquo; Turkish Journal of Computer and Mathematics Education (TURCOMAT), vol. 12, pp. 5950\u0026ndash;5962, 2021.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e Dhanalakshmi and Radha, \u0026ldquo;Discretized linear regression and multiclass support vector-based air pollution forecasting technique,\u0026rdquo; Int. J. Eng. Trends Technol., vol. 70, pp. 315\u0026ndash;323, Oct. 2022.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e B. Zhang, Y. Rong, R. Yong, D. Qin, M. Li, G. Zou, and J. Pan, \u0026ldquo;Deep learning for air pollutant concentration prediction: A review,\u0026rdquo; Atmos.Environ. (1994), vol. 290, p. 119347, Dec. 2022.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e B. Ghose and Z. Rehena, \u0026ldquo;Real-time air pollution monitoring system employing IoT,\u0026rdquo; in 2024 2nd International Conference on Intelligent Data Communication Technologies and Internet of Things (IDCIoT), pp. 53\u0026ndash;58, IEEE, Jan. 2024.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e S. Sachdeva, H. Singh, S. Bhatia, and P. Goswami, \u0026ldquo;An integrated framework for predicting air quality index using pollutant concentration and meteorological data,\u0026rdquo; Multimed.Tools Appl., vol. 83, pp. 46967\u0026ndash; 46996, Oct. 2023.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e A. Mishra and Y. Gupta, \u0026ldquo;Comparative analysis of air quality index prediction using deep learning algorithms,\u0026rdquo; Spat. Inf. Res., vol. 32, pp. 63\u0026ndash;72, Feb. 2024.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e A. Barthwal and A. K. Goel, \u0026ldquo;Advancing air quality prediction models in urban India: a deep learning approach integrating DCNN and LSTM architectures for AQI time-series classification,\u0026rdquo; Model. Earth Syst. Environ., Feb. 2024.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e Y. Lee, I. Na, and Y. Son, \u0026ldquo;Evaluation of machine learning application on the prediction of particulate matter concentrations in small/medium-sized city,\u0026rdquo; Journal of Korean Society of Environmental Engineers, 2024.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e R. Janarthanan, P. Partheeban, K. Somasundaram, and P. Navin Elam-parity, \u0026ldquo;A deep learning approach for prediction of air quality index in a metropolitan city,\u0026rdquo; Sustainable Cities and Society, vol. 67, p. 102720, 2021.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e H. Liu, Q. Li, D. Yu, and Y. Gu, \u0026ldquo;Air quality index and air pollutant concentration prediction based on machine learning algorithms,\u0026rdquo; Applied Sciences, vol. 9, no. 19, 2019.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e X. Liu and H. Guo, \u0026ldquo;Air quality indicators and aqi prediction coupling long-short term memory (lstm) and sparrow search algorithm (SSA): A case study of Shanghai,\u0026rdquo; Atmospheric Pollution Research, vol. 13, no. 10, p. 101551, 2022\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e Y. Sun and J. Liu, \u0026ldquo;Aqi prediction based on ceemdan-arma-lstm,\u0026rdquo; Sustainability, vol. 14, no. 19, 2022.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e S. Zhu, X. Lian, H. Liu, J. Hu, Y. Wang, and J. Che, \u0026ldquo;Daily air quality index forecasting with hybrid models: A China case,\u0026rdquo; Environmental Pollution, vol. 231, pp. 1232\u0026ndash;1244, 2017.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e M. T. Udristioiu, Y. EL Mghouchi, and H. Yildizhan, \u0026ldquo;Prediction, modeling, and forecasting of pm and aqi using hybrid machine learning,\u0026rdquo; Journal of Cleaner Production, vol. 421, p. 138496, 2023.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e S. Tinku Singh, Nikhil Sharma, and M. Kumar, \u0026ldquo;Analysis and forecasting of air quality index based on satellite data,\u0026rdquo; Inhalation Toxicology, vol. 35, no. 1\u0026ndash;2, pp. 24\u0026ndash;39, 2023. PMID: 36602767.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Air Quality Index, Particulate Matter, Gaseous Pollutant, Graded Response Action Plan(GRAP), Commission for Air Quality Management (CAQM)","lastPublishedDoi":"10.21203/rs.3.rs-6358438/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6358438/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eClean air is vital for sustaining life, and its quality directly impacts health. As industrialization progresses and populations grow, air pollution has become increasingly prevalent, emerging as a significant societal challenge. Air pollution has wide-ranging negative impacts on human health, including increased risk of early death and various ailments such as skin irritation., pulmonary infections, respiratory ailments, pneumonia, lung cancer, and cardiac complications. The purpose of this study is to use machine learning algorithms to predict the Taj Trapezium Zone\u0026rsquo;s air quality index. Such predictions can inform preventive action to mitigate air pollution. The study compares four algorithmic approaches: Light Gradient Boosting Machine (LightGBM), categorical boosting (Catboost), adaptive boosting (AdaBoost), and extreme gradient boosting (XGBoost). This algorithm\u0026rsquo;s performance is evaluated based on several parameters, including R-SQUARE, Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and ((MAE).\u003c/p\u003e","manuscriptTitle":"Machine Learning Algorithms for Forecasting Air Quality Index: A Predictive Analysis in theTaj Trapezium Zone (TTZ) of Agra","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-04-09 06:46:59","doi":"10.21203/rs.3.rs-6358438/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"1bc1ee22-3784-41dd-a9e2-fd795baf0141","owner":[],"postedDate":"April 9th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-04-11T13:21:49+00:00","versionOfRecord":[],"versionCreatedAt":"2025-04-09 06:46:59","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6358438","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6358438","identity":"rs-6358438","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00