Computation of Water Quality Index and Its Estimation Using Machine Learning Techniques | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Computation of Water Quality Index and Its Estimation Using Machine Learning Techniques Kakoli Banerjee, Pradeep Kumar, Ajay Kumar, Kamal Upreti, Shubham Mahajan, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5859247/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Water quality is an essential measure for maintaining health and quality of life. In this study, the water quality index has been computed for Gautam Buddha Nagar using the Brown et al. method by collection of 51 groundwater samples. For this purpose, six physicochemical water parameters were analysed, namely pH, Hardness, Turbidity, C.O.D., D.O. and B.O.D. The readings indicate that the groundwater condition of Gautam Buddha Nagar is extremely poor. The computation of the Water Quality Index is a complex task. The study found that when all input variables are available, Machine Learning Techniques can be employed to vastly reduce the complexity in the computation while giving a high accuracy of 99.99% using Linear Regression, followed by an accuracy of 99.97% using Support Vector Regressor. Collection of input data is a time-consuming and costly process, therefore, the dimensionality of the input data was reduced through correlation analysis in an attempt to compute the water quality index by using just a single parameter. The best score of 81.05% was obtained using Linear Regression when Turbidity was used as the only feature, due to its high correlation with the target variable. The algorithms used for this analysis are: Linear Regression, Support Vector Regressor, Decision Tree and Random Forest Regressor. Water Quality Index Groundwater Supervised Machine Learning Regression environmental pollutants data acquisition Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 Figure 10 Figure 11 Figure 12 Figure 13 Figure 14 1 Introduction Water is a key natural resource and asset (S. Tyagi et. al. 2020), which is essential for human existence as well as the sustainability of the planet (2022). Water quality is integral in determining health, main-taining gender equality and food security as well as ensuring economic development and social growth (M. K. Jhaet. Al. 2020). Traditionally, the water quality is calculated by assessing and analysing the physical, chemical and bi-ological properties of water (2022), with minimum value requirements for water constituents or quality indicators (2022). One of the primary sources of water is groundwater, which is available almost everywhere under the surface of the earth in the form of thousands of local aquifer systems and not as one single unit (S. Singh and A. Hussian, 2016). In developing nations, such as India, it plays a vital role in strengthening the country’s economic growth (P. Ravikumar et. al. 2011). The Central Pollution Control Board has 870 monitoring stations across the country which monitors surface water quality on a monthly or quarterly basis and groundwater on a half yearly basis, which indicates that organic pollutants predominantly cause contamination (R. M. Bhardwaj, 2005). Apart from India, it has been estimated that over 50 countries in the world are treated with polluted or partially treated polluted water (J. Mateo-Sagasta et. al. 2017). In developing countries, over 80 percent of polluted water is used for irrigation purposes (D. D. Mara and S. Cairncross, 2019). Other diseases related to excreta are extremely common in developing countries as the wastewater contains pathogens such as viruses and bacterias, which leads to a risk to public safety and health. Therefore the contaminated wa-ter which contains plant nutrients may benefit soil but it also contains soluble salts and heavy metals in high concentrations (A. Gafoor et. al. 1994). This water is used by farmers who want to cut on their ex-penses, but over long periods of time extensive use of such methods for irrigation leads to groundwater contamination, the harmful effects of which can last for several years (M. Ibrahim and S. Salmon, 2020). Due to overgrowing population and urban expansion numerous water quality issues are becoming a concern, making water quality an important tool for measuring health related issues and safety measures in both humans and aquatic life (O. Tunc Dede et. al. 2013). Although some waste disposal practices might not affect groundwater, it directly contaminates surface level water bodies (E. Sánchez et al.,2007). Other agents such as poisonous metals may be mobile enough to seep into the ground and mix with the drinking water supplies making it unfit for human consumption– for example Silver metals in form of AgNPs. To estimate the water grade, a metric called the Water Quality Index (WQI) is used which is represented through a single number which indicates the water quality through aggregation of individual water quality parameters, with the score representing the water quality wherein a lower score indicates poor quality and a higher score indicates better quality (A. Lumb, et. al. 2011). WQI is a rating which reflects the compound impact of various water quality parameters and is considered a powerful tool to evaluate both surface and groundwater pollution, as well as in quality upgradation programmes. Its importance can be understood by the fact that it is a single term which still gives a unique rating for depicting over-all water quality (K. A. Shah and G. S. Joshi, 2017). There are various indices available which make use of different parameters depending on the intended objectives for its estimation (C. E. Canada and Canada. 2003). Although the WQI estimation might provide us with crucial information about the state of the water and its quality, there is no universal method of calculating WQI as of now moreover, WQI calculations consume not only resources such as time as it is a complex and a lengthy process, it also has multiple inconsistencies as WQI often uses different equations for calculation (D. T. Bui et. al. 2020). Another such drawback of WQI estimation is the way in which the magnitude of weights are calculated for each water quality parameter, most of which are subjective in nature i.e. they do not take into account the factor that it may cause hazards (R. Roy, 2018). Therefore, the primary objective of our study is to minimize the number of input variables from the feature set in order to calculate WQI, which should ideally be esti-mated by just using a single parameter. There are majorly two broad areas of work in this study: calculation of WQI and development of machine learning modes for its prediction. Our approach uses the Brown et. al. method for WQI calculation that first selects the parameters, namely pH, Turbidity (NTU), Hardness (mg/L), D.O. (mg/L), B.O.D. (mg/L), C.O.D. (mg/L), then these parameters are assigned a weight factor, and the Q value is calculated. Finally the WQI is computed which gives the grade for the water quality. Once the WQI is found, the dataset is split into training and testing sets and various Supervised Machine Learning algorithms are applied such as Linear Regression, Decision Tree Regressor, Random Forest Regressor and Support Vector Regressor. The motive for this approach was to reduce the dimensionality of the input variables required to compute WQI, while maintaining a high accuracy. 2 Study Area Due to high anthropogenic activities like the development and setup of industries and factories in Gau-tam Buddha Nagar, the contamination of groundwater has gone up significantly which requires more investigation for the fulfilment of current and future potable water requirements (R. K. Horton, 2015, Murthy et al. 2004). Gautam Buddha District has a soil classification that consists of a combination of excessively drained, well drained and moderately well drained soil (R. K. Horton, 2015). Gautum Buddha Nagar was selected as the field of study and was used as a site for the collection of samples and data. Noida is a district in Uttar Pradesh which comes under Gautam Buddha Nagar and is close to the Na-tional Capital Territory of Delhi. The city spans across the latitudes 28° 26′ 39 ′′ N to 28° 38′ 10′′ N and longitudes 77° 29′ 53′′ E to 77° 17′ 29′′ E covering a total area of 203 sq. km. 3. History of WQI The WQI was developed in 1965 in United States by Horton, it was based on common parameters used to determine water quality (such as pH, coliforms, dissolved oxygen, sewage treatment, specific con-ductance, carbon chloroform extract, alkalinity and chlorides). While most WQI models operate on a scale of zero to hundred, the Horton model implements a scale in which zero represents excellent water quality and values higher than hundred represent extremely poor water quality (Uddin, et. al. 2020). The criteria selected by Horton primarily focused on three factors is listed below in Table 1 (T. Abbasi and S. A. 2012). Variables used for the computation the water quality index should be limited Variables used should significant in most areas Variables used should have available and reliable data Following Horton, a large number of WQI models were formulated by several different national and in-ternational bodies. In 1970, Brown et al. developed the ‘Weighted Arithmetic Index Method’, with the enlisted steps: Calculation of sub-index of quality rating Calculation of unit weight Aggregation of quality rating with unit weight resulting in WQI Table 1 Water Quality Index range for its classification WQI Range Water Quality Status Possible Usage 0–5 Excellent Water Quality Drinking, irrigation and industrial 26–50 Good Water Quality Drinking, irrigation and industrial 51–75 Poor Water Quality Irrigation and industrial 76–100 Very Poor Water Quality Irrigation Above 100 Unsuitable for drinking and propagation of fish Proper treatment required before use 3.1 Methods for finding sub-indices The following are the three most commonly followed methods for calculating sub-indices: Expert Opinion For several WQIs one of the challenges is to select important parameters to calculate the final index ag-gregation as subjective assessments are a part of the initially selected parameters (N. Akhtar et al., 2021). To reduce the subjectivity and the uncertainty, several experts in parameters selection have been used (T. D. Banda and M. V. Kumarasamy, 2020). The expert judgement process is paired with parameter selec-tion using three different methods (M. A. Meyer and J. M. Booker, 2001). Individual interviews- Individual interviews are conducted through mail or other methods of commu-nication where experts (people who possess high knowledge of a subject) give their remarks on a topic or subject (J. A. G. M. Van Dijk, 2090). Virtual groups-: The virtual groups are similar to interviews in the way that in this method too experts give opinions but instead of taking individual remarks, there is a panel of expert individuals (C. Powell, 2003). Delphi method-: The Delphi method is basically a questionnaire sequence or rounds which is distrib-uted over a controlled feedback system in order to maximise the most reliable consensus of a group of experts H. A. Linstone and M. Turoff, 1975). Statistical Methods Using historical parameter data, rating curves’ key points are developed which are dependent on statis-tical characteristics such as mean values and various other quantities calculated over a long period of time. This method does not include the expert opinions. Use of Water Quality Standards- Another way for the development of the subindex functions is through the standard for water quality. Here, each water quality attribute is assigned a value denoting its rating between a scale of 0 and 100, based on national standards (India, Malaysia etc) or international standards such as WHO, United States Environmental Protection Agency (USEPA) (Standard, I. 2013, Ponsadailakshmi et. al. 2018). 3.2 CCME The CCME (Canadian Council of Ministers of the Environment) WQI gives an adaptable tool which provides a flexible index template (T. Hurley, R. Sadiq, and A. Mazumder, 2012) and is mainly composed of three factors (A. Lumb, et. al. 2006). The CCME is well documented and provides detailed procedures for index calculation and application (2022). The three factors are the following: Scope Frequency Amplitude. 4 Related Work WQI is an important indicator of water quality, and multiple studies have been conducted to compute its value. Kangabam et al. chose Loktak Lake as the area of study to assess the WQI from physicochem-ical water parameters (R. Das Kangabam et. al. 2017). Judran and Kumar mathematically developed the WQI for Al-Gharraf River Thi Qar of Iran, and found the quality of the river water to be poor for drinking (N. H. Judran and A. Kumar, 2020). Şener et al. analysed WQI of Aksu River, located in Turkey, to evaluate it for drinking purposes using thirteen water quality parameters and found that environmental pollu-tants affect the river’s quality negatively (Ş. Şener, E. Şener, and A. Davraz, 2017). Kamboj and Kamboj as-sessed the groundwater quality in the Indian district of Haridwar through the calculation of WQI using arithmetic mean method, and found the quality to be suitable for drinking (N. Kamboj and V. Kamboj, 2019). Wagh et al. classified the groundwater suitability of Kadava river of India using the CCME WQI method using eleven physicochemical parameters, and found the samples to fall in the marginal cate-gory for drinking purposes (V. M. Wagh et. al. 2017). Numerous studies have been conducted to study the quality of water using WQI. In recent times, Super-vised Machine Learning approaches and Artificial Intelligence methods have been employed to predict WQI. Ahmed et al. collected data from Pakistan Council of Research in Water Resources to estimate WQI using eight regression algorithms, as well as classify samples into predefined Water Quality Clas-ses using ten classification algorithms (U. Ahmed et. al. 2020). When they used four parameters (temper-ature, turbidity, pH and total dissolved solids (TDS)) to predict water quality, the best performing algo-rithm was found to be Gradient Boosting while Polynomial Regression was more efficient when TDS was dropped and only three parameters were used. Multi-layer perceptron performed the best in pre-dicting the WQC. Sillberg et al. classifies quality of water for Chao Phraya River, which is the largest river in Thailand, using a machine learning-based approach coined AR-SVM, which integrates support vector machine with attribute-realisation (C. Sillberg et. al. 2019). The AR-SVM approach resulted in an accuracy of 0.94. Hameed et al. applied Artificial Neural Network algorithms (back propagation neural networks models and radial basis function neural network (RBFNN)) for WQI prediction in Malaysia, finding that RBFNN proved to be effective for assessing the quality of water (M. Hameed, et. al. 2017). This domain of research has also been explored substantively in the Indian subcontinent. Hassan et al. used Indian water quality data collected from Kaggle and applied five machine learning classifiers, Neural Network, Random Forest, Multinomial Logistic Regression, Support Vector Machine, and Bagged Tree Model, to predict water quality (M. M. Hassan et al., 2021, Citaristi, I. 2022). Multinomial Logistic Regression gave the best accuracy score of 99.83%, while Support Vector Machine gave the lowest accuracy score of 96.98%. Aldyhani et al. implemented NARNET which is a nonlinear autoregressive neural network, along with long short-term memory (LSTM) deep learning algorithm for WQI prediction, while support vector machine (SVM), K-nearest neighbour (K-NN), and Naive Bayes, were used for forecasting WQC, on a dataset from Kaggle collected in India (T. H. H. Aldhyani et. al. 2020). The results showed that on the basis of R value, NARNET performed better than LSTM while SVM achieved highest accuracy for WQC prediction. Kadam et al. applied artificial neural network (ANN) and multiple linear regression (MLR) for prediction of WQI on data collected from Shivganga watershed, and found a higher level of precision in the ANN model (A. K. Kadam et. al. 2019). 5 Methodology 5.1 Data Acquisition The data used in Fig. 1 as shown in study was collected from different locations of Noida and some parts of Delhi by the Department of Computer Science, JSS Academy of Technical Education, Noida. A total number of 51 samples were collected. One litre of water sample was collected from each site and further tests were conducted in the laboratories of JSS Academy of Technical Education, Noida to find the properties of samples (K. Banerjee et al.,2022). For the collection of samples, the methods prescribed in the American Public Health Association manual were followed, and samples were collected from different sources such Government hand pumps, General hand pumps and Borwells, also depicted in Fig. 2 , (K. Banerjee et. al. 2021). The acquired data has different parameters related to water, also shown in Fig. 3 and the standard values are provided in Table 2 and Table 3 . These parameters are: PH pH denotes hydrogen ion concentration and is known as the “potential of hydrogen” or the “power of hydrogen” for aqueous solutions. It is the quantitative measure of the acidity or alkalinity of liquid solu-tions on a range of 0 to 14, with 7 being neutral. pH reading of less than 7 indicates acidity, while a reading greater than 7 indicates the solution is basic. Turbidity Turbidity is caused by suspension of fine matter (like organic matter, inorganic matter, clay and silt), and gives a measure to ascertain to what degree water loses its transparency because of the presence of such particles. It is measured in Nephelometric Turbidity Units or NTU through a nephelometer or a turbidimeter. It causes the cloudiness, haziness or milkiness of a fluid. Hardness Hardness is caused due to soluble bicarbonates, chlorides and sulphates of magnesium and calcium in the water molecules. It is water with a high amount of mineral content. This type of water does not easi-ly form lather with soap, and therefore hardness can also be defined as the capacity of water to precipi-tate soap. Dissolved Oxygen The dissolved oxygen, or D.O., in water is the amount of oxygen present in it. It gets mixed with water bodies through the atmosphere, water plants and algae such as phytoplankton. The presence of sun-light is important for these aquatic lives to produce oxygen through photosynthesis. The running water streams are found to have more dissolved oxygen than a stationary pond or a lake. Biochemical Oxygen Demand Biochemical Oxygen Demand, also known as B.O.D., is the oxygen amount microorganisms and dif-ferent bacterias require to break down organic matter in water. Chemical Oxygen Demand Chemical Oxygen Demand, also known as C.O.D., is the consumed oxygen amount to chemically oxi-dise inorganic and organic matter present in water. Table 2 BIS accepted and permissible values for water quality parameters Parameters Requirement (Acceptable Limit) Permissible Limit in the Absence of Alternate Source pH value 6.5–8.5 No relaxation Turbidity (NTU) Max 1 5 Total Dissolved Solids (mg/l) Max 500 2000 Total Hardness (mg/l) Max 200 600 Total Alkalinity (mg/l) Max 200 600 Chloride (mg/l) Max 250 1000 SOURCE : BIS (IS:10500) Table 3 WHO standard values for water quality parameters Parameters Units WHO Standard T ℃ 12–25 pH – 6.5–8.5 EC ㎲/cm 400 Cl- mg/l 250 TDS mg/l 500–1000 TH mg/l 500 SOURCE : WHO Guidelines for drinking water quality 5.2 Data Preprocessing Data preprocessing is a critical step in data analysis, and prepares the data for further processing in Fig. 4 . To ready the dataset, the non-affecting parameters and redundant parameters were removed, missing val-ues were handled and the dataset was normalised. In our study, the data was extracted manually and therefore it had many missing and ambiguous val-ues. The missing values were removed before splitting the dataset into training and testing subsets. At-tributes related to location such as Latitude, Longitude and Elevation were not a part of the final dataset because these attributes did not contribute to the model performance as they had very little correlation to the target value. Any other technique to replace the missing values such as mean value substitution was not employed because values such as pH, Turbidity, Temperature etc. were unique to different locations and any generalisation such as mean would have made the model skewed. SOURCE: (Banerjee et al. 2021, (2012) 5.3 Processing Computation To calculate WQI the following four steps are followed, also shown in Fig. 5 (D. Sutadian et. al. 2016): Parameter Selection Assigning of Weights Quality Rating Aggregation of sub indices Parameter Selection- The parameters used for calculating WQI are pH, Turbidity (NTU), Hardness (mg/L), D.O. (mg/L), B.O.D. (mg/L), C.O.D. (mg/L). Other parameters which provided information regarding the location of the test samples were not considered because they did not directly affect the quality of water. Although, location could be a huge determining factor if we draw comparisons between water from different cities or national states, in our case the study area was limited to only a small part of Gautam Buddha Nagar. 2. Assigning of weights The first important step for calculating the wqi is to calculate the relative weights of the selected parameters. These parameters are then assigned these relative weights (RW) against their original weights as given in Eq. (1) (R. Shabbir and S. S. Ahmad, 2015). RW = \(\:\frac{AWI}{{\sum\:}_{i=1}^{n}\:AW1}\) In the formula above, RW: relative weight AW: assigned weight of each parameter n: number of parameters. The relative weight is the ratio of the assigned weight of the parameter to the total sum of all the weights of the parameters as shown in Table 4 . It presents standard parameter values and their corresponding relative weights for evaluating a system. Key factors such as pH, chloride (Cl⁻), and total dissolved solids (TDS) have the highest relative weights (0.2431), indicating their significant influence compared to other parameters like temperature (T) and electrical conductivity (EC). The standard f Suthaharanor the water quality parameters in this study was based on BIS (Indian Standard Specifi-cation for Drinking Water 2012) (2022) and World Health Organisation (2011) (2017) recommendations. Table 4 Standard values and relative weights for water quality parameters Parameters Standard value Relative Weight T 8.5 0.1430 pH 5 0.2431 EC 200 0.0060 Cl- 5 0.2431 TDS 5 0.2431 TH 10 0.1215 Quality Rating-: The quality rating for any parameter is calculated as the ratio of its concentration to the standard values. It can be calculated using the formula given in Eq. (2) (Srinivas, Kumar et. al. 2017) $$\:\text{Q}\text{i}\:=\:\left[\frac{Ci\:-\:Vi}{Si\:-\:Vi}\right]\:\times\:\:100$$ Here, Qi: quality rating, Ci: computed value of the water parameter, Si: standard WHO or BIS value, Vi: ideal value which is 0 except for pH = 7.0 and D.O. = 14.6. The above equation can be simplified and written as Eq. (3). $$\:\text{Q}\text{i}\:=\:\left(\frac{Ci}{Si}\right)\:\times\:\:100$$ It should be noted that for the parameters pH and D.O. The quality rating is calculated using Eq. (2) because for these parameters the value for Vi is 7.0 and 14.6 respectively. Therefore, we cannot sub-stitute its value by 0. Aggregation of sub-indices- This is the last step for calculating WQI. The sub-indices are products of the relative weights and the quality rating. The products of all the values are then summed together in Table 5 , to give WQI using Eq. (4) (R. E. Rabeiy, 2018, Singh & Khan 2019). SIi = W r \(\:\times\:\) Q i WQI = \(\:{\sum\:}_{\:}^{\:}\:\) SI i The WQI values computed are further classified based on the range and ratings. Table 5 Calculated WQI values using the acquired parameter data Location pH Turbidity Hardness C.O.D D.O B.O.D WQI Pump No. H-1, Sector 15A 8.74 9 295 16 6.4 18 188.9849 Asagarpur Jagir Vilage, Sector 128 8.2 0 400 10.66 7.6 15.3 117.7359 Hindustan Petroleum, Near Jaypee Hospital 8 0 465 10.66 6.8 18 131.1810 Balaji Temple, Sector 126 8.35 0 430 10.66 8.4 14.4 112.8551 Ankit Nursery, Sector 131 8.04 1 425 5.33 5.6 16.2 124.1108 Green Beauty Farm, Sector 135 8.36 3 425 10.66 6.8 18 149.0787 Yakootpur, Sector 167 8.68 3 190 10.66 6 15.3 140.3131 Gulavali, Sector 162 8.2 56 340 5.33 5.6 19.8 410.3118 Jhatta Village, Sector 159 8.6 4 205 16 5.6 19.8 173.8431 Badauli, Sector 154 8.23 2 410 16 6.4 17.1 146.0595 Kambuxpur Derin Village, Sector 155 8.63 2 295 16 7.6 16.2 142.1085 Gujjar Derin, Kambuxpur, Sector 155A 8.51 3 270 53.33 6.4 16.2 194.1677 Kondali Bangar, Sector 149 8.33 3 240 58.66 8 14.4 186.0353 Garhi Samastpur, Sector 150 8.06 5 470 53.33 3.2 26.1 256.4511 Momnathal, Sector 150 8.9 3 235 32 6.4 18.9 184.9794 Shafipur Village, Sector 148 7.95 17 265 42.6 5.6 23.4 280.1487 Mohiyapur Village,Sector 163 7.57 2 485 37.33 5.6 21.6 189.8300 Nalgadha, Sector 145 7.87 2 510 53.33 6.8 19.8 200.4244 Ideal Industrial Training Institute, Sector 143 7.52 38 515 21.33 4.8 23.4 355.8180 Shahdara, Sector 141 7.1 4 1780 16 6.4 23.4 179.8069 Hindon Flood Plain, Kulesara, Sector 140 8.27 6 335 37.33 6.8 18 194.9539 Allahabad, Sector 86 7.79 5 1645 10.66 6.8 20.7 170.2049 Sai Dham Colony, Sector 88 7.94 1 395 48 6.8 19.8 189.4008 Kakrala Village, Sector 80 8.11 9 370 42.66 6 19.8 225.3792 Gijhor Village, Sector 53 7.54 1 775 26.66 8.4 15.3 134.8688 Sarfabad, Sector 73 8.28 1 445 10.66 8 17.1 131.2370 Sorkha Village, Sector 118 7.67 19 580 21.33 6.4 18.9 239.1284 Pumping Station 3, Sector 71 7.6 1 535 16 5.2 22.5 164.8662 Pump House, Sector 122 7.65 0 600 32 6 20.7 169.3493 19, Block H,Sector 116 7.97 1 935 5.33 6 21.6 150.2370 Baraula Village, Sector 49 7.54 21 545 58.66 6.4 19.8 297.2613 Pumping Station, Sector 35 7.83 1 550 0 6.8 20.7 134.8510 Pumping Station 3, Sector 34 7.58 6 830 21.33 8 15.3 154.2630 Peerbabaji, Sector 144 8.45 20 765 10.66 6.8 18 233.6300 Dallupura Village, Sector 164 8.65 0 360 32 8 17.1 155.5846 Dostpur, Mangrauli, Sector 167 8.78 50 205 53.33 6.4 19.8 442.5794 Nangli Village, Sector 134 8.55 0 580 64 7.2 18 200.6006 Bakhtawarpur, Sector 127 8.31 6 325 5.33 6 20.7 171.5604 Sultanpur Village, sector 128 7.75 0 690 10.66 5.6 20.7 145.6486 Shahpur, Sector 131 7.91 0 860 16 6 19.8 148.7928 Sadarpur, Sector 45 8.01 7 585 0 5.2 21.6 174.2757 Chhalera, Sector 44 7.78 0 1055 21.33 7.2 16.2 134.0816 Sanatan Temple, Sector 41 7.97 1 1745 48 3.6 20.7 206.2695 Shiv Mandir, Sector 31 7.74 0 695 26.66 6 21.6 168.3809 Nagla Charan Dass, Noida Phase-2 7.95 4 840 32 6 21.6 196.7644 Nursery, Sector 104 7.85 1 685 21.33 4.4 23.4 180.5868 Pumping Station, Sector 80 8.12 0 290 10.66 4.4 25.2 172.8801 Shiv Mandir, Sector 93 7.8 1 1230 42.66 5.2 22.5 201.2927 Salarpur Village, Sector 102 7.48 0 750 32 4.8 23.4 184.3517 Garhi Chaukhandi, sector 121 8.01 0 560 37.33 5.2 21.6 185.5412 Pumping Station, Block-G, Sector 63 7.92 0 590 32 4.4 24.3 193.4495 Correlation- After computing the WQI, relevant features needed to be selected for further modelling to predict the WQI value in Fıg. 6. For this purpose, the correlation was calculated between different variables. WQI showed a strong correlation with Turbidity, whereas other parameters did not show strong correlation with the target variable. GridSearch- GridSearch is an exhaustive method for finding the best parameter values for a given algorithm by run-ning every combination and comparing the score for each parameter value (M. M. Ramadhan et. al. 2017). All the algorithms were run through the cross-validation method GridSearchCV to achieve the best accuracy through every combination of hyperparameters and the one with the best performance was chosen. The cross validation folds were set to 3 for this study. 5.4 Algorithms/Models Machine Learning is a rapidly growing field of interest in Computer Science that detects meaningful patterns in a given set of data to perform data analysis, data extraction and data mining (O. F.y et al., 2017). In our study we applied supervised machine learning algorithms as it categorises data using in-formation already available in the form of training data (A. Singh, et. al. 2016). The supervised machine learning algorithms for predicting a categorical outcome is termed as classification whereas when an algorithm predicts a continuous outcome it is called regression (T. Jiang, et. al. 2020). The aim of super-vised machine learning is the prediction of target variable y ∈ Y based on a set of independent variables x ∈ X (C. Crisci et. al. 2012). The dataset in our study was split into training and testing subsets in which the training set consisted of 75% of the total data and the rest 25% was the testing data. The split was performed using the train_test_split feature present in the sklearn library of python. Machine Learning approaches are rigorously applied in almost every sector and undoubtedly the most applicable area is ecology (P. A. Flach, 2001). Some other areas of application include healthcare (K. Shailaja, et. al. 2018), fatigue detection (N. Yadav et. al. 2020), DNA sequence compression (K. Banerjee and R. A. Prasad, 2017) and stress analysis (T. Sharma T. Sharma et. al. 2020). Before implementing re-gressors or classifiers, it is important to check for any noise in the data which can be handled by using statistical methods such as data interpolation (K. K. Hiran et. al. 2017). The machine learning algorithms employed in this study for the prediction of WQI are the following: Linear Regression- Some of the important questions addressed in this study explore the idea of any correlation between different water quality metrics, i.e. the dependent and the independent variables. To find this correlation in order to devise a model, which gives acceptable prediction and at the same time uses less parameters, Linear Regression was administered. Regression is a statistical technique in which a dependent variable, also known as a response variable (S. Puntanen, 2010), paired with one or more than one independent variable is used to determine the re-lationship between them (G. K. Uyanık and N. Güler, 2013). Regression can be broadly classified into two categories (E. de A. L et. al. 2004): Univariate regression– uses only one independent variable Multivariate regression– uses more than one independent variable In this study we aim to decrease the number of independent variables as much as possible while main-taining an acceptable accuracy. The general form of the linear equation for the algorithm is given in Eq. (5) (X. Su, et. al. 2012). y = β0 + β1 X1 + β2 X2 +…. + β3 X3 + βn Xn + ε where, y: dependent variable X: independent variable βn: regression coefficients ε: error. The input data is fed to the algorithm equation as training data such that X is entered as the independ-ent variable for a random value of β which maps it with the actual value of y. These obtained values are then later used for unseen test cases to make predictions. As an extension to normal linear regression, multivariable linear regression can be applied which has multiple iterations of the same linear regression data which improves the performance of the model (T. Reddy et. al. 2020). The mean square error in Linear Regression is geometrically the vertical distance between the regression line and the actual y-values which can be visualised, as depicted in Fig. 7 as follows: (2021) Source: Applied Linear Regression, Sanford Weisberg, University of Minnesota Decision Tree- As the name suggests, Decision Tree Regressors differ from the more common and popular Decision Trees or Decision Tree Classifiers but they have some similarities as well (F. Mola, 1998). The idea be-hind this model is to break down complex structures into simpler binary problems known as decisions (M. Xu et. al. 2005). It recursively splits the training set’s feature space to create a tree structure where the segmentation is applied based on some simple rules and the goal is to deduce the set of these decision rules for a reliable hierarchical model (J. R. Quinlan, 1996). Until the tree reaches a minimum sample split parameter, it continues to split the node into two, which is done keeping in mind that the depth of the tree is directly proportional to variance and inversely proportional to bias (Rahul, A. Gupta, et. al. 2021). The main benefit of using a decision tree over other modelling techniques is that it portrays more interpretable and logical rules (G. K. F. Tso and K. K. W. Yau, 2007). Another benefit of using the Decision Trees is that features with maximum information are selected automatically rejecting the remaining and thus increasing the efficiency which also takes care of the Hughes phenomenon (G. Hughes, 1998) also known as the curse of dimensionality (J. H. Cho and P. U. Kurup, 2011). In our study the predicted out-come is a continuous quantity and not a class for which the level-wise classified data is produced by using the gini impurity which is calculated using the information gain to find the optimal split location (S. Suthaharan et. al. 2016, Lerman & Yitzhaki 1984). The equation for the Gini index is given in Eq. (6) (Gastwirth, 2011). Gini index = 1 − Σp 2 Here, p i is the probability of the occurrence of the event p i . To calculate Gini index we need Information Gain and to calculate information gain we require Entropy which is calculated as given in Eq. (7) and Eq. (8)( Y. Zhou et. al. 2020): E(S) = \(\:{\sum\:}_{i=1}^{c}\:-p\) i log 2 p i IG = Entropy(before) \(\:-{\sum\:}_{j=1}^{k}\:Entropy(j,\:after)\) Random Forest- The collection of multiple decision tree predictors is known as a random forest {h(x, Θk ) k = 1,…,k} where x is the input covariate vector and Θk are the identically distributed independent random vectors, which in case of a regression problem is the unweighted average over the function as given in Eq. (8) (M. R. Segal, 2022). h(x) = (1/K) \(\:{\sum\:}_{k=1}^{k}\:h(x,\:\varTheta\:k\:)\) This equation resamples the data without deletion therefore some data may be used multiple times in the training and some data might never be observed at all which results in greater stability– this process is called bagging, a contraction of bootstrap-aggregating (L. Breiman, 2017). The samples which were ignored in the bagging process for the k-th tree are then included in the out-of-bag subset which could be used by the kth tree for performance evaluation (J. Peters et al., 2007). In this study our approach was to use the Random Forest algorithm present in the Scikit-Learn module of Python. The manual backward elimination of the feature set reduces the overall performance score. The node significance is calculated using the Gini importance, assuming that its a binary tree, formula for which is given by Eq. (10) (1984). f ii = \(\:\frac{{\sum\:}_{j:node\:j\:splits\:on\:feature\:i}^{\:}\:{ni}_{j}}{{\sum\:}_{k\in\:all\:nodes}^{\:}\:{ni}_{k}}\) Here, fi sub(i) = the importance of feature i ni sub(j) = the importance of node j Support Vector Regressor- SVM performs dichotomy functions for classification for multi-dimensional feature vectors (V. N. Vapnik and A. Chervonenkis, 1964). It was originally meant for linear classification, and updated to accommodate non-linear classification and finally extended to include regression problems as well (C. Cortes and V. Vapnik, 1995). For regression problems, Support Vector Regressor (SVR) transforms the input data into a higher dimension space in which the classes can be classified with the help of the hyperplane, such as (H. Talabani and E. Avci, 2018) {〖X_n^ }〗_n^N = 1 with N total samples, X is a vector of L input features {〖Y_n^ }〗_n^N = 1 and with yn ∈ {−1,1}, The SVM regression model can be defined as f(x) = Wt ϕ(x) where where ϕ : x → ϕ(x) ∈ ℝH (V. Rodriguez-Galiano et. al. 2015). The optimisation of the SVM model is subject to soft margin constraints therefore this optimisation problem can be overcome by using the kernels. There are many kinds of kernels that exist which are given from Eq. (11) to Eq. (14) (H. Talabani and E. Avci, 2018). K linear (x, x′) = x, x′ K Polynomial (x, x′) = (γxx′ + r)ρ K RBF (x, x′) = exp (–γ ||x – x′ ||2) K Sigmoid (x, x′) = tanh(γxx′ + r) 6. Results and Discussion The analysis of the physicochemical properties of the groundwater gives an insight of the quality of wa-ter in the Gautam Buddha Nagar. This study finds a suitable machine learning model which can esti-mate WQI with one input variable and identifies the most valuable parameters contributing to the WQI prediction. The Turbidity of the groundwater in the study area ranges from 0 to 56, indicating that the water under the ground is unpredictable by location. There is very little correlation between the different water qual-ity metrics, for instance the Turbidity was noted highest at Gulavali, Sector 162 with a value of 56 NTU, but the pH of the water at the same place was found to be 8.2. Similarly the site Pumping Station, Sector 80 has 0 Turbidity and the pH is 8.12. This indicates that there is very low correlation between the input parameters, and the same was confirmed through correlation analysis. The minimum observed value of pH in this study is 7.1 and the maximum was found to be 8.9, showing that the groundwater is mostly neutral or sub alkaline having values close to the BIS standard for pH. The average value of hardness, C.O.D. and B.O.D. in this study are 588.23, 25.93 and 19.64 respectively, which are all greater than the standard BIS values. These values directly affect the quality of water and thus result in poor water quality in the study area. The accuracy score of the model when all 6 parameters were used for the training dataset was 0.9999 for Linear Regression, 0.6637 for Decision Tree ('maximum depth': 50, 'minimum sample leaf': 2, 'mini-mum samples split': 9), 0.9013 for Random Forest ('maximum depth': 12, 'minimum sample leaf': 2, 'minimum samples split': 4, 'number of estimators': 20), and 0.9997 in case of Support Vector Re-gressor ('kernel': Linear). When all the input parameters are used, the accuracy is found to be the highest, with similar scores for both Linear Regression and Support Vector Regressor. Using this model will greatly reduce the com-plexity in computing the WQI when all water quality parameters are available as shown in Fig. 8 . As the cost of measuring some water quality parameters is expensive, reducing the parameter size will prove to be helpful in indicating groundwater quality in a reliable, feasible and cost-effective manner. After analysis of the correlations between WQI and the input parameter set, Machine Learning models were built using single features. The score in Fig. 9 , when only Turbidity was used as an input parameter was 0.810 for Linear Regression, 0.571 for Decision Tree ('maximum depth': 50, 'minimum sample leaf': 1, 'minimum samples split': 9), 0.670 for Random Forest ('maximum depth': 14, 'minimum sample leaf': 2, 'minimum samples split': 3, 'number of estimators': 19) and 0.769 in case of Support Vector Regressor ('kernel': Linear). Figure 10 in the score when only pH was used as an input parameter was − 0.172 for Linear Regression, -0.361 for Decision Tree ('maximum depth': 50, 'minimum sample leaf': 3, 'minimum samples split': 11), -0.093 for Random Forest ('maximum depth': 13, 'minimum sample leaf': 4, 'minimum samples split': 2, 'number of estimators': 19) and − 0.048 in case of Support Vector Regressor (with rbf kernel). According to Fig. 11 , the score when only B.O.D was used as an input parameter was − 0.003 for Linear Regression, 0.111 for Decision Tree ('maximum depth': 50, 'minimum sample leaf': 2, 'minimum samples split': 9), 0.245 for Random Forest ('maximum depth': 14, 'minimum sample leaf': 3, 'minimum samples split': 2, 'number of estimators': 21) and − 0.042 in case of Support Vector Regressor (with rbf kernel) Figure 12 in the score when only D.O. was used as an input parameter was − 0.069 for Linear Regression, -0.140 for Decision Tree ('maximum depth': 50, 'minimum sample leaf': 2, 'minimum samples split': 11), -0.022 for Random Forest ('maximum depth': 14, 'minimum sample leaf': 3, 'minimum samples split': 4, 'number of estimators': 21) and − 0.029 in case of Support Vector Regressor ('kernel': Linear). The score, according to Fig. 13 , when only Hardness was used as an input parameter was − 0.021 for Linear Regression, -0.141 for Decision Tree ('maximum depth': 50, 'minimum sample leaf': 2, 'minimum samples split': 11), -0.019 for Random Forest ('maximum depth': 14, 'minimum sample leaf': 3, 'minimum samples split': 4, 'number of estimators': 21) and − 0.049 in case of Support Vector Regressor ('kernel': Linear). The score in Fig. 14 , when only C.O.D. was used as an input parameter was − 0.021 for Linear Regression, -0.141 for Decision Tree ('maximum depth': 50, 'minimum sample leaf': 3, 'minimum samples split': 10), -0.002 for Random Forest ('maximum depth': 13, 'minimum sample leaf': 3, 'minimum samples split': 3, 'number of estimators': 21) and − 0.049 in case of Support Vector Regressor ('kernel': Linear). 7 Conclusion The study reviewed and calculated WQI values for 51 locations in Gautum Buddha Nagar and revealed that the groundwater is highly contaminated and not fit for human consumption. The calculated WQI values were far higher than the standard and permissible limits due to high sub-indices values of B.O.D., C.O.D. and Turbidity, which range from 112.855 to 442.579. The calculation of WQI is a lengthy and costly method therefore the aim of this study was to develop a Machine Learning model which can accurately determine the WQI using just one input parameter. Our findings showed that the best performing supervised learning algorithm was Linear Regression with a score of 0.8105 when using only Turbidity as the input variable and the worst performing algorithm was Decision Tree with a score of -0.141 using Hardness as the input variable. The Machine Learning models will not only help in estimating WQI with a single water quality parameter but also aid future studies to further develop in this area. The study strongly suggests tracking the pollution sources in order to prevent water in the study area from getting contaminated any further. Since Gautam Buddha Nagar is home to many large scale facto-ries and industries, their presence might be the primary source of contamination. Prevention from dete-rioration requires frequent inspections and monitoring of the water quality, for which WQI plays an important role. Declarations Funding: No Funding support is provided for this paper publication. Competing Interests: No benefits in any form have been received or will be received from a commercial party related directly or indirectly to the subject of this article. All authors declare no conflict of interest for this article. Author's Contribution: Kakoli Banerjee: Conceptualization, Literature Review, Drafting the manuscript. Pradeep Kumar: Methodology, Data Analysis, Review & Editing. Ajay Kumar: Literature Review, Proofreading. Kamal Upreti: Data Collection, Review & Editing, Final Approval. Shubham Mahajan: Visualization, Literature Review. Mohammad Shahnawaz Nasir: Methodology, Data Analysis, Review & Editing. Ethics Approval and Consent to participate: No Participation of humans take place in this implementation process. Consent to Publish: All authors are showing their consent to publish this paper . Data Availability Statement: The data that support the findings of this study are available from the corresponding author upon reasonable request. References Abbasi T. and Abbasi,S. A. (2012). Water-Quality Indices, Water Quality Indices . 353–356. doi: 10.1016/b978-0-444-54304-2.00016-6. Ahmed, U., Mumtaz, R., Anwar, H., Shah, A. A., Irfan, R., & García-Nieto, J. (2020). Efficient Water Quality Prediction Using Supervised Machine Learning, Water , 11(11), 2210. Akhtar, N., Ishak, M. I. S., Ahmad, M. I., Umar, K., Md Yusuff, M. S., Anees, M. T., ... & Ali Almanasir, Y. K. (2021). Modification of the Water Quality Index (WQI) Process for Simple Calculation Us-ing the Multi-Criteria Decision-Making (MCDM) Method: A Review, Water , 13(7). 905. doi: 10.3390/w13070905. Aldhyani T. H. H., Al-Yaari M., Alkahtaniv H., and Maashiv M., (2020). Water Quality Prediction Using Artifi-cial Intelligence Algorithms, Appl. Bionics Biomech., 6659314, Dec. 2020. Banda T. D. and Kumarasamy M. V., (2020). Development of Water Quality Indices (WQIs): A Review, Polish Journal of Environmental Studies, 29(3). 2011–2021. doi: 10.15244/pjoes/110526. Banerjee K. and Prasad R. A., (2017). Reference based inter chromosomal similarity-based DNA sequence compression algorithm, in 2017 International Conference on Computing, Communication and Automation (ICCCA) , 234–238. Banerjee K., Santhosh Kumar M. B., Tilak L. N., and S. Vashistha, (2021). Analysis of Groundwater Quality Using GIS-Based Water Quality Index in Noida, Gautam Buddh Nagar, Uttar Pradesh (UP), India, Lecture Notes in Electrical Engineering . 171–187. doi: 10.1007/978-981-16-3067-5_14. Banerjee, K., Bali, V., Nawaz, N., Bali, S., Mathur, S., Mishra, R. K., & Rani, S. (2022). A Machine-Learning Approach for Prediction of Water Contamination Using Lat-itude, Longitude, and Elevation, Water, 14(5), 728. Bhardwaj R. M., (2005). Water Quality Monitoring in India - Achievements and Constraints, IWG-Env Joint Work Session on Water Statistics . Breiman, L. (2017). Classification and Regression Trees. Routledge . Brown M., McCLELLAND N. I., Deininger R. A., and O’connor M. F., (2013). A water quality index – crashing the psychological barrier, Advances in Water Pollution Research . 787–797. doi: 10.1016/b978-0-08-017005-3.50067-0. Bui, D. T., Khosravi, K., Tiefenbacher, J., Nguyen, H., & Kazakis, N. (2020). Improving prediction of water quality indices using novel hybrid machine-learning algorithms, Sci. Total Environ., 721, 137612. Canada C. E. and Canada. (2003). National Guidelines and Standards Office, Canadian Water Quality Guide-lines for the Protection of Aquatic Life: Nitrate Ion. National Guidelines and Standards Office. Cho, J. H., & Kurup, P. U. (2011). Decision tree approach for classification and dimensionality reduction of electronic nose data, Sensors and Actuators B: Chemical, 160(1). 542–548. doi: 10.1016/j.snb.2011.08.027. Citaristi, I. (2022). United Nations Office for Outer Space Affairs—UNOOSA. In The Europa Directory of International Organizations, 247-248. Routledge. Cortes C. and Vapnik, V. (1995). Support-vector networks, Mach. Learn ., 20(3), 273–297. Crisci, C., Ghattas, B., & Perera, G. (2012). A review of supervised machine learning algorithms and their applications to ecological data, Ecological Modelling , 240. 113–122. doi: 10.1016/j.ecolmodel.2012.03.001. Das Kangabam, R., Bhoominathan, S. D., Kanagaraj, S., & Govindaraju, M., (2017). Development of a water quality index (WQI) for the Loktak Lake in India, Applied Water Science, 7(6), 2907–2918. de A. Lima Neto, E., de Carvalho, F. A., & Tenorio, C. P. (2004). Univariate and Multi-variate Linear Regression Methods to Predict Interval-Valued Features, Lecture Notes in Computer Science. 526–537. doi: 10.1007/978-3-540-30549-1_46. Edition, F. (2011). Guidelines for drinking-water quality. WHO chronicle, 38(4), 104-8. Edition, F. (2011). Guidelines for drinking-water quality. WHO chronicle, 38(4), 104-8.. Flach P. A., (2001). On the state of the art in machine learning: A personal review, Artificial Intelligence, 131(1–2). 199–222. doi: 10.1016/s0004-3702(01)00125-4. Gafoor, A., Rauf, A., Arif, M., & Muzaffar, W. (1994). Chemical composition of effluents from different industries of the Faisalabad Pakistan. J. Agri. Sci, 33, 73-74. Gastwirth, (2011). The estimation of the Lorenz curve and Gini index, Rev. Econ. Stat . Hameed M., Sharqi S. S., Yaseen Z. M., Afan H. A., A. Hussain, and A. Elshafie, (2017). Application of artifi-cial intelligence (AI) techniques in water quality index prediction: a case study in tropical region, Malaysia, Neural Comput. Appl ., 28(1), 893–905.00 Hassan, M. M., Hassan, M. M., Akter, L., Rahman, M. M., Zaman, S., Hasib, K. M., ... & Mollick, S, (2021). Efficient prediction of water quality index (WQI) using machine learning algo-rithms, Human-Centric Intelligent Systems , 1(3–4), 86. Hiran, K. K., Khazanchi, D., Vyas, A. K., & Padmanaban, S., (2017). Machine Learning for Sustainable De-velopment. Horton R. K., (2015). An Index Number System for Rating Water Quality, Journal of the Water Pollution Control Federation, 37(3), 300–306. Hughes, G. (1998). On the mean accuracy of statistical pattern recognizers,” IEEE Transactions on Infor-mation Theory, 14(1). 55–63. doi: 10.1109/tit.1968.1054102. Hurley T., Sadiq R., and Mazumder A., (2012). Adaptation and evaluation of the Canadian Council of Min-isters of the Environment Water Quality Index (CCME WQI) for use as an effective tool to characterize drinking source water quality, Water Res., 46, 11, 3544–3552. Ibrahim and Salmon S., (2020). Chemical composition of Faisalabad - city [Pakistan] sewage effluent: I. Nitrogen, phosphorus and potassium contents, J. Agric. Res . Jha M. K., Shekhar A., and Jenifer M. A., (2020). Assessing groundwater quality for drinking water supply using hybrid fuzzy-GIS-based water quality index, Water Res., 179, 115867. Jiang T., Gradus J. L., and Rosellini A. J., (2020). Supervised Machine Learning: A Brief Primer, Behav. Ther., 51(5), 675–687. Judran H. and Kumar A., (2020). Evaluation of water quality of Al-Gharraf River using the water quality index (WQI), Modeling Earth Systems and Environment, 6(3), 1581–1588. Kadam, A. K., Wagh, V. M., Muley, A. A., Umrikar, B. N., & Sankhua, R. N. (2019). Prediction of water quality index using artificial neural network and multiple linear regression modelling approach in Shivganga River basin, India. Modeling Earth Systems and Environment, 5, 951-962. Kamboj N. and Kamboj V., (2019). Water quality assessment using overall index of pollution in riv-erbed-mining area of Ganga-River Haridwar, India, Water Science, 33(1). 65–74. doi: 10.1080/11104929.2019.1626631. Lerman, R. I., & Yitzhaki, S. (1984). A note on the calculation and interpretation of the Gini index. Economics Letters, 15(3-4), 363-368. Linstone H. A. and Turoff M. (1975). The Delphi Method: Techniques and Applications. Westview Press . Lumb A., A. L., Halliwell, D., & Tribeni Sharma, T. S., (2006). Application of CCME Water Quality Index to monitor water quality: a case study of the Mackenzie River Basin, Canada,” Environ. Monit. Assess., 113(1–3), 411–429. Lumb A., A. L., Halliwell, D., & Tribeni Sharma, T. S.,, (2011). A Review of Genesis and Evolution of Water Quality In-dex (WQI) and Some Future Directions, Water Quality, Exposure and Health, 3(1). 11–24. doi: 10.1007/s12403-011-0040-0. Mara, D. D., Cairncross, S., & World Health Organization. (1989). Guidelines for the safe use of wastewater and excreta in agriculture and aquaculture: measures for public health protection. World Health Organization . . Mateo-Sagasta, J., Zadeh, S. M., Turral, H., & Burke, J. (2017). Water pollution from agriculture: a global review. Executive summary. Rome, Italy: FAO Colombo, Sri Lanka: International Water Management Institute (IWMI). CGIAR Research Program on Water, Land and Ecosystems (WLE).,. Meyer M. A. and Booker J. M., (2001). Eliciting and Analyzing Expert Judgment: A Practical Guide. SIAM. Mola, F. (1998). Classification and Regression Trees Software and New Developments, Studies in Classifi-cation, Data Analysis, and Knowledge Organization. 311–318. doi: 10.1007/978-3-642-72253-0_42. Murthy, D. N. P., Xie, M., & Jiang, R. (2004). Weibull models. John Wiley & Sons. Osisanwo, F. Y., J. E. T. Akinsola, O. Awodele, J. O. Hinmikaiye, O. Olakanmi, and J. Akinjobi, (2017). Supervised Machine Learning Algorithms: Classification and Comparison,” Interna-tional Journal of Computer Trends and Technology, 48(3). 128–138. doi: 10.14445/22312803/ijctt-v48p126. Peters, J., De Baets, B., Verhoest, N. E., Samson, R., Degroeve, S., De Becker, P., & Huybrechts, W., (2007). Random forests as a tool for ecohydrological distribution modelling, Ecol. Modell. , 207(2–4), 304–318. Ponsadailakshmi S., Ganapathy Sankari S., Mythili Prasanna S., and Madhurambal G., (2018). Evaluation of water quality suitability for drinking using drinking water quality index in Nagapattinam district, Tamil Nadu in Southern India, Groundwater for Sustainable Development, 6. 43–49. doi: 10.1016/j.gsd.2017.10.005. Powell, C. (2003). The Delphi technique: myths and realities, J. Adv. Nurs ., 41(4), 376–382. Puntanen S., (2010). Linear Regression Analysis: Theory and Computing by Xin Yan, Xiao Gang Su, Inter-national Statistical Review, 78(1). 144–144. doi: 10.1111/j.1751-5823.2010.00109. Quinlan J. R., (1996). Learning decision tree classifiers, ACM Computing Surveys , 28(1). 71–72. doi: 10.1145/234313.234346. Rabeiy R. E., (2018). Assessment and modeling of groundwater quality using WQI and GIS in Upper Egypt area, Environ. Sci. Pollut. Res. Int., 25(31), 30808–30817. Rahul, A. Gupta, A. Bansal, and K. Roy, (2021). Solar Energy Prediction using Decision Tree Regres-sor, 2021 5th International Conference on Intelligent Computing and Control Systems (ICICCS) . doi: 10.1109/iciccs51141.2021.9432322. Ramadhan M. M., Sitanggang I. S., Nasution F. R., and Ghifari, A. (2017). Parameter Tuning in Random Forest Based on Grid Search Method for Gender Classification Based on Voice Frequency, DEStech Transac-tions on Computer Science and Engineering , 2017. doi: 10.12783/dtcse/cece2017/14611. Ravikumar P., Somashekar R. K., and Angami M., (2011). Hydrochemistry and evaluation of groundwater suita-bility for irrigation and drinking purposes in the Markandeya River basin, Belgaum District, Karnataka State, India, Environ. Monit. Assess ., 173(1–4), 459–487. Reddy T., Sewpersad Y., and Yen Y.-C. J., (2020). Characterisation of the primary heat replacement element event for a horizontal electric water heater, 2020 International SAUPEC/RobMech/PRASA Conference. doi: 10.1109/saupec/robmech/prasa48453.2020.9041057. Rodriguez-Galiano, V., Sanchez-Castillo, M., Chica-Olmo, M., & Chica-Rivas, M. J. O. G. R., (2015). Machine learning predictive models for mineral prospectivity: An evaluation of neural networks, random forest, regression trees and support vector machines, Ore Geology Reviews , 71, 804–818. doi: 10.1016/j.oregeorev.2015.01.001. Roy R., (2018). An Approach to Develop an Alternative Water Quality Index Using FLDM. Water Re-sources Development and Management. 51–68. doi: 10.1007/978-981-10-6205-6_3. Saffran, K., Cash, K., & Hallard, K. (2017). CCME Water Quality Index User’s Manual 2017 Update. Can. Water Qual. Guidel. Prot. Aquat. Life, 1-5. https://ccme.ca/en/res/wqimanualen.pdf (accessed. Sánchez, E., Colmenarejo, M. F., Vicente, J., Rubio, A., García, M. G., Travieso, L., & Borja, R., (2007). Use of the water quality index and dissolved oxygen deficit as simple indicators of watersheds pollution, Ecological Indicators, 7(2). 315–328. doi: 10.1016/j.ecolind.2006.02.005. Segal M. R., (2022). Machine Learning Benchmarks and Random Forest Regression. Accessed: Ma. [Online]. Available: https://escholarship.org/content/qt35x3v9t4/qt35x3v9t4.pdf?t=krnmaw Shabbir R. and Ahmad S. S., (2015). Use of Geographic Information System and Water Quality Index to As-sess Groundwater Quality in Rawalpindi and Islamabad, Arab. J. Sci. Eng., 40(7), 2033–2047. Shah K. A. and Joshi G. S., (2017). Evaluation of water quality index for River Sabarmati, Gujarat, India, Applied Water Science , 7(3). 1349–1358. doi: 10.1007/s13201-015-0318-7. Shailaja, K., Seetharamulu, B., & Jabbar, M. A. (2018, March). Machine learning in healthcare: A review. In 2018 Second international conference on electronics, communication and aerospace technology (ICECA) (pp. 910-914). IEEE.. Sharma T., Banerjee K., Mathur S., and Bali V., (2020). Stress Analysis Using Machine Learning Techniques, International Journal of Advanced Science and Technology , 29(3), 14654–14665. Sillberg, C., Kullavanijaya P., and O. Chavalparit, (2019). Water quality classification by integration of at-tribute-realization and support vector machine for the Chao phraya river, Inż. Ekol., 22(9), 70–86. Singh and Khan, (2019). Ground water quality assessment of Dhankawadi ward of Pune by using GIS, Int. j. geomat. geosci . Singh S. and Hussian A., (2016). Water quality index development for groundwater quality assessment of Greater Noida sub-basin. Uttar Pradesh, India , Cogent Engineering , 3, (1). 1177155. doi: 10.1080/23311916.2016.1177155. Singh, A., Thakur, N., & Sharma, A. (2016). A review of supervised machine learning algorithms,” 2016 3rd International Conference on Computing for Sustainable Global Development (INDIACom). [Online]. Available: https://ieeexplore.ieee.org/abstract/document/7724478 Srinivas, P., Kumar, G. P., Prasad, A. S., & Hemalatha, T. (2017). Generation of groundwater quality index map-A case study, Res. Rep. Dept. Civ. Environ . Eng. Fac. Eng Saitama Univ. Standard, I. (2013). Drinking water-specification. Ethiopian Standards Agency, Addis Ababa, Ethiopia. Su X., Yan X., and Tsai C.-L., (2012). Linear regression, Wiley Interdiscip. Rev. Comput. Stat ., 4(3), 275–294. Sutadian, A. D., Muttil, N., Yilmaz, A. G., & Perera, B. J. C. (2016). Development of river water quality indices—a review. Environmental monitoring and assessment , 188, 1-29. doi: 10.1007/s10661-015-5050-0. Suthaharan S. (2016) “Machine Learning Models and Algorithms for Big Data Classification,” Integrated Series in Information Systems. doi: 10.1007/978-1-4899-7641-3. Şener Ş., Şener E., and Davraz A., (2017). Evaluation of water quality using water quality index (WQI) method and GIS in Aksu River (SW-Turkey), Sci. Total Environ., 584–585, 131–144. Talabani H. and Avci E. (2018). Performance Comparison of SVM Kernel Types on Child Autism Disease Database, in 2018 International Conference on Artificial Intelligence and Data Processing (IDAP), 1–5. Talabani H. and Avci E., (2018). Performance Comparison of SVM Kernel Types on Child Autism Disease Database, 2018 International Conference on Artificial Intelligence and Data Processing (IDAP). doi: 10.1109/idap.2018.8620924. Tso, G. K., & Yau, K. K. (2007). Predicting electricity energy consumption: A comparison of regres-sion analysis, decision tree and neural networks, Energy The International Journal, 1761–1768. Tunc Dede O., Dede O. T., Telci I. T., and Aral M. M., (2013). The Use of Water Quality Index Models for the Evaluation of Surface Water Quality: A Case Study for Kirmir Basin, Ankara, Turkey, Water Quality, Exposure and Health , 5(1). 41–56. doi: 10.1007/s12403-013-0085-3. Tyagi S., Sharma B., Singh v, and Dobhal R., (2020). Water Quality Assessment in Terms of Water Quality Index,” American Journal of Water Resources , 1, (3). 34–38, 2020. doi: 10.12691/ajwr-1-3-3. Uddin, M. G., Olbert, A. I., & Nash, S., (2020). Assessment of water quality using Water Quality Index (WQI), Ecol. Indic., 85, 966–982. UN-Water. https://www.unwater.org/water-facts/ (accessed Mar. 26, 2022). Uyanık G. K. and Güler, N. (2013) A study on multiple linear regression analysis,” 4th International Con-ference on New Horizons in Education, vol. 106, p. Pages 234–240, Dec. 2013. Van Dijk, J. A. (2090). Delphi questionnaires versus individual and group interviews: A comparison case, Technol. Forecast. Soc. Change , 37(3), 293–304. Vapnik V. N. and Chervonenkis A., (1964). A note on one class of perceptrons, Automation and Remote Control 25(1). Wagh, V. M., Panaskar, D. B., Muley, A. A., & Mukate, S. V., (2017). Groundwater suitability evaluation by CCME WQI model for Kadava River Basin, Nashik, Maharashtra, India, Modeling Earth Systems and Envi-ronment, 3(2), 557–565. Watanachaturaporn M. Xu, Varshney P. P., and Arora M., (2005). Decision tree regression for soft classifica-tion of remote sensing data, Remote Sensing of Environment, 97(3). 322–336. doi: 10.1016/j.rse.2005.05.008. Yadav, N., Banerjee K., and Bali V., (2020). A Survey on Fatigue Detection of Workers Using Machine Learning, IJEHMC , 11(3), 1–8. Zhou Y., Cheng G., Jiang S., and Dai M., (2020). Building an efficient intrusion detection system based on feature selection and ensemble classifier, Computer Networks, 174. 107247. doi: 10.1016/j.comnet.2020.107247. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5859247","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":407713665,"identity":"8c264598-905b-41ac-a725-9bf74c546257","order_by":0,"name":"Kakoli Banerjee","email":"","orcid":"","institution":"JSS University: JSS Academy of Higher Education and Research","correspondingAuthor":false,"prefix":"","firstName":"Kakoli","middleName":"","lastName":"Banerjee","suffix":""},{"id":407713666,"identity":"55ca147f-dbf1-4e08-bbb5-ed61728f42da","order_by":1,"name":"Pradeep Kumar","email":"","orcid":"","institution":"JSS University: JSS Academy of Higher Education and Research","correspondingAuthor":false,"prefix":"","firstName":"Pradeep","middleName":"","lastName":"Kumar","suffix":""},{"id":407713667,"identity":"cf9b816b-a308-4542-9783-770fdc1b2764","order_by":2,"name":"Ajay Kumar","email":"","orcid":"","institution":"JSS University: JSS Academy of Higher Education and Research","correspondingAuthor":false,"prefix":"","firstName":"Ajay","middleName":"","lastName":"Kumar","suffix":""},{"id":407713668,"identity":"9e8a65a6-2d93-454c-9e45-6dc9c2fc5675","order_by":3,"name":"Kamal Upreti","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABEElEQVRIiWNgGAWjYBACCRDxAMxkbDjwocJGDsQ88ICQlgQQwcbYeHDGmTRjsJYE4rQwMB/mbTuc2AAXwQEk208nfkhsu5PPL9/ccJjnzOH0+WGHHwJtsZPTbcCuRZond7NEYtszy5ltjA0H51Sk5268nWYA1JJsbHYAuxY5htwNQC2HDQyOAb3/5ox17sbZCSAtBxK34dLC/3bzD5AWe5AW3jbmdMPZ6R/wapGWyN0GsYUN6DDeNucEeekc/LZIzni7zSLh3DMDiWOJDaBANtwgnVNwIMEAt18kzuduvvGh7I4Bf/Pxxx+AUSkvPzt9M5BhJ4dLCxQgyRqA2QZ4laNpkW8gqHoUjIJRMApGGAAAPwtu2F4xECMAAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0003-0665-530X","institution":"Christ College: CHRIST (Deemed to be University)","correspondingAuthor":true,"prefix":"","firstName":"Kamal","middleName":"","lastName":"Upreti","suffix":""},{"id":407713669,"identity":"93aa2ab2-1bda-4b2c-bd7f-b26ee1ac6420","order_by":4,"name":"Shubham Mahajan","email":"","orcid":"","institution":"Amity University Haryana: Amity University Haryana-Gurugram","correspondingAuthor":false,"prefix":"","firstName":"Shubham","middleName":"","lastName":"Mahajan","suffix":""},{"id":407713670,"identity":"37054680-f6d1-46bf-b57c-17a76a736e41","order_by":5,"name":"Mohammad Shahnawaz Nasir","email":"","orcid":"","institution":"Jazan University","correspondingAuthor":false,"prefix":"","firstName":"Mohammad","middleName":"Shahnawaz","lastName":"Nasir","suffix":""}],"badges":[],"createdAt":"2025-01-19 11:52:29","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-5859247/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5859247/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":75006297,"identity":"652eb37d-71e3-48e2-82dc-bdfd001acfc2","added_by":"auto","created_at":"2025-01-29 10:48:18","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":59176,"visible":true,"origin":"","legend":"\u003cp\u003eLocations used for data acquisition\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-5859247/v1/e887264c9e70e7925fc2785c.png"},{"id":75006304,"identity":"aa342b63-ab7c-408b-9120-70085dfd6f7b","added_by":"auto","created_at":"2025-01-29 10:48:18","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":243614,"visible":true,"origin":"","legend":"\u003cp\u003eParameter values for the locations in the target zone\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-5859247/v1/7d1ec44a690d7a20d2467889.png"},{"id":75006302,"identity":"99754c7e-4558-43dc-ac3f-61cf2c13f600","added_by":"auto","created_at":"2025-01-29 10:48:18","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":141938,"visible":true,"origin":"","legend":"\u003cp\u003eValue distribution of water quality parameters\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-5859247/v1/4486acc8e9fe765bb6ff94ab.png"},{"id":75008102,"identity":"23bdbc78-139e-427c-9cb7-820ced79dd20","added_by":"auto","created_at":"2025-01-29 10:56:18","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":181651,"visible":true,"origin":"","legend":"\u003cp\u003eData preprocessing process flow diagram\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-5859247/v1/7ea65d6450a2fe379874fdf0.png"},{"id":75006298,"identity":"56ae2f32-1981-41c4-84b7-7f412298d623","added_by":"auto","created_at":"2025-01-29 10:48:18","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":110463,"visible":true,"origin":"","legend":"\u003cp\u003eWorkflow diagram\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-5859247/v1/9b79d356aa85bd42258461f5.png"},{"id":75006314,"identity":"ada418c1-eb9b-499e-9257-e2df47e22451","added_by":"auto","created_at":"2025-01-29 10:48:19","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":32043,"visible":true,"origin":"","legend":"\u003cp\u003eCorrelation heatmap for the water quality parameter\u003c/p\u003e","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-5859247/v1/a571b59f6a5b9c82c7b13bf5.png"},{"id":75006319,"identity":"154c8559-21d1-4a63-9fa0-35547ba90813","added_by":"auto","created_at":"2025-01-29 10:48:19","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":58702,"visible":true,"origin":"","legend":"\u003cp\u003eMean square error calculation\u003c/p\u003e","description":"","filename":"floatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-5859247/v1/2f66574d037b55dd1bbe7ae2.png"},{"id":75006307,"identity":"45e349b2-809c-44a5-82da-c7f5f343d8e2","added_by":"auto","created_at":"2025-01-29 10:48:18","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":26513,"visible":true,"origin":"","legend":"\u003cp\u003eAccuracy of algorithms when all parameters were used for training\u003c/p\u003e","description":"","filename":"floatimage8.png","url":"https://assets-eu.researchsquare.com/files/rs-5859247/v1/1d3b998b6f4523ff980a672d.png"},{"id":75008106,"identity":"0d07298a-27a1-4278-b9ed-d7b7362fab8f","added_by":"auto","created_at":"2025-01-29 10:56:19","extension":"png","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":27355,"visible":true,"origin":"","legend":"\u003cp\u003eAccuracy of algorithms when only Turbidity was used for training\u003c/p\u003e","description":"","filename":"floatimage9.png","url":"https://assets-eu.researchsquare.com/files/rs-5859247/v1/223b049e4891a06bb87d5994.png"},{"id":75006313,"identity":"8dafd4ea-b210-4cda-b472-8502c24853c3","added_by":"auto","created_at":"2025-01-29 10:48:19","extension":"png","order_by":10,"title":"Figure 10","display":"","copyAsset":false,"role":"figure","size":25145,"visible":true,"origin":"","legend":"\u003cp\u003eAccuracy of algorithms when only pH was used for training\u003c/p\u003e","description":"","filename":"floatimage10.png","url":"https://assets-eu.researchsquare.com/files/rs-5859247/v1/1cf6fd0f75a3ae02c2db2677.png"},{"id":75006315,"identity":"f6dbcbdf-3232-4b98-a3b5-eaa38c13e9d1","added_by":"auto","created_at":"2025-01-29 10:48:19","extension":"png","order_by":11,"title":"Figure 11","display":"","copyAsset":false,"role":"figure","size":25280,"visible":true,"origin":"","legend":"\u003cp\u003eAccuracy of algorithms when only B.O.D was used for training\u003c/p\u003e","description":"","filename":"floatimage11.png","url":"https://assets-eu.researchsquare.com/files/rs-5859247/v1/49319ab9d73d4f9a3bd7c478.png"},{"id":75008405,"identity":"be011d2d-a050-45ad-89ff-89d988acf105","added_by":"auto","created_at":"2025-01-29 11:04:18","extension":"png","order_by":12,"title":"Figure 12","display":"","copyAsset":false,"role":"figure","size":27354,"visible":true,"origin":"","legend":"\u003cp\u003eAccuracy of algorithms when only D.O was used for training\u003c/p\u003e","description":"","filename":"floatimage12.png","url":"https://assets-eu.researchsquare.com/files/rs-5859247/v1/c2c6b37fd350e749a0658efe.png"},{"id":75006332,"identity":"2bef8fc3-0263-485c-bbe8-33c591a8fde7","added_by":"auto","created_at":"2025-01-29 10:48:19","extension":"png","order_by":13,"title":"Figure 13","display":"","copyAsset":false,"role":"figure","size":27083,"visible":true,"origin":"","legend":"\u003cp\u003eAccuracy of algorithms when only Hardness was used for training\u003c/p\u003e","description":"","filename":"floatimage13.png","url":"https://assets-eu.researchsquare.com/files/rs-5859247/v1/cbfe6ac95a04f8fb94f10ad3.png"},{"id":75008107,"identity":"0cfd51e0-dc26-4984-8ce6-03044dd76e1e","added_by":"auto","created_at":"2025-01-29 10:56:19","extension":"png","order_by":14,"title":"Figure 14","display":"","copyAsset":false,"role":"figure","size":26192,"visible":true,"origin":"","legend":"\u003cp\u003eAccuracy of algorithms when only C.O.D was used for training\u003c/p\u003e","description":"","filename":"floatimage14.png","url":"https://assets-eu.researchsquare.com/files/rs-5859247/v1/86348839e4e84e3cca0747ed.png"},{"id":77742896,"identity":"60c46b63-ce69-4dd1-a88a-83a1048ebb19","added_by":"auto","created_at":"2025-03-05 06:07:52","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2313248,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5859247/v1/2a017902-cf7a-4e00-90a9-3e2b83efc333.pdf"}],"financialInterests":"","formattedTitle":"\u003cp\u003eComputation of Water Quality Index and Its Estimation Using Machine Learning Techniques\u003c/p\u003e","fulltext":[{"header":"1 Introduction","content":"\u003cp\u003eWater is a key natural resource and asset (S. Tyagi et. al. 2020), which is essential for human existence as well as the sustainability of the planet (2022). Water quality is integral in determining health, main-taining gender equality and food security as well as ensuring economic development and social growth (M. K. Jhaet. Al. 2020).\u003c/p\u003e \u003cp\u003eTraditionally, the water quality is calculated by assessing and analysing the physical, chemical and bi-ological properties of water (2022), with minimum value requirements for water constituents or quality indicators (2022).\u003c/p\u003e \u003cp\u003eOne of the primary sources of water is groundwater, which is available almost everywhere under the surface of the earth in the form of thousands of local aquifer systems and not as one single unit (S. Singh and A. Hussian, 2016). In developing nations, such as India, it plays a vital role in strengthening the country\u0026rsquo;s economic growth (P. Ravikumar et. al. 2011). The Central Pollution Control Board has 870 monitoring stations across the country which monitors surface water quality on a monthly or quarterly basis and groundwater on a half yearly basis, which indicates that organic pollutants predominantly cause contamination (R. M. Bhardwaj, 2005).\u003c/p\u003e \u003cp\u003eApart from India, it has been estimated that over 50 countries in the world are treated with polluted or partially treated polluted water (J. Mateo-Sagasta et. al. 2017). In developing countries, over 80 percent of polluted water is used for irrigation purposes (D. D. Mara and S. Cairncross, 2019). Other diseases related to excreta are extremely common in developing countries as the wastewater contains pathogens such as viruses and bacterias, which leads to a risk to public safety and health. Therefore the contaminated wa-ter which contains plant nutrients may benefit soil but it also contains soluble salts and heavy metals in high concentrations (A. Gafoor et. al. 1994). This water is used by farmers who want to cut on their ex-penses, but over long periods of time extensive use of such methods for irrigation leads to groundwater contamination, the harmful effects of which can last for several years (M. Ibrahim and S. Salmon, 2020).\u003c/p\u003e \u003cp\u003eDue to overgrowing population and urban expansion numerous water quality issues are becoming a concern, making water quality an important tool for measuring health related issues and safety measures in both humans and aquatic life (O. Tunc Dede et. al. 2013). Although some waste disposal practices might not affect groundwater, it directly contaminates surface level water bodies (E. S\u0026aacute;nchez et al.,2007). Other agents such as poisonous metals may be mobile enough to seep into the ground and mix with the drinking water supplies making it unfit for human consumption\u0026ndash; for example Silver metals in form of AgNPs.\u003c/p\u003e \u003cp\u003eTo estimate the water grade, a metric called the Water Quality Index (WQI) is used which is represented through a single number which indicates the water quality through aggregation of individual water quality parameters, with the score representing the water quality wherein a lower score indicates poor quality and a higher score indicates better quality (A. Lumb, et. al. 2011). WQI is a rating which reflects the compound impact of various water quality parameters and is considered a powerful tool to evaluate both surface and groundwater pollution, as well as in quality upgradation programmes. Its importance can be understood by the fact that it is a single term which still gives a unique rating for depicting over-all water quality (K. A. Shah and G. S. Joshi, 2017). There are various indices available which make use of different parameters depending on the intended objectives for its estimation (C. E. Canada and Canada. 2003).\u003c/p\u003e \u003cp\u003eAlthough the WQI estimation might provide us with crucial information about the state of the water and its quality, there is no universal method of calculating WQI as of now moreover, WQI calculations consume not only resources such as time as it is a complex and a lengthy process, it also has multiple inconsistencies as WQI often uses different equations for calculation (D. T. Bui et. al. 2020). Another such drawback of WQI estimation is the way in which the magnitude of weights are calculated for each water quality parameter, most of which are subjective in nature i.e. they do not take into account the factor that it may cause hazards (R. Roy, 2018). Therefore, the primary objective of our study is to minimize the number of input variables from the feature set in order to calculate WQI, which should ideally be esti-mated by just using a single parameter.\u003c/p\u003e \u003cp\u003eThere are majorly two broad areas of work in this study: calculation of WQI and development of machine learning modes for its prediction. Our approach uses the Brown et. al. method for WQI calculation that first selects the parameters, namely pH, Turbidity (NTU), Hardness (mg/L), D.O. (mg/L), B.O.D. (mg/L), C.O.D. (mg/L), then these parameters are assigned a weight factor, and the Q value is calculated. Finally the WQI is computed which gives the grade for the water quality. Once the WQI is found, the dataset is split into training and testing sets and various Supervised Machine Learning algorithms are applied such as Linear Regression, Decision Tree Regressor, Random Forest Regressor and Support Vector Regressor. The motive for this approach was to reduce the dimensionality of the input variables required to compute WQI, while maintaining a high accuracy.\u003c/p\u003e"},{"header":"2 Study Area","content":"\u003cp\u003eDue to high anthropogenic activities like the development and setup of industries and factories in Gau-tam Buddha Nagar, the contamination of groundwater has gone up significantly which requires more investigation for the fulfilment of current and future potable water requirements (R. K. Horton, 2015, Murthy et al. 2004). Gautam Buddha District has a soil classification that consists of a combination of excessively drained, well drained and moderately well drained soil (R. K. Horton, 2015). Gautum Buddha Nagar was selected as the field of study and was used as a site for the collection of samples and data.\u003c/p\u003e \u003cp\u003eNoida is a district in Uttar Pradesh which comes under Gautam Buddha Nagar and is close to the Na-tional Capital Territory of Delhi. The city spans across the latitudes 28\u0026deg; 26\u0026prime; 39 \u0026prime;\u0026prime; N to 28\u0026deg; 38\u0026prime; 10\u0026prime;\u0026prime; N and longitudes 77\u0026deg; 29\u0026prime; 53\u0026prime;\u0026prime; E to 77\u0026deg; 17\u0026prime; 29\u0026prime;\u0026prime; E covering a total area of 203 sq. km.\u003c/p\u003e"},{"header":"3. History of WQI","content":"\u003cp\u003eThe WQI was developed in 1965 in United States by Horton, it was based on common parameters used to determine water quality (such as pH, coliforms, dissolved oxygen, sewage treatment, specific con-ductance, carbon chloroform extract, alkalinity and chlorides). While most WQI models operate on a scale of zero to hundred, the Horton model implements a scale in which zero represents excellent water quality and values higher than hundred represent extremely poor water quality (Uddin, et. al. 2020). The criteria selected by Horton primarily focused on three factors is listed below in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e (T. Abbasi and S. A. 2012).\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eVariables used for the computation the water quality index should be limited\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eVariables used should significant in most areas\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eVariables used should have available and reliable data\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eFollowing Horton, a large number of WQI models were formulated by several different national and in-ternational bodies.\u003c/p\u003e \u003cp\u003eIn 1970, Brown et al. developed the \u0026lsquo;Weighted Arithmetic Index Method\u0026rsquo;, with the enlisted steps:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eCalculation of sub-index of quality rating\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eCalculation of unit weight\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eAggregation of quality rating with unit weight resulting in WQI\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eWater Quality Index range for its classification\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWQI Range\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eWater Quality Status\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePossible Usage\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e0\u0026ndash;5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eExcellent Water Quality\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDrinking, irrigation and industrial\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e26\u0026ndash;50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGood Water Quality\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDrinking, irrigation and industrial\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e51\u0026ndash;75\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePoor Water Quality\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eIrrigation and industrial\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e76\u0026ndash;100\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eVery Poor Water Quality\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eIrrigation\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAbove 100\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eUnsuitable for drinking and propagation of fish\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eProper treatment required before use\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Methods for finding sub-indices\u003c/h2\u003e \u003cp\u003eThe following are the three most commonly followed methods for calculating sub-indices:\u003c/p\u003e \u003cp\u003e \u003cb\u003eExpert Opinion\u003c/b\u003e \u003c/p\u003e \u003cp\u003eFor several WQIs one of the challenges is to select important parameters to calculate the final index ag-gregation as subjective assessments are a part of the initially selected parameters (N. Akhtar et al., 2021). To reduce the subjectivity and the uncertainty, several experts in parameters selection have been used (T. D. Banda and M. V. Kumarasamy, 2020). The expert judgement process is paired with parameter selec-tion using three different methods (M. A. Meyer and J. M. Booker, 2001).\u003c/p\u003e \u003cp\u003e \u003cstrong\u003eIndividual interviews-\u003c/strong\u003e \u003cp\u003eIndividual interviews are conducted through mail or other methods of commu-nication where experts (people who possess high knowledge of a subject) give their remarks on a topic or subject (J. A. G. M. Van Dijk, 2090).\u003c/p\u003e \u003c/p\u003e \u003cp\u003eVirtual groups-: The virtual groups are similar to interviews in the way that in this method too experts give opinions but instead of taking individual remarks, there is a panel of expert individuals (C. Powell, 2003).\u003c/p\u003e \u003cp\u003eDelphi method-: The Delphi method is basically a questionnaire sequence or rounds which is distrib-uted over a controlled feedback system in order to maximise the most reliable consensus of a group of experts H. A. Linstone and M. Turoff, 1975).\u003c/p\u003e \u003cp\u003e \u003cstrong\u003eStatistical Methods\u003c/strong\u003e \u003cp\u003eUsing historical parameter data, rating curves\u0026rsquo; key points are developed which are dependent on statis-tical characteristics such as mean values and various other quantities calculated over a long period of time. This method does not include the expert opinions.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eUse of Water Quality Standards-\u003c/strong\u003e \u003cp\u003eAnother way for the development of the subindex functions is through the standard for water quality. Here, each water quality attribute is assigned a value denoting its rating between a scale of 0 and 100, based on national standards (India, Malaysia etc) or international standards such as WHO, United States Environmental Protection Agency (USEPA) (Standard, I. 2013, Ponsadailakshmi et. al. 2018).\u003c/p\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e3.2 CCME\u003c/h2\u003e \u003cp\u003eThe CCME (Canadian Council of Ministers of the Environment) WQI gives an adaptable tool which provides a flexible index template (T. Hurley, R. Sadiq, and A. Mazumder, 2012) and is mainly composed of three factors (A. Lumb, et. al. 2006). The CCME is well documented and provides detailed procedures for index calculation and application (2022). The three factors are the following:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eScope\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eFrequency\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eAmplitude.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"4 Related Work","content":"\u003cp\u003eWQI is an important indicator of water quality, and multiple studies have been conducted to compute its value. Kangabam et al. chose Loktak Lake as the area of study to assess the WQI from physicochem-ical water parameters (R. Das Kangabam et. al. 2017). Judran and Kumar mathematically developed the WQI for Al-Gharraf River Thi Qar of Iran, and found the quality of the river water to be poor for drinking (N. H. Judran and A. Kumar, 2020). Şener et al. analysed WQI of Aksu River, located in Turkey, to evaluate it for drinking purposes using thirteen water quality parameters and found that environmental pollu-tants affect the river\u0026rsquo;s quality negatively (Ş. Şener, E. Şener, and A. Davraz, 2017). Kamboj and Kamboj as-sessed the groundwater quality in the Indian district of Haridwar through the calculation of WQI using arithmetic mean method, and found the quality to be suitable for drinking (N. Kamboj and V. Kamboj, 2019). Wagh et al. classified the groundwater suitability of Kadava river of India using the CCME WQI method using eleven physicochemical parameters, and found the samples to fall in the marginal cate-gory for drinking purposes (V. M. Wagh et. al. 2017).\u003c/p\u003e \u003cp\u003eNumerous studies have been conducted to study the quality of water using WQI. In recent times, Super-vised Machine Learning approaches and Artificial Intelligence methods have been employed to predict WQI. Ahmed et al. collected data from Pakistan Council of Research in Water Resources to estimate WQI using eight regression algorithms, as well as classify samples into predefined Water Quality Clas-ses using ten classification algorithms (U. Ahmed et. al. 2020). When they used four parameters (temper-ature, turbidity, pH and total dissolved solids (TDS)) to predict water quality, the best performing algo-rithm was found to be Gradient Boosting while Polynomial Regression was more efficient when TDS was dropped and only three parameters were used. Multi-layer perceptron performed the best in pre-dicting the WQC. Sillberg et al. classifies quality of water for Chao Phraya River, which is the largest river in Thailand, using a machine learning-based approach coined AR-SVM, which integrates support vector machine with attribute-realisation (C. Sillberg et. al. 2019). The AR-SVM approach resulted in an accuracy of 0.94. Hameed et al. applied Artificial Neural Network algorithms (back propagation neural networks models and radial basis function neural network (RBFNN)) for WQI prediction in Malaysia, finding that RBFNN proved to be effective for assessing the quality of water (M. Hameed, et. al. 2017).\u003c/p\u003e \u003cp\u003eThis domain of research has also been explored substantively in the Indian subcontinent. Hassan et al. used Indian water quality data collected from Kaggle and applied five machine learning classifiers, Neural Network, Random Forest, Multinomial Logistic Regression, Support Vector Machine, and Bagged Tree Model, to predict water quality (M. M. Hassan et al., 2021, Citaristi, I. 2022). Multinomial Logistic Regression gave the best accuracy score of 99.83%, while Support Vector Machine gave the lowest accuracy score of 96.98%. Aldyhani et al. implemented NARNET which is a nonlinear autoregressive neural network, along with long short-term memory (LSTM) deep learning algorithm for WQI prediction, while support vector machine (SVM), K-nearest neighbour (K-NN), and Naive Bayes, were used for forecasting WQC, on a dataset from Kaggle collected in India (T. H. H. Aldhyani et. al. 2020). The results showed that on the basis of R value, NARNET performed better than LSTM while SVM achieved highest accuracy for WQC prediction. Kadam et al. applied artificial neural network (ANN) and multiple linear regression (MLR) for prediction of WQI on data collected from Shivganga watershed, and found a higher level of precision in the ANN model (A. K. Kadam et. al. 2019).\u003c/p\u003e"},{"header":"5 Methodology","content":"\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e5.1 Data Acquisition\u003c/h2\u003e \u003cp\u003eThe data used in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e as shown in study was collected from different locations of Noida and some parts of Delhi by the Department of Computer Science, JSS Academy of Technical Education, Noida. A total number of 51 samples were collected. One litre of water sample was collected from each site and further tests were conducted in the laboratories of JSS Academy of Technical Education, Noida to find the properties of samples (K. Banerjee et al.,2022). For the collection of samples, the methods prescribed in the American Public Health Association manual were followed, and samples were collected from different sources such Government hand pumps, General hand pumps and Borwells, also depicted in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, (K. Banerjee et. al. 2021).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe acquired data has different parameters related to water, also shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e and the standard values are provided in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e and Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e. These parameters are:\u003c/p\u003e \u003cp\u003e \u003cb\u003ePH\u003c/b\u003e \u003c/p\u003e \u003cp\u003epH denotes hydrogen ion concentration and is known as the \u0026ldquo;potential of hydrogen\u0026rdquo; or the \u0026ldquo;power of hydrogen\u0026rdquo; for aqueous solutions. It is the quantitative measure of the acidity or alkalinity of liquid solu-tions on a range of 0 to 14, with 7 being neutral. pH reading of less than 7 indicates acidity, while a reading greater than 7 indicates the solution is basic.\u003c/p\u003e \u003cp\u003e \u003cb\u003eTurbidity\u003c/b\u003e \u003c/p\u003e \u003cp\u003eTurbidity is caused by suspension of fine matter (like organic matter, inorganic matter, clay and silt), and gives a measure to ascertain to what degree water loses its transparency because of the presence of such particles. It is measured in Nephelometric Turbidity Units or NTU through a nephelometer or a turbidimeter. It causes the cloudiness, haziness or milkiness of a fluid.\u003c/p\u003e \u003cp\u003e \u003cb\u003eHardness\u003c/b\u003e \u003c/p\u003e \u003cp\u003eHardness is caused due to soluble bicarbonates, chlorides and sulphates of magnesium and calcium in the water molecules. It is water with a high amount of mineral content. This type of water does not easi-ly form lather with soap, and therefore hardness can also be defined as the capacity of water to precipi-tate soap.\u003c/p\u003e \u003cp\u003e \u003cb\u003eDissolved Oxygen\u003c/b\u003e \u003c/p\u003e \u003cp\u003eThe dissolved oxygen, or D.O., in water is the amount of oxygen present in it. It gets mixed with water bodies through the atmosphere, water plants and algae such as phytoplankton. The presence of sun-light is important for these aquatic lives to produce oxygen through photosynthesis. The running water streams are found to have more dissolved oxygen than a stationary pond or a lake.\u003c/p\u003e \u003cp\u003e \u003cb\u003eBiochemical Oxygen Demand\u003c/b\u003e \u003c/p\u003e \u003cp\u003eBiochemical Oxygen Demand, also known as B.O.D., is the oxygen amount microorganisms and dif-ferent bacterias require to break down organic matter in water.\u003c/p\u003e \u003cp\u003e \u003cb\u003eChemical Oxygen Demand\u003c/b\u003e \u003c/p\u003e \u003cp\u003eChemical Oxygen Demand, also known as C.O.D., is the consumed oxygen amount to chemically oxi-dise inorganic and organic matter present in water.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eBIS accepted and permissible values for water quality parameters\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eParameters\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRequirement (Acceptable Limit)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePermissible Limit in the Absence of Alternate Source\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003epH value\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e6.5\u0026ndash;8.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo relaxation\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTurbidity (NTU) Max\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTotal Dissolved Solids (mg/l) Max\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e500\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e2000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTotal Hardness (mg/l) Max\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e200\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e600\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTotal Alkalinity (mg/l) Max\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e200\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e600\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eChloride (mg/l) Max\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e250\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"3\"\u003e\u003cb\u003eSOURCE\u003c/b\u003e: BIS (IS:10500)\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eWHO standard values for water quality parameters\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eParameters\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eUnits\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eWHO Standard\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e℃\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e12\u0026ndash;25\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003epH\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u0026ndash;\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e6.5\u0026ndash;8.5\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e㎲/cm\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e400\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCl-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003emg/l\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e250\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTDS\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003emg/l\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e500\u0026ndash;1000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTH\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003emg/l\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e500\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"3\"\u003e\u003cb\u003eSOURCE\u003c/b\u003e: WHO Guidelines for drinking water quality\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e5.2 Data Preprocessing\u003c/h2\u003e \u003cp\u003eData preprocessing is a critical step in data analysis, and prepares the data for further processing in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e. To ready the dataset, the non-affecting parameters and redundant parameters were removed, missing val-ues were handled and the dataset was normalised.\u003c/p\u003e \u003cp\u003eIn our study, the data was extracted manually and therefore it had many missing and ambiguous val-ues. The missing values were removed before splitting the dataset into training and testing subsets. At-tributes related to location such as Latitude, Longitude and Elevation were not a part of the final dataset because these attributes did not contribute to the model performance as they had very little correlation to the target value. Any other technique to replace the missing values such as mean value substitution was not employed because values such as pH, Turbidity, Temperature etc. were unique to different locations and any generalisation such as mean would have made the model skewed.\u003c/p\u003e \u003cp\u003eSOURCE: (Banerjee et al. 2021, (2012)\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003e5.3 Processing\u003c/h2\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cb\u003eComputation\u003c/b\u003e \u003c/p\u003e \u003cp\u003eTo calculate WQI the following four steps are followed, also shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e (D. Sutadian et. al. 2016):\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eParameter Selection\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eAssigning of Weights\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eQuality Rating\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eAggregation of sub indices\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eParameter Selection-\u003c/strong\u003e \u003cp\u003eThe parameters used for calculating WQI are pH, Turbidity (NTU), Hardness (mg/L), D.O. (mg/L), B.O.D. (mg/L), C.O.D. (mg/L). Other parameters which provided information regarding the location of the test samples were not considered because they did not directly affect the quality of water. Although, location could be a huge determining factor if we draw comparisons between water from different cities or national states, in our case the study area was limited to only a small part of Gautam Buddha Nagar.\u003c/p\u003e \u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003e2. Assigning of weights\u003c/h3\u003e\n\u003cp\u003eThe first important step for calculating the wqi is to calculate the relative weights of the selected parameters. These parameters are then assigned these relative weights (RW) against their original weights as given in Eq.\u0026nbsp;(1) (R. Shabbir and S. S. Ahmad, 2015).\u003c/p\u003e \u003cp\u003eRW = \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\frac{AWI}{{\\sum\\:}_{i=1}^{n}\\:AW1}\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003cp\u003eIn the formula above,\u003c/p\u003e \u003cp\u003eRW: relative weight\u003c/p\u003e \u003cp\u003eAW: assigned weight of each parameter\u003c/p\u003e \u003cp\u003en: number of parameters.\u003c/p\u003e \u003cp\u003eThe relative weight is the ratio of the assigned weight of the parameter to the total sum of all the weights of the parameters as shown in Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e. It presents standard parameter values and their corresponding relative weights for evaluating a system. Key factors such as pH, chloride (Cl⁻), and total dissolved solids (TDS) have the highest relative weights (0.2431), indicating their significant influence compared to other parameters like temperature (T) and electrical conductivity (EC).\u003c/p\u003e \u003cp\u003eThe standard f Suthaharanor the water quality parameters in this study was based on BIS (Indian Standard Specifi-cation for Drinking Water 2012) (2022) and World Health Organisation (2011) (2017) recommendations.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eStandard values and relative weights for water quality parameters\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eParameters\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eStandard value\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRelative Weight\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.1430\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003epH\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.2431\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e200\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.0060\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCl-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.2431\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTDS\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.2431\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTH\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.1215\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003cp\u003e\u003cstrong\u003eQuality Rating-:\u003c/strong\u003e The quality rating for any parameter is calculated as the ratio of its concentration to the standard values. It can be calculated using the formula given in Eq. (2) (Srinivas, Kumar et. al. 2017)\u003c/p\u003e\u003cdiv id=\"Equa\" class=\"Equation\"\u003e \u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$\\:\\text{Q}\\text{i}\\:=\\:\\left[\\frac{Ci\\:-\\:Vi}{Si\\:-\\:Vi}\\right]\\:\\times\\:\\:100$$\u003c/div\u003e \u003c/div\u003e \u003c/p\u003e \u003cp\u003eHere,\u003c/p\u003e \u003cp\u003eQi: quality rating, Ci: computed value of the water parameter, Si: standard WHO or BIS value, Vi: ideal value which is 0 except for pH\u0026thinsp;=\u0026thinsp;7.0 and D.O. = 14.6.\u003c/p\u003e \u003cp\u003eThe above equation can be simplified and written as Eq.\u0026nbsp;(3).\u003cdiv id=\"Equb\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e\n$$\\:\\text{Q}\\text{i}\\:=\\:\\left(\\frac{Ci}{Si}\\right)\\:\\times\\:\\:100$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eIt should be noted that for the parameters pH and D.O. The quality rating is calculated using Eq.\u0026nbsp;(2) because for these parameters the value for Vi is 7.0 and 14.6 respectively. Therefore, we cannot sub-stitute its value by 0.\u003c/p\u003e \u003cp\u003e \u003cstrong\u003eAggregation of sub-indices-\u003c/strong\u003e \u003cp\u003eThis is the last step for calculating WQI. The sub-indices are products of the relative weights and the quality rating. The products of all the values are then summed together in Table\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e, to give WQI using Eq.\u0026nbsp;(4) (R. E. Rabeiy, 2018, Singh \u0026amp; Khan 2019).\u003c/p\u003e \u003c/p\u003e \u003cp\u003eSIi\u0026thinsp;=\u0026thinsp;W\u003csub\u003er\u003c/sub\u003e \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\times\\:\\)\u003c/span\u003e\u003c/span\u003e Q\u003csub\u003ei\u003c/sub\u003e\u003c/p\u003e \u003cp\u003eWQI = \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\sum\\:}_{\\:}^{\\:}\\:\\)\u003c/span\u003e\u003c/span\u003eSI\u003csub\u003ei\u003c/sub\u003e\u003c/p\u003e \u003cp\u003eThe WQI values computed are further classified based on the range and ratings.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eCalculated WQI values using the acquired parameter data\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"8\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLocation\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003epH\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTurbidity\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eHardness\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eC.O.D\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eD.O\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eB.O.D\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003eWQI\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePump No. H-1, Sector 15A\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.74\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e295\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e16\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e188.9849\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAsagarpur Jagir Vilage, Sector 128\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e400\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e10.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e7.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e15.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e117.7359\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHindustan Petroleum, Near Jaypee Hospital\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e465\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e10.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e131.1810\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBalaji Temple, Sector 126\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.35\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e430\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e10.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e8.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e14.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e112.8551\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAnkit Nursery, Sector 131\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.04\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e425\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e5.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e5.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e16.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e124.1108\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGreen Beauty Farm, Sector 135\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.36\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e425\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e10.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e149.0787\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eYakootpur, Sector 167\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.68\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e190\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e10.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e15.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e140.3131\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGulavali, Sector 162\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e56\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e340\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e5.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e5.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e19.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e410.3118\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eJhatta Village, Sector 159\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e205\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e16\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e5.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e19.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e173.8431\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBadauli, Sector 154\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.23\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e410\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e16\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e17.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e146.0595\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eKambuxpur Derin Village, Sector 155\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.63\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e295\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e16\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e7.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e16.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e142.1085\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGujjar Derin, Kambuxpur, Sector 155A\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.51\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e270\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e53.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e16.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e194.1677\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eKondali Bangar, Sector 149\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e240\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e58.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e14.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e186.0353\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGarhi Samastpur, Sector 150\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.06\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e470\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e53.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e3.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e26.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e256.4511\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMomnathal, Sector 150\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e235\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e32\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e18.9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e184.9794\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eShafipur Village, Sector 148\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.95\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e17\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e265\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e42.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e5.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e23.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e280.1487\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMohiyapur Village,Sector 163\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.57\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e485\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e37.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e5.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e21.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e189.8300\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNalgadha, Sector 145\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.87\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e510\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e53.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e19.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e200.4244\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eIdeal Industrial Training Institute, Sector 143\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.52\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e38\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e515\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e21.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e4.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e23.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e355.8180\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eShahdara, Sector 141\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1780\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e16\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e23.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e179.8069\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHindon Flood Plain, Kulesara, Sector 140\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.27\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e335\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e37.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e194.9539\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAllahabad, Sector 86\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.79\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1645\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e10.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e20.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e170.2049\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSai Dham Colony, Sector 88\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.94\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e395\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e48\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e19.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e189.4008\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eKakrala Village, Sector 80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e370\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e42.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e19.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e225.3792\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGijhor Village, Sector 53\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.54\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e775\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e26.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e8.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e15.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e134.8688\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSarfabad, Sector 73\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.28\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e445\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e10.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e17.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e131.2370\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSorkha Village, Sector 118\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.67\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e19\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e580\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e21.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e18.9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e239.1284\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePumping Station 3, Sector 71\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e535\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e16\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e5.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e22.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e164.8662\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePump House, Sector 122\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.65\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e600\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e32\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e20.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e169.3493\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e19, Block H,Sector 116\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.97\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e935\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e5.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e21.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e150.2370\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBaraula Village, Sector 49\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.54\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e21\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e545\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e58.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e19.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e297.2613\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePumping Station, Sector 35\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e550\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e20.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e134.8510\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePumping Station 3, Sector 34\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.58\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e830\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e21.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e15.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e154.2630\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePeerbabaji, Sector 144\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.45\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e765\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e10.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e233.6300\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDallupura Village, Sector 164\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.65\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e360\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e32\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e17.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e155.5846\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDostpur, Mangrauli, Sector 167\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.78\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e205\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e53.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e19.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e442.5794\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNangli Village, Sector 134\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.55\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e580\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e64\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e7.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e200.6006\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBakhtawarpur, Sector 127\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.31\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e325\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e5.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e20.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e171.5604\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSultanpur Village, sector 128\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.75\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e690\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e10.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e5.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e20.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e145.6486\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eShahpur, Sector 131\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.91\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e860\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e16\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e19.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e148.7928\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSadarpur, Sector 45\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.01\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e585\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e5.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e21.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e174.2757\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eChhalera, Sector 44\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.78\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1055\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e21.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e7.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e16.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e134.0816\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSanatan Temple, Sector 41\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.97\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1745\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e48\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e3.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e20.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e206.2695\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eShiv Mandir, Sector 31\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.74\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e695\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e26.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e21.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e168.3809\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNagla Charan Dass, Noida Phase-2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.95\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e840\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e32\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e21.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e196.7644\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNursery, Sector 104\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.85\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e685\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e21.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e4.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e23.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e180.5868\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePumping Station, Sector 80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e290\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e10.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e4.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e25.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e172.8801\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eShiv Mandir, Sector 93\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1230\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e42.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e5.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e22.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e201.2927\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSalarpur Village, Sector 102\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.48\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e750\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e32\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e4.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e23.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e184.3517\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGarhi Chaukhandi, sector 121\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.01\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e560\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e37.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e5.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e21.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e185.5412\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePumping Station, Block-G, Sector 63\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.92\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e590\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e32\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e4.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e24.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e193.4495\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eCorrelation-\u003c/strong\u003e \u003cp\u003eAfter computing the WQI, relevant features needed to be selected for further modelling to predict the WQI value in Fıg. 6. For this purpose, the correlation was calculated between different variables. WQI showed a strong correlation with Turbidity, whereas other parameters did not show strong correlation with the target variable.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eGridSearch-\u003c/strong\u003e \u003cp\u003eGridSearch is an exhaustive method for finding the best parameter values for a given algorithm by run-ning every combination and comparing the score for each parameter value (M. M. Ramadhan et. al. 2017). All the algorithms were run through the cross-validation method GridSearchCV to achieve the best accuracy through every combination of hyperparameters and the one with the best performance was chosen. The cross validation folds were set to 3 for this study.\u003c/p\u003e \u003c/p\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003e5.4 Algorithms/Models\u003c/h2\u003e \u003cp\u003eMachine Learning is a rapidly growing field of interest in Computer Science that detects meaningful patterns in a given set of data to perform data analysis, data extraction and data mining (O. F.y et al., 2017). In our study we applied supervised machine learning algorithms as it categorises data using in-formation already available in the form of training data (A. Singh, et. al. 2016). The supervised machine learning algorithms for predicting a categorical outcome is termed as classification whereas when an algorithm predicts a continuous outcome it is called regression (T. Jiang, et. al. 2020). The aim of super-vised machine learning is the prediction of target variable y \u0026isin; Y based on a set of independent variables x \u0026isin; X (C. Crisci et. al. 2012). The dataset in our study was split into training and testing subsets in which the training set consisted of 75% of the total data and the rest 25% was the testing data. The split was performed using the train_test_split feature present in the sklearn library of python.\u003c/p\u003e \u003cp\u003eMachine Learning approaches are rigorously applied in almost every sector and undoubtedly the most applicable area is ecology (P. A. Flach, 2001). Some other areas of application include healthcare (K. Shailaja, et. al. 2018), fatigue detection (N. Yadav et. al. 2020), DNA sequence compression (K. Banerjee and R. A. Prasad, 2017) and stress analysis (T. Sharma T. Sharma et. al. 2020). Before implementing re-gressors or classifiers, it is important to check for any noise in the data which can be handled by using statistical methods such as data interpolation (K. K. Hiran et. al. 2017). The machine learning algorithms employed in this study for the prediction of WQI are the following:\u003c/p\u003e \u003cp\u003e \u003cstrong\u003eLinear Regression-\u003c/strong\u003e \u003cp\u003eSome of the important questions addressed in this study explore the idea of any correlation between different water quality metrics, i.e. the dependent and the independent variables. To find this correlation in order to devise a model, which gives acceptable prediction and at the same time uses less parameters, Linear Regression was administered.\u003c/p\u003e \u003c/p\u003e \u003cp\u003eRegression is a statistical technique in which a dependent variable, also known as a response variable (S. Puntanen, 2010), paired with one or more than one independent variable is used to determine the re-lationship between them (G. K. Uyanık and N. G\u0026uuml;ler, 2013). Regression can be broadly classified into two categories (E. de A. L et. al. 2004):\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eUnivariate regression\u0026ndash; uses only one independent variable\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eMultivariate regression\u0026ndash; uses more than one independent variable\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eIn this study we aim to decrease the number of independent variables as much as possible while main-taining an acceptable accuracy.\u003c/p\u003e \u003cp\u003eThe general form of the linear equation for the algorithm is given in Eq.\u0026nbsp;(5) (X. Su, et. al. 2012).\u003c/p\u003e \u003cp\u003ey\u0026thinsp;=\u0026thinsp;β0\u0026thinsp;+\u0026thinsp;β1 X1\u0026thinsp;+\u0026thinsp;β2 X2 +\u0026hellip;.\u0026thinsp;+\u0026thinsp;β3 X3\u0026thinsp;+\u0026thinsp;βn Xn\u0026thinsp;+\u0026thinsp;ε\u003c/p\u003e \u003cp\u003ewhere,\u003c/p\u003e \u003cp\u003ey: dependent variable\u003c/p\u003e \u003cp\u003eX: independent variable\u003c/p\u003e \u003cp\u003eβn: regression coefficients\u003c/p\u003e \u003cp\u003eε: error.\u003c/p\u003e \u003cp\u003eThe input data is fed to the algorithm equation as training data such that X is entered as the independ-ent variable for a random value of β which maps it with the actual value of y. These obtained values are then later used for unseen test cases to make predictions.\u003c/p\u003e \u003cp\u003eAs an extension to normal linear regression, multivariable linear regression can be applied which has multiple iterations of the same linear regression data which improves the performance of the model (T. Reddy et. al. 2020). The mean square error in Linear Regression is geometrically the vertical distance between the regression line and the actual y-values which can be visualised, as depicted in Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003e as follows: (2021)\u003c/p\u003e \u003cp\u003eSource: Applied Linear Regression, Sanford Weisberg, University of Minnesota\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eDecision Tree-\u003c/strong\u003e \u003cp\u003eAs the name suggests, Decision Tree Regressors differ from the more common and popular Decision Trees or Decision Tree Classifiers but they have some similarities as well (F. Mola, 1998). The idea be-hind this model is to break down complex structures into simpler binary problems known as decisions (M. Xu et. al. 2005). It recursively splits the training set\u0026rsquo;s feature space to create a tree structure where the segmentation is applied based on some simple rules and the goal is to deduce the set of these decision rules for a reliable hierarchical model (J. R. Quinlan, 1996). Until the tree reaches a minimum sample split parameter, it continues to split the node into two, which is done keeping in mind that the depth of the tree is directly proportional to variance and inversely proportional to bias (Rahul, A. Gupta, et. al. 2021). The main benefit of using a decision tree over other modelling techniques is that it portrays more interpretable and logical rules (G. K. F. Tso and K. K. W. Yau, 2007). Another benefit of using the Decision Trees is that features with maximum information are selected automatically rejecting the remaining and thus increasing the efficiency which also takes care of the Hughes phenomenon (G. Hughes, 1998) also known as the curse of dimensionality (J. H. Cho and P. U. Kurup, 2011). In our study the predicted out-come is a continuous quantity and not a class for which the level-wise classified data is produced by using the gini impurity which is calculated using the information gain to find the optimal split location (S. Suthaharan et. al. 2016, Lerman \u0026amp; Yitzhaki 1984).\u003c/p\u003e \u003c/p\u003e \u003cp\u003eThe equation for the Gini index is given in Eq.\u0026nbsp;(6) (Gastwirth, 2011).\u003c/p\u003e \u003cp\u003eGini index\u0026thinsp;=\u0026thinsp;1\u0026thinsp;\u0026minus;\u0026thinsp;Σp\u003csup\u003e2\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eHere, p\u003csub\u003ei\u003c/sub\u003e is the probability of the occurrence of the event p\u003csub\u003ei\u003c/sub\u003e.\u003c/p\u003e \u003cp\u003eTo calculate Gini index we need Information Gain and to calculate information gain we require Entropy which is calculated as given in Eq.\u0026nbsp;(7) and Eq.\u0026nbsp;(8)( Y. Zhou et. al. 2020):\u003c/p\u003e \u003cp\u003eE(S) = \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\sum\\:}_{i=1}^{c}\\:-p\\)\u003c/span\u003e\u003c/span\u003e\u003csub\u003ei\u003c/sub\u003e log\u003csub\u003e2\u003c/sub\u003e p\u003csub\u003ei\u003c/sub\u003e\u003c/p\u003e \u003cp\u003eIG\u0026thinsp;=\u0026thinsp;Entropy(before)\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:-{\\sum\\:}_{j=1}^{k}\\:Entropy(j,\\:after)\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003cp\u003e \u003cstrong\u003eRandom Forest-\u003c/strong\u003e \u003cp\u003eThe collection of multiple decision tree predictors is known as a random forest {h(x, Θk ) k\u0026thinsp;=\u0026thinsp;1,\u0026hellip;,k} where x is the input covariate vector and Θk are the identically distributed independent random vectors, which in case of a regression problem is the unweighted average over the function as given in Eq.\u0026nbsp;(8) (M. R. Segal, 2022).\u003c/p\u003e \u003c/p\u003e \u003cp\u003eh(x) = (1/K) \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\sum\\:}_{k=1}^{k}\\:h(x,\\:\\varTheta\\:k\\:)\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003cp\u003eThis equation resamples the data without deletion therefore some data may be used multiple times in the training and some data might never be observed at all which results in greater stability\u0026ndash; this process is called bagging, a contraction of bootstrap-aggregating (L. Breiman, 2017). The samples which were ignored in the bagging process for the k-th tree are then included in the out-of-bag subset which could be used by the kth tree for performance evaluation (J. Peters et al., 2007). In this study our approach was to use the Random Forest algorithm present in the Scikit-Learn module of Python. The manual backward elimination of the feature set reduces the overall performance score. The node significance is calculated using the Gini importance, assuming that its a binary tree, formula for which is given by Eq.\u0026nbsp;(10) (1984).\u003cdiv class=\"BlockQuote\"\u003e\u003cp\u003ef\u003csub\u003eii\u003c/sub\u003e= \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\frac{{\\sum\\:}_{j:node\\:j\\:splits\\:on\\:feature\\:i}^{\\:}\\:{ni}_{j}}{{\\sum\\:}_{k\\in\\:all\\:nodes}^{\\:}\\:{ni}_{k}}\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eHere, fi sub(i)\u0026thinsp;=\u0026thinsp;the importance of feature i\u003c/p\u003e \u003cp\u003eni sub(j)\u0026thinsp;=\u0026thinsp;the importance of node j\u003c/p\u003e \u003cp\u003e \u003cstrong\u003eSupport Vector Regressor-\u003c/strong\u003e \u003cp\u003eSVM performs dichotomy functions for classification for multi-dimensional feature vectors (V. N. Vapnik and A. Chervonenkis, 1964). It was originally meant for linear classification, and updated to accommodate non-linear classification and finally extended to include regression problems as well (C. Cortes and V. Vapnik, 1995). For regression problems, Support Vector Regressor (SVR) transforms the input data into a higher dimension space in which the classes can be classified with the help of the hyperplane, such as (H. Talabani and E. Avci, 2018)\u003c/p\u003e \u003c/p\u003e \u003cp\u003e{〖X_n^ }〗_n^N\u0026thinsp;=\u0026thinsp;1 with N total samples, X is a vector of L input features\u003c/p\u003e \u003cp\u003e{〖Y_n^ }〗_n^N\u0026thinsp;=\u0026thinsp;1 and with yn \u0026isin; {\u0026minus;1,1}, The SVM regression model can be defined as\u003c/p\u003e \u003cp\u003ef(x)\u0026thinsp;=\u0026thinsp;Wt ϕ(x) where where ϕ : x \u0026rarr; ϕ(x) \u0026isin; ℝH (V. Rodriguez-Galiano et. al. 2015).\u003c/p\u003e \u003cp\u003eThe optimisation of the SVM model is subject to soft margin constraints therefore this optimisation problem can be overcome by using the kernels. There are many kinds of kernels that exist which are given from Eq.\u0026nbsp;(11) to Eq.\u0026nbsp;(14) (H. Talabani and E. Avci, 2018).\u003c/p\u003e \u003cp\u003eK\u003csub\u003elinear\u003c/sub\u003e (x, x\u0026prime;)\u0026thinsp;=\u0026thinsp;x, x\u0026prime;\u003c/p\u003e \u003cp\u003eK\u003csub\u003ePolynomial\u003c/sub\u003e (x, x\u0026prime;) = (γxx\u0026prime; + r)ρ\u003c/p\u003e \u003cp\u003eK\u003csub\u003eRBF\u003c/sub\u003e (x, x\u0026prime;)\u0026thinsp;=\u0026thinsp;exp (\u0026ndash;γ ||x \u0026ndash; x\u0026prime; ||2)\u003c/p\u003e \u003cp\u003eK\u003csub\u003eSigmoid\u003c/sub\u003e (x, x\u0026prime;)\u0026thinsp;=\u0026thinsp;tanh(γxx\u0026prime; + r)\u003c/p\u003e \u003c/div\u003e"},{"header":"6. Results and Discussion","content":"\u003cp\u003eThe analysis of the physicochemical properties of the groundwater gives an insight of the quality of wa-ter in the Gautam Buddha Nagar. This study finds a suitable machine learning model which can esti-mate WQI with one input variable and identifies the most valuable parameters contributing to the WQI prediction.\u003c/p\u003e \u003cp\u003eThe Turbidity of the groundwater in the study area ranges from 0 to 56, indicating that the water under the ground is unpredictable by location. There is very little correlation between the different water qual-ity metrics, for instance the Turbidity was noted highest at Gulavali, Sector 162 with a value of 56 NTU, but the pH of the water at the same place was found to be 8.2. Similarly the site Pumping Station, Sector 80 has 0 Turbidity and the pH is 8.12. This indicates that there is very low correlation between the input parameters, and the same was confirmed through correlation analysis.\u003c/p\u003e \u003cp\u003eThe minimum observed value of pH in this study is 7.1 and the maximum was found to be 8.9, showing that the groundwater is mostly neutral or sub alkaline having values close to the BIS standard for pH. The average value of hardness, C.O.D. and B.O.D. in this study are 588.23, 25.93 and 19.64 respectively, which are all greater than the standard BIS values. These values directly affect the quality of water and thus result in poor water quality in the study area.\u003c/p\u003e \u003cp\u003eThe accuracy score of the model when all 6 parameters were used for the training dataset was 0.9999 for Linear Regression, 0.6637 for Decision Tree ('maximum depth': 50, 'minimum sample leaf': 2, 'mini-mum samples split': 9), 0.9013 for Random Forest ('maximum depth': 12, 'minimum sample leaf': 2, 'minimum samples split': 4, 'number of estimators': 20), and 0.9997 in case of Support Vector Re-gressor ('kernel': Linear).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWhen all the input parameters are used, the accuracy is found to be the highest, with similar scores for both Linear Regression and Support Vector Regressor. Using this model will greatly reduce the com-plexity in computing the WQI when all water quality parameters are available as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e.\u003c/p\u003e \u003cp\u003eAs the cost of measuring some water quality parameters is expensive, reducing the parameter size will prove to be helpful in indicating groundwater quality in a reliable, feasible and cost-effective manner. After analysis of the correlations between WQI and the input parameter set, Machine Learning models were built using single features.\u003c/p\u003e \u003cp\u003eThe score in Fig.\u0026nbsp;\u003cspan refid=\"Fig9\" class=\"InternalRef\"\u003e9\u003c/span\u003e, when only Turbidity was used as an input parameter was 0.810 for Linear Regression, 0.571 for Decision Tree ('maximum depth': 50, 'minimum sample leaf': 1, 'minimum samples split': 9), 0.670 for Random Forest ('maximum depth': 14, 'minimum sample leaf': 2, 'minimum samples split': 3, 'number of estimators': 19) and 0.769 in case of Support Vector Regressor ('kernel': Linear).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFigure\u0026nbsp;\u003cspan refid=\"Fig10\" class=\"InternalRef\"\u003e10\u003c/span\u003e in the score when only pH was used as an input parameter was \u0026minus;\u0026thinsp;0.172 for Linear Regression, -0.361 for Decision Tree ('maximum depth': 50, 'minimum sample leaf': 3, 'minimum samples split': 11), -0.093 for Random Forest ('maximum depth': 13, 'minimum sample leaf': 4, 'minimum samples split': 2, 'number of estimators': 19) and \u0026minus;\u0026thinsp;0.048 in case of Support Vector Regressor (with rbf kernel).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eAccording to Fig.\u0026nbsp;\u003cspan refid=\"Fig11\" class=\"InternalRef\"\u003e11\u003c/span\u003e, the score when only B.O.D was used as an input parameter was \u0026minus;\u0026thinsp;0.003 for Linear Regression, 0.111 for Decision Tree ('maximum depth': 50, 'minimum sample leaf': 2, 'minimum samples split': 9), 0.245 for Random Forest ('maximum depth': 14, 'minimum sample leaf': 3, 'minimum samples split': 2, 'number of estimators': 21) and \u0026minus;\u0026thinsp;0.042 in case of Support Vector Regressor (with rbf kernel)\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFigure\u0026nbsp;\u003cspan refid=\"Fig12\" class=\"InternalRef\"\u003e12\u003c/span\u003e in the score when only D.O. was used as an input parameter was \u0026minus;\u0026thinsp;0.069 for Linear Regression, -0.140 for Decision Tree ('maximum depth': 50, 'minimum sample leaf': 2, 'minimum samples split': 11), -0.022 for Random Forest ('maximum depth': 14, 'minimum sample leaf': 3, 'minimum samples split': 4, 'number of estimators': 21) and \u0026minus;\u0026thinsp;0.029 in case of Support Vector Regressor ('kernel': Linear).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe score, according to Fig.\u0026nbsp;\u003cspan refid=\"Fig13\" class=\"InternalRef\"\u003e13\u003c/span\u003e, when only Hardness was used as an input parameter was \u0026minus;\u0026thinsp;0.021 for Linear Regression, -0.141 for Decision Tree ('maximum depth': 50, 'minimum sample leaf': 2, 'minimum samples split': 11), -0.019 for Random Forest ('maximum depth': 14, 'minimum sample leaf': 3, 'minimum samples split': 4, 'number of estimators': 21) and \u0026minus;\u0026thinsp;0.049 in case of Support Vector Regressor ('kernel': Linear).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe score in Fig.\u0026nbsp;\u003cspan refid=\"Fig14\" class=\"InternalRef\"\u003e14\u003c/span\u003e, when only C.O.D. was used as an input parameter was \u0026minus;\u0026thinsp;0.021 for Linear Regression, -0.141 for Decision Tree ('maximum depth': 50, 'minimum sample leaf': 3, 'minimum samples split': 10), -0.002 for Random Forest ('maximum depth': 13, 'minimum sample leaf': 3, 'minimum samples split': 3, 'number of estimators': 21) and \u0026minus;\u0026thinsp;0.049 in case of Support Vector Regressor ('kernel': Linear).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e"},{"header":"7 Conclusion","content":"\u003cp\u003eThe study reviewed and calculated WQI values for 51 locations in Gautum Buddha Nagar and revealed that the groundwater is highly contaminated and not fit for human consumption. The calculated WQI values were far higher than the standard and permissible limits due to high sub-indices values of B.O.D., C.O.D. and Turbidity, which range from 112.855 to 442.579.\u003c/p\u003e \u003cp\u003eThe calculation of WQI is a lengthy and costly method therefore the aim of this study was to develop a Machine Learning model which can accurately determine the WQI using just one input parameter. Our findings showed that the best performing supervised learning algorithm was Linear Regression with a score of 0.8105 when using only Turbidity as the input variable and the worst performing algorithm was Decision Tree with a score of -0.141 using Hardness as the input variable. The Machine Learning models will not only help in estimating WQI with a single water quality parameter but also aid future studies to further develop in this area.\u003c/p\u003e \u003cp\u003eThe study strongly suggests tracking the pollution sources in order to prevent water in the study area from getting contaminated any further. Since Gautam Buddha Nagar is home to many large scale facto-ries and industries, their presence might be the primary source of contamination. Prevention from dete-rioration requires frequent inspections and monitoring of the water quality, for which WQI plays an important role.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eFunding:\u0026nbsp;\u003c/strong\u003eNo Funding support is provided for this paper publication.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting Interests:\u0026nbsp;\u003c/strong\u003eNo benefits in any form have been received or will be received from a commercial party related directly or indirectly to the subject of this article. All authors declare no conflict of interest for this article.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor\u0026apos;s Contribution:\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n \u003cli\u003eKakoli Banerjee: Conceptualization, Literature Review, Drafting the \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;manuscript.\u003c/li\u003e\n \u003cli\u003ePradeep Kumar: Methodology, Data Analysis, Review \u0026amp; Editing.\u003c/li\u003e\n \u003cli\u003eAjay Kumar: Literature Review, Proofreading.\u003c/li\u003e\n \u003cli\u003eKamal Upreti: Data Collection, Review \u0026amp; Editing, Final Approval.\u003c/li\u003e\n \u003cli\u003eShubham Mahajan: Visualization, Literature Review.\u003c/li\u003e\n \u003cli\u003eMohammad Shahnawaz Nasir: Methodology, Data Analysis, Review \u0026amp; Editing.\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003e\u003cstrong\u003eEthics Approval and Consent to participate:\u0026nbsp;\u003c/strong\u003eNo Participation of humans take place in this implementation process.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent to Publish:\u0026nbsp;\u003c/strong\u003eAll authors are showing their consent to publish this paper .\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Availability Statement:\u003c/strong\u003e The data that support the findings of this study are available from the corresponding author upon reasonable request.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eAbbasi T. and Abbasi,S. A. (2012). Water-Quality Indices, \u003cem\u003eWater Quality Indices\u003c/em\u003e. 353\u0026ndash;356. doi: 10.1016/b978-0-444-54304-2.00016-6.\u003c/li\u003e\n\u003cli\u003eAhmed, U., Mumtaz, R., Anwar, H., Shah, A. A., Irfan, R., \u0026amp; Garc\u0026iacute;a-Nieto, J. (2020). Efficient Water Quality Prediction Using Supervised Machine Learning, \u003cem\u003eWater\u003c/em\u003e, 11(11), 2210.\u003c/li\u003e\n\u003cli\u003eAkhtar, N., Ishak, M. I. S., Ahmad, M. I., Umar, K., Md Yusuff, M. S., Anees, M. T., ... \u0026amp; Ali Almanasir, Y. K. (2021). Modification of the Water Quality Index (WQI) Process for Simple Calculation Us-ing the Multi-Criteria Decision-Making (MCDM) Method: A Review, \u003cem\u003eWater\u003c/em\u003e, 13(7). 905. doi: 10.3390/w13070905.\u003c/li\u003e\n\u003cli\u003eAldhyani T. H. H., Al-Yaari M., Alkahtaniv H., and Maashiv M., (2020). Water Quality Prediction Using Artifi-cial Intelligence Algorithms, \u003cem\u003eAppl. Bionics Biomech.,\u003c/em\u003e 6659314, Dec. 2020.\u003c/li\u003e\n\u003cli\u003eBanda T. D. and Kumarasamy M. V., (2020). Development of Water Quality Indices (WQIs): A Review, Polish Journal of Environmental Studies, 29(3). 2011\u0026ndash;2021. doi: 10.15244/pjoes/110526.\u003c/li\u003e\n\u003cli\u003eBanerjee K. and Prasad R. A., (2017). Reference based inter chromosomal similarity-based DNA sequence compression algorithm, \u003cem\u003ein 2017 International Conference on Computing, Communication and Automation (ICCCA)\u003c/em\u003e, 234\u0026ndash;238.\u003c/li\u003e\n\u003cli\u003eBanerjee K., Santhosh Kumar M. B., Tilak L. N., and S. Vashistha, (2021). Analysis of Groundwater Quality Using GIS-Based Water Quality Index in Noida, Gautam Buddh Nagar, Uttar Pradesh (UP), India, \u003cem\u003eLecture Notes in Electrical Engineering\u003c/em\u003e. 171\u0026ndash;187. doi: 10.1007/978-981-16-3067-5_14.\u003c/li\u003e\n\u003cli\u003eBanerjee, K., Bali, V., Nawaz, N., Bali, S., Mathur, S., Mishra, R. K., \u0026amp; Rani, S. (2022). A Machine-Learning Approach for Prediction of Water Contamination Using Lat-itude, Longitude, and Elevation, \u003cem\u003eWater,\u003c/em\u003e 14(5), 728.\u003c/li\u003e\n\u003cli\u003eBhardwaj R. M., (2005). Water Quality Monitoring in India - Achievements and Constraints, \u003cem\u003eIWG-Env Joint Work Session on Water Statistics\u003c/em\u003e.\u003c/li\u003e\n\u003cli\u003eBreiman, L. (2017). Classification and Regression Trees. \u003cem\u003eRoutledge\u003c/em\u003e.\u003c/li\u003e\n\u003cli\u003eBrown M., McCLELLAND N. I., Deininger R. A., and O\u0026rsquo;connor M. F., (2013). A water quality index \u0026ndash; crashing the psychological barrier, \u003cem\u003eAdvances in Water Pollution Research\u003c/em\u003e. 787\u0026ndash;797. doi: 10.1016/b978-0-08-017005-3.50067-0.\u003c/li\u003e\n\u003cli\u003eBui, D. T., Khosravi, K., Tiefenbacher, J., Nguyen, H., \u0026amp; Kazakis, N. (2020). Improving prediction of water quality indices using novel hybrid machine-learning algorithms, Sci. Total Environ., 721, 137612.\u003c/li\u003e\n\u003cli\u003eCanada C. E. and Canada. (2003). National Guidelines and Standards Office, Canadian Water Quality Guide-lines for the Protection of Aquatic Life: Nitrate Ion. National Guidelines and Standards Office.\u003c/li\u003e\n\u003cli\u003eCho, J. H., \u0026amp; Kurup, P. U. (2011). Decision tree approach for classification and dimensionality reduction of electronic nose data, Sensors and Actuators B: Chemical, 160(1). 542\u0026ndash;548. doi: 10.1016/j.snb.2011.08.027.\u003c/li\u003e\n\u003cli\u003eCitaristi, I. (2022). United Nations Office for Outer Space Affairs\u0026mdash;UNOOSA. In The Europa Directory of International Organizations, 247-248. Routledge.\u003c/li\u003e\n\u003cli\u003eCortes C. and Vapnik, V. (1995). Support-vector networks, \u003cem\u003eMach. Learn\u003c/em\u003e., 20(3), 273\u0026ndash;297.\u003c/li\u003e\n\u003cli\u003eCrisci, C., Ghattas, B., \u0026amp; Perera, G. (2012). A review of supervised machine learning algorithms and their applications to ecological data, \u003cem\u003eEcological Modelling\u003c/em\u003e, 240. 113\u0026ndash;122. doi: 10.1016/j.ecolmodel.2012.03.001.\u003c/li\u003e\n\u003cli\u003eDas Kangabam, R., Bhoominathan, S. D., Kanagaraj, S., \u0026amp; Govindaraju, M., (2017). Development of a water quality index (WQI) for the Loktak Lake in India, Applied Water Science, 7(6), 2907\u0026ndash;2918.\u003c/li\u003e\n\u003cli\u003ede A. Lima Neto, E., de Carvalho, F. A., \u0026amp; Tenorio, C. P. (2004). Univariate and Multi-variate Linear Regression Methods to Predict Interval-Valued Features, \u003cem\u003eLecture Notes in Computer Science.\u003c/em\u003e 526\u0026ndash;537. doi: 10.1007/978-3-540-30549-1_46.\u003c/li\u003e\n\u003cli\u003eEdition, F. (2011). Guidelines for drinking-water quality. WHO chronicle, 38(4), 104-8.\u003c/li\u003e\n\u003cli\u003eEdition, F. (2011). Guidelines for drinking-water quality. WHO chronicle, 38(4), 104-8..\u003c/li\u003e\n\u003cli\u003eFlach P. A., (2001). On the state of the art in machine learning: A personal review, Artificial Intelligence, 131(1\u0026ndash;2). 199\u0026ndash;222. doi: 10.1016/s0004-3702(01)00125-4.\u003c/li\u003e\n\u003cli\u003eGafoor, A., Rauf, A., Arif, M., \u0026amp; Muzaffar, W. (1994). Chemical composition of effluents from different industries of the Faisalabad Pakistan. J. Agri. Sci, 33, 73-74.\u003c/li\u003e\n\u003cli\u003eGastwirth, (2011). The estimation of the Lorenz curve and Gini index, \u003cem\u003eRev. Econ. Stat\u003c/em\u003e.\u003c/li\u003e\n\u003cli\u003eHameed M., Sharqi S. S., Yaseen Z. M., Afan H. A., A. Hussain, and A. Elshafie, (2017). Application of artifi-cial intelligence (AI) techniques in water quality index prediction: a case study in tropical region, Malaysia, \u003cem\u003eNeural Comput. Appl\u003c/em\u003e., 28(1), 893\u0026ndash;905.00\u003c/li\u003e\n\u003cli\u003eHassan, M. M., Hassan, M. M., Akter, L., Rahman, M. M., Zaman, S., Hasib, K. M., ... \u0026amp; Mollick, S, (2021). Efficient prediction of water quality index (WQI) using machine learning algo-rithms, \u003cem\u003eHuman-Centric Intelligent Systems\u003c/em\u003e, 1(3\u0026ndash;4), 86.\u003c/li\u003e\n\u003cli\u003eHiran, K. K., Khazanchi, D., Vyas, A. K., \u0026amp; Padmanaban, S., (2017). Machine Learning for Sustainable De-velopment.\u003c/li\u003e\n\u003cli\u003eHorton R. K., (2015). An Index Number System for Rating Water Quality, \u003cem\u003eJournal of the Water Pollution Control Federation,\u003c/em\u003e 37(3), 300\u0026ndash;306.\u003c/li\u003e\n\u003cli\u003eHughes, G. (1998). On the mean accuracy of statistical pattern recognizers,\u0026rdquo; \u003cem\u003eIEEE Transactions on Infor-mation Theory,\u003c/em\u003e 14(1). 55\u0026ndash;63. doi: 10.1109/tit.1968.1054102.\u003c/li\u003e\n\u003cli\u003eHurley T., Sadiq R., and Mazumder A., (2012). Adaptation and evaluation of the Canadian Council of Min-isters of the Environment Water Quality Index (CCME WQI) for use as an effective tool to characterize drinking source water quality, \u003cem\u003eWater Res.,\u003c/em\u003e 46, 11, 3544\u0026ndash;3552.\u003c/li\u003e\n\u003cli\u003eIbrahim and Salmon S., (2020). Chemical composition of Faisalabad - city [Pakistan] sewage effluent: I. Nitrogen, phosphorus and potassium contents, \u003cem\u003eJ. Agric. Res\u003c/em\u003e.\u003c/li\u003e\n\u003cli\u003eJha M. K., Shekhar A., and Jenifer M. A., (2020). Assessing groundwater quality for drinking water supply using hybrid fuzzy-GIS-based water quality index, \u003cem\u003eWater Res.,\u003c/em\u003e 179, 115867.\u003c/li\u003e\n\u003cli\u003eJiang T., Gradus J. L., and Rosellini A. J., (2020). Supervised Machine Learning: A Brief Primer, \u003cem\u003eBehav. Ther.,\u003c/em\u003e 51(5), 675\u0026ndash;687.\u003c/li\u003e\n\u003cli\u003eJudran H. and Kumar A., (2020). Evaluation of water quality of Al-Gharraf River using the water quality index (WQI), Modeling Earth Systems and Environment, 6(3), 1581\u0026ndash;1588.\u003c/li\u003e\n\u003cli\u003eKadam, A. K., Wagh, V. M., Muley, A. A., Umrikar, B. N., \u0026amp; Sankhua, R. N. (2019). Prediction of water quality index using artificial neural network and multiple linear regression modelling approach in Shivganga River basin, India. Modeling Earth Systems and Environment, 5, 951-962.\u003c/li\u003e\n\u003cli\u003eKamboj N. and Kamboj V., (2019). Water quality assessment using overall index of pollution in riv-erbed-mining area of Ganga-River Haridwar, India, Water Science, 33(1). 65\u0026ndash;74. doi: 10.1080/11104929.2019.1626631.\u003c/li\u003e\n\u003cli\u003eLerman, R. I., \u0026amp; Yitzhaki, S. (1984). A note on the calculation and interpretation of the Gini index. Economics Letters, 15(3-4), 363-368.\u003c/li\u003e\n\u003cli\u003eLinstone H. A. and Turoff M. (1975). The Delphi Method: Techniques and Applications. \u003cem\u003eWestview Press\u003c/em\u003e.\u003c/li\u003e\n\u003cli\u003eLumb A., A. L., Halliwell, D., \u0026amp; Tribeni Sharma, T. S., (2006). Application of CCME Water Quality Index to monitor water quality: a case study of the Mackenzie River Basin, Canada,\u0026rdquo; Environ. Monit. Assess., 113(1\u0026ndash;3), 411\u0026ndash;429.\u003c/li\u003e\n\u003cli\u003eLumb A., A. L., Halliwell, D., \u0026amp; Tribeni Sharma, T. S.,, (2011). A Review of Genesis and Evolution of Water Quality In-dex (WQI) and Some Future Directions, Water Quality, Exposure and Health, 3(1). 11\u0026ndash;24. doi: 10.1007/s12403-011-0040-0.\u003c/li\u003e\n\u003cli\u003eMara, D. D., Cairncross, S., \u0026amp; World Health Organization. (1989). Guidelines for the safe use of wastewater and excreta in agriculture and aquaculture: measures for public health protection. \u003cem\u003eWorld Health Organization\u003c/em\u003e.\u003cem\u003e.\u003c/em\u003e\u003c/li\u003e\n\u003cli\u003eMateo-Sagasta, J., Zadeh, S. M., Turral, H., \u0026amp; Burke, J. (2017). Water pollution from agriculture: a global review. Executive summary. Rome, Italy: FAO Colombo, Sri Lanka: International Water Management Institute (IWMI). CGIAR Research Program on Water, Land and Ecosystems (WLE).,.\u003c/li\u003e\n\u003cli\u003eMeyer M. A. and Booker J. M., (2001). Eliciting and Analyzing Expert Judgment: A Practical Guide. SIAM.\u003c/li\u003e\n\u003cli\u003eMola, F. (1998). Classification and Regression Trees Software and New Developments, Studies in Classifi-cation, Data Analysis, and Knowledge Organization. 311\u0026ndash;318. doi: 10.1007/978-3-642-72253-0_42.\u003c/li\u003e\n\u003cli\u003eMurthy, D. N. P., Xie, M., \u0026amp; Jiang, R. (2004). Weibull models. John Wiley \u0026amp; Sons.\u003c/li\u003e\n\u003cli\u003eOsisanwo, F. Y., J. E. T. Akinsola, O. Awodele, J. O. Hinmikaiye, O. Olakanmi, and J. Akinjobi, (2017). Supervised Machine Learning Algorithms: Classification and Comparison,\u0026rdquo; Interna-tional Journal of Computer Trends and Technology, 48(3). 128\u0026ndash;138. doi: 10.14445/22312803/ijctt-v48p126.\u003c/li\u003e\n\u003cli\u003ePeters, J., De Baets, B., Verhoest, N. E., Samson, R., Degroeve, S., De Becker, P., \u0026amp; Huybrechts, W., (2007). Random forests as a tool for ecohydrological distribution modelling, \u003cem\u003eEcol. Modell.\u003c/em\u003e, 207(2\u0026ndash;4), 304\u0026ndash;318.\u003c/li\u003e\n\u003cli\u003ePonsadailakshmi S., Ganapathy Sankari S., Mythili Prasanna S., and Madhurambal G., (2018). Evaluation of water quality suitability for drinking using drinking water quality index in Nagapattinam district, Tamil Nadu in Southern India, \u003cem\u003eGroundwater for Sustainable Development,\u003c/em\u003e 6. 43\u0026ndash;49. doi: 10.1016/j.gsd.2017.10.005.\u003c/li\u003e\n\u003cli\u003ePowell, C. (2003). The Delphi technique: myths and realities, \u003cem\u003eJ. Adv. Nurs\u003c/em\u003e., 41(4), 376\u0026ndash;382.\u003c/li\u003e\n\u003cli\u003ePuntanen S., (2010). Linear Regression Analysis: Theory and Computing by Xin Yan, Xiao Gang Su, \u003cem\u003eInter-national Statistical Review,\u003c/em\u003e 78(1). 144\u0026ndash;144. doi: 10.1111/j.1751-5823.2010.00109.\u003c/li\u003e\n\u003cli\u003eQuinlan J. R., (1996). Learning decision tree classifiers, \u003cem\u003eACM Computing Surveys\u003c/em\u003e, 28(1). 71\u0026ndash;72. doi: 10.1145/234313.234346.\u003c/li\u003e\n\u003cli\u003eRabeiy R. E., (2018). Assessment and modeling of groundwater quality using WQI and GIS in Upper Egypt area, \u003cem\u003eEnviron. Sci. Pollut. Res. Int.,\u003c/em\u003e 25(31), 30808\u0026ndash;30817.\u003c/li\u003e\n\u003cli\u003eRahul, A. Gupta, A. Bansal, and K. Roy, (2021). Solar Energy Prediction using Decision Tree Regres-sor, \u003cem\u003e2021 5th International Conference on Intelligent Computing and Control Systems (ICICCS)\u003c/em\u003e. doi: 10.1109/iciccs51141.2021.9432322.\u003c/li\u003e\n\u003cli\u003eRamadhan M. M., Sitanggang I. S., Nasution F. R., and Ghifari, A. (2017). Parameter Tuning in Random Forest Based on Grid Search Method for Gender Classification Based on Voice Frequency, \u003cem\u003eDEStech Transac-tions on Computer Science and Engineering\u003c/em\u003e, 2017. doi: 10.12783/dtcse/cece2017/14611.\u003c/li\u003e\n\u003cli\u003eRavikumar P., Somashekar R. K., and Angami M., (2011). Hydrochemistry and evaluation of groundwater suita-bility for irrigation and drinking purposes in the Markandeya River basin, Belgaum District, Karnataka State, India, \u003cem\u003eEnviron. Monit. Assess\u003c/em\u003e., 173(1\u0026ndash;4), 459\u0026ndash;487.\u003c/li\u003e\n\u003cli\u003eReddy T., Sewpersad Y., and Yen Y.-C. J., (2020). Characterisation of the primary heat replacement element event for a horizontal electric water heater, \u003cem\u003e2020 International SAUPEC/RobMech/PRASA Conference.\u003c/em\u003edoi: 10.1109/saupec/robmech/prasa48453.2020.9041057.\u003c/li\u003e\n\u003cli\u003eRodriguez-Galiano, V., Sanchez-Castillo, M., Chica-Olmo, M., \u0026amp; Chica-Rivas, M. J. O. G. R., (2015). Machine learning predictive models for mineral prospectivity: An evaluation of neural networks, random forest, regression trees and support vector machines, \u003cem\u003eOre Geology Reviews\u003c/em\u003e, 71, 804\u0026ndash;818. doi: 10.1016/j.oregeorev.2015.01.001.\u003c/li\u003e\n\u003cli\u003eRoy R., (2018). An Approach to Develop an Alternative Water Quality Index Using FLDM. \u003cem\u003eWater Re-sources Development and Management.\u003c/em\u003e 51\u0026ndash;68. doi: 10.1007/978-981-10-6205-6_3.\u003c/li\u003e\n\u003cli\u003eSaffran, K., Cash, K., \u0026amp; Hallard, K. (2017). CCME Water Quality Index User\u0026rsquo;s Manual 2017 Update. Can. Water Qual. Guidel. Prot. Aquat. Life, 1-5. https://ccme.ca/en/res/wqimanualen.pdf (accessed.\u003c/li\u003e\n\u003cli\u003eS\u0026aacute;nchez, E., Colmenarejo, M. F., Vicente, J., Rubio, A., Garc\u0026iacute;a, M. G., Travieso, L., \u0026amp; Borja, R., (2007). Use of the water quality index and dissolved oxygen deficit as simple indicators of watersheds pollution, \u003cem\u003eEcological Indicators,\u003c/em\u003e 7(2). 315\u0026ndash;328. doi: 10.1016/j.ecolind.2006.02.005.\u003c/li\u003e\n\u003cli\u003eSegal M. R., (2022). Machine Learning Benchmarks and Random Forest Regression. Accessed: Ma. [Online]. Available: https://escholarship.org/content/qt35x3v9t4/qt35x3v9t4.pdf?t=krnmaw\u003c/li\u003e\n\u003cli\u003eShabbir R. and Ahmad S. S., (2015). Use of Geographic Information System and Water Quality Index to As-sess Groundwater Quality in Rawalpindi and Islamabad, \u003cem\u003eArab. J. Sci. Eng.,\u003c/em\u003e 40(7), 2033\u0026ndash;2047.\u003c/li\u003e\n\u003cli\u003eShah K. A. and Joshi G. S., (2017). Evaluation of water quality index for River Sabarmati, Gujarat, India, \u003cem\u003eApplied Water Science\u003c/em\u003e, 7(3). 1349\u0026ndash;1358. doi: 10.1007/s13201-015-0318-7.\u003c/li\u003e\n\u003cli\u003eShailaja, K., Seetharamulu, B., \u0026amp; Jabbar, M. A. (2018, March). Machine learning in healthcare: A review. In 2018 Second international conference on electronics, communication and aerospace technology (ICECA) (pp. 910-914). IEEE..\u003c/li\u003e\n\u003cli\u003eSharma T., Banerjee K., Mathur S., and Bali V., (2020). Stress Analysis Using Machine Learning Techniques, \u003cem\u003eInternational Journal of Advanced Science and Technology\u003c/em\u003e, 29(3), 14654\u0026ndash;14665.\u003c/li\u003e\n\u003cli\u003eSillberg, C., Kullavanijaya P., and O. Chavalparit, (2019). Water quality classification by integration of at-tribute-realization and support vector machine for the Chao phraya river, \u003cem\u003eInż. Ekol.,\u003c/em\u003e 22(9), 70\u0026ndash;86.\u003c/li\u003e\n\u003cli\u003eSingh and Khan, (2019). Ground water quality assessment of Dhankawadi ward of Pune by using GIS, \u003cem\u003eInt. j. geomat. geosci\u003c/em\u003e.\u003c/li\u003e\n\u003cli\u003eSingh S. and Hussian A., (2016). Water quality index development for groundwater quality assessment of Greater Noida sub-basin. Uttar Pradesh, India\u003cem\u003e, Cogent Engineering\u003c/em\u003e, 3, (1). 1177155. doi: 10.1080/23311916.2016.1177155.\u003c/li\u003e\n\u003cli\u003eSingh, A., Thakur, N., \u0026amp; Sharma, A. (2016). A review of supervised machine learning algorithms,\u0026rdquo; \u003cem\u003e2016 3rd International Conference on Computing for Sustainable Global Development (INDIACom).\u003c/em\u003e [Online]. Available: https://ieeexplore.ieee.org/abstract/document/7724478\u003c/li\u003e\n\u003cli\u003eSrinivas, P., Kumar, G. P., Prasad, A. S., \u0026amp; Hemalatha, T. (2017). Generation of groundwater quality index map-A case study, \u003cem\u003eRes. Rep. Dept. Civ. Environ\u003c/em\u003e. Eng. Fac. Eng Saitama Univ.\u003c/li\u003e\n\u003cli\u003eStandard, I. (2013). Drinking water-specification. Ethiopian Standards Agency, Addis Ababa, Ethiopia.\u003c/li\u003e\n\u003cli\u003eSu X., Yan X., and Tsai C.-L., (2012). Linear regression, \u003cem\u003eWiley Interdiscip. Rev. Comput. Stat\u003c/em\u003e., 4(3), 275\u0026ndash;294.\u003c/li\u003e\n\u003cli\u003eSutadian, A. D., Muttil, N., Yilmaz, A. G., \u0026amp; Perera, B. J. C. (2016). Development of river water quality indices\u0026mdash;a review. \u003cem\u003eEnvironmental monitoring and assessment\u003c/em\u003e, 188, 1-29. doi: 10.1007/s10661-015-5050-0.\u003c/li\u003e\n\u003cli\u003eSuthaharan S. (2016) \u0026ldquo;Machine Learning Models and Algorithms for Big Data Classification,\u0026rdquo; Integrated Series in Information Systems. doi: 10.1007/978-1-4899-7641-3.\u003c/li\u003e\n\u003cli\u003eŞener Ş., Şener E., and Davraz A., (2017). Evaluation of water quality using water quality index (WQI) method and GIS in Aksu River (SW-Turkey), \u003cem\u003eSci. Total Environ.,\u003c/em\u003e 584\u0026ndash;585, 131\u0026ndash;144.\u003c/li\u003e\n\u003cli\u003eTalabani H. and Avci E. (2018). Performance Comparison of SVM Kernel Types on Child Autism Disease Database, \u003cem\u003ein 2018 International Conference on Artificial Intelligence and Data Processing (IDAP),\u003c/em\u003e 1\u0026ndash;5.\u003c/li\u003e\n\u003cli\u003eTalabani H. and Avci E., (2018). Performance Comparison of SVM Kernel Types on Child Autism Disease Database, \u003cem\u003e2018 International Conference on Artificial Intelligence and Data Processing (IDAP).\u003c/em\u003e doi: 10.1109/idap.2018.8620924.\u003c/li\u003e\n\u003cli\u003eTso, G. K., \u0026amp; Yau, K. K. (2007). Predicting electricity energy consumption: A comparison of regres-sion analysis, decision tree and neural networks, \u003cem\u003eEnergy The International Journal,\u003c/em\u003e 1761\u0026ndash;1768.\u003c/li\u003e\n\u003cli\u003eTunc Dede O., Dede O. T., Telci I. T., and Aral M. M., (2013). The Use of Water Quality Index Models for the Evaluation of Surface Water Quality: A Case Study for Kirmir Basin, Ankara, Turkey, \u003cem\u003eWater Quality, Exposure and Health\u003c/em\u003e, 5(1). 41\u0026ndash;56. doi: 10.1007/s12403-013-0085-3.\u003c/li\u003e\n\u003cli\u003eTyagi S., Sharma B., Singh v, and Dobhal R., (2020). Water Quality Assessment in Terms of Water Quality Index,\u0026rdquo; \u003cem\u003eAmerican Journal of Water Resources\u003c/em\u003e, 1, (3). 34\u0026ndash;38, 2020. doi: 10.12691/ajwr-1-3-3.\u003c/li\u003e\n\u003cli\u003eUddin, M. G., Olbert, A. I., \u0026amp; Nash, S., (2020). Assessment of water quality using Water Quality Index (WQI), \u003cem\u003eEcol. Indic.,\u003c/em\u003e 85, 966\u0026ndash;982.\u003c/li\u003e\n\u003cli\u003eUN-Water. https://www.unwater.org/water-facts/ (accessed Mar. 26, 2022).\u003c/li\u003e\n\u003cli\u003eUyanık G. K. and G\u0026uuml;ler, N. (2013) A study on multiple linear regression analysis,\u0026rdquo; 4th International Con-ference on New Horizons in Education, vol. 106, p. Pages 234\u0026ndash;240, Dec. 2013.\u003c/li\u003e\n\u003cli\u003eVan Dijk, J. A. (2090). Delphi questionnaires versus individual and group interviews: A comparison case, \u003cem\u003eTechnol. Forecast. Soc. Change\u003c/em\u003e, 37(3), 293\u0026ndash;304.\u003c/li\u003e\n\u003cli\u003eVapnik V. N. and Chervonenkis A., (1964). A note on one class of perceptrons, Automation and Remote Control 25(1).\u003c/li\u003e\n\u003cli\u003eWagh, V. M., Panaskar, D. B., Muley, A. A., \u0026amp; Mukate, S. V., (2017). Groundwater suitability evaluation by CCME WQI model for Kadava River Basin, Nashik, Maharashtra, India, Modeling Earth Systems and Envi-ronment, 3(2), 557\u0026ndash;565.\u003c/li\u003e\n\u003cli\u003eWatanachaturaporn M. Xu, Varshney P. P., and Arora M., (2005). Decision tree regression for soft classifica-tion of remote sensing data, Remote Sensing of Environment, 97(3). 322\u0026ndash;336. doi: 10.1016/j.rse.2005.05.008.\u003c/li\u003e\n\u003cli\u003eYadav, N., Banerjee K., and Bali V., (2020). A Survey on Fatigue Detection of Workers Using Machine Learning, \u003cem\u003eIJEHMC\u003c/em\u003e, 11(3), 1\u0026ndash;8.\u003c/li\u003e\n\u003cli\u003eZhou Y., Cheng G., Jiang S., and Dai M., (2020). Building an efficient intrusion detection system based on feature selection and ensemble classifier, \u003cem\u003eComputer Networks,\u003c/em\u003e 174. 107247. doi: 10.1016/j.comnet.2020.107247.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Water Quality Index, Groundwater, Supervised Machine Learning, Regression, environmental pollutants, data acquisition","lastPublishedDoi":"10.21203/rs.3.rs-5859247/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5859247/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eWater quality is an essential measure for maintaining health and quality of life. In this study, the water quality index has been computed for Gautam Buddha Nagar using the Brown et al. method by collection of 51 groundwater samples. For this purpose, six physicochemical water parameters were analysed, namely pH, Hardness, Turbidity, C.O.D., D.O. and B.O.D. The readings indicate that the groundwater condition of Gautam Buddha Nagar is extremely poor. The computation of the Water Quality Index is a complex task. The study found that when all input variables are available, Machine Learning Techniques can be employed to vastly reduce the complexity in the computation while giving a high accuracy of 99.99% using Linear Regression, followed by an accuracy of 99.97% using Support Vector Regressor. Collection of input data is a time-consuming and costly process, therefore, the dimensionality of the input data was reduced through correlation analysis in an attempt to compute the water quality index by using just a single parameter. The best score of 81.05% was obtained using Linear Regression when Turbidity was used as the only feature, due to its high correlation with the target variable. The algorithms used for this analysis are: Linear Regression, Support Vector Regressor, Decision Tree and Random Forest Regressor.\u003c/p\u003e","manuscriptTitle":"Computation of Water Quality Index and Its Estimation Using Machine Learning Techniques","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-01-29 10:48:14","doi":"10.21203/rs.3.rs-5859247/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"119e742d-5c0c-4d6d-8b33-95ba6712d365","owner":[],"postedDate":"January 29th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-03-05T05:51:18+00:00","versionOfRecord":[],"versionCreatedAt":"2025-01-29 10:48:14","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-5859247","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5859247","identity":"rs-5859247","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.