Automating image classification by using multi-resolution spaceborne imagery, cloud-based machine learning algorithms and open-source software and data | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Automating image classification by using multi-resolution spaceborne imagery, cloud-based machine learning algorithms and open-source software and data Sunil Bhaskaran, Arvindh Sharma, Sanjiv Bhatia, Stuti Mishra This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6497131/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract This paper describes the extraction of critical terrestrial features from multispectral (MSS) spaceborne imagery. We present a methodology to extract terrestrial features from spatiotemporal datasets by using machine learning algorithms, open-source software, and datasets. We used the Amazon Web Services sage-maker framework to conduct full data life cycle (FDLC) projects and automate steps from data acquisition to data mining and visualization. The results demonstrate a robust and efficient model to conduct image analyses with spatio-temporal datasets to extract variables that may be useful for a wide range of applications. The results have significant implications to analyze time-series of spaceborne imagery for both near-real time and other applications that are driven by data from the archives. Urban Land Cover Classification Cloud-Based Geospatial Analytics Big Data in Remote Sensing Multispectral and Hyperspectral Imaging Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 Figure 10 Figure 11 Figure 12 Figure 13 Figure 14 Figure 15 Figure 16 Figure 17 Introduction One of the major applications of remote sensing data is deducing the land use and land cover (LULC) of earth’s land surface [ 1 ]. The land cover identifies the surface’s biophysical attributes while the land use identifies the human-intended purpose for the land cover [ 2 ]. Tracking changes in the land cover is crucial in estimating the impacts of anthropogenic stressors, the extent of hazards such as floods and forest fires, and monitoring environmental outcomes such as climate change [ 3 ] [ 4 ]. Earth observing (EO) satellite data, combined with improved techniques to estimate land cover classes, have paved the way for tracking changes in spatio-temporal datasets. Many of these large EO datasets are publicly available and can aid researchers and policy-makers to utilize land cover change information in near-real-time [ 5 ]. The major challenges in converting the raw EO data into useful land cover/land use information include the identification and development of techniques to process the data and predict classification of the surface, and the ability to scale the solution to ingest, process, and store the large datasets. One of the key techniques used in land cover classification is using multi-spectral satellite imagery with pixel-based image analysis [ 6 ]. This is made easier by the public availability of high-quality multispectral data from programs such as NASA’s Landsat [ 7 ] and ESA’s Sentinel-2 [ 8 ], which have good spatial coverage and temporal frequency of revisit over most of earth’s surface. In this method, each pixel contains information collected in multiple spectral bands. For example, Sentinel-2 (S-2) senses radiations from the earth’s surface in thirteen different bands in visible, near infrared and shortwave infrared wavelengths [ 9 ]. Different classes, for example, water and vegetation, would have different spectral signatures registered on the sensors. By identifying the spectral signatures of different land classes, each pixel in the image can be classified to one of the identified classes. Beyond pixel-based classification, object-based methods identify spectrally homogenous pixels and segment them into objects [ 10 ]. While object-based method removes salt-and-pepper noise that is pervasive in pixel-based classification, it is still prone to segmentation errors and is less simple compared to the latter. On the other hand, pixel-based multispectral classification can be enhanced by incorporating texture features derived from Synthetic Aperture Radar (SAR) data, such as those produced by the Sentinel-1 (S-1) program [ 11 ]. SAR actively transmits microwave signals towards the earth’s surface and constructs an image based on the portion of the signal scattered back to the sensor [ 12 ]. The advantage is that the microwaves are less affected by cloud cover and can be collected at any time of the day [ 1 ]. Texture refers to the inter-relationship between neighboring pixels in an image window and can be used to distinguish between classes since different classes typically have different textures [ 13 ]. Texture is quantified by constructing a Gray Level Covariance Matrix (GLCM), and techniques incorporating it in pixel-based methods have shown promise. Iyyappan, et al. [ 14 ] observed that integrating texture metrics into multi-spectral optical bands can improve classification accuracy by 15%. They attribute this improvement to the enhanced ability to distinguish between all classes, particularly plantation and rice crop areas, which is not well discriminated when using only optical spectral data. The choice of machine learning (ML) classifier used in the pixel-based technique can also have a significant impact on accuracy. Some of the often-used land-cover classifiers are Random Forest (RF), k-Nearest Neighbor (kNN), Support Vector Machine (SVM), and XGBoost [ 15 ]. Noi, et al. [ 16 ] compared RF, kNN, and SVM classifiers for a study area and found that SVM outperformed the other models on overall accuracy. XGBoost, short for eXtreme Gradient Boosting , is a powerful, tree-boosting-based, open-source machine learning algorithm developed by Chen and Guestrin at the University of Washington [ 17 ]. It has been shown to outperform other ML models in classification competitions. Dobrinić, et al. [ 1 ] compared the performance of Random Forest and XGBoost classifiers in pixel-based classification using combined S-1 VV and VH bands (without texture features) and S-2 data, and found that the overall accuracy improved from 85.51–91.09% with the latter. By assigning weights to the data so that wrong predictions on data with larger weights result in larger loss functions, Chen et al. [ 18 ] found that weighted-XGBoost outperformed other models such as SVM, relevant vector machine, and deep belief network on accuracy of classifying radar emitter data. Recent studies have also explored alternative approaches to enhance classification outcomes. For high-resolution urban imagery, rule-based classification methods have been applied to extract urban land cover features with high accuracy, yet they require significant manual intervention and predefined classification rules, limiting scalability [ 19 ]. In parallel, integrating demographic data with EO imagery has shown promise for assessing community vulnerabilities to extreme weather events. Sharma, et al. [ 20 ] derived vulnerability indices for global cities using multi-resolution spaceborne imagery and socio-economic data. Complementing this approach, Cian, et al.[ 21 ] developed a methodology that fuses EO data with census information to map a multi-temporal flood vulnerability index in Northeast Italy. Moreover, Samadzadegan, et al. [ 22 ]critically reviewed data fusion techniques in multisource remote sensing, highlighting gaps in automation and scalability—a key limitation that the proposed FDLC model aims to overcome This brings us to the other major problem in processing large EO datasets: finding a platform that can efficiently store the data, offer tools to test and implement the classification technique, and scale the method to the large datasets. In order to address these challenges, cloud-based services such as Amazon Web Services (AWS) [ 23 ] and Google Earth Engine (GEE) [ 24 ] offer end-user-friendly solutions. GEE offers free access to petabytes of remote sensing data, with a temporal range going back decades and spatial range covering most of the world. Further, it also offers an integrated development environment and the computational power to process the large datasets on the platform [ 25 ]. Cloud computing environments like AWS contain EO data collections as open data, including NASA’s Landsat-8 and ESA’s Sentinel-2, in addition to being scalable and offering a range of products for data storage, machine learning training models, and computing resources [ 26 ]. Implementing land cover change studies in cloud-based environments can enable users with limited local storage and computing resources to perform spatio-temporal analyses on large datasets using powerful machine learning algorithms. For instance, Brovelli, et al. [ 27 ] studied deforestation in the Amazon by performing time-series land cover analysis in Google Earth Engine. They used multi-spectral images along with a Random Forest supervised classifier to estimate land cover change between the years 2000 and 2019. Soulard, et al. [ 28 ] demonstrated a method to map surface water in Cambodia using multi-decade Landsat satellite imagery in Google Earth Engine. They demonstrated that using remote sensing data can compensate for the lack of in-field sensors in hard-to-access areas while producing valuable information about the surface cover and its implications to policymakers. For a persistent, near-real time analysis of EO data with built-in machine learning training support, Amazon SageMaker platform is a great resource [ 29 ]. It offers a way to seamlessly access data from Earth on AWS, train models in SageMaker compute instances, and store the results in S3 buckets. This work presents a working methodology to implement land cover change analysis in AWS’ SageMaker framework. The main motivation to create a land-cover classification methodology in SageMaker is to demonstrate a technique that can be scaled up to predict land cover change using very large datasets in near-real-time. This is done by choosing two cities from different geographic and socio-economic backgrounds as study sites, and classifying their land cover with five different models, involving different combinations of training data and XGBoost parameters. The training datasets involve combinations of data from the study sites and from sites not included in this work. By comparing the overall accuracies (OA) of classification implemented without any local data with models including local data, we make a case to create a global training dataset similar to EuroSAT [ 30 ]. Finally, images for the two sites acquired a few years earlier are also classified, and the land cover change estimation is presented as a potential application for this workflow. Study Area We chose Austin, Texas and Greater Noida, India as the study sites. The rationale to choose the two cities is that this work focuses on identifying LULC change in urban areas, and the choice of the two sites offers samples from different countries with different socio-economic backgrounds and stages of economic development. The true-color images in Fig. 1 show Austin, Texas, from 2017 and 2020. Note the increase in built-up area especially in the south-west quadrant between the years. There is also an apparent reduction in the road area running north-south in the eastern side of the 2020 map. The higher-resolution Pleiades images in Fig. 2 show the same geographical extent. The image from 2020 also corresponds to the S-2 image captured in 2020. However, the image on the left is from much earlier (2013) than the latest available S-2 image, and therefore not used further in this study. Similarly, Fig. 3 shows Greater Noida, India, from 2017 and 2020. Notable differences are the haziness in 2020 due to smog. Furthermore, the 2017 image shows significant vegetation cover that had been converted to other classes in 2020. The development of a major road running roughly west to north-east is also noticeable. The Pleiades images in Fig. 4 correspond in geographical extent to the S-2 images, and the 2020 images were acquired within a month of the S-2 image. The 2013 image will not be used further in this work since the earliest available S-2 images are from four years later. Data and acquisition We used Sentinel-1 and Sentinel-2 data in the study for land cover classification, and higher-resolution Pleiades images to assess the classification accuracy. The S-1 data is acquired from Google Earth Engine since it is available as ready-to-analyze ground-range-detected (GRD) data. S-1 data is obtained by a dual-polarization, C-band (5.405 GHz) SAR instrument [ 31 ]. We obtain the VV and VH bands for data analysis; here, the letters represent the polarization of transmitted and received radiation. For instance, VV refers to vertical transmit, vertical receive, while VH refers to vertical transmit, horizontal receive. It is already preprocessed for orbit correction, border noise removal, thermal noise removal, radiometric calibration, and terrain correction in GEE [ 32 ]. Sentinel-2 data is obtained from AWS’ Sentinel-2 repository as orthorectified L2A product [ 33 ]. While other open-source repositories exist for S-2 data, it is convenient to access the data from within AWS using Sentinel Hub API [ 34 ]. The two satellites of Sentinel-2 mission, dubbed S2A and S2B, acquire multispectral optical sensor data at wavelengths in visual, near infrared, and shortwave infrared bands, as listed in Table 1 . The table also provides the bandwidth for each band and its spatial resolution. The raw S-2 data is then pre-processed to mosaic and clip it to the study area in SageMaker. Table 1 Sentinel-2 Band Specifications [ 9 ]. Band number S2A S2B Spatial Resolution (m) Central wavelength (nm) Bandwidth (nm) Central wavelength (nm) Bandwidth (nm) 2 492.4 66 492.1 66 10 3 559.8 36 559.0 36 10 4 664.6 31 664.9 31 10 8 832.8 106 832.9 106 10 5 704.1 15 703.8 16 20 6 740.5 15 739.1 15 20 7 782.8 20 779.7 20 20 8a 864.7 21 864.0 22 20 11 1613.7 91 1610.4 94 20 12 2202.4 175 2185.7 185 20 1 442.7 21 442.2 21 60 9 945.1 20 943.2 21 60 10 1373.5 31 1376.9 30 60 Pleiades Imagery Pleiades is a French spaceborne imagery that delivers very-high optical resolution (0.5m) with a swath of 20 kms. The five bands onboard the Pleideas high resolution imager (HRI) sense in 480–915 nm channels. A series of Pleideas imagery were acquired with funding from an ongoing NASA-MISTC project on global land cover changes. The Pleiades imagery was used in the study to derive accurate ground truth samples that was essential to estimate the accuracy of the classification from the moderate resolution MSS L2A Orthorectified S-2 datasets. The Pleiades images are saved in UTM projection. Customizing workflow in the SageMaker framework A workflow developed in AWS SageMaker was utilized to predict LULC from satellite images. While SageMaker provides a variety of built-in environments for data analysis, additional customizations were necessary to perform specific GIS analyses in the framework. A SageMaker notebook instance was created on an ‘ml.t3.xlarge’ instance, which provides 4 virtual CPUs and 16 GB of memory. The notebook volume size was set to 10 GB since large images are temporarily stored while processing. The customization of SageMaker notebook instance involved the following major steps: Python libraries such as ‘rioxarray’, ‘rasterio’, ‘xgboost’, and other relevant modules were installed in a custom conda environment in the notebook instance [ 35 ]; a ‘Lifecycle configuration’ script was used to make the conda environment persistently available as a kernel in the notebook instance. A Sentinel Hub (SH) [ 34 ] account is also necessary to access Sentinel-2 data using SH’s API, so an account was created with SH for this purpose. In addition to using SageMaker, the workflow also involved ingesting Sentinel-1 data from Google Earth Engine. This was done by downloading the data to Google Drive and copying it to AWS’ S3. The ‘boto3’ module was used to transfer data between the AWS S3 environment and SageMaker. Methodology After creating the working environment in the AWS-SM framework, S-2 L2A, S-1 Synthetic Aperture Radar (SAR) imagery, and texture bands derived from S-1 were used to classify land use and land cover. Ground truth dataset was created using stratified random sampling accuracy assessment model and Pleiades high-resolution MSS imagery. A brief description of all the steps involved in the project is shown in the flowchart in Fig. 5 , and described in the sections below. Step 1: Acquiring S-1, S-2 data and calculating texture metrics from S-1 The Sentinel Hub API offers an easy way to browse and download Sentinel-2 images from custom date ranges and geographic bounds. In addition, the images can also be filtered based on a maximum cloud-cover estimate. As a first step, the already available Pleiades high-resolution images that would be used for establishing ground-truth were mosaicked to form a composite image covering the study site. Then, the Sentinel images available in the AWS Sentinel-2 repository [ 33 ] were filtered in the notebook instance using the Sentinel Hub module. Images with very low cloud cover of less than 1% are used in this work since cloud cover reduces the accuracy of LULC predictions, and using cloudy images could complicate the accuracy assessment of the machine learning models. The images closest in time to when the Pleiades images were acquired are downloaded to SageMaker and transferred to an S3 user bucket. In case the study area is covered by more than one Sentinel-2 tiles, they are mosaicked to cover the entire geometry. The mosaicked images are then clipped to the geographic bounds of the Pleiades images. The images are acquired in UTM projection. Land use classification can be done by solely utilizing the Sentinel-2 band data [ 16 ]. However, earlier research and heuristic trials combining the texture features calculated from S-1 images along with S-2 data showed promise in improving classification accuracy. This is likely since the texture features add more dimensions to distinguish the land cover classes since the texture of water surfaces, for example, would be very different from that of vegetation. The texture features are calculated by creating “Gray-Level Covariance Matrix”, or GLCM, for sample windows much smaller than the image’s dimensions and measuring the relationship between pixels in that window [ 13 ]. The calculation window of 5 pixels by 5 pixels is translated across the image to measure texture features and are assigned to the center of the moving window. The texture metrics of contrast, energy, and correlation were measured for each of the VV and VH bands, thus resulting in six additional bands. We choose to obtain S-1 images from Google Earth Engine (GEE), since it offers analysis-ready S-1 data that has been pre-processed for terrain correction, orbit corrections, thermal noise removal, border noise removal, and radiometric calibration [ 32 ]. The S-1 images acquired within a week from the S-2 L2A images are downloaded for the same geographic bounds as the S-2 and Pleiades images using a Python program. Figure 6 shows the VV and VH backscatter bands, the contrast, energy, and correlation textures of each backscatter band from S-1 data and the 12 multi-spectral bands from S-2 data. Enhance the images in the figure showing the bands. Step 2(a): Training data We manually collected the training data corresponding to the study sites as point-geometry shapefiles in QGIS by utilizing photo-interpretation of S-2 bands and ESRI satellite images. The classification involved predicting four land-cover classes: water, built-up, low-vegetation, and vegetation. The classes were chosen by observing the different land cover classes in the true-color study site images. To aid in photointerpretation, setting the bands 7, 3, and 2 as red, green, and blue channels shows areas with vegetation in dark-red and low-vegetation areas in yellow or light red. Since water absorbs most of the radiation, the pixels covered by water are typically darker. The spectral signatures, SAR backscattering intensity and texture features of the four classes are shown for Austin in Fig. 7 and Fig. 8 . There is notable separation in classes visible in the S-2 band data, whereas the separation is only significant to tell water/non-water classes using S-1 data. The training data overlaid on the true-color S-2 image of Austin is shown in Fig. 9 . The goal of the training data collection was to create a number of samples representing each class across the map area. This was an iterative process where the training samples were increased as necessary to improve classification accuracy. The breakdown for the number of training samples for each study site is given in Table 2 . Table 2 Overview of training sample data. Number of samples for the class Study site Water Built-up Low-vegetation/soil Vegetation Total Austin 145 710 420 401 1701 Greater Noida 70 325 344 141 905 In addition to generating “local” training data for each site, we also utilized a database of training data generated from other sites around the world. We created training samples for a wide range of study sites around the globe, including the cities of Amos and Trenton in the USA, Ladysmith in South Africa, Sur in Oman, Silchar in India, Kitchener in Canada, and Bethania in Australia. A “global” dataset was created by merging the training data generated in this study and the other larger datasets. The impetus for testing this “global” dataset is to test the utility in predicting LULC accurately with an already available dataset without the need to freshly generate training data each time LULC needs to be estimated. The site locations for the global dataset are shown in Fig. 10 . Step 2(b): Training Machine Learning Classifier We used the Extreme Gradient Boosting model, or XGBoost, to train with the data and implement a classifier. It is a powerful, tree-boosting-based, open-source machine learning algorithm that is available both as a Python module for local computing and as a ready-to-use training framework in SageMaker [ 17 ] [ 36 ]. In essence, estimating land cover and land use from satellite imagery is a multi-class classification problem in machine learning. XGBoost has been shown to outperform other machine learning models, such as decision tree and random forest. Five different XGBoost models were trained with different subsets of the global and local data in SageMaker, and the accuracy was estimated for each model. The model frameworks were determined heuristically to distinguish performance with only local site-specific data, or with only site-exclusive global data, or with a combined dataset equally weighting local and site-exclusive data, or with combined dataset but with higher weightage assigned to local data. local_model (ingest study-specific training data) global_wt2 (ingest local_model training data + training data from other sites, with weighting of 2 assigned to study-specific data) global_wt5 (ingest local_model training data + training data from other sites, with weighting of 5 assigned to study-specific data) global_model (ingest local_model training data + training data from other sites, with equal weighting of all data) local_starved_model (ingest training data from other sites only) The first model ingested only the local site training data for the classifier, and we will refer to it as “local_model” . The second and third models utilized the global data but weighted the local data to different extents. The second model assigned twice the sample weights of global data to local data, and we call this model “global_wt2” . The third model assigned five times the sample weights of global data to local data, and we call it “global_wt5” . The weights were chosen heuristically by trial and error to see a difference in performance. The fourth model used the combined datasets of the global data, and we call this the “global_model” . The fifth model was trained on pre-existing larger dataset from the prior project alone, i.e., the classifier did not see any of the local data. This would be useful, for instance, to predict LULC for a new study site without any local training data in real-time. We call this “local_starved_model” . Hyper-parameters are the values assigned to the machine learning model variables, such as the sample weight and the number of iterations. By tuning hyper-parameters, the accuracy of a model can be greatly improved since the parameters optimize how well the model learns from the training data. We first tuned the hyperparameters for the global model. Before tuning, the training dataset was split such that 25% of samples from each class was assigned to a validation set, similar to the 70%-30% split used by Dobrinić et al. [ 1 ]. The resulting parameters were stored and then reused for training local XGBoost models in the notebook instance for the first three models. The rationale for reusing the hyperparameters is that the tuning process is time and resource intensive, and the parameter ranges in different trials with different data subsets were similar to the one obtained with the global model. A separate hyper-parameter tuning was performed for the global datasets excluding the local datasets, i.e., for the local_starved_model dataset. A separate tuning is necessary in this case since the other models all contain some instances of the local data whereas the local_starved_model has none, and this could change the model parameters. Step 3: Image classification and accuracy assessment Utilizing the five models discussed above, we classified the images with each model. In order to estimate the accuracy and compare the different approaches, a stratified random sampling design was implemented based on the recommendations of Olofsson et al. [ 37 ], since random sampling provides a better estimate of accuracy compared to opportunistic sampling. The methodology for this assessment involved estimating a total number of random points where ground-truth will be assessed by photo-interpretation of Pleiades images. This is the recommended approach since the Pleiades images are of much greater resolution than the S-2 images used in training the classifier. The number of points depends on the estimated accuracy of classification and the proportions of the different land cover classes, or strata, determined from one of the classified images for each site. A total of 431 and 426 points were sampled for Austin and Greater Noida, respectively. It was ensured that a minimum of 35 points were allocated to each stratum (i.e., class), with the allocation being proportional to the class area in the classified map otherwise. Thus, of the 431 points for Austin, the allocation was as follows: water − 49, built-up − 127, low-vegetation − 167, and vegetation − 88. Similarly, the 426 points for Greater Noida were distributed as follows: water − 39, built-up − 117, low-vegetation − 216, vegetation − 117. Step 4: Land cover change estimation As a demonstration of the application for the different approaches, S-1 and S-2 images corresponding to the same area but from 2017 were also classified using the same methodology described above. A land cover change raster was created in each case by comparing the pixels in the before and after images and assigning them pixel values according to the change matrix shown in Table 3 . The numbers in the table are a lookup table for pixel values assigned to the land cover change image, when comparing the change from before (BF) to after (AF). For instance, if the classification of a pixel in 2017 was built-up and changed to vegetation in 2020, it will be assigned a value of 8. Table 3 Matrix showing indices for land-cover change comparison. AF - Water AF - Built-up AF - Low-vegetation/soil AF - Vegetation BF - Water 1 2 3 4 BF - Built-up 5 6 7 8 BF - Low-vegetation/soil 9 10 11 12 BF - Vegetation 13 14 15 16 This methodology is shown as a flowchart in Fig. 11 . A comparison raster was created to highlight the increase in built-up class area (by highlighting pixels with values of 2, 10, and 14), and the conversion of vegetation to low-vegetation areas (by highlighting pixels with value 15). Results and Discussion Figure 12 shows the land cover classification estimates using the five machine learning models for Austin and Greater Noida. The model outputs for Austin are comparable and differences could not be discerned by visual inspection. The outputs for Greater Noida, on the other hand, show a large variation from model to model. For instance, local_model seems more land covered by low-vegetation/soil compared to all the other models. Similarly, the the local_starved_model seems to predict that most areas are ‘built-up’. However, as the weightage of local data increases in the models, the built-up area begins to be classified as low-vegetation. This suggests that the local data is adding critical information to the models to distinguish between built-up and low-vegetation classes, and without the local input, the model tends to overestimate the built-up area for Greater Noida. We present the accuracy metrics for each class and site estimated for the different models in Table 4 - Table 8 . The metrics are shown in boldface if the accuracy is better than 75%. The precision and recall metrics are also referred to as user’s accuracy and producer’s accuracy, respectively [ 38 ]. The precision and recall for water are consistently high (minima of 70.59% and 79.59% respectively) for both sites under all models, suggesting that any of these methods could be a good candidate for use in flood hazard response. With a minimum precision of 81.13% for detecting low-vegetation, the models are also good in accurately identifying low-vegetation areas. However, the recall is not consistently high for low-vegetation detection, indicating that the models underpredict low-vegetation areas. At the same time, most of the models show low precision and high-recall for built-up areas, suggesting that the built-up area is being overpredicted. The F-1 scores for the vegetation classification is below 70% in all the models, and future training could target this particular class to improve its accuracy. Table 4 Accuracy metrics for local_model. Accuracy metrics Location Austin Greater Noida Water Precision (%) 85.71 70.59 Recall (%) 85.71 92.31 F1-score (%) 85.71 80.00 Built-up Precision (%) 66.27 68.18 Recall (%) 88.19 76.92 F1-score (%) 75.68 72.29 Low-vegetation Precision (%) 92.86 86.77 Recall (%) 62.28 75.93 F1-score (%) 74.55 80.99 Vegetation Precision (%) 64.36 64.81 Recall (%) 73.86 64.81 F1-score (%) 68.78 64.81 Overall accuracy (%) 74.94 76.29 Table 5 Accuracy metrics for global_wt2. Accuracy metrics Location Austin Greater Noida Water Precision (%) 91.11 77.78 Recall (%) 83.67 89.74 F1-score (%) 87.23 83.33 Built-up Precision (%) 78.13 52.48 Recall (%) 78.74 90.60 F1-score (%) 78.43 66.46 Low-vegetation Precision (%) 82.82 88.03 Recall (%) 80.84 57.87 F1-score (%) 81.82 69.83 Vegetation Precision (%) 65.26 75.68 Recall (%) 70.45 51.85 F1-score (%) 67.76 61.54 Overall accuracy (%) 78.42 69.01 Table 6 Accuracy metrics for global_wt5. Accuracy metrics Location Austin Greater Noida Water Precision (%) 87.23 72.92 Recall (%) 83.67 89.74 F1-score (%) 85.42 80.46 Built-up Precision (%) 73.33 54.79 Recall (%) 77.95 88.03 F1-score (%) 75.57 67.54 Low-vegetation Precision (%) 84.00 89.26 Recall (%) 75.45 61.57 F1-score (%) 79.50 72.88 Vegetation Precision (%) 61.62 68.29 Recall (%) 69.32 51.85 F1-score (%) 65.24 58.95 Overall accuracy (%) 75.87 70.19 Table 7 Accuracy metrics for global_model. Accuracy metrics Location Austin Greater Noida Water Precision (%) 84.78 75.00 Recall (%) 79.59 92.31 F1-score (%) 82.11 82.76 Built-up Precision (%) 75.00 50.00 Recall (%) 77.95 88.89 F1-score (%) 76.45 64.00 Low-vegetation Precision (%) 81.13 87.41 Recall (%) 77.25 54.63 F1-score (%) 79.14 67.24 Vegetation Precision (%) 64.89 65.71 Recall (%) 69.32 42.59 F1-score (%) 67.03 51.69 Overall accuracy (%) 76.10 65.96 Table 8 Accuracy metrics for local_starved_model. Accuracy metrics Location Austin Greater Noida Water Precision (%) 91.11 72.92 Recall (%) 83.67 89.74 F1-score (%) 87.23 80.46 Built-up Precision (%) 73.13 47.95 Recall (%) 77.17 89.74 F1-score (%) 75.10 62.50 Low-vegetation Precision (%) 83.33 87.69 Recall (%) 77.84 52.78 F1-score (%) 80.50 65.90 Vegetation Precision (%) 63.54 75.86 Recall (%) 69.32 40.74 F1-score (%) 66.30 53.01 Overall accuracy (%) 76.57 64.79 The overall accuracy estimates of the five different machine learning classifier modeling approaches for the two study sites are presented in Fig. 13 . In Austin’s case, the different methodologies consistently produce accuracies around 75%, with the maximum accuracy attained with the combined data model where local data is weighted twice the global data (global_wt2). On the other hand, the accuracies for Greater Noida are very sensitive to the training data and model. While the accuracy is highest when using only local data (76.29%), it falls to about 70% when combined with global data but still weighting local data higher. However, using either the local-data-starved model or the global-local integrated model results in lower accuracy around 65%. The local_model is highly attuned to the band variations in the site’s geography. The advantage of this approach is that the ML classifier can be trained better to identify local peculiarities in how the classes are expressed in the band values. However, the drawback with this method is that a reliable training dataset matching the spatial and temporal extent of satellite imagery is necessary for this to work. This can require a significant time and monetary resources in many cases. Another drawback is that the model could be less applicable to other study sites differing in local morphology and topography. The local_model had middling performance for Austin whereas it outperformed the other models for Noida. This could be attributed to the fact that the combined datasets are overrepresented by cities from developed countries similar to Austin (for example, Trenton, Kitchener, and Amos) whereas cities from developing countries such as Greater Noida are underrepresented. Therefore, using only local data aids in improving predictions for the latter while starving the former of relevant data that could have improved accuracy. The global_wt2 and global_wt3 models improve the classification accuracy in Austin’s case while decreasing the accuracy for Greater Noida. This affirms the aforementioned theory about the differences in applicability of the global data to the individual sites. The global_wt2 is the best-performing model for Austin; however, the differences between the different model performance are small for Austin. By removing the higher weight to local data in global_model, the accuracy falls further for Greater Noida while remaining relatively unchanged for Austin. Finally, by completely removing any local data, Austin’s LULC classification is quite satisfactory when compared to the other methods while it performs the worst for Greater Noida. Examining the confusion matrix for Noida in Fig. 14 reveals that the model is overpredicting the built-up land cover and significantly underpredicting low-vegetation/soil class. Despite this inter-class confusion, 65% accuracy might be acceptable for a rough estimation of land cover when local training data is lacking. It is also possible that the accuracy with a global model might improve for developing cities like Greater Noida if more representative samples are added to the training set. Finally, to demonstrate the application of estimating LULC using cloud-based resources, we present land cover change maps and metrics from the highest accuracy models for each city. The “before” image was acquired for each city in 2017 whereas the “after” image was acquired in 2020. Figure 15 shows the classifications for Austin’s images. There is evidence of urban expansion in the previously less urban area in the southern half of the map. Loss of vegetation is focused in a few areas of the map and is not widespread. Figure 16 shows the increase in built-up area in the southern and western regions of the Greater Noida map. Significantly, the new construction of a roughly southwest-northeast running road is captured. There are identifiable losses in vegetation at various places. The metrics related to land cover change from before- and after- images (2017 and 2022) are presented in Table 9 . The table rows denote the “before” classification while the columns denote the “after” classification. For example, an area of 7.84 sq.km was classified as water both in the 2017 and 2020 images for Austin, while 11.04 sq. km was vegetation in 2017 and became low-vegetation in 2020. While the maps capture the local increase in built-up area in Austin, the metrics indicate that the trend need not be true across the map. However, owing to the known confusion between built-up and low-vegetation areas, it is likely that the classification for 2017 was overestimating the built-up area. The results for Greater Noida reflect a marked increase in built-up area and a significant loss (around 56%) of forested area to low-vegetation areas. Table 9 Land cover change metrics for Austin and Noida from global_wt2 and local_model, respectively. All values are in sq. km. Austin Aft:Water Aft:Built-up Aft:Low-vegetation Aft:Vegetation Bef:Water 7.84 1.16 1.05 1.24 Bef:Built-up 2.45 104.36 19.61 22.39 Bef:Low-vegetation 0.75 10.94 86.14 6.11 Bef:Vegetation 0.69 7.29 11.04 63.88 Noida Aft:Water Aft:Built-up Aft:Low-vegetation Aft:Vegetation Bef:Water 1.47 0.49 0.21 0.16 Bef:Built-up 0.46 29.98 10.28 2.84 Bef:Low-vegetation 0.23 11.29 48.91 4.27 Bef:Vegetation 0.18 2.19 12.25 7.39 Conclusion This work presented a methodology for implementing land cover classification in Amazon SageMaker using XGBoost algorithm. The cloud-based approach to image analysis is expected to benefit near-real-time data acquisition, analysis, and dissemination of results. SageMaker is particularly beneficial owing to the open availability of Sentinel data in AWS and powerful machine learning tools to analyze the large datasets. Five different approaches were tested with this framework to classify the land cover in Austin and Greater Noida. High-resolution Pleiades images were used to establish ground-truth data through photo-interpretation, which were used to estimate accuracy. The first approach utilized only training data located within the study sites. The second and third approaches utilized a global dataset combining local data and data from other sites trained previously, with the local data assigned larger weights. The fourth approach simply assigned equal weight to the combined global dataset. Finally, the fifth approach utilized only the data external to the sites, with the motivation that this could be useful in cases when local data is difficult to obtain either through photo-interpretation (e.g., cloud cover, corrupted data) or through field visits (e.g., harsh or inaccessible terrain). The results show good overall accuracy (above 70%) using all the approaches for Austin. However, the results for Greater Noida were better with the local data and deteriorated as the out-of-site data was mixed into the training sample. Despite the lowered performance, a minimum accuracy of about 65% was attained with the global data without utilizing any local data. This augments the case for building global representative datasets from a variety of sites that could be deployed on cloud-based land use monitoring systems, with only periodic updates to the local data added as necessary. Declarations Author Contribution Sunil Bhaskaran: Conceptualized the study idea and designed the overall methodology framework. Led the project direction and ensured alignment with urban informatics research themes. Supervised the study and ensured alignment with urban informatics research priorities.Arvindh Sharma: Conducted the technical analysis, including model training, and result interpretation using multi-resolution spaceborne imagery and cloud-based machine learning tools.Sanjiv Bhatia: Contributed to shaping the manuscript’s focus, particularly emphasizing connections to urban informatics applications and broader implications for urban resilience and planning.Stuti Mishra: Conducted in-depth review and quality assurance of the analysis workflows, curated and synthesized the literature review, and critically revised the manuscript for technical rigor, urban relevance, and clarity. Contributed to refining the research narrative to better position the study within the domain of urban informatics. Acknowledgments: The research was conducted with funding from the NASA-MISTC and United States NSF-ATE programs. Authors acknowledge the funding agencies and infrastructure provided by the BCC Geospatial Center of the CUNY CREST Institute, a leading center of excellence in New York. References D. Dobrinić, D. Medak and M. Gašparović, "INTEGRATION OF MULTITEMPORAL SENTINEL-1 AND SENTINEL-2 IMAGERY FOR LAND-COVER CLASSIFICATION USING MACHINE LEARNING METHODS," in Int. Arch. Photogramm. Remote Sens. Spatial Inf. Sci. , Nice, France, 2020. E. Lambin, B. Turner, H. Geist, B. Agbola, A. Angelsen, J. Bruce, C. Oliver, R. Dirzo, G. F. Fischer, P. George, K. Homewood, J. Imbernon, R. Leemans, X. Li, E. Moran, M. Mortimore, P. Ramakrishnan, J. Richards and J. Xu, "The causes of land-use and land-cover change: moving beyond the myths," Global Environmental Change , vol. 11, pp. 261–269, 2001. J.-F. Mas, R. Lemoine-Rodríguez, R. González-López, J. López-Sánchez, A. Piña-Garduño and E. Herrera-Flores, "Land use/land cover change detection combining," European Journal of Remote Sensing , vol. 50, no. 1, pp. 626–635, 2017. B. Pradhan, M. S. Tehrany and M. N. Jebur, "A New Semiautomated Detection Mapping of Flood Extent From TerraSAR-X Satellite Image Using Rule-Based Classification and Taguchi Optimization Techniques," IEEE Transactions on Geoscience and Remote Sensing , vol. 54, no. 7, 2016. M. Appel, F. Lahn, W. Buytaert and E. Pebesma, "Open and scalable analytics of large Earth observation datasets: From scenes to multidimensional arrays using SciDB and GDAL," ISPRS Journal of Photogrammetry and Remote Sensing , vol. 138, pp. 47–56, 2018. K. Navulur, Multispectral Image Analysis Using the Object-Oriented Paradigm, CRC Press, 2006. "Landsat Data Access," [Online]. Available: https://www.usgs.gov/landsat-missions/landsat-data-access . [Accessed 23 July 2022]. "Sentinel-2," The European Space Agency, [Online]. Available: https://sentinels.copernicus.eu/web/sentinel/missions/sentinel-2 . [Accessed 23 July 2022]. "MultiSpectral Instrument (MSI) Overview," European Space Agency, [Online]. Available: https://sentinels.copernicus.eu/web/sentinel/technical-guides/sentinel-2-msi/msi-instrument . [Accessed 04 10 2022]. D. Liu and F. Xia, "Assessing object-based classification: advantages and limitations," Remote Sensing Letters , vol. 1, no. 4, pp. 187–194, 2010. "Sentinel-1," The European Space Agency, [Online]. Available: https://sentinels.copernicus.eu/web/sentinel/missions/sentinel-1 . [Accessed 23 July 2022]. "Geophysical Measurements," European Space Agency, [Online]. Available: https://sentinels.copernicus.eu/web/sentinel/user-guides/sentinel-1-sar/product-overview/geophysical-measurements . [Accessed 04 10 2022]. M. Hall-Beyer, "GLCM Texture: A Tutorial v. 1.0 through 2.7.," 2007. [Online]. Available: https://prism.ucalgary.ca/handle/1880/51900 . [Accessed 22 July 2022]. M. Iyyappan and S. S. Ramakrishnan, "hancing land cover classification for multispectral images using hybrid polarimetry SAR data," International Journal of Remote Sensing , vol. 41, no. 17, pp. 6718–6754, 2020. P. T. Noi and M. Kappas, "Comparison of Random Forest, k-Nearest Neighbor, and Support Vector Machine Classifiers for Land Cover Classification Using Sentinel-2 Imagery," Sensors , vol. 18, no. 18, 2017. P. T. Noi and M. Kappas, "Comparison of Random Forest, k-Nearest Neighbor, and Support Vector Machine Classifiers for Land Cover Classification Using Sentinel-2 Imagery," Sensors , 2017. T. Chen and C. Guestrin, "XGBoost: A Scalable Tree Boosting System," in KDD '16: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining , San Francisco, California, USA, 2016. W. Chen, K. Fu, J. Zuo, X. Zheng, T. Huang and W. Ren, "Radar emitter classification for large data set based on weighted-xgboost," IET Radar, Sonar & Navigation , vol. 11, no. 8, pp. 1203–1207, 2017. S. Bhaskaran, E. Nez, K. Jimenez and S. K. Bhatia, "Rule-based classification of high-resolution imagery over urban areas in New York City," Geocarto International , 2012. A. R. Sharma and S. Bhaskaran, "Deriving community vulnerability indices by analyzing multi-resolution space-borne data and demographic data for extreme weather events in global cities,," Remote Sensing Applications: Society and Environment , vol. 33, 2024. F. Cian, C. Giupponi and M. Marconcini, "Integration of earth observation and census data for mapping a multi-temporal flood vulnerability index: a case study on Northeast Italy," Natural Hazards , p. 106, 2021. F. Samadzadegan, A. Toosi and F. Dadrass Javan, "A critical review on multi-sensor and multi-platform remote sensing data fusion approaches: current status and prospects," International Journal of Remote Sensing , p. 76, 2024. "Earth on AWS," Amazon Web Services, [Online]. Available: https://aws.amazon.com/earth/ . [Accessed 04 10 2022]. N. Gorelick, M. Hancher, M. Dixon, S. Ilyushchenko, D. Thau and R. Moore, "Google Earth Engine: Planetary-scale geospatial analysis for everyone.," Remote Sensing of Environment , 2017. H. Tamiminia, B. Salehi, M. Mahdianpari, L. Quackenbush, S. Adeli and B. Brisco, "Google Earth Engine for geo-big data applications: A meta-analysis and systematic review," ISPRS Journal of Photogrammetry and Remote Sensing , vol. 164, pp. 152–170, 2020. V. C. F. Gomes, G. R. Queiroz and K. R. Ferreira, "An Overview of Platforms for Big Earth Observation Data Management and Analysis," Remote Sensing , vol. 12, 2020. M. A. Brovelli, Y. Sun and V. Yordanov, "Monitoring Forest Change in the Amazon Using Multi-Temporal Remote Sensing Data and Machine Learning Classification on Google Earth Engine," International Journal of Geo-Information , vol. 9, no. 580, 2020. C. E. Soulard, J. J. Walker and R. E. Petrakis, "Implementation of a Surface Water Extent Model in Cambodia using Cloud-Based Remote Sensing," Remote Sensing , vol. 12, no. 984, 2020. E. Liberty and others, "Elastic Machine Learning Algorithms in Amazon SageMaker," in SIGMOD '20: Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data , Portland, OR, USA, 2020. P. Helber, B. Bischke, A. Dengel and D. Borth, "EuroSAT: A Novel Dataset and Deep Learning Benchmark for Land Use and Land Cover Classification," IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing , vol. 12, no. 7, 2019. "Sentinel-1 SAR GRD: C-band Synthetic Aperture Radar Ground Range Detected, log scaling," Google Earth Engine, [Online]. Available: https://developers.google.com/earth-engine/datasets/catalog/COPERNICUS_S1_GRD#description . [Accessed 04 10 2022]. "Sentinel-1 Algorithms," Google Earth Engine, [Online]. Available: https://developers.google.com/earth-engine/guides/sentinel1 . [Accessed 22 July 2022]. "Sentinel-2," Amazon, [Online]. Available: https://registry.opendata.aws/sentinel-2/ . [Accessed 30 09 2022]. "Sentinel Hub," Sinergise Ltd., [Online]. Available: https://www.sentinel-hub.com/ . [Accessed 30 09 2022]. Amazon AWS, [Online]. Available: https://aws.amazon.com/premiumsupport/knowledge-center/sagemaker-lifecycle-script-timeout/ . [Accessed 17 08 2022]. "XGBoost Algorithm," Amazon Web Services, [Online]. Available: https://docs.aws.amazon.com/sagemaker/latest/dg/xgboost.html . [Accessed 03 10 2022]. P. Olofsson, G. M. Foody, M. Herold, S. V. Stehman, C. E. Woodcock and M. A. Wulder, "Good practices for estimating area and assessing accuracy of land change," Remote Sensing of Environment , vol. 148, pp. 42–57, 2014. J. Weaver, B. Moore, A. Reith, J. McKee and D. Lunga, "A Comparison of Machine Learning Techniques to Extract Human Settlements from High Resolution Imagery," in IGARSS 2018–2018 IEEE International Geoscience and Remote Sensing Symposium , Valencia, Spain, 2018. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6497131","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":451687955,"identity":"30af8bb9-3c9f-4a07-869e-824efd8a12dd","order_by":0,"name":"Sunil Bhaskaran","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA+UlEQVRIiWNgGAWjYBAC+Rk8QPIAiMl8ACHMg0eLwQ24FrYEIrVIwLXwGBCpRbr34McfZ+zy+Wf3fH7xg8EucbtEAuODt224tcjPOZcszXMj2XLGnbPbLHsYkhN3zkhgNpyLRwvDjRwDaYYPzAYMN3K3GQOdmLjhRgKbNC9+LcY/f3yoN5C/kfMMpoX9NwEtZhI8Nw4bGNzIYX4Ms4UZnxaDO2fMrHnOHDcwvJFmxthjkGy84czDZsk55/B4f3aP8c0fx6oN5G4kP/7wo8JOdsPx5IMf3pThcRgSYJNgAEWNQGIDceqBgPkDmOI/QLSOUTAKRsEoGBkAAIPiWWNoka7+AAAAAElFTkSuQmCC","orcid":"","institution":"City University of New York","correspondingAuthor":true,"prefix":"","firstName":"Sunil","middleName":"","lastName":"Bhaskaran","suffix":""},{"id":451687956,"identity":"6290b168-2a59-4c96-9c27-aef3f24d7571","order_by":1,"name":"Arvindh Sharma","email":"","orcid":"","institution":"City University of New York","correspondingAuthor":false,"prefix":"","firstName":"Arvindh","middleName":"","lastName":"Sharma","suffix":""},{"id":451687957,"identity":"ff67f074-58ba-4d15-b06b-c4adfe8842d0","order_by":2,"name":"Sanjiv Bhatia","email":"","orcid":"","institution":"University of Missouri","correspondingAuthor":false,"prefix":"","firstName":"Sanjiv","middleName":"","lastName":"Bhatia","suffix":""},{"id":451687958,"identity":"c813b230-a8ab-47e5-b9c8-d1ca649b008b","order_by":3,"name":"Stuti Mishra","email":"","orcid":"","institution":"New York University","correspondingAuthor":false,"prefix":"","firstName":"Stuti","middleName":"","lastName":"Mishra","suffix":""}],"badges":[],"createdAt":"2025-04-21 14:53:18","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6497131/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6497131/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":82165632,"identity":"ef887b92-bae1-483c-af94-e24e2df0f089","added_by":"auto","created_at":"2025-05-07 09:08:51","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":5129478,"visible":true,"origin":"","legend":"\u003cp\u003eTrue-color S-2 images of Austin from 2017(left) and 2020 (right).\u003c/p\u003e","description":"","filename":"image1.png","url":"https://assets-eu.researchsquare.com/files/rs-6497131/v1/55a6f86e8174dfb9f9313489.png"},{"id":82165634,"identity":"672ceeb9-6cd6-437a-9cf4-cf4db0b4c7cb","added_by":"auto","created_at":"2025-05-07 09:08:51","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":3617740,"visible":true,"origin":"","legend":"\u003cp\u003eHigh-resolution Pleiades images covering the same geographical extent of Austin from 2012 (left) and 2020 (right). Only the 2020 image will be used for accuracy assessment.\u003c/p\u003e","description":"","filename":"image2.png","url":"https://assets-eu.researchsquare.com/files/rs-6497131/v1/463542bf3d37a798fbf3cb44.png"},{"id":82166853,"identity":"6a0cbfaf-b05f-4d72-9ae4-623f24f21b03","added_by":"auto","created_at":"2025-05-07 09:16:55","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":5404683,"visible":true,"origin":"","legend":"\u003cp\u003eTrue-color S-2 images of Greater Noida from 2017(left) and 2020 (right).\u003c/p\u003e","description":"","filename":"image3.png","url":"https://assets-eu.researchsquare.com/files/rs-6497131/v1/64a4d83db1376314975fb105.png"},{"id":82165649,"identity":"4b49fcf1-e789-4e32-9a4b-ce0655a500ac","added_by":"auto","created_at":"2025-05-07 09:08:52","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":5006149,"visible":true,"origin":"","legend":"\u003cp\u003eHigh-resolution Pleiades images covering the same geographical extent of Greater Noida from 2013 (left) and 2020 (right). Only the 2020 image will be used for accuracy assessment.\u003c/p\u003e","description":"","filename":"image4.png","url":"https://assets-eu.researchsquare.com/files/rs-6497131/v1/e98b3503aaa47a6efafe4013.png"},{"id":82165630,"identity":"54900ce9-4d63-414a-a599-1a63ef5b3c1f","added_by":"auto","created_at":"2025-05-07 09:08:51","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":352679,"visible":true,"origin":"","legend":"\u003cp\u003eFlowchart of the methodology involved in classifying LULC using Sentinel images.\u003c/p\u003e","description":"","filename":"image5.png","url":"https://assets-eu.researchsquare.com/files/rs-6497131/v1/4363b478c32222116e3c4bb7.png"},{"id":82166842,"identity":"93abdcf2-7724-47ec-9a9d-d2795e67fd07","added_by":"auto","created_at":"2025-05-07 09:16:51","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":4601696,"visible":true,"origin":"","legend":"\u003cp\u003eS-1 backscatter intensity bands, associated texture bands, and S-2 bands for Austin.\u003c/p\u003e","description":"","filename":"image6.png","url":"https://assets-eu.researchsquare.com/files/rs-6497131/v1/28ba7a31dcbf7e1e032b7a15.png"},{"id":82165671,"identity":"7b3abcd4-fda6-4672-948f-804d3200866e","added_by":"auto","created_at":"2025-05-07 09:08:54","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":1886072,"visible":true,"origin":"","legend":"\u003cp\u003eSpectral signatures for the four land cover classes from S-2 images.\u003c/p\u003e","description":"","filename":"image7.png","url":"https://assets-eu.researchsquare.com/files/rs-6497131/v1/9966847e290d920afbf3afdf.png"},{"id":82165621,"identity":"2af9f417-a0cb-4080-8c87-a909a5e3f4fd","added_by":"auto","created_at":"2025-05-07 09:08:51","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":1912771,"visible":true,"origin":"","legend":"\u003cp\u003eSAR backscattering intensity and texture features for Austin are shown after normalization.\u003c/p\u003e","description":"","filename":"image8.png","url":"https://assets-eu.researchsquare.com/files/rs-6497131/v1/b56d4658eab47d4a26f970ac.png"},{"id":82166846,"identity":"d13dc18e-aeba-4df5-9035-30ad4ab5ea80","added_by":"auto","created_at":"2025-05-07 09:16:53","extension":"png","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":7463534,"visible":true,"origin":"","legend":"\u003cp\u003eA true-color image of Austin area showing the different classes' training data.\u003c/p\u003e","description":"","filename":"image9.png","url":"https://assets-eu.researchsquare.com/files/rs-6497131/v1/3504d6b6cd2fda26b43f2f5c.png"},{"id":82166839,"identity":"bfdff332-59d1-4c0c-a2fe-5f2f5818da0d","added_by":"auto","created_at":"2025-05-07 09:16:51","extension":"png","order_by":10,"title":"Figure 10","display":"","copyAsset":false,"role":"figure","size":139518,"visible":true,"origin":"","legend":"\u003cp\u003eSite locations corresponding to training data (created by photo-interpretating of S2 images) in the global dataset.\u003c/p\u003e","description":"","filename":"image10.png","url":"https://assets-eu.researchsquare.com/files/rs-6497131/v1/937d846dec420373cf7e00ad.png"},{"id":82165655,"identity":"ec2c6db7-d265-4422-abbf-2163ef30a37a","added_by":"auto","created_at":"2025-05-07 09:08:53","extension":"png","order_by":11,"title":"Figure 11","display":"","copyAsset":false,"role":"figure","size":388713,"visible":true,"origin":"","legend":"\u003cp\u003eFlowchart describing the process to compare LCC of two images.\u003c/p\u003e","description":"","filename":"image11.png","url":"https://assets-eu.researchsquare.com/files/rs-6497131/v1/b6ff357b1deaa72169e28512.png"},{"id":82165645,"identity":"478a874a-3e98-4f26-bdbb-3ce4a2691f7f","added_by":"auto","created_at":"2025-05-07 09:08:52","extension":"png","order_by":12,"title":"Figure 12","display":"","copyAsset":false,"role":"figure","size":7648118,"visible":true,"origin":"","legend":"\u003cp\u003eLand cover classifications for Austin and Greater Noida using the five different classifier models.\u003c/p\u003e","description":"","filename":"image12.png","url":"https://assets-eu.researchsquare.com/files/rs-6497131/v1/57ba2487f0730115831bd536.png"},{"id":82165683,"identity":"baa77b85-037b-4b34-a24a-bc54261a5462","added_by":"auto","created_at":"2025-05-07 09:08:56","extension":"png","order_by":13,"title":"Figure 13","display":"","copyAsset":false,"role":"figure","size":275615,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eComparison of overall accuracies for the five different models. Blue: Austin, orange: Greater Noida.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"image13.png","url":"https://assets-eu.researchsquare.com/files/rs-6497131/v1/50f06461f69cb58d48c3014b.png"},{"id":82166854,"identity":"68b56c6c-2a0a-4a29-95fb-c16797cf644a","added_by":"auto","created_at":"2025-05-07 09:16:56","extension":"png","order_by":14,"title":"Figure 14","display":"","copyAsset":false,"role":"figure","size":244447,"visible":true,"origin":"","legend":"\u003cp\u003eConfusion matrix for local_starved_model results of Greater Noida.\u003c/p\u003e","description":"","filename":"image14.png","url":"https://assets-eu.researchsquare.com/files/rs-6497131/v1/1d7d925ecdc6cc3d0190747f.png"},{"id":82165682,"identity":"9f58542c-2439-4ade-ab65-c1bdba6699a5","added_by":"auto","created_at":"2025-05-07 09:08:56","extension":"png","order_by":15,"title":"Figure 15","display":"","copyAsset":false,"role":"figure","size":6314615,"visible":true,"origin":"","legend":"\u003cp\u003eClassification outputs and LCC change for Austin. Years: 2017 and 2020.\u003c/p\u003e","description":"","filename":"image15.png","url":"https://assets-eu.researchsquare.com/files/rs-6497131/v1/0de2d1a6c68c9ed13142ff23.png"},{"id":82165687,"identity":"cdd6c682-344f-4661-b14c-3ec4ca4f1ba9","added_by":"auto","created_at":"2025-05-07 09:08:56","extension":"png","order_by":16,"title":"Figure 16","display":"","copyAsset":false,"role":"figure","size":4538614,"visible":true,"origin":"","legend":"\u003cp\u003eClassification outputs and LCC change for Greater Noida. Years: 2017 and 2020.\u003c/p\u003e","description":"","filename":"image16.png","url":"https://assets-eu.researchsquare.com/files/rs-6497131/v1/c29f08ea289e2ea453ae5d57.png"},{"id":82165653,"identity":"36363d53-4e4b-4576-bdfd-e8541cbd2118","added_by":"auto","created_at":"2025-05-07 09:08:53","extension":"png","order_by":17,"title":"Figure 17","display":"","copyAsset":false,"role":"figure","size":11317,"visible":true,"origin":"","legend":"\u003cp\u003eUnnumbered image in the Methodology section.\u003c/p\u003e","description":"","filename":"Unnumber.png","url":"https://assets-eu.researchsquare.com/files/rs-6497131/v1/b8e5a89ab6bceeee11f54aae.png"},{"id":82911193,"identity":"300cdc78-edc3-4c4c-a8bd-8338850bfb28","added_by":"auto","created_at":"2025-05-16 15:16:48","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":52363601,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6497131/v1/8124f61d-4277-418f-90f2-b4612bba3772.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Automating image classification by using multi-resolution spaceborne imagery, cloud-based machine learning algorithms and open-source software and data ","fulltext":[{"header":"Introduction","content":"\u003cp\u003eOne of the major applications of remote sensing data is deducing the land use and land cover (LULC) of earth\u0026rsquo;s land surface [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. The land cover identifies the surface\u0026rsquo;s biophysical attributes while the land use identifies the human-intended purpose for the land cover [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. Tracking changes in the land cover is crucial in estimating the impacts of anthropogenic stressors, the extent of hazards such as floods and forest fires, and monitoring environmental outcomes such as climate change [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e] [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. Earth observing (EO) satellite data, combined with improved techniques to estimate land cover classes, have paved the way for tracking changes in spatio-temporal datasets. Many of these large EO datasets are publicly available and can aid researchers and policy-makers to utilize land cover change information in near-real-time [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. The major challenges in converting the raw EO data into useful land cover/land use information include the identification and development of techniques to process the data and predict classification of the surface, and the ability to scale the solution to ingest, process, and store the large datasets.\u003c/p\u003e \u003cp\u003eOne of the key techniques used in land cover classification is using multi-spectral satellite imagery with pixel-based image analysis [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. This is made easier by the public availability of high-quality multispectral data from programs such as NASA\u0026rsquo;s Landsat [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e] and ESA\u0026rsquo;s Sentinel-2 [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e], which have good spatial coverage and temporal frequency of revisit over most of earth\u0026rsquo;s surface. In this method, each pixel contains information collected in multiple spectral bands. For example, Sentinel-2 (S-2) senses radiations from the earth\u0026rsquo;s surface in thirteen different bands in visible, near infrared and shortwave infrared wavelengths [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Different classes, for example, water and vegetation, would have different spectral signatures registered on the sensors. By identifying the spectral signatures of different land classes, each pixel in the image can be classified to one of the identified classes. Beyond pixel-based classification, object-based methods identify spectrally homogenous pixels and segment them into objects [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. While object-based method removes salt-and-pepper noise that is pervasive in pixel-based classification, it is still prone to segmentation errors and is less simple compared to the latter. On the other hand, pixel-based multispectral classification can be enhanced by incorporating texture features derived from Synthetic Aperture Radar (SAR) data, such as those produced by the Sentinel-1 (S-1) program [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. SAR actively transmits microwave signals towards the earth\u0026rsquo;s surface and constructs an image based on the portion of the signal scattered back to the sensor [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. The advantage is that the microwaves are less affected by cloud cover and can be collected at any time of the day [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. Texture refers to the inter-relationship between neighboring pixels in an image window and can be used to distinguish between classes since different classes typically have different textures [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. Texture is quantified by constructing a Gray Level Covariance Matrix (GLCM), and techniques incorporating it in pixel-based methods have shown promise. Iyyappan, et al. [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e] observed that integrating texture metrics into multi-spectral optical bands can improve classification accuracy by 15%. They attribute this improvement to the enhanced ability to distinguish between all classes, particularly plantation and rice crop areas, which is not well discriminated when using only optical spectral data.\u003c/p\u003e \u003cp\u003eThe choice of machine learning (ML) classifier used in the pixel-based technique can also have a significant impact on accuracy. Some of the often-used land-cover classifiers are Random Forest (RF), k-Nearest Neighbor (kNN), Support Vector Machine (SVM), and XGBoost [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. Noi, et al. [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e] compared RF, kNN, and SVM classifiers for a study area and found that SVM outperformed the other models on overall accuracy. XGBoost, short for \u003cem\u003eeXtreme Gradient Boosting\u003c/em\u003e, is a powerful, tree-boosting-based, open-source machine learning algorithm developed by Chen and Guestrin at the University of Washington [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. It has been shown to outperform other ML models in classification competitions. Dobrinić, et al. [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e] compared the performance of Random Forest and XGBoost classifiers in pixel-based classification using combined S-1 VV and VH bands (without texture features) and S-2 data, and found that the overall accuracy improved from 85.51\u0026ndash;91.09% with the latter. By assigning weights to the data so that wrong predictions on data with larger weights result in larger loss functions, Chen et al. [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e] found that weighted-XGBoost outperformed other models such as SVM, relevant vector machine, and deep belief network on accuracy of classifying radar emitter data.\u003c/p\u003e \u003cp\u003eRecent studies have also explored alternative approaches to enhance classification outcomes. For high-resolution urban imagery, rule-based classification methods have been applied to extract urban land cover features with high accuracy, yet they require significant manual intervention and predefined classification rules, limiting scalability [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. In parallel, integrating demographic data with EO imagery has shown promise for assessing community vulnerabilities to extreme weather events. Sharma, et al. [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e] derived vulnerability indices for global cities using multi-resolution spaceborne imagery and socio-economic data. Complementing this approach, Cian, et al.[\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e] developed a methodology that fuses EO data with census information to map a multi-temporal flood vulnerability index in Northeast Italy. Moreover, Samadzadegan, et al. [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]critically reviewed data fusion techniques in multisource remote sensing, highlighting gaps in automation and scalability\u0026mdash;a key limitation that the proposed FDLC model aims to overcome\u003c/p\u003e \u003cp\u003eThis brings us to the other major problem in processing large EO datasets: finding a platform that can efficiently store the data, offer tools to test and implement the classification technique, and scale the method to the large datasets. In order to address these challenges, cloud-based services such as Amazon Web Services (AWS) [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e] and Google Earth Engine (GEE) [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e] offer end-user-friendly solutions. GEE offers free access to petabytes of remote sensing data, with a temporal range going back decades and spatial range covering most of the world. Further, it also offers an integrated development environment and the computational power to process the large datasets on the platform [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. Cloud computing environments like AWS contain EO data collections as open data, including NASA\u0026rsquo;s Landsat-8 and ESA\u0026rsquo;s Sentinel-2, in addition to being scalable and offering a range of products for data storage, machine learning training models, and computing resources [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]. Implementing land cover change studies in cloud-based environments can enable users with limited local storage and computing resources to perform spatio-temporal analyses on large datasets using powerful machine learning algorithms. For instance, Brovelli, et al. [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e] studied deforestation in the Amazon by performing time-series land cover analysis in Google Earth Engine. They used multi-spectral images along with a Random Forest supervised classifier to estimate land cover change between the years 2000 and 2019. Soulard, et al. [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e] demonstrated a method to map surface water in Cambodia using multi-decade Landsat satellite imagery in Google Earth Engine. They demonstrated that using remote sensing data can compensate for the lack of in-field sensors in hard-to-access areas while producing valuable information about the surface cover and its implications to policymakers. For a persistent, near-real time analysis of EO data with built-in machine learning training support, Amazon SageMaker platform is a great resource [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]. It offers a way to seamlessly access data from Earth on AWS, train models in SageMaker compute instances, and store the results in S3 buckets.\u003c/p\u003e \u003cp\u003eThis work presents a working methodology to implement land cover change analysis in AWS\u0026rsquo; SageMaker framework. The main motivation to create a land-cover classification methodology in SageMaker is to demonstrate a technique that can be scaled up to predict land cover change using very large datasets in near-real-time. This is done by choosing two cities from different geographic and socio-economic backgrounds as study sites, and classifying their land cover with five different models, involving different combinations of training data and XGBoost parameters. The training datasets involve combinations of data from the study sites and from sites not included in this work. By comparing the overall accuracies (OA) of classification implemented without any local data with models including local data, we make a case to create a global training dataset similar to EuroSAT [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e]. Finally, images for the two sites acquired a few years earlier are also classified, and the land cover change estimation is presented as a potential application for this workflow.\u003c/p\u003e"},{"header":"Study Area","content":"\u003cp\u003eWe chose Austin, Texas and Greater Noida, India as the study sites. The rationale to choose the two cities is that this work focuses on identifying LULC change in urban areas, and the choice of the two sites offers samples from different countries with different socio-economic backgrounds and stages of economic development.\u003c/p\u003e \u003cp\u003eThe true-color images in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e show Austin, Texas, from 2017 and 2020. Note the increase in built-up area especially in the south-west quadrant between the years. There is also an apparent reduction in the road area running north-south in the eastern side of the 2020 map. The higher-resolution Pleiades images in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e show the same geographical extent. The image from 2020 also corresponds to the S-2 image captured in 2020. However, the image on the left is from much earlier (2013) than the latest available S-2 image, and therefore not used further in this study.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eSimilarly, Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e shows Greater Noida, India, from 2017 and 2020. Notable differences are the haziness in 2020 due to smog. Furthermore, the 2017 image shows significant vegetation cover that had been converted to other classes in 2020. The development of a major road running roughly west to north-east is also noticeable. The Pleiades images in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e correspond in geographical extent to the S-2 images, and the 2020 images were acquired within a month of the S-2 image. The 2013 image will not be used further in this work since the earliest available S-2 images are from four years later.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eData and acquisition\u003c/strong\u003e \u003cp\u003eWe used Sentinel-1 and Sentinel-2 data in the study for land cover classification, and higher-resolution Pleiades images to assess the classification accuracy.\u003c/p\u003e \u003c/p\u003e \u003cp\u003eThe S-1 data is acquired from Google Earth Engine since it is available as ready-to-analyze ground-range-detected (GRD) data. S-1 data is obtained by a dual-polarization, C-band (5.405 GHz) SAR instrument [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. We obtain the VV and VH bands for data analysis; here, the letters represent the polarization of transmitted and received radiation. For instance, VV refers to vertical transmit, vertical receive, while VH refers to vertical transmit, horizontal receive. It is already preprocessed for orbit correction, border noise removal, thermal noise removal, radiometric calibration, and terrain correction in GEE [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eSentinel-2 data is obtained from AWS\u0026rsquo; Sentinel-2 repository as orthorectified L2A product [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. While other open-source repositories exist for S-2 data, it is convenient to access the data from within AWS using Sentinel Hub API [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]. The two satellites of Sentinel-2 mission, dubbed S2A and S2B, acquire multispectral optical sensor data at wavelengths in visual, near infrared, and shortwave infrared bands, as listed in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. The table also provides the bandwidth for each band and its spatial resolution. The raw S-2 data is then pre-processed to mosaic and clip it to the study area in SageMaker.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eSentinel-2 Band Specifications [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eBand number\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eS2A\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e \u003cp\u003eS2B\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eSpatial\u003c/p\u003e \u003cp\u003eResolution\u003c/p\u003e \u003cp\u003e(m)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCentral wavelength (nm)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eBandwidth (nm)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eCentral wavelength (nm)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eBandwidth (nm)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e492.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e492.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e559.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e36\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e559.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e36\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e664.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e31\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e664.9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e31\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e832.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e106\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e832.9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e106\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e704.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e15\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e703.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e16\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e740.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e15\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e739.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e15\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e782.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e779.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e8a\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e864.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e21\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e864.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e22\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1613.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e91\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1610.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e94\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2202.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e175\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e2185.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e185\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e442.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e21\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e442.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e21\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e60\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e945.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e943.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e21\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e60\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1373.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e31\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1376.9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e30\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e60\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003ePleiades Imagery\u003c/strong\u003e \u003cp\u003ePleiades is a French spaceborne imagery that delivers very-high optical resolution (0.5m) with a swath of 20 kms. The five bands onboard the Pleideas high resolution imager (HRI) sense in 480\u0026ndash;915 nm channels. A series of Pleideas imagery were acquired with funding from an ongoing NASA-MISTC project on global land cover changes. The Pleiades imagery was used in the study to derive accurate ground truth samples that was essential to estimate the accuracy of the classification from the moderate resolution MSS L2A Orthorectified S-2 datasets. The Pleiades images are saved in UTM projection.\u003c/p\u003e \u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eCustomizing workflow in the SageMaker framework\u003c/h2\u003e \u003cp\u003eA workflow developed in AWS SageMaker was utilized to predict LULC from satellite images. While SageMaker provides a variety of built-in environments for data analysis, additional customizations were necessary to perform specific GIS analyses in the framework. A SageMaker notebook instance was created on an \u0026lsquo;ml.t3.xlarge\u0026rsquo; instance, which provides 4 virtual CPUs and 16 GB of memory. The notebook volume size was set to 10 GB since large images are temporarily stored while processing.\u003c/p\u003e \u003cp\u003eThe customization of SageMaker notebook instance involved the following major steps: Python libraries such as \u0026lsquo;rioxarray\u0026rsquo;, \u0026lsquo;rasterio\u0026rsquo;, \u0026lsquo;xgboost\u0026rsquo;, and other relevant modules were installed in a custom conda environment in the notebook instance [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e]; a \u0026lsquo;Lifecycle configuration\u0026rsquo; script was used to make the conda environment persistently available as a kernel in the notebook instance. A Sentinel Hub (SH) [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e] account is also necessary to access Sentinel-2 data using SH\u0026rsquo;s API, so an account was created with SH for this purpose. In addition to using SageMaker, the workflow also involved ingesting Sentinel-1 data from Google Earth Engine. This was done by downloading the data to Google Drive and copying it to AWS\u0026rsquo; S3. The \u0026lsquo;boto3\u0026rsquo; module was used to transfer data between the AWS S3 environment and SageMaker.\u003c/p\u003e \u003c/div\u003e"},{"header":"Methodology","content":"\u003cp\u003eAfter creating the working environment in the AWS-SM framework, S-2 L2A, S-1 Synthetic Aperture Radar (SAR) imagery, and texture bands derived from S-1 were used to classify land use and land cover. Ground truth dataset was created using stratified random sampling accuracy assessment model and Pleiades high-resolution MSS imagery. A brief description of all the steps involved in the project is shown in the flowchart in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e, and described in the sections below.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\n\u003ch3\u003eStep 1: Acquiring S-1, S-2 data and calculating texture metrics from S-1\u003c/h3\u003e\n\u003cp\u003eThe Sentinel Hub API offers an easy way to browse and download Sentinel-2 images from custom date ranges and geographic bounds. In addition, the images can also be filtered based on a maximum cloud-cover estimate. As a first step, the already available Pleiades high-resolution images that would be used for establishing ground-truth were mosaicked to form a composite image covering the study site. Then, the Sentinel images available in the AWS Sentinel-2 repository [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e] were filtered in the notebook instance using the Sentinel Hub module. Images with very low cloud cover of less than 1% are used in this work since cloud cover reduces the accuracy of LULC predictions, and using cloudy images could complicate the accuracy assessment of the machine learning models. The images closest in time to when the Pleiades images were acquired are downloaded to SageMaker and transferred to an S3 user bucket. In case the study area is covered by more than one Sentinel-2 tiles, they are mosaicked to cover the entire geometry. The mosaicked images are then clipped to the geographic bounds of the Pleiades images. The images are acquired in UTM projection.\u003c/p\u003e \u003cp\u003eLand use classification can be done by solely utilizing the Sentinel-2 band data [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. However, earlier research and heuristic trials combining the texture features calculated from S-1 images along with S-2 data showed promise in improving classification accuracy. This is likely since the texture features add more dimensions to distinguish the land cover classes since the texture of water surfaces, for example, would be very different from that of vegetation. The texture features are calculated by creating \u0026ldquo;Gray-Level Covariance Matrix\u0026rdquo;, or GLCM, for sample windows much smaller than the image\u0026rsquo;s dimensions and measuring the relationship between pixels in that window [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. The calculation window of 5 pixels by 5 pixels is translated across the image to measure texture features and are assigned to the center of the moving window. The texture metrics of contrast, energy, and correlation were measured for each of the VV and VH bands, thus resulting in six additional bands. We choose to obtain S-1 images from Google Earth Engine (GEE), since it offers analysis-ready S-1 data that has been pre-processed for terrain correction, orbit corrections, thermal noise removal, border noise removal, and radiometric calibration [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e]. The S-1 images acquired within a week from the S-2 L2A images are downloaded for the same geographic bounds as the S-2 and Pleiades images using a Python program. Figure\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e shows the VV and VH backscatter bands, the contrast, energy, and correlation textures of each backscatter band from S-1 data and the 12 multi-spectral bands from S-2 data. Enhance the images in the figure showing the bands.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\n\u003ch3\u003eStep 2(a): Training data\u003c/h3\u003e\n\u003cp\u003e \u003c/p\u003e \u003cp\u003eWe manually collected the training data corresponding to the study sites as point-geometry shapefiles in QGIS by utilizing photo-interpretation of S-2 bands and ESRI satellite images. The classification involved predicting four land-cover classes: water, built-up, low-vegetation, and vegetation. The classes were chosen by observing the different land cover classes in the true-color study site images. To aid in photointerpretation, setting the bands 7, 3, and 2 as red, green, and blue channels shows areas with vegetation in dark-red and low-vegetation areas in yellow or light red. Since water absorbs most of the radiation, the pixels covered by water are typically darker. The spectral signatures, SAR backscattering intensity and texture features of the four classes are shown for Austin in Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003e and Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e. There is notable separation in classes visible in the S-2 band data, whereas the separation is only significant to tell water/non-water classes using S-1 data. The training data overlaid on the true-color S-2 image of Austin is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig9\" class=\"InternalRef\"\u003e9\u003c/span\u003e. The goal of the training data collection was to create a number of samples representing each class across the map area. This was an iterative process where the training samples were increased as necessary to improve classification accuracy. The breakdown for the number of training samples for each study site is given in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eOverview of training sample data.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colspan=\"5\" nameend=\"c6\" namest=\"c2\"\u003e \u003cp\u003eNumber of samples for the class\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eStudy site\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eWater\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eBuilt-up\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eLow-vegetation/soil\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eVegetation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eTotal\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAustin\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e145\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e710\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e420\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e401\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1701\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGreater Noida\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e70\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e325\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e344\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e141\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e905\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eIn addition to generating \u0026ldquo;local\u0026rdquo; training data for each site, we also utilized a database of training data generated from other sites around the world. We created training samples for a wide range of study sites around the globe, including the cities of Amos and Trenton in the USA, Ladysmith in South Africa, Sur in Oman, Silchar in India, Kitchener in Canada, and Bethania in Australia. A \u0026ldquo;global\u0026rdquo; dataset was created by merging the training data generated in this study and the other larger datasets. The impetus for testing this \u0026ldquo;global\u0026rdquo; dataset is to test the utility in predicting LULC accurately with an already available dataset without the need to freshly generate training data each time LULC needs to be estimated. The site locations for the global dataset are shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig10\" class=\"InternalRef\"\u003e10\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\n\u003ch3\u003eStep 2(b): Training Machine Learning Classifier\u003c/h3\u003e\n\u003cp\u003eWe used the Extreme Gradient Boosting model, or XGBoost, to train with the data and implement a classifier. It is a powerful, tree-boosting-based, open-source machine learning algorithm that is available both as a Python module for local computing and as a ready-to-use training framework in SageMaker [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e] [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. In essence, estimating land cover and land use from satellite imagery is a multi-class classification problem in machine learning. XGBoost has been shown to outperform other machine learning models, such as decision tree and random forest.\u003c/p\u003e \u003cp\u003eFive different XGBoost models were trained with different subsets of the global and local data in SageMaker, and the accuracy was estimated for each model. The model frameworks were determined heuristically to distinguish performance with only local site-specific data, or with only site-exclusive global data, or with a combined dataset equally weighting local and site-exclusive data, or with combined dataset but with higher weightage assigned to local data.\u003c/p\u003e \u003cp\u003elocal_model (ingest study-specific training data)\u003c/p\u003e \u003cp\u003e \u003cem\u003eglobal_wt2 (ingest local_model training data\u0026thinsp;+\u0026thinsp;training data from other sites, with weighting of 2 assigned to study-specific data)\u003c/em\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eglobal_wt5 (ingest local_model training data\u0026thinsp;+\u0026thinsp;training data from other sites, with weighting of 5 assigned to study-specific data)\u003c/em\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eglobal_model (ingest local_model training data\u0026thinsp;+\u0026thinsp;training data from other sites, with equal weighting of all data)\u003c/em\u003e \u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003elocal_starved_model (ingest training data from other sites only)\u003c/h2\u003e \u003cp\u003eThe first model ingested only the local site training data for the classifier, and we will refer to it as \u003cem\u003e\u0026ldquo;local_model\u0026rdquo;\u003c/em\u003e. The second and third models utilized the global data but weighted the local data to different extents. The second model assigned twice the sample weights of global data to local data, and we call this model \u003cem\u003e\u0026ldquo;global_wt2\u0026rdquo;\u003c/em\u003e. The third model assigned five times the sample weights of global data to local data, and we call it \u003cem\u003e\u0026ldquo;global_wt5\u0026rdquo;\u003c/em\u003e. The weights were chosen heuristically by trial and error to see a difference in performance. The fourth model used the combined datasets of the global data, and we call this the \u003cem\u003e\u0026ldquo;global_model\u0026rdquo;\u003c/em\u003e. The fifth model was trained on pre-existing larger dataset from the prior project alone, i.e., the classifier did not see any of the local data. This would be useful, for instance, to predict LULC for a new study site without any local training data in real-time. We call this \u003cem\u003e\u0026ldquo;local_starved_model\u0026rdquo;\u003c/em\u003e.\u003c/p\u003e \u003cp\u003eHyper-parameters are the values assigned to the machine learning model variables, such as the sample weight and the number of iterations. By tuning hyper-parameters, the accuracy of a model can be greatly improved since the parameters optimize how well the model learns from the training data. We first tuned the hyperparameters for the global model. Before tuning, the training dataset was split such that 25% of samples from each class was assigned to a validation set, similar to the 70%-30% split used by Dobrinić et al. [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. The resulting parameters were stored and then reused for training local XGBoost models in the notebook instance for the first three models. The rationale for reusing the hyperparameters is that the tuning process is time and resource intensive, and the parameter ranges in different trials with different data subsets were similar to the one obtained with the global model. A separate hyper-parameter tuning was performed for the global datasets excluding the local datasets, i.e., for the local_starved_model dataset. A separate tuning is necessary in this case since the other models all contain some instances of the local data whereas the local_starved_model has none, and this could change the model parameters.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eStep 3: Image classification and accuracy assessment\u003c/h3\u003e\n\u003cp\u003eUtilizing the five models discussed above, we classified the images with each model. In order to estimate the accuracy and compare the different approaches, a stratified random sampling design was implemented based on the recommendations of Olofsson et al. [\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e], since random sampling provides a better estimate of accuracy compared to opportunistic sampling. The methodology for this assessment involved estimating a total number of random points where ground-truth will be assessed by photo-interpretation of Pleiades images. This is the recommended approach since the Pleiades images are of much greater resolution than the S-2 images used in training the classifier. The number of points depends on the estimated accuracy of classification and the proportions of the different land cover classes, or strata, determined from one of the classified images for each site. A total of 431 and 426 points were sampled for Austin and Greater Noida, respectively. It was ensured that a minimum of 35 points were allocated to each stratum (i.e., class), with the allocation being proportional to the class area in the classified map otherwise. Thus, of the 431 points for Austin, the allocation was as follows: water \u0026minus;\u0026thinsp;49, built-up \u0026minus;\u0026thinsp;127, low-vegetation \u0026minus;\u0026thinsp;167, and vegetation \u0026minus;\u0026thinsp;88. Similarly, the 426 points for Greater Noida were distributed as follows: water \u0026minus;\u0026thinsp;39, built-up \u0026minus;\u0026thinsp;117, low-vegetation \u0026minus;\u0026thinsp;216, vegetation \u0026minus;\u0026thinsp;117.\u003c/p\u003e\n\u003ch3\u003eStep 4: Land cover change estimation\u003c/h3\u003e\n\u003cp\u003eAs a demonstration of the application for the different approaches, S-1 and S-2 images corresponding to the same area but from 2017 were also classified using the same methodology described above. A land cover change raster was created in each case by comparing the pixels in the before and after images and assigning them pixel values according to the change matrix shown in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e. The numbers in the table are a lookup table for pixel values assigned to the land cover change image, when comparing the change from before (BF) to after (AF). For instance, if the classification of a pixel in 2017 was built-up and changed to vegetation in 2020, it will be assigned a value of 8.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eMatrix showing indices for land-cover change comparison.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAF - Water\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAF - Built-up\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAF - Low-vegetation/soil\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eAF - Vegetation\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBF - Water\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBF - Built-up\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBF - Low-vegetation/soil\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e12\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBF - Vegetation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e13\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e14\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e15\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e16\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThis methodology is shown as a flowchart in Fig.\u0026nbsp;\u003cspan refid=\"Fig11\" class=\"InternalRef\"\u003e11\u003c/span\u003e. A comparison raster was created to highlight the increase in built-up class area (by highlighting pixels with values of 2, 10, and 14), and the conversion of vegetation to low-vegetation areas (by highlighting pixels with value 15).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e"},{"header":"Results and Discussion","content":"\u003cp\u003eFigure \u003cspan refid=\"Fig12\" class=\"InternalRef\"\u003e12\u003c/span\u003e shows the land cover classification estimates using the five machine learning models for Austin and Greater Noida. The model outputs for Austin are comparable and differences could not be discerned by visual inspection. The outputs for Greater Noida, on the other hand, show a large variation from model to model. For instance, local_model seems more land covered by low-vegetation/soil compared to all the other models. Similarly, the the local_starved_model seems to predict that most areas are \u0026lsquo;built-up\u0026rsquo;. However, as the weightage of local data increases in the models, the built-up area begins to be classified as low-vegetation. This suggests that the local data is adding critical information to the models to distinguish between built-up and low-vegetation classes, and without the local input, the model tends to overestimate the built-up area for Greater Noida.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWe present the accuracy metrics for each class and site estimated for the different models in Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e - Table\u0026nbsp;\u003cspan refid=\"Tab8\" class=\"InternalRef\"\u003e8\u003c/span\u003e. The metrics are shown in boldface if the accuracy is better than 75%. The precision and recall metrics are also referred to as user\u0026rsquo;s accuracy and producer\u0026rsquo;s accuracy, respectively [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]. The precision and recall for water are consistently high (minima of 70.59% and 79.59% respectively) for both sites under all models, suggesting that any of these methods could be a good candidate for use in flood hazard response. With a minimum precision of 81.13% for detecting low-vegetation, the models are also good in accurately identifying low-vegetation areas. However, the recall is not consistently high for low-vegetation detection, indicating that the models underpredict low-vegetation areas. At the same time, most of the models show low precision and high-recall for built-up areas, suggesting that the built-up area is being overpredicted. The F-1 scores for the vegetation classification is below 70% in all the models, and future training could target this particular class to improve its accuracy.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eAccuracy metrics for local_model.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"2\" morerows=\"1\" nameend=\"c2\" namest=\"c1\" rowspan=\"2\"\u003e \u003cp\u003eAccuracy metrics\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c4\" namest=\"c3\"\u003e \u003cp\u003eLocation\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAustin\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eGreater Noida\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eWater\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e85.71\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e70.59\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRecall (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e85.71\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e92.31\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eF1-score (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e85.71\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e80.00\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eBuilt-up\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e66.27\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e68.18\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRecall (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e88.19\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e76.92\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eF1-score (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e75.68\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e72.29\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eLow-vegetation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e92.86\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e86.77\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRecall (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e62.28\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e75.93\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eF1-score (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e74.55\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e80.99\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eVegetation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e64.36\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e64.81\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRecall (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e73.86\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e64.81\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eF1-score (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e68.78\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e64.81\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003eOverall accuracy (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e74.94\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e76.29\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eAccuracy metrics for global_wt2.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"2\" morerows=\"1\" nameend=\"c2\" namest=\"c1\" rowspan=\"2\"\u003e \u003cp\u003eAccuracy metrics\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c4\" namest=\"c3\"\u003e \u003cp\u003eLocation\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAustin\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eGreater Noida\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eWater\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e91.11\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e77.78\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRecall (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e83.67\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e89.74\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eF1-score (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e87.23\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e83.33\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eBuilt-up\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e78.13\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e52.48\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRecall (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e78.74\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e90.60\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eF1-score (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e78.43\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e66.46\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eLow-vegetation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e82.82\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e88.03\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRecall (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e80.84\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e57.87\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eF1-score (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e81.82\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e69.83\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eVegetation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e65.26\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e75.68\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRecall (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e70.45\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e51.85\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eF1-score (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e67.76\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e61.54\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003eOverall accuracy (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e78.42\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e69.01\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab6\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 6\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eAccuracy metrics for global_wt5.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"2\" morerows=\"1\" nameend=\"c2\" namest=\"c1\" rowspan=\"2\"\u003e \u003cp\u003eAccuracy metrics\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c4\" namest=\"c3\"\u003e \u003cp\u003eLocation\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAustin\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eGreater Noida\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eWater\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e87.23\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e72.92\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRecall (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e83.67\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e89.74\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eF1-score (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e85.42\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e80.46\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eBuilt-up\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e73.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e54.79\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRecall (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e77.95\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e88.03\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eF1-score (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e75.57\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e67.54\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eLow-vegetation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e84.00\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e89.26\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRecall (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e75.45\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e61.57\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eF1-score (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e79.50\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e72.88\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eVegetation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e61.62\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e68.29\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRecall (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e69.32\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e51.85\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eF1-score (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e65.24\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e58.95\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003eOverall accuracy (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e75.87\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e70.19\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab7\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 7\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eAccuracy metrics for global_model.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"2\" morerows=\"1\" nameend=\"c2\" namest=\"c1\" rowspan=\"2\"\u003e \u003cp\u003eAccuracy metrics\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c4\" namest=\"c3\"\u003e \u003cp\u003eLocation\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAustin\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eGreater Noida\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eWater\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e84.78\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e75.00\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRecall (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e79.59\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e92.31\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eF1-score (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e82.11\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e82.76\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eBuilt-up\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e75.00\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e50.00\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRecall (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e77.95\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e88.89\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eF1-score (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e76.45\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e64.00\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eLow-vegetation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e81.13\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e87.41\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRecall (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e77.25\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e54.63\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eF1-score (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e79.14\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e67.24\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eVegetation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e64.89\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e65.71\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRecall (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e69.32\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e42.59\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eF1-score (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e67.03\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e51.69\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003eOverall accuracy (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e76.10\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e65.96\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab8\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 8\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eAccuracy metrics for local_starved_model.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"2\" morerows=\"1\" nameend=\"c2\" namest=\"c1\" rowspan=\"2\"\u003e \u003cp\u003eAccuracy metrics\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c4\" namest=\"c3\"\u003e \u003cp\u003eLocation\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAustin\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eGreater Noida\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eWater\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e91.11\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e72.92\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRecall (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e83.67\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e89.74\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eF1-score (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e87.23\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e80.46\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eBuilt-up\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e73.13\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e47.95\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRecall (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e77.17\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e89.74\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eF1-score (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e75.10\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e62.50\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eLow-vegetation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e83.33\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e87.69\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRecall (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e77.84\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e52.78\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eF1-score (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e80.50\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e65.90\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eVegetation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e63.54\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e75.86\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRecall (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e69.32\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e40.74\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eF1-score (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e66.30\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e53.01\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003eOverall accuracy (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e76.57\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e64.79\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe overall accuracy estimates of the five different machine learning classifier modeling approaches for the two study sites are presented in Fig.\u0026nbsp;\u003cspan refid=\"Fig13\" class=\"InternalRef\"\u003e13\u003c/span\u003e. In Austin\u0026rsquo;s case, the different methodologies consistently produce accuracies around 75%, with the maximum accuracy attained with the combined data model where local data is weighted twice the global data (global_wt2). On the other hand, the accuracies for Greater Noida are very sensitive to the training data and model. While the accuracy is highest when using only local data (76.29%), it falls to about 70% when combined with global data but still weighting local data higher. However, using either the local-data-starved model or the global-local integrated model results in lower accuracy around 65%.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe local_model is highly attuned to the band variations in the site\u0026rsquo;s geography. The advantage of this approach is that the ML classifier can be trained better to identify local peculiarities in how the classes are expressed in the band values. However, the drawback with this method is that a reliable training dataset matching the spatial and temporal extent of satellite imagery is necessary for this to work. This can require a significant time and monetary resources in many cases. Another drawback is that the model could be less applicable to other study sites differing in local morphology and topography. The local_model had middling performance for Austin whereas it outperformed the other models for Noida. This could be attributed to the fact that the combined datasets are overrepresented by cities from developed countries similar to Austin (for example, Trenton, Kitchener, and Amos) whereas cities from developing countries such as Greater Noida are underrepresented. Therefore, using only local data aids in improving predictions for the latter while starving the former of relevant data that could have improved accuracy.\u003c/p\u003e \u003cp\u003eThe global_wt2 and global_wt3 models improve the classification accuracy in Austin\u0026rsquo;s case while decreasing the accuracy for Greater Noida. This affirms the aforementioned theory about the differences in applicability of the global data to the individual sites. The global_wt2 is the best-performing model for Austin; however, the differences between the different model performance are small for Austin.\u003c/p\u003e \u003cp\u003eBy removing the higher weight to local data in global_model, the accuracy falls further for Greater Noida while remaining relatively unchanged for Austin. Finally, by completely removing any local data, Austin\u0026rsquo;s LULC classification is quite satisfactory when compared to the other methods while it performs the worst for Greater Noida. Examining the confusion matrix for Noida in Fig.\u0026nbsp;\u003cspan refid=\"Fig14\" class=\"InternalRef\"\u003e14\u003c/span\u003e reveals that the model is overpredicting the built-up land cover and significantly underpredicting low-vegetation/soil class. Despite this inter-class confusion, 65% accuracy might be acceptable for a rough estimation of land cover when local training data is lacking. It is also possible that the accuracy with a global model might improve for developing cities like Greater Noida if more representative samples are added to the training set.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFinally, to demonstrate the application of estimating LULC using cloud-based resources, we present land cover change maps and metrics from the highest accuracy models for each city. The \u0026ldquo;before\u0026rdquo; image was acquired for each city in 2017 whereas the \u0026ldquo;after\u0026rdquo; image was acquired in 2020.\u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig15\" class=\"InternalRef\"\u003e15\u003c/span\u003e shows the classifications for Austin\u0026rsquo;s images. There is evidence of urban expansion in the previously less urban area in the southern half of the map. Loss of vegetation is focused in a few areas of the map and is not widespread. Figure\u0026nbsp;\u003cspan refid=\"Fig16\" class=\"InternalRef\"\u003e16\u003c/span\u003e shows the increase in built-up area in the southern and western regions of the Greater Noida map. Significantly, the new construction of a roughly southwest-northeast running road is captured. There are identifiable losses in vegetation at various places.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe metrics related to land cover change from before- and after- images (2017 and 2022) are presented in Table\u0026nbsp;\u003cspan refid=\"Tab9\" class=\"InternalRef\"\u003e9\u003c/span\u003e. The table rows denote the \u0026ldquo;before\u0026rdquo; classification while the columns denote the \u0026ldquo;after\u0026rdquo; classification. For example, an area of 7.84 sq.km was classified as water both in the 2017 and 2020 images for Austin, while 11.04 sq. km was vegetation in 2017 and became low-vegetation in 2020. While the maps capture the local increase in built-up area in Austin, the metrics indicate that the trend need not be true across the map. However, owing to the known confusion between built-up and low-vegetation areas, it is likely that the classification for 2017 was overestimating the built-up area. The results for Greater Noida reflect a marked increase in built-up area and a significant loss (around 56%) of forested area to low-vegetation areas.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab9\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 9\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eLand cover change metrics for Austin and Noida from global_wt2 and local_model, respectively. All values are in sq. km.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAustin\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAft:Water\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAft:Built-up\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAft:Low-vegetation\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eAft:Vegetation\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBef:Water\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.84\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.16\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.05\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1.24\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBef:Built-up\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e2.45\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e104.36\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e19.61\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e22.39\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBef:Low-vegetation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.75\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e10.94\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e86.14\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e6.11\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBef:Vegetation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.69\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e7.29\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e11.04\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e63.88\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eNoida\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAft:Water\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAft:Built-up\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAft:Low-vegetation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eAft:Vegetation\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBef:Water\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1.47\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.49\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.21\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.16\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBef:Built-up\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.46\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e29.98\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e10.28\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e2.84\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBef:Low-vegetation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.23\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e11.29\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e48.91\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e4.27\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBef:Vegetation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e2.19\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e12.25\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e7.39\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eThis work presented a methodology for implementing land cover classification in Amazon SageMaker using XGBoost algorithm. The cloud-based approach to image analysis is expected to benefit near-real-time data acquisition, analysis, and dissemination of results. SageMaker is particularly beneficial owing to the open availability of Sentinel data in AWS and powerful machine learning tools to analyze the large datasets. Five different approaches were tested with this framework to classify the land cover in Austin and Greater Noida. High-resolution Pleiades images were used to establish ground-truth data through photo-interpretation, which were used to estimate accuracy. The first approach utilized only training data located within the study sites. The second and third approaches utilized a global dataset combining local data and data from other sites trained previously, with the local data assigned larger weights. The fourth approach simply assigned equal weight to the combined global dataset. Finally, the fifth approach utilized only the data external to the sites, with the motivation that this could be useful in cases when local data is difficult to obtain either through photo-interpretation (e.g., cloud cover, corrupted data) or through field visits (e.g., harsh or inaccessible terrain). The results show good overall accuracy (above 70%) using all the approaches for Austin. However, the results for Greater Noida were better with the local data and deteriorated as the out-of-site data was mixed into the training sample. Despite the lowered performance, a minimum accuracy of about 65% was attained with the global data without utilizing any local data. This augments the case for building global representative datasets from a variety of sites that could be deployed on cloud-based land use monitoring systems, with only periodic updates to the local data added as necessary.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eSunil Bhaskaran: Conceptualized the study idea and designed the overall methodology framework. Led the project direction and ensured alignment with urban informatics research themes. Supervised the study and ensured alignment with urban informatics research priorities.Arvindh Sharma: Conducted the technical analysis, including model training, and result interpretation using multi-resolution spaceborne imagery and cloud-based machine learning tools.Sanjiv Bhatia: Contributed to shaping the manuscript\u0026rsquo;s focus, particularly emphasizing connections to urban informatics applications and broader implications for urban resilience and planning.Stuti Mishra: Conducted in-depth review and quality assurance of the analysis workflows, curated and synthesized the literature review, and critically revised the manuscript for technical rigor, urban relevance, and clarity. Contributed to refining the research narrative to better position the study within the domain of urban informatics.\u003c/p\u003e\u003ch2\u003eAcknowledgments:\u003c/h2\u003e \u003cp\u003eThe research was conducted with funding from the NASA-MISTC and United States NSF-ATE programs. Authors acknowledge the funding agencies and infrastructure provided by the BCC Geospatial Center of the CUNY CREST Institute, a leading center of excellence in New York.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eD. Dobrinić, D. Medak and M. Gašparović, \"INTEGRATION OF MULTITEMPORAL SENTINEL-1 AND SENTINEL-2 IMAGERY FOR LAND-COVER CLASSIFICATION USING MACHINE LEARNING METHODS,\" in \u003cem\u003eInt. Arch. Photogramm. Remote Sens. Spatial Inf. Sci.\u003c/em\u003e, Nice, France, 2020.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eE. Lambin, B. Turner, H. Geist, B. Agbola, A. Angelsen, J. Bruce, C. Oliver, R. Dirzo, G. F. Fischer, P. George, K. Homewood, J. Imbernon, R. Leemans, X. Li, E. Moran, M. Mortimore, P. Ramakrishnan, J. Richards and J. Xu, \"The causes of land-use and land-cover change: moving beyond the myths,\" \u003cem\u003eGlobal Environmental Change\u003c/em\u003e, vol. 11, pp. 261\u0026ndash;269, 2001.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJ.-F. Mas, R. Lemoine-Rodr\u0026iacute;guez, R. Gonz\u0026aacute;lez-L\u0026oacute;pez, J. L\u0026oacute;pez-S\u0026aacute;nchez, A. Pi\u0026ntilde;a-Gardu\u0026ntilde;o and E. Herrera-Flores, \"Land use/land cover change detection combining,\" \u003cem\u003eEuropean Journal of Remote Sensing\u003c/em\u003e, vol. 50, no. 1, pp. 626\u0026ndash;635, 2017.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eB. Pradhan, M. S. Tehrany and M. N. Jebur, \"A New Semiautomated Detection Mapping of Flood Extent From TerraSAR-X Satellite Image Using Rule-Based Classification and Taguchi Optimization Techniques,\" \u003cem\u003eIEEE Transactions on Geoscience and Remote Sensing\u003c/em\u003e, vol. 54, no. 7, 2016.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eM. Appel, F. Lahn, W. Buytaert and E. Pebesma, \"Open and scalable analytics of large Earth observation datasets: From scenes to multidimensional arrays using SciDB and GDAL,\" \u003cem\u003eISPRS Journal of Photogrammetry and Remote Sensing\u003c/em\u003e, vol. 138, pp. 47\u0026ndash;56, 2018.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eK. Navulur, Multispectral Image Analysis Using the Object-Oriented Paradigm, CRC Press, 2006.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e\"Landsat Data Access,\" [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.usgs.gov/landsat-missions/landsat-data-access\u003c/span\u003e\u003cspan address=\"https://www.usgs.gov/landsat-missions/landsat-data-access\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. [Accessed 23 July 2022].\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e\"Sentinel-2,\" The European Space Agency, [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://sentinels.copernicus.eu/web/sentinel/missions/sentinel-2\u003c/span\u003e\u003cspan address=\"https://sentinels.copernicus.eu/web/sentinel/missions/sentinel-2\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. [Accessed 23 July 2022].\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e\"MultiSpectral Instrument (MSI) Overview,\" European Space Agency, [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://sentinels.copernicus.eu/web/sentinel/technical-guides/sentinel-2-msi/msi-instrument\u003c/span\u003e\u003cspan address=\"https://sentinels.copernicus.eu/web/sentinel/technical-guides/sentinel-2-msi/msi-instrument\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. [Accessed 04 10 2022].\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eD. Liu and F. Xia, \"Assessing object-based classification: advantages and limitations,\" \u003cem\u003eRemote Sensing Letters\u003c/em\u003e, vol. 1, no. 4, pp. 187\u0026ndash;194, 2010.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e\"Sentinel-1,\" The European Space Agency, [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://sentinels.copernicus.eu/web/sentinel/missions/sentinel-1\u003c/span\u003e\u003cspan address=\"https://sentinels.copernicus.eu/web/sentinel/missions/sentinel-1\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. [Accessed 23 July 2022].\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e\"Geophysical Measurements,\" European Space Agency, [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://sentinels.copernicus.eu/web/sentinel/user-guides/sentinel-1-sar/product-overview/geophysical-measurements\u003c/span\u003e\u003cspan address=\"https://sentinels.copernicus.eu/web/sentinel/user-guides/sentinel-1-sar/product-overview/geophysical-measurements\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. [Accessed 04 10 2022].\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eM. Hall-Beyer, \"GLCM Texture: A Tutorial v. 1.0 through 2.7.,\" 2007. [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://prism.ucalgary.ca/handle/1880/51900\u003c/span\u003e\u003cspan address=\"https://prism.ucalgary.ca/handle/1880/51900\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. [Accessed 22 July 2022].\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eM. Iyyappan and S. S. Ramakrishnan, \"hancing land cover classification for multispectral images using hybrid polarimetry SAR data,\" \u003cem\u003eInternational Journal of Remote Sensing\u003c/em\u003e, vol. 41, no. 17, pp. 6718\u0026ndash;6754, 2020.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eP. T. Noi and M. Kappas, \"Comparison of Random Forest, k-Nearest Neighbor, and Support Vector Machine Classifiers for Land Cover Classification Using Sentinel-2 Imagery,\" \u003cem\u003eSensors\u003c/em\u003e, vol. 18, no. 18, 2017.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eP. T. Noi and M. Kappas, \"Comparison of Random Forest, k-Nearest Neighbor, and Support Vector Machine Classifiers for Land Cover Classification Using Sentinel-2 Imagery,\" \u003cem\u003eSensors\u003c/em\u003e, 2017.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eT. Chen and C. Guestrin, \"XGBoost: A Scalable Tree Boosting System,\" in \u003cem\u003eKDD '16: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining\u003c/em\u003e, San Francisco, California, USA, 2016.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eW. Chen, K. Fu, J. Zuo, X. Zheng, T. Huang and W. Ren, \"Radar emitter classification for large data set based on weighted-xgboost,\" \u003cem\u003eIET Radar, Sonar\u003c/em\u003e \u0026amp; \u003cem\u003eNavigation\u003c/em\u003e, vol. 11, no. 8, pp. 1203\u0026ndash;1207, 2017.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eS. Bhaskaran, E. Nez, K. Jimenez and S. K. Bhatia, \"Rule-based classification of high-resolution imagery over urban areas in New York City,\" \u003cem\u003eGeocarto International\u003c/em\u003e, 2012.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eA. R. Sharma and S. Bhaskaran, \"Deriving community vulnerability indices by analyzing multi-resolution space-borne data and demographic data for extreme weather events in global cities,,\" \u003cem\u003eRemote Sensing Applications: Society and Environment\u003c/em\u003e, vol. 33, 2024.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eF. Cian, C. Giupponi and M. Marconcini, \"Integration of earth observation and census data for mapping a multi-temporal flood vulnerability index: a case study on Northeast Italy,\" \u003cem\u003eNatural Hazards\u003c/em\u003e, p. 106, 2021.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eF. Samadzadegan, A. Toosi and F. Dadrass Javan, \"A critical review on multi-sensor and multi-platform remote sensing data fusion approaches: current status and prospects,\" \u003cem\u003eInternational Journal of Remote Sensing\u003c/em\u003e, p. 76, 2024.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e\"Earth on AWS,\" Amazon Web Services, [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://aws.amazon.com/earth/\u003c/span\u003e\u003cspan address=\"https://aws.amazon.com/earth/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. [Accessed 04 10 2022].\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eN. Gorelick, M. Hancher, M. Dixon, S. Ilyushchenko, D. Thau and R. Moore, \"Google Earth Engine: Planetary-scale geospatial analysis for everyone.,\" \u003cem\u003eRemote Sensing of Environment\u003c/em\u003e, 2017.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eH. Tamiminia, B. Salehi, M. Mahdianpari, L. Quackenbush, S. Adeli and B. Brisco, \"Google Earth Engine for geo-big data applications: A meta-analysis and systematic review,\" \u003cem\u003eISPRS Journal of Photogrammetry and Remote Sensing\u003c/em\u003e, vol. 164, pp. 152\u0026ndash;170, 2020.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eV. C. F. Gomes, G. R. Queiroz and K. R. Ferreira, \"An Overview of Platforms for Big Earth Observation Data Management and Analysis,\" \u003cem\u003eRemote Sensing\u003c/em\u003e, vol. 12, 2020.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eM. A. Brovelli, Y. Sun and V. Yordanov, \"Monitoring Forest Change in the Amazon Using Multi-Temporal Remote Sensing Data and Machine Learning Classification on Google Earth Engine,\" \u003cem\u003eInternational Journal of Geo-Information\u003c/em\u003e, vol. 9, no. 580, 2020.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eC. E. Soulard, J. J. Walker and R. E. Petrakis, \"Implementation of a Surface Water Extent Model in Cambodia using Cloud-Based Remote Sensing,\" \u003cem\u003eRemote Sensing\u003c/em\u003e, vol. 12, no. 984, 2020.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eE. Liberty and others, \"Elastic Machine Learning Algorithms in Amazon SageMaker,\" in \u003cem\u003eSIGMOD '20: Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data\u003c/em\u003e, Portland, OR, USA, 2020.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eP. Helber, B. Bischke, A. Dengel and D. Borth, \"EuroSAT: A Novel Dataset and Deep Learning Benchmark for Land Use and Land Cover Classification,\" \u003cem\u003eIEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing\u003c/em\u003e, vol. 12, no. 7, 2019.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e\"Sentinel-1 SAR GRD: C-band Synthetic Aperture Radar Ground Range Detected, log scaling,\" Google Earth Engine, [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://developers.google.com/earth-engine/datasets/catalog/COPERNICUS_S1_GRD#description\u003c/span\u003e\u003cspan address=\"https://developers.google.com/earth-engine/datasets/catalog/COPERNICUS_S1_GRD#description\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. [Accessed 04 10 2022].\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e\"Sentinel-1 Algorithms,\" Google Earth Engine, [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://developers.google.com/earth-engine/guides/sentinel1\u003c/span\u003e\u003cspan address=\"https://developers.google.com/earth-engine/guides/sentinel1\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. [Accessed 22 July 2022].\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e\"Sentinel-2,\" Amazon, [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://registry.opendata.aws/sentinel-2/\u003c/span\u003e\u003cspan address=\"https://registry.opendata.aws/sentinel-2/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. [Accessed 30 09 2022].\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e\"Sentinel Hub,\" Sinergise Ltd., [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.sentinel-hub.com/\u003c/span\u003e\u003cspan address=\"https://www.sentinel-hub.com/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. [Accessed 30 09 2022].\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAmazon AWS, [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://aws.amazon.com/premiumsupport/knowledge-center/sagemaker-lifecycle-script-timeout/\u003c/span\u003e\u003cspan address=\"https://aws.amazon.com/premiumsupport/knowledge-center/sagemaker-lifecycle-script-timeout/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. [Accessed 17 08 2022].\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e\"XGBoost Algorithm,\" Amazon Web Services, [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://docs.aws.amazon.com/sagemaker/latest/dg/xgboost.html\u003c/span\u003e\u003cspan address=\"https://docs.aws.amazon.com/sagemaker/latest/dg/xgboost.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. [Accessed 03 10 2022].\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eP. Olofsson, G. M. Foody, M. Herold, S. V. Stehman, C. E. Woodcock and M. A. Wulder, \"Good practices for estimating area and assessing accuracy of land change,\" \u003cem\u003eRemote Sensing of Environment\u003c/em\u003e, vol. 148, pp. 42\u0026ndash;57, 2014.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJ. Weaver, B. Moore, A. Reith, J. McKee and D. Lunga, \"A Comparison of Machine Learning Techniques to Extract Human Settlements from High Resolution Imagery,\" in \u003cem\u003eIGARSS 2018\u0026ndash;2018 IEEE International Geoscience and Remote Sensing Symposium\u003c/em\u003e, Valencia, Spain, 2018.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Urban Land Cover Classification, Cloud-Based Geospatial Analytics, Big Data in Remote Sensing Multispectral and Hyperspectral Imaging","lastPublishedDoi":"10.21203/rs.3.rs-6497131/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6497131/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThis paper describes the extraction of critical terrestrial features from multispectral (MSS) spaceborne imagery. We present a methodology to extract terrestrial features from spatiotemporal datasets by using machine learning algorithms, open-source software, and datasets. We used the Amazon Web Services sage-maker framework to conduct full data life cycle (FDLC) projects and automate steps from data acquisition to data mining and visualization. The results demonstrate a robust and efficient model to conduct image analyses with spatio-temporal datasets to extract variables that may be useful for a wide range of applications. The results have significant implications to analyze time-series of spaceborne imagery for both near-real time and other applications that are driven by data from the archives.\u003c/p\u003e","manuscriptTitle":"Automating image classification by using multi-resolution spaceborne imagery, cloud-based machine learning algorithms and open-source software and data ","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-05-07 09:08:21","doi":"10.21203/rs.3.rs-6497131/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"55a24e0f-f2ba-406d-afd7-184989f2fdb7","owner":[],"postedDate":"May 7th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-05-16T15:08:10+00:00","versionOfRecord":[],"versionCreatedAt":"2025-05-07 09:08:21","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6497131","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6497131","identity":"rs-6497131","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.