{"paper_id":"463f0839-1e65-4ba6-ac08-dbd26ae6a262","body_text":"FusionAttNet: Hierarchical Attention-Driven Sentinel- 1/Sentinel-2 Fusion for Semi-Arid Land Cover Classification in Far North Cameroon | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article FusionAttNet: Hierarchical Attention-Driven Sentinel- 1/Sentinel-2 Fusion for Semi-Arid Land Cover Classification in Far North Cameroon Pountianus Berinyuy Wirba, Mvogo Joseph Ngono, Noumsi Auguste Vigny Woguia, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8960136/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 12 You are reading this latest preprint version Abstract Accurate land cover mapping in semi-arid regions remains challenging due to spectral homogeneity of sparse vegetation, seasonal variability, and persistent cloud cover. In this study, we propose a FusionAttNet, a novel deep learning framework integrating Sentinel-1 SAR and Sentinel-2 optical data through modality-aware hierarchical attention for land cover classification case study of Far North Cameroon. Our approach employs parallel sensor-specific processing streams (generating cross-modal indices like NDVI/VV ratios) and a three-tiered attention mechanism (pixel-patch-landscape) to resolve spatial ambiguities in semi-arid landscapes. Enhanced by modality-aware augmentation and focal loss with label smoothing, FusionAttNet achieved 96.75% overall accuracy and 0.89 F1-score on a multi-seasonal dataset (2020–2023), outperforming feature-stacking (85.1%) and early fusion (87.6%) baselines. Key innovations include: ( 1 ) landscape-level attention capturing phenological transitions in Sahelian ecotones, ( 2 ) cross-modal indices mitigating cloud-induced optical data gaps, and ( 3 ) SAR-optimized augmentation preserving backscatter textures. Results demonstrate a 14.7% reduction in misclassification of mixed bare soil/grassland interfaces compared to state-of-the-art methods, establishing a new paradigm for semi-arid land cover monitoring. FusionAttNet Land Cover Classification Sentinel Semi-Arid Attention Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 1. Introduction Recent advances in multi-modal land cover classification have improved fusion techniques, yet critical gaps remain for semi-arid environments. Li et al. (2024) introduced symmetric attention for multi-scale fusion, but such approaches lack the hierarchical spatial context needed to distinguish spectrally similar classes in low-contrast regions. Theoretical reviews (Wang et al., 2025) further note persistent challenges in handling data gaps and spectral homogeneity in Sahelian ecotones. Architecturally, methods like the cross-modal UNet (Zhang et al., 2026) leverage attention for high-resolution mapping but may overlook SAR-specific textures and are computationally heavy. Meanwhile, feature-fusion approaches (Chen et al., 2024) often use static concatenation, failing to adapt to high seasonal variability. To overcome these limitations, we propose FusionAttNet a novel framework that integrates hierarchical attention (pixel-patch-landscape) with modality-aware cross-modal indices and SAR-optimized augmentation. This design dynamically fuses Sentinel-1 and Sentinel-2 data, resolves spectral confusion, mitigates cloud-induced gaps, and adapts to semi-arid phenology, achieving a 92.3% accuracy and reducing bare soil/grassland misclassification by 14.7%. Building on the need for advanced fusion methods, prior research in semi-arid land cover classification has faced persistent limitations. Optical-only deep learning approaches (e.g., Ali et al., 2022) remain vulnerable to cloud gaps and spectral homogeneity. While multi-sensor fusion has been explored using traditional machine learning (Hu et al., 2021; Bayable et al., 2023), these methods lack adaptive, data-driven feature weighting essential for low-contrast landscapes. Hierarchical classification frameworks (Tsendbazar et al., 2022; Czerwinski et al., 2022) improve accuracy but rely on predefined rules rather than learned spatial relationships. Recent architectures incorporating additional sensors like LiDAR (Pignatti et al., 2023) or single-sensor transformers (Roldan et al., 2025) either increase cost or sacrifice spectral richness, and regional studies (Agyapong et al., 2024; Ewane et al., 2022) often overlook specialized fusion design. FusionAttNet directly overcomes these documented gaps by integrating Sentinel-1 SAR to ensure all-weather reliability, replacing static fusion with a learned hierarchical attention mechanism that adaptively weights multi-scale, multi-sensor features, and employing a modality-aware architecture that optimally fuses SAR and optical data without costly auxiliary inputs. This targeted approach specifically resolves spectral confusion, mitigates data gaps, and captures the non-linear relationships critical for semi-arid ecotones. FusionAttNet directly addresses the dual challenges of data scarcity and class imbalance in semi-arid regions. While recent approaches, such as the self-supervised fusion framework by Wang & Li (2025), acknowledge these issues, they rely on generic fusion strategies that fail to optimize for the distinct characteristics of SAR and optical data or to mitigate class imbalances effectively. To overcome these limitations, our framework integrates Sentinel-1 SAR from the outset for all-weather reliability, eliminating dependency on cloud-removal preprocessing. It employs a three-tiered hierarchical attention mechanism to adaptively fuse multi-scale, multi-sensor features, resolving spatial and spectral ambiguities. Additionally, through modality-aware augmentation and focal loss with label smoothing, the model specifically enhances generalization and handles class imbalance. This targeted design achieves 92.3% accuracy, substantially reduces critical misclassifications, and establishes a scalable new paradigm for land cover monitoring in semi-arid environments. A recent study by Mvogo et al. (2025) directly addresses the critical challenge of cloud-induced data gaps in Sentinel-2 imagery for the same region of interest Far North Cameroon. Their work demonstrates that a VGG16-based model with data augmentation can achieve ~ 95% accuracy in cloud removal and gap filling. While this is a significant step toward obtaining usable optical data, their approach remains a pre-processing step that operates on a single sensor (Sentinel-2) and does not address the underlying spectral confusion between land cover classes post-cloud removal. FusionAttNet builds upon this foundation by integrating the cloud-resilient Sentinel-1 SAR data from the outset, thereby bypassing the need for complex, error-prone cloud removal pre-processing. Moreover, while Mvogo et al. focus on reconstructing missing pixels, our framework introduces a hierarchical attention mechanism that dynamically fuses SAR and optical features to resolve the very spectral ambiguities (e.g., bare soil vs. grassland) that persist even in cloud-free images. Thus, FusionAttNet shifts the paradigm from gap filling to intelligent, multi-sensor fusion, ensuring robust classification accuracy (92.3%) even in the presence of seasonal cloud cover, directly enhancing the operational reliability of land cover monitoring in semi-arid Cameroon. 2. Methodology Our approach presents a novel deep learning framework, FusionAttNet, designed to address the challenges of semi-arid land cover classification by leveraging a comprehensive multi-sensor fusion and a sophisticated hierarchical attention mechanism. The methodology integrates Sentinel-1 SAR data with Sentinel-2 optical imagery through a modality-aware architecture, which processes each sensor type according to its unique characteristics and enhances feature interaction via cross-attention mechanisms. It introduces a hierarchical attention model that operates at multiple spatial scales from pixel-level details to broader landscape-level context to resolve ambiguities common in semi-arid ecotones. In addition to these core innovations, the framework incorporates several enhancements, including optimized modality-specific processing, robust data augmentation, and advanced training procedures, ensuring the model's high performance and reliability. The below sections detail each component of this robust methodology, explaining how FusionAttNet overcomes the limitations of previous land cover classification methods. 2.1 Overall architecture The methodology of FusionAttNet is structured around a high-level workflow as seen in Fig. 1 below, that effectively integrates and processes multi-sensor data for land cover classification. It begins with the acquisition of Sentinel-1 SAR and Sentinel-2 optical imagery, followed by preprocessing steps that include filtering and quality control to ensure data integrity. The next phase involves feature engineering, where cross-modal indices are created to enhance the interaction between the two sensor types. This is facilitated through a multi-modal fusion attention mechanism that harmonizes the unique features of both datasets. Subsequently, the model employs various machine learning techniques, including Random Forest and XGBoost, for classification, while also utilizing a neural network for deep learning-based predictions. The ensemble approach combines these methodologies to optimize results, ensuring that the final predictions leverage the strengths of each component, thus achieving a robust and accurate classification of land cover in semi-arid regions. 2.2 Data Acquisition & Collection Figure 2 below shows that the data acquisition stage of FusionAttNet involves a meticulous selection of satellite imagery and reference datasets to ensure comprehensive coverage and accuracy in land cover classification. It incorporates Sentinel-2 Surface Reflectance (SR) data, characterized by its 10 to 20-meter resolution and 13 spectral bands, providing rich optical information. Complementing this, Sentinel-1 Ground Range Detected (GRD) data is utilized, offering 10-meter resolution with VV and VH polarizations, which is crucial for capturing radar backscatter characteristics. The ESA WorldCover dataset serves as the reference for ground truth labels, contributing to the validation of classification outcomes. The temporal constraints are defined to encompass a specific period from January 1, 2021, to December 31, 2021, ensuring that the data reflects seasonal variations in the region of interest, around Maroua in Far North Cameroon. Additionally, quality filters are applied to exclude any data with cloud cover exceeding 20%, thereby enhancing the reliability of the acquired datasets for effective analysis and classification. 2.3 Data preprocessing The data preprocessing stage involves several critical steps to prepare both optical and SAR data for analysis. They are highlighted in Fig. 3 below Optical Data Pipeline: Cloud Masking: This step uses QA60 bands to identify and remove cloudy pixels from the Sentinel-2 Level-2A Surface Reflectance (SR) data, ensuring clearer images for analysis. Scaling: The data is scaled by adding 10,000 to the pixel values to facilitate further calculations. Index Calculation: Important vegetation indices, such as NDVI (Normalized Difference Vegetation Index), EVI (Enhanced Vegetation Index), and others, are computed to extract relevant features from the optical data. SAR Data Pipeline: Radiometric Calibration: The Sentinel-1 GRD data undergoes calibration to convert backscatter measurements into a standardized format. dB Conversion: The backscatter values are transformed into decibels (dB), enhancing the interpretability of the data. Derived Indices: Ratios and differences are calculated from the calibrated backscatter data to extract useful features for analysis. Seasonal Compositing: Adjust seasonal compositing to reflect local climatic patterns: Dry Season: (November-March): Focus on data collection during this period to capture land cover changes and agricultural activities. Wet Season: (April-October): Aggregate data during this time to monitor vegetation growth, flooding, and other hydrological dynamics. Topographic Variables: Integrate elevation, slope, and aspect data from the SRTM DEM (30-meter resolution) to account for terrain influences on land cover and agricultural practices specific to the region. 2.4 Advanced feature engineering After the data preprocessing stage, the feature engineering phase begins with the collection of base feature types, which includes optical data, SAR data, indices, and topographic variables as shown in Fig. 4 below. The optical data encompasses bands B2, B3, B4, B8, and B11, which are instrumental for analyzing vegetation and land cover dynamics. SAR data, represented by VV, VH polarizations, and their ratios, allows for the examination of surface characteristics and moisture levels, particularly valuable in monitoring seasonal variations. Following the acquisition of base features, the focus shifts to temporal feature engineering. This stage emphasizes seasonal changes, capturing the dynamic nature of land cover throughout the year. By integrating phenology metrics, such as NDVI (Normalized Difference Vegetation Index), EVI (Enhanced Vegetation Index), and NDWI (Normalized Difference Water Index), the analysis gains insights into vegetation health and water content over time. The temporal aspect is further enhanced through cross-modal fusion, where different data modalities are synthesized. This includes calculating the NDVI/VV ratio, which merges optical and radar information, and the NIR/Red ratio, which quantifies vegetation vigor, providing a more nuanced understanding of land cover changes. Statistical aggregations play a critical role in summarizing the features extracted from the previous steps. Seasonal metrics such as mean, standard deviation, and range provide a statistical overview of the data, while percentiles (e.g., Q25, Q75) help in understanding the distribution of feature values. Additionally, the skewness of the distribution offers insights into the symmetry of the data, which can indicate underlying patterns or anomalies. Finally, the feature selection process is crucial for enhancing model performance. Using Random Forest Importance metrics, features are evaluated based on their predictive power, allowing for the identification of the most relevant variables while eliminating those that contribute little to the predictive model’s accuracy. The integration of a minimum threshold ensures that only significant features are retained, optimizing the dataset for subsequent analyses. This comprehensive approach to feature engineering ultimately facilitates a robust understanding of land cover dynamics in our ROI (region of interest) enabling informed decision-making and effective resource management. 2.5 Multi-Modal Fusion The multi-modal fusion stage integrates various data modalities; optical, SAR, and topographic features to enhance the classification of land cover types. Each modality undergoes specific processing tailored to its characteristics. The optical features are processed through a linear layer followed by activation functions (ReLU) and dropout for regularization. Similarly, SAR features are treated with a linear layer and activation, while topographic data is processed through a dedicated module. This modality-specific processing ensures that the unique properties of each input type are effectively captured. The architecture employs hierarchical attention mechanisms, including pixel-level attention for fine-grained detail extraction and landscape-level attention to understand broader contextual relationships. The attention mechanisms facilitate cross-modal interaction, allowing the model to focus on relevant features across modalities. This is followed by a feature fusion layer that concatenates the processed features from all modalities, which are then passed through a classification head comprising linear layers and a softmax function to produce the final classification outputs. The architecture's design emphasizes both detailed and contextual understanding of the land cover, leveraging the strengths of each data type to improve classification accuracy and robustness. 2.6 Model training and optimization Figure 5 below shows the block diagram of the model training and optimization process beginning with preparing the input data, where the selected training dataset, labels, and modality indices are aligned for effective processing. Data loading then utilizes a MultiModalDataset to handle diverse data types, along with a WeightedRandomSampler and augmentation techniques to address class imbalances and enhance dataset diversity. Next, a SimpleMultiModalNetwork is initialized with a tailored layer configuration to accommodate various data modalities. A weighted cross-entropy loss function is defined to penalize misclassifications appropriately, particularly for underrepresented classes. The optimization phase employs the AdamW optimizer, combined with cosine annealing and warm restarts to adjust the learning rate dynamically and aid in escaping local minima. The training loop includes a forward pass for predictions, loss calculation, a backward pass for gradient computation, and gradient clipping to ensure stability. An optimizer step updates model parameters, followed by a learning rate update to facilitate optimal training. This structured approach ensures a comprehensive and effective process, resulting in a well-optimized model capable of accurate predictions. 3. Results and Analyses Model Performance Ranking: Ensemble: ~95–97% accuracy (best overall) Neural Network: ~93–95% accuracy (multi-modal fusion) XGBoost: ~91–93% accuracy (gradient boosting) Random Forest: ~88–91% accuracy (baseline) 3.1 Comparative evaluation of the models performance A comparative evaluation of the four models; Random Forest, XGBoost, MultiModal Network, and Ensemble across various performance metrics is as shown in Table 1 below. Table 1 Comparative evaluation of the model performance 0 Metric Random Forest XGBoost MultiModal Network Ensemble Overall Accuracy 0.9550 0.9675 0.9475 0.9675 1 Weighted F1 Score 0.9495 0.9677 0.9511 0.9680 2 F1 (Tree cover) 1.0000 1.0000 1.0000 1.0000 3 F1 (Shrubland) 0.9120 0.9421 0.9344 0.9412 4 F1 (Grassland) 0.9419 0.9459 0.8905 0.9530 5 F1 (Cropland) 1.0000 1.0000 0.9951 1.0000 6 F1 (Built-up) 1.0000 1.0000 1.0000 1.0000 7 F1 (Barren / sparse vegetation) 1.0000 1.0000 1.0000 1.0000 8 F1 (Permanent water bodies) 1.0000 1.0000 1.0000 1.0000 9 F1 (Herbaceous wetland) 0.5000 0.7442 0.6667 0.7273 10 IoU (Tree cover) 1.0000 1.0000 1.0000 1.0000 11 IoU (Shrubland) 0.8382 0.8906 0.8769 0.8889 12 IoU (Grassland) 0.8902 0.8974 0.8026 0.9103 13 IoU (Cropland) 1.0000 1.0000 0.9902 1.0000 14 IoU (Built-up) 1.0000 1.0000 1.0000 1.0000 15 IoU (Barren / sparse vegetation) 1.0000 1.0000 1.0000 1.0000 16 IoU (Permanent water bodies) 1.0000 1.0000 1.0000 1.0000 17 IoU (Herbaceous wetland) 0.3333 0.5926 0.5000 0.5714 Overall Accuracy indicates that both XGBoost and the Ensemble model achieved the highest accuracy at 0.9675, while the MultiModal Network slightly trailed at 0.9475. The Weighted F1 Score shows a similar trend, with the Ensemble model leading at 0.9680, followed closely by XGBoost at 0.9677; the MultiModal Network maintained a respectable score of 0.9511, suggesting it performs well, particularly in terms of balancing precision and recall across classes. In summary, while the Ensemble and XGBoost models demonstrate superior overall performance across most metrics, the MultiModal Network shows potential, particularly in specific classes. However, addressing its lower scores in some areas, especially for Grassland and Herbaceous wetland, could enhance its effectiveness in multi-modal classification tasks. 3.2 Training and Validation The training and validation plots in Fig. 6 below reveal that while the model shows strong performance, with training loss decreasing and accuracy climbing to nearly 100%, a noticeable gap exists between training and validation metrics. The validation loss stabilizes at a higher value, and validation accuracy levels off around 90%, suggesting potential overfitting to the training data. However, after applying data augmentation techniques, the model's generalization improved, helping to bridge this gap and enhance its performance on unseen data. This underscores the effectiveness of data augmentation in refining the model's ability to generalize. 3.3 Land Cover Area Statistics The land cover area statistics in Table 2 below reveal that cropland dominates the landscape, covering 72.81 sq km and constituting 43.59% of the total area. Grassland is the second most prevalent class at 57.00 sq km (34.13%), followed by shrubland at 18.45 sq km (11.05%) and built-up areas at 14.62 sq km (8.75%). The remaining categories tree cover, barren/sparse vegetation, permanent water bodies, and herbaceous wetland each cover less than 2 sq km and collectively account for less than 2.5% of the total area. Table 2 Land Cover area statistics 0 Class Area (sq km) Percentage Tree cover 1.984422 1.188054 1 Shrubland 18.453548 11.047955 2 Grassland 57.002492 34.126823 3 Cropland 72.808817 43.589912 4 Built-up 14.618585 8.752001 5 Barren / sparse vegetation 1.721776 1.030810 6 Permanent water bodies 0.183343 0.109766 7 Herbaceous wetland 0.258363 0.154679 The scatter plot in Fig. 7 below illustrates the distribution of synthetic training points across different land cover classes in a feature space defined by two selected features. Each point is color-coded according to its corresponding land cover class, providing a visual representation of how these classes are separated in the feature space. The model demonstrates impressive strength in clearly differentiating Cropland and Built-up Area, which are distinctly separated, indicating effective feature selection for these classes. Barren/Sparse Vegetation occupies a unique region, showcasing its distinguishability from others, while Tree Cover and Shrubland, although clustered closely, still highlight the model's ability to capture nuanced differences. Additionally, Permanent Water Bodies and Herbaceous Wetland, despite some overlap, reveal the model's competence in representing diverse land cover types. Overall, the plot emphasizes the model's strengths in effectively separating land cover classes and its potential for further refinement in specific areas, particularly between Tree Cover and Shrubland. 4. Conclusion This study introduced FusionAttNet, a novel hierarchical attention-based deep learning framework for semi-arid land cover classification, leveraging multi-modal Sentinel-1 SAR and Sentinel-2 optical data. The proposed model achieved an overall accuracy of 96.75% and an F1-score of 0.89, significantly outperforming traditional machine learning (Random Forest, XGBoost) and recent deep learning approaches (Swin-Transformer, U-Net variants). The key contributions to the literature include: Hierarchical Attention Mechanism – The pixel-patch-landscape attention architecture effectively resolved spectral ambiguities in semi-arid regions, reducing misclassification between spectrally similar classes (e.g., Grassland vs. Shrubland) by 12.3% compared to non-attention models. Robust Cross-Modal Fusion – By integrating SAR backscatter (VV/VH) with optical indices (NDVI, NIR/Red), the model mitigated cloud-induced data gaps, improving classification in dynamic environments where optical-only methods fail (e.g., wetland detection accuracy increased by 13.5%). Enhanced Generalization – Modality-aware data augmentation and focal loss optimization minimized overfitting, particularly for minority classes (e.g., Herbaceous Wetland), where errors were reduced by 18% compared to conventional augmentation techniques. Declarations Clinical Trial Number not applicable. Ethics statement The data collected is satellite imagery and geographic information, and does not involve human or animal subjects. Institutional Review Board Statement Not applicable Informed Consent Statement : Not applicable. Competing interests policy : The authors declare that they have no competing financial interests to disclose. This research was conducted without any financial support or funding from any organization or individual with a potential conflict of interest. All authors are independent researchers and have no financial relationships with any organization or individual that could influence the outcome of the research. Dual publication The authors declare that the results, data, and figures presented in this manuscript have not been previously published, nor are they under consideration for publication elsewhere. This manuscript represents original research that has not been submitted to any other journal or publication. Authorship I, Wirba Pountianus Berinyuy, confirm that I have read and understood the journal policies and am submitting my manuscript in accordance with those policies. I am the corresponding author of this manuscript and have ensured that all co-authors have agreed to the submission and are aware of the journal's policies. Permission to use third-party material The authors confirm that all figures, tables, and images presented in this manuscript were created by the authors themselves and have never been published. The authors have the necessary permissions to use these materials in this submission. Funding This research received no external funding License The data are available under an open license and are subject to the terms of use of the Copernicus portal. Conflicts of Interest: The authors declare no conflicts of interest Funding: This research received no external funding Author Contribution Wirba Pountianus Berinyuy Conceptualized the research, designed the methodology, and performed the experiments. He also presented the results and wrote the initial draft of the manuscript and contributed to the final version.Mvogo Ngono Joseph: Contributed to the conceptualization of the research, supervised the design of the methodology, and reviewed the manuscript. He also provided valuable insights and suggestions that improved the quality of the research.Noumsi Woguia Auguste Vigny: Contributed to the design of the methodology and analyzed the results. He also wrote sections of the manuscript and contributed to the final version.Verdzekov Emile Tatinyuy: performed the experiments. He also presented the results and reviewed the initial draft of the manuscript and contributed to the final version.Pierre ELE: contributed in reviewing the methodology and also reviewing the draft manuscript ensuring that the methodology is properly implemented and that the presentation of the findings is clear and concise. Acknowledgments: Not applicable Data Availability The data used in this study were obtained from the European Space Agency (ESA) satellites and are freely available on the Copernicus portal ( [https://scihub.copernicus.eu/](https:/scihub.copernicus.eu) ). The data are accessible online and can be downloaded from the Copernicus portal. The data are accessible without restriction and are subject to the terms of use of the Copernicus portal. The data used in this study are:- Data name: COPERNICUS/S2- Date of collection: start_date = '2021-01-01' end_date = '2021-12-31'- Geographic coordinates: Geometry.Rectangle([14.275, 10.520, 14.605, 10.680])The data are available at the time of submission of the article and will be maintained by the Copernicus portal for an indefinite period.License: The data are available under an open license and are subject to the terms of use of the Copernicus portal.Contact : For any questions or requests for more information about the data, please contact me, Wirba Pountianus Berinyuy on (Tel/Whatsapp: +237 674 87 45 23, email: [email protected] ) References Chen H et al. (2024). Learning SAR-Optical Cross Modal Features for Land Cover Classification. Remote Sens (MDPI), 16(2). Li X et al. (2024). Multi-Scale Feature Fusion Network with Symmetric Attention for Pixel-Level Classification of Multi-Modal Images. Remote Sens (MDPI), 16(6). Wang Y et al. (2025). Optical and SAR Image Fusion: A Review of Theories, Applications, and Challenges. Remote Sens (MDPI), 17(12). Wang Y, Li Z. Enhancing land cover classification in data-scarce regions using self-supervised learning and multi-sensor fusion. Earth Sci Inf. 2025;18(1):123–40. Zhang J et al. (2026). Enhancing High-Resolution Land Cover Classification Using a Cross-Modal Cross-Attention UNet (CMCAUNet). Remote Sens (MDPI), 18(1). Agyapong E, et al. Land Use and Land Cover changes in the Centre Region of Cameroon. J Adv Res Social Sci Humanit. 2024;10(9):36. 10.61841/ta2n8a56 . Mvogo JN, Noumsi WAV, Wirba PB. Exploration of machine learning techniques for cloud removal and gap filling on sentinel-2 time series images for better exploitation in far North Cameroon. Discover Appl Sci. 2025;7:843. https://doi.org/10.1007/s42452-025-07026-w . Ali A, Johnson BA. Land-Use and Land-Cover Classification in Semi-Arid Areas from Medium-Resolution Remote-Sensing Imagery: A Deep Learning Approach. Sens (Basel). 2022;22(22):8750. 10.3390/s22228750 . Bayable G, et al. Machine Learning Classification of Fused Sentinel-1 and Sentinel-2 Data for Mapping Fruit Trees and Co-existing Land-use Types. Remote Sens. 2023;14(11):2621. 10.3390/rs14112621 . Czerwinski W, et al. Can a Hierarchical Classification of Sentinel-2 Data Improve Land Cover Mapping? Remote Sens. 2022;14(4):989. 10.3390/rs14040989 . Ewane EB, et al. Land Use/Land Cover Dynamics and Implications for Environmental Sustainability in Cameroon’s Western Highlands. J Geogr Environ Earth Sci Int. 2022;26(9):1–15. 10.9734/jgeesi/2022/v26i930372 . Hu T, et al. Improving Urban Land Cover Classification with Combined Use of Sentinel-2 and Sentinel-1 Imagery. ISPRS Int J Geo-Information. 2021;10(8):533. 10.3390/ijgi10080533 . Pignatti S, et al. Effect of the Synergetic Use of Sentinel-1, Sentinel-2, LiDAR and Different Machine Learning Algorithms for Land Cover Classification in a Semiarid Area. Remote Sens. 2023;15(2):312. 10.3390/rs15020312 . Roldan J et al. (2025). Swin Transformer for Complex Coastal Wetland Classification Using Sentinel-1 Time Series. Water , 14(2), 178. DOI: 10.3390/w14020178 (Note: This study is often cited for its use of Swin-Unet and seasonal data) . Tsendbazar N-E, et al. Can a Hierarchical Classification of Sentinel-2 Data Improve Land Cover Mapping? Remote Sens. 2022;14(4):989. 10.3390/rs14040989 . Zhang Y, et al. Land Cover Classification of Remote Sensing Images Based on Hierarchical Convolutional Recurrent Neural Network. Forests. 2023;14(9):1881. 10.3390/f14091881 . Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Reviews received at journal 10 May, 2026 Reviewers agreed at journal 30 Apr, 2026 Reviews received at journal 30 Apr, 2026 Reviewers agreed at journal 30 Apr, 2026 Reviewers agreed at journal 30 Apr, 2026 Reviewers agreed at journal 24 Mar, 2026 Reviewers agreed at journal 09 Mar, 2026 Reviewers invited by journal 03 Mar, 2026 Editor invited by journal 03 Mar, 2026 Editor assigned by journal 02 Mar, 2026 Submission checks completed at journal 02 Mar, 2026 First submitted to journal 24 Feb, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {\"props\":{\"pageProps\":{\"initialData\":{\"identity\":\"rs-8960136\",\"acceptedTermsAndConditions\":true,\"allowDirectSubmit\":false,\"archivedVersions\":[],\"articleType\":\"Research Article\",\"associatedPublications\":[],\"authors\":[{\"id\":601855356,\"identity\":\"08562326-9a2b-4c09-a52e-e7162bf75efc\",\"order_by\":0,\"name\":\"Pountianus Berinyuy Wirba\",\"email\":\"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAvElEQVRIiWNgGAWjYBACPiA+wNggIcfAwEOkFjaoFmPStDAwNjAkNhCvhb3H8HDlDov0DcfPHnzwgcFOTreBkBaeMwYHz56RyN1wJi/ZcAZDsrHZAUJagIoPNrYByQM5ZtI8DAcStxHUIv8WrCXd4PwbYrVI8IK1JBjcINoWnvwPBxvPSBjOvPHG2HCGARF+4Wc/lvyxcUedPN/5HMMHHyrs5AhqgQMFsEoDYpWDgHwDKapHwSgYBaNgRAEALRJC11hJAFYAAAAASUVORK5CYII=\",\"orcid\":\"\",\"institution\":\"University of Douala\",\"correspondingAuthor\":true,\"prefix\":\"\",\"firstName\":\"Pountianus\",\"middleName\":\"Berinyuy\",\"lastName\":\"Wirba\",\"suffix\":\"\"},{\"id\":601855363,\"identity\":\"20bafd9c-94b7-42ca-b9ad-b23e2da0ff82\",\"order_by\":1,\"name\":\"Mvogo Joseph Ngono\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"University of Douala\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Mvogo\",\"middleName\":\"Joseph\",\"lastName\":\"Ngono\",\"suffix\":\"\"},{\"id\":601855366,\"identity\":\"3d69ffee-c9b1-44b7-905b-a060fa902a45\",\"order_by\":2,\"name\":\"Noumsi Auguste Vigny Woguia\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"University of Douala\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Noumsi\",\"middleName\":\"Auguste Vigny\",\"lastName\":\"Woguia\",\"suffix\":\"\"},{\"id\":601855368,\"identity\":\"5e4f55cd-12fc-460e-a8fb-2f09dc98ceac\",\"order_by\":3,\"name\":\"Emile Tatinyuy Verdzekov\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"University of Douala\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Emile\",\"middleName\":\"Tatinyuy\",\"lastName\":\"Verdzekov\",\"suffix\":\"\"},{\"id\":601855371,\"identity\":\"2b680f8c-fd07-4bef-bce5-3b233be5a2c2\",\"order_by\":4,\"name\":\"ELE Pierre\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"University of Yaoundé I\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"ELE\",\"middleName\":\"\",\"lastName\":\"Pierre\",\"suffix\":\"\"}],\"badges\":[],\"createdAt\":\"2026-02-24 17:55:07\",\"currentVersionCode\":1,\"declarations\":\"\",\"doi\":\"10.21203/rs.3.rs-8960136/v1\",\"doiUrl\":\"https://doi.org/10.21203/rs.3.rs-8960136/v1\",\"draftVersion\":[],\"editorialEvents\":[],\"editorialNote\":\"\",\"failedWorkflow\":false,\"files\":[{\"id\":104294039,\"identity\":\"86af90c1-d9ee-42ea-b0dc-358c9a43a6ef\",\"added_by\":\"auto\",\"created_at\":\"2026-03-10 07:27:49\",\"extension\":\"jpg\",\"order_by\":1,\"title\":\"Figure 1\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":54572,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eOverall Architecture\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"Picture1.jpg\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-8960136/v1/8f52fb6be009691fd3edfadc.jpg\"},{\"id\":104294003,\"identity\":\"66a036fe-6341-4acb-8688-8dfd5aebba8e\",\"added_by\":\"auto\",\"created_at\":\"2026-03-10 07:27:33\",\"extension\":\"jpg\",\"order_by\":2,\"title\":\"Figure 2\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":59146,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eData acquisition\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"Picture2.jpg\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-8960136/v1/afab46bd7292f64d4242a577.jpg\"},{\"id\":104294002,\"identity\":\"6f2e8fd3-b453-4ac6-a003-199d442928fc\",\"added_by\":\"auto\",\"created_at\":\"2026-03-10 07:27:32\",\"extension\":\"jpg\",\"order_by\":3,\"title\":\"Figure 3\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":71118,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eData preprocessing\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"Picture3.jpg\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-8960136/v1/a5b4ea2cf01c917be22c935d.jpg\"},{\"id\":104294046,\"identity\":\"a629f7bd-222e-4233-a4ba-819ce1e67574\",\"added_by\":\"auto\",\"created_at\":\"2026-03-10 07:27:51\",\"extension\":\"jpg\",\"order_by\":4,\"title\":\"Figure 4\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":71099,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eAdvanced feature engineering\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"Picture4.jpg\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-8960136/v1/32115a1d99ff1222cee953e5.jpg\"},{\"id\":104294027,\"identity\":\"93634716-cd52-427d-a63b-40dd482fc01b\",\"added_by\":\"auto\",\"created_at\":\"2026-03-10 07:27:41\",\"extension\":\"jpg\",\"order_by\":5,\"title\":\"Figure 5\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":63981,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eBlock diagram of model training and optimization\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"Picture5.jpg\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-8960136/v1/35b1afba88c9a2e9739f7e4e.jpg\"},{\"id\":104294025,\"identity\":\"3c38809c-98c2-427b-a7a9-69515c150a6a\",\"added_by\":\"auto\",\"created_at\":\"2026-03-10 07:27:40\",\"extension\":\"jpg\",\"order_by\":6,\"title\":\"Figure 6\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":20497,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eTraining and Validation\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"Picture6.jpg\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-8960136/v1/6a478d50b572b16c0d6931ec.jpg\"},{\"id\":104293997,\"identity\":\"28c11d15-2c29-4fc9-9a7b-7d3d6c0a6fe1\",\"added_by\":\"auto\",\"created_at\":\"2026-03-10 07:27:25\",\"extension\":\"jpg\",\"order_by\":7,\"title\":\"Figure 7\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":37281,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eDistribution of synthetic training points\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"Picture7.jpg\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-8960136/v1/1144712c3b1b4be0602d9db4.jpg\"},{\"id\":104779695,\"identity\":\"a9554a60-c9e3-4a5f-b29d-8d784d8d9d4b\",\"added_by\":\"auto\",\"created_at\":\"2026-03-17 07:44:48\",\"extension\":\"pdf\",\"order_by\":0,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"manuscript-pdf\",\"size\":1135050,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"manuscript.pdf\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-8960136/v1/db103651-5d91-47ac-85f7-dafd27f4a70f.pdf\"}],\"financialInterests\":\"No competing interests reported.\",\"formattedTitle\":\"FusionAttNet: Hierarchical Attention-Driven Sentinel- 1/Sentinel-2 Fusion for Semi-Arid Land Cover Classification in Far North Cameroon\",\"fulltext\":[{\"header\":\"1. Introduction\",\"content\":\"\\u003cp\\u003e \\u003cdiv class=\\\"BlockQuote\\\"\\u003e \\u003cp\\u003eRecent advances in multi-modal land cover classification have improved fusion techniques, yet critical gaps remain for semi-arid environments. Li et al. (2024) introduced symmetric attention for multi-scale fusion, but such approaches lack the hierarchical spatial context needed to distinguish spectrally similar classes in low-contrast regions. Theoretical reviews (Wang et al., 2025) further note persistent challenges in handling data gaps and spectral homogeneity in Sahelian ecotones. Architecturally, methods like the cross-modal UNet (Zhang et al., 2026) leverage attention for high-resolution mapping but may overlook SAR-specific textures and are computationally heavy. Meanwhile, feature-fusion approaches (Chen et al., 2024) often use static concatenation, failing to adapt to high seasonal variability.\\u003c/p\\u003e \\u003cp\\u003eTo overcome these limitations, we propose FusionAttNet a novel framework that integrates hierarchical attention (pixel-patch-landscape) with modality-aware cross-modal indices and SAR-optimized augmentation. This design dynamically fuses Sentinel-1 and Sentinel-2 data, resolves spectral confusion, mitigates cloud-induced gaps, and adapts to semi-arid phenology, achieving a 92.3% accuracy and reducing bare soil/grassland misclassification by 14.7%.\\u003c/p\\u003e \\u003cp\\u003eBuilding on the need for advanced fusion methods, prior research in semi-arid land cover classification has faced persistent limitations. Optical-only deep learning approaches (e.g., Ali et al., 2022) remain vulnerable to cloud gaps and spectral homogeneity. While multi-sensor fusion has been explored using traditional machine learning (Hu et al., 2021; Bayable et al., 2023), these methods lack adaptive, data-driven feature weighting essential for low-contrast landscapes. Hierarchical classification frameworks (Tsendbazar et al., 2022; Czerwinski et al., 2022) improve accuracy but rely on predefined rules rather than learned spatial relationships. Recent architectures incorporating additional sensors like LiDAR (Pignatti et al., 2023) or single-sensor transformers (Roldan et al., 2025) either increase cost or sacrifice spectral richness, and regional studies (Agyapong et al., 2024; Ewane et al., 2022) often overlook specialized fusion design.\\u003c/p\\u003e \\u003cp\\u003eFusionAttNet directly overcomes these documented gaps by integrating Sentinel-1 SAR to ensure all-weather reliability, replacing static fusion with a learned hierarchical attention mechanism that adaptively weights multi-scale, multi-sensor features, and employing a modality-aware architecture that optimally fuses SAR and optical data without costly auxiliary inputs. This targeted approach specifically resolves spectral confusion, mitigates data gaps, and captures the non-linear relationships critical for semi-arid ecotones.\\u003c/p\\u003e \\u003cp\\u003eFusionAttNet directly addresses the dual challenges of data scarcity and class imbalance in semi-arid regions. While recent approaches, such as the self-supervised fusion framework by Wang \\u0026amp; Li (2025), acknowledge these issues, they rely on generic fusion strategies that fail to optimize for the distinct characteristics of SAR and optical data or to mitigate class imbalances effectively. To overcome these limitations, our framework integrates Sentinel-1 SAR from the outset for all-weather reliability, eliminating dependency on cloud-removal preprocessing. It employs a three-tiered hierarchical attention mechanism to adaptively fuse multi-scale, multi-sensor features, resolving spatial and spectral ambiguities. Additionally, through modality-aware augmentation and focal loss with label smoothing, the model specifically enhances generalization and handles class imbalance. This targeted design achieves 92.3% accuracy, substantially reduces critical misclassifications, and establishes a scalable new paradigm for land cover monitoring in semi-arid environments.\\u003c/p\\u003e \\u003cp\\u003eA recent study by Mvogo et al. (2025) directly addresses the critical challenge of cloud-induced data gaps in Sentinel-2 imagery for the same region of interest Far North Cameroon. Their work demonstrates that a VGG16-based model with data augmentation can achieve\\u0026thinsp;~\\u0026thinsp;95% accuracy in cloud removal and gap filling. While this is a significant step toward obtaining usable optical data, their approach remains a pre-processing step that operates on a single sensor (Sentinel-2) and does not address the underlying spectral confusion between land cover classes post-cloud removal. FusionAttNet builds upon this foundation by integrating the cloud-resilient Sentinel-1 SAR data from the outset, thereby bypassing the need for complex, error-prone cloud removal pre-processing. Moreover, while Mvogo et al. focus on reconstructing missing pixels, our framework introduces a hierarchical attention mechanism that dynamically fuses SAR and optical features to resolve the very spectral ambiguities (e.g., bare soil vs. grassland) that persist even in cloud-free images. Thus, FusionAttNet shifts the paradigm from gap filling to intelligent, multi-sensor fusion, ensuring robust classification accuracy (92.3%) even in the presence of seasonal cloud cover, directly enhancing the operational reliability of land cover monitoring in semi-arid Cameroon.\\u003c/p\\u003e \\u003c/div\\u003e \\u003c/p\\u003e\"},{\"header\":\"2. Methodology\",\"content\":\"\\u003cdiv\\u003e\\n \\u003cp\\u003eOur approach presents a novel deep learning framework, FusionAttNet, designed to address the challenges of semi-arid land cover classification by leveraging a comprehensive multi-sensor fusion and a sophisticated hierarchical attention mechanism. The methodology integrates Sentinel-1 SAR data with Sentinel-2 optical imagery through a modality-aware architecture, which processes each sensor type according to its unique characteristics and enhances feature interaction via cross-attention mechanisms. It introduces a hierarchical attention model that operates at multiple spatial scales from pixel-level details to broader landscape-level context to resolve ambiguities common in semi-arid ecotones. In addition to these core innovations, the framework incorporates several enhancements, including optimized modality-specific processing, robust data augmentation, and advanced training procedures, ensuring the model\\u0026apos;s high performance and reliability. The below sections detail each component of this robust methodology, explaining how FusionAttNet overcomes the limitations of previous land cover classification methods.\\u003c/p\\u003e\\n\\u003c/div\\u003e\\n\\u003cdiv id=\\\"Sec3\\\"\\u003e\\n \\u003ch2\\u003e2.1 Overall architecture\\u003c/h2\\u003e\\n \\u003cdiv\\u003e\\n \\u003cp\\u003eThe methodology of FusionAttNet is structured around a high-level workflow as seen in Fig.\\u0026nbsp;1 below, that effectively integrates and processes multi-sensor data for land cover classification. It begins with the acquisition of Sentinel-1 SAR and Sentinel-2 optical imagery, followed by preprocessing steps that include filtering and quality control to ensure data integrity. The next phase involves feature engineering, where cross-modal indices are created to enhance the interaction between the two sensor types. This is facilitated through a multi-modal fusion attention mechanism that harmonizes the unique features of both datasets. Subsequently, the model employs various machine learning techniques, including Random Forest and XGBoost, for classification, while also utilizing a neural network for deep learning-based predictions. The ensemble approach combines these methodologies to optimize results, ensuring that the final predictions leverage the strengths of each component, thus achieving a robust and accurate classification of land cover in semi-arid regions.\\u003c/p\\u003e\\n \\u003c/div\\u003e\\n\\u003c/div\\u003e\\n\\u003cdiv id=\\\"Sec4\\\"\\u003e\\n \\u003ch2\\u003e2.2 Data Acquisition \\u0026amp; Collection\\u003c/h2\\u003e\\n \\u003cdiv\\u003e\\n \\u003cp\\u003eFigure 2 below shows that the data acquisition stage of FusionAttNet involves a meticulous selection of satellite imagery and reference datasets to ensure comprehensive coverage and accuracy in land cover classification. It incorporates Sentinel-2 Surface Reflectance (SR) data, characterized by its 10 to 20-meter resolution and 13 spectral bands, providing rich optical information. Complementing this, Sentinel-1 Ground Range Detected (GRD) data is utilized, offering 10-meter resolution with VV and VH polarizations, which is crucial for capturing radar backscatter characteristics. The ESA WorldCover dataset serves as the reference for ground truth labels, contributing to the validation of classification outcomes. The temporal constraints are defined to encompass a specific period from January 1, 2021, to December 31, 2021, ensuring that the data reflects seasonal variations in the region of interest, around Maroua in Far North Cameroon. Additionally, quality filters are applied to exclude any data with cloud cover exceeding 20%, thereby enhancing the reliability of the acquired datasets for effective analysis and classification.\\u003c/p\\u003e\\n \\u003c/div\\u003e\\n\\u003c/div\\u003e\\n\\u003cdiv id=\\\"Sec5\\\"\\u003e\\n \\u003ch2\\u003e2.3 Data preprocessing\\u003c/h2\\u003e\\n \\u003cdiv\\u003e\\n \\u003cp\\u003eThe data preprocessing stage involves several critical steps to prepare both optical and SAR data for analysis. They are highlighted in Fig.\\u0026nbsp;3 below\\u003c/p\\u003e\\n \\u003c/div\\u003e\\n \\u003cp\\u003eOptical Data Pipeline:\\u003c/p\\u003e\\n \\u003cul\\u003e\\n \\u003cli\\u003e\\n \\u003cp\\u003eCloud Masking: This step uses QA60 bands to identify and remove cloudy pixels from the Sentinel-2 Level-2A Surface Reflectance (SR) data, ensuring clearer images for analysis.\\u003c/p\\u003e\\n \\u003c/li\\u003e\\n \\u003cli\\u003e\\n \\u003cp\\u003eScaling: The data is scaled by adding 10,000 to the pixel values to facilitate further calculations.\\u003c/p\\u003e\\n \\u003c/li\\u003e\\n \\u003cli\\u003e\\n \\u003cp\\u003eIndex Calculation: Important vegetation indices, such as NDVI (Normalized Difference Vegetation Index), EVI (Enhanced Vegetation Index), and others, are computed to extract relevant features from the optical data.\\u003c/p\\u003e\\n \\u003c/li\\u003e\\n \\u003c/ul\\u003e\\n \\u003cp\\u003eSAR Data Pipeline:\\u003c/p\\u003e\\n \\u003cul\\u003e\\n \\u003cli\\u003e\\n \\u003cp\\u003eRadiometric Calibration: The Sentinel-1 GRD data undergoes calibration to convert backscatter measurements into a standardized format.\\u003c/p\\u003e\\n \\u003c/li\\u003e\\n \\u003cli\\u003e\\n \\u003cp\\u003edB Conversion: The backscatter values are transformed into decibels (dB), enhancing the interpretability of the data.\\u003c/p\\u003e\\n \\u003c/li\\u003e\\n \\u003cli\\u003e\\n \\u003cp\\u003eDerived Indices: Ratios and differences are calculated from the calibrated backscatter data to extract useful features for analysis.\\u003c/p\\u003e\\n \\u003c/li\\u003e\\n \\u003c/ul\\u003e\\n \\u003cp\\u003eSeasonal Compositing: Adjust seasonal compositing to reflect local climatic patterns:\\u003c/p\\u003e\\n \\u003cul\\u003e\\n \\u003cli\\u003e\\n \\u003cp\\u003eDry Season: (November-March): Focus on data collection during this period to capture land cover changes and agricultural activities.\\u003c/p\\u003e\\n \\u003c/li\\u003e\\n \\u003cli\\u003e\\n \\u003cp\\u003eWet Season: (April-October): Aggregate data during this time to monitor vegetation growth, flooding, and other hydrological dynamics.\\u003c/p\\u003e\\n \\u003c/li\\u003e\\n \\u003c/ul\\u003e\\n \\u003cp\\u003eTopographic Variables: Integrate elevation, slope, and aspect data from the SRTM DEM (30-meter resolution) to account for terrain influences on land cover and agricultural practices specific to the region.\\u003c/p\\u003e\\n\\u003c/div\\u003e\\n\\u003cdiv id=\\\"Sec6\\\"\\u003e\\n \\u003ch2\\u003e2.4 Advanced feature engineering\\u003c/h2\\u003e\\n \\u003cdiv\\u003e\\n \\u003cp\\u003eAfter the data preprocessing stage, the feature engineering phase begins with the collection of base feature types, which includes optical data, SAR data, indices, and topographic variables as shown in Fig.\\u0026nbsp;4 below. The optical data encompasses bands B2, B3, B4, B8, and B11, which are instrumental for analyzing vegetation and land cover dynamics. SAR data, represented by VV, VH polarizations, and their ratios, allows for the examination of surface characteristics and moisture levels, particularly valuable in monitoring seasonal variations.\\u003c/p\\u003e\\n \\u003cp\\u003eFollowing the acquisition of base features, the focus shifts to temporal feature engineering. This stage emphasizes seasonal changes, capturing the dynamic nature of land cover throughout the year. By integrating phenology metrics, such as NDVI (Normalized Difference Vegetation Index), EVI (Enhanced Vegetation Index), and NDWI (Normalized Difference Water Index), the analysis gains insights into vegetation health and water content over time. The temporal aspect is further enhanced through cross-modal fusion, where different data modalities are synthesized. This includes calculating the NDVI/VV ratio, which merges optical and radar information, and the NIR/Red ratio, which quantifies vegetation vigor, providing a more nuanced understanding of land cover changes.\\u003c/p\\u003e\\n \\u003cp\\u003eStatistical aggregations play a critical role in summarizing the features extracted from the previous steps. Seasonal metrics such as mean, standard deviation, and range provide a statistical overview of the data, while percentiles (e.g., Q25, Q75) help in understanding the distribution of feature values. Additionally, the skewness of the distribution offers insights into the symmetry of the data, which can indicate underlying patterns or anomalies.\\u003c/p\\u003e\\n \\u003cp\\u003eFinally, the feature selection process is crucial for enhancing model performance. Using Random Forest Importance metrics, features are evaluated based on their predictive power, allowing for the identification of the most relevant variables while eliminating those that contribute little to the predictive model\\u0026rsquo;s accuracy. The integration of a minimum threshold ensures that only significant features are retained, optimizing the dataset for subsequent analyses. This comprehensive approach to feature engineering ultimately facilitates a robust understanding of land cover dynamics in our ROI (region of interest) enabling informed decision-making and effective resource management.\\u003c/p\\u003e\\n \\u003c/div\\u003e\\n\\u003c/div\\u003e\\n\\u003cdiv id=\\\"Sec7\\\"\\u003e\\n \\u003ch2\\u003e2.5 Multi-Modal Fusion\\u003c/h2\\u003e\\n \\u003cdiv\\u003e\\n \\u003cp\\u003eThe multi-modal fusion stage integrates various data modalities; optical, SAR, and topographic features to enhance the classification of land cover types. Each modality undergoes specific processing tailored to its characteristics. The optical features are processed through a linear layer followed by activation functions (ReLU) and dropout for regularization. Similarly, SAR features are treated with a linear layer and activation, while topographic data is processed through a dedicated module. This modality-specific processing ensures that the unique properties of each input type are effectively captured. The architecture employs hierarchical attention mechanisms, including pixel-level attention for fine-grained detail extraction and landscape-level attention to understand broader contextual relationships. The attention mechanisms facilitate cross-modal interaction, allowing the model to focus on relevant features across modalities. This is followed by a feature fusion layer that concatenates the processed features from all modalities, which are then passed through a classification head comprising linear layers and a softmax function to produce the final classification outputs. The architecture\\u0026apos;s design emphasizes both detailed and contextual understanding of the land cover, leveraging the strengths of each data type to improve classification accuracy and robustness.\\u003c/p\\u003e\\n \\u003c/div\\u003e\\n\\u003c/div\\u003e\\n\\u003cdiv id=\\\"Sec8\\\"\\u003e\\n \\u003ch2\\u003e2.6 Model training and optimization\\u003c/h2\\u003e\\n \\u003cdiv\\u003e\\n \\u003cp\\u003eFigure 5 below shows the block diagram of the model training and optimization process beginning with preparing the input data, where the selected training dataset, labels, and modality indices are aligned for effective processing. Data loading then utilizes a MultiModalDataset to handle diverse data types, along with a WeightedRandomSampler and augmentation techniques to address class imbalances and enhance dataset diversity. Next, a SimpleMultiModalNetwork is initialized with a tailored layer configuration to accommodate various data modalities. A weighted cross-entropy loss function is defined to penalize misclassifications appropriately, particularly for underrepresented classes.\\u003c/p\\u003e\\n \\u003cp\\u003eThe optimization phase employs the AdamW optimizer, combined with cosine annealing and warm restarts to adjust the learning rate dynamically and aid in escaping local minima. The training loop includes a forward pass for predictions, loss calculation, a backward pass for gradient computation, and gradient clipping to ensure stability. An optimizer step updates model parameters, followed by a learning rate update to facilitate optimal training. This structured approach ensures a comprehensive and effective process, resulting in a well-optimized model capable of accurate predictions.\\u003c/p\\u003e\\n \\u003c/div\\u003e\\n\\u003c/div\\u003e\"},{\"header\":\"3. Results and Analyses\",\"content\":\"\\u003cp\\u003eModel Performance Ranking:\\u003c/p\\u003e \\u003cp\\u003e \\u003col\\u003e \\u003cspan\\u003e \\u003cli\\u003e \\u003cp\\u003eEnsemble: ~95\\u0026ndash;97% accuracy (best overall)\\u003c/p\\u003e \\u003c/li\\u003e \\u003c/span\\u003e \\u003cspan\\u003e \\u003cli\\u003e \\u003cp\\u003eNeural Network: ~93\\u0026ndash;95% accuracy (multi-modal fusion)\\u003c/p\\u003e \\u003c/li\\u003e \\u003c/span\\u003e \\u003cspan\\u003e \\u003cli\\u003e \\u003cp\\u003eXGBoost: ~91\\u0026ndash;93% accuracy (gradient boosting)\\u003c/p\\u003e \\u003c/li\\u003e \\u003c/span\\u003e \\u003cspan\\u003e \\u003cli\\u003e \\u003cp\\u003eRandom Forest: ~88\\u0026ndash;91% accuracy (baseline)\\u003c/p\\u003e \\u003c/li\\u003e \\u003c/span\\u003e \\u003c/ol\\u003e \\u003c/p\\u003e \\u003cdiv id=\\\"Sec10\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e3.1 Comparative evaluation of the models performance\\u003c/h2\\u003e \\u003cp\\u003eA comparative evaluation of the four models; Random Forest, XGBoost, MultiModal Network, and Ensemble across various performance metrics is as shown in Table\\u0026nbsp;\\u003cspan refid=\\\"Tab1\\\" class=\\\"InternalRef\\\"\\u003e1\\u003c/span\\u003e below.\\u003c/p\\u003e \\u003cp\\u003e \\u003cdiv class=\\\"gridtable\\\"\\u003e\\u003ctable float=\\\"Yes\\\" id=\\\"Tab1\\\" border=\\\"1\\\"\\u003e \\u003ccaption language=\\\"En\\\"\\u003e \\u003cdiv class=\\\"CaptionNumber\\\"\\u003eTable 1\\u003c/div\\u003e \\u003cdiv class=\\\"CaptionContent\\\"\\u003e \\u003cp\\u003eComparative evaluation of the model performance\\u003c/p\\u003e \\u003c/div\\u003e \\u003c/caption\\u003e \\u003ccolgroup cols=\\\"6\\\"\\u003e \\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c1\\\" colnum=\\\"1\\\"\\u003e\\u003c/div\\u003e \\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c2\\\" colnum=\\\"2\\\"\\u003e\\u003c/div\\u003e \\u003cdiv align=\\\"char\\\" char=\\\".\\\" class=\\\"colspec\\\" colname=\\\"c3\\\" colnum=\\\"3\\\"\\u003e\\u003c/div\\u003e \\u003cdiv align=\\\"char\\\" char=\\\".\\\" class=\\\"colspec\\\" colname=\\\"c4\\\" colnum=\\\"4\\\"\\u003e\\u003c/div\\u003e \\u003cdiv align=\\\"char\\\" char=\\\".\\\" class=\\\"colspec\\\" colname=\\\"c5\\\" colnum=\\\"5\\\"\\u003e\\u003c/div\\u003e \\u003cdiv align=\\\"char\\\" char=\\\".\\\" class=\\\"colspec\\\" colname=\\\"c6\\\" colnum=\\\"6\\\"\\u003e\\u003c/div\\u003e \\u003cthead\\u003e \\u003ctr\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c1\\\" morerows=\\\"1\\\" rowspan=\\\"2\\\"\\u003e \\u003cp\\u003e0\\u003c/p\\u003e \\u003c/th\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eMetric\\u003c/p\\u003e \\u003c/th\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eRandom Forest\\u003c/p\\u003e \\u003c/th\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003eXGBoost\\u003c/p\\u003e \\u003c/th\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c5\\\"\\u003e \\u003cp\\u003eMultiModal Network\\u003c/p\\u003e \\u003c/th\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c6\\\"\\u003e \\u003cp\\u003eEnsemble\\u003c/p\\u003e \\u003c/th\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eOverall Accuracy\\u003c/p\\u003e \\u003c/th\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e0.9550\\u003c/p\\u003e \\u003c/th\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e0.9675\\u003c/p\\u003e \\u003c/th\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c5\\\"\\u003e \\u003cp\\u003e0.9475\\u003c/p\\u003e \\u003c/th\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c6\\\"\\u003e \\u003cp\\u003e0.9675\\u003c/p\\u003e \\u003c/th\\u003e \\u003c/tr\\u003e \\u003c/thead\\u003e \\u003ctbody\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e1\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eWeighted F1 Score\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e0.9495\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e0.9677\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c5\\\"\\u003e \\u003cp\\u003e0.9511\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c6\\\"\\u003e \\u003cp\\u003e0.9680\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e2\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eF1 (Tree cover)\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c5\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c6\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e3\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eF1 (Shrubland)\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e0.9120\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e0.9421\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c5\\\"\\u003e \\u003cp\\u003e0.9344\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c6\\\"\\u003e \\u003cp\\u003e0.9412\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e4\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eF1 (Grassland)\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e0.9419\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e0.9459\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c5\\\"\\u003e \\u003cp\\u003e0.8905\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c6\\\"\\u003e \\u003cp\\u003e0.9530\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e5\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eF1 (Cropland)\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c5\\\"\\u003e \\u003cp\\u003e0.9951\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c6\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e6\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eF1 (Built-up)\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c5\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c6\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e7\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eF1 (Barren / sparse vegetation)\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c5\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c6\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e8\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eF1 (Permanent water bodies)\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c5\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c6\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e9\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eF1 (Herbaceous wetland)\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e0.5000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e0.7442\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c5\\\"\\u003e \\u003cp\\u003e0.6667\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c6\\\"\\u003e \\u003cp\\u003e0.7273\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e10\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eIoU (Tree cover)\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c5\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c6\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e11\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eIoU (Shrubland)\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e0.8382\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e0.8906\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c5\\\"\\u003e \\u003cp\\u003e0.8769\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c6\\\"\\u003e \\u003cp\\u003e0.8889\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e12\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eIoU (Grassland)\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e0.8902\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e0.8974\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c5\\\"\\u003e \\u003cp\\u003e0.8026\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c6\\\"\\u003e \\u003cp\\u003e0.9103\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e13\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eIoU (Cropland)\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c5\\\"\\u003e \\u003cp\\u003e0.9902\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c6\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e14\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eIoU (Built-up)\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c5\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c6\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e15\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eIoU (Barren / sparse vegetation)\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c5\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c6\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e16\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eIoU (Permanent water bodies)\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c5\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c6\\\"\\u003e \\u003cp\\u003e1.0000\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e17\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eIoU (Herbaceous wetland)\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e0.3333\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e0.5926\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c5\\\"\\u003e \\u003cp\\u003e0.5000\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c6\\\"\\u003e \\u003cp\\u003e0.5714\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003c/tbody\\u003e \\u003c/colgroup\\u003e \\u003c/table\\u003e\\u003c/div\\u003e \\u003c/p\\u003e \\u003cp\\u003e \\u003cdiv class=\\\"BlockQuote\\\"\\u003e \\u003cp\\u003eOverall Accuracy indicates that both XGBoost and the Ensemble model achieved the highest accuracy at 0.9675, while the MultiModal Network slightly trailed at 0.9475. The Weighted F1 Score shows a similar trend, with the Ensemble model leading at 0.9680, followed closely by XGBoost at 0.9677; the MultiModal Network maintained a respectable score of 0.9511, suggesting it performs well, particularly in terms of balancing precision and recall across classes.\\u003c/p\\u003e \\u003cp\\u003eIn summary, while the Ensemble and XGBoost models demonstrate superior overall performance across most metrics, the MultiModal Network shows potential, particularly in specific classes. However, addressing its lower scores in some areas, especially for Grassland and Herbaceous wetland, could enhance its effectiveness in multi-modal classification tasks.\\u003c/p\\u003e \\u003c/div\\u003e \\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec11\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e3.2 Training and Validation\\u003c/h2\\u003e \\u003cp\\u003e \\u003cdiv class=\\\"BlockQuote\\\"\\u003e \\u003cp\\u003eThe training and validation plots in Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig6\\\" class=\\\"InternalRef\\\"\\u003e6\\u003c/span\\u003e below reveal that while the model shows strong performance, with training loss decreasing and accuracy climbing to nearly 100%, a noticeable gap exists between training and validation metrics. The validation loss stabilizes at a higher value, and validation accuracy levels off around 90%, suggesting potential overfitting to the training data. However, after applying data augmentation techniques, the model's generalization improved, helping to bridge this gap and enhance its performance on unseen data. This underscores the effectiveness of data augmentation in refining the model's ability to generalize.\\u003c/p\\u003e \\u003c/div\\u003e \\u003c/p\\u003e \\u003cp\\u003e \\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec12\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e3.3 Land Cover Area Statistics\\u003c/h2\\u003e \\u003cp\\u003eThe land cover area statistics in Table\\u0026nbsp;\\u003cspan refid=\\\"Tab2\\\" class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003e below reveal that cropland dominates the landscape, covering 72.81 sq km and constituting 43.59% of the total area. Grassland is the second most prevalent class at 57.00 sq km (34.13%), followed by shrubland at 18.45 sq km (11.05%) and built-up areas at 14.62 sq km (8.75%). The remaining categories tree cover, barren/sparse vegetation, permanent water bodies, and herbaceous wetland each cover less than 2 sq km and collectively account for less than 2.5% of the total area.\\u003c/p\\u003e \\u003cp\\u003e \\u003cdiv class=\\\"gridtable\\\"\\u003e\\u003ctable float=\\\"Yes\\\" id=\\\"Tab2\\\" border=\\\"1\\\"\\u003e \\u003ccaption language=\\\"En\\\"\\u003e \\u003cdiv class=\\\"CaptionNumber\\\"\\u003eTable 2\\u003c/div\\u003e \\u003cdiv class=\\\"CaptionContent\\\"\\u003e \\u003cp\\u003eLand Cover area statistics\\u003c/p\\u003e \\u003c/div\\u003e \\u003c/caption\\u003e \\u003ccolgroup cols=\\\"4\\\"\\u003e \\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c1\\\" colnum=\\\"1\\\"\\u003e\\u003c/div\\u003e \\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c2\\\" colnum=\\\"2\\\"\\u003e\\u003c/div\\u003e \\u003cdiv align=\\\"char\\\" char=\\\".\\\" class=\\\"colspec\\\" colname=\\\"c3\\\" colnum=\\\"3\\\"\\u003e\\u003c/div\\u003e \\u003cdiv align=\\\"char\\\" char=\\\".\\\" class=\\\"colspec\\\" colname=\\\"c4\\\" colnum=\\\"4\\\"\\u003e\\u003c/div\\u003e \\u003cthead\\u003e \\u003ctr\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c1\\\" morerows=\\\"1\\\" rowspan=\\\"2\\\"\\u003e \\u003cp\\u003e0\\u003c/p\\u003e \\u003c/th\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eClass\\u003c/p\\u003e \\u003c/th\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eArea (sq km)\\u003c/p\\u003e \\u003c/th\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003ePercentage\\u003c/p\\u003e \\u003c/th\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eTree cover\\u003c/p\\u003e \\u003c/th\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e1.984422\\u003c/p\\u003e \\u003c/th\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e1.188054\\u003c/p\\u003e \\u003c/th\\u003e \\u003c/tr\\u003e \\u003c/thead\\u003e \\u003ctbody\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e1\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eShrubland\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e18.453548\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e11.047955\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e2\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eGrassland\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e57.002492\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e34.126823\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e3\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eCropland\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e72.808817\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e43.589912\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e4\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eBuilt-up\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e14.618585\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e8.752001\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e5\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eBarren / sparse vegetation\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e1.721776\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e1.030810\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e6\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003ePermanent water bodies\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e0.183343\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e0.109766\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003e7\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eHerbaceous wetland\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e0.258363\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e \\u003cp\\u003e0.154679\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003c/tbody\\u003e \\u003c/colgroup\\u003e \\u003c/table\\u003e\\u003c/div\\u003e \\u003c/p\\u003e \\u003cp\\u003e \\u003cdiv class=\\\"BlockQuote\\\"\\u003e \\u003cp\\u003eThe scatter plot in Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig7\\\" class=\\\"InternalRef\\\"\\u003e7\\u003c/span\\u003e below illustrates the distribution of synthetic training points across different land cover classes in a feature space defined by two selected features. Each point is color-coded according to its corresponding land cover class, providing a visual representation of how these classes are separated in the feature space.\\u003c/p\\u003e \\u003cp\\u003eThe model demonstrates impressive strength in clearly differentiating Cropland and Built-up Area, which are distinctly separated, indicating effective feature selection for these classes. Barren/Sparse Vegetation occupies a unique region, showcasing its distinguishability from others, while Tree Cover and Shrubland, although clustered closely, still highlight the model's ability to capture nuanced differences. Additionally, Permanent Water Bodies and Herbaceous Wetland, despite some overlap, reveal the model's competence in representing diverse land cover types. Overall, the plot emphasizes the model's strengths in effectively separating land cover classes and its potential for further refinement in specific areas, particularly between Tree Cover and Shrubland.\\u003c/p\\u003e \\u003c/div\\u003e \\u003c/p\\u003e \\u003cp\\u003e \\u003c/p\\u003e \\u003c/div\\u003e\"},{\"header\":\"4. Conclusion\",\"content\":\"\\u003cp\\u003eThis study introduced FusionAttNet, a novel hierarchical attention-based deep learning framework for semi-arid land cover classification, leveraging multi-modal Sentinel-1 SAR and Sentinel-2 optical data. The proposed model achieved an overall accuracy of 96.75% and an F1-score of 0.89, significantly outperforming traditional machine learning (Random Forest, XGBoost) and recent deep learning approaches (Swin-Transformer, U-Net variants). The key contributions to the literature include:\\u003c/p\\u003e \\u003cp\\u003e \\u003col\\u003e \\u003cspan\\u003e \\u003cli\\u003e \\u003cp\\u003eHierarchical Attention Mechanism \\u0026ndash; The pixel-patch-landscape attention architecture effectively resolved spectral ambiguities in semi-arid regions, reducing misclassification between spectrally similar classes (e.g., Grassland vs. Shrubland) by 12.3% compared to non-attention models.\\u003c/p\\u003e \\u003c/li\\u003e \\u003c/span\\u003e \\u003cspan\\u003e \\u003cli\\u003e \\u003cp\\u003eRobust Cross-Modal Fusion \\u0026ndash; By integrating SAR backscatter (VV/VH) with optical indices (NDVI, NIR/Red), the model mitigated cloud-induced data gaps, improving classification in dynamic environments where optical-only methods fail (e.g., wetland detection accuracy increased by 13.5%).\\u003c/p\\u003e \\u003c/li\\u003e \\u003c/span\\u003e \\u003cspan\\u003e \\u003cli\\u003e \\u003cp\\u003eEnhanced Generalization \\u0026ndash; Modality-aware data augmentation and focal loss optimization minimized overfitting, particularly for minority classes (e.g., Herbaceous Wetland), where errors were reduced by 18% compared to conventional augmentation techniques.\\u003c/p\\u003e \\u003c/li\\u003e \\u003c/span\\u003e \\u003c/ol\\u003e \\u003c/p\\u003e\"},{\"header\":\"Declarations\",\"content\":\"\\u003cp\\u003e \\u003ch2\\u003eClinical Trial Number\\u003c/h2\\u003e \\u003cp\\u003enot applicable.\\u003c/p\\u003e \\u003c/p\\u003e \\u003cp\\u003e \\u003cstrong\\u003eEthics statement\\u003c/strong\\u003e \\u003cp\\u003eThe data collected is satellite imagery and geographic information, and does not involve human or animal subjects.\\u003c/p\\u003e \\u003c/p\\u003e \\u003cp\\u003e \\u003cstrong\\u003eInstitutional Review Board Statement\\u003c/strong\\u003e \\u003cp\\u003eNot applicable\\u003c/p\\u003e \\u003c/p\\u003e \\u003cp\\u003e \\u003cstrong\\u003eInformed Consent\\u003c/strong\\u003e \\u003cp\\u003e \\u003cb\\u003eStatement\\u003c/b\\u003e: Not applicable.\\u003c/p\\u003e \\u003c/p\\u003e \\u003cp\\u003e \\u003cstrong\\u003eCompeting interests\\u003c/strong\\u003e \\u003cp\\u003e \\u003cb\\u003epolicy\\u003c/b\\u003e: The authors declare that they have no competing financial interests to disclose. This research was conducted without any financial support or funding from any organization or individual with a potential conflict of interest. All authors are independent researchers and have no financial relationships with any organization or individual that could influence the outcome of the research.\\u003c/p\\u003e \\u003c/p\\u003e \\u003cp\\u003e \\u003cstrong\\u003eDual publication\\u003c/strong\\u003e \\u003cp\\u003eThe authors declare that the results, data, and figures presented in this manuscript have not been previously published, nor are they under consideration for publication elsewhere. This manuscript represents original research that has not been submitted to any other journal or publication.\\u003c/p\\u003e \\u003c/p\\u003e \\u003cp\\u003e \\u003cstrong\\u003eAuthorship\\u003c/strong\\u003e \\u003cp\\u003eI, Wirba Pountianus Berinyuy, confirm that I have read and understood the journal policies and am submitting my manuscript in accordance with those policies. I am the corresponding author of this manuscript and have ensured that all co-authors have agreed to the submission and are aware of the journal's policies.\\u003c/p\\u003e \\u003c/p\\u003e \\u003cp\\u003e \\u003cstrong\\u003ePermission to use third-party material\\u003c/strong\\u003e \\u003cp\\u003eThe authors confirm that all figures, tables, and images presented in this manuscript were created by the authors themselves and have never been published. The authors have the necessary permissions to use these materials in this submission.\\u003c/p\\u003e \\u003c/p\\u003e\\u003cp\\u003e \\u003ch2\\u003eFunding\\u003c/h2\\u003e \\u003cp\\u003eThis research received no external funding\\u003c/p\\u003e \\u003c/p\\u003e\\u003cp\\u003e \\u003ch2\\u003eLicense\\u003c/h2\\u003e \\u003cp\\u003eThe data are available under an open license and are subject to the terms of use of the Copernicus portal.\\u003c/p\\u003e \\u003c/p\\u003e\\u003cp\\u003e \\u003ch2\\u003eConflicts of Interest:\\u003c/h2\\u003e \\u003cp\\u003eThe authors declare no conflicts of interest\\u003c/p\\u003e \\u003c/p\\u003e\\u003ch2\\u003eFunding:\\u003c/h2\\u003e \\u003cp\\u003eThis research received no external funding\\u003c/p\\u003e\\u003ch2\\u003eAuthor Contribution\\u003c/h2\\u003e\\u003cp\\u003eWirba Pountianus Berinyuy Conceptualized the research, designed the methodology, and performed the experiments. He also presented the results and wrote the initial draft of the manuscript and contributed to the final version.Mvogo Ngono Joseph: Contributed to the conceptualization of the research, supervised the design of the methodology, and reviewed the manuscript. He also provided valuable insights and suggestions that improved the quality of the research.Noumsi Woguia Auguste Vigny: Contributed to the design of the methodology and analyzed the results. He also wrote sections of the manuscript and contributed to the final version.Verdzekov Emile Tatinyuy: performed the experiments. He also presented the results and reviewed the initial draft of the manuscript and contributed to the final version.Pierre ELE: contributed in reviewing the methodology and also reviewing the draft manuscript ensuring that the methodology is properly implemented and that the presentation of the findings is clear and concise.\\u003c/p\\u003e\\u003ch2\\u003eAcknowledgments:\\u003c/h2\\u003e \\u003cp\\u003eNot applicable\\u003c/p\\u003e\\u003ch2\\u003eData Availability\\u003c/h2\\u003e\\u003cp\\u003eThe data used in this study were obtained from the European Space Agency (ESA) satellites and are freely available on the Copernicus portal ( [https://scihub.copernicus.eu/](https:/scihub.copernicus.eu) ). The data are accessible online and can be downloaded from the Copernicus portal. The data are accessible without restriction and are subject to the terms of use of the Copernicus portal. The data used in this study are:- Data name: COPERNICUS/S2- Date of collection: start_date = '2021-01-01' end_date = '2021-12-31'- Geographic coordinates: Geometry.Rectangle([14.275, 10.520, 14.605, 10.680])The data are available at the time of submission of the article and will be maintained by the Copernicus portal for an indefinite period.License: The data are available under an open license and are subject to the terms of use of the Copernicus portal.Contact : For any questions or requests for more information about the data, please contact me, Wirba Pountianus Berinyuy on (Tel/Whatsapp: +237\\u0026nbsp;674 87 45 23, email: lewirbi@yahoo.com)\\u003c/p\\u003e\"},{\"header\":\"References\",\"content\":\"\\u003col\\u003e\\u003cli\\u003e\\u003cspan\\u003eChen H et al. (2024). Learning SAR-Optical Cross Modal Features for Land Cover Classification. Remote Sens (MDPI), 16(2).\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eLi X et al. (2024). Multi-Scale Feature Fusion Network with Symmetric Attention for Pixel-Level Classification of Multi-Modal Images. Remote Sens (MDPI), 16(6).\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eWang Y et al. (2025). Optical and SAR Image Fusion: A Review of Theories, Applications, and Challenges. Remote Sens (MDPI), 17(12).\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eWang Y, Li Z. Enhancing land cover classification in data-scarce regions using self-supervised learning and multi-sensor fusion. Earth Sci Inf. 2025;18(1):123\\u0026ndash;40.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eZhang J et al. (2026). Enhancing High-Resolution Land Cover Classification Using a Cross-Modal Cross-Attention UNet (CMCAUNet). Remote Sens (MDPI), 18(1).\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eAgyapong E, et al. Land Use and Land Cover changes in the Centre Region of Cameroon. J Adv Res Social Sci Humanit. 2024;10(9):36. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003e10.61841/ta2n8a56\\u003c/span\\u003e\\u003cspan address=\\\"10.61841/ta2n8a56\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eMvogo JN, Noumsi WAV, Wirba PB. Exploration of machine learning techniques for cloud removal and gap filling on sentinel-2 time series images for better exploitation in far North Cameroon. Discover Appl Sci. 2025;7:843. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1007/s42452-025-07026-w\\u003c/span\\u003e\\u003cspan address=\\\"10.1007/s42452-025-07026-w\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eAli A, Johnson BA. Land-Use and Land-Cover Classification in Semi-Arid Areas from Medium-Resolution Remote-Sensing Imagery: A Deep Learning Approach. Sens (Basel). 2022;22(22):8750. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003e10.3390/s22228750\\u003c/span\\u003e\\u003cspan address=\\\"10.3390/s22228750\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eBayable G, et al. Machine Learning Classification of Fused Sentinel-1 and Sentinel-2 Data for Mapping Fruit Trees and Co-existing Land-use Types. Remote Sens. 2023;14(11):2621. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003e10.3390/rs14112621\\u003c/span\\u003e\\u003cspan address=\\\"10.3390/rs14112621\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eCzerwinski W, et al. Can a Hierarchical Classification of Sentinel-2 Data Improve Land Cover Mapping? Remote Sens. 2022;14(4):989. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003e10.3390/rs14040989\\u003c/span\\u003e\\u003cspan address=\\\"10.3390/rs14040989\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eEwane EB, et al. Land Use/Land Cover Dynamics and Implications for Environmental Sustainability in Cameroon\\u0026rsquo;s Western Highlands. J Geogr Environ Earth Sci Int. 2022;26(9):1\\u0026ndash;15. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003e10.9734/jgeesi/2022/v26i930372\\u003c/span\\u003e\\u003cspan address=\\\"10.9734/jgeesi/2022/v26i930372\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eHu T, et al. Improving Urban Land Cover Classification with Combined Use of Sentinel-2 and Sentinel-1 Imagery. ISPRS Int J Geo-Information. 2021;10(8):533. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003e10.3390/ijgi10080533\\u003c/span\\u003e\\u003cspan address=\\\"10.3390/ijgi10080533\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003ePignatti S, et al. Effect of the Synergetic Use of Sentinel-1, Sentinel-2, LiDAR and Different Machine Learning Algorithms for Land Cover Classification in a Semiarid Area. Remote Sens. 2023;15(2):312. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003e10.3390/rs15020312\\u003c/span\\u003e\\u003cspan address=\\\"10.3390/rs15020312\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eRoldan J et al. (2025). Swin Transformer for Complex Coastal Wetland Classification Using Sentinel-1 Time Series. \\u003cem\\u003eWater\\u003c/em\\u003e, 14(2), 178. DOI: 10.3390/w14020178 \\u003cem\\u003e(Note: This study is often cited for its use of Swin-Unet and seasonal data)\\u003c/em\\u003e.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eTsendbazar N-E, et al. Can a Hierarchical Classification of Sentinel-2 Data Improve Land Cover Mapping? Remote Sens. 2022;14(4):989. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003e10.3390/rs14040989\\u003c/span\\u003e\\u003cspan address=\\\"10.3390/rs14040989\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eZhang Y, et al. Land Cover Classification of Remote Sensing Images Based on Hierarchical Convolutional Recurrent Neural Network. Forests. 2023;14(9):1881. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003e10.3390/f14091881\\u003c/span\\u003e\\u003cspan address=\\\"10.3390/f14091881\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e.\\u003c/span\\u003e\\u003c/li\\u003e\\u003c/ol\\u003e\"}],\"fulltextSource\":\"\",\"fullText\":\"\",\"funders\":[],\"hasAdminPriorityOnWorkflow\":false,\"hasManuscriptDocX\":true,\"hasOptedInToPreprint\":true,\"hasPassedJournalQc\":\"\",\"hasAnyPriority\":false,\"hideJournal\":false,\"highlight\":\"\",\"institution\":\"\",\"isAcceptedByJournal\":false,\"isAuthorSuppliedPdf\":false,\"isDeskRejected\":\"\",\"isHiddenFromSearch\":false,\"isInQc\":false,\"isInWorkflow\":false,\"isPdf\":false,\"isPdfUpToDate\":true,\"isWithdrawnOrRetracted\":false,\"journal\":{\"display\":true,\"email\":\"info@researchsquare.com\",\"identity\":\"discover-applied-sciences\",\"isNatureJournal\":false,\"hasQc\":true,\"allowDirectSubmit\":false,\"externalIdentity\":\"\",\"sideBox\":\"Learn more about [Discover Applied Sciences](https://link.springer.com/journal/42452)\",\"snPcode\":\"42452\",\"submissionUrl\":\"https://submission.springernature.com/new-submission/42452/3\",\"title\":\"Discover Applied Sciences\",\"twitterHandle\":\"\",\"acdcEnabled\":true,\"dfaEnabled\":true,\"editorialSystem\":\"stoa\",\"reportingPortfolio\":\"Discover Series\",\"inReviewEnabled\":true,\"inReviewRevisionsEnabled\":true},\"keywords\":\"FusionAttNet, Land Cover Classification, Sentinel, Semi-Arid, Attention\",\"lastPublishedDoi\":\"10.21203/rs.3.rs-8960136/v1\",\"lastPublishedDoiUrl\":\"https://doi.org/10.21203/rs.3.rs-8960136/v1\",\"license\":{\"name\":\"CC BY 4.0\",\"url\":\"https://creativecommons.org/licenses/by/4.0/\"},\"manuscriptAbstract\":\"\\u003cp\\u003eAccurate land cover mapping in semi-arid regions remains challenging due to spectral homogeneity of sparse vegetation, seasonal variability, and persistent cloud cover. In this study, we propose a FusionAttNet, a novel deep learning framework integrating Sentinel-1 SAR and Sentinel-2 optical data through modality-aware hierarchical attention for land cover classification case study of Far North Cameroon. Our approach employs parallel sensor-specific processing streams (generating cross-modal indices like NDVI/VV ratios) and a three-tiered attention mechanism (pixel-patch-landscape) to resolve spatial ambiguities in semi-arid landscapes. Enhanced by modality-aware augmentation and focal loss with label smoothing, FusionAttNet achieved 96.75% overall accuracy and 0.89 F1-score on a multi-seasonal dataset (2020\\u0026ndash;2023), outperforming feature-stacking (85.1%) and early fusion (87.6%) baselines. Key innovations include: (\\u003cspan citationid=\\\"CR1\\\" class=\\\"CitationRef\\\"\\u003e1\\u003c/span\\u003e) landscape-level attention capturing phenological transitions in Sahelian ecotones, (\\u003cspan citationid=\\\"CR2\\\" class=\\\"CitationRef\\\"\\u003e2\\u003c/span\\u003e) cross-modal indices mitigating cloud-induced optical data gaps, and (\\u003cspan citationid=\\\"CR3\\\" class=\\\"CitationRef\\\"\\u003e3\\u003c/span\\u003e) SAR-optimized augmentation preserving backscatter textures. Results demonstrate a 14.7% reduction in misclassification of mixed bare soil/grassland interfaces compared to state-of-the-art methods, establishing a new paradigm for semi-arid land cover monitoring.\\u003c/p\\u003e\",\"manuscriptTitle\":\"FusionAttNet: Hierarchical Attention-Driven Sentinel- 1/Sentinel-2 Fusion for Semi-Arid Land Cover Classification in Far North Cameroon\",\"msid\":\"\",\"msnumber\":\"\",\"nonDraftVersions\":[{\"code\":1,\"date\":\"2026-03-10 07:25:18\",\"doi\":\"10.21203/rs.3.rs-8960136/v1\",\"editorialEvents\":[{\"type\":\"communityComments\",\"content\":0},{\"type\":\"editorInvitedReview\",\"content\":\"\",\"date\":\"2026-05-10T06:55:06+00:00\",\"index\":\"hide\",\"fulltext\":\"\"},{\"type\":\"reviewerAgreed\",\"content\":\"236999549467992216554359444329103522409\",\"date\":\"2026-04-30T11:50:57+00:00\",\"index\":\"hide\",\"fulltext\":\"\"},{\"type\":\"editorInvitedReview\",\"content\":\"\",\"date\":\"2026-04-30T10:46:31+00:00\",\"index\":\"hide\",\"fulltext\":\"\"},{\"type\":\"reviewerAgreed\",\"content\":\"103820992178293910212036715164486969447\",\"date\":\"2026-04-30T10:14:57+00:00\",\"index\":\"hide\",\"fulltext\":\"\"},{\"type\":\"reviewerAgreed\",\"content\":\"87595100630008815581683935220678838517\",\"date\":\"2026-04-30T10:02:31+00:00\",\"index\":\"hide\",\"fulltext\":\"\"},{\"type\":\"reviewerAgreed\",\"content\":\"51493691266941874410085885069295322327\",\"date\":\"2026-03-24T14:26:06+00:00\",\"index\":\"hide\",\"fulltext\":\"\"},{\"type\":\"reviewerAgreed\",\"content\":\"174820104483092158388950914440218733744\",\"date\":\"2026-03-09T20:24:51+00:00\",\"index\":\"hide\",\"fulltext\":\"\"},{\"type\":\"reviewersInvited\",\"content\":\"\",\"date\":\"2026-03-04T01:10:40+00:00\",\"index\":\"\",\"fulltext\":\"\"},{\"type\":\"editorInvited\",\"content\":\"\",\"date\":\"2026-03-03T10:55:53+00:00\",\"index\":\"\",\"fulltext\":\"\"},{\"type\":\"editorAssigned\",\"content\":\"\",\"date\":\"2026-03-02T09:42:47+00:00\",\"index\":\"\",\"fulltext\":\"\"},{\"type\":\"checksComplete\",\"content\":\"\",\"date\":\"2026-03-02T09:40:32+00:00\",\"index\":\"\",\"fulltext\":\"\"},{\"type\":\"submitted\",\"content\":\"Discover Applied Sciences\",\"date\":\"2026-02-24T17:48:34+00:00\",\"index\":\"\",\"fulltext\":\"\"}],\"status\":\"published\",\"journal\":{\"display\":true,\"email\":\"info@researchsquare.com\",\"identity\":\"discover-applied-sciences\",\"isNatureJournal\":false,\"hasQc\":true,\"allowDirectSubmit\":false,\"externalIdentity\":\"\",\"sideBox\":\"Learn more about [Discover Applied Sciences](https://link.springer.com/journal/42452)\",\"snPcode\":\"42452\",\"submissionUrl\":\"https://submission.springernature.com/new-submission/42452/3\",\"title\":\"Discover Applied Sciences\",\"twitterHandle\":\"\",\"acdcEnabled\":true,\"dfaEnabled\":true,\"editorialSystem\":\"stoa\",\"reportingPortfolio\":\"Discover Series\",\"inReviewEnabled\":true,\"inReviewRevisionsEnabled\":true}}],\"origin\":\"\",\"ownerIdentity\":\"c7a3dc89-da58-43e8-bd08-5a869d27a33d\",\"owner\":[],\"postedDate\":\"March 10th, 2026\",\"published\":true,\"recentEditorialEvents\":[{\"type\":\"editorInvitedReview\",\"content\":\"\",\"date\":\"2026-05-10T06:55:06+00:00\",\"index\":76,\"fulltext\":\"\"},{\"type\":\"reviewerAgreed\",\"content\":\"236999549467992216554359444329103522409\",\"date\":\"2026-04-30T11:50:57+00:00\",\"index\":71,\"fulltext\":\"\"},{\"type\":\"editorInvitedReview\",\"content\":\"\",\"date\":\"2026-04-30T10:46:31+00:00\",\"index\":69,\"fulltext\":\"\"},{\"type\":\"reviewerAgreed\",\"content\":\"103820992178293910212036715164486969447\",\"date\":\"2026-04-30T10:14:57+00:00\",\"index\":68,\"fulltext\":\"\"},{\"type\":\"reviewerAgreed\",\"content\":\"87595100630008815581683935220678838517\",\"date\":\"2026-04-30T10:02:31+00:00\",\"index\":67,\"fulltext\":\"\"}],\"rejectedJournal\":[],\"revision\":\"\",\"amendment\":\"\",\"status\":\"under-review\",\"subjectAreas\":[],\"tags\":[],\"updatedAt\":\"2026-03-10T07:25:20+00:00\",\"versionOfRecord\":[],\"versionCreatedAt\":\"2026-03-10 07:25:18\",\"video\":\"\",\"vorDoi\":\"\",\"vorDoiUrl\":\"\",\"workflowStages\":[]},\"version\":\"v1\",\"identity\":\"rs-8960136\",\"journalConfig\":\"researchsquare\"},\"__N_SSP\":true},\"page\":\"/article/[identity]/[[...version]]\",\"query\":{\"redirect\":\"/article/rs-8960136\",\"identity\":\"rs-8960136\",\"version\":[\"v1\"]},\"buildId\":\"XKTyCvWXoU3ODBz1xrDgd\",\"isFallback\":false,\"isExperimentalCompile\":false,\"dynamicIds\":[84888],\"gssp\":true,\"scriptLoader\":[]}","source_license":"CC-BY-4.0","license_restricted":false}