Novel Distribution-Free Eigenspace Framework via Truncated Singular Value Decomposition for High-Resolution Dengue Clusters | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Novel Distribution-Free Eigenspace Framework via Truncated Singular Value Decomposition for High-Resolution Dengue Clusters Muhammad Fayyaz, Alamgir Khalil, Hameed Ali, Zeineb Klai, Sami Ullah, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7484169/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract The development of efficient and reliable algorithm to identify risk spots is vital in real life domain, specifically in epidemiology. This study aims to design a novel eigenspace based framework for detecting multiple spatiotemporal clusters of vector borne diseases, considering dengue data from Khyber Pakhtunkhwa, Pakistan, as a case study. Unlike traditional scan-based methods constrained by strict distributional assumptions and regular cluster shape, our approach effectively detects spatiotemporal clusters of arbitrary shape and under any data distribution. Initiated by refined expected case matrix construction, the proposed method integrates truncated singular value decomposition and control chart to precisely discover multiple simultaneous clusters. In a direct comparison with classical eigenspace techniques such as EigenSpot and MultiEigenSpot, our framework consistently outperforms in identification of clusters across both spatial and temporal dimensions. Further, our approach also utilizes vectorized implementation making it computationally economic against competing methods. To facilitate clear interpretation and rapid public health response, the framework includes interactive visual analytics, such as heatmaps of the relative risk (RR) matrix. Our work offers significant contribution to the practitioners, specifically epidemiologist in making informed decision in real life domains. While motivated by epidemiological applications, the proposed framework can seamlessly be adapted in other scenario such environmental monitoring, crime hotspot mapping, and disaster impact assessment. Biological sciences/Computational biology and bioinformatics Physical sciences/Mathematics and computing Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 Figure 10 1. Introduction Vector Borne Diseases (VBD) such as dengue, malaria, and chikungunya are caused by pathogens transmitted through vectors like mosquitoes and ticks, causes numerous causalities globally [ 1 ]. These diseases remain a persistent threat to public health, particularly in low and middle income countries, where environmental, climatic, and socioeconomic conditions further severe their spread [ 2 ]. In regions such as Khyber Pakhtunkhwa (KP), Pakistan, VBD has a rising trend in frequency, intensity, and geographic range, causes by seasonal variation, urbanization, population density, and deficiencies in healthcare infrastructure [ 3 ]. Effective disease surveillance is essential for identifying outbreaks, tracking spatial transmission dynamics, and implementing timely public health interventions [ 4 ]. To this end, spatiotemporal data capturing disease occurrence over space and time is utilized to detect abnormal patterns of elevated disease risk. Disease clusters are typically defined as subregions or time intervals within the spatiotemporal domain that shows significantly higher observed cases than those of expected cases [ 5 ], [ 6 ]. Identifying such clusters enables public health authorities to prioritize regions for further epidemiological investigation and intervention. Numerous statistical methodologies have been developed to detect disease clusters, such as space-time scan statistics, which are mostly designed to detect regularly shaped clusters [ 7 ], [ 8 ], [ 9 ], [ 10 ]. However, these approaches often fail to accurately identify irregular shaped clusters, especially in geographically complex or terrain influenced regions. This limitation becomes particularly problematic in areas where disease spread follows natural or infrastructural features such as rivers or roads. To address these challenges, researchers have developed more flexible methods for detecting irregular shaped clusters, such as space-time scan statistic [ 11 ], grid based [ 12 ], [ 13 ], and permutation based scan techniques [ 14 ]. However, all these approaches rely on strong parametric assumptions, which limits their effectiveness when applied to nontraditional, noisy, or sparse datasets [ 15 ], [ 16 ]. EigenSpot detects spatiotemporal clusters without distributional or shape assumptions by analyzing matrix structure deviations, but it is limited to a single cluster and therefore misses multiple concurrent outbreaks [ 17 ]. To overcome this limitation, Multi-EigenSpot [ 18 ] was introduced, this method uses the expected case matrix as the baseline rather than the population-at-risk matrix and iteratively replaces detected clusters with expected values to identify multiple clusters. While this algorithm extends the capacity of EigenSpot to detect more than one spatiotemporal cluster, it suffers from several drawbacks, including poor sensitivity to rare disease clusters, high computational cost on large-scale datasets and miss some important cluster in sparce data. Building upon this line of research, the present study introduces a novel enhancement to the eigenspace based framework for detecting multiple, irregularly shaped, and low-incidence spatiotemporal clusters. The key methodological contributions of the proposed approach are as follows: Refined Expected Case Matrix Construction: The method incorporates an improved mechanism for constructing the expected case matrix to enhance sensitivity and detect more clusters. Improved Numerical Stability: Rather than employing Singular Value Decomposition (SVD), the truncated form of SVD is used. Enhanced Computational Efficiency: The algorithm utilizes vectorized operations and parallel processing techniques to significantly reduce computational time, thereby improving scalability to high resolution spatiotemporal datasets. To support interpretation and facilitate real time decision making, the proposed framework incorporates advanced visual analytics, including heatmaps of the relative risk (RR) matrix, which allow for intuitive representation of spatial and temporal variations in disease risk. Such visual tools are dynamic for aiding public health professionals, epidemiologists, and decision makers in quickly detecting and responding to emerging hotspots. The rest of this paper is organized as follows: Section 2 presents the study area, data sources, and the proposed methodology, detailing the novel EigenSpace algorithm based on truncated SVD, Z-control charts, and heatmap visualization. Section 3 discusses the results and findings from the application of the proposed method to vector borne disease data in Khyber Pakhtunkhwa, highlighting detected clusters. Section 4 evaluates the computational efficiency of the proposed method and comparing performance with existing approaches. Section 5 concludes the study with key insights, limitations, and future directions. 2. Methodology 2.1. Materials and Methods Approach The analysis was conducted using MATLAB R2017a, comparing the proposed Novel EigenSpace method with conventional EigenSpot and Multi EigenSpot techniques. Additionally, spatial mapping and cluster visualization were performed using QGIS, a freely available and open-source geographic information system (Version 3.34 ‘Firenze’, https://qgis.org ), an open-source geographic information system. Study Area This study focuses on Khyber Pakhtunkhwa (KP), a northwestern province of Pakistan located between 31.0°N–36.5°N latitude and 69.0°E–74.0°E longitude. The province spans an area of 74,521 km² and had an estimated population of 30.5 million as reported in the 2017 national census (Pakistan Bureau of Statistics, 2017) [ 19 ]. KP shares international borders with Afghanistan to the northwest and provincial borders with Gilgit Baltistan (northeast), Azad Jammu and Kashmir (east), Ex-FATA (west and south), Balochistan and Punjab (south), and the Islamabad Capital Territory (southeast). As the third most populous region in the country, KP plays a critical role in spatiotemporal disease surveillance and vector borne disease control, particularly due to its demographic scale, cross-border connectivity, and ecological diversity. Figure 1 displays the study area map of Khyber Pakhtunkhwa, generated by the authors using QGIS software with administrative boundary data. 2.2. Data Collection The dataset used in this study contains of district level records of dengue cases for the year 2024, collected from the Integrated Disease Surveillance and Response (IDSR) system, accessible through the Public Health Bulletin of Pakistan https://phb.nih.org.pk/integratedisease-surveillance-and-response a freely accessible online resource for research. To support demographic normalization and the estimation of expected case counts, population data from the 2017 national census were incorporated [ 19 ]. The data were curated to capture both spatial distribution and temporal trends of disease outbreaks, forming a solid foundation for spatiotemporal cluster analysis of dengue in the region. 2.3. Computationally Efficiency The algorithms proposed in [ 17 ], [ 18 ] implement the EigenSpace framework by leveraging conventional SVD techniques for dimensionality reduction and matrix factorization within the spatiotemporal data structure. However, standard SVD is computationally expensive, particularly for large and sparse datasets. To address the computational and structural limitations essential in EigenSpot and Multi-EigenSpot, the proposed methods integrate an advanced modified SVDs (Sparse Singular Value Decomposition). Design for efficient processing of high dimensional and sparse matrices, SVDs significantly improves decomposition accuracy and scalability, making it particularly effective for analyzing rare disease datasets with limited cases. The proposed algorithm integrates the following three methodological components: SVDs : Employed to extract the principal left and right singular vectors (LSV and RSV) from the K and E matrices for dimensionality reduction. Z-Control Chart : Employed to find abnormal components in differences vectors. Visualization : A heatmap is used to display the final Relative Risk (RR) matrix, highlighting potential cluster regions through colour intensity variations. 2.4. Proposed Algorithm The proposed algorithm is based on the Eigen Space framework technique to identify multiple abnormal spatiotemporal clusters. It functions on matrix formatted surveillance data and utilizes statistical control measures to isolate significant deviations in disease distribution. The methodology is systematically outlined as follows: 1. Total observed cases matrix is denoted by K and the population at risk matrix is denoted by P. $$\:K=\left[\begin{array}{ccc}{k}_{11}&\:\cdots\:&\:{k}_{1n}\\\:⋮&\:\ddots\:&\:⋮\\\:{k}_{m1}&\:\cdots\:&\:{k}_{mn}\end{array}\right]\:and\:P=\left[\begin{array}{ccc}{p}_{11}&\:\cdots\:&\:{p}_{1n}\\\:⋮&\:\ddots\:&\:⋮\\\:{p}_{m1}&\:\cdots\:&\:{p}_{mn}\end{array}\right]$$ Where \(\:{k}_{11}\) is the total disease in the first region first-time point \(\:{p}_{11}\) is the total population at risk in the first region and first-time point \(\:\:m\) is total spatial dimensions and \(\:\:n\:\) total time points. 2. The expected disease cases matrix \(\:{E}_{ij}\) and relative risks matrix \(\:{R}_{ij}\) are computed using the observed cases \(\:{K}_{ij}\) and population at risk \(\:{P}_{ij}\:\) matrices. Where \(\:ith\) represent spatial unit (region) and \(\:jth\) time point. $$\:E=\left[\begin{array}{ccc}{E}_{11}&\:\cdots\:&\:{E}_{1n}\\\:⋮&\:\ddots\:&\:⋮\\\:{E}_{m1}&\:\cdots\:&\:{E}_{mn}\end{array}\right]\:and\:R=\left[\begin{array}{ccc}{R}_{11}&\:\cdots\:&\:{R}_{1n}\\\:⋮&\:\ddots\:&\:⋮\\\:{R}_{m1}&\:\cdots\:&\:{R}_{mn}\end{array}\right]$$ Once the relative risk matrix \(\:R\) is calculated, it serves as the basis for visualizing potential clusters on a heatmap. 3. The one-rank SVDs is used to find the LSV and RSV for matrices \(\:K\) and \(\:E\) . Our approach only requires the principal singular vector corresponding to the highest eigenvalue, as the first principal singular vector explains the majority of variance in the data [ 20 ]. While full-rank SVD decomposes a matrix into a combination of orthogonal vectors, one-rank SVDs captures the most significant singular value and corresponding singular vectors, efficiently representing the matrix with a single dominant direction. For matrix K, the principal LSV is denoted as \(\:SK=({sk}_{1},\:{sk}_{2},\:\dots\:,\:{sk}_{m})\) and the principal RSV is denoted as \(\:TK=({tk}_{1},\:{tk}_{2},\:\dots\:,\:{tk}_{m})\) . Similarly, for matrix E, the principal LSV is denoted as \(\:SE=({se}_{1},\:{se}_{2},\:\dots\:,\:{se}_{m})\) , and the principal RSV is denoted as \(\:TE=({te}_{1},\:{te}_{2},\:\dots\:,\:{te}_{m})\) . The elements in the principal LSV link to the components in the spatial dimension, while the elements in the principal RSV link to the components in the temporal dimension. 4. Difference vectors DS and DT are calculated for LSV pair as DS = SK − SE, and the difference vector of the RSV pair as DT = TK − TE. 5. Calculate z-scores vectors from the difference vectors \(\:DS\) and \(\:DT\) . Apply a z-control chart to each vector, using a predefined significance level of α. Elements in the difference vectors DS and DT that yield a p-value less than α in the left tail are considered statistically out of control, indicating potential anomalies in the spatial and temporal dimensions, respectively. 6. If both vectors \(\:DS\) and \(\:DT\) shows combine abnormal elements update the observed case matrix \(\:K\) by replacing the corresponding spatial and temporal elements with their respective expected values from matrix \(\:E\) . Simultaneously, update the relative risk matrix \(\:R\) by substituting those elements with the median value of \(\:R\) . 7. Search again for any new abnormal elements in spatiotemporal dimensions. Keep applying steps (1) through (6) until no more anomalies are detected in either. 8. For all elements in the updated matrix \(\:R\) that are not classified as abnormal, assign a default value of 1, indicating regions and time points without unusual risk elevation. 9. Display the final R on a heat map, using different colors for detected clusters. The flowchart in Fig. 2 visually summarizes the successive steps of the proposed algorithm for spatiotemporal cluster detection, offering a clearer understanding of the methodological framework. 3. Results and Discussions Dengue fever, a mosquito borne viral disease caused by the Dengue virus (DENV) [ 21 ], poses a growing global health threat, infecting over 100 million individuals annually and resulting in approximately 30,000 deaths worldwide (WHO, 2013). DENV, a single stranded positive sense RNA virus of the Flaviviridae family and Flavivirus genus, comprises four antigenically distinct serotypes (DENV-1 to DENV-4), with a potential fifth serotype (DENV-5) recently identified [ 22 ], [ 23 ]. In Pakistan, DENV-2, DENV-3, and DENV-4 have been reported predominantly in Punjab, while DENV-2 and DENV-3 have caused major outbreaks in Swat, Khyber Pakhtunkhwa (KP), with high morbidity and mortality [ 24 ] In Pakistan, especially in KP, frequent outbreaks are driven by inadequate control, poor sanitation, urban crowding, and limited access to clean water. This study investigates dengue case data from Khyber Pakhtunkhwa, revealing several spatiotemporal hotspots with significantly elevated case concentrations. Detecting these clusters is serious for early risk identification, enabling targeted interventions, optimized resource allocation, and public health awareness efforts. The application of spatiotemporal cluster detection techniques offers actionable, evidence based insights for public health authorities to strengthen vector surveillance, implement vaccination or fogging programs, and reduce the morbidity and mortality associated with dengue outbreaks in endemic settings. Moreover, this study identifies high risk regions/periods that may warrant further investigation by researchers and public health experts to explore fundamental environmental, social, and infrastructural factors contributing to disease transmission. The district wise population distribution in Khyber Pakhtunkhwa (KP), based on the 2017 census, is shown in Fig. 3 . Peshawar stands out as the record populated district with nearly 4.76 million people, followed by Mardan (2.74 million), Swat (2.68 million), and Swabi (1.89 million). In contrast, the districts such as Tor Ghar (200,445), Chitral Lower (320,407), and Kohistan Lower (340,017) report the lowest populations. This wide gap highlights the geographic and demographic diversity within KP, which is essential to consider in public health planning, resource distribution, and disease surveillance policies. The temporal distribution of recorded dengue cases over the year 2024 is illustrated in Fig. 4. The outbreak appears to initiate in August, escalates sharply to a peak in October indicating the height of transmission and then gradually declines through November, dropping by December. This seasonal pattern highlights October as the critical intervention period for disease control and prevention strategies. The highest incidence of dengue cases in 2024 was reported in the districts of Peshawar, Mardan, and Lakki Marwat as shown in Fig. 5 . This spatial concentration may be attributed to a combination of favorable climatic conditions such as high humidity and persistent warmth along with urban congestion, poor waste management, and stagnant water bodies that create ideal breeding surroundings for Aedes mosquitoes. The alpha threshold was set at 0.10 because, in common diseases, observed cases typically exceed expected cases in most regions. Setting the alpha threshold at 0.05 or 0.01 could limit the detection of multiple high-risk regions/periods or lead to undetected hotspots. To verify the results, we plotted the observed and expected dengue cases for all 29 districts, month by month from January to December 2024. These comparisons, shown in Fig. 6, help illustrate how closely the model captures the actual disease trends. The heatmap presents the spatiotemporal distribution of the RR matrix for dengue incidence. As shown in Figs. 7, 8 and 9 , the most prominent and statistically significant cluster emerged in October, including the districts of Peshawar and Lower Kohistan, with a notably high average RR of 1.925, shown by a red colour. A secondary likely cluster was detected in the districts of Mardan and Laki Marwat during September and November, while Mansehra, Abbottabad, Kohat, Nowshera, and Hangu showed high risks during September, November, and December. These districts recorded a RR of 1.519, shown in green. An RR value of 1, by contrast, shows the absence of abnormal case counts, suggesting no unusual disease activity in those regions during specific periods. Importantly, the persistence of high RR values across multiple months in Peshawar, Mardan, Mansehra, Abbottabad, Kohat, Nowshera, Hangu, and Laki Marwat suggests these areas acted as repeated hotspots. These spatial and temporal patterns highlight a continuous risk of transmission rather than isolated outbreaks. Such findings are consistent with prior studies indicating that densely populated urban centres with inadequate sanitation, stagnant water sources, and climatic suitability provide favourable environments for dengue spread. 4. Performance evaluation A comparative evaluation of the proposed method against the EigenSpot and Multi-EigenSpot algorithms for detecting spatiotemporal clusters is presented in Fig. 10 . The results demonstrate notable differences in performance. EigenSpot fails to detect any clusters, while Multi-EigenSpot detects only a single cluster and overlooks several other significant areas of activity. In contrast, the proposed algorithm accurately detects multiple disease clusters across both spatial and temporal dimensions. The map highlights its ability to capture real, high-risk areas with greater precision, even in the presence of sparse data. This not only addresses the limitations of previous approaches but also supports the proposed method’s strength in identifying meaningful patterns that conventional techniques often miss. From Table 1 it is cleared that the proposed algorithm is computationally efficient. Overall, the visual and analytical outcomes confirm that the proposed framework offers improved sensitivity and accuracy in cluster detection, making it a more reliable tool for epidemiological surveillance and timely public health response. Table 1 Computational time of the Multi-EigenSpot Algorithm vs the Novel Multi-EigenSpace Algorithm Features Multi EigenSpace Algorithm Novel Multi-EigenSpace Algorithm Data (Matrix Size) 29 by 12 29 by 12 Detected clusters Method Spatial: 2 Temporal: 1 Spatiotemporal: 1 SVD Spatial: 8 Temporal: 4 Spatiotemporal: 2 SVDs (Truncated form of SVD) Loop Nested Loops Vectorized Allocation Grow in Loop Pre-Allocated Relative Efficiency Baseline 5 to 10 times faster Estimated Computation time 1 to 3 seconds 0.1 to 0.5 seconds 5. Discussion and Conclusion The proposed Novel Multi-EigenSpot method represents a considerable advancement in spatiotemporal cluster detection, addressing two key limitations of prior approaches. Firstly, the inability to detect multiple and rare clusters, and secondly, the computational inefficiency inherent in loop-based implementations. Applied to 2024 dengue surveillance data from Khyber Pakhtunkhwa, our proposed framework demonstrated strong empirical performance by detecting several statistically significant clusters ignored by both EigenSpot and Multi-EigenSpot. In particular, while EigenSpot failed to identify any clusters and Multi-EigenSpot detected only two districts, Peshawar and Kohistan Lower, as active in October, our algorithm has power to detect more clusters across both spatial and temporal dimensions. These included: Primary significant clusters: Peshawar and Kohistan Lower in October. Secondary significant clusters: Mardan and Lakki Marwat in September and November, and a series of irregular hotspots in Abbottabad, Kohat, Mansehra, Hangu, and Nowshera between September and December. In summary, this enhanced eigenspace framework delivers broader detection capability, detectd irregular shape cluster, and significantly reduces processing time. It presents a more reliable and scalable tool for epidemiological surveillance of emerging or low incidence diseases. Despite its strengths, the method has several limitations. The current heatmap visualization of the final relative risk matrix effectively highlights spatial and temporal hotspots, but it does not provide insight into the directional spread or transmission dynamics of the disease. Moreover, by aggregating data into fixed administrative regions and discrete time intervals (e.g., monthly), the approach may obscure patterns that transcend district boundaries or overlap across time windows. Future work will aim to enhance the model's inferential depth and adaptability. This includes incorporating spatiotemporal network structures to explore potential transmission pathways, adopting rolling window analysis for more fluid detection over continuous time, and integrating auxiliary covariates such as environmental, demographic, or mobility data to better contextualize underlying outbreak drivers. These developments will extend the Novel EigenSpace paradigm into a more comprehensive, adaptive platform for high resolution disease monitoring and targeted intervention planning. Declarations Acknowledgement The authors extend their appreciation to the Deanship of Scientific Research at Northern Border University, Arar, KSA for funding this research work through the project number "NBU-FPEJ-2025-2942-XX" Declaration of generative AI and AI-assisted technologies in the writing process During the preparation of this work the author(s) used Grammarly in order to readability and avoid grammatical mistakes. After using this tool/service, the author(s) reviewed and edited the content as needed and take(s) full responsibility for the content of the publication. Data Availability The dataset used in this study is publicly available at the National Institute of Health Pakistan website: https://www.nih.org.pk/phb/weekly-bulletin. Researchers may use this dataset freely for replication and validation purposes. Conflict of interest The authors have no conflict of interest. Credit authorship contribution Writing - original draft: Conceptualization: Muhammad Fayyaz Supervision : Alamgir Khalil Writing–review & editing, Investigation, Validation: Hameed Ali Project administration, Funding acquisition: Zeineb Klai Solution Methodology, Software, Formal analysis: Sami Ullah Project administration, Funding acquisition; Solution methodology: Abdulrahman Obaid Alshammari Funding : No Funding available for this research References W. Socha, M. Kwasnik, M. Larska, J. Rola, and W. Rozek, “Vector-Borne Viral Diseases as a Current Threat for Human and Animal Health—One Health Perspective,” Journal of Clinical Medicine , vol. 11, no. 11, Art. no. 11, Jan. 2022, doi: 10.3390/jcm11113026. R. E. Baker et al. , “Infectious disease in an era of global change,” Nat Rev Microbiol , vol. 20, no. 4, pp. 193–205, Apr. 2022, doi: 10.1038/s41579-021-00639-z. N. C. Nieto, K. Khan, G. Uhllah, and M. B. Teglas, “The Emergence and Maintenance of Vector-Borne Diseases in the Khyber Pakhtunkhwa Province, and the Federally Administered Tribal Areas of Pakistan,” Front. Physiol. , vol. 3, July 2012, doi: 10.3389/fphys.2012.00250. B. Ingelbeen et al. , “Embedding risk monitoring in infectious disease surveillance for timely and effective outbreak prevention and control,” BMJ Glob Health , vol. 10, no. 2, Feb. 2025, doi: 10.1136/bmjgh-2024-016870. H. Wang and A. Rodríguez, “Identifying pediatric cancer clusters in Florida using loglinear models and generalized lasso penalties,” Stat Public Policy (Phila) , vol. 1, no. 1, pp. 86–96, 2014, doi: 10.1080/2330443X.2014.960120. R. Amin, A. Bohnert, L. Holmes, A. Rajasekaran, and C. Assanasen, “Epidemiologic mapping of Florida childhood cancer clusters,” Pediatric Blood & Cancer , vol. 54, no. 4, pp. 511–518, Apr. 2010, doi: 10.1002/pbc.22403. M. Kulldorff, W. F. Athas, E. J. Feurer, B. A. Miller, and C. R. Key, “Evaluating cluster alarms: a space-time scan statistic and brain cancer in Los Alamos, New Mexico.,” Am J Public Health , vol. 88, no. 9, pp. 1377–1380, Sept. 1998, doi: 10.2105/AJPH.88.9.1377. M. Kulldorff, “Prospective Time Periodic Geographical Disease Surveillance Using a Scan Statistic,” Journal of the Royal Statistical Society Series A: Statistics in Society , vol. 164, no. 1, pp. 61–72, Jan. 2001, doi: 10.1111/1467-985x.00186. V. S. Iyengar, “Space-time clusters with flexible shapes,” MMWR Suppl , vol. 54, pp. 71–6, 2005. R. Assunção and T. Correa, “Surveillance to detect emerging space–time clusters,” Computational Statistics & Data Analysis , vol. 53, no. 8, pp. 2817–2830, June 2009, doi: 10.1016/j.csda.2008.10.032. K. Takahashi, M. Kulldorff, T. Tango, and K. Yih, “A flexibly shaped space-time scan statistic for disease outbreak detection and monitoring,” Int J Health Geogr , vol. 7, no. 1, p. 14, 2008, doi: 10.1186/1476-072X-7-14. W. Dong, X. Zhang, Z. Jiang, W. Sun, L. Xie, and A. Hampapur, “Detect irregularly shaped spatio-temporal clusters for decision support,” in Proceedings of 2011 IEEE International Conference on Service Operations, Logistics and Informatics , IEEE, 2011, pp. 231–236. Accessed: July 01, 2025. [Online]. Available: https://ieeexplore.ieee.org/abstract/document/5986561/ W. Dong, X. Zhang, L. Li, C. Sun, L. Shi, and W. Sun, “Detecting Irregularly Shaped Significant Spatial and Spatio-Temporal Clusters,” in Proceedings of the 2012 SIAM International Conference on Data Mining , Society for Industrial and Applied Mathematics, Apr. 2012, pp. 732–743. doi: 10.1137/1.9781611972825.63. M. A. Costa and M. Kulldorff, “Maximum linkage space-time permutation scan statistics for disease outbreak detection,” Int J Health Geogr , vol. 13, no. 1, p. 20, 2014, doi: 10.1186/1476-072X-13-20. D. B. Neill, Detection of spatial and spatio-temporal clusters . Carnegie Mellon University, 2006. Accessed: July 01, 2025. [Online]. Available: https://search.proquest.com/openview/9e4841184f5e31120f86231108ff15c8/1?pq-origsite=gscholar&cbl=18750&diss=y D. B. Neill, “An empirical comparison of spatial scan statistics for outbreak detection,” Int J Health Geogr , vol. 8, no. 1, p. 20, 2009, doi: 10.1186/1476-072X-8-20. H. Fanaee-T and J. Gama, “Eigenspace method for spatiotemporal hotspot detection,” Expert Systems , vol. 32, no. 3, pp. 454–464, 2015, doi: 10.1111/exsy.12088. S. Ullah, H. Daud, S. C. Dass, H. Fanaee-T, and A. Khalil, “An Eigenspace approach for detecting multiple space-time disease clusters: Application to measles hotspots detection in Khyber-Pakhtunkhwa, Pakistan,” Plos one , vol. 13, no. 6, p. e0199176, 2018. “Final Results (Census-2017) | Pakistan Bureau of Statistics.” Accessed: July 25, 2025. [Online]. Available: https://www.pbs.gov.pk/content/final-results-census-2017 H. Abdi and L. J. Williams, “Principal component analysis,” WIREs Computational Statistics , vol. 2, no. 4, pp. 433–459, 2010, doi: 10.1002/wics.101. X. Chen and P. Moraga, “Forecasting dengue across Brazil with LSTM neural networks and SHAP-driven lagged climate and spatial effects,” BMC Public Health , vol. 25, no. 1, p. 973, Mar. 2025, doi: 10.1186/s12889-025-22106-7. S. K. Roy and S. Bhattacharjee, “Dengue virus: epidemiology, biology, and disease aetiology,” Can. J. Microbiol. , vol. 67, no. 10, pp. 687–702, Oct. 2021, doi: 10.1139/cjm-2020-0572. M. F. Lee, Y. S. Wu, and C. L. Poh, “Molecular Mechanisms of Antiviral Agents against Dengue Virus,” Viruses , vol. 15, no. 3, Art. no. 3, Mar. 2023, doi: 10.3390/v15030705. J. Khan, A. Ghaffar, and S. A. Khan, “The changing epidemiological pattern of Dengue in Swat, Khyber Pakhtunkhwa,” PLOS ONE , vol. 13, no. 4, p. e0195706, Apr. 2018, doi: 10.1371/journal.pone.0195706. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7484169","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":523910931,"identity":"cc2f65b2-b718-4240-8a2a-07d58b6cbf56","order_by":0,"name":"Muhammad Fayyaz","email":"","orcid":"","institution":"University of Peshawar","correspondingAuthor":false,"prefix":"","firstName":"Muhammad","middleName":"","lastName":"Fayyaz","suffix":""},{"id":523910932,"identity":"707b5214-3f59-46c0-b0cb-fe541c5a88d5","order_by":1,"name":"Alamgir Khalil","email":"","orcid":"","institution":"University of Peshawar","correspondingAuthor":false,"prefix":"","firstName":"Alamgir","middleName":"","lastName":"Khalil","suffix":""},{"id":523910933,"identity":"9d17c238-4c14-49de-90bd-623bf8cba8e9","order_by":2,"name":"Hameed Ali","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA50lEQVRIiWNgGAWjYBACAzDJxiDDxt4AZjICKWaitPCw8RwgVQuDRAKRWszZTyc++FBmx8Mn+fiZdAGDjeyGA7yHDfBpsezJ3Ww441wyD5t0mpn0DIY04w0H+JIT8DrsQO42ad42ZqCWBDNpHobDiRsO8BgfwKvl/Nvtv/+21fOwSR7/BtTynwgtN3K3MTO2HeZhk+AB2XIArAW/w2683SzZc+44MJBziq1nGCQbzzzMl4zX+wbnczd++FFWLSfffnzj7YIKO9m+472HJfBpQQHM4Ghi5iFaAzwKSdEyCkbBKBgFIwEAAJFLRXqCYaomAAAAAElFTkSuQmCC","orcid":"","institution":"The University of Agriculture","correspondingAuthor":true,"prefix":"","firstName":"Hameed","middleName":"","lastName":"Ali","suffix":""},{"id":523910934,"identity":"a9b5c04a-3c10-4d50-87cd-d28b3ae39446","order_by":3,"name":"Zeineb Klai","email":"","orcid":"","institution":"Northern Border University","correspondingAuthor":false,"prefix":"","firstName":"Zeineb","middleName":"","lastName":"Klai","suffix":""},{"id":523910935,"identity":"34835936-6b00-45f4-a279-af13dd4a6d32","order_by":4,"name":"Sami Ullah","email":"","orcid":"","institution":"Beijing University of Technology","correspondingAuthor":false,"prefix":"","firstName":"Sami","middleName":"","lastName":"Ullah","suffix":""},{"id":523910936,"identity":"6e5f49a0-a3b2-491e-993d-d726ff280c9e","order_by":5,"name":"Abdulrahman Obaid Alshammari","email":"","orcid":"","institution":"Jouf University","correspondingAuthor":false,"prefix":"","firstName":"Abdulrahman","middleName":"Obaid","lastName":"Alshammari","suffix":""}],"badges":[],"createdAt":"2025-08-29 02:38:22","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7484169/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7484169/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":93004985,"identity":"555284f4-1754-4136-9837-5377a80c4d00","added_by":"auto","created_at":"2025-10-08 06:30:27","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1383827,"visible":true,"origin":"","legend":"","description":"","filename":"Manuscript..docx","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/b19e633168c72455c5256b99.docx"},{"id":93006229,"identity":"fc18a7ea-2356-4fb0-a396-383e251ce8f5","added_by":"auto","created_at":"2025-10-08 06:46:27","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":7728,"visible":true,"origin":"","legend":"","description":"","filename":"0fe6778fdd69495190167b495a5fc087.json","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/af59c951485df88f8bef013a.json"},{"id":93004987,"identity":"1fc7ec51-c126-47f8-95d6-a4a1d5135fb9","added_by":"auto","created_at":"2025-10-08 06:30:27","extension":"xml","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":78673,"visible":true,"origin":"","legend":"","description":"","filename":"0fe6778fdd69495190167b495a5fc0871enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/3e7a60338508cb721c185e3d.xml"},{"id":93006230,"identity":"40a48ab3-7862-4efd-b7ec-7ff2183e8bc6","added_by":"auto","created_at":"2025-10-08 06:46:27","extension":"jpeg","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":229123,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/b50da81ec52a089ea4acf07b.jpeg"},{"id":93005496,"identity":"6ab405f6-9a6c-4c83-a3dd-7b3f85d1320b","added_by":"auto","created_at":"2025-10-08 06:38:28","extension":"jpeg","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":93169,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/bfd27a567fa1f05e781aa063.jpeg"},{"id":93004978,"identity":"0470f736-6f49-4d5d-9e90-0d2f84dbfeae","added_by":"auto","created_at":"2025-10-08 06:30:27","extension":"png","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":14184,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/dc87da405353107260bddcdb.png"},{"id":93005481,"identity":"db9dac23-0d84-4529-a2f8-8b13cefd6bea","added_by":"auto","created_at":"2025-10-08 06:38:27","extension":"jpeg","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":142760,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage4.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/de4a6042f382585ed48dcfe2.jpeg"},{"id":93004988,"identity":"89686f59-1b2c-4eae-9e36-a0a1be73a460","added_by":"auto","created_at":"2025-10-08 06:30:27","extension":"jpeg","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":174898,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage5.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/8b53b4231d71401b896ba0cb.jpeg"},{"id":93005007,"identity":"6d4ebf04-2f7e-4b2f-b939-de5f769eb1e7","added_by":"auto","created_at":"2025-10-08 06:30:28","extension":"jpeg","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":186417,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage6.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/146e279a59469606f2febc0a.jpeg"},{"id":93005010,"identity":"d3f8dfac-a403-4149-b4b6-7ebc482e5c86","added_by":"auto","created_at":"2025-10-08 06:30:28","extension":"jpeg","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":211779,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage7.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/99733b3619d73fab1007365e.jpeg"},{"id":93005012,"identity":"958f8dff-d695-424a-86b2-9c5c19eb4a0a","added_by":"auto","created_at":"2025-10-08 06:30:28","extension":"jpeg","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":260174,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage8.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/407f643bf6435bfd8da83cd3.jpeg"},{"id":93004973,"identity":"5387d437-4306-4e76-8c49-49a24171928b","added_by":"auto","created_at":"2025-10-08 06:30:27","extension":"png","order_by":13,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":99400,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/139c6218b4e2321c1caa46dc.png"},{"id":93005009,"identity":"ddfc4f61-47bf-4964-91e7-9021b217d916","added_by":"auto","created_at":"2025-10-08 06:30:28","extension":"png","order_by":14,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":32398,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/25dd47b315391c63fa14b8d9.png"},{"id":93004989,"identity":"5d40b9cc-e02e-441f-9da6-24f5f6c71316","added_by":"auto","created_at":"2025-10-08 06:30:27","extension":"png","order_by":15,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":6084,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/15cf3633e5e90994ea3691c4.png"},{"id":93004991,"identity":"2ebe657c-38aa-4972-8f09-3fb950cc62cb","added_by":"auto","created_at":"2025-10-08 06:30:27","extension":"png","order_by":16,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":33820,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/1231f1f817a4f7342b7b9324.png"},{"id":93005013,"identity":"1f197485-887b-4f51-ba07-c8e185875e59","added_by":"auto","created_at":"2025-10-08 06:30:29","extension":"png","order_by":17,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":81467,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/a2fb9e186e06ac92bb0af21a.png"},{"id":93005502,"identity":"0a10b2a7-0777-49d1-b64f-81a261dbb4a4","added_by":"auto","created_at":"2025-10-08 06:38:28","extension":"png","order_by":18,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":83084,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/1a827c45321aa3321117833d.png"},{"id":93005500,"identity":"c51062db-b4c4-4873-83af-7076dba1a569","added_by":"auto","created_at":"2025-10-08 06:38:28","extension":"png","order_by":19,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":80738,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/8f547593bd85da433c7940a1.png"},{"id":93005001,"identity":"fec5ec16-8aad-4ded-8220-b9d37ac1c25c","added_by":"auto","created_at":"2025-10-08 06:30:28","extension":"png","order_by":20,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":100127,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage8.png","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/488af80da27f1145cc0e20ee.png"},{"id":93005495,"identity":"d4f9cac7-5bf1-4176-9042-8a9669c2f3a7","added_by":"auto","created_at":"2025-10-08 06:38:28","extension":"xml","order_by":21,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":78073,"visible":true,"origin":"","legend":"","description":"","filename":"0fe6778fdd69495190167b495a5fc0871structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/da0626ea40476a3c4d0513d7.xml"},{"id":93004975,"identity":"a710b7a2-598c-4659-9401-b3ac4e0c3799","added_by":"auto","created_at":"2025-10-08 06:30:27","extension":"html","order_by":22,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":91066,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/a425d3101ad4f9dd99f4afa0.html"},{"id":93004974,"identity":"923a29ee-f8b4-4031-a2e9-84da3ad59eaa","added_by":"auto","created_at":"2025-10-08 06:30:27","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":145705,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eStudy Area Map of Khyber Pakhtunkhwa for 2024 Dengue Cluster Detection\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/f4709e03f97671878d734020.jpg"},{"id":93004983,"identity":"0b40930c-bcfc-4f67-8159-29f24402d3ed","added_by":"auto","created_at":"2025-10-08 06:30:27","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":92554,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFlowchart frameworks the main steps of the proposed algorithm for methodological clarity.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/f166db3e83e0cfcb10036d2b.jpg"},{"id":93004972,"identity":"749f7be1-0d02-45d7-86fa-9067df3f6b9a","added_by":"auto","created_at":"2025-10-08 06:30:26","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":127877,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eDistrict wise population distribution in Khyber Pakhtunkhwa\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/a9d0268ee2c631161da0ac15.jpg"},{"id":93004981,"identity":"79603bbb-9e39-4b52-a411-c8ede84b5ffa","added_by":"auto","created_at":"2025-10-08 06:30:27","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":56731,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eMonthly distribution of observed dengue cases\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/67d0ad2839219d6a4433b6e5.jpg"},{"id":93004994,"identity":"5ced2092-f4c0-4af2-b86f-3ff265cd8a8e","added_by":"auto","created_at":"2025-10-08 06:30:28","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":95846,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eDistrict wise distribution of recorded dengue cases in 2024\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"5.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/fc7cfe92fc69120ce2fcba0d.jpg"},{"id":93005480,"identity":"6f5326f5-51ff-41b7-be3a-9a15fbcafe7b","added_by":"auto","created_at":"2025-10-08 06:38:27","extension":"jpg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":122910,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eObserved and Expected dengue Cases of KP 2024\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"6.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/6b79cfd37134cd8c841097f4.jpg"},{"id":93005483,"identity":"9f392744-79d8-41a7-8c50-54a782fb455a","added_by":"auto","created_at":"2025-10-08 06:38:27","extension":"jpg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":194937,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSpatiotemporal heatmap of dengue cases using the Multi-EigenSpot method\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"7.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/7e6286a04d02a03c6bf86097.jpg"},{"id":93004986,"identity":"6b0a2c25-6733-41a4-8c28-64187f1937a2","added_by":"auto","created_at":"2025-10-08 06:30:27","extension":"jpg","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":208091,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSpatiotemporal heatmap of dengue cases using proposed Novel Multi-EigenSpot method\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"8.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/705cfe2820c6b30d6d525528.jpg"},{"id":93005498,"identity":"31595258-6cd4-4d39-826d-83d178e6eac0","added_by":"auto","created_at":"2025-10-08 06:38:28","extension":"jpg","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":130881,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eStatistically significant spatiotemporal clusters detected using the proposed methodology\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"9.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/92f2101ecb1860569ddc2fc8.jpg"},{"id":93004996,"identity":"52e4616d-78d2-4842-87e8-fc22afdd8d7e","added_by":"auto","created_at":"2025-10-08 06:30:28","extension":"jpg","order_by":10,"title":"Figure 10","display":"","copyAsset":false,"role":"figure","size":155359,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eComparative mapping of dengue clusters detected by EigenSpot, Multi-EigenSpot, and the Proposed Method\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"10.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/7f4c4148020f1339e8725a36.jpg"},{"id":93364074,"identity":"883d9250-69bc-4f3b-9287-dcf1c85496db","added_by":"auto","created_at":"2025-10-13 04:09:03","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2263331,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7484169/v1/27d29550-8feb-4e51-8441-985c440f4707.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Novel Distribution-Free Eigenspace Framework via Truncated Singular Value Decomposition for High-Resolution Dengue Clusters","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eVector Borne Diseases (VBD) such as dengue, malaria, and chikungunya are caused by pathogens transmitted through vectors like mosquitoes and ticks, causes numerous causalities globally [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. These diseases remain a persistent threat to public health, particularly in low and middle income countries, where environmental, climatic, and socioeconomic conditions further severe their spread [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. In regions such as Khyber Pakhtunkhwa (KP), Pakistan, VBD has a rising trend in frequency, intensity, and geographic range, causes by seasonal variation, urbanization, population density, and deficiencies in healthcare infrastructure [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. Effective disease surveillance is essential for identifying outbreaks, tracking spatial transmission dynamics, and implementing timely public health interventions [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. To this end, spatiotemporal data capturing disease occurrence over space and time is utilized to detect abnormal patterns of elevated disease risk. Disease clusters are typically defined as subregions or time intervals within the spatiotemporal domain that shows significantly higher observed cases than those of expected cases [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e], [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. Identifying such clusters enables public health authorities to prioritize regions for further epidemiological investigation and intervention. Numerous statistical methodologies have been developed to detect disease clusters, such as space-time scan statistics, which are mostly designed to detect regularly shaped clusters [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e], [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e], [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e], [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. However, these approaches often fail to accurately identify irregular shaped clusters, especially in geographically complex or terrain influenced regions. This limitation becomes particularly problematic in areas where disease spread follows natural or infrastructural features such as rivers or roads. To address these challenges, researchers have developed more flexible methods for detecting irregular shaped clusters, such as space-time scan statistic [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e], grid based [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e], [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e], and permutation based scan techniques [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. However, all these approaches rely on strong parametric assumptions, which limits their effectiveness when applied to nontraditional, noisy, or sparse datasets [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e], [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. EigenSpot detects spatiotemporal clusters without distributional or shape assumptions by analyzing matrix structure deviations, but it is limited to a single cluster and therefore misses multiple concurrent outbreaks [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. To overcome this limitation, Multi-EigenSpot [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e] was introduced, this method uses the expected case matrix as the baseline rather than the population-at-risk matrix and iteratively replaces detected clusters with expected values to identify multiple clusters. While this algorithm extends the capacity of EigenSpot to detect more than one spatiotemporal cluster, it suffers from several drawbacks, including poor sensitivity to rare disease clusters, high computational cost on large-scale datasets and miss some important cluster in sparce data.\u003c/p\u003e\u003cp\u003eBuilding upon this line of research, the present study introduces a novel enhancement to the eigenspace based framework for detecting multiple, irregularly shaped, and low-incidence spatiotemporal clusters. The key methodological contributions of the proposed approach are as follows:\u003c/p\u003e\u003cp\u003e\u003col\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eRefined Expected Case Matrix Construction: The method incorporates an improved mechanism for constructing the expected case matrix to enhance sensitivity and detect more clusters.\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eImproved Numerical Stability: Rather than employing Singular Value Decomposition (SVD), the truncated form of SVD is used.\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eEnhanced Computational Efficiency: The algorithm utilizes vectorized operations and parallel processing techniques to significantly reduce computational time, thereby improving scalability to high resolution spatiotemporal datasets.\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003c/ol\u003e\u003c/p\u003e\u003cp\u003eTo support interpretation and facilitate real time decision making, the proposed framework incorporates advanced visual analytics, including heatmaps of the relative risk (RR) matrix, which allow for intuitive representation of spatial and temporal variations in disease risk. Such visual tools are dynamic for aiding public health professionals, epidemiologists, and decision makers in quickly detecting and responding to emerging hotspots.\u003c/p\u003e\u003cp\u003eThe rest of this paper is organized as follows:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eSection 2 presents the study area, data sources, and the proposed methodology, detailing the novel EigenSpace algorithm based on truncated SVD, Z-control charts, and heatmap visualization.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eSection \u003cspan refid=\"Sec8\" class=\"InternalRef\"\u003e3\u003c/span\u003e discusses the results and findings from the application of the proposed method to vector borne disease data in Khyber Pakhtunkhwa, highlighting detected clusters.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eSection 4 evaluates the computational efficiency of the proposed method and comparing performance with existing approaches.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eSection 5 concludes the study with key insights, limitations, and future directions.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e"},{"header":"2. Methodology","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\u003ch2\u003e2.1. Materials and Methods\u003c/h2\u003e\u003cp\u003e\u003cstrong\u003eApproach\u003c/strong\u003e\u003cp\u003eThe analysis was conducted using MATLAB R2017a, comparing the proposed Novel EigenSpace method with conventional EigenSpot and Multi EigenSpot techniques.\u003c/p\u003e\u003c/p\u003e\u003cp\u003eAdditionally, spatial mapping and cluster visualization were performed using QGIS, a freely available and open-source geographic information system (Version 3.34 \u0026lsquo;Firenze\u0026rsquo;, \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://qgis.org\u003c/span\u003e\u003cspan address=\"https://qgis.org\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), an open-source geographic information system.\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eStudy Area\u003c/strong\u003e\u003cp\u003eThis study focuses on Khyber Pakhtunkhwa (KP), a northwestern province of Pakistan located between 31.0\u0026deg;N\u0026ndash;36.5\u0026deg;N latitude and 69.0\u0026deg;E\u0026ndash;74.0\u0026deg;E longitude. The province spans an area of 74,521 km\u0026sup2; and had an estimated population of 30.5\u0026nbsp;million as reported in the 2017 national census (Pakistan Bureau of Statistics, 2017) [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. KP shares international borders with Afghanistan to the northwest and provincial borders with Gilgit Baltistan (northeast), Azad Jammu and Kashmir (east), Ex-FATA (west and south), Balochistan and Punjab (south), and the Islamabad Capital Territory (southeast).\u003c/p\u003e\u003c/p\u003e\u003cp\u003eAs the third most populous region in the country, KP plays a critical role in spatiotemporal disease surveillance and vector borne disease control, particularly due to its demographic scale, cross-border connectivity, and ecological diversity.\u003c/p\u003e\u003cp\u003eFigure\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e displays the study area map of Khyber Pakhtunkhwa, generated by the authors using QGIS software with administrative boundary data.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e\u003ch2\u003e2.2. Data Collection\u003c/h2\u003e\u003cp\u003eThe dataset used in this study contains of district level records of dengue cases for the year 2024, collected from the Integrated Disease Surveillance and Response (IDSR) system, accessible through the Public Health Bulletin of Pakistan \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://phb.nih.org.pk/integratedisease-surveillance-and-response\u003c/span\u003e\u003cspan address=\"https://phb.nih.org.pk/integratedisease-surveillance-and-response\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e a freely accessible online resource for research. To support demographic normalization and the estimation of expected case counts, population data from the 2017 national census were incorporated [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. The data were curated to capture both spatial distribution and temporal trends of disease outbreaks, forming a solid foundation for spatiotemporal cluster analysis of dengue in the region.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec5\" class=\"Section2\"\u003e\u003ch2\u003e2.3. Computationally Efficiency\u003c/h2\u003e\u003cp\u003eThe algorithms proposed in [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e], [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e] implement the EigenSpace framework by leveraging conventional SVD techniques for dimensionality reduction and matrix factorization within the spatiotemporal data structure. However, standard SVD is computationally expensive, particularly for large and sparse datasets. To address the computational and structural limitations essential in EigenSpot and Multi-EigenSpot, the proposed methods integrate an advanced modified SVDs (Sparse Singular Value Decomposition). Design for efficient processing of high dimensional and sparse matrices, SVDs significantly improves decomposition accuracy and scalability, making it particularly effective for analyzing rare disease datasets with limited cases.\u003c/p\u003e\u003cp\u003eThe proposed algorithm integrates the following three methodological components:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003e\u003cb\u003eSVDs\u003c/b\u003e: Employed to extract the principal left and right singular vectors (LSV and RSV) from the K and E matrices for dimensionality reduction.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003e\u003cb\u003eZ-Control Chart\u003c/b\u003e: Employed to find abnormal components in differences vectors.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003e\u003cb\u003eVisualization\u003c/b\u003e: A heatmap is used to display the final Relative Risk (RR) matrix, highlighting potential cluster regions through colour intensity variations.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\u003ch2\u003e2.4. Proposed Algorithm\u003c/h2\u003e\u003cp\u003eThe proposed algorithm is based on the Eigen Space framework technique to identify multiple abnormal spatiotemporal clusters. It functions on matrix formatted surveillance data and utilizes statistical control measures to isolate significant deviations in disease distribution. The methodology is systematically outlined as follows:\u003c/p\u003e\u003cp\u003e1. Total observed cases matrix is denoted by K and the population at risk matrix is denoted by P.\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003c/ol\u003e\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$\\:K=\\left[\\begin{array}{ccc}{k}_{11}\u0026amp;\\:\\cdots\\:\u0026amp;\\:{k}_{1n}\\\\\\:⋮\u0026amp;\\:\\ddots\\:\u0026amp;\\:⋮\\\\\\:{k}_{m1}\u0026amp;\\:\\cdots\\:\u0026amp;\\:{k}_{mn}\\end{array}\\right]\\:and\\:P=\\left[\\begin{array}{ccc}{p}_{11}\u0026amp;\\:\\cdots\\:\u0026amp;\\:{p}_{1n}\\\\\\:⋮\u0026amp;\\:\\ddots\\:\u0026amp;\\:⋮\\\\\\:{p}_{m1}\u0026amp;\\:\\cdots\\:\u0026amp;\\:{p}_{mn}\\end{array}\\right]$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"BlockQuote\"\u003e\u003cp\u003eWhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{k}_{11}\\)\u003c/span\u003e\u003c/span\u003e is the total disease in the first region first-time point \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{p}_{11}\\)\u003c/span\u003e\u003c/span\u003e is the total population at risk in the first region and first-time point\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\:m\\)\u003c/span\u003e\u003c/span\u003e is total spatial dimensions and\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\:n\\:\\)\u003c/span\u003e\u003c/span\u003etotal time points.\u003c/p\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e2. The expected disease cases matrix \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{E}_{ij}\\)\u003c/span\u003e\u003c/span\u003e and relative risks matrix \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{R}_{ij}\\)\u003c/span\u003e\u003c/span\u003e are computed using the observed cases \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{K}_{ij}\\)\u003c/span\u003e\u003c/span\u003e and population at risk \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{P}_{ij}\\:\\)\u003c/span\u003e\u003c/span\u003ematrices.\u003c/p\u003e\u003cp\u003eWhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:ith\\)\u003c/span\u003e\u003c/span\u003e represent spatial unit (region) and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:jth\\)\u003c/span\u003e\u003c/span\u003e time point.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Equb\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e\n$$\\:E=\\left[\\begin{array}{ccc}{E}_{11}\u0026amp;\\:\\cdots\\:\u0026amp;\\:{E}_{1n}\\\\\\:⋮\u0026amp;\\:\\ddots\\:\u0026amp;\\:⋮\\\\\\:{E}_{m1}\u0026amp;\\:\\cdots\\:\u0026amp;\\:{E}_{mn}\\end{array}\\right]\\:and\\:R=\\left[\\begin{array}{ccc}{R}_{11}\u0026amp;\\:\\cdots\\:\u0026amp;\\:{R}_{1n}\\\\\\:⋮\u0026amp;\\:\\ddots\\:\u0026amp;\\:⋮\\\\\\:{R}_{m1}\u0026amp;\\:\\cdots\\:\u0026amp;\\:{R}_{mn}\\end{array}\\right]$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"BlockQuote\"\u003e\u003cp\u003eOnce the relative risk matrix \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:R\\)\u003c/span\u003e\u003c/span\u003e is calculated, it serves as the basis for visualizing potential clusters on a heatmap.\u003c/p\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e3. The one-rank SVDs is used to find the LSV and RSV for matrices \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:K\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:E\\)\u003c/span\u003e\u003c/span\u003e. Our approach only requires the principal singular vector corresponding to the highest eigenvalue, as the first principal singular vector explains the majority of variance in the data [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. While full-rank SVD decomposes a matrix into a combination of orthogonal vectors, one-rank SVDs captures the most significant singular value and corresponding singular vectors, efficiently representing the matrix with a single dominant direction. For matrix K, the principal LSV is denoted as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:SK=({sk}_{1},\\:{sk}_{2},\\:\\dots\\:,\\:{sk}_{m})\\)\u003c/span\u003e\u003c/span\u003e and the principal RSV is denoted as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:TK=({tk}_{1},\\:{tk}_{2},\\:\\dots\\:,\\:{tk}_{m})\\)\u003c/span\u003e\u003c/span\u003e. Similarly, for matrix E, the principal LSV is denoted as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:SE=({se}_{1},\\:{se}_{2},\\:\\dots\\:,\\:{se}_{m})\\)\u003c/span\u003e\u003c/span\u003e, and the principal RSV is denoted as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:TE=({te}_{1},\\:{te}_{2},\\:\\dots\\:,\\:{te}_{m})\\)\u003c/span\u003e\u003c/span\u003e. The elements in the principal LSV link to the components in the spatial dimension, while the elements in the principal RSV link to the components in the temporal dimension.\u003c/p\u003e\u003cp\u003e4. Difference vectors DS and DT are calculated for LSV pair as DS\u0026thinsp;=\u0026thinsp;SK\u0026thinsp;\u0026minus;\u0026thinsp;SE, and the difference vector of the RSV pair as DT\u0026thinsp;=\u0026thinsp;TK\u0026thinsp;\u0026minus;\u0026thinsp;TE.\u003c/p\u003e\u003cp\u003e5. Calculate z-scores vectors from the difference vectors \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:DS\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:DT\\)\u003c/span\u003e\u003c/span\u003e. Apply a z-control chart to each vector, using a predefined significance level of α. Elements in the difference vectors DS and DT that yield a p-value less than α in the left tail are considered statistically out of control, indicating potential anomalies in the spatial and temporal dimensions, respectively.\u003c/p\u003e\u003cp\u003e6. If both vectors \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:DS\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:DT\\)\u003c/span\u003e\u003c/span\u003e shows combine abnormal elements update the observed case matrix \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:K\\)\u003c/span\u003e\u003c/span\u003e by replacing the corresponding spatial and temporal elements with their respective expected values from matrix \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:E\\)\u003c/span\u003e\u003c/span\u003e. Simultaneously, update the relative risk matrix \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:R\\)\u003c/span\u003e\u003c/span\u003e by substituting those elements with the median value of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:R\\)\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e\u003cp\u003e7. Search again for any new abnormal elements in spatiotemporal dimensions. Keep applying steps (1) through (6) until no more anomalies are detected in either.\u003c/p\u003e\u003cp\u003e8. For all elements in the updated matrix \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:R\\)\u003c/span\u003e\u003c/span\u003e that are not classified as abnormal, assign a default value of 1, indicating regions and time points without unusual risk elevation.\u003c/p\u003e\n\u003cp\u003e9. Display the final R on a heat map, using different colors for detected clusters.\u003c/p\u003e\n\u003cp\u003eThe flowchart in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e visually summarizes the successive steps of the proposed algorithm for spatiotemporal cluster detection, offering a clearer understanding of the methodological framework.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e"},{"header":"3. Results and Discussions","content":"\u003cp\u003eDengue fever, a mosquito borne viral disease caused by the Dengue virus (DENV) [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e], poses a growing global health threat, infecting over 100\u0026nbsp;million individuals annually and resulting in approximately 30,000 deaths worldwide (WHO, 2013). DENV, a single stranded positive sense RNA virus of the Flaviviridae family and Flavivirus genus, comprises four antigenically distinct serotypes (DENV-1 to DENV-4), with a potential fifth serotype (DENV-5) recently identified [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e], [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]. In Pakistan, DENV-2, DENV-3, and DENV-4 have been reported predominantly in Punjab, while DENV-2 and DENV-3 have caused major outbreaks in Swat, Khyber Pakhtunkhwa (KP), with high morbidity and mortality [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]\u003c/p\u003e\u003cp\u003eIn Pakistan, especially in KP, frequent outbreaks are driven by inadequate control, poor sanitation, urban crowding, and limited access to clean water. This study investigates dengue case data from Khyber Pakhtunkhwa, revealing several spatiotemporal hotspots with significantly elevated case concentrations. Detecting these clusters is serious for early risk identification, enabling targeted interventions, optimized resource allocation, and public health awareness efforts. The application of spatiotemporal cluster detection techniques offers actionable, evidence based insights for public health authorities to strengthen vector surveillance, implement vaccination or fogging programs, and reduce the morbidity and mortality associated with dengue outbreaks in endemic settings.\u003c/p\u003e\u003cp\u003eMoreover, this study identifies high risk regions/periods that may warrant further investigation by researchers and public health experts to explore fundamental environmental, social, and infrastructural factors contributing to disease transmission.\u003c/p\u003e\u003cp\u003eThe district wise population distribution in Khyber Pakhtunkhwa (KP), based on the 2017 census, is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e. Peshawar stands out as the record populated district with nearly 4.76\u0026nbsp;million people, followed by Mardan (2.74\u0026nbsp;million), Swat (2.68\u0026nbsp;million), and Swabi (1.89\u0026nbsp;million). In contrast, the districts such as Tor Ghar (200,445), Chitral Lower (320,407), and Kohistan Lower (340,017) report the lowest populations. This wide gap highlights the geographic and demographic diversity within KP, which is essential to consider in public health planning, resource distribution, and disease surveillance policies.\u003c/p\u003e\u003cp\u003eThe temporal distribution of recorded dengue cases over the year 2024 is illustrated in Fig.\u0026nbsp;4. The outbreak appears to initiate in August, escalates sharply to a peak in October indicating the height of transmission and then gradually declines through November, dropping by December. This seasonal pattern highlights October as the critical intervention period for disease control and prevention strategies.\u003c/p\u003e\u003cp\u003eThe highest incidence of dengue cases in 2024 was reported in the districts of Peshawar, Mardan, and Lakki Marwat as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e5\u003c/span\u003e. This spatial concentration may be attributed to a combination of favorable climatic conditions such as high humidity and persistent warmth along with urban congestion, poor waste management, and stagnant water bodies that create ideal breeding surroundings for Aedes mosquitoes.\u003c/p\u003e\u003cp\u003eThe alpha threshold was set at 0.10 because, in common diseases, observed cases typically exceed expected cases in most regions. Setting the alpha threshold at 0.05 or 0.01 could limit the detection of multiple high-risk regions/periods or lead to undetected hotspots. To verify the results, we plotted the observed and expected dengue cases for all 29 districts, month by month from January to December 2024. These comparisons, shown in Fig.\u0026nbsp;6, help illustrate how closely the model captures the actual disease trends.\u003c/p\u003e\u003cp\u003eThe heatmap presents the spatiotemporal distribution of the RR matrix for dengue incidence. As shown in Figs.\u0026nbsp;7, 8 and \u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e9\u003c/span\u003e, the most prominent and statistically significant cluster emerged in October, including the districts of Peshawar and Lower Kohistan, with a notably high average RR of 1.925, shown by a red colour. A secondary likely cluster was detected in the districts of Mardan and Laki Marwat during September and November, while Mansehra, Abbottabad, Kohat, Nowshera, and Hangu showed high risks during September, November, and December. These districts recorded a RR of 1.519, shown in green. An RR value of 1, by contrast, shows the absence of abnormal case counts, suggesting no unusual disease activity in those regions during specific periods.\u003c/p\u003e\u003cp\u003eImportantly, the persistence of high RR values across multiple months in Peshawar, Mardan, Mansehra, Abbottabad, Kohat, Nowshera, Hangu, and Laki Marwat suggests these areas acted as repeated hotspots. These spatial and temporal patterns highlight a continuous risk of transmission rather than isolated outbreaks. Such findings are consistent with prior studies indicating that densely populated urban centres with inadequate sanitation, stagnant water sources, and climatic suitability provide favourable environments for dengue spread.\u003c/p\u003e"},{"header":"4. Performance evaluation","content":"\u003cp\u003eA comparative evaluation of the proposed method against the EigenSpot and Multi-EigenSpot algorithms for detecting spatiotemporal clusters is presented in Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e10\u003c/span\u003e. The results demonstrate notable differences in performance. EigenSpot fails to detect any clusters, while Multi-EigenSpot detects only a single cluster and overlooks several other significant areas of activity.\u003c/p\u003e\u003cp\u003eIn contrast, the proposed algorithm accurately detects multiple disease clusters across both spatial and temporal dimensions. The map highlights its ability to capture real, high-risk areas with greater precision, even in the presence of sparse data. This not only addresses the limitations of previous approaches but also supports the proposed method\u0026rsquo;s strength in identifying meaningful patterns that conventional techniques often miss. From Table \u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e it is cleared that the proposed algorithm is computationally efficient.\u003c/p\u003e\u003cp\u003eOverall, the visual and analytical outcomes confirm that the proposed framework offers improved sensitivity and accuracy in cluster detection, making it a more reliable tool for epidemiological surveillance and timely public health response.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eComputational time of the Multi-EigenSpot Algorithm vs the Novel Multi-EigenSpace Algorithm\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"3\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eFeatures\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eMulti EigenSpace Algorithm\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eNovel Multi-EigenSpace Algorithm\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eData (Matrix Size)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e29 by 12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e29 by 12\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDetected clusters\u003c/p\u003e\u003cp\u003eMethod\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eSpatial: 2\u003c/p\u003e\u003cp\u003eTemporal: 1\u003c/p\u003e\u003cp\u003eSpatiotemporal: 1\u003c/p\u003e\u003cp\u003eSVD\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eSpatial: 8\u003c/p\u003e\u003cp\u003eTemporal: 4\u003c/p\u003e\u003cp\u003eSpatiotemporal: 2\u003c/p\u003e\u003cp\u003eSVDs (Truncated form of SVD)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLoop\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNested Loops\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eVectorized\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAllocation\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eGrow in Loop\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003ePre-Allocated\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRelative Efficiency\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eBaseline\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e5 to 10 times faster\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eEstimated Computation time\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e1 to 3 seconds\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.1 to 0.5 seconds\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e"},{"header":"5. Discussion and Conclusion","content":"\u003cp\u003eThe proposed Novel Multi-EigenSpot method represents a considerable advancement in spatiotemporal cluster detection, addressing two key limitations of prior approaches. Firstly, the inability to detect multiple and rare clusters, and secondly, the computational inefficiency inherent in loop-based implementations. Applied to 2024 dengue surveillance data from Khyber Pakhtunkhwa, our proposed framework demonstrated strong empirical performance by detecting several statistically significant clusters ignored by both EigenSpot and Multi-EigenSpot. In particular, while EigenSpot failed to identify any clusters and Multi-EigenSpot detected only two districts, Peshawar and Kohistan Lower, as active in October, our algorithm has power to detect more clusters across both spatial and temporal dimensions.\u003c/p\u003e\u003cp\u003eThese included:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003ePrimary significant clusters: Peshawar and Kohistan Lower in October.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eSecondary significant clusters: Mardan and Lakki Marwat in September and November, and a series of irregular hotspots in Abbottabad, Kohat, Mansehra, Hangu, and Nowshera between September and December.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eIn summary, this enhanced eigenspace framework delivers broader detection capability, detectd irregular shape cluster, and significantly reduces processing time. It presents a more reliable and scalable tool for epidemiological surveillance of emerging or low incidence diseases.\u003c/p\u003e\u003cp\u003eDespite its strengths, the method has several limitations. The current heatmap visualization of the final relative risk matrix effectively highlights spatial and temporal hotspots, but it does not provide insight into the directional spread or transmission dynamics of the disease. Moreover, by aggregating data into fixed administrative regions and discrete time intervals (e.g., monthly), the approach may obscure patterns that transcend district boundaries or overlap across time windows.\u003c/p\u003e\u003cp\u003eFuture work will aim to enhance the model's inferential depth and adaptability. This includes incorporating spatiotemporal network structures to explore potential transmission pathways, adopting rolling window analysis for more fluid detection over continuous time, and integrating auxiliary covariates such as environmental, demographic, or mobility data to better contextualize underlying outbreak drivers. These developments will extend the Novel EigenSpace paradigm into a more comprehensive, adaptive platform for high resolution disease monitoring and targeted intervention planning.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003eAcknowledgement\u003c/p\u003e\n\u003cp\u003eThe authors extend their appreciation to the Deanship of Scientific Research at Northern Border University, Arar, KSA for funding this research work through the project number \"NBU-FPEJ-2025-2942-XX\"\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDeclaration of generative AI and AI-assisted technologies in the writing process\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eDuring the preparation of this work the author(s) used Grammarly in order to readability and avoid grammatical mistakes. After using this tool/service, the author(s) reviewed and edited the content as needed and take(s) full responsibility for the content of the publication.\u003c/p\u003e\n\u003cp\u003eData Availability\u003c/p\u003e\n\u003cp\u003eThe dataset used in this study is publicly available at the National Institute of Health Pakistan website: https://www.nih.org.pk/phb/weekly-bulletin. Researchers may use this dataset freely for replication and validation purposes.\u003c/p\u003e\n\u003cp\u003eConflict of interest\u003c/p\u003e\n\u003cp\u003eThe authors have no conflict of interest.\u003c/p\u003e\n\u003cp\u003eCredit authorship contribution\u003c/p\u003e\n\u003cp\u003eWriting - original draft: Conceptualization: \u003cstrong\u003eMuhammad Fayyaz\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eSupervision\u003cstrong\u003e:\u0026nbsp;\u003c/strong\u003e\u003cstrong\u003eAlamgir\u003c/strong\u003e\u003cstrong\u003e\u0026nbsp;Khalil\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWriting–review \u0026amp; editing, Investigation, Validation:\u003cstrong\u003e\u0026nbsp;Hameed Ali\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eProject administration, Funding acquisition: \u003cstrong\u003eZeineb Klai\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eSolution Methodology, Software, Formal analysis:\u003cstrong\u003e\u0026nbsp;Sami Ullah\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eProject administration, Funding acquisition; Solution methodology: \u003cstrong\u003eAbdulrahman Obaid Alshammari\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e: No Funding available for this research\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eW. Socha, M. Kwasnik, M. Larska, J. Rola, and W. Rozek, \u0026ldquo;Vector-Borne Viral Diseases as a Current Threat for Human and Animal Health\u0026mdash;One Health Perspective,\u0026rdquo; \u003cem\u003eJournal of Clinical Medicine\u003c/em\u003e, vol. 11, no. 11, Art. no. 11, Jan. 2022, doi: 10.3390/jcm11113026.\u003c/li\u003e\n\u003cli\u003eR. E. Baker \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Infectious disease in an era of global change,\u0026rdquo; \u003cem\u003eNat Rev Microbiol\u003c/em\u003e, vol. 20, no. 4, pp. 193\u0026ndash;205, Apr. 2022, doi: 10.1038/s41579-021-00639-z.\u003c/li\u003e\n\u003cli\u003eN. C. Nieto, K. Khan, G. Uhllah, and M. B. Teglas, \u0026ldquo;The Emergence and Maintenance of Vector-Borne Diseases in the Khyber Pakhtunkhwa Province, and the Federally Administered Tribal Areas of Pakistan,\u0026rdquo; \u003cem\u003eFront. Physiol.\u003c/em\u003e, vol. 3, July 2012, doi: 10.3389/fphys.2012.00250.\u003c/li\u003e\n\u003cli\u003eB. Ingelbeen \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Embedding risk monitoring in infectious disease surveillance for timely and effective outbreak prevention and control,\u0026rdquo; \u003cem\u003eBMJ Glob Health\u003c/em\u003e, vol. 10, no. 2, Feb. 2025, doi: 10.1136/bmjgh-2024-016870.\u003c/li\u003e\n\u003cli\u003eH. Wang and A. Rodr\u0026iacute;guez, \u0026ldquo;Identifying pediatric cancer clusters in Florida using loglinear models and generalized lasso penalties,\u0026rdquo; \u003cem\u003eStat Public Policy (Phila)\u003c/em\u003e, vol. 1, no. 1, pp. 86\u0026ndash;96, 2014, doi: 10.1080/2330443X.2014.960120.\u003c/li\u003e\n\u003cli\u003eR. Amin, A. Bohnert, L. Holmes, A. Rajasekaran, and C. Assanasen, \u0026ldquo;Epidemiologic mapping of Florida childhood cancer clusters,\u0026rdquo; \u003cem\u003ePediatric Blood \u0026amp; Cancer\u003c/em\u003e, vol. 54, no. 4, pp. 511\u0026ndash;518, Apr. 2010, doi: 10.1002/pbc.22403.\u003c/li\u003e\n\u003cli\u003eM. Kulldorff, W. F. Athas, E. J. Feurer, B. A. Miller, and C. R. Key, \u0026ldquo;Evaluating cluster alarms: a space-time scan statistic and brain cancer in Los Alamos, New Mexico.,\u0026rdquo; \u003cem\u003eAm J Public Health\u003c/em\u003e, vol. 88, no. 9, pp. 1377\u0026ndash;1380, Sept. 1998, doi: 10.2105/AJPH.88.9.1377.\u003c/li\u003e\n\u003cli\u003eM. Kulldorff, \u0026ldquo;Prospective Time Periodic Geographical Disease Surveillance Using a Scan Statistic,\u0026rdquo; \u003cem\u003eJournal of the Royal Statistical Society Series A: Statistics in Society\u003c/em\u003e, vol. 164, no. 1, pp. 61\u0026ndash;72, Jan. 2001, doi: 10.1111/1467-985x.00186.\u003c/li\u003e\n\u003cli\u003eV. S. Iyengar, \u0026ldquo;Space-time clusters with flexible shapes,\u0026rdquo; \u003cem\u003eMMWR Suppl\u003c/em\u003e, vol. 54, pp. 71\u0026ndash;6, 2005.\u003c/li\u003e\n\u003cli\u003eR. Assun\u0026ccedil;\u0026atilde;o and T. Correa, \u0026ldquo;Surveillance to detect emerging space\u0026ndash;time clusters,\u0026rdquo; \u003cem\u003eComputational Statistics \u0026amp; Data Analysis\u003c/em\u003e, vol. 53, no. 8, pp. 2817\u0026ndash;2830, June 2009, doi: 10.1016/j.csda.2008.10.032.\u003c/li\u003e\n\u003cli\u003eK. Takahashi, M. Kulldorff, T. Tango, and K. Yih, \u0026ldquo;A flexibly shaped space-time scan statistic for disease outbreak detection and monitoring,\u0026rdquo; \u003cem\u003eInt J Health Geogr\u003c/em\u003e, vol. 7, no. 1, p. 14, 2008, doi: 10.1186/1476-072X-7-14.\u003c/li\u003e\n\u003cli\u003eW. Dong, X. Zhang, Z. Jiang, W. Sun, L. Xie, and A. Hampapur, \u0026ldquo;Detect irregularly shaped spatio-temporal clusters for decision support,\u0026rdquo; in \u003cem\u003eProceedings of 2011 IEEE International Conference on Service Operations, Logistics and Informatics\u003c/em\u003e, IEEE, 2011, pp. 231\u0026ndash;236. Accessed: July 01, 2025. [Online]. Available: https://ieeexplore.ieee.org/abstract/document/5986561/\u003c/li\u003e\n\u003cli\u003eW. Dong, X. Zhang, L. Li, C. Sun, L. Shi, and W. Sun, \u0026ldquo;Detecting Irregularly Shaped Significant Spatial and Spatio-Temporal Clusters,\u0026rdquo; in \u003cem\u003eProceedings of the 2012 SIAM International Conference on Data Mining\u003c/em\u003e, Society for Industrial and Applied Mathematics, Apr. 2012, pp. 732\u0026ndash;743. doi: 10.1137/1.9781611972825.63.\u003c/li\u003e\n\u003cli\u003eM. A. Costa and M. Kulldorff, \u0026ldquo;Maximum linkage space-time permutation scan statistics for disease outbreak detection,\u0026rdquo; \u003cem\u003eInt J Health Geogr\u003c/em\u003e, vol. 13, no. 1, p. 20, 2014, doi: 10.1186/1476-072X-13-20.\u003c/li\u003e\n\u003cli\u003eD. B. Neill, \u003cem\u003eDetection of spatial and spatio-temporal clusters\u003c/em\u003e. Carnegie Mellon University, 2006. Accessed: July 01, 2025. [Online]. Available: https://search.proquest.com/openview/9e4841184f5e31120f86231108ff15c8/1?pq-origsite=gscholar\u0026amp;cbl=18750\u0026amp;diss=y\u003c/li\u003e\n\u003cli\u003eD. B. Neill, \u0026ldquo;An empirical comparison of spatial scan statistics for outbreak detection,\u0026rdquo; \u003cem\u003eInt J Health Geogr\u003c/em\u003e, vol. 8, no. 1, p. 20, 2009, doi: 10.1186/1476-072X-8-20.\u003c/li\u003e\n\u003cli\u003eH. Fanaee-T and J. Gama, \u0026ldquo;Eigenspace method for spatiotemporal hotspot detection,\u0026rdquo; \u003cem\u003eExpert Systems\u003c/em\u003e, vol. 32, no. 3, pp. 454\u0026ndash;464, 2015, doi: 10.1111/exsy.12088.\u003c/li\u003e\n\u003cli\u003eS. Ullah, H. Daud, S. C. Dass, H. Fanaee-T, and A. Khalil, \u0026ldquo;An Eigenspace approach for detecting multiple space-time disease clusters: Application to measles hotspots detection in Khyber-Pakhtunkhwa, Pakistan,\u0026rdquo; \u003cem\u003ePlos one\u003c/em\u003e, vol. 13, no. 6, p. e0199176, 2018.\u003c/li\u003e\n\u003cli\u003e\u0026ldquo;Final Results (Census-2017) | Pakistan Bureau of Statistics.\u0026rdquo; Accessed: July 25, 2025. [Online]. Available: https://www.pbs.gov.pk/content/final-results-census-2017\u003c/li\u003e\n\u003cli\u003eH. Abdi and L. J. Williams, \u0026ldquo;Principal component analysis,\u0026rdquo; \u003cem\u003eWIREs Computational Statistics\u003c/em\u003e, vol. 2, no. 4, pp. 433\u0026ndash;459, 2010, doi: 10.1002/wics.101.\u003c/li\u003e\n\u003cli\u003eX. Chen and P. Moraga, \u0026ldquo;Forecasting dengue across Brazil with LSTM neural networks and SHAP-driven lagged climate and spatial effects,\u0026rdquo; \u003cem\u003eBMC Public Health\u003c/em\u003e, vol. 25, no. 1, p. 973, Mar. 2025, doi: 10.1186/s12889-025-22106-7.\u003c/li\u003e\n\u003cli\u003eS. K. Roy and S. Bhattacharjee, \u0026ldquo;Dengue virus: epidemiology, biology, and disease aetiology,\u0026rdquo; \u003cem\u003eCan. J. Microbiol.\u003c/em\u003e, vol. 67, no. 10, pp. 687\u0026ndash;702, Oct. 2021, doi: 10.1139/cjm-2020-0572.\u003c/li\u003e\n\u003cli\u003eM. F. Lee, Y. S. Wu, and C. L. Poh, \u0026ldquo;Molecular Mechanisms of Antiviral Agents against Dengue Virus,\u0026rdquo; \u003cem\u003eViruses\u003c/em\u003e, vol. 15, no. 3, Art. no. 3, Mar. 2023, doi: 10.3390/v15030705.\u003c/li\u003e\n\u003cli\u003eJ. Khan, A. Ghaffar, and S. A. Khan, \u0026ldquo;The changing epidemiological pattern of Dengue in Swat, Khyber Pakhtunkhwa,\u0026rdquo; \u003cem\u003ePLOS ONE\u003c/em\u003e, vol. 13, no. 4, p. e0195706, Apr. 2018, doi: 10.1371/journal.pone.0195706.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-7484169/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7484169/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe development of efficient and reliable algorithm to identify risk spots is vital in real life domain, specifically in epidemiology. This study aims to design a novel eigenspace based framework for detecting multiple spatiotemporal clusters of vector borne diseases, considering dengue data from Khyber Pakhtunkhwa, Pakistan, as a case study. Unlike traditional scan-based methods constrained by strict distributional assumptions and regular cluster shape, our approach effectively detects spatiotemporal clusters of arbitrary shape and under any data distribution. Initiated by refined expected case matrix construction, the proposed method integrates truncated singular value decomposition and control chart to precisely discover multiple simultaneous clusters. In a direct comparison with classical eigenspace techniques such as EigenSpot and MultiEigenSpot, our framework consistently outperforms in identification of clusters across both spatial and temporal dimensions. Further, our approach also utilizes vectorized implementation making it computationally economic against competing methods. To facilitate clear interpretation and rapid public health response, the framework includes interactive visual analytics, such as heatmaps of the relative risk (RR) matrix. Our work offers significant contribution to the practitioners, specifically epidemiologist in making informed decision in real life domains. While motivated by epidemiological applications, the proposed framework can seamlessly be adapted in other scenario such environmental monitoring, crime hotspot mapping, and disaster impact assessment.\u003c/p\u003e","manuscriptTitle":"Novel Distribution-Free Eigenspace Framework via Truncated Singular Value Decomposition for High-Resolution Dengue Clusters","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-10-08 06:30:21","doi":"10.21203/rs.3.rs-7484169/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"6b64048f-2223-40a1-ab86-3fb1f3bd92fe","owner":[],"postedDate":"October 8th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":55684111,"name":"Biological sciences/Computational biology and bioinformatics"},{"id":55684112,"name":"Physical sciences/Mathematics and computing"}],"tags":[],"updatedAt":"2025-10-13T04:08:49+00:00","versionOfRecord":[],"versionCreatedAt":"2025-10-08 06:30:21","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7484169","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7484169","identity":"rs-7484169","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.