LandScan HD: A High-Resolution Gridded Ambient Population Methodology for the World

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Population datasets accounting for the full range of routine human activities are needed to address many global human security challenges, including disasters, conflict, and infrastructure demand. LandScan High Definition (HD) supports this need through gridded ambient population estimates that measure average human presence between daytime and nighttime at a high spatial resolution of 3 arcseconds (roughly 90 m). Although LandScan HD has traditionally been produced on a country-specific basis, advances in global foundational data and computational resources now enable scaling its methodology to the world. Combining aspects of top-down and bottom-up gridded population methods, LandScan HD allocates subnational population totals from authoritative statistics to built-up areas based on occupancy estimates for multiple facility types (e.g., residential, commercial) and then reaggregates these estimates to a global population grid. We scale this approach by organizing the LandScan HD data stack into a 1° resolution tileset of vector analytic features, enabling an efficient and repeatable workflow for all countries worldwide. Examining the Philippines as an output of the global LandScan HD baseline dataset, we contrast the ambient population with a gridded population representing residential activities (WorldPop) by (1) exploring a practical application for flood risk assessment and (2) evaluating congruence with outcomes of collective human activities (subnational CO2 emissions). Finally, we discuss confronting current LandScan HD limitations through data/modeling and uncertainty quantification improvements and provide outlook for workflow automation and extending the model to social, demographic and economic population characteristics.
Full text 164,581 characters · extracted from preprint-html · click to expand
LandScan HD: A High-Resolution Gridded Ambient Population Methodology for the World | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Method Article LandScan HD: A High-Resolution Gridded Ambient Population Methodology for the World Joseph V. Tuccillo, Jessica Moehl, Daniel Adams, Angela R. Cunningham, and 14 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6396722/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Population datasets accounting for the full range of routine human activities are needed to address many global human security challenges, including disasters, conflict, and infrastructure demand. LandScan High Definition (HD) supports this need through gridded ambient population estimates that measure average human presence between daytime and nighttime at a high spatial resolution of 3 arcseconds (roughly 90 m). Although LandScan HD has traditionally been produced on a country-specific basis, advances in global foundational data and computational resources now enable scaling its methodology to the world. Combining aspects of top-down and bottom-up gridded population methods, LandScan HD allocates subnational population totals from authoritative statistics to built-up areas based on occupancy estimates for multiple facility types (e.g., residential, commercial) and then reaggregates these estimates to a global population grid. We scale this approach by organizing the LandScan HD data stack into a 1° resolution tileset of vector analytic features, enabling an efficient and repeatable workflow for all countries worldwide. Examining the Philippines as an output of the global LandScan HD baseline dataset, we contrast the ambient population with a gridded population representing residential activities (WorldPop) by (1) exploring a practical application for flood risk assessment and (2) evaluating congruence with outcomes of collective human activities (subnational CO 2 emissions). Finally, we discuss confronting current LandScan HD limitations through data/modeling and uncertainty quantification improvements and provide outlook for workflow automation and extending the model to social, demographic and economic population characteristics. Geographic Information Systems population distribution data fusion building morphology building occupancy Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 1 Introduction Measuring the footprint of human activities on Earth is fundamental to addressing global challenges of climate-driven and technological hazards, conflict, demand on infrastructure, and resource scarcity (Pörtner et al., 2022). Central to this task are gridded population datasets, which estimate human presence at high spatial resolution based on the suitability of built-up areas for inhabitation and overcome challenges of limited data availability when estimating global population from census data alone (Bhaduri et al., 2002; Dobson et al., 2000 and 2003; Wardrop et al., 2018; Kugler et al., 2019; Leyk et al., 2019; Archila Bustos et al., 2020). Most gridded population models focus on residential population distributions—a means of “meeting people where they are”—to promote global equity in health services and public resource delivery (Tatem, 2017; Wardrop et al., 2018). However, population changes caused by natural disasters, environmental exposures, demand on critical water and transportation infrastructure, and access to vital services such as nutrition and healthcare cannot be fully addressed without knowledge of human presence both at home and performing routine activities (Bhaduri et al., 2007; Rose & Bright, 2014). Acknowledging the difference in population distribution during nonresidential hours, there exists a complementary need to estimate the ambient 24 h population distribution, which accounts for activities unobserved in census data (e.g., school, work, services). To address this issue, Oak Ridge National Laboratory (ORNL) developed LandScan Global (Dobson et al., 2000; Bhaduri et al., 2002), which provides ambient population distribution at 30 arcsecond (approximately 1 km) resolution. Subsequently, ORNL developed the LandScan USA model (Bhaduri et al., 2007), which provides nighttime as well as daytime gridded population estimates for the United States at 3 arcseconds (approximately 90 m) resolution. Complementing the worldwide coverage offered by LandScan Global, the LandScan High Definition (HD) model emerged to develop gridded populations for select countries that offer increased spatial resolution (3 arcseconds) and temporal resolution (nighttime, daytime) and enhance knowledge of human security challenges, including vaccine distribution and conflict (Weber et al., 2018, Urban et al., 2023b). LandScan HD began as an extension of the LandScan USA model, initially combining dasymetric modeling with settlement detection from high-resolution satellite images (Cheriyadat et al., 2007; Vijayaraj et al., 2007 and 2008) and subsequently building occupancy models (Stewart et al., 2016; Urban et al., 2023a). Although LandScan HD has traditionally been produced on a country-specific basis, recent advances in the spatial coverage and precision of foundational building and land-use datasets, as well as computational resources for gridded population modeling, now enable production of the first global implementation of the model. This paper describes the core LandScan HD methodology and how it is scaled to produce a high-resolution gridded ambient population for the world. Using the LandScan HD global baseline, we provide several illustrations that demonstrate practical application of high spatial resolution ambient populations. We also reflect on limitations of the current approach and strategies for improving the methodological components of our framework as it continues to develop. 2 Background Since 1999, ORNL’s LandScan Global gridded population model and annual datasets have reported ambient distributions that estimate where populations may be found on average across a 24 h period (Leyk et al., 2019 ). This approach primarily supports consequence assessment of unwarned population related to rapid onset events, such as critical infrastructure failures, conflict-driven violence, or natural disasters, for which the routine activities of the affected population hold critical importance beyond residential locations alone (Dobson et al., 2000 ; Bhaduri et al., 2007; Bhaduri 2008 ; Rose & Bright, 2014 ). New censuses, Earth observation, and digital trace data coupled with advancing technologies and theoretical sophistication have ushered in a period of unprecedented development in predictive modeling of human–environment relationships to place populations more precisely and accurately (Darin et al., 2022 ). Based on the assumption that the visible traces people leave on the surface of the earth—cultivated and impermeable land cover, roads and other infrastructure, buildings in their different shapes and distributions, nighttime lights—accurately reflect the people themselves, ever more granular data about these “reasonable prox[ies] of human activity” have been processed through ever more powerful algorithms to produce population estimates of ever higher resolution in the years since (Dobson et al., 2000 and 2002; Rose and Bright, 2014 ; Leyk et al., 2019 ). Producing a 24 hour ambient population based on these features requires balancing authoritative population totals with country and region-specific characterization of activity spaces. To achieve this balance, LandScan HD combines elements of the top-down and bottom-up taxonomy of gridded population models (Wardrop et al., 2018 ; Weber et al., 2018 ; Leyk et al., 2019 ; Dahmm et al., 2020 ). Top-down methods disaggregate authoritative population counts from coarser spatial resolutions to regular spatial grids (often ranging in resolution from 100 m to 1 km). Although a basic areal interpolation would simply evenly divide the coarser population, ignoring how features in the natural and built environments influence people’s activity spaces, all the currently available gridded population datasets with global coverage account in some way for those constraints, and nearly all do so through dasymetric methods with varying levels of complexity. Much of the development of these top-down methods has focused on refinement of both spatial resolution and model constraints. The Gridded Population of the World (GPW), produced by the Center for International Earth Science Information Network, forms the basis for many top-down gridded population products. GPW is characterized as a “minimally modeled” product that distributes administrative population counts to grid cells (v1: 5 arcminutes/8.5 km, v2-3: 2.5 arcminutes/5 km, v4: 30 arcseconds/1 km) based on available land area per cell, excluding water bodies and protected areas (Lloyd et al., 2017 ; Dahmm et al., 2020 ). The Global Rural Urban Mapping Project (GRUMP), an early enhancement of GPWv3, increases GPW’s spatial resolution of 30 arcseconds, bases urban and rural delineation on factors such as nighttime lights, and uses known settlement locations to disaggregate population counts (Balk et al., 2005a , b ; Balk et al., 2010 ; Leyk et al., 2019 ). The Global Human Settlement Layer Population Grid (GHS-POP) refines GPW’s spatial resolution further to 250 m built-up area presence/absence derived from Landsat imagery (Friere et al., 2016). With continued emphasis on spatial refinement of GPW, the University of Southampton’s WorldPop dataset (Tatem, 2017 ) employs GPWv4 as its population baseline and uses a machine learning (ML) approach to reapportion it to a 3 arcsecond (approximately 100 m) grid using satellite imagery and a suite of geospatial covariates (Leyk et al., 2019 ; Dahmm et al., 2020 ), with adjustments made for country-official or UNPD estimates and building footprint data (Stevens et al., 2015 ; WorldPop, 2024 ). LandScan Global, another top-down method developed in parallel to GPW and available at 30 arcsecond resolution, takes subnational census and administrative data as a starting point and reapportions populations through country-specific multivariable dasymetric models with weights for land use/land cover, infrastructure, urban areas, satellite imagery, and natural features such as slope, elevation, and bodies of water (Dobson et al., 2000 ; Bhaduri et al., 2002 ; Rose & Bright, 2014 ). The resulting ambient population distribution accounts for a greater range of potential activity locations than WorldPop’s more direct residential focus. Bottom-up methods model relationships between geolocated population surveys and environmental covariates. The resulting point estimates of population may be reaggregated to any resolution appropriate to the problem of interest. For example, Weber et al. ( 2018 ) use microcensus surveys—direct measures of building occupancy, typically conducted for small, populated areas—to model population densities at 1 km resolution relative to different settlement types in Nigeria. An added advantage of the bottom-up approach is that is can be applied wherever covariate predictors are available, overcoming challenges for contexts in which data are generally sparse, infrequently collected, or otherwise unreliable (Weber et al., 2018 ). The bottom-up approach does not supersede the top-down approach: the two are complementary and can be used together in a two-step process to fill data gaps (Darin et al., 2022 ) or add details only present in the survey data (Bosco et al., 2017 ). LandScan HD follows the top-down approach of allocating authoritative population counts to subnational locations, but it supports this process using bottom-up estimates from ORNL’s Population Density Tables (PDT) model. PDT’s Bayesian ML approach incorporates observations and subject matter expertise on facility occupancy throughout the world (Stewart et al., 2016 ; Urban et al., 2023 a). Combining volumetric building characteristics (floor area, height) with PDT occupancy estimates for multiple facility types at daytime and nighttime, LandScan HD designates proportional population weights to redistribute population counts to distinct point locations. The resulting estimates are easily aggregated to a 3 arcsecond grid that preserves the official subnational population counts. 3 Data and Methods The core LandScan HD methodology begins with an account of local sociocultural and economic activities (population dynamics) at the building level. Human activities play out across a variety of facility types (e.g., residential, commercial, medical, industrial, transportation)—each with a plausible range of occupancy characteristics—that are distinct from one another based on activity purpose, time of day, and country/region. By combining the expected occupancy characteristics of different facility types with physical characteristics of buildings (floor area, height), the development of localized population weights enables distribution of authoritative population totals to built-up areas indexed by administrative boundaries and grid cell location. 3.1 Data Sources We curate several datasets as inputs to the LandScan HD data stack: Foundational information on populated areas : Building footprint and settlement data that indicate where people reside and where normal patterns of activity occur. Built-up area attribution : Estimation of available interior space that can account for available for human occupancy during residential and nonresidential facility use. Occupancy attribution : Probabilistic estimation of daytime and nighttime average occupant density by facility type (e.g., residential, commercial, industrial, recreational) in units of people per 1000 ft 2 . Population statistics : Proportional reweighting of raw building-level population applied at an appropriate spatial scale to ensure alignment with authoritative population summaries (e.g., census). We default to the administrative level 1 (Adm1) because these are assumed to be self-contained zones where activities primarily take place. 3.1.1 Built-up areas and facility occupancy Identifying built-up areas An organizing principle of LandScan HD is that populations most likely occur in built-up areas, as represented by building footprint and settlement data. Curation of the LandScan HD data stack begins with a catalog of 2D building footprint data from ORNL’s building feature extraction (BFE) product. BFE’s core engine uses convolutional neural network (CNN)-based semantic segmentation algorithms along with a series of advances made for continuous performance improvements (Yang et al., 2018 and 2024 ) as LandScan HD has progressed toward global coverage. The most recent operational-ready version of the BFE model is trained using the Oak Ridge Building Image and Training Label Net (Swan et al., 2024 ), which contains over 130,000 labeled tiles of 500 × 500 pixels drawn from thousands of high-resolution (average resolution 0.47 m/px ) commercial satellite imagery from Maxar WorldView (Maxar, 2024 ) (BGR + near-infrared) covering multiple regions of the globe. Model training uses a boundary-aware loss function as per Yang et al. ( 2018 ) to improve separability of building instances. On average, evaluation at the pixel level across held-out tile samples indicates precision and recall rates on the order of 90%. Image preprocessing for BFE is handled by ORNL’s Parallel Integration and Processing Engine (PIPE), a workflow designed to transform raw satellite imagery into high-resolution, analysis-ready datasets crucial for geospatial intelligence applications, including modules for imagery ingest, orthorectification to correct geometric distortions, image resolution enhancement, and cloud detection, ensuring clear visibility of ground features (Reith et al., 2023 ). When BFE is unavailable, we instead leverage Microsoft and Google building footprints (Sirko et al., 2021; Microsoft, 2024 ). However, each of these datasets can lack coverage in areas of the world where human settlements are known to exist. We identify such areas using the Global Human Settlement Layer (GHSL) Built-S 10 m raster, which provides continuous global measurement of built-up surface area (Pesaresi & Politis, 2023 ). Areas containing nonzero values in Built-S but lacking coverage in the BFE, Microsoft, and Google datasets are converted to a polygonal “no-data” mask. Using the no-data mask, we substitute vectorized Built-S cells as a failsafe for built-up areas (Fig. 2 ). For failsafe cases, Built-S must be normalized to have the same density relationship with building footprints relative to the country of interest (Fig. 3 ). This normalization is accomplished by measuring the total planar area of Built-S for the country of interest that intersects with other built-up area sources and then calculating the ratio between the two sources using Eq. (1): $$\:\begin{array}{c}FGRatio=\:\frac{\sum\:F{B}_{a}}{\sum\:G{B}_{a}},\#\left(1\right)\end{array}$$ where FB a represents the footprint-based built-up area, and GB a represents the Built-S built-up area. Built-up area attribution To support occupancy estimation, we attribute built-up areas by land use (e.g., residential, commercial, industrial) and volumetric properties (area and floor count). Most openly available global land-use data come from OpenStreetMap (OSM), a community-driven mapping platform that provides foundational geospatial data throughout the world (OpenStreetMap Contributors, 2024). OSM land-use polygons are attributed by key:value pairs, providing many categories (e.g., landuse:residential, building:school) that are assignable to built-up areas via spatial overlay. OSM provides extensive user guidance on feature labeling and a robust review cycle for contributions. However, this capability results in large volumes of continually updated land-use tags that complicate matching built-up area with PDT facility use types to estimate occupancy. As a solution, we apply a crosswalk between OSM land-use tags and PDT facility types, accounting for over 110,000 category mappings, to accommodate new or updated OSM features in most situations (Adams et al., 2023 ). For areas lacking coverage and detail in OSM, we supplement land-use attribution using the Urban Tactical Planner (UTP), a program run by the US Army Geospatial Center to facilitate the mapping and display of geospatial information in urban areas (US Army Corps of Engineers Army Geospatial Center, 2013 ) and the Multinational Geospatial Co-production Program (MGCP) (József & Olívia, 2009). Following standards from the Defence Geospatial Information Working Group (DGIWG), these datasets have been continually updated by at least 28 nations since 2003. As with OSM, we crosswalk DGIWG land-use definitions with PDT categories to facilitate occupancy estimation. As a failsafe for missing land-use labels, our default assumption is general residential use. Whereas floor area is an intrinsic property of building footprints, floor count requires manual calculation. We convert building height to floor count using estimates provided by OSM, UTP, and MGCP. If raw building height is available, then we manually estimate floor counts assuming 3 m per floor and then round up as needed to arrive at an integer number of floors. For other cases in which the data are structured to indicate a range of floors, we use the average floor count between the upper and lower bounds. As with land use, we apply a default floor count value (two floors) when height information is completely missing from the source datasets. Facility occupancy To produce bottom-up population (BUP) weights, facility occupancy estimates from PDT (Stewart et al., 2016 ; Urban et al., 2023 a) are combined with building volumetric properties based on use type. PDT provides probabilistic estimates and ranges for daytime and nighttime densities (people per 1000 ft 2 ) for over 60 facility types representing the spectrum of human activity spaces present over much of the globe. Conflating PDT facility classes with building use enables us to transfer the density estimates to each built-up area specific to different activity spaces. PDT overcomes data sparsity using observation models and sociocultural knowledge. Observation models infer occupancy from observable data via rigorously proven theoretical linkages (Stewart et al., 2016 ; Morton, 2013). These observation models can accommodate qualitative information such as culturally informed activity patterns in places like cemeteries where facility use is not typically observed and/or recorded (Lunga et al., 2022 ), or through earth observation to understand measure structural properties of residential facilities (Woody & Frazier, 2023 ). PDT further reinforces these data by encoding sociocultural knowledge from experts and from open-source data on economic practices and cultural norms. These forms of data (knowledge encoding and modeled observational data) are dynamically updated in PDT via a Bayesian ML framework that produces final occupancy estimates and their associated uncertainty (Stewart et al., 2016 ). Built-up areas are attributed by PDT point estimates (default is 50th percentile) based on matching facility classes. This strategy propagates PDT uncertainty from occupant density to population weights via volumetric characteristics, allowing us to estimate the proportion of an area’s total population at each built-up area location. 3.1.2 Administrative boundaries and population totals Population models are often informed by authoritative population counts reported from the country of interest and organized around the boundaries of recognized subnational political units. LandScan HD relies on subnational population statistics available at each country’s Adm1 zones (e.g., states, provinces, regions) to produce an ambient population distribution. Population movements tend to occur at the scale of larger, socially and economically self-contained administrative zones rather than smaller areas like neighborhoods (Brelsford et al., 2022 ). Thus we use the Adm1 level as a common basis for catchment zones of daily routine activities throughout the world. The four key population components of the LandScan HD data stack are (1) authoritative international boundaries, (2) authoritative national population totals, which provide a reference for subnational population distribution, (3) Adm1 boundaries for geolocating buildings, and (4) Adm1-level population statistics for distributing subnational counts to buildings based on daytime/nighttime population weights. National Population Data We aim to use each country’s most current census data to represent population totals, particularly from surveys conducted in the 2010s or 2020s. Where these are not available or are significantly dated, we use statistical estimates and projections instead. To standardize the population counts, we use the United States Census Bureau’s International Database (U.S Census Bureau, 2024) for each country to minimize inconsistencies in collection procedures and to ensure that countries without direct information maintain a dependable source. We apply this count to the national total and prorate and subnational population counts. Administrative Boundaries Producing a complete global catalog of Adm1 zones to match population totals involves harmonizing multiple representations of international political boundaries. Whenever possible, we source subnational boundaries directly from government agencies. Otherwise, we obtain them from a variety of open data sources including OSM, the Humanitarian Data Exchange (United Nations Office for the Coordination of Humanitarian Affairs, 2024 ), and GeoBoundaries (William & Mary geoLab, 2024). We perform comparisons across the various boundary sources to determine the most accurate source. Manual edits may be required to update names or derive feature mapping, and ancillary sources may be used to apply new reported changes in the administrative hierarchy. We clip all sourced spatial boundaries to the Large Scale International Boundaries (LSIB) dataset, which is produced by the US Department of State and provides annual updates on political boundaries. Specifically, we use a modified version of the LSIB that accounts for coastlines and sovereign political regions, provided jointly by the National Geospatial Intelligence Agency’s Office of Geography’s Geographic Boundaries Branch and the Department of State’s Office of the Geographer and Global Issues (US Department of State, 2024 ). As a final verification step, we perform detailed random checks on political boundaries to assure their alignment with topological features such as rivers or manufactured features such as roads and highways. 3.2 Methods 3.2.1 Harmonizing the data stack A challenge associated with the data curation process described in Section 3.1 is limited spatial conformity among the various layers in the data stack. Following Moehl et al. ( 2021 ), we overcome this challenge through a vector analytical framework approach. We represent built-up areas as a table of model leverageable cut-up labeled entities (MOLECULEs) that merges built-up area geometries, attributes, and geographic identifiers (Adm1 and grid cell index). As shown in Fig. 3 , each MOLECULE is a portion of a building footprint intersected by attribution source polygons and boundaries, encoded in the attribute table as a unique combination of data stack features. For cases in which multiple building use types overlap a MOLECULE, we select the labels associated with the smallest area polygons, assuming that smaller polygons represent more specific activity types associated with that space. 3.2.2 Gridded population estimation Converting the data stack to MOLECULEs enables us to distribute total Adm1 population to every available population unit. First, expected populations for each MOLECULE, B , are estimated using Eq. (2): $$\:\begin{array}{c}{B}_{p}={B}_{a}{B}_{f}{B}_{ld},\#\left(2\right)\end{array}$$ where p is the expected population based on footprint area a , floor count f , and PDT density d as defined by the land use l . Then, BUP weights are estimated by normalizing the expected populations of each MOLECULE by their sum. The product of the Adm1 population total P Z and BUP weights (Eq. [3]) gives an updated baseline population estimate, B pop , of subnational population per MOLECULE: $$\:\begin{array}{c}{B}_{pop}=\frac{{B}_{p}}{\sum\:{B}_{p}}\times\:{P}_{z}.\#\left(3\right)\end{array}$$ We perform this procedure for both daytime and nighttime PDT estimates and average the results to obtain a 24 hour ambient population estimate. Because the MOLECULEs are indexed by grid cell, we simply sum the ambient population counts by unique row/column identifier to produce gridded population counts. Finally, the cell estimates are converted to whole numbers using greatest mantissa rounding, which redistributes the total fractional portion of the estimates among the largest integer values (Coleman, 2015 ). 3.2.3 Global implementation Advances in the global coverage of data on the built environment and activity spaces (PDT) have enabled us to build a global data stack to support gridded ambient population estimates (Section 3.1) for every built-up area on Earth. The large volume of information in the global LandScan HD data stack is organized into a centralized PostGIS database. We index the data stack by a 1° resolution tileset, which forms the basis for worldwide generation of MOLECULEs datasets. The tile MOLECULE tables are attached to a parent table per built-up area source (e.g., level 0 and level 1 administrative boundaries). This design supports country-specific model runs by querying the MOLECULEs in tiles (1) that cover the country’s level 0 boundaries and (2) that have a presence flag for that country. Joining the MOLECULEs to country-specific PDT densities by land use then yields BUP weights, which are combined with Adm1 totals and converted to ambient population estimates and aggregated to the 3 arcsecond global grid, as discussed in Section 3.2.2. In rare cases, border grid cells may contain built-up areas from multiple countries. We address these instances by updating the cell counts to the grand sum of the country-level estimates. 4 Demonstration 4.1 LandScan HD Baseline for the Philippines We explore several illustrations of the Philippines as an example result from the global LandScan HD implementation. For this dataset, 98% of the built-up area comes from Microsoft building footprints, and the remaining 2% comes from GHSL Built-S. Figure 4 shows the distribution of area by general category type. Approximately 75% of the density assignments in the model come from the land-use data stack. Most buildings in the Philippines are labeled; therefore, the model accounts for more variation in daytime/nighttime activity types than a purely residential model. However, this example is limited by a lack of height and floor counts (this information was available for less than 1% of the building area), so the model in this case relies heavily on default assumption of two floors per building. In updates to LandScan HD, we plan to address this deficiency and improve coverage of land-use labels by using predictive ML models of missing attributes (Section 5). 4.2 Using LandScan HD to Assess Population Risk We examine a practical application of LandScan HD using ambient estimates of individuals susceptible to flooding in Iloilo City, Philippines, over a 24 hour period. Our aim is to illustrate how ambient population data can be leveraged for urban planning and disaster risk management, providing more actionable insights than estimates derived from residential-only gridded population datasets. We compare LandScan HD to a residential-only population grid, WorldPop, specifically its latest 2020 top-down constrained product, which adjusts for UN World Population Prospects national population estimates (United Nations Department of Economic and Social Affairs, 2022 ; WorldPop, 2024 ). Iloilo City is the principal city and capital of Iloilo province in Western Visayas (Philippines Region VI), a region prone to perennial flooding and severe typhoon risk (Bagsit et al., 2014 ; Subade et al., 2014 ). This scenario is ideal for showcasing the utility of the LandScan HD dataset. In this demonstration, a flood hazard dataset obtained from the Philippine government via the country’s geoportal (Geoportal Philippines, 2024) encompasses nearly the entire municipality. This dataset delineates the flood hazards, scaled at 1:10,000, into three susceptibility categories: low, moderate, and high. For this analysis, we use the extent of the flood hazards layer to compute estimates of the population susceptible to each level of flood risk. 4.2.1 Scenario Methodology We employ both LandScan HD and WorldPop’s gridded population datasets, P , where each grid cell i contains a population count p i . Both datasets have a common spatial resolution of 3 arcseconds. To assess flood risk, the flood hazards dataset delineates three flood zones for Iloilo City, R = { r low, r moderate, r high}. The total population within the boundaries of Iloilo City, P total , is calculated using Eq. (4): $$\:\begin{array}{c}{P}_{\text{t}\text{o}\text{t}\text{a}\text{l}}=\sum\:_{i\in\:B}{p}_{i}\#\left(4\right)\end{array}$$ where B represents all grid cells within the city limits. For each susceptibility zone r s in R , where s is either low, moderate, or high, A s is the subset of grid cells that fall within the boundaries of r s . The population susceptible within each zone is computed using zonal statistics and Eq. (5): $$\:\begin{array}{c}{P}_{\text{s}}=\sum\:_{i\in\:{A}_{\text{s}}}{p}_{i}\#\left(5\right)\end{array}$$ We used ArcGIS Pro version 3.3.0 to match the geographic projection of the flood hazards dataset and with LandScan HD, then employed its zonal statistics tool to compute the population counts for Iloilo City and its susceptibility zones. The percentage of the population susceptible within each flood zone r s , denoted as pct s , is given by Eq. (6): $$\:\begin{array}{c}{pct}_{\text{s}}=\frac{{P}_{\text{s}}}{{P}_{\text{t}\text{o}\text{t}\text{a}\text{l}}}\times\:100\%.\#\left(6\right)\end{array}$$ This analysis presents spatial distributions for both the LandScan HD and WorldPop datasets relative to flood susceptibility exposure, as shown in Fig. 5 and summarized in Table 1 , by populations within each susceptibility zone. Although WorldPop estimates a greater proportion of Iloilo City’s residential population resides in high flood risk zones, more daily activities captured by the LandScan HD ambient population occur in moderate risk zones. However, we also observe localized differences in the concentration of high and moderate risk populations: whereas residential locations estimated by WorldPop tend to concentrate in coastal and riverine areas, the LandScan HD ambient population is more spatially dispersed, including within areas of increased flood risk without a large residential presence. This result is perhaps most notable along the Iloilo River near the geographic center of the city, which features a high density of commercial, business, and academic activities. We also observe a much higher ambient population estimate within Iloilo City than the residential one (659,990 vs. 430,457, respectively), suggesting that people move from adjacent suburban areas within the larger Adm1 zone (Philippines Region VI). Table 1 Ambient and residential-only population model estimates delineated by Iloilo City flood susceptibility, contrasting WorldPop (c. 2020) and LandScan HD (c. 2023). Pop. Count % of Pop. Total Pop. Pop. Count % of Pop. Total Pop. Low Moderate 89662 182927 20.83% 42.50% 430457 150331 319135 22.78% 48.35% 659990 High 157868 36.67% 190524 28.87% 4.3 Construct Validity Assessment In addition to providing a practical application of the LandScan HD global baseline (Section 4.2), we aim to demonstrate the construct validity of the ambient population estimates, defined broadly as their congruence with the construct (24 hour activities) they are designed to represent (Strauss and Smith, 2009 ). Because LandScan HD’s ambient population measures human presence at both residential and nonresidential locations, an important criterion for the actionability of information it provides is to match observational data reflecting collective human activities (infrastructure demand, atmospheric emissions). Furthermore, as examined in Section 4.2, the relationship between LandScan HD and human activity observations should be distinct from a purely residential representation of population. With these criteria in mind, we assess the global baseline LandScan HD release with respect to subnational CO 2 emissions estimates. This prevalent atmospheric greenhouse gas results from a variety of anthropogenic activities, including buildings, industry, agriculture, and transportation (Solomon et al., 2009 ; Crippa et al., 2023 ; Kabir et al., 2023 ). Subnational CO 2 estimates enable assessment of how LandScan HD aligns with the cumulative outcomes of human activities at roughly the scale of urban subdistricts, rural communities, and major transportation and industrial sites. For this exercise, subnational gridded CO 2 emissions available globally were obtained from the European Commission’s Emissions Database for Global Atmospheric Research (EDGAR) (Crippa et al., 2023 ). We compare EDGAR’s association with LandScan HD to residential population estimates provided by WorldPop. Specifically, we assess EDGAR global gridded CO 2 estimates available at 0.1° resolution for 2022, the most recent year in its historical archive. As in Section 4.2, we compare ambient to residential population estimates via the current LandScan HD global baseline and WorldPop’s most recent 2020 constrained (building footprint based), UN-adjusted population counts. We summed LandScan HD and WorldPop population counts to the coarser EDGAR grid cells based on their containment of 3 arcsecond grid cell centers. To account for areas with neither a persistent residential nor ambient population, grid cells with zero population counts were excluded for both LandScan HD and WorldPop, resulting in a total of 3596 EDGAR grid cells for analysis. We compared EDGAR to the aggregated LandScan HD and WorldPop grids using Pearson’s product moment correlation (Sedgwick, 2012 ), which measures the strength of linear association between estimated population and CO 2 emissions (tons released annually). To account for outlying high values in both population and CO 2 emissions totals, which can distort Pearson correlation statistics, the variables were log-transformed as \(\:\text{log}1p\left(x\right)=\text{log}(1+x)\) , such that instances of zero residential population counts and nonzero ambient population counts (e.g., shipping ports with no permanent residential population) were preserved. Although both aggregated population grids feature somewhat strong positive correlation with EDGAR, the ambient population estimates provided by LandScan HD are related more closely to CO 2 emissions ( ρ = 0.761) than those of WorldPop ( ρ = 0.682). Furthermore, the 95% confidence intervals between the two correlation statistics do not overlap (LandScan HD lower-bound ρ = 0.746; WorldPop upper-bound ρ = 0.699), suggesting a distinctly closer association between LandScan HD and CO 2 emissions in the Philippines than that of WorldPop. Figure 6 compares EDGAR with the aggregated population grids. At a high level, the most notable difference between LandScan HD and WorldPop is the extent of unpopulated or very sparsely populated areas in the former. In LandScan HD, these areas tend to contain some ambient population, which corresponds to light to moderate annual emissions in EDGAR and in many cases appears to be linked to human settlements or economic activities captured by the LandScan HD data stack. 5 Limitations and Improvements The LandScan HD baseline model provides a worldwide representation of ambient population at a high spatial resolution of 3 arcseconds, at which only gridded residential populations (WorldPop) were previously available. As demonstrated in Section 4, this approach broadens the scope of data available to researchers and practitioners for making precise population estimates, particularly when human presence and activities are of chief importance. This approach marks the first attempt to scale this methodology globally, and as such, it features several limitations. These largely involve how we address multiple challenges of incompleteness or sparsity of layers in the data stack—particularly with respect to building floor area and activity type—using basic default modeling assumptions. However, these defaults will become less necessary as several new methods for building feature extraction, building land-use, and height prediction become available to fill labeling gaps in the LandScan HD data stack. The first of these issues pertains to detection of built-up areas. Although the problem of missing building footprints in populated areas has been addressed by using GHSL Built-S as a failsafe, an added challenge is that the building footprints themselves also feature errors of omission or commission, including inaccurate spatial geometries (Gonzales, 2023 ). As a solution, we plan to expand ORNL BFE coverage to more areas of the world by developing several AI model improvements that incorporate modern transformer-based architectures (Dosovitskiy et al., 2020 ) as well as foundation models for high-resolution imagery with self-supervised learning schemes for model pretraining (Dias et al., 2023 ). In automation of BFE, ongoing efforts to update PIPE focus on developing additional processing steps (e.g., atmospheric correction), implementing each processing function as its own plug-and-play module, and enhancing workflow scalability and efficiency. Missing or incomplete information about building height and use poses another challenge in the LandScan HD model. In terms of physical characteristics, the lack of a global inventory of building heights currently hinders the development of accurate measures of building floor area in some regions of the world. Inaccurate representation of building floor area in turn leads to disproportionately large or small BUP weights. The current default assumption of two floors for buildings with missing height attribution does not adequately represent local building heights at a global scale. Toward building use attribution, we also acknowledge that the default assumption of unlabeled buildings from OSM and other data layers being residential is likely too broad. Incorrectly categorizing residential/nonresidential uses can significantly compromise the accuracy of BUP estimates by assigning excessive or insufficient occupancy levels. To address these limitations, we plan to build upon recent advancements in ML models for predicting building height and use these models at scale by leveraging the intrinsic relationship between building morphology, function, and urban form (Biljecki et al., 2017 ; Milojevic-Dupont et al., 2020 ; Adams et al., 2023 ; Nachtigall et al., 2023 ; Hartmann et al., 2024 ; Stipek et al., 2024 ). Furthermore, to support more robust characterization of nonresidential activities, particularly in data-poor regions of the world, we next plan to incorporate MapSpace, a high spatial resolution inventory of land-use labels built on road network tessellations (Thakur & Fan, 2021 ; Fan & Thakur, 2023 ). MapSpace extends ORNL’s PlanetSense platform, a comprehensive catalog of points of interest (Thakur et al., 2015 and 2016 ) leveraging a semantic labeling system, SONET, to conflate these categories with PDT facility types (Palumbo et al., 2019 ; Fan et al., 2023 ). A final challenge for LandScan HD is quantifying uncertainty related to modeling decisions. Enhancements to the data stack may improve modeling decisions, but their effects on the end product still need direct evaluation. The LandScan HD model’s deterministic nature eliminates traditional uncertainty measurements because it consistently produces the same output from identical inputs without variability. Nevertheless, the individual inputs of the LandScan HD model contain various levels of uncertainty that can be quantified. Examples include assigning building commercial use when multiple land-use labels are present or labeling a building with three floors when it may plausibly have four. Thus, future work will involve providing measures of confidence in our population estimates by developing a confidence index that is driven by the uncertainties of the data used to produce the population estimate (Section 3.1). 6 Conclusion and Outlook To our knowledge, the global LandScan HD baseline model provides the first 24 hour ambient population for the world at 3 arcsecond (roughly 90 m) spatial resolution. The methodology detailed in this paper requires curating and fusing large volumes of data, including population and occupancy statistics as well as the physical and activity-specific properties of buildings, in development of a scalable workflow. Our approach demonstrates the importance of ambient population estimates (combined residential/nonresidential activities) in terms of risk to unwarned populations in the face of climate-driven hazards (Section 4.2), which may be extended to a variety of related human security challenges. Furthermore, the alignment of LandScan HD with the collective effects of human activities (CO 2 emissions) in the Philippines (Section 4.3) demonstrates the underlying model’s congruence with 24 hour population dynamics. In each of these illustrations, LandScan HD provides insights about the location and presence of the unwarned population in ways that purely residential/nighttime gridded populations (WorldPop) that are often used in practice do not. In addition to enhancing the LandScan HD model and developing a global confidence index for quantifying uncertainty in our population estimates (Section 5), automation of the LandScan HD workflow can support rapid population estimates in the face of global humanitarian crises (Urban et al., 2023 b). To support this objective, we plan to introduce techniques such as building instance segmentation/counting, building damage assessment for conflict/natural disaster scenarios (Urban et al., 2023 b), and feature extraction for informal settlements (e.g., refugee camps). The development of a global LandScan HD baseline model also opens possibilities for exploring how human activity spaces align with sociodemographic and livelihood characteristics throughout the world to address problems like climate risk and vulnerability, spatial accessibility to vital resources, and infrastructure use. Although methods for exploring such problems have been established for the United States (Tuccillo & Gaboardi, 2022 ; Tuccillo et al., 2023 ), they require further development for other regions of the world, particularly low- and middle-income countries (LMICs), for which sparse population data often inhibit decisions about emergency and resource planning. To address this shortcoming, we are now using LandScan HD to build out methods for characterizing populations in LMICs, expanding upon an approach established by Urban et. al ( 2022 ) to downscale social, demographic, and economic information from global microdata sources such as the Demographic and Health Surveys to built-up areas. Declarations Acknowledgements The authors would like to thank Edward Bright, Amy Rose, Jacob McKee, and Melanie Laverdiere for their foundational contributions to the Landscan HD project. References Adams, D.S., Hauser, T., and Moehl, J. (2023). Decoding Ethiopian Abodes: Towards Classifying Buildings by Occupancy Type Using Footprint Morphology. Proceedings of the 2023 International Conference on Machine Learning and Applications (ICMLA) (pp. 210-217). https://doi.org/10.1109/ICMLA58977.2023.00037 Archila Bustos, M.F., Hall, O., Niedomysl, T., Ernstson, U. (2020). A pixel level evaluation of five multitemporal global gridded population datasets: a case study in sweden, 1990–2015. Population and Environment , 42, 255–277. Bagsit, F.U., Suyo, J.G.B., Subade, R.F., Basco, J., et al. (2014). Do adaptation and coping mechanisms to extreme climate events differ by gender? the case of flood-affected households in dumangas, iloilo, philippines. Asian Fisheries Science , 25, 111–118. Balk, D., Brickman, M., Anderson, B., Pozzi, F., & Yetman, G. (2005b). Estimates of future global population distribution to 2015. Mapping global urban and rural population distributions . Food and Agriculture Organization of the United Nations (pp. 53–73). Balk, D., Pozzi, F., Yetman, G., Deichmann, U., & Nelson, A. (2005a). The distribution of people and the dimension of place: methodologies to improve the global estimation of urban extents. International Society for Photogrammetry and Remote Sensing Proceedings of the Urban Remote Sensing Conference (pp. 14–16). Balk, D., Yetman, G., & De Sherbinin, A. (2010). Construction of gridded population and poverty data sets from different data sources. Proceedings of European Forum for Geostatistics Conference (pp. 12–20). Bhaduri, B. (2008). Population Distribution During the Day. In: Shekhar, S. and Hui Xiong, (Eds). Encyclopedia of GIS (pp. 1377). Springer-Verlag. Bhaduri, B., Bright, E., Coleman, P., and Dobson. J. (2002). LandScan: Locating People is What Matters. Geoinformatics , 5(2): 4–37. Biljecki, F., Ledoux, H., Stoter, J. (2017). Generating 3d city models without elevation data. Computers, Environment and Urban Systems , 64, 1–18. https://doi.org/10. 1016/j.compenvurbsys.2017.01.001 Bosco, C., Alegana, V., Bird, T., Pezzulo, C., Bengtsson, L., Sorichetta, A., Steele, J., Hornby, G., Ruktanonchai, C., Ruktanonchai, N., et al. (2017). Exploring the high resolution mapping of gender-disaggregated development indicators. Journal of The Royal Society Interface, 14(129). Brelsford, C., Moehl, J., Weber, E., Sparks, K., Rose, A. (2022). Improving the LandScan USA non-obligate population estimate (nope). University of California Santa Barbara Spatial Data Science Symposium 2022 Short Paper Proceedings . https://doi. org/10.25436/E2X307 Cheriyadat, A., Bright, E., Bhaduri, B, and Potere, D. (2007). Mapping of Settlements in High Resolution Satellite Imagery Using High Performance Computing. GeoJournal , 69(1–2), 119–129. Coleman, Charles (2015). SAS® Macros for Constraining Arrays of Numbers . Southeast SAS Users Group 2015 Proceedings (pp. 1-9). https://www2.census.gov/library/working-papers/2015/econ/2015-coleman.pdf Crippa, M., Guizzardi, D., Schaaf, E., Monforti-Ferrario, F., Quadrelli, R., Risquez Martin, A., Rossi, S., Vignati, E., Muntean, M., Brandao De Melo, J., Oom, D., Pagani, F., Banja, M., Taghavi, Moharamli, P., Köykkä, J., Grassi, G., Branco, A., San-Miguel, J. (2023). GHG emissions of all world countries – 2023. Publications Office of the European Union . https://doi.org/10.2760/953322 Dahmm, H., Rabiee, M., Espey, J., Adamo, S. (2020). Leaving no one off the map: a guide for gridded population data for sustainable development. SDSN TReNDS. https://files.unsdsn.org/Leaving%2Bno%2Bone%2Boff%2Bthe%2Bmap-4.pdf Darin, E., Ku´epi´e, M., Bassinga, H., Boo, G., Tatem, A.J., Reeve, P. (2022). The population seen from space: When satellite images come to the rescue of the census. Population , 77(3), 437–464. Dias, P., Potnis, A., Guggilam, S., Yang, L., Tsaris, A., Medeiros, H., Lunga, D. (2023). An agenda for multimodal foundation models for earth observation. In IGARSS 2023 IEEE International Geoscience and Remote Sensing Symposium (pp. 1237– 1240). Dobson, J.E., Bright, E.A., Coleman, P.R., Durfee, R.C., Worley, B.A. (2000). LandScan: a global population database for estimating populations at risk. Photogrammetric engineering and remote sensing , 66(7), 849–857. Dobson, J., Bright, E., Coleman, P., and Bhaduri, B., (2003). LandScan: A Global population database for estimating population at risk. In Victor Mesev (Ed.) Remotely-Sensed Cities (pp. 378). Taylor and Francis. Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv:2010.11929 Fan, J., Bentley, J., Thakur, G.M. (2023). Sonet++: A knowledge graph of geographic categories based on osm tag representation. https://www.osti.gov/servlets/purl/2000381 Fan, J., Thakur, G. (2023). Towards poi-based large-scale land use modeling: spatial scale, semantic granularity, and geographic context. International Journal of Digital Earth , 16(1), 430–445. Freire S., MacManus K., Pesaresi M., Doxsey-Whitfield E., Mills J. (2016). Development of new open and free multi-temporal global population grids at 250 m resolution. Geospatial Data in a Changing World; Association of Geographic Information Laboratories in Europe (AGILE) 2016 . Gonzales, J.J. (2023). Building-level comparison of microsoft and google open building footprints datasets. In Beecham, R., Long, J.A., Smith, D., Zhao, Q., Wise, S. (eds.) Proceedings of the 12th International Conference on Geographic Information Science (GIScience 2023) (pp. 35:1 – 35:6). https://doi.org/10.4230/LIPIcs.GIScience.2023.35 Hartmann, A., Behnisch, M., Hecht, R., Meinel, G. (2024). Prediction of residential and nonresidential building usage in Germany based on a novel nationwide reference data set. Environment and Planning B: Urban Analytics and City Science , 51(1), 216–233. https://doi.org/10.1177/23998083231175680 József, C., & Olıvia, M. (2009). Multinational Geospatial Co-production Program (Mgcp). Geodezia Es Kartografia , 61, 15–18. Kabir, M., Habiba, U.E., Khan, W., Shah, A., Rahim, S., Patricio, R., Ali, L., Shafiq, M., et al. (2023). Climate change due to increasing concentration of carbon dioxide and its impacts on environment in 21st century; a mini review. Journal of King Saud University-Science , 35(5). Kugler, T.A., Grace, K., Wrathall, D.J., Sherbinin, A., Van Riper, D., Aubrecht, C., Comer, D., Adamo, S.B., Cervone, G., Engstrom, R., et al. (2019). People and pixels 20 years later: the current data landscape and research trends blending population and environmental data. Population and Environment , 41, 209–234. Leyk, S., Gaughan, A.E., Adamo, S.B., Sherbinin, A., Balk, D., Freire, S., Rose, A., Stevens, F.R., Blankespoor, B., Frye, C., Comenetz, J., Sorichetta, A., MacManus, K., Pistolesi, L., Levy, M., Tatem, A.J., Pesaresi, M. (2019). The spatial allocation of population: a review of large-scale gridded population data products and their fitness for use. Earth System Science Data , 11(3), 1385–1409. https://doi.org/10. 5194/essd-11-1385-2019 Lloyd, C. T., Sorichetta, A., & Tatem, A. J. (2017). High resolution global gridded data for use in population studies. Scientific data , 4(1), 1–17. Lunga, D., Dhamdhere, R., Walters, S., Bragg, L., Makkar, N., Urban, M. (2022). Learning to count grave sites for cemetery observation models with satellite imagery. IEEE Geoscience and Remote Sensing Letters, 19, 1–5. https://doi.org/10.1109/ LGRS.2020.3022328 Maxar (2024). WorldView-3. https://resources.maxar.com/data-sheets/worldview-3. Microsoft (2024). Global ML Building Footprints [Data set]. https://github.com/ microsoft/GlobalMLBuildingFootprints Milojevic-Dupont, N., Hans, N., Kaack, L.H., Zumwald, M., Andrieux, F., Barros Soares, D., Lohrey, S., Pichler, P.-P., Creutzig, F. (2020). Learning from urban form to predict building heights. PLOS ONE , 15(12), 1–22. https://doi.org/10.1371/ journal.pone.0242010 Moehl, J., Weber, E., McKee, J. (2021). A vector analytical framework for population modeling. In The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences (pp. 103–108). https://doi.org/10.5194/ isprs-archives-XLVI-4-W2-2021-103-2021 Nachtigall, F., Milojevic-Dupont, N., Wagner, F., Creutzig, F. (2023). Predicting building age from urban form at large scale. Computers, Environment and Urban Systems , 105. https://doi.org/10.1016/j.compenvurbsys.2023.102010 OpenStreetMap Contributors (2024). Planet OSM [Data set]. https://planet.osm.org Palumbo, R., Thompson, L., Thakur, G. (2019). Sonet: a semantic ontological network graph for managing points of interest data heterogeneity. In Proceedings of the 3rd ACM SIGSPATIAL International Workshop on Geospatial Humanities (pp. 1–6). Pesaresi, M., Politis, P. (2023). GHS-BUILT-S R2023A - GHS built-up surface grid, derived from Sentinel2 composite and Landsat, multitemporal (1975-2030) [Data set]. http://data.europa.eu/89h/9f06f36f-4b11-47ec-abb0-4f8b7b1d72ea,doi:10.2905/9F06F36F-4B11-47EC-ABB0-4F8B7B1D72EA Pörtner, H.-O., Roberts, D.C., Adams, H., Adelekan, I., Adler, C., Adrian, R., Aldunce, P., Ali, E., Ara Begum, R., Bednar-Friedl, B. (2022). Technical summary. IPCC Sixth Assessment Report (pp. 37–118). Reith, A., McKee, J., Rose, A., Laverdiere, M., Swan, B., Hughes, D., ... & Lunga, D. (2023). Providing geospatial intelligence through a scalable imagery pipeline. In Advances in Scalable and Intelligent Geospatial Analytics: Challenges and Applications (pp. 153-168). CRC Press. Rose, A.N., Bright, E. (2014). The LandScan Global Population Distribution Project: Current State of the Art and Prospective Innovation. Proceedings of the Population Association of America 2014 Annual Meeting . https://paa2014.populationassociation.org/papers/ 143242 Sedgwick, P. (2012). Pearson’s correlation coefficient. BMJ , 345, e4483.https://doi.org/10.1136/bmj.e4483 Sirko, W., Kashubin, S., Ritter, M., Annkah, A., Bouchareb, Y.S.E., Dauphin, Y.N., Keysers, D., Neumann, M., Ciss´e, M., Quinn, J. (2023). Continental-scale building detection from high resolution satellite imagery. https://doi.org/10.48550/arXiv.2107.12283 Solomon, S., Plattner, G.-K., Knutti, R., Friedlingstein, P. (2009). Irreversible climate change due to carbon dioxide emissions. Proceedings of the national academy of sciences , 106(6), 1704–1709. Stevens, F.R., Gaughan, A.E., Linard, C., Tatem, A.J. (2015). Disaggregating census data for population mapping using random forests with remotely-sensed and ancillary data. PloS one, 10(2). Stewart, R., Urban, M., Duchscherer, S., Kaufman, J., Morton, A., Thakur, G., Piburn, J., Moehl, J. (2016). A bayesian machine learning model for estimating building occupancy from open source data. Natural Hazards , 81(3), 1929–1956. https://doi.org/ 10.1007/s11069-016-2164-9 Stipek, C., Hauser, T., Adams, D., Epting, J., Brelsford, C., Moehl, J., ... & Stewart, R. (2024). Inferring building height from footprint morphology data. Scientific Reports , 14 (1). Strauss, M.E., Smith, G.T. (2009). Construct validity: Advances in theory and methodology. Annual review of clinical psychology, 5, 1–25. Subade, R., Suyo, J., Ebay, J., Lozada, E., Dator-Bercilla, J., Tionko, A., et al. (2014). Adaptation and coping strategies to extreme climate conditions: impact of Typhoon Frank in selected sites in Iloilo, Philippines. http://www.eepseapartners.org/post-2476 Swan, B., Pyle, J., Roddy, D., Rose, A., Yang, H.L., Laverdiere, M. (2024). ORBITaL-Net training library for building extraction [Data set]. https://doi.org/10.25452/figshare. plus.25282225.v1 Tatem, A.J. (2017). Worldpop, open data for spatial demography. Scientific Data, 4(1), 1–4. Thakur, G.S., Bhaduri, B.L., Piburn, J.O., Sims, K.M., Stewart, R.N., Urban, M.L. (2015). Planetsense: a real-time streaming and spatio-temporal analytics platform for gathering geo-spatial intelligence from open source data. In Proceedings of the 23rd SIGSPATIAL International Conference on Advances in Geographic Information Systems (pp. 1–4). Thakur, G., Fan, J. (2021). Mapspace: POI-based multi-scale global land use modeling. In GIScience 2023 Short Paper Proceedings. https://doi.org/10.25436/E2Z59N Thakur, G.S., Sparks, K., Li, R., Stewart, R.N., Urban, M.L. (2016). Demonstrating PlanetSense: gathering geo-spatial intelligence from crowd-sourced and social-media data. In: Proceedings of the 24th ACM SIGSPATIAL International Conference on Advances in Geographic Information Systems (pp. 1–4). Tuccillo, J.V., Gaboardi, J.D. (2022). Likeness: a toolkit for connecting the social fabric of place to human dynamics. In Agarwal, M., Calloway, C., Niederhut, D., Shupe, D. (eds.) Proceedings of the 21 st Python in Science Conference (pp. 125 – 135). https: //doi.org/110.25080/majora-212e5952-014 Tuccillo, J., Stewart, R., Rose, A., Trombley, N., Moehl, J., Nagle, N., Bhaduri, B. (2023). UrbanPop: A spatial microsimulation framework for exploring demographic influences on human dynamics. Applied Geography, 151. https://doi.org/10.1016/j.apgeog.2022.102844 United Nations Department of Economic and Social Affairs (2022). Methodology of the United Nations population estimates and projections. https: //population.un.org/wpp/Publications/Files/WPP2022 Methodology.pdf United Nations Office for the Coordination of Humanitarian Affairs (2024). Humanitarian Data Exchange v1.85.9 [Data set]. https://data.humdata.org/ Urban, M., Moehl, J., Dias, P., Tuccillo, J., Reith, A., Sims, K., Walters, S., Arndt, J., Potnis, A., Lunga, D. (2023). Towards rapid response updates of populations at risk. Proceedings of the IEEE International Geoscience and Remote Sensing Symposium (pp. 907–910). https://doi.org/10.1109/IGARSS52108.2023. 10282319 Urban, M., Stewart, R., Basford, S., Palmer, Z., Kaufman, J. (2023). Estimating building occupancy: a machine learning system for day, night, and episodic events. Natural Hazards , 116(2), 2417–2436. https://doi.org/10.1007/s11069-022-05772-3 Urban, M., Tuccillo, J., Frazier, T., Cunningham, A., Fan, J., Dias, P., Arndt, J., Bowman, J., Gaboardi, J. (2022). New Insights into the Tightly Coupled Social Fabric of the Built Environment [Conference presentation]. American Geophysical Union 2022 Annual Meeting. US Army Corps of Engineers Army Geospatial Center (2013). Urban Tactical Planner Factsheet. https://www.agc.army.mil/Media/Fact-Sheets/Fact-Sheet-Article-View/Article/480931/urban-tactical-planner US Census Bureau (2024). International Database (IDB) [Data set]. https://www.census.gov/programs-surveys/international-programs/about/idb.html US Department of State (2024). Large Scale International Boundaries [Data set]. https://geodata.state.gov/geonetwork/srv/eng/catalog.search#/metadata/3bdb81a0-c1b9-439a-a0b1-85dac30c59b2 Vijayaraj, V., Bright E., and Bhaduri, B., (2007). High Resolution Urban Feature Extraction for Global Population Mapping using High Performance Computing, Proceedings of the IEEE International geosciences and remote sensing symposium (IGARSS) 2007 . https://doi.org/10.1109/IGARSS10946.2007 Vijayaraj, V., Bright E., and Bhaduri, B. (2008). Rapid Damage Assessment from High Resolution Imagery, Proceedings of the IEEE International Geosciences and Remote Sensing Symposium (IGARSS) 2008 . https://doi.org/10.1109/IGARSS10663.2008 Wardrop, N., Jochem, W., Bird, T., Chamberlain, H., Clarke, D., Kerr, D., Bengtsson, L., Juran, S., Seaman, V., Tatem, A. (2018). Spatially disaggregated population estimates in the absence of national population and housing census data. Proceedings of the National Academy of Sciences , 115(14), 3529–3537. Weber, E.M., Seaman, V.Y., Stewart, R.N., Bird, T.J., Tatem, A.J., McKee, J.J., Bhaduri, B.L., Moehl, J.J., Reith, A.E. (2018). Census-independent population mapping in northern Nigeria. Remote Sensing of Environment , 204, 786–798. https: //doi.org/10.1016/j.rse.2017.09.024 William & Mary geoLab (2024). geoBoundaries [Data set]. https://www.geoboundaries.org. Woody, C., Frazier, T. (2023). Waffle Homes: Utilizing Aerial Imagery of Unfinished Buildings to Determine Average Room Size. In Beecham, R., Long, J.A., Smith, D., Zhao, Q., Wise, S. (eds.) Proceedings of the 12th International Conference on Geographic Information Science (GIScience 2023) (pp. 85:1–85:6). https://doi.org/10.4230/LIPIcs.GIScience.2023. 85. WorldPop (2024). Top-down estimation modelling: Constrained vs Unconstrained. https://www.worldpop.org/methods/top_down_constrained_vs_unconstrained Yang, H.L., Yuan, J., Lunga, D., Laverdiere, M., Rose, A., Bhaduri, B. (2018). Building extraction at scale using convolutional neural network: Mapping of the united states. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing , 11(8), 2600–2614. https://doi.org/10.1109/JSTARS.2018.2835377 Yang, H.L., Laverdiere, M., Hauser, T., Swan, B., Schmidt, E., Moehl, J., Reith, A., Adams, D., Morris, B., McKee, J., Whitehead, M., Tuttle, M. (2024). A baseline structure inventory with critical attribution for the us and its territories. Scientific Data , 11. https://doi.org/10.1038/s41597-024-03219-x Yang, L., Varma, L., Neunsinger, L., Lunga, D. (2023). Providing Geospatial Intelligence Through a Scalabale Imagery Pipeline. In Advances in Scalable and Intelligent Geospatial Analytics: Challenges and Applications . CRC Press. https: //doi.org/10.1201/9781003270928-11 Additional Declarations The authors declare no competing interests. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6396722","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Method Article","associatedPublications":[],"authors":[{"id":440213980,"identity":"fcbb446c-8e4b-45a7-b03a-fb836e09a5ed","order_by":0,"name":"Joseph V. Tuccillo","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAyUlEQVRIiWNgGAWjYLCChAILBsYG5gNApgVh1TxgLQYSQC1sCUCmBJFaGAxAKnkMiNNiL5H88MMDAwk55hk5Hx9X7pBgkHfvMcBvi0SasQTQYcaMM3I3G549I8FgeOYMIS05bCC/JDb2nN0m2dgG1DIjLYFYLWeekaqlvYcNrEVeIvkAfi1nnkH90t5mbAjUwmPAcxi/Fvb25Icff1TYyBk2Mz982NhmIyff3tiAVwscGELV8RjgtwMJyMMZRNoxCkbBKBgFIwcAAJcWPNptAkW7AAAAAElFTkSuQmCC","orcid":"https://orcid.org/0000-0002-5930-0943","institution":"Oak Ridge National Laboratory","correspondingAuthor":true,"prefix":"","firstName":"Joseph","middleName":"V.","lastName":"Tuccillo","suffix":""},{"id":440213981,"identity":"faa40ffc-8d22-46e3-aa1b-e3bab0e02a7c","order_by":1,"name":"Jessica Moehl","email":"","orcid":"https://orcid.org/0000-0001-9579-2562","institution":"Oak Ridge National Laboratory","correspondingAuthor":false,"prefix":"","firstName":"Jessica","middleName":"","lastName":"Moehl","suffix":""},{"id":440213982,"identity":"458c61d8-2e48-4894-a4dd-204587d037aa","order_by":2,"name":"Daniel Adams","email":"","orcid":"https://orcid.org/0000-0001-9695-0577","institution":"Oak Ridge National Laboratory","correspondingAuthor":false,"prefix":"","firstName":"Daniel","middleName":"","lastName":"Adams","suffix":""},{"id":440213983,"identity":"1b2da5a4-a8b6-4a5d-a4da-dc6394652f25","order_by":3,"name":"Angela R. Cunningham","email":"","orcid":"https://orcid.org/0000-0001-7379-5383","institution":"Oak Ridge National Laboratory","correspondingAuthor":false,"prefix":"","firstName":"Angela","middleName":"R.","lastName":"Cunningham","suffix":""},{"id":440213984,"identity":"b1256ac9-14ee-4ad1-b6da-8c5ec8ee5c9d","order_by":4,"name":"Marie Urban","email":"","orcid":"https://orcid.org/0000-0001-9571-832X","institution":"Oak Ridge National Laboratory","correspondingAuthor":false,"prefix":"","firstName":"Marie","middleName":"","lastName":"Urban","suffix":""},{"id":440213985,"identity":"1526f1f1-d15f-4fae-b514-03beb4d2cc63","order_by":5,"name":"Sarah Walters","email":"","orcid":"https://orcid.org/0000-0002-3318-8543","institution":"Oak Ridge National Laboratory","correspondingAuthor":false,"prefix":"","firstName":"Sarah","middleName":"","lastName":"Walters","suffix":""},{"id":440213986,"identity":"ee8814fe-a229-4e95-8d2d-9bc96673602f","order_by":6,"name":"Carson Woody","email":"","orcid":"https://orcid.org/0000-0003-2365-1159","institution":"Oak Ridge National Laboratory","correspondingAuthor":false,"prefix":"","firstName":"Carson","middleName":"","lastName":"Woody","suffix":""},{"id":440213987,"identity":"eb5427fc-af44-4845-9d54-80e325777e6b","order_by":7,"name":"Andrew Reith","email":"","orcid":"https://orcid.org/0000-0002-6205-8473","institution":"Oak Ridge National Laboratory","correspondingAuthor":false,"prefix":"","firstName":"Andrew","middleName":"","lastName":"Reith","suffix":""},{"id":440213988,"identity":"465249fc-5616-47f0-b073-e2bc91f007d7","order_by":8,"name":"Jason Kaufman","email":"","orcid":"https://orcid.org/0000-0001-5482-6914","institution":"Oak Ridge National Laboratory","correspondingAuthor":false,"prefix":"","firstName":"Jason","middleName":"","lastName":"Kaufman","suffix":""},{"id":440213989,"identity":"fe504293-3f95-4083-b6dd-f6900546419c","order_by":9,"name":"Justin Epting","email":"","orcid":"https://orcid.org/0009-0004-9305-4813","institution":"Bechtel","correspondingAuthor":false,"prefix":"","firstName":"Justin","middleName":"","lastName":"Epting","suffix":""},{"id":440213990,"identity":"89856245-4101-4321-9a32-2a9839b6311f","order_by":10,"name":"Jack Gonzales","email":"","orcid":"https://orcid.org/0000-0002-9343-438X","institution":"Oak Ridge National Laboratory","correspondingAuthor":false,"prefix":"","firstName":"Jack","middleName":"","lastName":"Gonzales","suffix":""},{"id":440213991,"identity":"8692fdcf-0aeb-4b2c-afa7-d7cf2e8de7e5","order_by":11,"name":"Philipe Ambrozio Dias","email":"","orcid":"https://orcid.org/0000-0001-9427-7112","institution":"Oak Ridge National Laboratory","correspondingAuthor":false,"prefix":"","firstName":"Philipe","middleName":"Ambrozio","lastName":"Dias","suffix":""},{"id":440213992,"identity":"a4ef8c71-cffc-4ea6-88c2-20c88360c0f2","order_by":12,"name":"Cecilia Clark","email":"","orcid":"https://orcid.org/0000-0002-4826-2152","institution":"Oak Ridge National Laboratory","correspondingAuthor":false,"prefix":"","firstName":"Cecilia","middleName":"","lastName":"Clark","suffix":""},{"id":440213993,"identity":"1b7df12c-cee5-4b4a-a21f-d2f2af25ef46","order_by":13,"name":"Hsuihan Lexie Yang","email":"","orcid":"https://orcid.org/0000-0003-2252-6778","institution":"Oak Ridge National Laboratory","correspondingAuthor":false,"prefix":"","firstName":"Hsuihan","middleName":"Lexie","lastName":"Yang","suffix":""},{"id":440213994,"identity":"7166cd28-2735-4ba8-b26f-67a2ec0c6e49","order_by":14,"name":"Robert Stewart","email":"","orcid":"https://orcid.org/0000-0002-8186-7559","institution":"Oak Ridge National Laboratory","correspondingAuthor":false,"prefix":"","firstName":"Robert","middleName":"","lastName":"Stewart","suffix":""},{"id":440213995,"identity":"7771b4e2-bcd8-4646-90dd-0f676e8faf1e","order_by":15,"name":"Dalton Lunga","email":"","orcid":"https://orcid.org/0000-0003-0054-1141","institution":"Oak Ridge National Laboratory","correspondingAuthor":false,"prefix":"","firstName":"Dalton","middleName":"","lastName":"Lunga","suffix":""},{"id":440213996,"identity":"825a4a42-86e3-425b-879b-2b72c3d8013b","order_by":16,"name":"Eric Weber","email":"","orcid":"https://orcid.org/0000-0002-0098-3874","institution":"HDR, Inc.","correspondingAuthor":false,"prefix":"","firstName":"Eric","middleName":"","lastName":"Weber","suffix":""},{"id":440213997,"identity":"41437f69-7e7f-441e-8eea-667c65566532","order_by":17,"name":"Budhendra Bhaduri","email":"","orcid":"https://orcid.org/0000-0003-1555-1377","institution":"Oak Ridge National Laboratory","correspondingAuthor":false,"prefix":"","firstName":"Budhendra","middleName":"","lastName":"Bhaduri","suffix":""}],"badges":[],"createdAt":"2025-04-07 18:14:51","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-6396722/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6396722/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":80200105,"identity":"840542d4-c299-4822-97e6-6f62ead7363f","added_by":"auto","created_at":"2025-04-09 06:36:36","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":406465,"visible":true,"origin":"","legend":"\u003cp\u003eLandScan HD data stack, demonstrating inputs for the dual bottom-up weighting and top-down population components of the methodology.\u003c/p\u003e","description":"","filename":"figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-6396722/v1/f6f25df28b60eb13c82327e4.png"},{"id":80200102,"identity":"4792abd6-6bde-4c26-af71-892b38b20cdf","added_by":"auto","created_at":"2025-04-09 06:36:36","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":54989,"visible":true,"origin":"","legend":"\u003cp\u003eExample scenario with two building footprint layers A and B, along with failsafe layer C. Model leverageable cut-up labeled entities (MOLECULEs) for all three built-up area sources are intersected with each mask layer and attributed such that MOLECULEs can be selected and combined in order of preference into a final complete built layer.\u003c/p\u003e","description":"","filename":"figure2.png","url":"https://assets-eu.researchsquare.com/files/rs-6396722/v1/e4a5f1bd77edc9c1fac18861.png"},{"id":80200103,"identity":"14106418-0305-45f5-814c-4faa3d5c7274","added_by":"auto","created_at":"2025-04-09 06:36:36","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":49792,"visible":true,"origin":"","legend":"\u003cp\u003eTo normalize GHSL area to a given footprint layer, we calculate a ratio of the total area of footprint layer A and failsafe layer C only where both are present.\u003c/p\u003e","description":"","filename":"figure3.png","url":"https://assets-eu.researchsquare.com/files/rs-6396722/v1/e3a6106b405cf8321aa026e9.png"},{"id":80200108,"identity":"7897ec41-3726-4520-8314-d01a18964edc","added_by":"auto","created_at":"2025-04-09 06:36:36","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":4479530,"visible":true,"origin":"","legend":"\u003cp\u003eEach building (A) is cut by polygonal source features (OSM, UTP, MGCP), administrative boundaries, and gridlines (B), resulting in a table of Model leverageable cut-up labeled entities (MOLECULEs) (C). An individual MOLECULE (D) has attribution for each information source (E).\u003c/p\u003e","description":"","filename":"figure4.png","url":"https://assets-eu.researchsquare.com/files/rs-6396722/v1/e47a84d817d0489e9840e024.png"},{"id":80200107,"identity":"c031af60-1b39-4902-b641-140cc851a9c6","added_by":"auto","created_at":"2025-04-09 06:36:36","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":34343,"visible":true,"origin":"","legend":"\u003cp\u003eLand-use labels from the LandScan HD data stack provide over 75% of the labeling for built-up area in the Philippines. The remaining 25% of unlabeled buildings are assumed to be general residential.\u003c/p\u003e","description":"","filename":"figure5.png","url":"https://assets-eu.researchsquare.com/files/rs-6396722/v1/7e10b35c1aa64c2f621f78f4.png"},{"id":80200111,"identity":"c650b6ef-cdb7-43f7-a3fb-6625d9b7d4a9","added_by":"auto","created_at":"2025-04-09 06:36:37","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":4309453,"visible":true,"origin":"","legend":"\u003cp\u003eLandScan HD and WorldPop gridded population distributions delineated by flood susceptibility in Iloilo City, Philippines. The populations that fall within each flood susceptibility zone are color coded by the susceptibility level: red for high susceptibility, purple for moderate susceptibility, and blue for low susceptibility. The hue of the color and the height correspond to a higher population estimate. Service layer credit to ESRI, TomTom, FAO, NOAA, USGS, Foursquare, NASA, Government of Philippines.\u003c/p\u003e","description":"","filename":"figure6.png","url":"https://assets-eu.researchsquare.com/files/rs-6396722/v1/e28a0070738178fcf84a3b95.png"},{"id":80200115,"identity":"d8899d2c-be28-49f2-84f8-3d5677b97464","added_by":"auto","created_at":"2025-04-09 06:36:37","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":490901,"visible":true,"origin":"","legend":"\u003cp\u003eComparison of EDGAR gridded CO\u003csub\u003e2\u003c/sub\u003e emissions for 2022 with the most current LandScan HD (ambient) population and WorldPop (residential) population products for the Philippines.\u003c/p\u003e","description":"","filename":"figure7.png","url":"https://assets-eu.researchsquare.com/files/rs-6396722/v1/a8997c440dc3d19d46b879a1.png"},{"id":80201386,"identity":"de575ac6-01b0-4dbe-af21-da304bc76f9f","added_by":"auto","created_at":"2025-04-09 06:52:42","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":9720655,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6396722/v1/1a6fcced-11d3-4626-a8f8-5e2e6467efe0.pdf"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003eLandScan HD: A High-Resolution Gridded Ambient Population Methodology for the World\u003c/p\u003e","fulltext":[{"header":"1\tIntroduction","content":"\u003cp\u003eMeasuring the footprint of human activities on Earth is fundamental to addressing global challenges of climate-driven and technological hazards, conflict, demand on infrastructure, and resource scarcity (P\u0026ouml;rtner et al., 2022). Central to this task are gridded population datasets, which estimate human presence at high spatial resolution based on the suitability of built-up areas for inhabitation and overcome challenges of limited data availability when estimating global population from census data alone (Bhaduri et al., 2002; Dobson et al., 2000 and 2003; Wardrop et al., 2018; Kugler et al., 2019; Leyk et al., 2019; Archila Bustos et al., 2020). Most gridded population models focus on residential population distributions\u0026mdash;a means of \u0026ldquo;meeting people where they are\u0026rdquo;\u0026mdash;to promote global equity in health services and public resource delivery (Tatem, 2017; Wardrop et al., 2018). However, population changes caused by natural disasters, environmental exposures, demand on critical water and transportation infrastructure, and access to vital services such as nutrition and healthcare cannot be fully addressed without knowledge of human presence both at home and performing routine activities (Bhaduri et al., 2007; Rose \u0026amp; Bright, 2014). Acknowledging the difference in population distribution during nonresidential hours, there exists a complementary need to estimate the \u003cem\u003eambient\u0026nbsp;\u003c/em\u003e24 h population distribution, which accounts for activities unobserved in census data (e.g., school, work, services). To address this issue, Oak Ridge National Laboratory (ORNL) developed LandScan Global (Dobson et al., 2000; Bhaduri et al., 2002), which provides ambient population distribution at 30 arcsecond (approximately 1 km) resolution. Subsequently, ORNL developed the LandScan USA model (Bhaduri et al., 2007), which provides nighttime as well as daytime gridded population estimates for the United States at 3 arcseconds (approximately 90 m) resolution.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eComplementing the worldwide coverage offered by LandScan Global, the LandScan High Definition (HD) model emerged to develop gridded populations for select countries that offer increased spatial resolution (3 arcseconds) and temporal resolution (nighttime, daytime) and enhance knowledge of human security challenges, including vaccine distribution and conflict (Weber et al., 2018, Urban et al., 2023b). LandScan HD began as an extension of the LandScan USA model, initially combining dasymetric modeling with settlement detection from high-resolution satellite images (Cheriyadat et al., 2007; Vijayaraj et al., 2007 and 2008) and subsequently building occupancy models (Stewart et al., 2016; Urban et al., 2023a).\u003c/p\u003e\n\u003cp\u003eAlthough LandScan HD has traditionally been produced on a country-specific basis, recent advances in the spatial coverage and precision of foundational building and land-use datasets, as well as computational resources for gridded population modeling, now enable production of the first global implementation of the model. This paper describes the core LandScan HD methodology and how it is scaled to produce a high-resolution gridded ambient population for the world. Using the LandScan HD global baseline, we provide several illustrations that demonstrate practical application of high spatial resolution ambient populations. We also reflect on limitations of the current approach and strategies for improving the methodological components of our framework as it continues to develop.\u003c/p\u003e"},{"header":"2 Background","content":"\u003cp\u003eSince 1999, ORNL\u0026rsquo;s LandScan Global gridded population model and annual datasets have reported ambient distributions that estimate where populations may be found on average across a 24 h period (Leyk et al., \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). This approach primarily supports consequence assessment of unwarned population related to rapid onset events, such as critical infrastructure failures, conflict-driven violence, or natural disasters, for which the routine activities of the affected population hold critical importance beyond residential locations alone (Dobson et al., \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2000\u003c/span\u003e; Bhaduri et al., 2007; Bhaduri \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2008\u003c/span\u003e; Rose \u0026amp; Bright, \u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e2014\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eNew censuses, Earth observation, and digital trace data coupled with advancing technologies and theoretical sophistication have ushered in a period of unprecedented development in predictive modeling of human\u0026ndash;environment relationships to place populations more precisely and accurately (Darin et al., \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). Based on the assumption that the visible traces people leave on the surface of the earth\u0026mdash;cultivated and impermeable land cover, roads and other infrastructure, buildings in their different shapes and distributions, nighttime lights\u0026mdash;accurately reflect the people themselves, ever more granular data about these \u0026ldquo;reasonable prox[ies] of human activity\u0026rdquo; have been processed through ever more powerful algorithms to produce population estimates of ever higher resolution in the years since (Dobson et al., \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2000\u003c/span\u003e and 2002; Rose and Bright, \u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e2014\u003c/span\u003e; Leyk et al., \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). Producing a 24 hour ambient population based on these features requires balancing authoritative population totals with country and region-specific characterization of activity spaces. To achieve this balance, LandScan HD combines elements of the top-down and bottom-up taxonomy of gridded population models (Wardrop et al., \u003cspan citationid=\"CR68\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; Weber et al., \u003cspan citationid=\"CR69\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; Leyk et al., \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Dahmm et al., \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2020\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eTop-down methods disaggregate authoritative population counts from coarser spatial resolutions to regular spatial grids (often ranging in resolution from 100 m to 1 km). Although a basic areal interpolation would simply evenly divide the coarser population, ignoring how features in the natural and built environments influence people\u0026rsquo;s activity spaces, all the currently available gridded population datasets with global coverage account in some way for those constraints, and nearly all do so through dasymetric methods with varying levels of complexity. Much of the development of these top-down methods has focused on refinement of both spatial resolution and model constraints.\u003c/p\u003e \u003cp\u003eThe Gridded Population of the World (GPW), produced by the Center for International Earth Science Information Network, forms the basis for many top-down gridded population products. GPW is characterized as a \u0026ldquo;minimally modeled\u0026rdquo; product that distributes administrative population counts to grid cells (v1: 5 arcminutes/8.5 km, v2-3: 2.5 arcminutes/5 km, v4: 30 arcseconds/1 km) based on available land area per cell, excluding water bodies and protected areas (Lloyd et al., \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e2017\u003c/span\u003e; Dahmm et al., \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). The Global Rural Urban Mapping Project (GRUMP), an early enhancement of GPWv3, increases GPW\u0026rsquo;s spatial resolution of 30 arcseconds, bases urban and rural delineation on factors such as nighttime lights, and uses known settlement locations to disaggregate population counts (Balk et al., \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2005a\u003c/span\u003e,\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003eb\u003c/span\u003e; Balk et al., \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2010\u003c/span\u003e; Leyk et al., \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). The Global Human Settlement Layer Population Grid (GHS-POP) refines GPW\u0026rsquo;s spatial resolution further to 250 m built-up area presence/absence derived from Landsat imagery (Friere et al., 2016). With continued emphasis on spatial refinement of GPW, the University of Southampton\u0026rsquo;s WorldPop dataset (Tatem, \u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e2017\u003c/span\u003e) employs GPWv4 as its population baseline and uses a machine learning (ML) approach to reapportion it to a 3 arcsecond (approximately 100 m) grid using satellite imagery and a suite of geospatial covariates (Leyk et al., \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Dahmm et al., \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2020\u003c/span\u003e), with adjustments made for country-official or UNPD estimates and building footprint data (Stevens et al., \u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e2015\u003c/span\u003e; WorldPop, \u003cspan citationid=\"CR72\" class=\"CitationRef\"\u003e2024\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eLandScan Global, another top-down method developed in parallel to GPW and available at 30 arcsecond resolution, takes subnational census and administrative data as a starting point and reapportions populations through country-specific multivariable dasymetric models with weights for land use/land cover, infrastructure, urban areas, satellite imagery, and natural features such as slope, elevation, and bodies of water (Dobson et al., \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2000\u003c/span\u003e; Bhaduri et al., \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2002\u003c/span\u003e; Rose \u0026amp; Bright, \u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e2014\u003c/span\u003e). The resulting ambient population distribution accounts for a greater range of potential activity locations than WorldPop\u0026rsquo;s more direct residential focus.\u003c/p\u003e \u003cp\u003eBottom-up methods model relationships between geolocated population surveys and environmental covariates. The resulting point estimates of population may be reaggregated to any resolution appropriate to the problem of interest. For example, Weber et al. (\u003cspan citationid=\"CR69\" class=\"CitationRef\"\u003e2018\u003c/span\u003e) use microcensus surveys\u0026mdash;direct measures of building occupancy, typically conducted for small, populated areas\u0026mdash;to model population densities at 1 km resolution relative to different settlement types in Nigeria. An added advantage of the bottom-up approach is that is can be applied wherever covariate predictors are available, overcoming challenges for contexts in which data are generally sparse, infrequently collected, or otherwise unreliable (Weber et al., \u003cspan citationid=\"CR69\" class=\"CitationRef\"\u003e2018\u003c/span\u003e). The bottom-up approach does not supersede the top-down approach: the two are complementary and can be used together in a two-step process to fill data gaps (Darin et al., \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2022\u003c/span\u003e) or add details only present in the survey data (Bosco et al., \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2017\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eLandScan HD follows the top-down approach of allocating authoritative population counts to subnational locations, but it supports this process using bottom-up estimates from ORNL\u0026rsquo;s Population Density Tables (PDT) model. PDT\u0026rsquo;s Bayesian ML approach incorporates observations and subject matter expertise on facility occupancy throughout the world (Stewart et al., \u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e2016\u003c/span\u003e; Urban et al., \u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e2023\u003c/span\u003ea). Combining volumetric building characteristics (floor area, height) with PDT occupancy estimates for multiple facility types at daytime and nighttime, LandScan HD designates proportional population weights to redistribute population counts to distinct point locations. The resulting estimates are easily aggregated to a 3 arcsecond grid that preserves the official subnational population counts.\u003c/p\u003e"},{"header":"3 Data and Methods","content":"\u003cp\u003eThe core LandScan HD methodology begins with an account of local sociocultural and economic activities (population dynamics) at the building level. Human activities play out across a variety of facility types (e.g., residential, commercial, medical, industrial, transportation)\u0026mdash;each with a plausible range of occupancy characteristics\u0026mdash;that are distinct from one another based on activity purpose, time of day, and country/region. By combining the expected occupancy characteristics of different facility types with physical characteristics of buildings (floor area, height), the development of localized population weights enables distribution of authoritative population totals to built-up areas indexed by administrative boundaries and grid cell location.\u003c/p\u003e\n\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\n \u003ch2\u003e3.1 Data Sources\u003c/h2\u003e\n \u003cp\u003eWe curate several datasets as inputs to the LandScan HD data stack:\u003c/p\u003e\n \u003cul\u003e\n \u003cli\u003e\n \u003cp\u003e\u003cstrong\u003eFoundational information on populated areas\u003c/strong\u003e: Building footprint and settlement data that indicate where people reside and where normal patterns of activity occur.\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003e\u003cstrong\u003eBuilt-up area attribution\u003c/strong\u003e: Estimation of available interior space that can account for available for human occupancy during residential and nonresidential facility use.\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003e\u003cstrong\u003eOccupancy attribution\u003c/strong\u003e: Probabilistic estimation of daytime and nighttime average occupant density by facility type (e.g., residential, commercial, industrial, recreational) in units of people per 1000 ft\u003csup\u003e2\u003c/sup\u003e.\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003e\u003cstrong\u003ePopulation statistics\u003c/strong\u003e: Proportional reweighting of raw building-level population applied at an appropriate spatial scale to ensure alignment with authoritative population summaries (e.g., census). We default to the administrative level 1 (Adm1) because these are assumed to be self-contained zones where activities primarily take place.\u003c/p\u003e\n \u003c/li\u003e\n \u003c/ul\u003e\n \u003cdiv id=\"Sec4\" class=\"Section3\"\u003e\n \u003ch2\u003e3.1.1 Built-up areas and facility occupancy\u003c/h2\u003e\n \u003cp\u003e\u003cem\u003eIdentifying built-up areas\u003c/em\u003e\u003c/p\u003e\n \u003cp\u003eAn organizing principle of LandScan HD is that populations most likely occur in built-up areas, as represented by building footprint and settlement data. Curation of the LandScan HD data stack begins with a catalog of 2D building footprint data from ORNL\u0026rsquo;s building feature extraction (BFE) product. BFE\u0026rsquo;s core engine uses convolutional neural network (CNN)-based semantic segmentation algorithms along with a series of advances made for continuous performance improvements (Yang et al., \u003cspan class=\"CitationRef\"\u003e2018\u003c/span\u003e and \u003cspan class=\"CitationRef\"\u003e2024\u003c/span\u003e) as LandScan HD has progressed toward global coverage. The most recent operational-ready version of the BFE model is trained using the Oak Ridge Building Image and Training Label Net (Swan et al., \u003cspan class=\"CitationRef\"\u003e2024\u003c/span\u003e), which contains over 130,000 labeled tiles of 500 \u0026times; 500 pixels drawn from thousands of high-resolution (average resolution 0.47 \u003cem\u003em/px\u003c/em\u003e) commercial satellite imagery from Maxar WorldView (Maxar, \u003cspan class=\"CitationRef\"\u003e2024\u003c/span\u003e) (BGR\u0026thinsp;+\u0026thinsp;near-infrared) covering multiple regions of the globe. Model training uses a boundary-aware loss function as per Yang et al. (\u003cspan class=\"CitationRef\"\u003e2018\u003c/span\u003e) to improve separability of building instances. On average, evaluation at the pixel level across held-out tile samples indicates precision and recall rates on the order of 90%. Image preprocessing for BFE is handled by ORNL\u0026rsquo;s Parallel Integration and Processing Engine (PIPE), a workflow designed to transform raw satellite imagery into high-resolution, analysis-ready datasets crucial for geospatial intelligence applications, including modules for imagery ingest, orthorectification to correct geometric distortions, image resolution enhancement, and cloud detection, ensuring clear visibility of ground features (Reith et al., \u003cspan class=\"CitationRef\"\u003e2023\u003c/span\u003e).\u003c/p\u003e\n \u003cp\u003eWhen BFE is unavailable, we instead leverage Microsoft and Google building footprints (Sirko et al., 2021; Microsoft, \u003cspan class=\"CitationRef\"\u003e2024\u003c/span\u003e). However, each of these datasets can lack coverage in areas of the world where human settlements are known to exist. We identify such areas using the Global Human Settlement Layer (GHSL) Built-S 10 m raster, which provides continuous global measurement of built-up surface area (Pesaresi \u0026amp; Politis, \u003cspan class=\"CitationRef\"\u003e2023\u003c/span\u003e). Areas containing nonzero values in Built-S but lacking coverage in the BFE, Microsoft, and Google datasets are converted to a polygonal \u0026ldquo;no-data\u0026rdquo; mask. Using the no-data mask, we substitute vectorized Built-S cells as a failsafe for built-up areas (Fig. \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e).\u003c/p\u003e\n \u003cp\u003eFor failsafe cases, Built-S must be normalized to have the same density relationship with building footprints relative to the country of interest (Fig. \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e). This normalization is accomplished by measuring the total planar area of Built-S for the country of interest that intersects with other built-up area sources and then calculating the ratio between the two sources using Eq.\u0026nbsp;(1):\u003c/p\u003e\n \u003cdiv id=\"Equa\" class=\"Equation\"\u003e\n \u003cdiv class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e$$\\:\\begin{array}{c}FGRatio=\\:\\frac{\\sum\\:F{B}_{a}}{\\sum\\:G{B}_{a}},\\#\\left(1\\right)\\end{array}$$\u003c/div\u003e\n \u003c/div\u003e\n \u003cp\u003ewhere \u003cem\u003eFB\u003c/em\u003e\u003csub\u003e\u003cem\u003ea\u003c/em\u003e\u003c/sub\u003e represents the footprint-based built-up area, and \u003cem\u003eGB\u003c/em\u003e\u003csub\u003e\u003cem\u003ea\u003c/em\u003e\u003c/sub\u003e represents the Built-S built-up area.\u003c/p\u003e\n \u003cp\u003e\u003cem\u003eBuilt-up area attribution\u003c/em\u003e\u003c/p\u003e\n \u003cp\u003eTo support occupancy estimation, we attribute built-up areas by land use (e.g., residential, commercial, industrial) and volumetric properties (area and floor count). Most openly available global land-use data come from OpenStreetMap (OSM), a community-driven mapping platform that provides foundational geospatial data throughout the world (OpenStreetMap Contributors, 2024). OSM land-use polygons are attributed by key:value pairs, providing many categories (e.g., landuse:residential, building:school) that are assignable to built-up areas via spatial overlay. OSM provides extensive user guidance on feature labeling and a robust review cycle for contributions. However, this capability results in large volumes of continually updated land-use tags that complicate matching built-up area with PDT facility use types to estimate occupancy. As a solution, we apply a crosswalk between OSM land-use tags and PDT facility types, accounting for over 110,000 category mappings, to accommodate new or updated OSM features in most situations (Adams et al., \u003cspan class=\"CitationRef\"\u003e2023\u003c/span\u003e).\u003c/p\u003e\n \u003cp\u003eFor areas lacking coverage and detail in OSM, we supplement land-use attribution using the Urban Tactical Planner (UTP), a program run by the US Army Geospatial Center to facilitate the mapping and display of geospatial information in urban areas (US Army Corps of Engineers Army Geospatial Center, \u003cspan class=\"CitationRef\"\u003e2013\u003c/span\u003e) and the Multinational Geospatial Co-production Program (MGCP) (J\u0026oacute;zsef \u0026amp; Ol\u0026iacute;via, 2009). Following standards from the Defence Geospatial Information Working Group (DGIWG), these datasets have been continually updated by at least 28 nations since 2003. As with OSM, we crosswalk DGIWG land-use definitions with PDT categories to facilitate occupancy estimation. As a failsafe for missing land-use labels, our default assumption is general residential use.\u003c/p\u003e\n \u003cp\u003eWhereas floor area is an intrinsic property of building footprints, floor count requires manual calculation. We convert building height to floor count using estimates provided by OSM, UTP, and MGCP. If raw building height is available, then we manually estimate floor counts assuming 3 m per floor and then round up as needed to arrive at an integer number of floors. For other cases in which the data are structured to indicate a range of floors, we use the average floor count between the upper and lower bounds. As with land use, we apply a default floor count value (two floors) when height information is completely missing from the source datasets.\u003c/p\u003e\n \u003cp\u003e\u003cem\u003eFacility occupancy\u003c/em\u003e\u003c/p\u003e\n \u003cp\u003eTo produce bottom-up population (BUP) weights, facility occupancy estimates from PDT (Stewart et al., \u003cspan class=\"CitationRef\"\u003e2016\u003c/span\u003e; Urban et al., \u003cspan class=\"CitationRef\"\u003e2023\u003c/span\u003ea) are combined with building volumetric properties based on use type. PDT provides probabilistic estimates and ranges for daytime and nighttime densities (people per 1000 ft\u003csup\u003e2\u003c/sup\u003e) for over 60 facility types representing the spectrum of human activity spaces present over much of the globe. Conflating PDT facility classes with building use enables us to transfer the density estimates to each built-up area specific to different activity spaces.\u003c/p\u003e\n \u003cp\u003ePDT overcomes data sparsity using observation models and sociocultural knowledge. Observation models infer occupancy from observable data via rigorously proven theoretical linkages (Stewart et al., \u003cspan class=\"CitationRef\"\u003e2016\u003c/span\u003e; Morton, 2013). These observation models can accommodate qualitative information such as culturally informed activity patterns in places like cemeteries where facility use is not typically observed and/or recorded (Lunga et al., \u003cspan class=\"CitationRef\"\u003e2022\u003c/span\u003e), or through earth observation to understand measure structural properties of residential facilities (Woody \u0026amp; Frazier, \u003cspan class=\"CitationRef\"\u003e2023\u003c/span\u003e). PDT further reinforces these data by encoding sociocultural knowledge from experts and from open-source data on economic practices and cultural norms. These forms of data (knowledge encoding and modeled observational data) are dynamically updated in PDT via a Bayesian ML framework that produces final occupancy estimates and their associated uncertainty (Stewart et al., \u003cspan class=\"CitationRef\"\u003e2016\u003c/span\u003e). Built-up areas are attributed by PDT point estimates (default is 50th percentile) based on matching facility classes. This strategy propagates PDT uncertainty from occupant density to population weights via volumetric characteristics, allowing us to estimate the proportion of an area\u0026rsquo;s total population at each built-up area location.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv id=\"Sec5\" class=\"Section3\"\u003e\n \u003ch2\u003e3.1.2 Administrative boundaries and population totals\u003c/h2\u003e\n \u003cp\u003ePopulation models are often informed by authoritative population counts reported from the country of interest and organized around the boundaries of recognized subnational political units. LandScan HD relies on subnational population statistics available at each country\u0026rsquo;s Adm1 zones (e.g., states, provinces, regions) to produce an ambient population distribution. Population movements tend to occur at the scale of larger, socially and economically self-contained administrative zones rather than smaller areas like neighborhoods (Brelsford et al., \u003cspan class=\"CitationRef\"\u003e2022\u003c/span\u003e). Thus we use the Adm1 level as a common basis for catchment zones of daily routine activities throughout the world.\u003c/p\u003e\n \u003cp\u003eThe four key population components of the LandScan HD data stack are (1) authoritative international boundaries, (2) authoritative national population totals, which provide a reference for subnational population distribution, (3) Adm1 boundaries for geolocating buildings, and (4) Adm1-level population statistics for distributing subnational counts to buildings based on daytime/nighttime population weights.\u003c/p\u003e\n \u003cp\u003e\u003cem\u003eNational Population Data\u003c/em\u003e\u003c/p\u003e\n \u003cp\u003eWe aim to use each country\u0026rsquo;s most current census data to represent population totals, particularly from surveys conducted in the 2010s or 2020s. Where these are not available or are significantly dated, we use statistical estimates and projections instead. To standardize the population counts, we use the United States Census Bureau\u0026rsquo;s International Database (U.S Census Bureau, 2024) for each country to minimize inconsistencies in collection procedures and to ensure that countries without direct information maintain a dependable source. We apply this count to the national total and prorate and subnational population counts.\u003c/p\u003e\n \u003cp\u003e\u003cem\u003eAdministrative Boundaries\u003c/em\u003e\u003c/p\u003e\n \u003cp\u003eProducing a complete global catalog of Adm1 zones to match population totals involves harmonizing multiple representations of international political boundaries. Whenever possible, we source subnational boundaries directly from government agencies. Otherwise, we obtain them from a variety of open data sources including OSM, the Humanitarian Data Exchange (United Nations Office for the Coordination of Humanitarian Affairs, \u003cspan class=\"CitationRef\"\u003e2024\u003c/span\u003e), and GeoBoundaries (William \u0026amp; Mary geoLab, 2024). We perform comparisons across the various boundary sources to determine the most accurate source. Manual edits may be required to update names or derive feature mapping, and ancillary sources may be used to apply new reported changes in the administrative hierarchy. We clip all sourced spatial boundaries to the Large Scale International Boundaries (LSIB) dataset, which is produced by the US Department of State and provides annual updates on political boundaries. Specifically, we use a modified version of the LSIB that accounts for coastlines and sovereign political regions, provided jointly by the National Geospatial Intelligence Agency\u0026rsquo;s Office of Geography\u0026rsquo;s Geographic Boundaries Branch and the Department of State\u0026rsquo;s Office of the Geographer and Global Issues (US Department of State, \u003cspan class=\"CitationRef\"\u003e2024\u003c/span\u003e). As a final verification step, we perform detailed random checks on political boundaries to assure their alignment with topological features such as rivers or manufactured features such as roads and highways.\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\n \u003ch2\u003e3.2 Methods\u003c/h2\u003e\n \u003cdiv id=\"Sec7\" class=\"Section3\"\u003e\n \u003ch2\u003e3.2.1 Harmonizing the data stack\u003c/h2\u003e\n \u003cp\u003eA challenge associated with the data curation process described in Section 3.1 is limited spatial conformity among the various layers in the data stack. Following Moehl et al. (\u003cspan class=\"CitationRef\"\u003e2021\u003c/span\u003e), we overcome this challenge through a vector analytical framework approach. We represent built-up areas as a table of model leverageable cut-up labeled entities (MOLECULEs) that merges built-up area geometries, attributes, and geographic identifiers (Adm1 and grid cell index). As shown in Fig. \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e, each MOLECULE is a portion of a building footprint intersected by attribution source polygons and boundaries, encoded in the attribute table as a unique combination of data stack features. For cases in which multiple building use types overlap a MOLECULE, we select the labels associated with the smallest area polygons, assuming that smaller polygons represent more specific activity types associated with that space.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv id=\"Sec8\" class=\"Section3\"\u003e\n \u003ch2\u003e3.2.2 Gridded population estimation\u003c/h2\u003e\n \u003cp\u003eConverting the data stack to MOLECULEs enables us to distribute total Adm1 population to every available population unit. First, expected populations for each MOLECULE, \u003cem\u003eB\u003c/em\u003e, are estimated using Eq.\u0026nbsp;(2):\u003c/p\u003e\n \u003cdiv id=\"Equb\" class=\"Equation\"\u003e\n \u003cdiv class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e$$\\:\\begin{array}{c}{B}_{p}={B}_{a}{B}_{f}{B}_{ld},\\#\\left(2\\right)\\end{array}$$\u003c/div\u003e\n \u003c/div\u003e\n \u003cp\u003ewhere \u003cem\u003ep\u003c/em\u003e is the expected population based on footprint area \u003cem\u003ea\u003c/em\u003e, floor count \u003cem\u003ef\u003c/em\u003e, and PDT density \u003cem\u003ed\u003c/em\u003e as defined by the land use \u003cem\u003el\u003c/em\u003e.\u003c/p\u003e\n \u003cp\u003eThen, BUP weights are estimated by normalizing the expected populations of each MOLECULE by their sum. The product of the Adm1 population total \u003cem\u003eP\u003c/em\u003e\u003csub\u003e\u003cem\u003eZ\u003c/em\u003e\u003c/sub\u003e and BUP weights (Eq.\u0026nbsp;[3]) gives an updated baseline population estimate, \u003cem\u003eB\u003c/em\u003e\u003csub\u003e\u003cem\u003epop\u003c/em\u003e\u003c/sub\u003e, of subnational population per MOLECULE:\u003c/p\u003e\n \u003cdiv id=\"Equc\" class=\"Equation\"\u003e\n \u003cdiv class=\"mathdisplay\" id=\"FileID_Equc\" name=\"EquationSource\"\u003e$$\\:\\begin{array}{c}{B}_{pop}=\\frac{{B}_{p}}{\\sum\\:{B}_{p}}\\times\\:{P}_{z}.\\#\\left(3\\right)\\end{array}$$\u003c/div\u003e\n \u003c/div\u003e\n \u003cp\u003eWe perform this procedure for both daytime and nighttime PDT estimates and average the results to obtain a 24 hour ambient population estimate. Because the MOLECULEs are indexed by grid cell, we simply sum the ambient population counts by unique row/column identifier to produce gridded population counts. Finally, the cell estimates are converted to whole numbers using greatest mantissa rounding, which redistributes the total fractional portion of the estimates among the largest integer values (Coleman, \u003cspan class=\"CitationRef\"\u003e2015\u003c/span\u003e).\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv id=\"Sec9\" class=\"Section3\"\u003e\n \u003ch2\u003e3.2.3 Global implementation\u003c/h2\u003e\n \u003cp\u003eAdvances in the global coverage of data on the built environment and activity spaces (PDT) have enabled us to build a global data stack to support gridded ambient population estimates (Section 3.1) for every built-up area on Earth. The large volume of information in the global LandScan HD data stack is organized into a centralized PostGIS database. We index the data stack by a 1\u0026deg; resolution tileset, which forms the basis for worldwide generation of MOLECULEs datasets. The tile MOLECULE tables are attached to a parent table per built-up area source (e.g., level 0 and level 1 administrative boundaries). This design supports country-specific model runs by querying the MOLECULEs in tiles (1) that cover the country\u0026rsquo;s level 0 boundaries and (2) that have a presence flag for that country. Joining the MOLECULEs to country-specific PDT densities by land use then yields BUP weights, which are combined with Adm1 totals and converted to ambient population estimates and aggregated to the 3 arcsecond global grid, as discussed in Section 3.2.2. In rare cases, border grid cells may contain built-up areas from multiple countries. We address these instances by updating the cell counts to the grand sum of the country-level estimates.\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e"},{"header":"4 Demonstration","content":"\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\n \u003ch2\u003e4.1 LandScan HD Baseline for the Philippines\u003c/h2\u003e\n \u003cp\u003eWe explore several illustrations of the Philippines as an example result from the global LandScan HD implementation. For this dataset, 98% of the built-up area comes from Microsoft building footprints, and the remaining 2% comes from GHSL Built-S. Figure \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e shows the distribution of area by general category type. Approximately 75% of the density assignments in the model come from the land-use data stack. Most buildings in the Philippines are labeled; therefore, the model accounts for more variation in daytime/nighttime activity types than a purely residential model. However, this example is limited by a lack of height and floor counts (this information was available for less than 1% of the building area), so the model in this case relies heavily on default assumption of two floors per building. In updates to LandScan HD, we plan to address this deficiency and improve coverage of land-use labels by using predictive ML models of missing attributes (Section 5).\u003c/p\u003e\n \u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e\n \u003ch2\u003e4.2 Using LandScan HD to Assess Population Risk\u003c/h2\u003e\n \u003cp\u003eWe examine a practical application of LandScan HD using ambient estimates of individuals susceptible to flooding in Iloilo City, Philippines, over a 24 hour period. Our aim is to illustrate how ambient population data can be leveraged for urban planning and disaster risk management, providing more actionable insights than estimates derived from residential-only gridded population datasets. We compare LandScan HD to a residential-only population grid, WorldPop, specifically its latest 2020 top-down constrained product, which adjusts for UN World Population Prospects national population estimates (United Nations Department of Economic and Social Affairs, \u003cspan class=\"CitationRef\"\u003e2022\u003c/span\u003e; WorldPop, \u003cspan class=\"CitationRef\"\u003e2024\u003c/span\u003e).\u003c/p\u003e\n \u003cp\u003eIloilo City is the principal city and capital of Iloilo province in Western Visayas (Philippines Region VI), a region prone to perennial flooding and severe typhoon risk (Bagsit et al., \u003cspan class=\"CitationRef\"\u003e2014\u003c/span\u003e; Subade et al., \u003cspan class=\"CitationRef\"\u003e2014\u003c/span\u003e). This scenario is ideal for showcasing the utility of the LandScan HD dataset. In this demonstration, a flood hazard dataset obtained from the Philippine government via the country\u0026rsquo;s geoportal (Geoportal Philippines, 2024) encompasses nearly the entire municipality. This dataset delineates the flood hazards, scaled at 1:10,000, into three susceptibility categories: low, moderate, and high. For this analysis, we use the extent of the flood hazards layer to compute estimates of the population susceptible to each level of flood risk.\u003c/p\u003e\n \u003cdiv id=\"Sec13\" class=\"Section3\"\u003e\n \u003ch2\u003e4.2.1 Scenario Methodology\u003c/h2\u003e\n \u003cp\u003eWe employ both LandScan HD and WorldPop\u0026rsquo;s gridded population datasets, \u003cem\u003eP\u003c/em\u003e, where each grid cell \u003cem\u003ei\u003c/em\u003e contains a population count \u003cem\u003ep\u003c/em\u003e\u003csub\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sub\u003e. Both datasets have a common spatial resolution of 3 arcseconds. To assess flood risk, the flood hazards dataset delineates three flood zones for Iloilo City, \u003cem\u003eR\u003c/em\u003e = {\u003cem\u003er\u003c/em\u003elow,\u003cem\u003er\u003c/em\u003emoderate,\u003cem\u003er\u003c/em\u003ehigh}. The total population within the boundaries of Iloilo City, \u003cem\u003eP\u003c/em\u003e\u003csub\u003etotal\u003c/sub\u003e, is calculated using Eq.\u0026nbsp;(4):\u003c/p\u003e\n \u003cdiv id=\"Equd\" class=\"Equation\"\u003e\n \u003cdiv class=\"mathdisplay\" id=\"FileID_Equd\" name=\"EquationSource\"\u003e$$\\:\\begin{array}{c}{P}_{\\text{t}\\text{o}\\text{t}\\text{a}\\text{l}}=\\sum\\:_{i\\in\\:B}{p}_{i}\\#\\left(4\\right)\\end{array}$$\u003c/div\u003e\n \u003c/div\u003e\n \u003cp\u003ewhere \u003cem\u003eB\u003c/em\u003e represents all grid cells within the city limits.\u003c/p\u003e\n \u003cp\u003eFor each susceptibility zone \u003cem\u003er\u003c/em\u003e\u003csub\u003es\u003c/sub\u003e in \u003cem\u003eR\u003c/em\u003e, where s is either low, moderate, or high, \u003cem\u003eA\u003c/em\u003e\u003csub\u003es\u003c/sub\u003e is the subset of grid cells that fall within the boundaries of \u003cem\u003er\u003c/em\u003e\u003csub\u003es\u003c/sub\u003e. The population susceptible within each zone is computed using zonal statistics and Eq.\u0026nbsp;(5):\u003c/p\u003e\n \u003cdiv id=\"Eque\" class=\"Equation\"\u003e\n \u003cdiv class=\"mathdisplay\" id=\"FileID_Eque\" name=\"EquationSource\"\u003e$$\\:\\begin{array}{c}{P}_{\\text{s}}=\\sum\\:_{i\\in\\:{A}_{\\text{s}}}{p}_{i}\\#\\left(5\\right)\\end{array}$$\u003c/div\u003e\n \u003c/div\u003e\n \u003cp\u003eWe used ArcGIS Pro version 3.3.0 to match the geographic projection of the flood hazards dataset and with LandScan HD, then employed its zonal statistics tool to compute the population counts for Iloilo City and its susceptibility zones. The percentage of the population susceptible within each flood zone \u003cem\u003er\u003c/em\u003e\u003csub\u003es\u003c/sub\u003e, denoted as \u003cem\u003epct\u003c/em\u003e\u003csub\u003es\u003c/sub\u003e, is given by Eq.\u0026nbsp;(6):\u003c/p\u003e\n \u003cdiv id=\"Equf\" class=\"Equation\"\u003e\n \u003cdiv class=\"mathdisplay\" id=\"FileID_Equf\" name=\"EquationSource\"\u003e$$\\:\\begin{array}{c}{pct}_{\\text{s}}=\\frac{{P}_{\\text{s}}}{{P}_{\\text{t}\\text{o}\\text{t}\\text{a}\\text{l}}}\\times\\:100\\%.\\#\\left(6\\right)\\end{array}$$\u003c/div\u003e\n \u003c/div\u003e\n \u003cp\u003eThis analysis presents spatial distributions for both the LandScan HD and WorldPop datasets relative to flood susceptibility exposure, as shown in Fig. \u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e and summarized in Table \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e, by populations within each susceptibility zone. Although WorldPop estimates a greater proportion of Iloilo City\u0026rsquo;s residential population resides in high flood risk zones, more daily activities captured by the LandScan HD ambient population occur in moderate risk zones. However, we also observe localized differences in the concentration of high and moderate risk populations: whereas residential locations estimated by WorldPop tend to concentrate in coastal and riverine areas, the LandScan HD ambient population is more spatially dispersed, including within areas of increased flood risk without a large residential presence. This result is perhaps most notable along the Iloilo River near the geographic center of the city, which features a high density of commercial, business, and academic activities. We also observe a much higher ambient population estimate within Iloilo City than the residential one (659,990 vs. 430,457, respectively), suggesting that people move from adjacent suburban areas within the larger Adm1 zone (Philippines Region VI).\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eTable 1\u0026nbsp;\u003c/strong\u003eAmbient and residential-only population model estimates delineated by Iloilo City flood susceptibility, contrasting WorldPop (c. 2020) and LandScan HD (c. 2023).\u003c/p\u003e\n \u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\" width=\"520\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 13.6276%;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.547%;\"\u003e\n \u003cp\u003e\u003cstrong\u003ePop. Count\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.8196%;\"\u003e\n \u003cp\u003e\u003cstrong\u003e% of Pop.\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.5873%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTotal Pop.\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.547%;\"\u003e\n \u003cp\u003e\u003cstrong\u003ePop. Count\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.8196%;\"\u003e\n \u003cp\u003e\u003cstrong\u003e% of Pop.\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.0518%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTotal Pop.\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 13.6276%;\"\u003e\n \u003cp\u003eLow\u003c/p\u003e\n \u003cp\u003eModerate\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.547%;\"\u003e\n \u003cp\u003e89662\u003c/p\u003e\n \u003cp\u003e182927\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.8196%;\"\u003e\n \u003cp\u003e20.83%\u003c/p\u003e\n \u003cp\u003e42.50%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 14.5873%;\"\u003e\n \u003cp\u003e430457\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.547%;\"\u003e\n \u003cp\u003e150331\u003c/p\u003e\n \u003cp\u003e319135\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.8196%;\"\u003e\n \u003cp\u003e22.78%\u003c/p\u003e\n \u003cp\u003e48.35%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 13.0518%;\"\u003e\n \u003cp\u003e659990\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 13.6276%;\"\u003e\n \u003cp\u003eHigh\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.547%;\"\u003e\n \u003cp\u003e157868\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.8196%;\"\u003e\n \u003cp\u003e36.67%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.5873%;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.547%;\"\u003e\n \u003cp\u003e190524\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.8196%;\"\u003e\n \u003cp\u003e28.87%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.0518%;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec14\" class=\"Section2\"\u003e\n \u003ch2\u003e4.3 Construct Validity Assessment\u003c/h2\u003e\n \u003cp\u003eIn addition to providing a practical application of the LandScan HD global baseline (Section 4.2), we aim to demonstrate the \u003cem\u003econstruct validity\u003c/em\u003e of the ambient population estimates, defined broadly as their congruence with the construct (24 hour activities) they are designed to represent (Strauss and Smith, \u003cspan class=\"CitationRef\"\u003e2009\u003c/span\u003e). Because LandScan HD\u0026rsquo;s ambient population measures human presence at both residential and nonresidential locations, an important criterion for the actionability of information it provides is to match observational data reflecting collective human activities (infrastructure demand, atmospheric emissions). Furthermore, as examined in Section 4.2, the relationship between LandScan HD and human activity observations should be distinct from a purely residential representation of population. With these criteria in mind, we assess the global baseline LandScan HD release with respect to subnational CO\u003csub\u003e2\u003c/sub\u003e emissions estimates. This prevalent atmospheric greenhouse gas results from a variety of anthropogenic activities, including buildings, industry, agriculture, and transportation (Solomon et al., \u003cspan class=\"CitationRef\"\u003e2009\u003c/span\u003e; Crippa et al., \u003cspan class=\"CitationRef\"\u003e2023\u003c/span\u003e; Kabir et al., \u003cspan class=\"CitationRef\"\u003e2023\u003c/span\u003e). Subnational CO\u003csub\u003e2\u003c/sub\u003e estimates enable assessment of how LandScan HD aligns with the cumulative outcomes of human activities at roughly the scale of urban subdistricts, rural communities, and major transportation and industrial sites. For this exercise, subnational gridded CO\u003csub\u003e2\u003c/sub\u003e emissions available globally were obtained from the European Commission\u0026rsquo;s Emissions Database for Global Atmospheric Research (EDGAR) (Crippa et al., \u003cspan class=\"CitationRef\"\u003e2023\u003c/span\u003e).\u003c/p\u003e\n \u003cp\u003eWe compare EDGAR\u0026rsquo;s association with LandScan HD to residential population estimates provided by WorldPop. Specifically, we assess EDGAR global gridded CO\u003csub\u003e2\u003c/sub\u003e estimates available at 0.1\u0026deg; resolution for 2022, the most recent year in its historical archive. As in Section 4.2, we compare ambient to residential population estimates via the current LandScan HD global baseline and WorldPop\u0026rsquo;s most recent 2020 constrained (building footprint based), UN-adjusted population counts.\u003c/p\u003e\n \u003cp\u003eWe summed LandScan HD and WorldPop population counts to the coarser EDGAR grid cells based on their containment of 3 arcsecond grid cell centers. To account for areas with neither a persistent residential nor ambient population, grid cells with zero population counts were excluded for both LandScan HD and WorldPop, resulting in a total of 3596 EDGAR grid cells for analysis. We compared EDGAR to the aggregated LandScan HD and WorldPop grids using Pearson\u0026rsquo;s product moment correlation (Sedgwick, \u003cspan class=\"CitationRef\"\u003e2012\u003c/span\u003e), which measures the strength of linear association between estimated population and CO\u003csub\u003e2\u003c/sub\u003e emissions (tons released annually). To account for outlying high values in both population and CO\u003csub\u003e2\u003c/sub\u003e emissions totals, which can distort Pearson correlation statistics, the variables were log-transformed as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\text{log}1p\\left(x\\right)=\\text{log}(1+x)\\)\u003c/span\u003e\u003c/span\u003e, such that instances of zero residential population counts and nonzero ambient population counts (e.g., shipping ports with no permanent residential population) were preserved.\u003c/p\u003e\n \u003cp\u003eAlthough both aggregated population grids feature somewhat strong positive correlation with EDGAR, the ambient population estimates provided by LandScan HD are related more closely to CO\u003csub\u003e2\u003c/sub\u003e emissions (\u003cem\u003e\u0026rho;\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.761) than those of WorldPop (\u003cem\u003e\u0026rho;\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.682). Furthermore, the 95% confidence intervals between the two correlation statistics do not overlap (LandScan HD lower-bound \u003cem\u003e\u0026rho;\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.746; WorldPop upper-bound \u003cem\u003e\u0026rho;\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.699), suggesting a distinctly closer association between LandScan HD and CO\u003csub\u003e2\u003c/sub\u003e emissions in the Philippines than that of WorldPop. Figure \u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e compares EDGAR with the aggregated population grids. At a high level, the most notable difference between LandScan HD and WorldPop is the extent of unpopulated or very sparsely populated areas in the former. In LandScan HD, these areas tend to contain some ambient population, which corresponds to light to moderate annual emissions in EDGAR and in many cases appears to be linked to human settlements or economic activities captured by the LandScan HD data stack.\u003c/p\u003e\n\u003c/div\u003e"},{"header":"5 Limitations and Improvements","content":"\u003cp\u003eThe LandScan HD baseline model provides a worldwide representation of ambient population at a high spatial resolution of 3 arcseconds, at which only gridded residential populations (WorldPop) were previously available. As demonstrated in Section 4, this approach broadens the scope of data available to researchers and practitioners for making precise population estimates, particularly when human presence and activities are of chief importance. This approach marks the first attempt to scale this methodology globally, and as such, it features several limitations. These largely involve how we address multiple challenges of incompleteness or sparsity of layers in the data stack\u0026mdash;particularly with respect to building floor area and activity type\u0026mdash;using basic default modeling assumptions. However, these defaults will become less necessary as several new methods for building feature extraction, building land-use, and height prediction become available to fill labeling gaps in the LandScan HD data stack.\u003c/p\u003e \u003cp\u003eThe first of these issues pertains to detection of built-up areas. Although the problem of missing building footprints in populated areas has been addressed by using GHSL Built-S as a failsafe, an added challenge is that the building footprints themselves also feature errors of omission or commission, including inaccurate spatial geometries (Gonzales, \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). As a solution, we plan to expand ORNL BFE coverage to more areas of the world by developing several AI model improvements that incorporate modern transformer-based architectures (Dosovitskiy et al., \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2020\u003c/span\u003e) as well as foundation models for high-resolution imagery with self-supervised learning schemes for model pretraining (Dias et al., \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). In automation of BFE, ongoing efforts to update PIPE focus on developing additional processing steps (e.g., atmospheric correction), implementing each processing function as its own plug-and-play module, and enhancing workflow scalability and efficiency.\u003c/p\u003e \u003cp\u003eMissing or incomplete information about building height and use poses another challenge in the LandScan HD model. In terms of physical characteristics, the lack of a global inventory of building heights currently hinders the development of accurate measures of building floor area in some regions of the world. Inaccurate representation of building floor area in turn leads to disproportionately large or small BUP weights. The current default assumption of two floors for buildings with missing height attribution does not adequately represent local building heights at a global scale. Toward building use attribution, we also acknowledge that the default assumption of unlabeled buildings from OSM and other data layers being residential is likely too broad. Incorrectly categorizing residential/nonresidential uses can significantly compromise the accuracy of BUP estimates by assigning excessive or insufficient occupancy levels. To address these limitations, we plan to build upon recent advancements in ML models for predicting building height and use these models at scale by leveraging the intrinsic relationship between building morphology, function, and urban form (Biljecki et al., \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2017\u003c/span\u003e; Milojevic-Dupont et al., \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Adams et al., \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Nachtigall et al., \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Hartmann et al., \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e2024\u003c/span\u003e; Stipek et al., \u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). Furthermore, to support more robust characterization of nonresidential activities, particularly in data-poor regions of the world, we next plan to incorporate MapSpace, a high spatial resolution inventory of land-use labels built on road network tessellations (Thakur \u0026amp; Fan, \u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e2021\u003c/span\u003e; Fan \u0026amp; Thakur, \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). MapSpace extends ORNL\u0026rsquo;s PlanetSense platform, a comprehensive catalog of points of interest (Thakur et al., \u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e2015\u003c/span\u003e and \u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e2016\u003c/span\u003e) leveraging a semantic labeling system, SONET, to conflate these categories with PDT facility types (Palumbo et al., \u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Fan et al., \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2023\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eA final challenge for LandScan HD is quantifying uncertainty related to modeling decisions. Enhancements to the data stack may improve modeling decisions, but their effects on the end product still need direct evaluation. The LandScan HD model\u0026rsquo;s deterministic nature eliminates traditional uncertainty measurements because it consistently produces the same output from identical inputs without variability. Nevertheless, the individual inputs of the LandScan HD model contain various levels of uncertainty that can be quantified. Examples include assigning building commercial use when multiple land-use labels are present or labeling a building with three floors when it may plausibly have four. Thus, future work will involve providing measures of confidence in our population estimates by developing a confidence index that is driven by the uncertainties of the data used to produce the population estimate (Section 3.1).\u003c/p\u003e"},{"header":"6 Conclusion and Outlook","content":"\u003cp\u003eTo our knowledge, the global LandScan HD baseline model provides the first 24 hour ambient population for the world at 3 arcsecond (roughly 90 m) spatial resolution. The methodology detailed in this paper requires curating and fusing large volumes of data, including population and occupancy statistics as well as the physical and activity-specific properties of buildings, in development of a scalable workflow. Our approach demonstrates the importance of ambient population estimates (combined residential/nonresidential activities) in terms of risk to unwarned populations in the face of climate-driven hazards (Section 4.2), which may be extended to a variety of related human security challenges. Furthermore, the alignment of LandScan HD with the collective effects of human activities (CO\u003csub\u003e2\u003c/sub\u003e emissions) in the Philippines (Section 4.3) demonstrates the underlying model\u0026rsquo;s congruence with 24 hour population dynamics. In each of these illustrations, LandScan HD provides insights about the location and presence of the unwarned population in ways that purely residential/nighttime gridded populations (WorldPop) that are often used in practice do not.\u003c/p\u003e \u003cp\u003eIn addition to enhancing the LandScan HD model and developing a global confidence index for quantifying uncertainty in our population estimates (Section 5), automation of the LandScan HD workflow can support rapid population estimates in the face of global humanitarian crises (Urban et al., \u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e2023\u003c/span\u003eb). To support this objective, we plan to introduce techniques such as building instance segmentation/counting, building damage assessment for conflict/natural disaster scenarios (Urban et al., \u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e2023\u003c/span\u003eb), and feature extraction for informal settlements (e.g., refugee camps).\u003c/p\u003e \u003cp\u003eThe development of a global LandScan HD baseline model also opens possibilities for exploring how human activity spaces align with sociodemographic and livelihood characteristics throughout the world to address problems like climate risk and vulnerability, spatial accessibility to vital resources, and infrastructure use. Although methods for exploring such problems have been established for the United States (Tuccillo \u0026amp; Gaboardi, \u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Tuccillo et al., \u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e2023\u003c/span\u003e), they require further development for other regions of the world, particularly low- and middle-income countries (LMICs), for which sparse population data often inhibit decisions about emergency and resource planning. To address this shortcoming, we are now using LandScan HD to build out methods for characterizing populations in LMICs, expanding upon an approach established by Urban et. al (\u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e2022\u003c/span\u003e) to downscale social, demographic, and economic information from global microdata sources such as the Demographic and Health Surveys to built-up areas.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAcknowledgements\u003c/h2\u003e \u003cp\u003eThe authors would like to thank Edward Bright, Amy Rose, Jacob McKee, and Melanie Laverdiere for their foundational contributions to the Landscan HD project.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eAdams, D.S., Hauser, T., and Moehl, J. (2023). Decoding Ethiopian Abodes: Towards Classifying Buildings by Occupancy Type Using Footprint Morphology. \u003cem\u003eProceedings of the 2023 International Conference on Machine Learning and Applications (ICMLA)\u003c/em\u003e (pp. 210-217). https://doi.org/10.1109/ICMLA58977.2023.00037 \u003c/li\u003e\n\u003cli\u003eArchila Bustos, M.F., Hall, O., Niedomysl, T., Ernstson, U. (2020). A pixel level evaluation of five multitemporal global gridded population datasets: a case study in sweden, 1990\u0026ndash;2015. \u003cem\u003ePopulation and Environment\u003c/em\u003e, 42, 255\u0026ndash;277.\u003c/li\u003e\n\u003cli\u003eBagsit, F.U., Suyo, J.G.B., Subade, R.F., Basco, J., \u003cem\u003eet al.\u003c/em\u003e (2014). Do adaptation and coping mechanisms to extreme climate events differ by gender? the case of flood-affected households in dumangas, iloilo, philippines. \u003cem\u003eAsian Fisheries Science\u003c/em\u003e, 25, 111\u0026ndash;118.\u003c/li\u003e\n\u003cli\u003eBalk, D., Brickman, M., Anderson, B., Pozzi, F., \u0026amp; Yetman, G. (2005b). Estimates of future global population distribution to 2015. \u003cem\u003eMapping global urban and rural population distributions\u003c/em\u003e. Food and Agriculture Organization of the United Nations (pp. 53\u0026ndash;73).\u003c/li\u003e\n\u003cli\u003eBalk, D., Pozzi, F., Yetman, G., Deichmann, U., \u0026amp; Nelson, A. (2005a). The distribution of people and the dimension of place: methodologies to improve the global estimation of urban extents. \u003cem\u003eInternational Society for Photogrammetry and Remote Sensing Proceedings of the Urban Remote Sensing Conference\u003c/em\u003e (pp. 14\u0026ndash;16).\u003c/li\u003e\n\u003cli\u003eBalk, D., Yetman, G., \u0026amp; De Sherbinin, A. (2010). Construction of gridded population and poverty data sets from different data sources. \u003cem\u003eProceedings of European Forum for Geostatistics Conference\u003c/em\u003e (pp. 12\u0026ndash;20).\u003c/li\u003e\n\u003cli\u003eBhaduri, B. (2008). Population Distribution During the Day. In: Shekhar, S. and Hui Xiong, (Eds). \u003cem\u003eEncyclopedia of GIS \u003c/em\u003e(pp. 1377). Springer-Verlag.\u003c/li\u003e\n\u003cli\u003eBhaduri, B., Bright, E., Coleman, P., and Dobson. J. (2002). LandScan: Locating People is What Matters. \u003cem\u003eGeoinformatics\u003c/em\u003e, 5(2): 4\u0026ndash;37.\u003c/li\u003e\n\u003cli\u003eBiljecki, F., Ledoux, H., Stoter, J. (2017). Generating 3d city models without elevation data. \u003cem\u003eComputers, Environment and Urban Systems\u003c/em\u003e, 64, 1\u0026ndash;18. https://doi.org/10. 1016/j.compenvurbsys.2017.01.001\u003c/li\u003e\n\u003cli\u003eBosco, C., Alegana, V., Bird, T., Pezzulo, C., Bengtsson, L., Sorichetta, A., Steele, J., Hornby, G., Ruktanonchai, C., Ruktanonchai, N., \u003cem\u003eet al.\u003c/em\u003e (2017). Exploring the high resolution mapping of gender-disaggregated development indicators. \u003cem\u003eJournal of The Royal Society Interface,\u003c/em\u003e 14(129).\u003c/li\u003e\n\u003cli\u003eBrelsford, C., Moehl, J., Weber, E., Sparks, K., Rose, A. (2022). Improving the LandScan USA non-obligate population estimate (nope). \u003cem\u003eUniversity of California Santa Barbara Spatial Data Science Symposium 2022 Short Paper Proceedings\u003c/em\u003e.\u003cem\u003e \u003c/em\u003ehttps://doi. org/10.25436/E2X307\u003c/li\u003e\n\u003cli\u003eCheriyadat, A., Bright, E., Bhaduri, B, and Potere, D. (2007). Mapping of Settlements in High Resolution Satellite Imagery Using High Performance Computing. \u003cem\u003eGeoJournal\u003c/em\u003e, 69(1\u0026ndash;2), 119\u0026ndash;129.\u003c/li\u003e\n\u003cli\u003eColeman, Charles (2015). SAS\u0026reg; Macros for Constraining Arrays of Numbers\u003cem\u003e.\u003c/em\u003e \u003cem\u003eSoutheast SAS Users Group 2015 Proceedings\u003c/em\u003e (pp. 1-9). https://www2.census.gov/library/working-papers/2015/econ/2015-coleman.pdf \u003c/li\u003e\n\u003cli\u003eCrippa, M., Guizzardi, D., Schaaf, E., Monforti-Ferrario, F., Quadrelli, R., Risquez Martin, A., Rossi, S., Vignati, E., Muntean, M., Brandao De Melo, J., Oom, D., Pagani, F., Banja, M., Taghavi, Moharamli, P., K\u0026ouml;ykk\u0026auml;, J., Grassi, G., Branco, A., San-Miguel, J. (2023). GHG emissions of all world countries \u0026ndash; 2023. \u003cem\u003ePublications Office of the European Union\u003c/em\u003e. https://doi.org/10.2760/953322\u003c/li\u003e\n\u003cli\u003eDahmm, H., Rabiee, M., Espey, J., Adamo, S. (2020). Leaving no one off the map: a guide for gridded population data for sustainable development. SDSN TReNDS. https://files.unsdsn.org/Leaving%2Bno%2Bone%2Boff%2Bthe%2Bmap-4.pdf \u003c/li\u003e\n\u003cli\u003eDarin, E., Ku\u0026acute;epi\u0026acute;e, M., Bassinga, H., Boo, G., Tatem, A.J., Reeve, P. (2022). The population seen from space: When satellite images come to the rescue of the census. \u003cem\u003ePopulation\u003c/em\u003e, 77(3), 437\u0026ndash;464.\u003c/li\u003e\n\u003cli\u003eDias, P., Potnis, A., Guggilam, S., Yang, L., Tsaris, A., Medeiros, H., Lunga, D. (2023). An agenda for multimodal foundation models for earth observation. \u003cem\u003eIn IGARSS 2023 IEEE International Geoscience and Remote Sensing Symposium\u003c/em\u003e (pp. 1237\u0026ndash; 1240).\u003c/li\u003e\n\u003cli\u003eDobson, J.E., Bright, E.A., Coleman, P.R., Durfee, R.C., Worley, B.A. (2000). LandScan: a global population database for estimating populations at risk. \u003cem\u003ePhotogrammetric engineering and remote sensing\u003c/em\u003e, 66(7), 849\u0026ndash;857.\u003c/li\u003e\n\u003cli\u003eDobson, J., Bright, E., Coleman, P., and Bhaduri, B., (2003). LandScan: A Global population database for estimating population at risk. In Victor Mesev (Ed.) \u003cem\u003eRemotely-Sensed Cities\u003c/em\u003e (pp. 378). Taylor and Francis.\u003c/li\u003e\n\u003cli\u003eDosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv:2010.11929\u003c/li\u003e\n\u003cli\u003eFan, J., Bentley, J., Thakur, G.M. (2023). Sonet++: A knowledge graph of geographic categories based on osm tag representation. https://www.osti.gov/servlets/purl/2000381 \u003c/li\u003e\n\u003cli\u003eFan, J., Thakur, G. (2023). Towards poi-based large-scale land use modeling: spatial scale, semantic granularity, and geographic context. \u003cem\u003eInternational Journal of Digital Earth\u003c/em\u003e, 16(1), 430\u0026ndash;445.\u003c/li\u003e\n\u003cli\u003eFreire S., MacManus K., Pesaresi M., Doxsey-Whitfield E., Mills J. (2016). Development of new open and free multi-temporal global population grids at 250 m resolution. \u003cem\u003eGeospatial Data in a Changing World; Association of Geographic Information Laboratories in Europe (AGILE) 2016\u003c/em\u003e.\u003c/li\u003e\n\u003cli\u003eGonzales, J.J. (2023). Building-level comparison of microsoft and google open building footprints datasets. In Beecham, R., Long, J.A., Smith, D., Zhao, Q., Wise, S. (eds.) \u003cem\u003eProceedings of the 12th International Conference on Geographic Information Science (GIScience 2023)\u003c/em\u003e (pp. 35:1 \u0026ndash; 35:6). https://doi.org/10.4230/LIPIcs.GIScience.2023.35 \u003c/li\u003e\n\u003cli\u003eHartmann, A., Behnisch, M., Hecht, R., Meinel, G. (2024). Prediction of residential and nonresidential building usage in Germany based on a novel nationwide reference data set. \u003cem\u003eEnvironment and Planning B: Urban Analytics and City Science\u003c/em\u003e, 51(1), 216\u0026ndash;233. https://doi.org/10.1177/23998083231175680\u003c/li\u003e\n\u003cli\u003eJ\u0026oacute;zsef, C., \u0026amp; Olıvia, M. (2009). Multinational Geospatial Co-production Program (Mgcp). \u003cem\u003eGeodezia Es Kartografia\u003c/em\u003e, 61, 15\u0026ndash;18.\u003c/li\u003e\n\u003cli\u003eKabir, M., Habiba, U.E., Khan, W., Shah, A., Rahim, S., Patricio, R., Ali, L., Shafiq, M., \u003cem\u003eet al.\u003c/em\u003e (2023). Climate change due to increasing concentration of carbon dioxide and its impacts on environment in 21st century; a mini review. \u003cem\u003eJournal of King Saud University-Science\u003c/em\u003e, 35(5).\u003c/li\u003e\n\u003cli\u003eKugler, T.A., Grace, K., Wrathall, D.J., Sherbinin, A., Van Riper, D., Aubrecht, C., Comer, D., Adamo, S.B., Cervone, G., Engstrom, R., \u003cem\u003eet al.\u003c/em\u003e (2019). People and pixels 20 years later: the current data landscape and research trends blending population and environmental data. \u003cem\u003ePopulation and Environment\u003c/em\u003e, 41, 209\u0026ndash;234.\u003c/li\u003e\n\u003cli\u003eLeyk, S., Gaughan, A.E., Adamo, S.B., Sherbinin, A., Balk, D., Freire, S., Rose, A., Stevens, F.R., Blankespoor, B., Frye, C., Comenetz, J., Sorichetta, A., MacManus, K., Pistolesi, L., Levy, M., Tatem, A.J., Pesaresi, M. (2019). The spatial allocation of population: a review of large-scale gridded population data products and their fitness for use. \u003cem\u003eEarth System Science Data\u003c/em\u003e, 11(3), 1385\u0026ndash;1409. https://doi.org/10. 5194/essd-11-1385-2019\u003c/li\u003e\n\u003cli\u003eLloyd, C. T., Sorichetta, A., \u0026amp; Tatem, A. J. (2017). High resolution global gridded data for use in population studies. \u003cem\u003eScientific data\u003c/em\u003e, 4(1), 1\u0026ndash;17.\u003c/li\u003e\n\u003cli\u003eLunga, D., Dhamdhere, R., Walters, S., Bragg, L., Makkar, N., Urban, M. (2022). Learning to count grave sites for cemetery observation models with satellite imagery. \u003cem\u003eIEEE Geoscience and Remote Sensing Letters, \u003c/em\u003e19, 1\u0026ndash;5. https://doi.org/10.1109/ LGRS.2020.3022328\u003c/li\u003e\n\u003cli\u003eMaxar (2024). WorldView-3. https://resources.maxar.com/data-sheets/worldview-3.\u003c/li\u003e\n\u003cli\u003eMicrosoft (2024). \u003cem\u003eGlobal ML Building Footprints\u003c/em\u003e [Data set]. https://github.com/ microsoft/GlobalMLBuildingFootprints\u003c/li\u003e\n\u003cli\u003eMilojevic-Dupont, N., Hans, N., Kaack, L.H., Zumwald, M., Andrieux, F., Barros Soares, D., Lohrey, S., Pichler, P.-P., Creutzig, F. (2020). Learning from urban form to predict building heights. \u003cem\u003ePLOS ONE\u003c/em\u003e, 15(12), 1\u0026ndash;22. https://doi.org/10.1371/ journal.pone.0242010\u003c/li\u003e\n\u003cli\u003eMoehl, J., Weber, E., McKee, J. (2021). A vector analytical framework for population modeling. In The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences (pp. 103\u0026ndash;108). https://doi.org/10.5194/ isprs-archives-XLVI-4-W2-2021-103-2021\u003c/li\u003e\n\u003cli\u003eNachtigall, F., Milojevic-Dupont, N., Wagner, F., Creutzig, F. (2023). Predicting building age from urban form at large scale. \u003cem\u003eComputers, Environment and Urban Systems\u003c/em\u003e, 105. https://doi.org/10.1016/j.compenvurbsys.2023.102010\u003c/li\u003e\n\u003cli\u003eOpenStreetMap Contributors (2024). Planet OSM [Data set]. https://planet.osm.org \u003c/li\u003e\n\u003cli\u003ePalumbo, R., Thompson, L., Thakur, G. (2019). Sonet: a semantic ontological network graph for managing points of interest data heterogeneity. In \u003cem\u003eProceedings of the 3rd ACM SIGSPATIAL International Workshop on Geospatial Humanities\u003c/em\u003e (pp. 1\u0026ndash;6).\u003c/li\u003e\n\u003cli\u003ePesaresi, M., Politis, P. (2023). \u003cem\u003eGHS-BUILT-S R2023A - GHS built-up surface grid, derived from Sentinel2 composite and Landsat, multitemporal (1975-2030)\u003c/em\u003e [Data set]. http://data.europa.eu/89h/9f06f36f-4b11-47ec-abb0-4f8b7b1d72ea,doi:10.2905/9F06F36F-4B11-47EC-ABB0-4F8B7B1D72EA\u003c/li\u003e\n\u003cli\u003eP\u0026ouml;rtner, H.-O., Roberts, D.C., Adams, H., Adelekan, I., Adler, C., Adrian, R., Aldunce, P., Ali, E., Ara Begum, R., Bednar-Friedl, B. (2022). Technical summary. \u003cem\u003eIPCC Sixth Assessment Report\u003c/em\u003e (pp. 37\u0026ndash;118).\u003c/li\u003e\n\u003cli\u003eReith, A., McKee, J., Rose, A., Laverdiere, M., Swan, B., Hughes, D., ... \u0026amp; Lunga, D. (2023). Providing geospatial intelligence through a scalable imagery pipeline. In \u003cem\u003eAdvances in Scalable and Intelligent Geospatial Analytics: Challenges and Applications\u003c/em\u003e (pp. 153-168). CRC Press.\u003c/li\u003e\n\u003cli\u003eRose, A.N., Bright, E. (2014). The LandScan Global Population Distribution Project: Current State of the Art and Prospective Innovation. \u003cem\u003eProceedings of the Population Association of America 2014 Annual Meeting\u003c/em\u003e. https://paa2014.populationassociation.org/papers/ 143242\u003c/li\u003e\n\u003cli\u003eSedgwick, P. (2012). Pearson\u0026rsquo;s correlation coefficient. \u003cem\u003eBMJ\u003c/em\u003e, 345, e4483.https://doi.org/10.1136/bmj.e4483\u003c/li\u003e\n\u003cli\u003eSirko, W., Kashubin, S., Ritter, M., Annkah, A., Bouchareb, Y.S.E., Dauphin, Y.N., Keysers, D., Neumann, M., Ciss\u0026acute;e, M., Quinn, J. (2023). Continental-scale building detection from high resolution satellite imagery. https://doi.org/10.48550/arXiv.2107.12283 \u003c/li\u003e\n\u003cli\u003eSolomon, S., Plattner, G.-K., Knutti, R., Friedlingstein, P. (2009). Irreversible climate change due to carbon dioxide emissions. \u003cem\u003eProceedings of the national academy of sciences\u003c/em\u003e, 106(6), 1704\u0026ndash;1709.\u003c/li\u003e\n\u003cli\u003eStevens, F.R., Gaughan, A.E., Linard, C., Tatem, A.J. (2015). Disaggregating census data for population mapping using random forests with remotely-sensed and ancillary data. \u003cem\u003ePloS one,\u003c/em\u003e 10(2).\u003c/li\u003e\n\u003cli\u003eStewart, R., Urban, M., Duchscherer, S., Kaufman, J., Morton, A., Thakur, G., Piburn, J., Moehl, J. (2016). A bayesian machine learning model for estimating building occupancy from open source data. \u003cem\u003eNatural Hazards\u003c/em\u003e, 81(3), 1929\u0026ndash;1956. https://doi.org/ 10.1007/s11069-016-2164-9\u003c/li\u003e\n\u003cli\u003eStipek, C., Hauser, T., Adams, D., Epting, J., Brelsford, C., Moehl, J., ... \u0026amp; Stewart, R. (2024). Inferring building height from footprint morphology data. \u003cem\u003eScientific Reports\u003c/em\u003e, \u003cem\u003e14\u003c/em\u003e(1).\u003c/li\u003e\n\u003cli\u003eStrauss, M.E., Smith, G.T. (2009). Construct validity: Advances in theory and methodology. Annual review of clinical psychology, 5, 1\u0026ndash;25.\u003c/li\u003e\n\u003cli\u003eSubade, R., Suyo, J., Ebay, J., Lozada, E., Dator-Bercilla, J., Tionko, A., et al. (2014). Adaptation and coping strategies to extreme climate conditions: impact of Typhoon Frank in selected sites in Iloilo, Philippines. http://www.eepseapartners.org/post-2476 \u003c/li\u003e\n\u003cli\u003eSwan, B., Pyle, J., Roddy, D., Rose, A., Yang, H.L., Laverdiere, M. (2024). ORBITaL-Net training library for building extraction [Data set]. https://doi.org/10.25452/figshare. plus.25282225.v1\u003c/li\u003e\n\u003cli\u003eTatem, A.J. (2017). Worldpop, open data for spatial demography. \u003cem\u003eScientific Data,\u003c/em\u003e 4(1), 1\u0026ndash;4.\u003c/li\u003e\n\u003cli\u003eThakur, G.S., Bhaduri, B.L., Piburn, J.O., Sims, K.M., Stewart, R.N., Urban, M.L. (2015). Planetsense: a real-time streaming and spatio-temporal analytics platform for gathering geo-spatial intelligence from open source data. In \u003cem\u003eProceedings of the 23rd SIGSPATIAL International Conference on Advances in Geographic Information Systems\u003c/em\u003e (pp. 1\u0026ndash;4).\u003c/li\u003e\n\u003cli\u003eThakur, G., Fan, J. (2021). Mapspace: POI-based multi-scale global land use modeling. In \u003cem\u003eGIScience 2023 Short Paper Proceedings. \u003c/em\u003ehttps://doi.org/10.25436/E2Z59N \u003c/li\u003e\n\u003cli\u003eThakur, G.S., Sparks, K., Li, R., Stewart, R.N., Urban, M.L. (2016). Demonstrating PlanetSense: gathering geo-spatial intelligence from crowd-sourced and social-media data. In: \u003cem\u003eProceedings of the 24th ACM SIGSPATIAL International Conference on Advances in Geographic Information Systems\u003c/em\u003e (pp. 1\u0026ndash;4).\u003c/li\u003e\n\u003cli\u003eTuccillo, J.V., Gaboardi, J.D. (2022). Likeness: a toolkit for connecting the social fabric of place to human dynamics. In Agarwal, M., Calloway, C., Niederhut, D., Shupe, D. (eds.) \u003cem\u003eProceedings of the 21\u003csup\u003est \u003c/sup\u003ePython in Science Conference \u003c/em\u003e(pp. 125 \u0026ndash; 135). https: //doi.org/110.25080/majora-212e5952-014\u003c/li\u003e\n\u003cli\u003eTuccillo, J., Stewart, R., Rose, A., Trombley, N., Moehl, J., Nagle, N., Bhaduri, B. (2023). UrbanPop: A spatial microsimulation framework for exploring demographic influences on human dynamics. Applied Geography, 151. https://doi.org/10.1016/j.apgeog.2022.102844 \u003c/li\u003e\n\u003cli\u003eUnited Nations Department of Economic and Social Affairs (2022). Methodology of the United Nations population estimates and projections. https: //population.un.org/wpp/Publications/Files/WPP2022 Methodology.pdf\u003c/li\u003e\n\u003cli\u003eUnited Nations Office for the Coordination of Humanitarian Affairs (2024). Humanitarian Data Exchange v1.85.9 [Data set]. https://data.humdata.org/\u003c/li\u003e\n\u003cli\u003eUrban, M., Moehl, J., Dias, P., Tuccillo, J., Reith, A., Sims, K., Walters, S., Arndt, J., Potnis, A., Lunga, D. (2023). Towards rapid response updates of populations at risk. \u003cem\u003eProceedings of the IEEE International Geoscience and Remote Sensing Symposium\u003c/em\u003e (pp. 907\u0026ndash;910). https://doi.org/10.1109/IGARSS52108.2023. 10282319\u003c/li\u003e\n\u003cli\u003eUrban, M., Stewart, R., Basford, S., Palmer, Z., Kaufman, J. (2023). Estimating building occupancy: a machine learning system for day, night, and episodic events. \u003cem\u003eNatural Hazards\u003c/em\u003e, 116(2), 2417\u0026ndash;2436. https://doi.org/10.1007/s11069-022-05772-3\u003c/li\u003e\n\u003cli\u003eUrban, M., Tuccillo, J., Frazier, T., Cunningham, A., Fan, J., Dias, P., Arndt, J., Bowman, J., Gaboardi, J. (2022). \u003cem\u003eNew Insights into the Tightly Coupled Social Fabric of the Built Environment\u003c/em\u003e [Conference presentation]. American Geophysical Union 2022 Annual Meeting.\u003c/li\u003e\n\u003cli\u003eUS Army Corps of Engineers Army Geospatial Center (2013). Urban Tactical Planner Factsheet. https://www.agc.army.mil/Media/Fact-Sheets/Fact-Sheet-Article-View/Article/480931/urban-tactical-planner \u003c/li\u003e\n\u003cli\u003eUS Census Bureau (2024). \u003cem\u003eInternational Database (IDB)\u003c/em\u003e [Data set]. https://www.census.gov/programs-surveys/international-programs/about/idb.html\u003c/li\u003e\n\u003cli\u003eUS Department of State (2024). Large Scale International Boundaries [Data set]. https://geodata.state.gov/geonetwork/srv/eng/catalog.search#/metadata/3bdb81a0-c1b9-439a-a0b1-85dac30c59b2\u003c/li\u003e\n\u003cli\u003eVijayaraj, V., Bright E., and Bhaduri, B., (2007). High Resolution Urban Feature Extraction for Global Population Mapping using High Performance Computing, \u003cem\u003eProceedings of the IEEE International geosciences and remote sensing symposium (IGARSS) 2007\u003c/em\u003e. https://doi.org/10.1109/IGARSS10946.2007\u003c/li\u003e\n\u003cli\u003eVijayaraj, V., Bright E., and Bhaduri, B. (2008). Rapid Damage Assessment from High Resolution Imagery, \u003cem\u003eProceedings of the IEEE International Geosciences and Remote Sensing Symposium (IGARSS) 2008\u003c/em\u003e. https://doi.org/10.1109/IGARSS10663.2008\u003c/li\u003e\n\u003cli\u003eWardrop, N., Jochem, W., Bird, T., Chamberlain, H., Clarke, D., Kerr, D., Bengtsson, L., Juran, S., Seaman, V., Tatem, A. (2018). Spatially disaggregated population estimates in the absence of national population and housing census data. \u003cem\u003eProceedings of the National Academy of Sciences\u003c/em\u003e, 115(14), 3529\u0026ndash;3537.\u003c/li\u003e\n\u003cli\u003eWeber, E.M., Seaman, V.Y., Stewart, R.N., Bird, T.J., Tatem, A.J., McKee, J.J., Bhaduri, B.L., Moehl, J.J., Reith, A.E. (2018). Census-independent population mapping in northern Nigeria. \u003cem\u003eRemote Sensing of Environment\u003c/em\u003e, 204, 786\u0026ndash;798. https: //doi.org/10.1016/j.rse.2017.09.024\u003c/li\u003e\n\u003cli\u003eWilliam \u0026amp; Mary geoLab (2024). geoBoundaries [Data set]. https://www.geoboundaries.org.\u003c/li\u003e\n\u003cli\u003eWoody, C., Frazier, T. (2023). Waffle Homes: Utilizing Aerial Imagery of Unfinished Buildings to Determine Average Room Size. In Beecham, R., Long, J.A., Smith, D., Zhao, Q., Wise, S. (eds.) \u003cem\u003eProceedings of the 12th International Conference on Geographic Information Science (GIScience 2023)\u003c/em\u003e (pp. 85:1\u0026ndash;85:6). https://doi.org/10.4230/LIPIcs.GIScience.2023. 85.\u003c/li\u003e\n\u003cli\u003eWorldPop (2024). Top-down estimation modelling: Constrained vs Unconstrained. https://www.worldpop.org/methods/top_down_constrained_vs_unconstrained \u003c/li\u003e\n\u003cli\u003eYang, H.L., Yuan, J., Lunga, D., Laverdiere, M., Rose, A., Bhaduri, B. (2018). Building extraction at scale using convolutional neural network: Mapping of the united states. \u003cem\u003eIEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing\u003c/em\u003e, 11(8), 2600\u0026ndash;2614. https://doi.org/10.1109/JSTARS.2018.2835377\u003c/li\u003e\n\u003cli\u003eYang, H.L., Laverdiere, M., Hauser, T., Swan, B., Schmidt, E., Moehl, J., Reith, A., Adams, D., Morris, B., McKee, J., Whitehead, M., Tuttle, M. (2024). A baseline structure inventory with critical attribution for the us and its territories. \u003cem\u003eScientific Data\u003c/em\u003e, 11. https://doi.org/10.1038/s41597-024-03219-x\u003c/li\u003e\n\u003cli\u003eYang, L., Varma, L., Neunsinger, L., Lunga, D. (2023). Providing Geospatial Intelligence Through a Scalabale Imagery Pipeline. In \u003cem\u003eAdvances in Scalable and Intelligent Geospatial Analytics: Challenges and Applications\u003c/em\u003e. CRC Press. https: //doi.org/10.1201/9781003270928-11\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[{"identity":"48ce5342-c022-4687-b0f9-9cc1e22208d9","identifier":"10.13039/100000005","name":"U.S. Department of Defense","awardNumber":"NA","order_by":0}],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"Oak Ridge National Laboratory","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"population distribution, data fusion, building morphology, building occupancy","lastPublishedDoi":"10.21203/rs.3.rs-6396722/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6396722/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003ePopulation datasets accounting for the full range of routine human activities are needed to address many global human security challenges, including disasters, conflict, and infrastructure demand. LandScan High Definition (HD) supports this need through gridded ambient population estimates that measure average human presence between daytime and nighttime at a high spatial resolution of 3 arcseconds (roughly 90\u0026nbsp;m). Although LandScan HD has traditionally been produced on a country-specific basis, advances in global foundational data and computational resources now enable scaling its methodology to the world. Combining aspects of top-down and bottom-up gridded population methods, LandScan HD allocates subnational population totals from authoritative statistics to built-up areas based on occupancy estimates for multiple facility types (e.g., residential, commercial) and then reaggregates these estimates to a global population grid. We scale this approach by organizing the LandScan HD data stack into a 1° resolution tileset of vector analytic features, enabling an efficient and repeatable workflow for all countries worldwide. Examining the Philippines as an output of the global LandScan HD baseline dataset, we contrast the ambient population with a gridded population representing residential activities (WorldPop) by (1) exploring a practical application for flood risk assessment and (2) evaluating congruence with outcomes of collective human activities (subnational CO\u003csub\u003e2\u003c/sub\u003e emissions). Finally, we discuss confronting current LandScan HD limitations through data/modeling and uncertainty quantification improvements and provide outlook for workflow automation and extending the model to social, demographic and economic population characteristics.\u003c/p\u003e","manuscriptTitle":"LandScan HD: A High-Resolution Gridded Ambient Population Methodology for the World","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-04-09 06:36:32","doi":"10.21203/rs.3.rs-6396722/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"769ee644-a98a-4df8-917a-5602dabfdc1c","owner":[],"postedDate":"April 9th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":46866088,"name":"Geographic Information Systems"}],"tags":[],"updatedAt":"2025-04-09T06:36:32+00:00","versionOfRecord":[],"versionCreatedAt":"2025-04-09 06:36:32","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6396722","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6396722","identity":"rs-6396722","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-06-02T02:00:03.124865+00:00
License: CC-BY-4.0