{"paper_id":"0425d0cf-2689-405d-a4a2-1ac096107d4d","body_text":"A database of disaster impacts in the Global South using Red Cross reports and Large Language Models | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article A database of disaster impacts in the Global South using Red Cross reports and Large Language Models Laura Hasbini, Luca G. Severino, Mariana Madruga de Brito, Gabriela C. Gesualdo, and 5 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8778674/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Damage from natural hazards exacts a heavy toll on society and is expected to increase under climate change. Yet, existing impact datasets remain limited and often biased toward Northern countries and monetary losses. To help address these gaps, we present ROUGE; a new socio-economic impact database obtained using textual operational reports from the International Federation of Red Cross and Red Crescent Societies (IFRC). These reports are systematically collected and provide broad coverage of regions that are commonly underrepresented in existing sources. Using large language models, we extract qualitative and quantitative information on a wide range of non-monetary impacts at national and sub-national scales. The resulting dataset documents socio-economic impacts of natural hazards on the population and the built environment with a spatial detail reaching the subregional level, capturing impacts that are rarely included in conventional databases. This resource is designed to support research and applications that require geographically explicit information on socio-economic impacts of disasters, enabling more precise and inclusive analyses of socio-economic consequences of natural hazards across the world. Luca G. Severino and Laura Hasbini are both first authors and contributed equally to the work. Climate Analysis and Modeling Artificial Intelligence and Machine Learning Natural hazards socio-economic impacts disaster impact database IFRC large language models text mining impact assessment Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Background & Summary Natural hazards exact a heavy toll on society, and have caused over 2 million fatalities and 3.6 trillion Euros in damages worldwide in the past 50 years 1 . This trend is expected to continue in the future, as many disasters become more frequent and severe due to climate change and exposure shifts 2 . Within this context, understanding how natural hazards lead to impacts is crucial for improving risk management and reducing damage. Gaining such understanding demands a comprehensive view of where and when disasters occur, as well as the magnitude of their impacts. While several databases recording the impacts of natural hazards at the global 1 or regional scale 3 do exist, they suffer from geographical biases 4 , which can lead to under-reporting for the Global South 5 and result in overlooking certain risks in the most vulnerable regions. A second gap is that most impact databases focus on monetary losses, disregarding human, social, and environmental losses 6 . This focus undermines the importance of non-monetary, intangible losses and further amplifies the bias towards impacts in the Global North, where more financial and insured assets are concentrated. Finally, only a few databases report impacts at the subnational scale, while information at fine spatial resolution would improve impact modelling precision, especially because hazard information does exist at much higher resolution than the national scale. With the advent of big data and recent advances in computer science and artificial intelligence, a larger volume of data has become exploitable. Artificial intelligence and machine learning tools have emerged as powerful tools for improving the understanding and characterisation of natural hazards 7 and their impacts 8 – 10 . In particular, natural language processing has demonstrated strong potential for extracting information from scattered and unstructured data sources, where manual extraction is too time-demanding. For instance, recent work has shown how natural language processing can be used to automatically identify drought impacts in large collections of news articles, extracting details about the affected sectors, locations, and reported consequences through rule-based and machine learning classifiers 8 . However, the performance of keyword search and rule-based approaches is limited, requiring highly detailed and context-specific dictionaries 11 . Moreover, traditional classification models require large volumes of annotated training data, which are time-intensive to produce, especially for rare impact types 8 . In such a context, large language models (LLM) represent a more flexible and extensive tool for data extraction. Latest studies demonstrate the effectiveness of LLM-based approaches in systematically retrieving and analysing impact information 10 , 12 . Additionally, the use of such approaches in the field of disaster risk management is still at an early stage, and challenges remain in developing methodologies and evaluating the applicability of these methods. In particular, validating the extracted information is crucial, given LLMs' propensity to hallucinate (i.e. produce inaccurate or unsupported information) 13 . This validation task remains time-consuming, especially when reference data is scarce and the data extracted is complex. Against this background, the International Federation of the Red Cross' Societies (IFRC) reports impacts from natural hazards since 1919 14 in the form of written operation reports. Part of the data that can be found in these reports is available in machine-readable format through the IFRC GO platform 14 and the IFRC’s Montandon initiative 15 . These databases only include figures on the number of fatalities, people affected, injured, displaced, and missing at the national level. However, the reports provide information on a wider range of impacts, such as health and infrastructure damage, and at a finer spatial resolution. Hence, these written reports represent a valuable and largely underexplored resource for creating impact datasets. Another advantage of this data is its explicit focus on disasters, which reduces the biases that often affect other impact datasets derived from news 8 , or Wikipedia articles 10 . Moreover, the IFRC provides systematic reporting on events affecting Global South countries, which are frequently underrepresented in global impact datasets. As such, the IFRC reports can provide a more inclusive and balanced perspective on disaster impacts. The combination of societal urgency, new artificial intelligence tools, and increasing volumes of text data provides a unique window of opportunity for studying the impacts of climate extremes. In this study, we leverage data from IFRC's textual reports and LLMs to produce a novel socio-economic impact database: ROUGE (Redcross Operations Unified Global Emergency database). Specifically, we focus on unconventional, non-monetary, qualitative, and quantitative impacts at the national and sub-national levels. Our database records the impacts of 716 disasters in 145 countries for 14 different hazards from 2016 to 2025. The database covers impacts for 20 subtypes of impacts from three main types of impacts, including Human , Infrastructure and Service Access , and Economy and Culture . We expect this database to help the modelling and investigation of disasters and their risks, by providing new data on natural hazards and their impacts in regions which suffer from under-reporting in other conventional databases. Methods The extraction and processing of text data from IFRC reports followed four main steps (Fig. 1 ). First, we extract the reports and convert them to text format. Second, we extract and clean the text to prepare it for automatic extraction (Subsection Preprocessing). Third, we prompt an LLM to identify and extract impact data from the cleaned text (Subsection Data extraction). Finally, during post-processing (Subsection Post-processing), we standardise the extracted outputs, remove inconsistent elements, and geolocate the reported locations (Subsection Geocoding). Additionally, for technical validation (Section Technical Validation), we compare the extracted information with manually labelled data and other established impact databases. Data We use operational reports from the Disaster Response Emergency Fund (DREF) and the Emergency Appeal Fund of the IFRC as source data to extract the impacts of natural hazards. The reports are issued by National Societies of the IFRC in cases of humanitarian crises or disasters, covering small-scale crises (relatively narrow geographical scope and a relatively low number of affected people) to large-scale disasters (a large area of a country or countries is concerned and a very large number of people are affected) 16 . These reports describe the extent of the crisis or disaster, including the area and populations affected, as well as the drivers that lead to the emergency situation (including natural hazards, epidemics, and conflicts). The reports also include detailed descriptions of the impacts on the population, the infrastructure and the environment, as well as the planned response measures for disaster relief. The IFRC identifies events using appealCode that uniquely identify each crisis. For each appealCode , the IFRC can issue several reports that provide updates on the operation's status, including updates and a mandatory final closing report when the emergency appeal has been terminated 17 . In order to avoid extracting duplicate impacts, we select a single report per event from the multiple reports issued. Specifically, we keep the longest available report for each event to minimise information loss, as later updates or final reports may contain more summarised versions of the original details. Furthermore, we only consider reports from April 2016 to ensure structural consistency throughout the reports. Next, we select reports for which the associated disaster type in the IFRC metadata corresponds to a natural hazard from the following list: Drought, Flood, Glacial lake outburst, Cyclone, Hurricane, Typhoon, Storm, Tornado, Heatwave, Coldwave, Mass movement, Earthquake, Volcano, Tidal Wave, Wildfire. As an additional check, we verify whether any of these hazard terms are mentioned in the report title. We scrape the reports from the IFRC website 14 and download them as PDFs. The text is extracted using the PyPDF2 Python package 18 . All images are discarded, and only text, including tables, is retained. This results in 717 reports, describing unique disaster events in 141 countries from April 2016 to July 2025. For technical validation, we used the EMDAT 1 and IFRC Go 14,19 impact datasets. EM-DAT is a widely used impact data provided by the Centre for Research on the Epidemiology of Disasters in open access 1 . It contains impact information from more than 27000 events worldwide from 1900 to the present day. The dataset gathers information from various sources such as UN agencies, non-governmental organisations, reinsurance companies, research institutes, and press agencies. It targets only major events that have resulted in at least 10 deaths and/or 100 affected people and/or were declared as an emergency by the state and/or were associated with an international assistance call. Precise information about the location impacted can be provided, but the impact is only given at the country level. To support systematic impact tracking, the IFRC has developed several initiatives aimed at collecting and structuring disaster impact information, such as the IFRC GO. This database compiles hazard events and associated impact data from IFRC DREF reports. It provides structured impact information for 1476 events recorded between 1997 and 2025. While this dataset offers valuable opportunities for impact comparison, its scope is limited to a narrow set of human and monetary impact categories. Preprocessing In order to ensure machine-readability and facilitate the use of LLMs, we conduct several preprocessing steps on the raw text from the reports using the Spacy python library 20 . First, we identify the language of each report using a language detection pipeline 21 . Results show that only four reports are not written in English. These reports are discarded to avoid needing to adapt our extraction pipeline to other languages. Next, we remove unwanted text and characters such as URLs, ligatured characters, line breaks and multiple spaces. Based on the observed repeated structure of the reports (see supplementary material), we extract only the sections of the report that contain information about the hazard and its impacts. This limits the text we process, from an average of 244 sentences per report before selection down to an average of 78 sentences per report after selection, making extraction faster and more focused. Impact data typology and database structure We define a socio-economic impact typology comprising 3 impact types and 20 subtypes as detailed in Table 1 . The typology was defined using a combination of inductive and deductive approaches, incorporating results from a data-driven topic analysis of the reports and partly adapting the impact types and subtypes of the EM-DAT and DESInventar frameworks. Besides the impacts, we also extract information such as the hazard type, location, and dates. The hazard typology has been adapted from EM-DAT hazard typology 1 and is described in Table 2 . Table 1 Impact types and description Impact Maintype Impact Subtype Description Human Affected People Total number of individuals impacted by the hazard event. Injured People Number of people injured, including those hospitalised or admitted. Displaced People Number of people forcefully displaced or evacuated before or following the event. Homeless People Number of people losing housing following the event. Missing People Number of people unaccounted for following the event. Human Deaths Number of fatalities caused by the hazard. Infected and Ill People Number of people contaminated (cases) by an infectious disease. Human Health and Wellbeing Generic impacts on human health (physical, mental) or wellbeing not necessarily associated with the spread of an infectious disease. Infrastructure and Service access Transportation - Roads and transportation infrastructure (e.g. bridges, highways..) impacted by a hazard. - People losing the capacity to move safely and efficiently between locations. Water, Sanitation, and Hygiene - Number of water, sanitation, and hygiene infrastructure such as sewage networks, drainage systems, wastewater treatment plants, etc. impacted by a hazard. - People losing safe, clean, and consistent supply of water for drinking, sanitation, hygiene services and other essential uses. Healthcare - Number of healthcare infrastructure such as hospitals, healthcare centers, pharmacies, clinics, etc. impacted by a hazard. - People losing the ability to obtain needed medical services, including preventative care, emergency services … IT and Communication - Number of IT and communication infrastructure such as data centers, communication towers, and cables impacted by a hazard. - People losing communication access Residential Buildings - Number of residential buildings (e.g. houses) impacted by a hazard. Informal settlements - Number of informal settlements such as refugee camps, slums, tents, etc. impacted by a hazard. Education - Number of education infrastructures such as schools, universities, etc. impacted by a hazard. - People losing the ability to attend educational institutions and receive instruction. Power and Energy Production - Number of energy production infrastructures such as power plants, turbines, grids, pipelines, etc., impacted by a hazard. - People losing access to electricity and services related to power production Agriculture and Access to Food - Number of agricultural infrastructures such as farms, warehouses, greenhouses, fisheries, etc. impacted by a hazard. - Number of crops, agricultural production and forest impacted by a hazard - Total number of animals, including terrestrial and aquatic species, impacted by the hazard. (e.g. perished animals or fishes) - People losing access to a secure food supply Undefined Infrastructure and Service Access - Any identified impact on infrastructure where the type of infrastructure impacted is not clearly defined, e.g. critical infrastructure, public infrastructure, access to basic services Economy & Culture Recreation, Tourism, and Culture - Tourist attractions and cultural sites impacted by a hazard. Economic and Livelihood - Any identified impact on the economy or living conditions resulting from a natural hazard. Table 2 Hazard types and description Hazard Description Drought Prolonged lack of precipitation Wildfire Uncontrolled natural fires Earthquake Sudden tectonic shifting. Includes as well tsunami Mass movement Any type of downslope movement of earth materials. Includes Landslides and rockfalls Volcanic activity Eruptions and related phenomena Flood River, coastal, flash and ice jam flooding Wave action Wind-generated surface waves (e.g. Rogue waves, seiche) Extreme warm temperature Prolonged, abnormally high heat Extreme cold temperature Prolonged, abnormally low cold Tropical storm Tropical cyclonic storms (e.g tropical cyclone, tropical storm, hurricanes, typhoons) Convective storm Convective-related storms (e.g. severe convective storms, derechos, hail, tornado) Other storm Any type of storm which does not correspond to a tropical cyclone or convective storm (e.g. extra-tropical storm, snow storm…) Epidemic Sudden outbreak of disease Conflicts Disagreements or disputes between different groups, organizations, or states Impact data extraction We select the model meta-llama/llama-4-scout-17b-16e-instruct 22 , to identify and extract natural hazard impacts, as it is open-source and offers the best performance-to-price ratio compared to other models tested (in supplementary material). We use the GROQ cloud platform 23 to run the extraction. The extraction of the impact data is done using a sequence of prompts 24 . For each report, we prompt the LLM with a series of queries. This allows us to better guide the extraction by providing the LLM with more precise information at each extraction step, to have better control over the outputs and to limit the length of the output context windows. The extraction sequence is designed as follows: Identify subtypes of impacts from the text, given a pre-defined list of impact subtypes and their definitions . The output is the list of identified subtypes. For each identified subtype of impacts, extract the values and units of those impacts . Extracting all impacts at once prevents the reuse of the same values for different identified subtypes. The output is a list of dictionaries. For each identified impact (defined as a subtype, a value, and a unit), identify the affected locations . We allow a single impact to have multiple affected locations to account for aggregated reports of impacts. The output is a dictionary. For each identified impact (defined as a subtype, a value, a unit, and a location), identify the starting and ending dates . The output is a dictionary. For each identified impact (defined as a subtype, a value, a unit, a location, and a date), identify the hazards driving the impacts . We allow for multiple hazards to account for compound events. At each step, we use the pydantic library 25 to validate the output. This validation allows us to control that each desired field is (i) present in the extracted json, (ii) of the correct data type e.g. string, integer and (iii) contains only allowed values when value restrictions apply. The validation on the value constraints is only applied to the hazards field, to ensure that only hazards matching the hazards lists are extracted. The validation of value constraints for the impactSubtype field was disabled to avoid the LLM relabelling unwanted data. If validation fails, we reprompt the LLM once with the validation error so it can correct the invalid output. Furthermore, we purposely extract unwanted “control” subtypes, such as the money raised by the appeal or the number of assisted people, as this prevents these figures from being wrongly extracted and classified as other similar impact subtypes, such as Economy and Livelihood or Affected People . These control subtypes are later discarded from the database. To ensure traceability and for validation purposes, we also query the LLM to obtain the exact text excerpt it used for data extraction at each step. The extracted data is stored in a CSV file with the structure defined in the Data Descriptor section. Post-processing We conduct the following post-processing steps. The impactValue is evaluated to ensure that it falls within the defined bounds, such that \\(\\:impactValueMin\\:\\le\\:\\:impactValue\\:\\le\\:\\:impactValueMax\\) . We reclassify and standardise the impact subtypes and hazard types using regular expressions. We format the impactUnit with the following four sub-steps; (i) we detect numbers from the unit field and move them to the impactValue field (e.g. 10 thousand people becomes 10’000 people ), (ii) we convert monetary units to EUR using the price parser 26 and currency converter 27 packages, (iii) we then standardise metric units (i.e. distance, mass, surface, volume) to SI units, and (iv) we standardise “measured” units using regular expressions and pre-defined conversion factors (e.g. 4 households becomes 12 people , 2 maternities becomes 2 healthcare structures ). Finally, we perform the following final cleaning steps. We (i) remove duplicated rows, (ii) filter out unknown impact subtypes and the control subtypes introduced during the extraction, and (iii) filter out unknown impact units that could not be standardised in the previous post-processing steps. Geocoding We convert textual location descriptions into georeferenced polygons. This step is done not only to check the correctness of the location identified by the LLM but also to enrich it with geographic boundaries. For administrative boundaries, we rely on the geoBoundaries dataset 28 , which compiles information from governmental and open-license sources. It provides Polygons or Multi-Polygons geometries for all countries, with administrative levels available up to admin 5. Since administrative levels beyond 2 (e.g. subdistricts, villages, neighbourhoods) are not consistently available across all countries, we restrict our analysis to level 2. The textual locations are first converted to geographic information using the OpenStreetMap (OSM) Nominatim tool 29 . This service returns an OSM address with a geometry in the form of a Point, Polygon, or MultiPolygon object. An OSM address may include multiple address details (also known as address keys 30 ), such as city, suburb, hamlet, street, and district. In order to convert textual location strings to polygons we first compute the Levenshtein similarity 31 between the input location string and all OSM address keys, with and without administrative descriptors and across rotated word orders. If a perfect match (1.0) is found, the search stops; otherwise, we keep the address key with the highest similarity above a 0.2 threshold, which identifies the most likely location name and administrative level, correcting spelling variations introduced by the LLM. The corrected name is then matched to the corresponding geoBoundaries polygon at the identified administrative level. If matching with a polygon fails, we apply several fallback strategies. First, the location name may differ from the official geoBoundaries designation, so we iterate over other OSM address keys at the same administrative level and attempt to match each to a polygon. If this fails, we translate the location name into the country’s official languages and widely-spoken languages (French, Spanish, German) and retry the match. If no textual match is found, we use the geographic object returned by OSM and search for a corresponding geoBoundaries polygon by testing for intersections at the identified administrative level. If all attempts fail, we repeat the procedure at progressively coarser administrative levels until a match is found. This process yields a polygon for each textual location. Because reported impacts may reference multiple locations at different administrative levels, the polygons of each resolution identified are kept in the final geometry. For each post-processing step, we introduce quality flags to ensure traceability and data quality and allow users to flexibly select or filter out data according to their needs. The flags and their descriptions are detailed in the database structure table (Table 3 ) Table 3 Structure of the impact database Variable Format Description impactType str Impact type from predefined list impactSubtype str Impact subtype from predefined list impactValue float or None Numerical value quantifying the impact impactValueMin float or None Lower bound estimate of the impactValue impactValueMax float or None Upper bound estimate of the impactValue impactUnit str Unit of impactValue startYear int or None Start year of impact startMonth int or None Start month of impact startDay int or None Start day of impact endYear int or None End year of impact endMonth int or None End month of impact endDay int or None End day of impact hazards list List of hazards associated with the described impact location list List of locations associated with the described impact locationLowestAdmin str Lowest administrative level identified with the geocoding locationPolygon str Name of the location polygon identified with the geocoding geometry Polygon or MutiPolygon MultiPolygon of the locations identified with the geocoding country_iso3 list List of iso3 codes of the country identified with the geocoding continent list List of continents identified with the geocoding Flags : impactPrecision str Flag describing whether the impactValue is exact or approximate valid_errors_impactValue int Validation error impactValue extraction. Count of validation errors during the extraction of the impactValue. valid_errors_loc int Validation error location extraction. Count of validation errors during the extraction of the location. valid_errors_dates int Validation error dates extraction. Count of validation errors during the extraction of the dates. valid_errors_haz int Validation error hazards extraction. Count of validation errors during the extraction of the hazards. flag_value_not_in_text bool ImpactValue should be in the original text. True if impactValue cannot be found in original text flag_pop_cntry bool ImpactValue with “people” unit should be smaller than country’s population. True if impactValue with people unit is higher than country population flag_impactSubtype_reclass bool Inferred impactSubtype must be in the allowed list. True if extracted impactSubtype was not in the allowed list and has been reclassified via regular expression flag_hazards_reclass bool Inferred hazardType must be in the allowed list. True if extracted hazardType was not in allowed list and has been reclassified via regular expression flag_unit_conversion bool Unit conversion. True if a non-metric unit has been converted. These include families -> 3*people, households -> 3*people, communities -> 100*people, villages -> 1000*people flag_non-SI_unit_standardization bool Standardization of non-metric unit. True if a unit has been reclassified using regular expression, e.g. schools -> education structures, bridges -> roads flag_SI_unit_standardization bool Standardization of metric unit to SI. True if the extracted impactUnit has been standardized to a SI unit. flag_remove_number_unit bool True if a number has been identified from the extracted impactUnit and has been moved to the impactValue from the impactUnit (e.g. impactValue = 10, impactUnit=”milion people” => impactValue = 1000000, impactUnit=”people”) flag_currency_conversion bool True if a monetary unit has been identified but has been failed to standardized to the default currency flag_unit_processing_error bool True if there has been an issue in any of the following stages of the postprocessing of the impact unit: remove_number_unit, unit_conversion, SI_unit_standardization, non-SI_unit_standardization, currency_conversion. flag_unit_nonstd bool True if the extracted impactUnit is not part of the “standard units” flag_reclass_subtype_from_unit bool True if the extracted impactSubtype has been reclassified based on the impactUnit flag_value_no_unit bool True if an impactValue has no unit flag_partial_unit bool True if the impactUnit lacks the measured unit (e.g. “hectares” instead of “hectares of crops”) flag_geocoding_osm int Number of locations for which osm point/polygon has been used to retrieve the GAUL/NE polygon, with spatial matching flag_geocoding_country int Number of locations downgraded to the country to find a polygon Metadata : appealCode str Appeal code of IFRC report reportDate str Date of the report in YYYY/MM/DD format reportLink str Url link of the pdf report disasterType_IFRC str Disaster type identified by the IFRC sourceExcerpts str Text excerpts used for the information extraction Data Records We extract a database with the structure described in Table 3 . Each row of the database represents one impact record, specified by an impact subtype, a value, a unit, a location (geometry), a date, and a driver (associated hazards). Each impact record is also complemented with metadata (Table 3 ), informing on the source (i.e. the appeal code uniquely identifying an IFRC appeal, the report date, the sentences that were used for the extraction of the impacts), the precision of the impact value, and several quality flags, which track issues and possible inconsistencies. The data will be made accessible via Zenodo under the Creative Commons Attribution 4.0 International license. The database is available in two formats: as a global csv file without geometries and as Parquet files with geometries, available globally and split by continents. A notebook detailing how to open and access the data is available in the code archive, which can be downloaded anonymously by the reviewers at https://figshare.com/s/99da5ff800fc56ae10c0 . The data can be downloaded anonymously by reviewers at the same link. Data Overview The LLM-based impact dataset is created from IFRC reports and primarily captures impact in the Global South, as illustrated in Fig. 2 . The regions most affected are Africa and Asia, with 5760 and 4525 impacts reports (Fig. 2 ), respectively, covering 707 and 796 different natural hazards (Fig. 3 ). Figure 3 shows that floods, along with tropical and convective storms, are the hazards most frequently associated with the extracted impacts. Additionally, mass movements and epidemics also contribute substantially to the overall recorded impacts. 54% of the impacts extracted are geocoded at the country level (ADMIN0), 36%at the first-order administrative division (ADMIN1), and 10% at the second-order administrative division (ADMIN2). Technical Validation Comparison with manually labelled data (internal validation) We validate our database internally by comparing LLM-extracted impacts with manually extracted impacts from 17 randomly selected reports (listed in the supplementary). Manual extraction followed strict, pre-defined guidelines (see supplementary Labelling Guidelines ), and was performed by the authors with cross-reviewing. The labelled data acts as a control and represents the reference standard to which the LLM-extracted data is expected to adhere. From the 17 reports selected, we manually labelled 361 quantitative and 169 qualitative impacts, spanning over 20 impact subtypes and 13 hazards. In comparison, the automatic extraction using the LLM extracted 243 quantitative and 131 qualitative impacts, spanning over 20 impact subtypes and 9 hazards. Comparing the LLM-extracted with the manually-labelled impacts allows us to derive similarity scores for the different variables extracted, as well as the precision and recall of the extraction. The similarity scores are defined as a function of the type of the examined variable and range from 0 for perfect dissimilarity to 1 for perfect similarity. Similarity of categorical variables is expressed using cosine similarity 32 , similarity of numerical variables ( impactValue ) with a normalised relative difference (Eq. 2 in the supplementary Matching algorithm for internal validation ), and geometries with the Intersection-over-Union (Eq. 1 in the supplementary Matching algorithm for internal validation ). Precision and recall represent, respectively, the percentage of correctly extracted data over the total extracted and the percentage of correctly extracted data over the total manually-labelled. In order to compute the similarity scores for each extracted impact, we use an algorithm that matches each LLM-extracted row to the most similar row in the manually-labelled data. The matching procedure is described in the supplementary Matching algorithm for internal validation . Figure 4 shows the average similarity per matching variables, separated by the quantitative and qualitative extraction. The match similarity is the weighted mean of the variables’ similarities for the matched data. We explain the lower match similarity in the qualitative data by the lower similarity of the impact subtypes, with some qualitative impact subtypes being more abstract, which may cause the LLM to perform worse at classifying the data into the correct subtypes. The similarity scores for the remaining variables are generally high, with scores above 80% for the impact unit, the iso3 code of the countries, the start year, and the impact values. The similarity of the dates is lower (around 50%). We attribute this lower performance to the complexity of clearly identifying event dates, which the human labellers encountered during manual labelling. Next, we compute the Precision and Recall using the matched data to estimate the predictive performance of our extraction: $$\\:Precision=\\frac{TP}{TP+FP}$$ 1 $$\\:Recall=\\frac{TP}{TP+FN}$$ 2 where TP represents the True Positives, i.e. the number of LLM-extracted impacts for which we find a corresponding manually-extracted impacts, FP the False Positives, i.e. the number of LLM-extracted impacts for which we do not find a corresponding manually-extracted impacts, and FN the False Negatives, i.e. the number of manually-extracted impacts for which we did not find a corresponding LLM-extracted impacts. A perfect association should result in a precision and a recall score of 1. Lower precision corresponds to a higher tendency of the LLM to hallucinate, while lower recall indicates that the LLM is missing a larger portion of the true impacts. To quantify the sampling uncertainty, we approximate the 95% confidence intervals of the precision and recall metrics using a bootstrapping approach. For each sample from which the metrics are estimated, we sample with replacement 5000 random samples of true positives, false positives, and false negatives. We then compute the metrics from each random sample to obtain the bootstrapped distribution of the metrics. Finally, we use the 2.5th and 97.5th quantiles of the bootstrapped distribution as the bounds of the confidence interval. Figure 5 a shows precision and recall , and their estimated 95% confidence intervals for the total number of quantitative and qualitative impacts extracted. We find a precision of 0.71 and 0.62 for the quantitative and qualitative data, respectively. These results suggest that quantitative data can be extracted more precisely than qualitative data. This could indicate that the model is better at extracting numbers than at identifying more abstract entities, such as qualitative descriptions of impacts. Regarding recall, we obtain 0.37 for the quantitative data and 0.52 for the qualitative data. This suggests that only one third of the quantitative impacts that are described in the reports can be extracted with our approach. The extraction of qualitative impacts performs slightly better. We explain this by the higher number of quantitative (18 on average per report) than qualitative (10 on average per report) impacts present in the reports. Figure 5 b shows the precision and recall metrics per impact subtype. The 95% confidence intervals indicate that sampling uncertainty is high due to the smaller sample sizes when grouping by impact subtype. For 7 of the 20 impact subtypes ( Infected and Ill People, Homeless People, Missing People, Power and Energy Production, Recreation, Tourism, and Culture, IT and Communication , and Undefined Infrastructure and Service Access ), the sampling uncertainty is too large to exclude the precision and recall to be 0. For the remaining subtypes, their performance does not differ significantly from the average precision and recall. In addition, we ran robustness checks to assess whether our approach is sensitive to the choice of the LLM used and if it is deterministic (see supplementary Sensitivity analyses ). Comparison with external data (external validation) We perform external validation by comparing quantitative impacts extracted by the LLM at the national level to EMDAT and IFRC-Go datasets. From our impactSubtypes (Table 1 ) we can compare the Affected People , Human Death and Homeless People categories with EMDAT and the Affected people , Injured people , Displaced People , Human Deaths and Missing People with IFRC-Go. As EM-DAT uses the IFRC to source their records of the number of people affected in 65% of cases 1 , we expect this category to match well across the three data sources. From the LLM-extracted data, impacts are aggregated at the ADMIN0 level by retaining, for each record, the highest impact value reported at the coarsest administrative level. Square symbols in Fig. 6 illustrate the percentage of events for which a quantitative impact was identified across the previously mentioned common impact subtypes. The three databases show differing proportions of events with recorded quantitative impacts. Counts of Affected People and Human Deaths are the most frequently reported, with EM-DAT consistently exhibiting the highest percentage of reports containing quantitative impact values. For the Injured People , Displaced People , and Missing People subtypes, IFRC Go reports quantitative impacts for approximately 40% of events on average. In contrast, the corresponding proportions derived from LLM-based extraction are lower, dropping to 24% for Injured People and 18% for Missing People. Overall, this indicates that the LLM captures quantitative impacts more effectively for widely reported subtypes such as Affected People and Human Deaths , while it appears to struggle with less consistently documented subtypes, which are also under-documented in EM-DAT and IFRC-Go, including Injured People , Displaced People , and Missing People . We then match the reports containing quantitative information to the external datasets in order to compare the values of the reported impacts. Because EM-DAT and IFRC do not share a common unique event identifier, we match events using a stepwise procedure. We first link each IFRC report to EM-DAT events occurring in the same year and country. When multiple hazards are listed in an IFRC report, we select the hazard associated with the greatest number of reported impacts and retain only the EM-DAT events whose DisasterType corresponds to this dominant hazard. If multiple EM-DAT events still remain after this filtering, we select the event whose entry date is closest to the IFRC report date. This systematic procedure associates an event from EMDAT to each of the 594 events for which we extract quantitative information with the LLM. For IFRC-Go, matching can be done directly using the appeal code. We extract 566 events from the IFRC-Go data API over the period 2016–2025, which are associated with natural disasters. We successfully link 383 of these events with events which we extracted. Bar plots from Fig. 6 show the percentage of events from our database that we can match to the two external databases. The colours of the bars represent the percentage of matching against one of the two external databases. The brightness of the colours represent the percentage of matches for which we find (i) no impact value, (ii) an impact value differing from the LLM-extracted value by more than 10% and (iii) an impact value differing from the LLM-extracted value by less than 10% (see section Comparison with manually labelled data ). We do not find a match within the IFRC-Go database for approximately 40% of the LLM-extracted events. This implies that our database retrieved impacts for appealCode which are not available within the IFRC-Go data API. Over the events matched, we do not find an impact record for the different impact subtypes considered in about 30% of the cases for both the IFRC-Go and EM-DAT. This suggests that, for the events considered, the LLM-based extraction is more robust in automatically detecting impacts for the considered impact subtypes. However, the degree of agreement between the impact values we extract and the external databases remains variable across the different impact subtypes. Overall, we observe a higher percentage of matched events with quantitative impact figures in EM-DAT compared to IFRC-Go, which is inline with the higher percentage of total events with impacts present in EM-DAT (square symbols). Interestingly, exact matches between LLM-extracted values and EM-DAT occur more often than between LLM-extracted values and IFRC-Go. On average across the different subtypes, we find exact value matches in less than 10% of the cases with IFRC-Go, and less than 30% of the cases with EM-DAT. Possible explanations include differences in the specific report versions used in the case of IFRC-Go, limitation of the matching procedure for EM-DAT or inaccuracies in the extraction process in either the LLM, IFRC-Go or EM-DAT databases. Nonetheless, these discrepancies across sources and impact categories further highlight the inherent challenges in estimating quantitative impact information from natural disasters. Declarations Competing Interests The authors declare no competing interests. Funding This work was supported by the French ANRT, which funded LH’s PhD. L.H. is employed by Generali France. Author Contributions LH: Conceptualisation, Methodology, Software, Validation, Formal analysis, Data Curation, Writing - Original Draft; Visualisation L.S.: Conceptualisation, Methodology, Software, Validation, Formal analysis, Data Curation, Writing - Original Draft; Visualisation MMdB: Conceptualisation, Writing - Review & Editing; GCG : Conceptualisation, Data Curation, Writing - Review & Editing AMR : Conceptualisation, Data Curation, Writing - Review & Editing DNB : Supervision, Funding acquisition, Writing - Review & Editing EM : Supervision, Writing - Review & Editing JW : Conceptualisation TMNC: Conceptualisation, Data Curation, Supervision, Writing - Review & Editing. Acknowledgements The authors thank Pascal Yiou (LSCE, UMR 8212CEA-CNRS-UVSQ) for reviewing the manuscript. The authors are grateful to the organisers and participants of the 2024 Como Training School on Compound climate-related Events for providing the structure and motivation which started this work. Data Availability The data will be made accessible via Figshare in csv format without the geometries and in Parquet format with the geometries, under the Creative Commons Attribution 4.0 International license. The data can be downloaded anonymously by reviewers at: https://figshare.com/s/99da5ff800fc56ae10c0 Code Availability The code to produce the database, perform the technical validation and make the analysis plots is available anonymously at: https://figshare.com/s/99da5ff800fc56ae10c0 . References Delforge D et al (2025) EM-DAT: the Emergency Events Database. Int J Disaster Risk Reduct 124:105509 Intergovernmental Panel On Climate Change (Ipcc) (2023) Climate Change 2022 – Impacts, Adaptation and Vulnerability: Working Group II Contribution to the Sixth Assessment Report of the Intergovernmental Panel on Climate Change. Cambridge University Press. 10.1017/9781009325844 Reduction UNO (2022) for D. R. Global Assessment Report on Disaster Risk Reduction 2022: Our World at Risk: Transforming Governance for a Resilient Future United Nations Publications, Bloomfield Gall M, Borden KA, Cutter SL (2009) When Do Losses Count? Six Fallacies of Natural Hazards Loss Data. Bull Am Meteorol Soc 90:799–810 Osuteye E, Johnson C, Brown D (2017) The data gap: An analysis of data availability on disaster losses in sub-Saharan African cities. Int J Disaster Risk Reduct 26:24–33 De Brito MM et al (2024) Uncovering the Dynamics of Multi-Sector Impacts of Hydrological Extremes: A Methods Overview. Earths Future 12:e2023EF003906 Camps-Valls G et al (2025) Artificial intelligence for modeling and understanding extreme weather and climate events. Nat Commun 16:1919 Sodoge J, Kuhlicke C, De Brito MM (2023) Automatized spatio-temporal detection of drought impacts from newspaper articles using natural language processing and machine learning. Weather Clim Extrem 41:100574 De Madruga M, Sodoge J, Kreibich H, Kuhlicke C (2024) Comprehensive Assessment of Flood Socioeconomic Impacts Through Text-Mining. Water Resour. Res. 61, eWR037813 (2025) Li N et al (2025) Wikimpacts 1.0: A new global climate impact database based on automated information extraction from Wikipedia. EGUsphere 1–43 10.5194/egusphere-2025-4891 Adnan K, Akbar R (2019) An analytical study of information extraction from unstructured and multidimensional big data. J Big Data 6:91 Nunes Carvalho TM, de Souza Filho F (2024) Madruga de Brito, M. Unveiling water allocation dynamics: a text analysis of 25 years of stakeholder meetings. Environ Res Lett 19:044066 Huang L et al (2025) A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions. ACM Trans Inf Syst 43(1–42):55 International Federation of Red Cross Crescent(IFRC) (2025) IFRC GO Monty Extension Documentation https://ifrcgo.org/monty-stac-extension/ International Federation of Red Cross Crescent(IFRC). Emergency Response Framework (2025) DREF Guidelines (2020) | IFRC. https://www.ifrc.org/document/dref-guidelines-2020 (2021) Fenniak M et al (2022) The PyPDF2 library Wiki GO GO wiki https ://go-wiki.ifrc.org/en/home Montani I et al (2023) explosion/spaCy: v3.7.2: Fixes for APIs and requirements. Zenodo https://doi.org/10.5281/zenodo.10009823 Spacy, FastLang · spaCy Universe. Spacy FastLang https://spacy.io/universe/project/spacy_fastlang meta-llama/Llama (2025) -4-Scout-17B-16E-Instruct · Hugging Face. https://huggingface.co/meta-llama/Llama-4 -Scout-17B-16E-Instruct GroqCloud - Build Fast. https://console.groq.com Tonmoy SMTI et al (2024) A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models. Preprint at https://doi.org/10.48550/arXiv.2401.01313 Release v2 11.9 2025-09-13 · pydantic/pydantic. https://github.com/pydantic/pydantic/releases/tag/v2.11.9 scrapinghub/price- parser at 0.4.0. https://github.com/scrapinghub/price-parser/tree/0.4.0 alexprengere/currencyconverter at v0.18.9 https://github.com/alexprengere/currencyconverter/tree/v0.18.9 Runfola D et al (2020) geoBoundaries: A global database of political administrative boundaries. PLoS ONE 15:e0231866 OpenStreetMap OpenStreetMap https://www.openstreetmap.org/ Place Output Formats - Nominatim 5.1.0 Manual. https://nominatim.org/release-docs/latest/api/Output/#addressdetails Levenshtein V (1966) Binary codes capable of correcting deletions, insertions, and reversals. in Soviet physics-doklady vol. 10 Salton G, Wong A, Yang CS (1975) A vector space model for automatic indexing. Commun ACM 18:613–620 Additional Declarations The authors declare no competing interests. Supplementary Files SupplementaryMaterial.docx Supplementary Material ImpactLabellingrules.docx Impact Labelling rules SupplementaryMaterial.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {\"props\":{\"pageProps\":{\"initialData\":{\"identity\":\"rs-8778674\",\"acceptedTermsAndConditions\":true,\"allowDirectSubmit\":true,\"archivedVersions\":[],\"articleType\":\"Research Article\",\"associatedPublications\":[],\"authors\":[{\"id\":585208194,\"identity\":\"e3ad148f-0e70-4e21-8fb0-2de51552cded\",\"order_by\":0,\"name\":\"Laura Hasbini\",\"email\":\"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABFElEQVRIiWNgGAWjYBACPgh1gIFBgrGBIaHCAsa1YGBvwK6FDVXLGQk4l4HnAEEtQIqxTQImgUcL+xmzDz8Y7sjLz25ue/BwnoQ9f3sD4+GCP0At0tj1sPHkGM/sYXhmuOHOwXaDxG0SiTPOHGA4PBNoHQ9fAg6H5Rgz8DAcZtwgkdgmAdSSYCCRwHCYt0GCwZ4Hh8P43xgz/mE4bD9/BkjLHAl7A/kHDId5QA7DpUUix5gZaEtiww2QlgYJoHUMQC1s+LQ8K2aWMTicDPRLm0TCMZBfEhtAfuHBpYWfP3kz45uKw7bzZ7c/k/xRYwMMscOHPxf8sZHDpQUCDFB4jA3MQBKvBkzATJryUTAKRsEoGOYAANgSVVScRm+YAAAAAElFTkSuQmCC\",\"orcid\":\"https://orcid.org/0009-0000-6875-1386\",\"institution\":\"1Laboratoire des Sciences du Climat et de l’Environnement, UMR 8212CEA-CNRS-UVSQ, Université Paris-Saclay, Gif-sur-Yvette, France 2Generali France SAS, 93210, Saint Denis, France\",\"correspondingAuthor\":true,\"prefix\":\"\",\"firstName\":\"Laura\",\"middleName\":\"\",\"lastName\":\"Hasbini\",\"suffix\":\"\"},{\"id\":585211207,\"identity\":\"41719069-40d3-4179-81f6-23ce50f2f7ea\",\"order_by\":1,\"name\":\"Luca G. Severino\",\"email\":\"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABEElEQVRIiWNgGAWjYBACPgSTsYGBoQKJzSCBXQsbqpYzUPZB4rSAdLURo0Ui+eEHhj82if3Sh9sefJx32N7g/AHGxx931DHwz27AoSXNWIKxLS1xZl9iu+HMbYcTN9xIYDY4eIaNQeLOARxachgkGBsOGxucYWyT5t12OMHgBlDwYBsPg4FEAi4tzD8Y/vw3tgdrmQN2GPuPg20S+LSwSTCwHZAz4AFpaTjMuOFAAhvDwTYD3Fp4nplZJLYly0kAbZGccSw9ceaNxGaJs20JPBI3sGvhZ09+fOPDHzse/h72ZxIfaqzt+c4fPvihsq1Ojn8Gdi1ggCKlcAAc9Qw8uNWjA/kG4tWOglEwCkbByAAA+YdaKX+/9jYAAAAASUVORK5CYII=\",\"orcid\":\"https://orcid.org/0009-0008-4002-6535\",\"institution\":\"3Institute for Environmental Decisions, ETH Zurich, Universitätstr. 22, 8092 Zurich, Switzerland 4Federal Office of Meteorology and Climatology MeteoSwiss, Operation Center 1, P.O. Box 257, 8058 Zurich-Airport, Switzerland\",\"correspondingAuthor\":true,\"prefix\":\"\",\"firstName\":\"Luca\",\"middleName\":\"G.\",\"lastName\":\"Severino\",\"suffix\":\"\"},{\"id\":585211208,\"identity\":\"76bec02c-e69e-4403-9da5-4f04bc82b83f\",\"order_by\":2,\"name\":\"Mariana Madruga de Brito\",\"email\":\"\",\"orcid\":\"https://orcid.org/0000-0003-4191-1647\",\"institution\":\"5 Helmholtz-Centre for Environmental Research, Department of Urban and Environmental Sociology, Leipzig, Germany\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Mariana\",\"middleName\":\"Madruga\",\"lastName\":\"de Brito\",\"suffix\":\"\"},{\"id\":585211209,\"identity\":\"c32bd5cb-a6e8-4417-9136-ef5193ab89a1\",\"order_by\":3,\"name\":\"Gabriela C. Gesualdo\",\"email\":\"\",\"orcid\":\"https://orcid.org/0000-0001-6589-3397\",\"institution\":\"6 Department of Geosciences, The Pennsylvania State University, State College, PA, United States\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Gabriela\",\"middleName\":\"C.\",\"lastName\":\"Gesualdo\",\"suffix\":\"\"},{\"id\":585211210,\"identity\":\"8a0bf9ec-4e71-4695-9fd4-9a43897db70d\",\"order_by\":4,\"name\":\"Ana Maria Rotaru\",\"email\":\"\",\"orcid\":\"https://orcid.org/0009-0008-8374-6348\",\"institution\":\"7 Department of Civil and Environmental Engineering, Politecnico di Milano, Milan, Italy\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Ana\",\"middleName\":\"Maria\",\"lastName\":\"Rotaru\",\"suffix\":\"\"},{\"id\":585211211,\"identity\":\"c799bff6-116a-4c47-bcc4-4c4c3d998e2a\",\"order_by\":5,\"name\":\"David N. Bresch\",\"email\":\"\",\"orcid\":\"https://orcid.org/0000-0002-8431-4263\",\"institution\":\"3Institute for Environmental Decisions, ETH Zurich, Universitätstr. 22, 8092 Zurich, Switzerland 4Federal Office of Meteorology and Climatology MeteoSwiss, Operation Center 1, P.O. Box 257, 8058 Zurich-Airport, Switzerland\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"David\",\"middleName\":\"N.\",\"lastName\":\"Bresch\",\"suffix\":\"\"},{\"id\":585211212,\"identity\":\"3c27de90-08b4-4bdb-8e29-45103d2f7575\",\"order_by\":6,\"name\":\"Evelyn Mühlhofer\",\"email\":\"\",\"orcid\":\"https://orcid.org/0000-0002-5587-9070\",\"institution\":\"4Federal Office of Meteorology and Climatology MeteoSwiss, Operation Center 1, P.O. Box 257, 8058 Zurich-Airport, Switzerland\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Evelyn\",\"middleName\":\"\",\"lastName\":\"Mühlhofer\",\"suffix\":\"\"},{\"id\":585211213,\"identity\":\"07913272-18a4-4687-8a57-c38838a6f5a6\",\"order_by\":7,\"name\":\"Jingxian Wang\",\"email\":\"\",\"orcid\":\"https://orcid.org/0000-0001-7650-9987\",\"institution\":\"8 University School for Advanced Studies IUSS Pavia, Pavia, Italy 9 Department of Electronics, Information and Bioengineering, Politecnico di Milano, Milan, Italy\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Jingxian\",\"middleName\":\"\",\"lastName\":\"Wang\",\"suffix\":\"\"},{\"id\":585211214,\"identity\":\"c0340b83-48fd-4d2d-a565-ddd59ebcf097\",\"order_by\":8,\"name\":\"Taís Maria Nunes Carvalho\",\"email\":\"\",\"orcid\":\"https://orcid.org/0000-0001-8658-9781\",\"institution\":\"5 Helmholtz-Centre for Environmental Research, Department of Urban and Environmental Sociology, Leipzig, Germany 10 Center for Scalable Data Analytics and Artificial Intelligence (ScaDS.AI), Universität Leipzig, Leipzig, Germany\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Taís\",\"middleName\":\"Maria Nunes\",\"lastName\":\"Carvalho\",\"suffix\":\"\"}],\"badges\":[],\"createdAt\":\"2026-02-03 17:05:48\",\"currentVersionCode\":1,\"declarations\":{\"humanSubjects\":false,\"vertebrateSubjects\":false,\"conflictsOfInterestStatement\":false,\"humanSubjectEthicalGuidelines\":false,\"humanSubjectConsent\":false,\"humanSubjectClinicalTrial\":false,\"humanSubjectCaseReport\":false,\"vertebrateSubjectEthicalGuidelines\":false},\"doi\":\"10.21203/rs.3.rs-8778674/v1\",\"doiUrl\":\"https://doi.org/10.21203/rs.3.rs-8778674/v1\",\"draftVersion\":[],\"editorialEvents\":[],\"editorialNote\":\"\",\"failedWorkflow\":false,\"files\":[{\"id\":102296173,\"identity\":\"66aa3c6f-8f35-4eaa-bcad-7f7ba7c0ee55\",\"added_by\":\"auto\",\"created_at\":\"2026-02-10 10:17:52\",\"extension\":\"png\",\"order_by\":1,\"title\":\"Figure 1\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":190539,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eFlowchart of the impact extraction pipeline.\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"figure1.png\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-8778674/v1/6111e1e23f35b0dd27bc059f.png\"},{\"id\":102295722,\"identity\":\"864e0812-af8b-420c-8c4c-a1871658866f\",\"added_by\":\"auto\",\"created_at\":\"2026-02-10 10:14:20\",\"extension\":\"png\",\"order_by\":2,\"title\":\"Figure 2\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":2621730,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eUpper panel (a) shows the spatial distribution of the number of reported impacts per country in colour. Circular bar plots show the share of \\u003cem\\u003eimpactType \\u003c/em\\u003eper continent, with the inner numbers indicating the total number of impacts. The Lower panel (b) is the histogram of the number of impacts in each \\u003cem\\u003eimpactSubtype \\u003c/em\\u003ecategory. Colours correspond to the \\u003cem\\u003eimpactType.\\u003c/em\\u003e\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"figure2.png\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-8778674/v1/f4a357ba25904248e1727d2a.png\"},{\"id\":102062386,\"identity\":\"fcf22223-1220-4f3f-ab80-594bdc12015b\",\"added_by\":\"auto\",\"created_at\":\"2026-02-06 17:19:41\",\"extension\":\"png\",\"order_by\":3,\"title\":\"Figure 3\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":959227,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eShare of hazards (in %) to which impacts are associated, for each continent. Colours are ordered according to the type of hazard.\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"figure3.png\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-8778674/v1/e3527bca3eac00854cc1c793.png\"},{\"id\":102296169,\"identity\":\"1a8b6c8c-317c-4556-bd04-ed0f9ae64238\",\"added_by\":\"auto\",\"created_at\":\"2026-02-10 10:17:49\",\"extension\":\"png\",\"order_by\":4,\"title\":\"Figure 4\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":1031912,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eSimilarity score per matching variable. The horizontal bars show the similarity score (%) for each of the variables considered for the matching of extracted and manually-labelled impacts. The bottom bar plot shows the average for all matches of the weighted mean of the similarity scores for all matching variables. Colours show results for the quantitative vs qualitative impacts.\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"figure4.png\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-8778674/v1/1f0c70ccbcf5eac665f6a895.png\"},{\"id\":102062382,\"identity\":\"9a356df3-e5f2-4b09-9dde-e2f2bc5ae27d\",\"added_by\":\"auto\",\"created_at\":\"2026-02-06 17:19:41\",\"extension\":\"png\",\"order_by\":5,\"title\":\"Figure 5\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":1720556,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003ePrecision and recall per quantitative vs. qualitative extraction (a), and per impact subtypes (b). Colorbars represent the value of the metrics estimated from the samples directly. Red error bars represent the 95% confidence interval estimated via bootstrapping.\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"figure5.png\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-8778674/v1/5370757823174d2baf02dcde.png\"},{\"id\":102062385,\"identity\":\"0969b849-d561-497b-a999-35c931081de2\",\"added_by\":\"auto\",\"created_at\":\"2026-02-06 17:19:41\",\"extension\":\"png\",\"order_by\":6,\"title\":\"Figure 6\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":986604,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003ePercentage of events from IFRC Go (grey) and EM-DAT (blue) that could be matched to an LLM-extracted event for the impact subtypes “\\u003cem\\u003eAffected People\\u003c/em\\u003e”, “\\u003cem\\u003eInjured People\\u003c/em\\u003e”, “\\u003cem\\u003eDisplaced People\\u003c/em\\u003e”, “\\u003cem\\u003eHuman Deaths\\u003c/em\\u003e”, “\\u003cem\\u003eMissing People\\u003c/em\\u003e”, and “\\u003cem\\u003eHomeless People\\u003c/em\\u003e”. Among the matched events, darker shades indicate cases where the quantitative impact value exactly matches the LLM extraction, mid-tones indicate a match with a different reported value, and lighter shades show matched events for which no quantitative impact is recorded in IFRC Go or EM-DAT. The red horizontal line marks the 100% level corresponding to the total number of LLM events with reported quantitative impacts. Grey, blue, and red squares summarise, for each impact subtype, the proportion of all events (not only matched ones) with quantitative impacts reported in IFRC Go, EM-DAT, and the LLM database, respectively.\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"figure6.png\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-8778674/v1/b1243dc351b4524287b90eea.png\"},{\"id\":104834897,\"identity\":\"c7047f3b-99c4-4e4c-9046-d20deac023d4\",\"added_by\":\"auto\",\"created_at\":\"2026-03-17 17:35:35\",\"extension\":\"pdf\",\"order_by\":0,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"manuscript-pdf\",\"size\":8295311,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"manuscript.pdf\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-8778674/v1/99b16b66-b9f3-4d57-949d-3c2e96ce06e7.pdf\"},{\"id\":102295755,\"identity\":\"88cb7b10-accc-4787-ad11-d79d7cca7e29\",\"added_by\":\"auto\",\"created_at\":\"2026-02-10 10:14:42\",\"extension\":\"docx\",\"order_by\":1,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"supplement\",\"size\":285904,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eSupplementary Material\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"SupplementaryMaterial.docx\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-8778674/v1/b1cf407a39f5913688f2d3b6.docx\"},{\"id\":102062381,\"identity\":\"2ea8b98d-3f1d-411d-8565-039af6939f38\",\"added_by\":\"auto\",\"created_at\":\"2026-02-06 17:19:41\",\"extension\":\"docx\",\"order_by\":1,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"supplement\",\"size\":23148,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eImpact Labelling rules\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"ImpactLabellingrules.docx\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-8778674/v1/60310f9ea1dff374f31a45b7.docx\"},{\"id\":102247455,\"identity\":\"1a1c99b8-1342-4d3d-a09a-977c6c1809f4\",\"added_by\":\"auto\",\"created_at\":\"2026-02-09 18:28:28\",\"extension\":\"docx\",\"order_by\":2,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"supplement\",\"size\":285904,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"SupplementaryMaterial.docx\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-8778674/v1/e6c1f62b6457d7510b8ffc2f.docx\"}],\"financialInterests\":\"The authors declare no competing interests.\",\"formattedTitle\":\"\\u003cp\\u003eA database of disaster impacts in the Global South using Red Cross reports and Large Language Models\\u003c/p\\u003e\",\"fulltext\":[{\"header\":\"Background \\u0026 Summary\",\"content\":\"\\u003cp\\u003eNatural hazards exact a heavy toll on society, and have caused over 2\\u0026nbsp;million fatalities and 3.6 trillion Euros in damages worldwide in the past 50 years \\u003csup\\u003e\\u003cspan citationid=\\\"CR1\\\" class=\\\"CitationRef\\\"\\u003e1\\u003c/span\\u003e\\u003c/sup\\u003e. This trend is expected to continue in the future, as many disasters become more frequent and severe due to climate change and exposure shifts \\u003csup\\u003e\\u003cspan citationid=\\\"CR2\\\" class=\\\"CitationRef\\\"\\u003e2\\u003c/span\\u003e\\u003c/sup\\u003e.\\u003c/p\\u003e \\u003cp\\u003eWithin this context, understanding how natural hazards lead to impacts is crucial for improving risk management and reducing damage. Gaining such understanding demands a comprehensive view of where and when disasters occur, as well as the magnitude of their impacts. While several databases recording the impacts of natural hazards at the global \\u003csup\\u003e\\u003cspan citationid=\\\"CR1\\\" class=\\\"CitationRef\\\"\\u003e1\\u003c/span\\u003e\\u003c/sup\\u003e or regional scale\\u003csup\\u003e\\u003cspan citationid=\\\"CR3\\\" class=\\\"CitationRef\\\"\\u003e3\\u003c/span\\u003e\\u003c/sup\\u003e do exist, they suffer from geographical biases \\u003csup\\u003e\\u003cspan citationid=\\\"CR4\\\" class=\\\"CitationRef\\\"\\u003e4\\u003c/span\\u003e\\u003c/sup\\u003e, which can lead to under-reporting for the Global South \\u003csup\\u003e\\u003cspan citationid=\\\"CR5\\\" class=\\\"CitationRef\\\"\\u003e5\\u003c/span\\u003e\\u003c/sup\\u003e and result in overlooking certain risks in the most vulnerable regions. A second gap is that most impact databases focus on monetary losses, disregarding human, social, and environmental losses \\u003csup\\u003e\\u003cspan citationid=\\\"CR6\\\" class=\\\"CitationRef\\\"\\u003e6\\u003c/span\\u003e\\u003c/sup\\u003e. This focus undermines the importance of non-monetary, intangible losses and further amplifies the bias towards impacts in the Global North, where more financial and insured assets are concentrated. Finally, only a few databases report impacts at the subnational scale, while information at fine spatial resolution would improve impact modelling precision, especially because hazard information does exist at much higher resolution than the national scale.\\u003c/p\\u003e \\u003cp\\u003eWith the advent of big data and recent advances in computer science and artificial intelligence, a larger volume of data has become exploitable. Artificial intelligence and machine learning tools have emerged as powerful tools for improving the understanding and characterisation of natural hazards\\u003csup\\u003e\\u003cspan citationid=\\\"CR7\\\" class=\\\"CitationRef\\\"\\u003e7\\u003c/span\\u003e\\u003c/sup\\u003e and their impacts\\u003csup\\u003e\\u003cspan additionalcitationids=\\\"CR9\\\" citationid=\\\"CR8\\\" class=\\\"CitationRef\\\"\\u003e8\\u003c/span\\u003e\\u0026ndash;\\u003cspan citationid=\\\"CR10\\\" class=\\\"CitationRef\\\"\\u003e10\\u003c/span\\u003e\\u003c/sup\\u003e. In particular, natural language processing has demonstrated strong potential for extracting information from scattered and unstructured data sources, where manual extraction is too time-demanding. For instance, recent work has shown how natural language processing can be used to automatically identify drought impacts in large collections of news articles, extracting details about the affected sectors, locations, and reported consequences through rule-based and machine learning classifiers \\u003csup\\u003e\\u003cspan citationid=\\\"CR8\\\" class=\\\"CitationRef\\\"\\u003e8\\u003c/span\\u003e\\u003c/sup\\u003e. However, the performance of keyword search and rule-based approaches is limited, requiring highly detailed and context-specific dictionaries\\u003csup\\u003e\\u003cspan citationid=\\\"CR11\\\" class=\\\"CitationRef\\\"\\u003e11\\u003c/span\\u003e\\u003c/sup\\u003e. Moreover, traditional classification models require large volumes of annotated training data, which are time-intensive to produce, especially for rare impact types \\u003csup\\u003e\\u003cspan citationid=\\\"CR8\\\" class=\\\"CitationRef\\\"\\u003e8\\u003c/span\\u003e\\u003c/sup\\u003e. In such a context, large language models (LLM) represent a more flexible and extensive tool for data extraction. Latest studies demonstrate the effectiveness of LLM-based approaches in systematically retrieving and analysing impact information \\u003csup\\u003e\\u003cspan citationid=\\\"CR10\\\" class=\\\"CitationRef\\\"\\u003e10\\u003c/span\\u003e,\\u003cspan citationid=\\\"CR12\\\" class=\\\"CitationRef\\\"\\u003e12\\u003c/span\\u003e\\u003c/sup\\u003e.\\u003c/p\\u003e \\u003cp\\u003eAdditionally, the use of such approaches in the field of disaster risk management is still at an early stage, and challenges remain in developing methodologies and evaluating the applicability of these methods. In particular, validating the extracted information is crucial, given LLMs' propensity to hallucinate (i.e. produce inaccurate or unsupported information)\\u003csup\\u003e\\u003cspan citationid=\\\"CR13\\\" class=\\\"CitationRef\\\"\\u003e13\\u003c/span\\u003e\\u003c/sup\\u003e. This validation task remains time-consuming, especially when reference data is scarce and the data extracted is complex.\\u003c/p\\u003e \\u003cp\\u003eAgainst this background, the International Federation of the Red Cross' Societies (IFRC) reports impacts from natural hazards since 1919\\u003csup\\u003e14\\u003c/sup\\u003e in the form of written operation reports. Part of the data that can be found in these reports is available in machine-readable format through the IFRC GO platform\\u003csup\\u003e\\u003cspan citationid=\\\"CR14\\\" class=\\\"CitationRef\\\"\\u003e14\\u003c/span\\u003e\\u003c/sup\\u003e and the IFRC\\u0026rsquo;s Montandon initiative\\u003csup\\u003e\\u003cspan citationid=\\\"CR15\\\" class=\\\"CitationRef\\\"\\u003e15\\u003c/span\\u003e\\u003c/sup\\u003e. These databases only include figures on the number of fatalities, people affected, injured, displaced, and missing at the national level. However, the reports provide information on a wider range of impacts, such as health and infrastructure damage, and at a finer spatial resolution. Hence, these written reports represent a valuable and largely underexplored resource for creating impact datasets. Another advantage of this data is its explicit focus on disasters, which reduces the biases that often affect other impact datasets derived from news\\u003csup\\u003e\\u003cspan citationid=\\\"CR8\\\" class=\\\"CitationRef\\\"\\u003e8\\u003c/span\\u003e\\u003c/sup\\u003e, or Wikipedia articles\\u003csup\\u003e\\u003cspan citationid=\\\"CR10\\\" class=\\\"CitationRef\\\"\\u003e10\\u003c/span\\u003e\\u003c/sup\\u003e. Moreover, the IFRC provides systematic reporting on events affecting Global South countries, which are frequently underrepresented in global impact datasets. As such, the IFRC reports can provide a more inclusive and balanced perspective on disaster impacts.\\u003c/p\\u003e \\u003cp\\u003eThe combination of societal urgency, new artificial intelligence tools, and increasing volumes of text data provides a unique window of opportunity for studying the impacts of climate extremes. In this study, we leverage data from IFRC's textual reports and LLMs to produce a novel socio-economic impact database: ROUGE (Redcross Operations Unified Global Emergency database). Specifically, we focus on unconventional, non-monetary, qualitative, and quantitative impacts at the national and sub-national levels. Our database records the impacts of 716 disasters in 145 countries for 14 different hazards from 2016 to 2025. The database covers impacts for 20 subtypes of impacts from three main types of impacts, including \\u003cem\\u003eHuman\\u003c/em\\u003e, \\u003cem\\u003eInfrastructure and Service Access\\u003c/em\\u003e, and \\u003cem\\u003eEconomy and Culture\\u003c/em\\u003e. We expect this database to help the modelling and investigation of disasters and their risks, by providing new data on natural hazards and their impacts in regions which suffer from under-reporting in other conventional databases.\\u003c/p\\u003e \"},{\"header\":\"Methods\",\"content\":\"\\u003cp\\u003eThe extraction and processing of text data from IFRC reports followed four main steps (Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig1\\\" class=\\\"InternalRef\\\"\\u003e1\\u003c/span\\u003e). First, we extract the reports and convert them to text format. Second, we extract and clean the text to prepare it for automatic extraction (Subsection Preprocessing). Third, we prompt an LLM to identify and extract impact data from the cleaned text (Subsection Data extraction). Finally, during post-processing (Subsection Post-processing), we standardise the extracted outputs, remove inconsistent elements, and geolocate the reported locations (Subsection Geocoding). Additionally, for technical validation (Section Technical Validation), we compare the extracted information with manually labelled data and other established impact databases.\\u003c/p\\u003e \\u003cp\\u003e \\u003c/p\\u003e \\u003cp\\u003eData\\u003c/p\\u003e \\u003cp\\u003eWe use operational reports from the Disaster Response Emergency Fund (DREF) and the Emergency Appeal Fund of the IFRC as source data to extract the impacts of natural hazards. The reports are issued by National Societies of the IFRC in cases of humanitarian crises or disasters, covering small-scale crises (relatively narrow geographical scope and a relatively low number of affected people) to large-scale disasters (a large area of a country or countries is concerned and a very large number of people are affected)\\u003csup\\u003e\\u003cspan citationid=\\\"CR16\\\" class=\\\"CitationRef\\\"\\u003e16\\u003c/span\\u003e\\u003c/sup\\u003e. These reports describe the extent of the crisis or disaster, including the area and populations affected, as well as the drivers that lead to the emergency situation (including natural hazards, epidemics, and conflicts). The reports also include detailed descriptions of the impacts on the population, the infrastructure and the environment, as well as the planned response measures for disaster relief. The IFRC identifies events using \\u003cem\\u003eappealCode\\u003c/em\\u003e that uniquely identify each crisis. For each \\u003cem\\u003eappealCode\\u003c/em\\u003e, the IFRC can issue several reports that provide updates on the operation's status, including updates and a mandatory final closing report when the emergency appeal has been terminated\\u003csup\\u003e\\u003cspan citationid=\\\"CR17\\\" class=\\\"CitationRef\\\"\\u003e17\\u003c/span\\u003e\\u003c/sup\\u003e.\\u003c/p\\u003e \\u003cp\\u003eIn order to avoid extracting duplicate impacts, we select a single report per event from the multiple reports issued. Specifically, we keep the longest available report for each event to minimise information loss, as later updates or final reports may contain more summarised versions of the original details. Furthermore, we only consider reports from April 2016 to ensure structural consistency throughout the reports.\\u003c/p\\u003e \\u003cp\\u003eNext, we select reports for which the associated disaster type in the IFRC metadata corresponds to a natural hazard from the following list: Drought, Flood, Glacial lake outburst, Cyclone, Hurricane, Typhoon, Storm, Tornado, Heatwave, Coldwave, Mass movement, Earthquake, Volcano, Tidal Wave, Wildfire. As an additional check, we verify whether any of these hazard terms are mentioned in the report title.\\u003c/p\\u003e \\u003cp\\u003eWe scrape the reports from the IFRC website\\u003csup\\u003e\\u003cspan citationid=\\\"CR14\\\" class=\\\"CitationRef\\\"\\u003e14\\u003c/span\\u003e\\u003c/sup\\u003e and download them as PDFs. The text is extracted using the PyPDF2 Python package\\u003csup\\u003e\\u003cspan citationid=\\\"CR18\\\" class=\\\"CitationRef\\\"\\u003e18\\u003c/span\\u003e\\u003c/sup\\u003e. All images are discarded, and only text, including tables, is retained. This results in 717 reports, describing unique disaster events in 141 countries from April 2016 to July 2025.\\u003c/p\\u003e \\u003cp\\u003eFor technical validation, we used the EMDAT\\u003csup\\u003e\\u003cspan citationid=\\\"CR1\\\" class=\\\"CitationRef\\\"\\u003e1\\u003c/span\\u003e\\u003c/sup\\u003e and IFRC Go\\u003csup\\u003e14,19\\u003c/sup\\u003e impact datasets. EM-DAT is a widely used impact data provided by the Centre for Research on the Epidemiology of Disasters in open access\\u003csup\\u003e\\u003cspan citationid=\\\"CR1\\\" class=\\\"CitationRef\\\"\\u003e1\\u003c/span\\u003e\\u003c/sup\\u003e. It contains impact information from more than 27000 events worldwide from 1900 to the present day. The dataset gathers information from various sources such as UN agencies, non-governmental organisations, reinsurance companies, research institutes, and press agencies. It targets only major events that have resulted in at least 10 deaths and/or 100 affected people and/or were declared as an emergency by the state and/or were associated with an international assistance call. Precise information about the location impacted can be provided, but the impact is only given at the country level. To support systematic impact tracking, the IFRC has developed several initiatives aimed at collecting and structuring disaster impact information, such as the IFRC GO. This database compiles hazard events and associated impact data from IFRC DREF reports. It provides structured impact information for 1476 events recorded between 1997 and 2025. While this dataset offers valuable opportunities for impact comparison, its scope is limited to a narrow set of human and monetary impact categories.\\u003c/p\\u003e \\u003cp\\u003ePreprocessing\\u003c/p\\u003e \\u003cp\\u003eIn order to ensure machine-readability and facilitate the use of LLMs, we conduct several preprocessing steps on the raw text from the reports using the Spacy python library\\u003csup\\u003e\\u003cspan citationid=\\\"CR20\\\" class=\\\"CitationRef\\\"\\u003e20\\u003c/span\\u003e\\u003c/sup\\u003e. First, we identify the language of each report using a language detection pipeline\\u003csup\\u003e\\u003cspan citationid=\\\"CR21\\\" class=\\\"CitationRef\\\"\\u003e21\\u003c/span\\u003e\\u003c/sup\\u003e. Results show that only four reports are not written in English. These reports are discarded to avoid needing to adapt our extraction pipeline to other languages. Next, we remove unwanted text and characters such as URLs, ligatured characters, line breaks and multiple spaces. Based on the observed repeated structure of the reports (see supplementary material), we extract only the sections of the report that contain information about the hazard and its impacts. This limits the text we process, from an average of 244 sentences per report before selection down to an average of 78 sentences per report after selection, making extraction faster and more focused.\\u003c/p\\u003e \\u003cp\\u003eImpact data typology and database structure\\u003c/p\\u003e \\u003cp\\u003eWe define a socio-economic impact typology comprising 3 impact types and 20 subtypes as detailed in Table\\u0026nbsp;\\u003cspan refid=\\\"Tab1\\\" class=\\\"InternalRef\\\"\\u003e1\\u003c/span\\u003e. The typology was defined using a combination of inductive and deductive approaches, incorporating results from a data-driven topic analysis of the reports and partly adapting the impact types and subtypes of the EM-DAT and DESInventar frameworks. Besides the impacts, we also extract information such as the hazard type, location, and dates. The hazard typology has been adapted from EM-DAT hazard typology\\u003csup\\u003e\\u003cspan citationid=\\\"CR1\\\" class=\\\"CitationRef\\\"\\u003e1\\u003c/span\\u003e\\u003c/sup\\u003e and is described in Table\\u0026nbsp;\\u003cspan refid=\\\"Tab2\\\" class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003e.\\u003c/p\\u003e \\u003cp\\u003e \\u003cdiv class=\\\"gridtable\\\"\\u003e\\u003ctable float=\\\"Yes\\\" id=\\\"Tab1\\\" border=\\\"1\\\"\\u003e \\u003ccaption language=\\\"En\\\"\\u003e \\u003cdiv class=\\\"CaptionNumber\\\"\\u003eTable 1\\u003c/div\\u003e \\u003cdiv class=\\\"CaptionContent\\\"\\u003e \\u003cp\\u003eImpact types and description\\u003c/p\\u003e \\u003c/div\\u003e \\u003c/caption\\u003e \\u003ccolgroup cols=\\\"3\\\"\\u003e \\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c1\\\" colnum=\\\"1\\\"\\u003e\\u003c/div\\u003e \\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c2\\\" colnum=\\\"2\\\"\\u003e\\u003c/div\\u003e \\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c3\\\" colnum=\\\"3\\\"\\u003e\\u003c/div\\u003e \\u003cthead\\u003e \\u003ctr\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eImpact Maintype\\u003c/p\\u003e \\u003c/th\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eImpact Subtype\\u003c/p\\u003e \\u003c/th\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eDescription\\u003c/p\\u003e \\u003c/th\\u003e \\u003c/tr\\u003e \\u003c/thead\\u003e \\u003ctbody\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\" morerows=\\\"7\\\" rowspan=\\\"8\\\"\\u003e \\u003cp\\u003eHuman\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eAffected People\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eTotal number of individuals impacted by the hazard event.\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eInjured People\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eNumber of people injured, including those hospitalised or admitted.\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eDisplaced People\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eNumber of people forcefully displaced or evacuated before or following the event.\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eHomeless People\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eNumber of people losing housing following the event.\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eMissing People\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eNumber of people unaccounted for following the event.\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eHuman Deaths\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eNumber of fatalities caused by the hazard.\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eInfected and Ill People\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eNumber of people contaminated (cases) by an infectious disease.\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eHuman Health and Wellbeing\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eGeneric impacts on human health (physical, mental) or wellbeing not necessarily associated with the spread of an infectious disease.\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\" morerows=\\\"9\\\" rowspan=\\\"10\\\"\\u003e \\u003cp\\u003eInfrastructure and Service access\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eTransportation\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e- Roads and transportation infrastructure (e.g. bridges, highways..) impacted by a hazard.\\u003c/p\\u003e \\u003cp\\u003e- People losing the capacity to move safely and efficiently between locations.\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eWater, Sanitation, and Hygiene\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e- Number of water, sanitation, and hygiene infrastructure such as sewage networks, drainage systems, wastewater treatment plants, etc. impacted by a hazard.\\u003c/p\\u003e \\u003cp\\u003e- People losing safe, clean, and consistent supply of water for drinking, sanitation, hygiene services and other essential uses.\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eHealthcare\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e- Number of healthcare infrastructure such as hospitals, healthcare centers, pharmacies, clinics, etc. impacted by a hazard.\\u003c/p\\u003e \\u003cp\\u003e- People losing the ability to obtain needed medical services, including preventative care, emergency services \\u0026hellip;\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eIT and Communication\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e- Number of IT and communication infrastructure such as data centers, communication towers, and cables impacted by a hazard.\\u003c/p\\u003e \\u003cp\\u003e- People losing communication access\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eResidential Buildings\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e- Number of residential buildings (e.g. houses) impacted by a hazard.\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eInformal settlements\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e- Number of informal settlements such as refugee camps, slums, tents, etc. impacted by a hazard.\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eEducation\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e- Number of education infrastructures such as schools, universities, etc. impacted by a hazard.\\u003c/p\\u003e \\u003cp\\u003e- People losing the ability to attend educational institutions and receive instruction.\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003ePower and Energy Production\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e- Number of energy production infrastructures such as power plants, turbines, grids, pipelines, etc., impacted by a hazard.\\u003c/p\\u003e \\u003cp\\u003e- People losing access to electricity and services related to power production\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eAgriculture and Access to Food\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e- Number of agricultural infrastructures such as farms, warehouses, greenhouses, fisheries, etc. impacted by a hazard.\\u003c/p\\u003e \\u003cp\\u003e- Number of crops, agricultural production and forest impacted by a hazard\\u003c/p\\u003e \\u003cp\\u003e- Total number of animals, including terrestrial and aquatic species, impacted by the hazard. (e.g. perished animals or fishes)\\u003c/p\\u003e \\u003cp\\u003e- People losing access to a secure food supply\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eUndefined Infrastructure and Service Access\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e- Any identified impact on infrastructure where the type of infrastructure impacted is not clearly defined, e.g. critical infrastructure, public infrastructure, access to basic services\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\" morerows=\\\"1\\\" rowspan=\\\"2\\\"\\u003e \\u003cp\\u003eEconomy \\u0026amp; Culture\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eRecreation, Tourism, and Culture\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e- Tourist attractions and cultural sites impacted by a hazard.\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eEconomic and Livelihood\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003e- Any identified impact on the economy or living conditions resulting from a natural hazard.\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003c/tbody\\u003e \\u003c/colgroup\\u003e \\u003c/table\\u003e\\u003c/div\\u003e \\u003c/p\\u003e \\u003cp\\u003e \\u003cdiv class=\\\"gridtable\\\"\\u003e\\u003ctable float=\\\"Yes\\\" id=\\\"Tab2\\\" border=\\\"1\\\"\\u003e \\u003ccaption language=\\\"En\\\"\\u003e \\u003cdiv class=\\\"CaptionNumber\\\"\\u003eTable 2\\u003c/div\\u003e \\u003cdiv class=\\\"CaptionContent\\\"\\u003e \\u003cp\\u003eHazard types and description\\u003c/p\\u003e \\u003c/div\\u003e \\u003c/caption\\u003e \\u003ccolgroup cols=\\\"2\\\"\\u003e \\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c1\\\" colnum=\\\"1\\\"\\u003e\\u003c/div\\u003e \\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c2\\\" colnum=\\\"2\\\"\\u003e\\u003c/div\\u003e \\u003cthead\\u003e \\u003ctr\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eHazard\\u003c/p\\u003e \\u003c/th\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eDescription\\u003c/p\\u003e \\u003c/th\\u003e \\u003c/tr\\u003e \\u003c/thead\\u003e \\u003ctbody\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eDrought\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eProlonged lack of precipitation\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eWildfire\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eUncontrolled natural fires\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eEarthquake\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eSudden tectonic shifting. Includes as well tsunami\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eMass movement\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eAny type of downslope movement of earth materials. Includes Landslides and rockfalls\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eVolcanic activity\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eEruptions and related phenomena\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eFlood\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eRiver, coastal, flash and ice jam flooding\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eWave action\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eWind-generated surface waves (e.g. Rogue waves, seiche)\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eExtreme warm temperature\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eProlonged, abnormally high heat\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eExtreme cold temperature\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eProlonged, abnormally low cold\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eTropical storm\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eTropical cyclonic storms (e.g tropical cyclone, tropical storm, hurricanes, typhoons)\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eConvective storm\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eConvective-related storms (e.g. severe convective storms, derechos, hail, tornado)\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eOther storm\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eAny type of storm which does not correspond to a tropical cyclone or convective storm (e.g. extra-tropical storm, snow storm\\u0026hellip;)\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eEpidemic\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eSudden outbreak of disease\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eConflicts\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eDisagreements or disputes between different groups, organizations, or states\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003c/tbody\\u003e \\u003c/colgroup\\u003e \\u003c/table\\u003e\\u003c/div\\u003e \\u003c/p\\u003e \\u003cp\\u003eImpact data extraction\\u003c/p\\u003e \\u003cp\\u003eWe select the model meta-llama/llama-4-scout-17b-16e-instruct \\u003csup\\u003e\\u003cspan citationid=\\\"CR22\\\" class=\\\"CitationRef\\\"\\u003e22\\u003c/span\\u003e\\u003c/sup\\u003e, to identify and extract natural hazard impacts, as it is open-source and offers the best performance-to-price ratio compared to other models tested (in supplementary material). We use the GROQ cloud platform\\u003csup\\u003e\\u003cspan citationid=\\\"CR23\\\" class=\\\"CitationRef\\\"\\u003e23\\u003c/span\\u003e\\u003c/sup\\u003e to run the extraction.\\u003c/p\\u003e \\u003cp\\u003eThe extraction of the impact data is done using a sequence of prompts\\u003csup\\u003e\\u003cspan citationid=\\\"CR24\\\" class=\\\"CitationRef\\\"\\u003e24\\u003c/span\\u003e\\u003c/sup\\u003e. For each report, we prompt the LLM with a series of queries. This allows us to better guide the extraction by providing the LLM with more precise information at each extraction step, to have better control over the outputs and to limit the length of the output context windows. The extraction sequence is designed as follows:\\u003c/p\\u003e \\u003cp\\u003e \\u003cem\\u003eIdentify subtypes of impacts from the text, given a pre-defined list of impact subtypes and their definitions\\u003c/em\\u003e. The output is the list of identified subtypes.\\u003cdiv class=\\\"BlockQuote\\\"\\u003e\\u003cp\\u003e \\u003cem\\u003eFor each identified subtype of impacts, extract the values and units of those impacts\\u003c/em\\u003e. Extracting all impacts at once prevents the reuse of the same values for different identified subtypes. The output is a list of dictionaries.\\u003c/p\\u003e\\u003cp\\u003e \\u003cem\\u003eFor each identified impact (defined as a subtype, a value, and a unit), identify the affected locations\\u003c/em\\u003e. We allow a single impact to have multiple affected locations to account for aggregated reports of impacts. The output is a dictionary.\\u003c/p\\u003e\\u003cp\\u003e \\u003cem\\u003eFor each identified impact (defined as a subtype, a value, a unit, and a location), identify the starting and ending dates\\u003c/em\\u003e. The output is a dictionary.\\u003c/p\\u003e\\u003cp\\u003e \\u003cem\\u003eFor each identified impact (defined as a subtype, a value, a unit, a location, and a date), identify the hazards driving the impacts\\u003c/em\\u003e. We allow for multiple hazards to account for compound events.\\u003c/p\\u003e\\u003c/div\\u003e\\u003c/p\\u003e \\u003cp\\u003eAt each step, we use the pydantic library\\u003csup\\u003e\\u003cspan citationid=\\\"CR25\\\" class=\\\"CitationRef\\\"\\u003e25\\u003c/span\\u003e\\u003c/sup\\u003e to validate the output. This validation allows us to control that each desired field is (i) present in the extracted json, (ii) of the correct data type e.g. string, integer and (iii) contains only allowed values when value restrictions apply.\\u003c/p\\u003e \\u003cp\\u003eThe validation on the value constraints is only applied to the \\u003cem\\u003ehazards\\u003c/em\\u003e field, to ensure that only hazards matching the hazards lists are extracted. The validation of value constraints for the \\u003cem\\u003eimpactSubtype\\u003c/em\\u003e field was disabled to avoid the LLM relabelling unwanted data. If validation fails, we reprompt the LLM once with the validation error so it can correct the invalid output. Furthermore, we purposely extract unwanted \\u0026ldquo;control\\u0026rdquo; subtypes, such as the money raised by the appeal or the number of assisted people, as this prevents these figures from being wrongly extracted and classified as other similar impact subtypes, such as \\u003cem\\u003eEconomy and Livelihood\\u003c/em\\u003e or \\u003cem\\u003eAffected People\\u003c/em\\u003e. These control subtypes are later discarded from the database.\\u003c/p\\u003e \\u003cp\\u003eTo ensure traceability and for validation purposes, we also query the LLM to obtain the exact text excerpt it used for data extraction at each step. The extracted data is stored in a CSV file with the structure defined in the Data Descriptor section.\\u003c/p\\u003e \\u003cp\\u003ePost-processing\\u003c/p\\u003e \\u003cp\\u003eWe conduct the following post-processing steps. The \\u003cem\\u003eimpactValue\\u003c/em\\u003e is evaluated to ensure that it falls within the defined bounds, such that \\u003cspan class=\\\"InlineEquation\\\"\\u003e\\u003cspan class=\\\"mathinline\\\"\\u003e\\\\(\\\\:impactValueMin\\\\:\\\\le\\\\:\\\\:impactValue\\\\:\\\\le\\\\:\\\\:impactValueMax\\\\)\\u003c/span\\u003e\\u003c/span\\u003e. We reclassify and standardise the impact subtypes and hazard types using regular expressions. We format the \\u003cem\\u003eimpactUnit\\u003c/em\\u003e with the following four sub-steps; (i) we detect numbers from the unit field and move them to the \\u003cem\\u003eimpactValue\\u003c/em\\u003e field (e.g. \\u003cem\\u003e10 thousand people\\u003c/em\\u003e becomes \\u003cem\\u003e10\\u0026rsquo;000 people\\u003c/em\\u003e), (ii) we convert monetary units to EUR using the price parser\\u003csup\\u003e\\u003cspan citationid=\\\"CR26\\\" class=\\\"CitationRef\\\"\\u003e26\\u003c/span\\u003e\\u003c/sup\\u003e and currency converter\\u003csup\\u003e\\u003cspan citationid=\\\"CR27\\\" class=\\\"CitationRef\\\"\\u003e27\\u003c/span\\u003e\\u003c/sup\\u003e packages, (iii) we then standardise metric units (i.e. distance, mass, surface, volume) to SI units, and (iv) we standardise \\u0026ldquo;measured\\u0026rdquo; units using regular expressions and pre-defined conversion factors (e.g. \\u003cem\\u003e4 households\\u003c/em\\u003e becomes \\u003cem\\u003e12 people\\u003c/em\\u003e, \\u003cem\\u003e2 maternities\\u003c/em\\u003e becomes \\u003cem\\u003e2 healthcare structures\\u003c/em\\u003e).\\u003c/p\\u003e \\u003cp\\u003eFinally, we perform the following final cleaning steps. We (i) remove duplicated rows, (ii) filter out unknown impact subtypes and the control subtypes introduced during the extraction, and (iii) filter out unknown impact units that could not be standardised in the previous post-processing steps.\\u003c/p\\u003e \\u003cp\\u003eGeocoding\\u003c/p\\u003e \\u003cp\\u003eWe convert textual location descriptions into georeferenced polygons. This step is done not only to check the correctness of the location identified by the LLM but also to enrich it with geographic boundaries.\\u003c/p\\u003e \\u003cp\\u003eFor administrative boundaries, we rely on the \\u003cem\\u003egeoBoundaries\\u003c/em\\u003e dataset\\u003csup\\u003e\\u003cspan citationid=\\\"CR28\\\" class=\\\"CitationRef\\\"\\u003e28\\u003c/span\\u003e\\u003c/sup\\u003e, which compiles information from governmental and open-license sources. It provides Polygons or Multi-Polygons geometries for all countries, with administrative levels available up to admin 5. Since administrative levels beyond 2 (e.g. subdistricts, villages, neighbourhoods) are not consistently available across all countries, we restrict our analysis to level 2.\\u003c/p\\u003e \\u003cp\\u003eThe textual locations are first converted to geographic information using the OpenStreetMap (OSM) Nominatim tool\\u003csup\\u003e\\u003cspan citationid=\\\"CR29\\\" class=\\\"CitationRef\\\"\\u003e29\\u003c/span\\u003e\\u003c/sup\\u003e. This service returns an OSM address with a geometry in the form of a Point, Polygon, or MultiPolygon object. An OSM address may include multiple address details (also known as \\u003cem\\u003eaddress keys\\u003c/em\\u003e \\u003csup\\u003e\\u003cspan citationid=\\\"CR30\\\" class=\\\"CitationRef\\\"\\u003e30\\u003c/span\\u003e\\u003c/sup\\u003e), such as city, suburb, hamlet, street, and district.\\u003c/p\\u003e \\u003cp\\u003eIn order to convert textual location strings to polygons we first compute the Levenshtein similarity\\u003csup\\u003e\\u003cspan citationid=\\\"CR31\\\" class=\\\"CitationRef\\\"\\u003e31\\u003c/span\\u003e\\u003c/sup\\u003e between the input location string and all OSM address keys, with and without administrative descriptors and across rotated word orders. If a perfect match (1.0) is found, the search stops; otherwise, we keep the address key with the highest similarity above a 0.2 threshold, which identifies the most likely location name and administrative level, correcting spelling variations introduced by the LLM. The corrected name is then matched to the corresponding \\u003cem\\u003egeoBoundaries\\u003c/em\\u003e polygon at the identified administrative level.\\u003c/p\\u003e \\u003cp\\u003eIf matching with a polygon fails, we apply several fallback strategies. First, the location name may differ from the official \\u003cem\\u003egeoBoundaries\\u003c/em\\u003e designation, so we iterate over other OSM address keys at the same administrative level and attempt to match each to a polygon. If this fails, we translate the location name into the country\\u0026rsquo;s official languages and widely-spoken languages (French, Spanish, German) and retry the match. If no textual match is found, we use the geographic object returned by OSM and search for a corresponding geoBoundaries polygon by testing for intersections at the identified administrative level. If all attempts fail, we repeat the procedure at progressively coarser administrative levels until a match is found.\\u003c/p\\u003e \\u003cp\\u003eThis process yields a polygon for each textual location. Because reported impacts may reference multiple locations at different administrative levels, the polygons of each resolution identified are kept in the final geometry.\\u003c/p\\u003e \\u003cp\\u003eFor each post-processing step, we introduce quality flags to ensure traceability and data quality and allow users to flexibly select or filter out data according to their needs. The flags and their descriptions are detailed in the database structure table (Table\\u0026nbsp;\\u003cspan refid=\\\"Tab3\\\" class=\\\"InternalRef\\\"\\u003e3\\u003c/span\\u003e)\\u003c/p\\u003e \\u003cp\\u003e \\u003cdiv class=\\\"gridtable\\\"\\u003e\\u003ctable float=\\\"Yes\\\" id=\\\"Tab3\\\" border=\\\"1\\\"\\u003e \\u003ccaption language=\\\"En\\\"\\u003e \\u003cdiv class=\\\"CaptionNumber\\\"\\u003eTable 3\\u003c/div\\u003e \\u003cdiv class=\\\"CaptionContent\\\"\\u003e \\u003cp\\u003eStructure of the impact database\\u003c/p\\u003e \\u003c/div\\u003e \\u003c/caption\\u003e \\u003ccolgroup cols=\\\"3\\\"\\u003e \\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c1\\\" colnum=\\\"1\\\"\\u003e\\u003c/div\\u003e \\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c2\\\" colnum=\\\"2\\\"\\u003e\\u003c/div\\u003e \\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c3\\\" colnum=\\\"3\\\"\\u003e\\u003c/div\\u003e \\u003cthead\\u003e \\u003ctr\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eVariable\\u003c/p\\u003e \\u003c/th\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eFormat\\u003c/p\\u003e \\u003c/th\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eDescription\\u003c/p\\u003e \\u003c/th\\u003e \\u003c/tr\\u003e \\u003c/thead\\u003e \\u003ctbody\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eimpactType\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003estr\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eImpact type from predefined list\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eimpactSubtype\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003estr\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eImpact subtype from predefined list\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eimpactValue\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003efloat or None\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eNumerical value quantifying the impact\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eimpactValueMin\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003efloat or None\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eLower bound estimate of the impactValue\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eimpactValueMax\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003efloat or None\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eUpper bound estimate of the impactValue\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eimpactUnit\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003estr\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eUnit of impactValue\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003estartYear\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eint or None\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eStart year of impact\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003estartMonth\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eint or None\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eStart month of impact\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003estartDay\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eint or None\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eStart day of impact\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eendYear\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eint or None\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eEnd year of impact\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eendMonth\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eint or None\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eEnd month of impact\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eendDay\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eint or None\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eEnd day of impact\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003ehazards\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003elist\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eList of hazards associated with the described impact\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003elocation\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003elist\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eList of locations associated with the described impact\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003elocationLowestAdmin\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003estr\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eLowest administrative level identified with the geocoding\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003elocationPolygon\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003estr\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eName of the location polygon identified with the geocoding\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003egeometry\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003ePolygon or MutiPolygon\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eMultiPolygon of the locations identified with the geocoding\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003ecountry_iso3\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003elist\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eList of iso3 codes of the country identified with the geocoding\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003econtinent\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003elist\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eList of continents identified with the geocoding\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c3\\\" namest=\\\"c1\\\"\\u003e \\u003cp\\u003e\\u003cb\\u003eFlags\\u003c/b\\u003e :\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eimpactPrecision\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003estr\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eFlag describing whether the impactValue is exact or approximate\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003evalid_errors_impactValue\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eint\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eValidation error impactValue extraction. Count of validation errors during the extraction of the impactValue.\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003evalid_errors_loc\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eint\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eValidation error location extraction. Count of validation errors during the extraction of the location.\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003evalid_errors_dates\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eint\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eValidation error dates extraction. Count of validation errors during the extraction of the dates.\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003evalid_errors_haz\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eint\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eValidation error hazards extraction. Count of validation errors during the extraction of the hazards.\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eflag_value_not_in_text\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003ebool\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eImpactValue should be in the original text. True if impactValue cannot be found in original text\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eflag_pop_cntry\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003ebool\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eImpactValue with \\u0026ldquo;people\\u0026rdquo; unit should be smaller than country\\u0026rsquo;s population. True if impactValue with people unit is higher than country population\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eflag_impactSubtype_reclass\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003ebool\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eInferred impactSubtype must be in the allowed list. True if extracted impactSubtype was not in the allowed list and has been reclassified via regular expression\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eflag_hazards_reclass\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003ebool\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eInferred hazardType must be in the allowed list. True if extracted hazardType was not in allowed list and has been reclassified via regular expression\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eflag_unit_conversion\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003ebool\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eUnit conversion. True if a non-metric unit has been converted. These include families -\\u0026gt; 3*people, households -\\u0026gt; 3*people, communities -\\u0026gt; 100*people, villages -\\u0026gt; 1000*people\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eflag_non-SI_unit_standardization\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003ebool\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eStandardization of non-metric unit. True if a unit has been reclassified using regular expression, e.g. schools -\\u0026gt; education structures, bridges -\\u0026gt; roads\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eflag_SI_unit_standardization\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003ebool\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eStandardization of metric unit to SI. True if the extracted impactUnit has been standardized to a SI unit.\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eflag_remove_number_unit\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003ebool\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eTrue if a number has been identified from the extracted impactUnit and has been moved to the impactValue from the impactUnit (e.g. impactValue\\u0026thinsp;=\\u0026thinsp;10, impactUnit=\\u0026rdquo;milion people\\u0026rdquo; =\\u0026gt; impactValue\\u0026thinsp;=\\u0026thinsp;1000000, impactUnit=\\u0026rdquo;people\\u0026rdquo;)\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eflag_currency_conversion\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003ebool\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eTrue if a monetary unit has been identified but has been failed to standardized to the default currency\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eflag_unit_processing_error\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003ebool\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eTrue if there has been an issue in any of the following stages of the postprocessing of the impact unit: remove_number_unit, unit_conversion, SI_unit_standardization, non-SI_unit_standardization, currency_conversion.\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eflag_unit_nonstd\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003ebool\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eTrue if the extracted impactUnit is not part of the \\u0026ldquo;standard units\\u0026rdquo;\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eflag_reclass_subtype_from_unit\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003ebool\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eTrue if the extracted impactSubtype has been reclassified based on the impactUnit\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eflag_value_no_unit\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003ebool\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eTrue if an impactValue has no unit\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eflag_partial_unit\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003ebool\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eTrue if the impactUnit lacks the measured unit (e.g. \\u0026ldquo;hectares\\u0026rdquo; instead of \\u0026ldquo;hectares of crops\\u0026rdquo;)\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eflag_geocoding_osm\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eint\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eNumber of locations for which osm point/polygon has been used to retrieve the GAUL/NE polygon, with spatial matching\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eflag_geocoding_country\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eint\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eNumber of locations downgraded to the country to find a polygon\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c3\\\" namest=\\\"c1\\\"\\u003e \\u003cp\\u003e\\u003cb\\u003eMetadata\\u003c/b\\u003e :\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eappealCode\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003estr\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eAppeal code of IFRC report\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003ereportDate\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003estr\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eDate of the report in YYYY/MM/DD format\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003ereportLink\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003estr\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eUrl link of the pdf report\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003edisasterType_IFRC\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003estr\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eDisaster type identified by the IFRC\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003esourceExcerpts\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003estr\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e \\u003cp\\u003eText excerpts used for the information extraction\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003c/tbody\\u003e \\u003c/colgroup\\u003e \\u003c/table\\u003e\\u003c/div\\u003e \\u003c/p\\u003e \\u003cp\\u003eData Records\\u003c/p\\u003e \\u003cp\\u003eWe extract a database with the structure described in Table\\u0026nbsp;\\u003cspan refid=\\\"Tab3\\\" class=\\\"InternalRef\\\"\\u003e3\\u003c/span\\u003e. Each row of the database represents one impact record, specified by an impact subtype, a value, a unit, a location (geometry), a date, and a driver (associated hazards). Each impact record is also complemented with metadata (Table\\u0026nbsp;\\u003cspan refid=\\\"Tab3\\\" class=\\\"InternalRef\\\"\\u003e3\\u003c/span\\u003e), informing on the source (i.e. the appeal code uniquely identifying an IFRC appeal, the report date, the sentences that were used for the extraction of the impacts), the precision of the impact value, and several quality flags, which track issues and possible inconsistencies. The data will be made accessible via Zenodo under the \\u003cem\\u003eCreative Commons Attribution 4.0 International\\u003c/em\\u003e license. The database is available in two formats: as a global csv file without geometries and as Parquet files with geometries, available globally and split by continents. A notebook detailing how to open and access the data is available in the code archive, which can be downloaded anonymously by the reviewers at \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://figshare.com/s/99da5ff800fc56ae10c0\\u003c/span\\u003e\\u003cspan address=\\\"https://figshare.com/s/99da5ff800fc56ae10c0\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e. The data can be downloaded anonymously by reviewers at the same link.\\u003c/p\\u003e \\u003cp\\u003eData Overview\\u003c/p\\u003e \\u003cp\\u003eThe LLM-based impact dataset is created from IFRC reports and primarily captures impact in the Global South, as illustrated in Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig2\\\" class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003e. The regions most affected are Africa and Asia, with 5760 and 4525 impacts reports (Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig2\\\" class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003e), respectively, covering 707 and 796 different natural hazards (Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig3\\\" class=\\\"InternalRef\\\"\\u003e3\\u003c/span\\u003e). Figure\\u0026nbsp;\\u003cspan refid=\\\"Fig3\\\" class=\\\"InternalRef\\\"\\u003e3\\u003c/span\\u003e shows that floods, along with tropical and convective storms, are the hazards most frequently associated with the extracted impacts. Additionally, mass movements and epidemics also contribute substantially to the overall recorded impacts. 54% of the impacts extracted are geocoded at the country level (ADMIN0), 36%at the first-order administrative division (ADMIN1), and 10% at the second-order administrative division (ADMIN2).\\u003c/p\\u003e \\u003cp\\u003e \\u003c/p\\u003e \\u003cp\\u003e \\u003c/p\\u003e \\u003cp\\u003eTechnical Validation\\u003c/p\\u003e \\u003cp\\u003eComparison with manually labelled data (internal validation)\\u003c/p\\u003e \\u003cp\\u003eWe validate our database internally by comparing LLM-extracted impacts with manually extracted impacts from 17 randomly selected reports (listed in the supplementary). Manual extraction followed strict, pre-defined guidelines (see supplementary \\u003cem\\u003eLabelling Guidelines\\u003c/em\\u003e), and was performed by the authors with cross-reviewing. The labelled data acts as a control and represents the reference standard to which the LLM-extracted data is expected to adhere. From the 17 reports selected, we manually labelled 361 quantitative and 169 qualitative impacts, spanning over 20 impact subtypes and 13 hazards. In comparison, the automatic extraction using the LLM extracted 243 quantitative and 131 qualitative impacts, spanning over 20 impact subtypes and 9 hazards.\\u003c/p\\u003e \\u003cp\\u003eComparing the LLM-extracted with the manually-labelled impacts allows us to derive similarity scores for the different variables extracted, as well as the \\u003cem\\u003eprecision\\u003c/em\\u003e and \\u003cem\\u003erecall\\u003c/em\\u003e of the extraction. The similarity scores are defined as a function of the type of the examined variable and range from 0 for perfect dissimilarity to 1 for perfect similarity. Similarity of categorical variables is expressed using cosine similarity \\u003csup\\u003e\\u003cspan citationid=\\\"CR32\\\" class=\\\"CitationRef\\\"\\u003e32\\u003c/span\\u003e\\u003c/sup\\u003e, similarity of numerical variables (\\u003cem\\u003eimpactValue\\u003c/em\\u003e) with a normalised relative difference (Eq.\\u0026nbsp;\\u003cspan refid=\\\"Equ2\\\" class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003e in the supplementary \\u003cem\\u003eMatching algorithm for internal validation\\u003c/em\\u003e), and geometries with the Intersection-over-Union (Eq.\\u0026nbsp;\\u003cspan refid=\\\"Equ1\\\" class=\\\"InternalRef\\\"\\u003e1\\u003c/span\\u003e in the supplementary \\u003cem\\u003eMatching algorithm for internal validation\\u003c/em\\u003e). Precision and recall represent, respectively, the percentage of correctly extracted data over the total extracted and the percentage of correctly extracted data over the total manually-labelled.\\u003c/p\\u003e \\u003cp\\u003eIn order to compute the similarity scores for each extracted impact, we use an algorithm that matches each LLM-extracted row to the most similar row in the manually-labelled data. The matching procedure is described in the supplementary \\u003cem\\u003eMatching algorithm for internal validation\\u003c/em\\u003e. Figure\\u0026nbsp;\\u003cspan refid=\\\"Fig4\\\" class=\\\"InternalRef\\\"\\u003e4\\u003c/span\\u003e shows the average similarity per matching variables, separated by the quantitative and qualitative extraction. The \\u003cem\\u003ematch similarity\\u003c/em\\u003e is the weighted mean of the variables\\u0026rsquo; similarities for the matched data. We explain the lower match similarity in the qualitative data by the lower similarity of the impact subtypes, with some qualitative impact subtypes being more abstract, which may cause the LLM to perform worse at classifying the data into the correct subtypes. The similarity scores for the remaining variables are generally high, with scores above 80% for the impact unit, the iso3 code of the countries, the start year, and the impact values. The similarity of the dates is lower (around 50%). We attribute this lower performance to the complexity of clearly identifying event dates, which the human labellers encountered during manual labelling.\\u003c/p\\u003e \\u003cp\\u003e \\u003c/p\\u003e \\u003cp\\u003eNext, we compute the \\u003cem\\u003ePrecision\\u003c/em\\u003e and \\u003cem\\u003eRecall\\u003c/em\\u003e using the matched data to estimate the predictive performance of our extraction:\\u003cdiv id=\\\"Equ1\\\" class=\\\"Equation\\\"\\u003e\\u003cdiv format=\\\"TEX\\\" class=\\\"mathdisplay\\\" id=\\\"FileID_Equ1\\\" name=\\\"EquationSource\\\"\\u003e\\n$$\\\\:Precision=\\\\frac{TP}{TP+FP}$$\\u003c/div\\u003e\\u003cdiv class=\\\"EquationNumber\\\"\\u003e1\\u003c/div\\u003e\\u003c/div\\u003e\\u003cdiv id=\\\"Equ2\\\" class=\\\"Equation\\\"\\u003e\\u003cdiv format=\\\"TEX\\\" class=\\\"mathdisplay\\\" id=\\\"FileID_Equ2\\\" name=\\\"EquationSource\\\"\\u003e\\n$$\\\\:Recall=\\\\frac{TP}{TP+FN}$$\\u003c/div\\u003e\\u003cdiv class=\\\"EquationNumber\\\"\\u003e2\\u003c/div\\u003e\\u003c/div\\u003e\\u003c/p\\u003e \\u003cp\\u003ewhere \\u003cem\\u003eTP\\u003c/em\\u003e represents the True Positives, i.e. the number of LLM-extracted impacts for which we find a corresponding manually-extracted impacts, \\u003cem\\u003eFP\\u003c/em\\u003e the False Positives, i.e. the number of LLM-extracted impacts for which we do not find a corresponding manually-extracted impacts, and \\u003cem\\u003eFN\\u003c/em\\u003e the False Negatives, i.e. the number of manually-extracted impacts for which we did not find a corresponding LLM-extracted impacts. A perfect association should result in a \\u003cem\\u003eprecision\\u003c/em\\u003e and a \\u003cem\\u003erecall\\u003c/em\\u003e score of 1. Lower \\u003cem\\u003eprecision\\u003c/em\\u003e corresponds to a higher tendency of the LLM to hallucinate, while lower \\u003cem\\u003erecall\\u003c/em\\u003e indicates that the LLM is missing a larger portion of the true impacts.\\u003c/p\\u003e \\u003cp\\u003eTo quantify the sampling uncertainty, we approximate the 95% confidence intervals of the precision and recall metrics using a bootstrapping approach. For each sample from which the metrics are estimated, we sample with replacement 5000 random samples of true positives, false positives, and false negatives. We then compute the metrics from each random sample to obtain the bootstrapped distribution of the metrics. Finally, we use the 2.5th and 97.5th quantiles of the bootstrapped distribution as the bounds of the confidence interval.\\u003c/p\\u003e \\u003cp\\u003eFigure \\u003cspan refid=\\\"Fig5\\\" class=\\\"InternalRef\\\"\\u003e5\\u003c/span\\u003ea shows \\u003cem\\u003eprecision\\u003c/em\\u003e and \\u003cem\\u003erecall\\u003c/em\\u003e, and their estimated 95% confidence intervals for the total number of quantitative and qualitative impacts extracted. We find a precision of 0.71 and 0.62 for the quantitative and qualitative data, respectively. These results suggest that quantitative data can be extracted more precisely than qualitative data. This could indicate that the model is better at extracting numbers than at identifying more abstract entities, such as qualitative descriptions of impacts. Regarding recall, we obtain 0.37 for the quantitative data and 0.52 for the qualitative data. This suggests that only one third of the quantitative impacts that are described in the reports can be extracted with our approach. The extraction of qualitative impacts performs slightly better. We explain this by the higher number of quantitative (18 on average per report) than qualitative (10 on average per report) impacts present in the reports.\\u003c/p\\u003e \\u003cp\\u003eFigure \\u003cspan refid=\\\"Fig5\\\" class=\\\"InternalRef\\\"\\u003e5\\u003c/span\\u003eb shows the precision and recall metrics per impact subtype. The 95% confidence intervals indicate that sampling uncertainty is high due to the smaller sample sizes when grouping by impact subtype. For 7 of the 20 impact subtypes (\\u003cem\\u003eInfected and Ill People, Homeless People, Missing People, Power and Energy Production, Recreation, Tourism, and Culture, IT and Communication\\u003c/em\\u003e, and \\u003cem\\u003eUndefined Infrastructure and Service Access\\u003c/em\\u003e), the sampling uncertainty is too large to exclude the precision and recall to be 0. For the remaining subtypes, their performance does not differ significantly from the average precision and recall.\\u003c/p\\u003e \\u003cp\\u003e \\u003c/p\\u003e \\u003cp\\u003eIn addition, we ran robustness checks to assess whether our approach is sensitive to the choice of the LLM used and if it is deterministic (see supplementary \\u003cem\\u003eSensitivity analyses\\u003c/em\\u003e).\\u003c/p\\u003e \\u003cp\\u003eComparison with external data (external validation)\\u003c/p\\u003e \\u003cp\\u003eWe perform external validation by comparing quantitative impacts extracted by the LLM at the national level to EMDAT and IFRC-Go datasets. From our \\u003cem\\u003eimpactSubtypes\\u003c/em\\u003e (Table\\u0026nbsp;\\u003cspan refid=\\\"Tab1\\\" class=\\\"InternalRef\\\"\\u003e1\\u003c/span\\u003e) we can compare the \\u003cem\\u003eAffected People\\u003c/em\\u003e, \\u003cem\\u003eHuman Death\\u003c/em\\u003e and \\u003cem\\u003eHomeless People\\u003c/em\\u003e categories with EMDAT and the \\u003cem\\u003eAffected people\\u003c/em\\u003e, \\u003cem\\u003eInjured people\\u003c/em\\u003e, \\u003cem\\u003eDisplaced People\\u003c/em\\u003e, \\u003cem\\u003eHuman Deaths\\u003c/em\\u003e and \\u003cem\\u003eMissing People\\u003c/em\\u003e with IFRC-Go. As EM-DAT uses the IFRC to source their records of the number of people affected in 65% of cases \\u003csup\\u003e\\u003cspan citationid=\\\"CR1\\\" class=\\\"CitationRef\\\"\\u003e1\\u003c/span\\u003e\\u003c/sup\\u003e, we expect this category to match well across the three data sources.\\u003c/p\\u003e \\u003cp\\u003eFrom the LLM-extracted data, impacts are aggregated at the ADMIN0 level by retaining, for each record, the highest impact value reported at the coarsest administrative level. Square symbols in Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig6\\\" class=\\\"InternalRef\\\"\\u003e6\\u003c/span\\u003e illustrate the percentage of events for which a quantitative impact was identified across the previously mentioned common impact subtypes. The three databases show differing proportions of events with recorded quantitative impacts. Counts of \\u003cem\\u003eAffected People\\u003c/em\\u003e and \\u003cem\\u003eHuman Deaths\\u003c/em\\u003e are the most frequently reported, with EM-DAT consistently exhibiting the highest percentage of reports containing quantitative impact values. For the \\u003cem\\u003eInjured People\\u003c/em\\u003e, \\u003cem\\u003eDisplaced People\\u003c/em\\u003e, and \\u003cem\\u003eMissing People\\u003c/em\\u003e subtypes, IFRC Go reports quantitative impacts for approximately 40% of events on average. In contrast, the corresponding proportions derived from LLM-based extraction are lower, dropping to 24% for \\u003cem\\u003eInjured People\\u003c/em\\u003e and 18% for \\u003cem\\u003eMissing People.\\u003c/em\\u003e Overall, this indicates that the LLM captures quantitative impacts more effectively for widely reported subtypes such as \\u003cem\\u003eAffected People\\u003c/em\\u003e and \\u003cem\\u003eHuman Deaths\\u003c/em\\u003e, while it appears to struggle with less consistently documented subtypes, which are also under-documented in EM-DAT and IFRC-Go, including \\u003cem\\u003eInjured People\\u003c/em\\u003e, \\u003cem\\u003eDisplaced People\\u003c/em\\u003e, and \\u003cem\\u003eMissing People\\u003c/em\\u003e.\\u003c/p\\u003e \\u003cp\\u003eWe then match the reports containing quantitative information to the external datasets in order to compare the values of the reported impacts. Because EM-DAT and IFRC do not share a common unique event identifier, we match events using a stepwise procedure. We first link each IFRC report to EM-DAT events occurring in the same year and country. When multiple hazards are listed in an IFRC report, we select the hazard associated with the greatest number of reported impacts and retain only the EM-DAT events whose DisasterType corresponds to this dominant hazard. If multiple EM-DAT events still remain after this filtering, we select the event whose entry date is closest to the IFRC report date. This systematic procedure associates an event from EMDAT to each of the 594 events for which we extract quantitative information with the LLM. For IFRC-Go, matching can be done directly using the appeal code. We extract 566 events from the IFRC-Go data API over the period 2016\\u0026ndash;2025, which are associated with natural disasters. We successfully link 383 of these events with events which we extracted.\\u003c/p\\u003e \\u003cp\\u003eBar plots from Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig6\\\" class=\\\"InternalRef\\\"\\u003e6\\u003c/span\\u003e show the percentage of events from our database that we can match to the two external databases. The colours of the bars represent the percentage of matching against one of the two external databases. The brightness of the colours represent the percentage of matches for which we find (i) no impact value, (ii) an impact value differing from the LLM-extracted value by more than 10% and (iii) an impact value differing from the LLM-extracted value by less than 10% (see section \\u003cem\\u003eComparison with manually labelled data\\u003c/em\\u003e). We do not find a match within the IFRC-Go database for approximately 40% of the LLM-extracted events. This implies that our database retrieved impacts for appealCode which are not available within the IFRC-Go data API. Over the events matched, we do not find an impact record for the different impact subtypes considered in about 30% of the cases for both the IFRC-Go and EM-DAT. This suggests that, for the events considered, the LLM-based extraction is more robust in automatically detecting impacts for the considered impact subtypes. However, the degree of agreement between the impact values we extract and the external databases remains variable across the different impact subtypes. Overall, we observe a higher percentage of matched events with quantitative impact figures in EM-DAT compared to IFRC-Go, which is inline with the higher percentage of total events with impacts present in EM-DAT (square symbols). Interestingly, exact matches between LLM-extracted values and EM-DAT occur more often than between LLM-extracted values and IFRC-Go. On average across the different subtypes, we find exact value matches in less than 10% of the cases with IFRC-Go, and less than 30% of the cases with EM-DAT. Possible explanations include differences in the specific report versions used in the case of IFRC-Go, limitation of the matching procedure for EM-DAT or inaccuracies in the extraction process in either the LLM, IFRC-Go or EM-DAT databases. Nonetheless, these discrepancies across sources and impact categories further highlight the inherent challenges in estimating quantitative impact information from natural disasters.\\u003c/p\\u003e \\u003cp\\u003e \\u003c/p\\u003e\"},{\"header\":\"Declarations\",\"content\":\"\\u003cp\\u003e \\u003ch2\\u003eCompeting Interests\\u003c/h2\\u003e \\u003cp\\u003eThe authors declare no competing interests.\\u003c/p\\u003e \\u003c/p\\u003e\\u003ch2\\u003eFunding\\u003c/h2\\u003e \\u003cp\\u003eThis work was supported by the French ANRT, which funded LH\\u0026rsquo;s PhD. L.H. is employed by Generali France.\\u003c/p\\u003e\\u003ch2\\u003eAuthor Contributions\\u003c/h2\\u003e \\u003cp\\u003eLH: Conceptualisation, Methodology, Software, Validation, Formal analysis, Data Curation, Writing - Original Draft; Visualisation\\u003c/p\\u003e\\n\\u003cp\\u003eL.S.: Conceptualisation, Methodology, Software, Validation, Formal analysis, Data Curation, Writing - Original Draft; Visualisation\\u003c/p\\u003e\\n\\u003cp\\u003eMMdB: Conceptualisation, Writing - Review \\u0026 Editing; \\u003c/p\\u003e\\n\\u003cp\\u003eGCG : Conceptualisation, Data Curation, Writing - Review \\u0026 Editing\\u003c/p\\u003e\\n\\u003cp\\u003eAMR : Conceptualisation, Data Curation, Writing - Review \\u0026 Editing\\u003c/p\\u003e\\n\\u003cp\\u003eDNB : Supervision, Funding acquisition, Writing - Review \\u0026 Editing\\u003c/p\\u003e\\n\\u003cp\\u003eEM : Supervision, Writing - Review \\u0026 Editing\\n\\u003cp\\u003eJW : Conceptualisation\\u003c/p\\u003e\\n\\u003cp\\u003eTMNC: Conceptualisation, Data Curation, Supervision, Writing - Review \\u0026 Editing.\\n\\u003c/p\\u003e\\u003ch2\\u003eAcknowledgements\\u003c/h2\\u003e \\u003cp\\u003eThe authors thank Pascal Yiou (LSCE, UMR 8212CEA-CNRS-UVSQ) for reviewing the manuscript. The authors are grateful to the organisers and participants of the 2024 \\u003cem\\u003eComo Training School on Compound climate-related Events\\u003c/em\\u003e for providing the structure and motivation which started this work.\\u003c/p\\u003e\\u003ch2\\u003eData Availability\\u003c/h2\\u003e \\u003cp\\u003eThe data will be made accessible via Figshare in csv format without the geometries and in Parquet format with the geometries, under the \\u003cem\\u003eCreative Commons Attribution 4.0 International\\u003c/em\\u003e license. The data can be downloaded anonymously by reviewers at: \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://figshare.com/s/99da5ff800fc56ae10c0\\u003c/span\\u003e\\u003cspan address=\\\"https://figshare.com/s/99da5ff800fc56ae10c0\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/p\\u003e \\u003cp\\u003eCode Availability\\u003c/p\\u003e \\u003cp\\u003eThe code to produce the database, perform the technical validation and make the analysis plots is available anonymously at: \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://figshare.com/s/99da5ff800fc56ae10c0\\u003c/span\\u003e\\u003cspan address=\\\"https://figshare.com/s/99da5ff800fc56ae10c0\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e.\\u003c/p\\u003e\"},{\"header\":\"References\",\"content\":\"\\u003col\\u003e\\u003cli\\u003e\\u003cspan\\u003eDelforge D et al (2025) EM-DAT: the Emergency Events Database. Int J Disaster Risk Reduct 124:105509\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eIntergovernmental Panel On Climate Change (Ipcc) (2023) Climate Change 2022 \\u0026ndash; Impacts, Adaptation and Vulnerability: Working Group II Contribution to the Sixth Assessment Report of the Intergovernmental Panel on Climate Change. Cambridge University Press. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003e10.1017/9781009325844\\u003c/span\\u003e\\u003cspan address=\\\"10.1017/9781009325844\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eReduction UNO (2022) for D. R. \\u003cem\\u003eGlobal Assessment Report on Disaster Risk Reduction 2022: Our World at Risk: Transforming Governance for a Resilient Future\\u003c/em\\u003eUnited Nations Publications, Bloomfield\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eGall M, Borden KA, Cutter SL (2009) When Do Losses Count? Six Fallacies of Natural Hazards Loss Data. Bull Am Meteorol Soc 90:799\\u0026ndash;810\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eOsuteye E, Johnson C, Brown D (2017) The data gap: An analysis of data availability on disaster losses in sub-Saharan African cities. Int J Disaster Risk Reduct 26:24\\u0026ndash;33\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eDe Brito MM et al (2024) Uncovering the Dynamics of Multi-Sector Impacts of Hydrological Extremes: A Methods Overview. Earths Future 12:e2023EF003906\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eCamps-Valls G et al (2025) Artificial intelligence for modeling and understanding extreme weather and climate events. Nat Commun 16:1919\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eSodoge J, Kuhlicke C, De Brito MM (2023) Automatized spatio-temporal detection of drought impacts from newspaper articles using natural language processing and machine learning. Weather Clim Extrem 41:100574\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eDe Madruga M, Sodoge J, Kreibich H, Kuhlicke C (2024) Comprehensive Assessment of Flood Socioeconomic Impacts Through Text-Mining. \\u003cem\\u003eWater Resour. Res.\\u003c/em\\u003e 61, eWR037813 (2025)\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eLi N et al (2025) Wikimpacts 1.0: A new global climate impact database based on automated information extraction from Wikipedia. \\u003cem\\u003eEGUsphere\\u003c/em\\u003e 1\\u0026ndash;43 \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003e10.5194/egusphere-2025-4891\\u003c/span\\u003e\\u003cspan address=\\\"10.5194/egusphere-2025-4891\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eAdnan K, Akbar R (2019) An analytical study of information extraction from unstructured and multidimensional big data. J Big Data 6:91\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eNunes Carvalho TM, de Souza Filho F (2024) Madruga de Brito, M. Unveiling water allocation dynamics: a text analysis of 25 years of stakeholder meetings. Environ Res Lett 19:044066\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eHuang L et al (2025) A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions. ACM Trans Inf Syst 43(1\\u0026ndash;42):55\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eInternational Federation of Red Cross Crescent(IFRC) (2025) IFRC GO\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eMonty Extension Documentation \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://ifrcgo.org/monty-stac-extension/\\u003c/span\\u003e\\u003cspan address=\\\"https://ifrcgo.org/monty-stac-extension/\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eInternational Federation of Red Cross Crescent(IFRC). Emergency Response Framework (2025)\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eDREF Guidelines (2020) | IFRC. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://www.ifrc.org/document/dref-guidelines-2020\\u003c/span\\u003e\\u003cspan address=\\\"https://www.ifrc.org/document/dref-guidelines-2020\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2021)\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eFenniak M et al (2022) The PyPDF2 library\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eWiki GO \\u003cem\\u003eGO wiki\\u003c/em\\u003e https\\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003e://go-wiki.ifrc.org/en/home\\u003c/span\\u003e\\u003cspan address=\\\"http://://go-wiki.ifrc.org/en/home\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eMontani I et al (2023) explosion/spaCy: v3.7.2: Fixes for APIs and requirements. Zenodo \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.5281/zenodo.10009823\\u003c/span\\u003e\\u003cspan address=\\\"10.5281/zenodo.10009823\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eSpacy, FastLang \\u0026middot; spaCy Universe. Spacy FastLang \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://spacy.io/universe/project/spacy_fastlang\\u003c/span\\u003e\\u003cspan address=\\\"https://spacy.io/universe/project/spacy_fastlang\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003emeta-llama/Llama (2025) -4-Scout-17B-16E-Instruct \\u0026middot; Hugging Face. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://huggingface.co/meta-llama/Llama-4\\u003c/span\\u003e\\u003cspan address=\\\"https://huggingface.co/meta-llama/Llama-4\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e-Scout-17B-16E-Instruct\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eGroqCloud - Build Fast. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://console.groq.com\\u003c/span\\u003e\\u003cspan address=\\\"https://console.groq.com\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eTonmoy SMTI et al (2024) A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models. Preprint at \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.48550/arXiv.2401.01313\\u003c/span\\u003e\\u003cspan address=\\\"10.48550/arXiv.2401.01313\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eRelease v2 11.9 2025-09-13 \\u0026middot; pydantic/pydantic. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://github.com/pydantic/pydantic/releases/tag/v2.11.9\\u003c/span\\u003e\\u003cspan address=\\\"https://github.com/pydantic/pydantic/releases/tag/v2.11.9\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003escrapinghub/price- parser at 0.4.0. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://github.com/scrapinghub/price-parser/tree/0.4.0\\u003c/span\\u003e\\u003cspan address=\\\"https://github.com/scrapinghub/price-parser/tree/0.4.0\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003ealexprengere/currencyconverter at v0.18.9 \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://github.com/alexprengere/currencyconverter/tree/v0.18.9\\u003c/span\\u003e\\u003cspan address=\\\"https://github.com/alexprengere/currencyconverter/tree/v0.18.9\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eRunfola D et al (2020) geoBoundaries: A global database of political administrative boundaries. PLoS ONE 15:e0231866\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eOpenStreetMap OpenStreetMap \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://www.openstreetmap.org/\\u003c/span\\u003e\\u003cspan address=\\\"https://www.openstreetmap.org/\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003ePlace Output Formats - Nominatim 5.1.0 Manual. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://nominatim.org/release-docs/latest/api/Output/#addressdetails\\u003c/span\\u003e\\u003cspan address=\\\"https://nominatim.org/release-docs/latest/api/Output/#addressdetails\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eLevenshtein V (1966) Binary codes capable of correcting deletions, insertions, and reversals. in \\u003cem\\u003eSoviet physics-doklady\\u003c/em\\u003e vol. 10\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eSalton G, Wong A, Yang CS (1975) A vector space model for automatic indexing. Commun ACM 18:613\\u0026ndash;620\\u003c/span\\u003e\\u003c/li\\u003e\\u003c/ol\\u003e\"}],\"fulltextSource\":\"\",\"fullText\":\"\",\"funders\":[],\"hasAdminPriorityOnWorkflow\":false,\"hasManuscriptDocX\":true,\"hasOptedInToPreprint\":true,\"hasPassedJournalQc\":\"\",\"hasAnyPriority\":true,\"hideJournal\":true,\"highlight\":\"\",\"institution\":\"ETH Zurich\",\"isAcceptedByJournal\":false,\"isAuthorSuppliedPdf\":false,\"isDeskRejected\":\"\",\"isHiddenFromSearch\":false,\"isInQc\":false,\"isInWorkflow\":true,\"isPdf\":false,\"isPdfUpToDate\":true,\"isWithdrawnOrRetracted\":false,\"journal\":{\"display\":true,\"email\":\"info@researchsquare.com\",\"identity\":\"researchsquare\",\"isNatureJournal\":false,\"hasQc\":true,\"allowDirectSubmit\":true,\"externalIdentity\":\"\",\"sideBox\":\"\",\"snPcode\":\"\",\"submissionUrl\":\"/submission\",\"title\":\"Research Square\",\"twitterHandle\":\"researchsquare\",\"acdcEnabled\":true,\"dfaEnabled\":false,\"editorialSystem\":\"\",\"reportingPortfolio\":\"\",\"inReviewEnabled\":false,\"inReviewRevisionsEnabled\":true},\"keywords\":\"Natural hazards, socio-economic impacts, disaster impact database, IFRC, large language models, text mining, impact assessment\",\"lastPublishedDoi\":\"10.21203/rs.3.rs-8778674/v1\",\"lastPublishedDoiUrl\":\"https://doi.org/10.21203/rs.3.rs-8778674/v1\",\"license\":{\"name\":\"CC BY 4.0\",\"url\":\"https://creativecommons.org/licenses/by/4.0/\"},\"manuscriptAbstract\":\"\\u003cp\\u003eDamage from natural hazards exacts a heavy toll on society and is expected to increase under climate change. Yet, existing impact datasets remain limited and often biased toward Northern countries and monetary losses. To help address these gaps, we present ROUGE; a new socio-economic impact database obtained using textual operational reports from the International Federation of Red Cross and Red Crescent Societies (IFRC). These reports are systematically collected and provide broad coverage of regions that are commonly underrepresented in existing sources. Using large language models, we extract qualitative and quantitative information on a wide range of non-monetary impacts at national and sub-national scales. The resulting dataset documents socio-economic impacts of natural hazards on the population and the built environment with a spatial detail reaching the subregional level, capturing impacts that are rarely included in conventional databases. This resource is designed to support research and applications that require geographically explicit information on socio-economic impacts of disasters, enabling more precise and inclusive analyses of socio-economic consequences of natural hazards across the world.\\u003c/p\\u003e\\n\\u003cp\\u003e\\u003cbr\\u003e\\u003c/p\\u003e\\n\\u003cp\\u003eLuca G. Severino and Laura Hasbini are both first authors and contributed equally to the work.\\u003c/p\\u003e\",\"manuscriptTitle\":\"A database of disaster impacts in the Global South using Red Cross reports and Large Language Models\",\"msid\":\"\",\"msnumber\":\"\",\"nonDraftVersions\":[{\"code\":1,\"date\":\"2026-02-06 17:19:36\",\"doi\":\"10.21203/rs.3.rs-8778674/v1\",\"editorialEvents\":[{\"type\":\"communityComments\",\"content\":0}],\"status\":\"published\",\"journal\":{\"display\":true,\"email\":\"info@researchsquare.com\",\"identity\":\"researchsquare\",\"isNatureJournal\":false,\"hasQc\":true,\"allowDirectSubmit\":true,\"externalIdentity\":\"\",\"sideBox\":\"\",\"snPcode\":\"\",\"submissionUrl\":\"/submission\",\"title\":\"Research Square\",\"twitterHandle\":\"researchsquare\",\"acdcEnabled\":true,\"dfaEnabled\":false,\"editorialSystem\":\"\",\"reportingPortfolio\":\"\",\"inReviewEnabled\":false,\"inReviewRevisionsEnabled\":true}}],\"origin\":\"\",\"ownerIdentity\":\"c4adddfa-c0de-4ce9-b443-875da5791cc1\",\"owner\":[],\"postedDate\":\"February 6th, 2026\",\"published\":true,\"recentEditorialEvents\":[],\"rejectedJournal\":[],\"revision\":\"\",\"amendment\":\"\",\"status\":\"posted\",\"subjectAreas\":[{\"id\":62342713,\"name\":\"Climate Analysis and Modeling\"},{\"id\":62342714,\"name\":\"Artificial Intelligence and Machine Learning\"}],\"tags\":[],\"updatedAt\":\"2026-02-06T17:19:36+00:00\",\"versionOfRecord\":[],\"versionCreatedAt\":\"2026-02-06 17:19:36\",\"video\":\"\",\"vorDoi\":\"\",\"vorDoiUrl\":\"\",\"workflowStages\":[]},\"version\":\"v1\",\"identity\":\"rs-8778674\",\"journalConfig\":\"researchsquare\"},\"__N_SSP\":true},\"page\":\"/article/[identity]/[[...version]]\",\"query\":{\"redirect\":\"/article/rs-8778674\",\"identity\":\"rs-8778674\",\"version\":[\"v1\"]},\"buildId\":\"XKTyCvWXoU3ODBz1xrDgd\",\"isFallback\":false,\"isExperimentalCompile\":false,\"dynamicIds\":[84888],\"gssp\":true,\"scriptLoader\":[]}","source_license":"CC-BY-4.0","license_restricted":false}