A Named Entity Recognition and Topic Modeling-based Solution for Locating and Better Assessment of Natural Disasters in Social Media | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article A Named Entity Recognition and Topic Modeling-based Solution for Locating and Better Assessment of Natural Disasters in Social Media Ayaz Mehmood, Muhammad Asif Ayub, Muhammad Tayyab Zamir, Nasir Ahmad, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4658218/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Over the last decade, similar to other application domains, social media content has been widely explored for disaster informatics. However, due to the unstructured nature of the data, several challenges are associated with disaster analysis in social media content. For example, social media content is generally very noisy containing disaster-related words (e.g., floods, wildfires) used in a different context. Similarly, most of social media posts either do not contain geo-location information or do not represent the actual disaster location. To fully explore the potential of social media content in disaster informatics, access to relevant content and the correct geo-location information is very critical. In this paper, we propose a three-step solution to tackling these challenges. Firstly, the proposed solution aims at the classification of social media posts into relevant and irrelevant posts followed by the automatic extraction of location information from the posts' text itself through Named Entity Recognition (NER) analysis. Finally, to quickly analyze the topics covered in large volumes of social media posts, we perform topic modeling resulting in a list of top keywords, that highlight the issues discussed in the tweet. For the classification of the tweets, we proposed a merit-based fusion framework combining the capabilities of multiple Large Language Models (LLMs), obtaining the highest F1-score of 0.933 on a benchmark dataset. For the Location Extraction from Twitter Text (LETT), we evaluated four LLMs obtaining the highest F1-score of 0.960. For topic modeling, we used the BERTopic library to discover the hidden topic patterns in the relevant tweets. The experimental results of all the components of the proposed end-to-end solution are encouraging and hint at the effectiveness of social media content and NLP in disaster management. Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4658218","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":320997978,"identity":"98053945-6010-4fb8-91e8-512dee7baa31","order_by":0,"name":"Ayaz Mehmood","email":"","orcid":"","institution":"University of Engineering and Technology Peshawar","correspondingAuthor":false,"prefix":"","firstName":"Ayaz","middleName":"","lastName":"Mehmood","suffix":""},{"id":320997979,"identity":"c3d1f846-f665-4b6b-bab4-9b76d67050be","order_by":1,"name":"Muhammad Asif Ayub","email":"","orcid":"","institution":"University of Engineering and Technology Peshawar","correspondingAuthor":false,"prefix":"","firstName":"Muhammad","middleName":"Asif","lastName":"Ayub","suffix":""},{"id":320997980,"identity":"6943bab4-8ff3-44b9-88bc-45a76d971508","order_by":2,"name":"Muhammad Tayyab Zamir","email":"","orcid":"","institution":"Abbasyn University, Islamabad Campus, Pakistan.","correspondingAuthor":false,"prefix":"","firstName":"Muhammad","middleName":"Tayyab","lastName":"Zamir","suffix":""},{"id":320997981,"identity":"d335b86a-7fb5-407d-b561-0b5e93d1777a","order_by":3,"name":"Nasir Ahmad","email":"","orcid":"","institution":"University of Engineering and Technology Peshawar","correspondingAuthor":false,"prefix":"","firstName":"Nasir","middleName":"","lastName":"Ahmad","suffix":""},{"id":320997982,"identity":"5249998e-9ba5-4259-970f-06e5c5388896","order_by":4,"name":"Kashif Ahmad","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA3klEQVRIiWNgGAWjYBACCQYGNoYEGwYGAxDnA0iImSgtaRAtkjNAFFFaGKBapHnAFAEg2X782YMHCQzy5mKHH962zfmT2MDOfACvFmmeHHODhAQGw52z04ytc7cZJDYwsyXg1SLHkMMmkfiDgXHD7QQzaYgWHvyuk+N//kwCaIv9htvp36QtwVr4P+B3mESCGUhL4obbOWbSjBBb8OoAhusbkBaJZKCWYsvebcbGbcxs+B0mcT79meSPBBtboMM23vi5TU62n//wA/zWQHUimGzEqB8Fo2AUjIJRgB8AAPrcQExpQls1AAAAAElFTkSuQmCC","orcid":"","institution":"Munster Technological University","correspondingAuthor":true,"prefix":"","firstName":"Kashif","middleName":"","lastName":"Ahmad","suffix":""}],"badges":[],"createdAt":"2024-06-29 08:08:17","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4658218/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4658218/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":83153369,"identity":"854db103-24b7-47b0-bc56-91e0bd2a0db3","added_by":"auto","created_at":"2025-05-20 14:09:54","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":509737,"visible":true,"origin":"","legend":"","description":"","filename":"DisasterAnalysisNER.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4658218/v1_covered_5f698d95-f238-452a-8a74-5b176611acd0.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"A Named Entity Recognition and Topic Modeling-based Solution for Locating and Better Assessment of Natural Disasters in Social Media","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-4658218/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4658218/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Over the last decade, similar to other application domains, social media content has been widely explored for disaster informatics. However, due to the unstructured nature of the data, several challenges are associated with disaster analysis in social media content. For example, social media content is generally very noisy containing disaster-related words (e.g., floods, wildfires) used in a different context. Similarly, most of social media posts either do not contain geo-location information or do not represent the actual disaster location. To fully explore the potential of social media content in disaster informatics, access to relevant content and the correct geo-location information is very critical. In this paper, we propose a three-step solution to tackling these challenges. Firstly, the proposed solution aims at the classification of social media posts into relevant and irrelevant posts followed by the automatic extraction of location information from the posts' text itself through Named Entity Recognition (NER) analysis. Finally, to quickly analyze the topics covered in large volumes of social media posts, we perform topic modeling resulting in a list of top keywords, that highlight the issues discussed in the tweet. For the classification of the tweets, we proposed a merit-based fusion framework combining the capabilities of multiple Large Language Models (LLMs), obtaining the highest F1-score of 0.933 on a benchmark dataset. For the Location Extraction from Twitter Text (LETT), we evaluated four LLMs obtaining the highest F1-score of 0.960. For topic modeling, we used the BERTopic library to discover the hidden topic patterns in the relevant tweets. The experimental results of all the components of the proposed end-to-end solution are encouraging and hint at the effectiveness of social media content and NLP in disaster management.","manuscriptTitle":"A Named Entity Recognition and Topic Modeling-based Solution for Locating and Better Assessment of Natural Disasters in Social Media","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-07-22 18:52:57","doi":"10.21203/rs.3.rs-4658218/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"4d704509-231f-4808-a863-4a1684feac0d","owner":[],"postedDate":"July 22nd, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-05-20T14:01:47+00:00","versionOfRecord":[],"versionCreatedAt":"2024-07-22 18:52:57","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-4658218","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4658218","identity":"rs-4658218","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.