Machine Learning has the Capability to Monitor the Advancement of Climate Technology Innovation Using Climate-Related Texts | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Machine Learning has the Capability to Monitor the Advancement of Climate Technology Innovation Using Climate-Related Texts Meenal B This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5942954/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Part one : In recent years, significant advancements in the field of natural language processing (NLP) have been driven by the emergence of large pre-trained language models (LM). Nevertheless, although pre-training on general language has demonstrated effectiveness for conventional language, challenges arise when applied to specialised language varieties. Specifically, texts pertaining to climate issues contain specialised vocabulary that cannot be accurately captured by mainstream language models. We posit that the deficiency in current Language Models restricts the potential use of cutting-edge Natural Language Processing in the extensive domain of text analysis related to climate change. In response to this issue, we introduce CLIMATE-BERT, a transformer-driven language model that receives additional pre-training on a dataset comprising more than 2 million paragraphs of climate-focused content sourced from a diverse range of outlets, including mainstream news, scholarly articles, and corporate climate reports. Our findings indicate that CLIMATE-BERT results in a 48% enhancement in performance on a masked language model task. This improvement subsequently translates to a reduction in error rates ranging from 3.57% to 35.71% across a range of downstream climate-related tasks such as text classification, sentiment analysis, and fact-checking. Part two : In order to adhere to international climate goals, expeditious advancement in climate technologies is imperative. Both governmental bodies and businesses are allocating significant resources to expedite this procedure. However, monitoring the progress and expansion of climate technologies poses difficulties. In this study, we demonstrate the utility of machine learning in monitoring and analysing advancements in climate technology. Our study leverages extensive language models and data from LinkedIn to establish a comprehensive network encompassing major public and private entities engaged in collaboration towards climate-tech innovation at a global scale. The network that emerges comprises 134,727 organisations spanning 189 countries. It encompasses a range of collaborative initiatives, encompassing research and development partnerships and equity investments, across 19 distinct climate technologies. The data we have collected demonstrate a significant focus on a select few climate technologies, as approximately 60% of emerging climate technology startups and innovation efforts are concentrated in just three areas: solar, electric vehicles, and hydrogen. However, this specific focus gives rise to apprehensions regarding the insufficient progress in various other essential climate technologies, namely heat pumps, biofuels, and carbon capture and storage. These technologies demand increased levels of innovation and entrepreneurial activities to ensure alignment with global climate objectives. Furthermore, our findings indicate that certain governmental entities perform a vital function within networks of innovation related to climate technology, especially in the context of commercialisation. Our study reveals that governmental organisations have the capability to foster pioneering innovation clusters, even in areas with minimal existing industry presence. Our findings indicate a crucial necessity for governmental entities to encourage increased technological diversification within regions, by fostering innovation in a broader spectrum of climate technologies. In subsequent periods, our approach in machine learning holds potential for ongoing monitoring of the evolution of climate-technology innovation, presenting an opportunity to evaluate the impact of governmental interventions. This framework is transferable to various other sectors, including pharmaceuticals, chemicals, and information and communication technologies. Machine learning large language models monitoring innovation networks climate technologies industrial policy Full Text Additional Declarations The authors declare no competing interests. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5942954","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":410041622,"identity":"47aca368-041b-4bac-bdfe-d6ec4916b6f7","order_by":0,"name":"Meenal B","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABCElEQVRIiWNgGAWjYNCCAoYEBoY0hgMMBjYMDBJQQTa8WgxgWgrSSNTCwPDhMEILLqDb3nzswQ8Dhjz+9rREoMPOJ/bPbj74gKHGhoFPugGrFrMzx9INewwYiiXOPDsA1HI7ccadY8kGDMfSGNhkDmDXciPHTILHgCGx4UZ6A1hLA0iEseEwA5tEAnYt999/k/wD1DIfouUckEFIyw0eNmmQLRtupIEcdgDIIKTlTJqZtAzQL4ZnniUAtSQbb7yRlmyQcCyNB6eW44efSb6pYMiTO55m/IHhj53svBvJBx98qLGRk5+BXQsU/AeTzH8YGBwbQCygYh586lGAPdEqR8EoGAWjYMQAAITQYOruV+viAAAAAElFTkSuQmCC","orcid":"","institution":"","correspondingAuthor":true,"prefix":"","firstName":"Meenal","middleName":"","lastName":"B","suffix":""}],"badges":[],"createdAt":"2025-02-01 17:33:00","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-5942954/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5942954/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":75380056,"identity":"fe50706b-6082-4e11-a272-7536b1675770","added_by":"auto","created_at":"2025-02-04 02:31:13","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":5288270,"visible":true,"origin":"","legend":"","description":"","filename":"jklm.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5942954/v1_covered_4a926268-dbad-473b-a114-14bc2a5a843b.pdf"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003e\u003cstrong\u003eMachine Learning has the Capability to Monitor the Advancement of Climate Technology Innovation Using Climate-Related Texts\u003c/strong\u003e\u003c/p\u003e","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Machine learning, large language models, monitoring, innovation networks, climate technologies, industrial policy","lastPublishedDoi":"10.21203/rs.3.rs-5942954/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5942954/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003ePart one : \u003c/strong\u003eIn recent years, significant advancements in the field of natural language processing (NLP) have been driven by the emergence of large pre-trained language models (LM). Nevertheless, although pre-training on general language has demonstrated effectiveness for conventional language, challenges arise when applied to specialised language varieties. Specifically, texts pertaining to climate issues contain specialised vocabulary that cannot be accurately captured by mainstream language models. We posit that the deficiency in current Language Models restricts the potential use of cutting-edge Natural Language Processing in the extensive domain of text analysis related to climate change. In response to this issue, we introduce CLIMATE-BERT, a transformer-driven language model that receives additional pre-training on a dataset comprising more than 2 million paragraphs of climate-focused content sourced from a diverse range of outlets, including mainstream news, scholarly articles, and corporate climate reports. Our findings indicate that CLIMATE-BERT results in a 48% enhancement in performance on a masked language model task. This improvement subsequently translates to a reduction in error rates ranging from 3.57% to 35.71% across a range of downstream climate-related tasks such as text classification, sentiment analysis, and fact-checking.\u003cstrong\u003e\u003cbr\u003e\nPart two : \u003c/strong\u003eIn order to adhere to international climate goals, expeditious advancement in climate technologies is imperative. Both governmental bodies and businesses are allocating significant resources to expedite this procedure. However, monitoring the progress and expansion of climate technologies poses difficulties. In this study, we demonstrate the utility of machine learning in monitoring and analysing advancements in climate technology. Our study leverages extensive language models and data from LinkedIn to establish a comprehensive network encompassing major public and private entities engaged in collaboration towards climate-tech innovation at a global scale. The network that emerges comprises 134,727 organisations spanning 189 countries. It encompasses a range of collaborative initiatives, encompassing research and development partnerships and equity investments, across 19 distinct climate technologies. The data we have collected demonstrate a significant focus on a select few climate technologies, as approximately 60% of emerging climate technology startups and innovation efforts are concentrated in just three areas: solar, electric vehicles, and hydrogen. However, this specific focus gives rise to apprehensions regarding the insufficient progress in various other essential climate technologies, namely heat pumps, biofuels, and carbon capture and storage. These technologies demand increased levels of innovation and entrepreneurial activities to ensure alignment with global climate objectives. Furthermore, our findings indicate that certain governmental entities perform a vital function within networks of innovation related to climate technology, especially in the context of commercialisation. Our study reveals that governmental organisations have the capability to foster pioneering innovation clusters, even in areas with minimal existing industry presence. Our findings indicate a crucial necessity for governmental entities to encourage increased technological diversification within regions, by fostering innovation in a broader spectrum of climate technologies. In subsequent periods, our approach in machine learning holds potential for ongoing monitoring of the evolution of climate-technology innovation, presenting an opportunity to evaluate the impact of governmental interventions. This framework is transferable to various other sectors, including pharmaceuticals, chemicals, and information and communication technologies.\u003c/p\u003e","manuscriptTitle":"Machine Learning has the Capability to Monitor the Advancement of Climate Technology Innovation Using Climate-Related Texts","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-02-04 02:23:02","doi":"10.21203/rs.3.rs-5942954/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"26ae5978-7f65-40bf-a08a-db929be08459","owner":[],"postedDate":"February 4th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-02-04T02:23:02+00:00","versionOfRecord":[],"versionCreatedAt":"2025-02-04 02:23:02","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-5942954","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5942954","identity":"rs-5942954","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.