Exploration of the relationship between SDGs and CSR reports with text mining techniques for stock exchange companies over Taiwan

preprint OA: closed
Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-17

This study used text mining on Taiwanese CSR reports to analyze their similarity and relationship with SDGs, identifying key feature words and their alignment with specific SDG sub-items.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

Abstract

Abstract This research collects corporate social responsibility (CSR) reports from stock exchange companies in Taiwan and employs text mining technologies to analyze the relationship and document similarity between CSR reports from various industries and the Sustainable Development Goals (SDGs). The methods used include natural language processing (NLP), TF-IDF weighting, principal component analysis (PCA), and document similarity assessment. The study applies sub-items of selected SDG terms to screen feature words, build the TF-IDF matrix, analyze the CSR report contents using PCA, and utilize cosine similarity to compare the similarity between CSR reports and SDG sub-items. A total of 225 feature words were identified based on SDG sub-items, with the top 60 feature words (26.7%) accounting for 77.9% of the total TF-IDF weights, aligning with the Pareto principle. Analyzing 370 CSR reports from selected stock exchange companies (0050 ETF), unique and representative feature words and explained variations were identified. Each rotated principal component allowed the identification of corresponding SDG sub-items through specific feature words. The high diversity of feature words resulted in low and unique explained variance for each rotated principal component. Document similarity comparisons between CSR reports and SDG sub-items revealed confidence levels indicating the degree of alignment between CSR reports and SDG sub-items. For the natural language segmentation process and automatic document classification of CSR reports, the assistance of domain experts is recommended to ensure accurate and consistent segmentation and classification results.
Full text 11,940 characters · extracted from preprint-html · click to expand
Exploration of the relationship between SDGs and CSR reports with text mining techniques for stock exchange companies over Taiwan | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Exploration of the relationship between SDGs and CSR reports with text mining techniques for stock exchange companies over Taiwan Tai-Yi yu, Jeou-Shyan Horng, I-Cheng Chang, Tai-Kuei Yu, Chih-Hsing Liu, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4894913/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract This research collects corporate social responsibility (CSR) reports from stock exchange companies in Taiwan and employs text mining technologies to analyze the relationship and document similarity between CSR reports from various industries and the Sustainable Development Goals (SDGs). The methods used include natural language processing (NLP), TF-IDF weighting, principal component analysis (PCA), and document similarity assessment. The study applies sub-items of selected SDG terms to screen feature words, build the TF-IDF matrix, analyze the CSR report contents using PCA, and utilize cosine similarity to compare the similarity between CSR reports and SDG sub-items. A total of 225 feature words were identified based on SDG sub-items, with the top 60 feature words (26.7%) accounting for 77.9% of the total TF-IDF weights, aligning with the Pareto principle. Analyzing 370 CSR reports from selected stock exchange companies (0050 ETF), unique and representative feature words and explained variations were identified. Each rotated principal component allowed the identification of corresponding SDG sub-items through specific feature words. The high diversity of feature words resulted in low and unique explained variance for each rotated principal component. Document similarity comparisons between CSR reports and SDG sub-items revealed confidence levels indicating the degree of alignment between CSR reports and SDG sub-items. For the natural language segmentation process and automatic document classification of CSR reports, the assistance of domain experts is recommended to ensure accurate and consistent segmentation and classification results. Text mining principal component analysis TF-IDF weight document similarity Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4894913","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":354674684,"identity":"0b3689b9-1f47-4275-a285-b1d0a05dcd06","order_by":0,"name":"Tai-Yi yu","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAxklEQVRIiWNgGAWjYFACxsYHDAwSMlAeMxE62BibDYBaeBBa2AhqYWCTAFIkaDG439xWdaPGgodB7PAzCYYK68QG+R4D/FqOMbbdzjkGdJh0mpkEw5n0xAY2HmK0sIG0AEnGtsNALbwbCGopzvkH0/KPSC3MuW0wLQ1EaJE8ltgsndsnwcMmnWZskXAs3biNLf8DXi18h48//JzzrU6OXzr54Y0PNday/czHEvBqUTgAZYBjI4GBiJiUbyCkYhSMglEwCkYBAGKsOtzF8zOtAAAAAElFTkSuQmCC","orcid":"","institution":"Ming Chuan University","correspondingAuthor":true,"prefix":"","firstName":"Tai-Yi","middleName":"","lastName":"yu","suffix":""},{"id":354674685,"identity":"5b8710a0-ee0c-4a42-b8dc-80c6294d9b63","order_by":1,"name":"Jeou-Shyan Horng","email":"","orcid":"","institution":"Shih Chien University","correspondingAuthor":false,"prefix":"","firstName":"Jeou-Shyan","middleName":"","lastName":"Horng","suffix":""},{"id":354674686,"identity":"e8a74320-7b99-43a9-a209-bc6ea31b2659","order_by":2,"name":"I-Cheng Chang","email":"","orcid":"","institution":"National Ilan University","correspondingAuthor":false,"prefix":"","firstName":"I-Cheng","middleName":"","lastName":"Chang","suffix":""},{"id":354674688,"identity":"fe177312-a6fe-4cd8-ab79-2c25b55ddd11","order_by":3,"name":"Tai-Kuei Yu","email":"","orcid":"","institution":"National Quemoy University","correspondingAuthor":false,"prefix":"","firstName":"Tai-Kuei","middleName":"","lastName":"Yu","suffix":""},{"id":354674690,"identity":"2b3a15ec-9b36-48be-9db6-75e994953e9c","order_by":4,"name":"Chih-Hsing Liu","email":"","orcid":"","institution":"National Kaohsiung University of Science and Technology","correspondingAuthor":false,"prefix":"","firstName":"Chih-Hsing","middleName":"","lastName":"Liu","suffix":""},{"id":354674691,"identity":"8aa1bccd-d7a0-4116-a6c4-18e61aaac757","order_by":5,"name":"Sheng-Fang Chou","email":"","orcid":"","institution":"Ming Chuan University","correspondingAuthor":false,"prefix":"","firstName":"Sheng-Fang","middleName":"","lastName":"Chou","suffix":""}],"badges":[],"createdAt":"2024-08-11 11:15:05","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4894913/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4894913/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":75116321,"identity":"8d3577a9-a27c-4f07-9e71-0fa1cb647de9","added_by":"auto","created_at":"2025-01-30 16:16:47","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":792755,"visible":true,"origin":"","legend":"","description":"","filename":"paper20240810discomp.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4894913/v1_covered_0781d58e-b7cd-4d5c-85ea-911432748611.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Exploration of the relationship between SDGs and CSR reports with text mining techniques for stock exchange companies over Taiwan","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Text mining, principal component analysis, TF-IDF weight, document similarity","lastPublishedDoi":"10.21203/rs.3.rs-4894913/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4894913/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThis research collects corporate social responsibility (CSR) reports from stock exchange companies in Taiwan and employs text mining technologies to analyze the relationship and document similarity between CSR reports from various industries and the Sustainable Development Goals (SDGs). The methods used include natural language processing (NLP), TF-IDF weighting, principal component analysis (PCA), and document similarity assessment. The study applies sub-items of selected SDG terms to screen feature words, build the TF-IDF matrix, analyze the CSR report contents using PCA, and utilize cosine similarity to compare the similarity between CSR reports and SDG sub-items. A total of 225 feature words were identified based on SDG sub-items, with the top 60 feature words (26.7%) accounting for 77.9% of the total TF-IDF weights, aligning with the Pareto principle. Analyzing 370 CSR reports from selected stock exchange companies (0050 ETF), unique and representative feature words and explained variations were identified. Each rotated principal component allowed the identification of corresponding SDG sub-items through specific feature words. The high diversity of feature words resulted in low and unique explained variance for each rotated principal component. Document similarity comparisons between CSR reports and SDG sub-items revealed confidence levels indicating the degree of alignment between CSR reports and SDG sub-items. For the natural language segmentation process and automatic document classification of CSR reports, the assistance of domain experts is recommended to ensure accurate and consistent segmentation and classification results.\u003c/p\u003e","manuscriptTitle":"Exploration of the relationship between SDGs and CSR reports with text mining techniques for stock exchange companies over Taiwan","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-09-20 06:26:07","doi":"10.21203/rs.3.rs-4894913/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"42ac36b7-f59c-4726-8e59-5ad9046be32b","owner":[],"postedDate":"September 20th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-01-30T16:08:35+00:00","versionOfRecord":[],"versionCreatedAt":"2024-09-20 06:26:07","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-4894913","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4894913","identity":"rs-4894913","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00