{"paper_id":"19ba6a5f-4c98-4d5b-a915-ea052b98ae72","body_text":"Few shot learning with fine-tuned language model for suicidal text detection | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Few shot learning with fine-tuned language model for suicidal text detection Sandeep Varma, Shivam Shivam, Biswarup Ray, Ankita Banerjee This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-2392230/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract The most effective way to prevent potential suicide attempts is early identification followed by prompt treatment. Nowadays, online communication channels especially social media is becoming a way of expressing suicidal tendencies. However, the availability of properly tagged textual data is a major issue for researchers to perform identification or classification of such tendencies through texts. In light of the abovementioned facts, in the present work, we have proposed an approach to determine early detection with the goal of early diagnosis of suicidal behaviour through text posted on social media via supervised learning using few shot learning process. Therefore, to detect suicidal behaviour, we extract embeddings by finetuning a large pre-trained language model. The contextual embeddings have then been used for binary classification of the text for suicidal behaviour, in which a comparison of various classifiers has been performed. This includes traditional supervised classifiers and neural network models. A comparison of the model’s performance with and without an outlier detection and removal step has also been performed to highlight the importance of the outlier detection step in the pipeline. The feasibility and practicality of the approach have been demonstrated by generating results for user-generated content scraped from Reddit platform posts made on subreddits ‘SuicideWatch’ and ‘depression’. The results generated by the model are also seen to outperform various state-of-the-art models. Suicidal Behavior Identification Language Models Outlier Detection Classification and Word Embedding Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {\"props\":{\"pageProps\":{\"initialData\":{\"identity\":\"rs-2392230\",\"acceptedTermsAndConditions\":true,\"allowDirectSubmit\":true,\"archivedVersions\":[],\"articleType\":\"Research Article\",\"associatedPublications\":[],\"authors\":[{\"id\":161767407,\"identity\":\"884d27f4-abe7-46fc-854f-145c300ded47\",\"order_by\":0,\"name\":\"Sandeep Varma\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"ZS Associates\",\"correspondingAuthor\":false,\"submittingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Sandeep\",\"middleName\":\"\",\"lastName\":\"Varma\",\"suffix\":\"\"},{\"id\":161767408,\"identity\":\"42e84f2b-a931-4ec3-be38-f9c610ca7d70\",\"order_by\":1,\"name\":\"Shivam Shivam\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"ZS Associates\",\"correspondingAuthor\":false,\"submittingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Shivam\",\"middleName\":\"\",\"lastName\":\"Shivam\",\"suffix\":\"\"},{\"id\":161767409,\"identity\":\"5f9a4524-a2ae-423b-b037-c01d95eb99a8\",\"order_by\":2,\"name\":\"Biswarup Ray\",\"email\":\"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA+ElEQVRIiWNgGAWjYHACNoYECIPxwQcQl50ELcyGM0BcZmK0wBjSPGCdBNSbt58xe/Dgz2F58xm5x6Rtfm2T52NmYPzwMQe3FpkzOeYGiW2HDefcyEu2zu27bdjGzMAsOXMbbi0SDDlmEokNtxlnSOQY3s7tuc0I1MLGzItPC/8bM4mEP7ftgVoMpC17btsT1iIBtCWB7XYiUIuRNMOP24lEaHlWDvTL/+QZPO+SDXsbbie3MTM24/cLf/K2hz/+pNnOYM89+ODHn9u289ubD374iEcLAwOHAZQBjBTGNhCDsQGfeiBgf4DQwvCHgOJRMApGwSgYkQAAAtFPFh5b9b8AAAAASUVORK5CYII=\",\"orcid\":\"\",\"institution\":\"ZS Associates\",\"correspondingAuthor\":true,\"submittingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Biswarup\",\"middleName\":\"\",\"lastName\":\"Ray\",\"suffix\":\"\"},{\"id\":161767410,\"identity\":\"cc3b3f00-9606-44b7-ba24-398ec9db3238\",\"order_by\":3,\"name\":\"Ankita Banerjee\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"ZS Associates\",\"correspondingAuthor\":false,\"submittingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Ankita\",\"middleName\":\"\",\"lastName\":\"Banerjee\",\"suffix\":\"\"}],\"badges\":[],\"createdAt\":\"2022-12-19 06:44:18\",\"currentVersionCode\":1,\"declarations\":\"\",\"doi\":\"10.21203/rs.3.rs-2392230/v1\",\"doiUrl\":\"https://doi.org/10.21203/rs.3.rs-2392230/v1\",\"draftVersion\":[],\"editorialEvents\":[],\"editorialNote\":\"\",\"failedWorkflow\":false,\"files\":[{\"id\":32470210,\"identity\":\"e713036f-5541-4dbb-bfe2-c3bbc656a93b\",\"added_by\":\"auto\",\"created_at\":\"2023-02-04 11:14:39\",\"extension\":\"pdf\",\"order_by\":0,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"manuscript-pdf\",\"size\":404919,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"sxqzjfvvccnmntnqvynrqnngzbmwjzny.pdf\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-2392230/v1_covered.pdf\"},{\"id\":30695937,\"identity\":\"13c0d1fc-b78b-48cc-bde3-b3ccc0dcd9a4\",\"added_by\":\"auto\",\"created_at\":\"2022-12-23 03:38:17\",\"extension\":\"pdf\",\"order_by\":0,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"manuscript-pdf\",\"size\":420483,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"sxqzjfvvccnmntnqvynrqnngzbmwjzny.pdf\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-2392230/v1/39f5355cf3eb87ed474e6cb9.pdf\"}],\"financialInterests\":\"No competing interests reported.\",\"formattedTitle\":\"Few shot learning with fine-tuned language model for suicidal text detection\",\"fulltext\":[],\"fulltextSource\":\"\",\"fullText\":\"\",\"funders\":[],\"hasAdminPriorityOnWorkflow\":false,\"hasManuscriptDocX\":false,\"hasOptedInToPreprint\":true,\"hasPassedJournalQc\":\"\",\"hasAnyPriority\":false,\"hideJournal\":true,\"highlight\":\"\",\"institution\":\"\",\"isAcceptedByJournal\":false,\"isAuthorSuppliedPdf\":true,\"isDeskRejected\":\"\",\"isHiddenFromSearch\":false,\"isInQc\":false,\"isInWorkflow\":false,\"isPdf\":true,\"isPdfUpToDate\":true,\"isWithdrawnOrRetracted\":false,\"journal\":{\"display\":true,\"email\":\"info@researchsquare.com\",\"identity\":\"researchsquare\",\"isNatureJournal\":false,\"hasQc\":true,\"allowDirectSubmit\":true,\"externalIdentity\":\"\",\"sideBox\":\"\",\"snPcode\":\"\",\"submissionUrl\":\"/submission\",\"title\":\"Research Square\",\"twitterHandle\":\"researchsquare\",\"acdcEnabled\":true,\"dfaEnabled\":false,\"editorialSystem\":\"\",\"reportingPortfolio\":\"\",\"inReviewEnabled\":false,\"inReviewRevisionsEnabled\":true},\"keywords\":\"Suicidal Behavior Identification, Language Models, Outlier Detection, Classification and Word Embedding\",\"lastPublishedDoi\":\"10.21203/rs.3.rs-2392230/v1\",\"lastPublishedDoiUrl\":\"https://doi.org/10.21203/rs.3.rs-2392230/v1\",\"license\":{\"name\":\"CC BY 4.0\",\"url\":\"https://creativecommons.org/licenses/by/4.0/\"},\"manuscriptAbstract\":\"The most effective way to prevent potential suicide attempts is early identification followed by prompt treatment. Nowadays, online communication channels especially social media is becoming a way of expressing suicidal tendencies. However, the availability of properly tagged textual data is a major issue for researchers to perform identification or classification of such tendencies through texts. In light of the abovementioned facts, in the present work, we have proposed an approach to determine early detection with the goal of early diagnosis of suicidal behaviour through text posted on social media via supervised learning using few shot learning process. Therefore, to detect suicidal behaviour, we extract embeddings by finetuning a large pre-trained language model. The contextual embeddings have then been used for binary classification of the text for suicidal behaviour, in which a comparison of various classifiers has been performed. This includes traditional supervised classifiers and neural network models. A comparison of the model’s performance with and without an outlier detection and removal step has also been performed to highlight the importance of the outlier detection step in the pipeline. The feasibility and practicality of the approach have been demonstrated by generating results for user-generated content scraped from Reddit platform posts made on subreddits ‘SuicideWatch’ and ‘depression’. The results generated by the model are also seen to outperform various state-of-the-art models.\",\"manuscriptTitle\":\"Few shot learning with fine-tuned language model for suicidal text detection\",\"msid\":\"\",\"msnumber\":\"\",\"nonDraftVersions\":[{\"code\":1,\"date\":\"2022-12-23 03:38:12\",\"doi\":\"10.21203/rs.3.rs-2392230/v1\",\"editorialEvents\":[{\"type\":\"communityComments\",\"content\":0}],\"status\":\"published\",\"journal\":{\"display\":true,\"email\":\"info@researchsquare.com\",\"identity\":\"researchsquare\",\"isNatureJournal\":false,\"hasQc\":true,\"allowDirectSubmit\":true,\"externalIdentity\":\"\",\"sideBox\":\"\",\"snPcode\":\"\",\"submissionUrl\":\"/submission\",\"title\":\"Research Square\",\"twitterHandle\":\"researchsquare\",\"acdcEnabled\":true,\"dfaEnabled\":false,\"editorialSystem\":\"\",\"reportingPortfolio\":\"\",\"inReviewEnabled\":false,\"inReviewRevisionsEnabled\":true}}],\"origin\":\"\",\"ownerIdentity\":\"3880678a-98f5-4528-9c49-45dc8c9d39ac\",\"owner\":[],\"postedDate\":\"December 23rd, 2022\",\"published\":true,\"recentEditorialEvents\":[],\"rejectedJournal\":[],\"revision\":\"\",\"amendment\":\"\",\"status\":\"posted\",\"subjectAreas\":[],\"tags\":[],\"updatedAt\":\"2023-02-04T11:14:24+00:00\",\"versionOfRecord\":[],\"versionCreatedAt\":\"2022-12-23 03:38:12\",\"video\":\"\",\"vorDoi\":\"\",\"vorDoiUrl\":\"\",\"workflowStages\":[]},\"version\":\"v1\",\"identity\":\"rs-2392230\",\"journalConfig\":\"researchsquare\"},\"__N_SSP\":true},\"page\":\"/article/[identity]/[[...version]]\",\"query\":{\"redirect\":\"/article/rs-2392230\",\"identity\":\"rs-2392230\",\"version\":[\"v1\"]},\"buildId\":\"WrCJVZZCHTDjtuVLN7oU0\",\"isFallback\":false,\"isExperimentalCompile\":false,\"dynamicIds\":[84888],\"gssp\":true,\"scriptLoader\":[]}","source_license":"CC-BY-4.0","license_restricted":false}