DiLLaB: Discussion Labeling with LLMs for Building Datasets | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article DiLLaB: Discussion Labeling with LLMs for Building Datasets Ludimila Gonçalves, Márcia Lima, André Carvalho, Walter Nakamura, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8620355/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 5 You are reading this latest preprint version Abstract GitHub Discussions has emerged as a prominent platform for collaborative knowledge exchange in open-source software (OSS) development. However, as participation increases, the platform faces challenges common to Programming Community-based Question and Answering (PCQA) environments, particularly the proliferation of duplicate and semantically related questions, which can fragment knowledge and reduce retrieval effectiveness. Although in-context links are often shared between related threads, no labeled dataset or automated method currently exists for identifying semantic relatedness in this setting. We present \texttt{DiLLaB}, a framework that leverages these in-context links for high-precision candidate selection and uses prompt-based Large Language Models (LLMs) to label discussion pairs as \emph{related} or \emph{unrelated}, supporting a graph-based, leakage-free pipeline for dataset construction. We evaluate \texttt{DiLLaB} across seven distinct labeling configurations, spanning from basic prompting to zero- and few-shot strategies, with examples drawn from within or across repositories and using either full or summarized input, to analyze how different prompting setups affect labeling effectiveness. A feasibility study confirms the reliability of link-based signals, and evaluation across five repositories shows that zero-shot prompting achieves strong labeling performance (F1-score $ > 0.90$). The resulting dataset enables effective fine-tuning of a \texttt{RoBERTa} classifier, achieving a 48% improvement in F1-score over a transfer learning baseline. Our results offer a scalable alternative to manual annotation, enables related post recommendation in GitHub Discussions, and lays a foundation for future research in discussion understanding within NLP for Software Engineering. Community-based Q&A Programming Community-based Question and Answering (PCQA) GitHub Discussions Large Language Model (LLM) Prompting Strategies Question relatedness Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Reviewers agreed at journal 18 May, 2026 Reviewers invited by journal 27 Jan, 2026 Editor assigned by journal 17 Jan, 2026 Submission checks completed at journal 17 Jan, 2026 First submitted to journal 16 Jan, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8620355","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":581266772,"identity":"cc340969-53de-463c-84e9-97eb4c86f5d4","order_by":0,"name":"Ludimila Gonçalves","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABTklEQVRIie3RPUvDQBgH8CcEziVN1kdK61e4EiiC1X6VhoO4XGvBpYNDpk59WdvZLxARnE8P4hJwDWQpFDpXBQkoapKmEm3rLJj/cHfP3f3uOA6gSJE/GB2ACqCH6zoZqEKsCmWWtg5oeUJWBNd1MiCtjKh0F8l2fhGNwq8E2dVd0MUDMC69WXSB1tjwH0XUu+kAMuUpgEbFFSVvmSd2V3KKNQcXp7WBh9Z00nZvB354HhO1zME2XaGzSZ5wmhDFQb9ejkuTBiVXKv3QcvBMxERartBM2CTNFXlHs/ngzzLC1FcOH7uI5RiDelnpY5VCfHJGSHyL2CDaIn0L6yOx94cjrGJg0/QtRJuTI06ZOZU6yxFjj10/87fG8diQHkYvDc0Yy/ky6oWdeEkNee+kMrofStgSgq3vEy2SdsnXqNtAcp/4QXbsK1KkSJH/l0/AHXtmPvTNlwAAAABJRU5ErkJggg==","orcid":"","institution":"Federal University of Amazonas","correspondingAuthor":true,"prefix":"","firstName":"Ludimila","middleName":"","lastName":"Gonçalves","suffix":""},{"id":581266773,"identity":"ea626720-6bdb-4400-a08e-ba8860e67c45","order_by":1,"name":"Márcia Lima","email":"","orcid":"","institution":"University of the State of Amazonas","correspondingAuthor":false,"prefix":"","firstName":"Márcia","middleName":"","lastName":"Lima","suffix":""},{"id":581266774,"identity":"f10be744-3336-4f71-886c-4c8d2332e019","order_by":2,"name":"André Carvalho","email":"","orcid":"","institution":"Federal University of Amazonas","correspondingAuthor":false,"prefix":"","firstName":"André","middleName":"","lastName":"Carvalho","suffix":""},{"id":581266775,"identity":"0fb91618-a72e-41b0-a832-4ce6330fe4dc","order_by":3,"name":"Walter Nakamura","email":"","orcid":"","institution":"Federal University of Technology – Paraná","correspondingAuthor":false,"prefix":"","firstName":"Walter","middleName":"","lastName":"Nakamura","suffix":""},{"id":581266776,"identity":"d6afb9e1-a741-4fd0-ada6-0af15f660f26","order_by":4,"name":"Igor Steinmacher","email":"","orcid":"","institution":"Northern Arizona University","correspondingAuthor":false,"prefix":"","firstName":"Igor","middleName":"","lastName":"Steinmacher","suffix":""},{"id":581266777,"identity":"277dc086-2b58-43e0-968a-127595a9bcc6","order_by":5,"name":"Tayana Conte","email":"","orcid":"","institution":"Federal University of Amazonas","correspondingAuthor":false,"prefix":"","firstName":"Tayana","middleName":"","lastName":"Conte","suffix":""}],"badges":[],"createdAt":"2026-01-16 15:08:43","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8620355/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8620355/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":101301840,"identity":"6b70afd9-f1ec-4bf3-942b-3b2ac0500d65","added_by":"auto","created_at":"2026-01-28 09:52:41","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1054348,"visible":true,"origin":"","legend":"","description":"","filename":"ASEDiLLaB.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8620355/v1_covered_28911ac1-c6b7-4bc6-b817-e785bcf138a5.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"DiLLaB: Discussion Labeling with LLMs for Building Datasets","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"automated-software-engineering","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"ause","sideBox":"Learn more about [Automated Software Engineering](http://link.springer.com/journal/10515)","snPcode":"10515","submissionUrl":"https://submission.nature.com/new-submission/10515/3","title":"Automated Software Engineering","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Community-based Q\u0026A, Programming Community-based Question and Answering (PCQA), GitHub Discussions, Large Language Model (LLM), Prompting Strategies, Question relatedness","lastPublishedDoi":"10.21203/rs.3.rs-8620355/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8620355/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"GitHub Discussions has emerged as a prominent platform for collaborative knowledge exchange in open-source software (OSS) development. However, as participation increases, the platform faces challenges common to Programming Community-based Question and Answering (PCQA) environments, particularly the proliferation of duplicate and semantically related questions, which can fragment knowledge and reduce retrieval effectiveness. Although in-context links are often shared between related threads, no labeled dataset or automated method currently exists for identifying semantic relatedness in this setting. We present \\texttt{DiLLaB}, a framework that leverages these in-context links for high-precision candidate selection and uses prompt-based Large Language Models (LLMs) to label discussion pairs as \\emph{related} or \\emph{unrelated}, supporting a graph-based, leakage-free pipeline for dataset construction. We evaluate \\texttt{DiLLaB} across seven distinct labeling configurations, spanning from basic prompting to zero- and few-shot strategies, with examples drawn from within or across repositories and using either full or summarized input, to analyze how different prompting setups affect labeling effectiveness. A feasibility study confirms the reliability of link-based signals, and evaluation across five repositories shows that zero-shot prompting achieves strong labeling performance (F1-score $ \u003e 0.90$). The resulting dataset enables effective fine-tuning of a \\texttt{RoBERTa} classifier, achieving a 48\\% improvement in F1-score over a transfer learning baseline. Our results offer a scalable alternative to manual annotation, enables related post recommendation in GitHub Discussions, and lays a foundation for future research in discussion understanding within NLP for Software Engineering.","manuscriptTitle":"DiLLaB: Discussion Labeling with LLMs for Building Datasets","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-01-28 09:38:23","doi":"10.21203/rs.3.rs-8620355/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"reviewerAgreed","content":"253018669803157961295615010567304506376","date":"2026-05-18T23:36:54+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-01-27T06:33:31+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-01-17T14:08:30+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-01-17T14:07:19+00:00","index":"","fulltext":""},{"type":"submitted","content":"Automated Software Engineering","date":"2026-01-16T14:58:20+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"automated-software-engineering","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"ause","sideBox":"Learn more about [Automated Software Engineering](http://link.springer.com/journal/10515)","snPcode":"10515","submissionUrl":"https://submission.nature.com/new-submission/10515/3","title":"Automated Software Engineering","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"20dcbec5-6689-4111-aa5f-00b3ce6a6b9b","owner":[],"postedDate":"January 28th, 2026","published":true,"recentEditorialEvents":[{"type":"reviewerAgreed","content":"253018669803157961295615010567304506376","date":"2026-05-18T23:36:54+00:00","index":32,"fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-01-28T09:38:23+00:00","versionOfRecord":[],"versionCreatedAt":"2026-01-28 09:38:23","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8620355","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8620355","identity":"rs-8620355","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.