Multiclass Classification and Prioritisation of Static Analysis Warnings Using Developer-Labelled Industrial Data | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Multiclass Classification and Prioritisation of Static Analysis Warnings Using Developer-Labelled Industrial Data Benedikt Fein, Vibhash Kumar Singh, Gordon Fraser This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6469376/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Automatic static code analysis tools are used to identify code quality issues like vulnerabilities or performance problems. In practice, the high number of irrelevant warnings created by such tools is problematic, but can be addressed by pre-filtering and ranking these warnings before they are shown to the developer. Since ground truth labelled data is rarely available, existing research tends to heuristically construct training data from unlabelled open-source data by assigning clearly separated binary categories like fixed and irrelevant to the warnings. However, this labelling approach cannot capture distinctions such as between relevant, but not yet fixed and not fixed, because irrelevant warnings. The Teamscale software developed by CQSE provides a unique opportunity to investigate this concern, since its developers have adopted the practice of meticulously labelling all static analysis warnings as either accepted, tolerated, or false-positive. Using this dataset, we adapt previously proposed models for a new multiclass classification task to evaluate both their ability to classify warnings in this new setting and their ability to prioritise important warnings. Our experiments show that especially the mentioned subtle difference between categories of unresolved warnings is more challenging for the models compared to a binary prediction. Nevertheless, training the models for the multiclass task rather than the binary one results in a statistically significant improvement in the prioritisation of the warnings. static code analysis warning prioritisation software quality Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6469376","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":448918695,"identity":"a58610f5-2270-4e4a-8f14-48b327c37c52","order_by":0,"name":"Benedikt Fein","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABJklEQVRIie2RPUvEQBCG5whsmtmkzRHk/sIegYTjwN+SsBCb4Ac2B4pulSpiLQT9F1cnLGhzWGvlgXBVijtsIqYwFxEVFz86wX2KKXbm4Z1hATSaPwhbF2QhgFFAARMA8tYMv1QcIGGrzH6swIsCvfRDU60E5kk5r3bjIyBGKffO5bYFtFzVzWYkTF6olFF2zYc5S9oUEsqzqdwnYHGXpjwSuFDGsJvEd5FNHBhkTNKpjNJBxdyeMCLhJEyp3FXBU6cQeylp3iqA3mPdHLfKzlKdgr6B3WIIkopO8R0kcp2iPH+UJV4/Z3E/JYRJvNxqFSse0/TKS3GhXCwwZ8Nl1XDbJsb9Ax6OowtB5W3dHGycmnyujHnl3Q/ip5dvwV/MajQazX/gGTCSWtxuRSbKAAAAAElFTkSuQmCC","orcid":"","institution":"University of Passau","correspondingAuthor":true,"prefix":"","firstName":"Benedikt","middleName":"","lastName":"Fein","suffix":""},{"id":448918696,"identity":"a87e53c5-9fcf-4afc-80ff-3f77f828ce6f","order_by":1,"name":"Vibhash Kumar Singh","email":"","orcid":"","institution":"CQSE","correspondingAuthor":false,"prefix":"","firstName":"Vibhash","middleName":"Kumar","lastName":"Singh","suffix":""},{"id":448918697,"identity":"f329ffbf-0c81-4971-b867-c047f028cae9","order_by":2,"name":"Gordon Fraser","email":"","orcid":"","institution":"University of Passau","correspondingAuthor":false,"prefix":"","firstName":"Gordon","middleName":"","lastName":"Fraser","suffix":""}],"badges":[],"createdAt":"2025-04-17 07:53:22","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-6469376/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6469376/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":84577889,"identity":"9eecebe5-f155-4996-a198-b41314d1be15","added_by":"auto","created_at":"2025-06-13 17:31:36","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":584414,"visible":true,"origin":"","legend":"","description":"","filename":"paper.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6469376/v1_covered_f6fdf95c-235e-448b-886b-f4eeddc96f99.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Multiclass Classification and Prioritisation of Static Analysis Warnings Using Developer-Labelled Industrial Data","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":true,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"static code analysis warning prioritisation, software quality","lastPublishedDoi":"10.21203/rs.3.rs-6469376/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6469376/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Automatic static code analysis tools are used to identify code quality issues like\nvulnerabilities or performance problems. In practice, the high number of irrelevant\nwarnings created by such tools is problematic, but can be addressed by pre-filtering\nand ranking these warnings before they are shown to the developer. Since ground\ntruth labelled data is rarely available, existing research tends to heuristically\nconstruct training data from unlabelled open-source data by assigning clearly\nseparated binary categories like fixed and irrelevant to the warnings. However,\nthis labelling approach cannot capture distinctions such as between relevant, but\nnot yet fixed and not fixed, because irrelevant warnings. The Teamscale software\ndeveloped by CQSE provides a unique opportunity to investigate this concern,\nsince its developers have adopted the practice of meticulously labelling all static\nanalysis warnings as either accepted, tolerated, or false-positive. Using this dataset,\nwe adapt previously proposed models for a new multiclass classification task\nto evaluate both their ability to classify warnings in this new setting and their\nability to prioritise important warnings. Our experiments show that especially\nthe mentioned subtle difference between categories of unresolved warnings is\nmore challenging for the models compared to a binary prediction. Nevertheless,\ntraining the models for the multiclass task rather than the binary one results in\na statistically significant improvement in the prioritisation of the warnings.","manuscriptTitle":"Multiclass Classification and Prioritisation of Static Analysis Warnings Using Developer-Labelled Industrial Data","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-04-30 06:32:09","doi":"10.21203/rs.3.rs-6469376/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"18c88bd3-53ce-4577-841a-3bc4cd067688","owner":[],"postedDate":"April 30th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-06-13T17:23:22+00:00","versionOfRecord":[],"versionCreatedAt":"2025-04-30 06:32:09","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6469376","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6469376","identity":"rs-6469376","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.