Software testing with performance profiles clustering using unsupervised learning | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Software testing with performance profiles clustering using unsupervised learning Vitor Silva Montes, Diego Braga, Thiago de Jesus Oliveira Durães, and 3 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-1857904/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract The functional testing technique has been widely applied to reveal unknown faults in the software caused by programmer mistakes. Nevertheless, in autonomous data processing systems with highly variable inputs and outputs, such as embedded applications, data streams, and machine learning algorithms, the non-functional testing helps to reveal unknown-nontrivial faults in the deployment of software products. This paper addresses the detection of unknown code-faults by employing performance analysis as a nonfunctional requirement without predefined test cases, test oracles, or source-code analyses. The premise is that codefaults change demands for hardware resources during software execution, and a novel testing methodology can automatically detect them. This paper proposes the Tricorder testing methodology for automating workload characterization and detecting potential performance anomalies, caused by code-faults in autonomous data processing systems. Tricorder evaluates the performance profiles of hardware regarding the detection of the source code-faults. DAMICORE, a non-parametric multipurpose clustering methodology, enables Tricorder to group performance profiles of the software under testing and identify performance anomalies using non-parametric data in unsupervised learning based on Normalized Compression Distance (NCD). Tricorder reveals unknown source code faults, with no specialist to determine standards for input and output data, previous models inherent to architecture, test case creation, or workload characterization. We evaluate the capability of Tricorder in revealing faults through experiments based on three benchmarks: cryptography system, machine learning algorithm, and data stream processing server. Tricorder detects faults even under various workloads for different applications in our experiments. Tricorder helps the maintenance phase of the software development life cycle, providing additional information regarding the proper functioning of the application release before and after the updating process. This work contributes to the cost reduction of regression testing during the maintenance phase of autonomous data processing applications and can be used as a complementary testing technique. Functional Software Testing Performance Analysis Test Automation Unsupervised Machine Learning DAMICORE Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-1857904","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":122733790,"identity":"1cd96827-3e07-4026-92ac-1cf2cd75920e","order_by":0,"name":"Vitor Silva Montes","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAuklEQVRIiWNgGAWjYBADORBx4AHR6g8wMBiD6QRStCQ2gBhEaTG4dvjY4w8Vdenzww4/BNpiJ6fbQEjL7bR0gwNnDuduvJ1mANSSbGx2gIAWydk5ZhIH2w7kbpydANJyIHEbcVr+1aUbzk7/QJwWfmmQlgbmBHnpHCJt4ZdOS5M4c+yw4QbpnIIDCQZE+IVNOvmYREVNnbz87PTNHz5U2MkR1AIHBmCVBsQqBwH5BlJUj4JRMApGwYgCAJmrRqPkpLBjAAAAAElFTkSuQmCC","orcid":"","institution":"University of São Paulo","correspondingAuthor":true,"prefix":"","firstName":"Vitor","middleName":"Silva","lastName":"Montes","suffix":""},{"id":122733791,"identity":"b7a53bfa-8b11-42da-b08b-dba1a492a0d2","order_by":1,"name":"Diego Braga","email":"","orcid":"","institution":"University of São Paulo","correspondingAuthor":false,"prefix":"","firstName":"Diego","middleName":"","lastName":"Braga","suffix":""},{"id":122733792,"identity":"2697df48-e89f-4f4c-b728-abc0b2911a62","order_by":2,"name":"Thiago de Jesus Oliveira Durães","email":"","orcid":"","institution":"University of São Paulo","correspondingAuthor":false,"prefix":"","firstName":"Thiago","middleName":"de Jesus Oliveira","lastName":"Durães","suffix":""},{"id":122733793,"identity":"8c9c9507-3b27-4b3f-a70d-544fc9d6dc75","order_by":3,"name":"Alexandre Claudio Delbem","email":"","orcid":"","institution":"University of São Paulo","correspondingAuthor":false,"prefix":"","firstName":"Alexandre","middleName":"Claudio","lastName":"Delbem","suffix":""},{"id":122733794,"identity":"fb993458-e410-441f-b822-46b745f297ac","order_by":4,"name":"Simone Senger Souza","email":"","orcid":"","institution":"University of São Paulo","correspondingAuthor":false,"prefix":"","firstName":"Simone","middleName":"Senger","lastName":"Souza","suffix":""},{"id":122733795,"identity":"7ce8e9d5-b3ab-4b45-81a4-2a7e7411212c","order_by":5,"name":"Paulo Sergio Lopes Souza","email":"","orcid":"","institution":"University of São Paulo","correspondingAuthor":false,"prefix":"","firstName":"Paulo","middleName":"Sergio Lopes","lastName":"Souza","suffix":""}],"badges":[],"createdAt":"2022-07-14 12:14:17","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-1857904/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-1857904/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":24215387,"identity":"7f7ef91b-2c95-4555-92c2-96108fcca8a3","added_by":"auto","created_at":"2022-07-22 18:02:33","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":522455,"visible":true,"origin":"","legend":"","description":"","filename":"TricorderSQJ20220713.pdf","url":"https://assets-eu.researchsquare.com/files/rs-1857904/v1_covered.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Software testing with performance profiles clustering using unsupervised learning","fulltext":[{"header":"Full Text","content":"This preprint is available for \u003ca href='/article/rs-1857904/latest.pdf' target='_blank'\u003edownload as a PDF\u003c/a\u003e."}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Functional Software Testing, Performance Analysis, Test Automation, Unsupervised Machine Learning, DAMICORE","lastPublishedDoi":"10.21203/rs.3.rs-1857904/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-1857904/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"The functional testing technique has been widely applied to reveal unknown faults in the software caused by programmer mistakes. Nevertheless, in autonomous data processing systems with highly variable inputs and outputs, such as embedded applications, data streams, and machine learning algorithms, the non-functional testing helps to reveal unknown-nontrivial faults in the deployment of software products. This paper addresses the detection of unknown code-faults by employing performance analysis as a nonfunctional requirement without predefined test cases, test oracles, or source-code analyses. The premise is that codefaults change demands for hardware resources during software execution, and a novel testing methodology can automatically detect them. This paper proposes the Tricorder testing methodology for automating workload characterization and detecting potential performance anomalies, caused by code-faults in autonomous data processing systems. Tricorder evaluates the performance profiles of hardware regarding the detection of the source code-faults. DAMICORE, a non-parametric multipurpose clustering methodology, enables Tricorder to group performance profiles of the software under testing and identify performance anomalies using non-parametric data in unsupervised learning based on Normalized Compression Distance (NCD). Tricorder reveals unknown source code faults, with no specialist to determine standards for input and output data, previous models inherent to architecture, test case creation, or workload characterization. We evaluate the capability of Tricorder in revealing faults through experiments based on three benchmarks: cryptography system, machine learning algorithm, and data stream processing server. Tricorder detects faults even under various workloads for different applications in our experiments. Tricorder helps the maintenance phase of the software development life cycle, providing additional information regarding the proper functioning of the application release before and after the updating process. This work contributes to the cost reduction of regression testing during the maintenance phase of autonomous data processing applications and can be used as a complementary testing technique.","manuscriptTitle":"Software testing with performance profiles clustering using unsupervised learning","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2022-07-22 18:02:26","doi":"10.21203/rs.3.rs-1857904/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"96d71c11-6cb6-4a58-aba2-4e3bd3f6f3b5","owner":[],"postedDate":"July 22nd, 2022","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2023-10-08T22:59:14+00:00","versionOfRecord":[],"versionCreatedAt":"2022-07-22 18:02:26","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-1857904","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-1857904","identity":"rs-1857904","version":["v1"]},"buildId":"_2-kVJe1T_tPrBINL-cwx","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.