Conditional Diffusion to Enhance Performance of Object Detection in Unbalanced Data Engineering Drawings

preprint OA: closed
Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-15

This study developed a conditional diffusion model that generates synthetic engineering drawings to improve object detection performance on unbalanced datasets, achieving a 12.54% higher mean average precision.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-15 · read from full text

The paper studies how to improve deep-learning object detection for digitizing relay-based railway interlocking system engineering drawings, where rare symbols create unbalanced training datasets. The authors propose a conditional diffusion–based synthetic drawing generator that follows RRIS schematic rules and probabilities, iteratively producing new samples targeted to classes/symbols where the model underperforms on original data. They report that the approach yields a mean average precision 12.54% higher than training using only the original data, while compared generators produced less variable samples and struggled more with unbalanced classes. The work is presented as a Research Square preprint and notes it has not undergone peer review. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract To train deep learning models, it is necessary to have a large amount of data available, preferably with balanced classes. Models often struggle to achieve acceptable performance because unbalanced classes drastically reduce the model's generalization capacity. In the case of applying deep learning for digitizing engineering drawings (EDs) of the rail networks, rare symbols are common as some drawings are used to describe specific scenarios, resulting in unbalanced datasets. In this paper, a synthetic EDs generator is proposed to augment the database of relay-based railway interlocking systems (RRIS). The synthetic EDs are created based on RRIS schematic rules and probabilities considering the performance of the model using original data, new samples are created based on the results that need to be improved. Considering conditional diffusion, the proposed method achieves a mean average precision 12.54\% higher than the results of the model using only original data. Other compared models had limited variability in the generated samples and struggled to deal with unbalanced classes. The proposed approach overcame these problems and proved to be a promising technique for generating synthetic drawings.
Full text 10,863 characters · extracted from preprint-html · click to expand
Conditional Diffusion to Enhance Performance of Object Detection in Unbalanced Data Engineering Drawings | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Conditional Diffusion to Enhance Performance of Object Detection in Unbalanced Data Engineering Drawings Stefano Frizzo Stefenon, Marco Cristoforetti, Alessandro Cimatti This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6204987/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract To train deep learning models, it is necessary to have a large amount of data available, preferably with balanced classes. Models often struggle to achieve acceptable performance because unbalanced classes drastically reduce the model's generalization capacity. In the case of applying deep learning for digitizing engineering drawings (EDs) of the rail networks, rare symbols are common as some drawings are used to describe specific scenarios, resulting in unbalanced datasets. In this paper, a synthetic EDs generator is proposed to augment the database of relay-based railway interlocking systems (RRIS). The synthetic EDs are created based on RRIS schematic rules and probabilities considering the performance of the model using original data, new samples are created based on the results that need to be improved. Considering conditional diffusion, the proposed method achieves a mean average precision 12.54\% higher than the results of the model using only original data. Other compared models had limited variability in the generated samples and struggled to deal with unbalanced classes. The proposed approach overcame these problems and proved to be a promising technique for generating synthetic drawings. Artificial Intelligence and Machine Learning Diffusion Models Engineering Drawings Unbalanced Data Synthetic Data Full Text Additional Declarations The authors declare no competing interests. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6204987","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":427359802,"identity":"0fa84060-7cf1-49c0-8ea1-ae1430ef1bc5","order_by":0,"name":"Stefano Frizzo Stefenon","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA8UlEQVRIie3QMUsDMRTA8RcCyfJuD5w0XyGlq3BfJUXoTUUn6VgQ0iUfoH4a3xG46XB2cKgU3EuXggomdFGwqW4O+Q+XW36X9w6gVPqPcQI6vrFNeo7kWpwhwh4JAjfpnODTLwh8JVN/jmhXvXSH98umkXdse7N4bn3tLOzcaWJ6aUjhbOox8Ml6eJ37i57YfY5oB2RUsKiu+7pyYf6g2iWvMkQ7AWRNaFBv5VskLSbykSHQR0I2MK+4iB9P182Is+wuArolxV2GK17jEMZxF9v5x+xgfJf+mFwFtsdF0Fi78eZwmxnsx+ivoFQqlUrf+wRtPU6uvJ8CtwAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0000-0002-3723-616X","institution":"ondazione Bruno Kessler, Via Sommarive 18, Trento, 38123, TN, Italy","correspondingAuthor":true,"prefix":"","firstName":"Stefano","middleName":"Frizzo","lastName":"Stefenon","suffix":""},{"id":427359803,"identity":"2ff0a55c-2ff8-4b1d-a25b-9c5eeab62b20","order_by":1,"name":"Marco Cristoforetti","email":"","orcid":"https://orcid.org/0000-0002-0127-1342","institution":"ondazione Bruno Kessler, Via Sommarive 18, Trento, 38123, TN, Italy","correspondingAuthor":false,"prefix":"","firstName":"Marco","middleName":"","lastName":"Cristoforetti","suffix":""},{"id":427359804,"identity":"dc84c28a-25fb-477d-ac1f-6fbbb8bd0b40","order_by":2,"name":"Alessandro Cimatti","email":"","orcid":"https://orcid.org/0000-0002-1315-6990","institution":"ondazione Bruno Kessler, Via Sommarive 18, Trento, 38123, TN, Italy","correspondingAuthor":false,"prefix":"","firstName":"Alessandro","middleName":"","lastName":"Cimatti","suffix":""}],"badges":[],"createdAt":"2025-03-11 16:16:56","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":true,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":true},"doi":"10.21203/rs.3.rs-6204987/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6204987/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":78417189,"identity":"2425f484-c7d0-4e42-b0c2-1f54896efa30","added_by":"auto","created_at":"2025-03-13 04:57:16","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2471107,"visible":true,"origin":"","legend":"","description":"","filename":"PaperNeuralComputingandApplications.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6204987/v1_covered_7e39347c-c503-451f-892f-77d2e4f35a86.pdf"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003eConditional Diffusion to Enhance Performance of Object Detection in Unbalanced Data Engineering Drawings\u003c/p\u003e","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"Fondazione Bruno Kessler","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Diffusion Models, Engineering Drawings, Unbalanced Data, Synthetic Data","lastPublishedDoi":"10.21203/rs.3.rs-6204987/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6204987/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eTo train deep learning models, it is necessary to have a large amount of data available, preferably with balanced classes. Models often struggle to achieve acceptable performance because unbalanced classes drastically reduce the model's generalization capacity. In the case of applying deep learning for digitizing engineering drawings (EDs) of the rail networks, rare symbols are common as some drawings are used to describe specific scenarios, resulting in unbalanced datasets. In this paper, a synthetic EDs generator is proposed to augment the database of relay-based railway interlocking systems (RRIS). The synthetic EDs are created based on RRIS schematic rules and probabilities considering the performance of the model using original data, new samples are created based on the results that need to be improved. Considering conditional diffusion, the proposed method achieves a mean average precision 12.54\\% higher than the results of the model using only original data. Other compared models had limited variability in the generated samples and struggled to deal with unbalanced classes. The proposed approach overcame these problems and proved to be a promising technique for generating synthetic drawings.\u003c/p\u003e","manuscriptTitle":"Conditional Diffusion to Enhance Performance of Object Detection in Unbalanced Data Engineering Drawings","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-03-13 04:49:07","doi":"10.21203/rs.3.rs-6204987/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"c05b1019-041a-4592-b78e-33b39dc32410","owner":[],"postedDate":"March 13th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":45538060,"name":"Artificial Intelligence and Machine Learning"}],"tags":[],"updatedAt":"2025-03-13T04:49:07+00:00","versionOfRecord":[],"versionCreatedAt":"2025-03-13 04:49:07","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6204987","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6204987","identity":"rs-6204987","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00