CSegNet:A Crack Segmentation Network Combining CNN and Transformer | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article CSegNet:A Crack Segmentation Network Combining CNN and Transformer Hao Dong, Yinlai Du, Dong Feng, Qingyuan Hu, Mingzhu Zhou, Jun Xing, and 3 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3925781/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Detecting cracks from images plays a crucial role in road maintenance. Road cracks exhibit significant diversity and complexity in terms of shape, size, texture, and road images may contain various noises and interferences such as lighting variations, shadows, and different appearances due to varying perspectives and scales. To address these challenges, we constructed a comprehensive dataset called the Comprehensive Road Crack Dataset (CRCrack Dataset), which encompasses various crack characteristics. In this study, we propose a road crack segmentation network called CSegNet (Crack Segmentation Network), which combines convolutional neural networks (CNNs) and Transformers. The network adopts an encoder-decoder framework, like DeepLab V3+. In the encoder, leveraging the flexibility of Transformers in modeling long-term dependencies and the ability of CNNs to capture local contextual information through local receptive fields, weight sharing, and spatial subsampling, we design a ResNeXTR (ResNeXt-Transformer) feature extraction module as the backbone network to enhance the feature extraction capability for road crack images. To reduce the computational cost in self-attention computation of transformer, we introduce an average pooling layer to downsample the dimensions of the encoded features. In the decoder, to focus on the key information of road cracks under diverse environmental conditions and interferences, we combine the Efficient Channel Attention Module (ECAM) and the Spatial Attention Module (SAM) to design an Efficient Convolutional Block Attention Module (ECBAM) attention module to further optimize feature representation. Additionally, we employ the ReLU activation function, SGD gradient descent, and a hybrid loss function of Binary Cross Entropy with Logits to accelerate convergence speed and improve segmentation accuracy. Through comparative experiments on the CRCrack dataset, the results demonstrate that our proposed method outperforms classic networks such as U-Net and DeepLab V3 + in terms of IoU, Dice, and AUROC evaluation metrics. It exhibits good adaptability to ground crack images from different sources, providing a basis for estimating the degree of road damage. Physical sciences/Mathematics and computing/Information technology Physical sciences/Mathematics and computing/Computer science Road Crack Crack Segmentation Transformer CNN Full Text Additional Declarations No competing interests reported. Supplementary Files CRCrack.rar Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3925781","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":273506320,"identity":"54f75ea7-80e1-4b83-a02f-67e1caaeb86a","order_by":0,"name":"Hao Dong","email":"","orcid":"","institution":"Science Island Branch of Graduate School University of Science and Technology of China","correspondingAuthor":false,"prefix":"","firstName":"Hao","middleName":"","lastName":"Dong","suffix":""},{"id":273506321,"identity":"9e3c0e26-f2c2-4bb2-86e3-a4d5b019a542","order_by":1,"name":"Yinlai Du","email":"","orcid":"","institution":"Anhui University","correspondingAuthor":false,"prefix":"","firstName":"Yinlai","middleName":"","lastName":"Du","suffix":""},{"id":273506322,"identity":"c72fb6b5-7e58-4911-9353-b693f3fab840","order_by":2,"name":"Dong Feng","email":"","orcid":"","institution":"Hefei Institutes of Physical Science","correspondingAuthor":false,"prefix":"","firstName":"Dong","middleName":"","lastName":"Feng","suffix":""},{"id":273506323,"identity":"606a59b8-244b-448b-a66a-833a0eaeb4c1","order_by":3,"name":"Qingyuan Hu","email":"","orcid":"","institution":"Science Island Branch of Graduate School University of Science and Technology of China","correspondingAuthor":false,"prefix":"","firstName":"Qingyuan","middleName":"","lastName":"Hu","suffix":""},{"id":273506324,"identity":"b590bc30-db73-40d3-be9b-c9bbb3fcd421","order_by":4,"name":"Mingzhu Zhou","email":"","orcid":"","institution":"China National Tobacco Quality Supervision and Test Center","correspondingAuthor":false,"prefix":"","firstName":"Mingzhu","middleName":"","lastName":"Zhou","suffix":""},{"id":273506325,"identity":"c64eac4c-8de4-4af7-854d-f615251d44bc","order_by":5,"name":"Jun Xing","email":"","orcid":"","institution":"China National Tobacco Quality Supervision and Test Center","correspondingAuthor":false,"prefix":"","firstName":"Jun","middleName":"","lastName":"Xing","suffix":""},{"id":273506326,"identity":"fa138098-00ff-4b36-8f08-af4ce157fd19","order_by":6,"name":"Long Zhang","email":"","orcid":"","institution":"Science Island Branch of Graduate School University of Science and Technology of China","correspondingAuthor":false,"prefix":"","firstName":"Long","middleName":"","lastName":"Zhang","suffix":""},{"id":273506327,"identity":"fbf578ab-5ab2-4304-978e-90e24c4832cb","order_by":7,"name":"Shu Wang","email":"","orcid":"","institution":"Hefei Institutes of Physical Science","correspondingAuthor":false,"prefix":"","firstName":"Shu","middleName":"","lastName":"Wang","suffix":""},{"id":273506328,"identity":"71800d12-e5e6-43a0-8408-5034c610ac3a","order_by":8,"name":"Yong Liu","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAw0lEQVRIiWNgGAWjYBACPmYIzcPA3tj48AMxWtigWmQYeA43G0sQpQVK2zBIpLcJ8BClhZ3HTOLjjloe+ZkP2xgkGOzkdBsIOozHTHLmmeM8BrcT2x4UMCQbmx0gQos0b9sxHgPpxHYDCYYDiduI1iI/82CbBA8JWmp4GG4wEq2FrdhyZtsBHoMzicBANiDCL/z8hzfe+NhWZy/ffvzhww8VdnIEtQABCzACD0PZBoSVgwAzMJnUEad0FIyCUTAKRiYAAJ3zOPG0RRPSAAAAAElFTkSuQmCC","orcid":"","institution":"Science Island Branch of Graduate School University of Science and Technology of China","correspondingAuthor":true,"prefix":"","firstName":"Yong","middleName":"","lastName":"Liu","suffix":""}],"badges":[],"createdAt":"2024-02-04 01:29:08","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3925781/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3925781/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":56798314,"identity":"fada587e-77f3-47aa-a84d-bba7f0d57f1e","added_by":"auto","created_at":"2024-05-20 15:14:23","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":832787,"visible":true,"origin":"","legend":"","description":"","filename":"CSegNetACrackSegmentationNetworkCombiningCNNandTransformer.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3925781/v1_covered_e7e3fe71-3623-4241-87b0-a6226c70f581.pdf"},{"id":51386587,"identity":"cce8e054-b73e-4d66-9514-d392ebf2877e","added_by":"auto","created_at":"2024-02-20 17:48:25","extension":"rar","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":1062818274,"visible":true,"origin":"","legend":"","description":"","filename":"CRCrack.rar","url":"https://assets-eu.researchsquare.com/files/rs-3925781/v1/c451cc9bc308e3c15d768308.rar"}],"financialInterests":"No competing interests reported.","formattedTitle":"CSegNet:A Crack Segmentation Network Combining CNN and Transformer","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Road Crack, Crack Segmentation, Transformer, CNN","lastPublishedDoi":"10.21203/rs.3.rs-3925781/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3925781/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eDetecting cracks from images plays a crucial role in road maintenance. Road cracks exhibit significant diversity and complexity in terms of shape, size, texture, and road images may contain various noises and interferences such as lighting variations, shadows, and different appearances due to varying perspectives and scales. To address these challenges, we constructed a comprehensive dataset called the Comprehensive Road Crack Dataset (CRCrack Dataset), which encompasses various crack characteristics. In this study, we propose a road crack segmentation network called CSegNet (Crack Segmentation Network), which combines convolutional neural networks (CNNs) and Transformers. The network adopts an encoder-decoder framework, like DeepLab V3+. In the encoder, leveraging the flexibility of Transformers in modeling long-term dependencies and the ability of CNNs to capture local contextual information through local receptive fields, weight sharing, and spatial subsampling, we design a ResNeXTR (ResNeXt-Transformer) feature extraction module as the backbone network to enhance the feature extraction capability for road crack images. To reduce the computational cost in self-attention computation of transformer, we introduce an average pooling layer to downsample the dimensions of the encoded features. In the decoder, to focus on the key information of road cracks under diverse environmental conditions and interferences, we combine the Efficient Channel Attention Module (ECAM) and the Spatial Attention Module (SAM) to design an Efficient Convolutional Block Attention Module (ECBAM) attention module to further optimize feature representation. Additionally, we employ the ReLU activation function, SGD gradient descent, and a hybrid loss function of Binary Cross Entropy with Logits to accelerate convergence speed and improve segmentation accuracy. Through comparative experiments on the CRCrack dataset, the results demonstrate that our proposed method outperforms classic networks such as U-Net and DeepLab V3\u0026thinsp;+\u0026thinsp;in terms of IoU, Dice, and AUROC evaluation metrics. It exhibits good adaptability to ground crack images from different sources, providing a basis for estimating the degree of road damage.\u003c/p\u003e","manuscriptTitle":"CSegNet:A Crack Segmentation Network Combining CNN and Transformer","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-02-20 17:47:08","doi":"10.21203/rs.3.rs-3925781/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"7c2bdec3-57bf-41c0-8c02-c410d4d3b952","owner":[],"postedDate":"February 20th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":28823596,"name":"Physical sciences/Mathematics and computing/Information technology"},{"id":28823597,"name":"Physical sciences/Mathematics and computing/Computer science"}],"tags":[],"updatedAt":"2024-05-20T15:06:14+00:00","versionOfRecord":[],"versionCreatedAt":"2024-02-20 17:47:08","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-3925781","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3925781","identity":"rs-3925781","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.