{"paper_id":"33ff1500-ccde-4076-8f55-77a07f55a9f9","body_text":"Van-DETR: Enhanced Real-Time Object Detection with VanillaNet and Advanced Feature Fusion | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Van-DETR: Enhanced Real-Time Object Detection with VanillaNet and Advanced Feature Fusion Xinbiao Lu, Gaofan Zhan, Wen Wu, Wentao Zhang, Xiaolong Wu, Changjiang Han This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4814787/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 21 Oct, 2024 Read the published version in The Visual Computer → Version 1 posted 10 You are reading this latest preprint version Abstract Recently, end-to-end detectors based on transformer (DETRs) have made remarkable progress. However, their high computational cost still limits the performance of the DETRs series as real-time object detectors. In order to solve this problem, we introduce the Van-DETR model, which enhances the first real-time end-to-end object detector, RT-DETR. Specifically, we innovatively introduce a new, more lightweight backbone—VanillaNet, replacing the former backbone—ResNet. To address its weak nonlinearity and poor local analysis, we combine large kernel convolutions with small kernel convolutions to integrate global and local information, significantly enhancing feature extraction capabilities. Secondly, in the hybrid encoder, we cascade group process the features extracted by the backbone and design a gated linear unit with a star-shaped connection for intra-scale feature interaction. During the cross-scale feature fusion stage, we propose a high-low frequency feature fusion module with strong feature representation capabilities. To verify the effectiveness of the model, we conduct experiments on two public object detection datasets—visdrone dataset and a people dataset from roboflow. Experimental results show that the proposed Van-DETR model achieves MAP 50 of 0.471 and 0.730 on two object detection datasets, respectively, representing improvements of 4.5% and 2.8% over the original RT-DETR model. Source code is available at https://github.com/vangoghzz/Van-DETR . Real-Time Object Detection Transformer Feature Fusion Self-Attention Mechanism Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Published Journal Publication published 21 Oct, 2024 Read the published version in The Visual Computer → Version 1 posted Editorial decision: Revision requested 27 Aug, 2024 Reviews received at journal 21 Aug, 2024 Reviews received at journal 12 Aug, 2024 Reviewers agreed at journal 05 Aug, 2024 Reviewers agreed at journal 02 Aug, 2024 Reviewers agreed at journal 31 Jul, 2024 Reviewers invited by journal 31 Jul, 2024 Editor assigned by journal 28 Jul, 2024 Submission checks completed at journal 28 Jul, 2024 First submitted to journal 27 Jul, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {\"props\":{\"pageProps\":{\"initialData\":{\"identity\":\"rs-4814787\",\"acceptedTermsAndConditions\":true,\"allowDirectSubmit\":false,\"archivedVersions\":[],\"articleType\":\"Research Article\",\"associatedPublications\":[],\"authors\":[{\"id\":343253488,\"identity\":\"07eeb80d-4e2f-4f3a-8e56-a2e6af3aaa5b\",\"order_by\":0,\"name\":\"Xinbiao Lu\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"Hohai University\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Xinbiao\",\"middleName\":\"\",\"lastName\":\"Lu\",\"suffix\":\"\"},{\"id\":343253489,\"identity\":\"d30a7dff-6473-4783-babf-b568f7dd467c\",\"order_by\":1,\"name\":\"Gaofan Zhan\",\"email\":\"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABCklEQVRIie3Rv0vEMBTA8VcCmXJ36xOl9U94EogO95+45FDulgrnIjcUPSjEseuJP/4GQXBWApkK/gtxcRM6Od1gvHOTtqtDvkOGkg9pXgBisf8Y/ixUpHxUvvqG8KpizPp+MndyiO7kYLU4Sm6u+ZT6ScMmD5CrXVEvksc3sY9dIrstPzwQlxxqKQcGmbQCCIrxcRtJ7twhAe2lnJnz93uDXNnBiwc3PVu2EIZa4fYU+0SfBoWyQ03J0rYSjrOvQNjEQLDhx1CWgrCLCMzVLzlVO6JGItZDEPOLQJzk2yGjRhuGrDvukq1mzwjrIs2qzVNe6lFlrW+KcSvZjGD955Pu2B6LxWKx/r4BvvtTPg72/ZEAAAAASUVORK5CYII=\",\"orcid\":\"\",\"institution\":\"Hohai University\",\"correspondingAuthor\":true,\"prefix\":\"\",\"firstName\":\"Gaofan\",\"middleName\":\"\",\"lastName\":\"Zhan\",\"suffix\":\"\"},{\"id\":343253490,\"identity\":\"39d1431d-1c6c-4931-8f86-aad1c79f2f1f\",\"order_by\":2,\"name\":\"Wen Wu\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"Hohai University\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Wen\",\"middleName\":\"\",\"lastName\":\"Wu\",\"suffix\":\"\"},{\"id\":343253491,\"identity\":\"4197826c-db6f-4279-9b35-aec9f673915c\",\"order_by\":3,\"name\":\"Wentao Zhang\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"Hohai University\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Wentao\",\"middleName\":\"\",\"lastName\":\"Zhang\",\"suffix\":\"\"},{\"id\":343253492,\"identity\":\"dea718dc-84a1-4efa-bb39-94312f031adc\",\"order_by\":4,\"name\":\"Xiaolong Wu\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"Hohai University\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Xiaolong\",\"middleName\":\"\",\"lastName\":\"Wu\",\"suffix\":\"\"},{\"id\":343253493,\"identity\":\"5fa2a702-0631-46d2-9958-6553eb51b8dd\",\"order_by\":5,\"name\":\"Changjiang Han\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"Hohai University\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Changjiang\",\"middleName\":\"\",\"lastName\":\"Han\",\"suffix\":\"\"}],\"badges\":[],\"createdAt\":\"2024-07-28 02:08:20\",\"currentVersionCode\":1,\"declarations\":\"\",\"doi\":\"10.21203/rs.3.rs-4814787/v1\",\"doiUrl\":\"https://doi.org/10.21203/rs.3.rs-4814787/v1\",\"draftVersion\":[],\"editorialEvents\":[{\"content\":\"https://doi.org/10.1007/s00371-024-03656-0\",\"type\":\"published\",\"date\":\"2024-10-21T15:57:33+00:00\"}],\"editorialNote\":\"\",\"failedWorkflow\":false,\"files\":[{\"id\":67681879,\"identity\":\"53116001-89ba-47eb-8c25-cf45a52065ae\",\"added_by\":\"auto\",\"created_at\":\"2024-10-28 16:10:48\",\"extension\":\"pdf\",\"order_by\":1,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"manuscript-pdf\",\"size\":1196129,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"VanDETR.pdf\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-4814787/v1_covered_b567f88b-d9f2-471a-9d88-f5befabbc0b7.pdf\"}],\"financialInterests\":\"No competing interests reported.\",\"formattedTitle\":\"Van-DETR: Enhanced Real-Time Object Detection with VanillaNet and Advanced Feature Fusion\",\"fulltext\":[],\"fulltextSource\":\"\",\"fullText\":\"\",\"funders\":[],\"hasAdminPriorityOnWorkflow\":false,\"hasManuscriptDocX\":false,\"hasOptedInToPreprint\":true,\"hasPassedJournalQc\":\"\",\"hasAnyPriority\":false,\"hideJournal\":false,\"highlight\":\"\",\"institution\":\"\",\"isAcceptedByJournal\":true,\"isAuthorSuppliedPdf\":true,\"isDeskRejected\":\"\",\"isHiddenFromSearch\":false,\"isInQc\":false,\"isInWorkflow\":false,\"isPdf\":true,\"isPdfUpToDate\":true,\"isWithdrawnOrRetracted\":false,\"journal\":{\"display\":true,\"email\":\"info@researchsquare.com\",\"identity\":\"the-visual-computer\",\"isNatureJournal\":false,\"hasQc\":true,\"allowDirectSubmit\":false,\"externalIdentity\":\"tvcj\",\"sideBox\":\"Learn more about [The Visual Computer](http://link.springer.com/journal/371)\",\"snPcode\":\"371\",\"submissionUrl\":\"https://submission.nature.com/new-submission/371/3\",\"title\":\"The Visual Computer\",\"twitterHandle\":\"\",\"acdcEnabled\":true,\"dfaEnabled\":true,\"editorialSystem\":\"stoa\",\"reportingPortfolio\":\"Springer Hybrid\",\"inReviewEnabled\":true,\"inReviewRevisionsEnabled\":false},\"keywords\":\"Real-Time Object Detection, Transformer, Feature Fusion, Self-Attention Mechanism\",\"lastPublishedDoi\":\"10.21203/rs.3.rs-4814787/v1\",\"lastPublishedDoiUrl\":\"https://doi.org/10.21203/rs.3.rs-4814787/v1\",\"license\":{\"name\":\"CC BY 4.0\",\"url\":\"https://creativecommons.org/licenses/by/4.0/\"},\"manuscriptAbstract\":\"\\u003cp\\u003eRecently, end-to-end detectors based on transformer (DETRs) have made remarkable progress. However, their high computational cost still limits the performance of the DETRs series as real-time object detectors. In order to solve this problem, we introduce the Van-DETR model, which enhances the first real-time end-to-end object detector, RT-DETR. Specifically, we innovatively introduce a new, more lightweight backbone\\u0026mdash;VanillaNet, replacing the former backbone\\u0026mdash;ResNet. To address its weak nonlinearity and poor local analysis, we combine large kernel convolutions with small kernel convolutions to integrate global and local information, significantly enhancing feature extraction capabilities. Secondly, in the hybrid encoder, we cascade group process the features extracted by the backbone and design a gated linear unit with a star-shaped connection for intra-scale feature interaction. During the cross-scale feature fusion stage, we propose a high-low frequency feature fusion module with strong feature representation capabilities. To verify the effectiveness of the model, we conduct experiments on two public object detection datasets\\u0026mdash;visdrone dataset and a people dataset from roboflow. Experimental results show that the proposed Van-DETR model achieves MAP\\u003csub\\u003e50\\u003c/sub\\u003e of 0.471 and 0.730 on two object detection datasets, respectively, representing improvements of 4.5% and 2.8% over the original RT-DETR model. Source code is available at \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://github.com/vangoghzz/Van-DETR\\u003c/span\\u003e\\u003cspan address=\\\"https://github.com/vangoghzz/Van-DETR\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e.\\u003c/p\\u003e\",\"manuscriptTitle\":\"Van-DETR: Enhanced Real-Time Object Detection with VanillaNet and Advanced Feature Fusion\",\"msid\":\"\",\"msnumber\":\"\",\"nonDraftVersions\":[{\"code\":1,\"date\":\"2024-08-22 18:26:29\",\"doi\":\"10.21203/rs.3.rs-4814787/v1\",\"editorialEvents\":[{\"type\":\"communityComments\",\"content\":0},{\"type\":\"decision\",\"content\":\"Revision requested\",\"date\":\"2024-08-27T15:24:11+00:00\",\"index\":\"\",\"fulltext\":\"\"},{\"type\":\"editorInvitedReview\",\"content\":\"\",\"date\":\"2024-08-21T14:13:28+00:00\",\"index\":\"hide\",\"fulltext\":\"\"},{\"type\":\"editorInvitedReview\",\"content\":\"\",\"date\":\"2024-08-12T12:02:21+00:00\",\"index\":\"hide\",\"fulltext\":\"\"},{\"type\":\"reviewerAgreed\",\"content\":\"104705247326265047925845325705061065889\",\"date\":\"2024-08-05T08:59:23+00:00\",\"index\":\"hide\",\"fulltext\":\"\"},{\"type\":\"reviewerAgreed\",\"content\":\"49645443876516143626597866639660123622\",\"date\":\"2024-08-02T04:03:14+00:00\",\"index\":\"hide\",\"fulltext\":\"\"},{\"type\":\"reviewerAgreed\",\"content\":\"251907286183558236177290010481502412210\",\"date\":\"2024-07-31T07:40:38+00:00\",\"index\":\"hide\",\"fulltext\":\"\"},{\"type\":\"reviewersInvited\",\"content\":\"\",\"date\":\"2024-07-31T07:37:50+00:00\",\"index\":\"\",\"fulltext\":\"\"},{\"type\":\"editorAssigned\",\"content\":\"\",\"date\":\"2024-07-28T09:20:32+00:00\",\"index\":\"\",\"fulltext\":\"\"},{\"type\":\"checksComplete\",\"content\":\"\",\"date\":\"2024-07-28T08:26:49+00:00\",\"index\":\"\",\"fulltext\":\"\"},{\"type\":\"submitted\",\"content\":\"The Visual Computer\",\"date\":\"2024-07-28T02:04:52+00:00\",\"index\":\"\",\"fulltext\":\"\"}],\"status\":\"published\",\"journal\":{\"display\":true,\"email\":\"info@researchsquare.com\",\"identity\":\"the-visual-computer\",\"isNatureJournal\":false,\"hasQc\":true,\"allowDirectSubmit\":false,\"externalIdentity\":\"tvcj\",\"sideBox\":\"Learn more about [The Visual Computer](http://link.springer.com/journal/371)\",\"snPcode\":\"371\",\"submissionUrl\":\"https://submission.nature.com/new-submission/371/3\",\"title\":\"The Visual Computer\",\"twitterHandle\":\"\",\"acdcEnabled\":true,\"dfaEnabled\":true,\"editorialSystem\":\"stoa\",\"reportingPortfolio\":\"Springer Hybrid\",\"inReviewEnabled\":true,\"inReviewRevisionsEnabled\":false}}],\"origin\":\"\",\"ownerIdentity\":\"dd0d34fa-6aba-42c4-937c-59637a3f3e49\",\"owner\":[],\"postedDate\":\"August 22nd, 2024\",\"published\":true,\"recentEditorialEvents\":[],\"rejectedJournal\":[],\"revision\":\"\",\"amendment\":\"\",\"status\":\"published-in-journal\",\"subjectAreas\":[],\"tags\":[],\"updatedAt\":\"2024-10-28T16:01:51+00:00\",\"versionOfRecord\":{\"articleIdentity\":\"rs-4814787\",\"link\":\"https://doi.org/10.1007/s00371-024-03656-0\",\"journal\":{\"identity\":\"the-visual-computer\",\"isVorOnly\":false,\"title\":\"The Visual Computer\"},\"publishedOn\":\"2024-10-21 15:57:33\",\"publishedOnDateReadable\":\"October 21st, 2024\"},\"versionCreatedAt\":\"2024-08-22 18:26:29\",\"video\":\"\",\"vorDoi\":\"10.1007/s00371-024-03656-0\",\"vorDoiUrl\":\"https://doi.org/10.1007/s00371-024-03656-0\",\"workflowStages\":[]},\"version\":\"v1\",\"identity\":\"rs-4814787\",\"journalConfig\":\"researchsquare\"},\"__N_SSP\":true},\"page\":\"/article/[identity]/[[...version]]\",\"query\":{\"redirect\":\"/article/rs-4814787\",\"identity\":\"rs-4814787\",\"version\":[\"v1\"]},\"buildId\":\"qtupq5eGEP_6zYnWcrvyt\",\"isFallback\":false,\"isExperimentalCompile\":false,\"dynamicIds\":[84888],\"gssp\":true,\"scriptLoader\":[]}","source_license":"CC-BY-4.0","license_restricted":false}