Special attribute-based cross-modal interactive fusion network for RGBT tracking | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Special attribute-based cross-modal interactive fusion network for RGBT tracking Xiaoqiang Shao, Hao Li, Zhiyue Lyu, Bo Ma, Zehui HAN, Mingqian Liu This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4527385/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Due to the strong robustness of RGBT object tracking, which is less susceptible to the effects of illumination and occlusion, it has been widely used in the fields of video surveillance and automated driving. In this paper, an effective tracking network is constructed by fully interacting with both modal information using challenging attributes in infrared and visible images. The network consists of three parts: the special attribute fusion(SAF) module, the common attribute fusion(CAF) module, and the cross-modal interaction(CMI) module. The SAF module enables the network to extract unique challeng attribute information from two modalities, fully leveraging the advantages of different modal information. The CAF module extracts features from attributes matched in both modalities, and adaptively aggregates them, assigning corresponding weights to each challenging attribute to enhance the tracker's adaptability. The CMI module facilitates modal interaction between the infrared and visible image modalities, integrating common modal information with the specific modal information of each, thereby enhancing the network's robustness. The proposed network is tested on GTOT, RGBT234, and LasHer datasets, respectively. The results show that our tracker outperforms other trackers, proving the superiority of our method. Object tracking Thermal infrared tracking Challenge attributes Modal interaction Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4527385","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":313811053,"identity":"1c1d54e2-8de5-4d9c-b2d1-420e046df883","order_by":0,"name":"Xiaoqiang Shao","email":"","orcid":"","institution":"Xi'an University of Science and Technology","correspondingAuthor":false,"prefix":"","firstName":"Xiaoqiang","middleName":"","lastName":"Shao","suffix":""},{"id":313811054,"identity":"11c46123-cd6e-4836-a71d-15f021cb5609","order_by":1,"name":"Hao Li","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA2UlEQVRIiWNgGAWjYBACfvb+Dwc+GPyz42dvPkCcFsmeA4YPZxQcSJbsOZZAnBaDGw7GxjwfDjBuuJFjQKTLbjCkSc4wuMNscCPn4403DHZyug0EdDDObjgm8cHgGZ/kmbebLecwJBubHSCghVnmYBvQFmZmvuO526R5GA4kbiOkhU0imU2ax4CZseFAzjPitPBIpDEb8xgcZpxwIoeNOC0SPGcYH84wSAMFsrHlHAMi/GJ/vIfhwIc/NqCofHjjTYWdHEEtaFYSGzVIWkjVMQpGwSgYBSMCAAC6iUhyIUxDSwAAAABJRU5ErkJggg==","orcid":"","institution":"Xi'an University of Science and Technology","correspondingAuthor":true,"prefix":"","firstName":"Hao","middleName":"","lastName":"Li","suffix":""},{"id":313811055,"identity":"1f86364c-47a6-4e92-afb0-3c8871b55686","order_by":2,"name":"Zhiyue Lyu","email":"","orcid":"","institution":"Xi'an University of Science and Technology","correspondingAuthor":false,"prefix":"","firstName":"Zhiyue","middleName":"","lastName":"Lyu","suffix":""},{"id":313811056,"identity":"17d3c84c-fbde-4d92-ab55-d9ca20a58a5d","order_by":3,"name":"Bo Ma","email":"","orcid":"","institution":"Xi'an University of Science and Technology","correspondingAuthor":false,"prefix":"","firstName":"Bo","middleName":"","lastName":"Ma","suffix":""},{"id":313811057,"identity":"72970b6f-4c29-4a6e-b44d-2e87e2686171","order_by":4,"name":"Zehui HAN","email":"","orcid":"","institution":"Xi'an University of Science and Technology","correspondingAuthor":false,"prefix":"","firstName":"Zehui","middleName":"","lastName":"HAN","suffix":""},{"id":313811058,"identity":"91a5c71d-b90b-48c6-91c1-9a7dc479c503","order_by":5,"name":"Mingqian Liu","email":"","orcid":"","institution":"Xi'an University of Science and Technology","correspondingAuthor":false,"prefix":"","firstName":"Mingqian","middleName":"","lastName":"Liu","suffix":""}],"badges":[],"createdAt":"2024-06-04 10:42:32","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4527385/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4527385/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":79015234,"identity":"74f38fa8-c324-4a88-906f-8e7381592726","added_by":"auto","created_at":"2025-03-22 13:16:39","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":3381608,"visible":true,"origin":"","legend":"","description":"","filename":"SAFNet.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4527385/v1_covered_afb86bc7-dddf-4c86-93c8-c96ae06b669c.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Special attribute-based cross-modal interactive fusion network for RGBT tracking","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Object tracking, Thermal infrared tracking, Challenge attributes, Modal interaction","lastPublishedDoi":"10.21203/rs.3.rs-4527385/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4527385/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Due to the strong robustness of RGBT object tracking, which is less susceptible to the effects of illumination and occlusion, it has been widely used in the fields of video surveillance and automated driving. In this paper, an effective tracking network is constructed by fully interacting with both modal information using challenging attributes in infrared and visible images. The network consists of three parts: the special attribute fusion(SAF) module, the common attribute fusion(CAF) module, and the cross-modal interaction(CMI) module. The SAF module enables the network to extract unique challeng attribute information from two modalities, fully leveraging the advantages of different modal information. The CAF module extracts features from attributes matched in both modalities, and adaptively aggregates them, assigning corresponding weights to each challenging attribute to enhance the tracker's adaptability. The CMI module facilitates modal interaction between the infrared and visible image modalities, integrating common modal information with the specific modal information of each, thereby enhancing the network's robustness. The proposed network is tested on GTOT, RGBT234, and LasHer datasets, respectively. The results show that our tracker outperforms other trackers, proving the superiority of our method.","manuscriptTitle":"Special attribute-based cross-modal interactive fusion network for RGBT tracking","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-06-24 02:28:51","doi":"10.21203/rs.3.rs-4527385/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"438feb84-fcbb-4395-ac60-10b8d782f737","owner":[],"postedDate":"June 24th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-03-22T13:08:31+00:00","versionOfRecord":[],"versionCreatedAt":"2024-06-24 02:28:51","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-4527385","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4527385","identity":"rs-4527385","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.