Independent Benchmarking of Prompt-Based Medical Segmentation Models | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Independent Benchmarking of Prompt-Based Medical Segmentation Models Ayhan Can Erdur, Daniel Scholz, Josef A. Buchner, Denise Bernhardt, and 4 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7593773/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Medical image segmentation rapidly shifts toward vision(-language) foundation models that unify diverse modalities and tasks within a single framework. In this work, we systematically benchmark high-impact vision-language and segment-anything-based architectures across multiple clinically relevant CT and MRI tasks. We show that while these models achieve strong performance, each comes with specific (dis)advantages. Non-3D models are highly flexible but require substantial user guidance and are prone to over- or under-detection. 3D architectures offer overall more reliable volumetric consistency, but can still have detection problems. Vision-language models appear sensitive to the coverage of training data, whereas click-prompted SAM-based models are more universal, with a, though limited, ability to address zero-shot targets. When tested with more complex text prompts, most vision-language models exhibit missing semantic language understanding. Overall, these models hold considerable promise but still express limitations. Our work highlights key areas where future research is needed to advance vision(-language) foundation models. Biological sciences/Computational biology and bioinformatics Health sciences/Health care Physical sciences/Mathematics and computing Health sciences/Medical research Segmentation Deep Learning Vision-Language Models Foundation Models 1 Full Text Additional Declarations Competing interest reported. - SEC: Grants, Consulting fees and Honoraria: Roche, AstraZeneca, Medac, Dr. Sennewald Medizintechnik, Elekta, Accuray, BMS, Brainlab, Daiichi Sankyo, Icotec AG, Carl Zeiss Meditec AG, HMG Systems Engineering, Janssen; Safety monitoring/advisory board: CureVac DSMB Member; Leadership: NOA Board Member, DEGRO Board Member DR: Grants: DFG, ERC, EPSRC, BMBF, Alexander von Humboldt Stiftung; Consulting fees: ERC BW: Grants: DFG, NIH, Deutsche Krebshilfe, BMWi; Consulting fees and Stock: Need; Honoraria: Philips, Novartis JCP: Honoraria: AstraZeneca, Support for current manuscript: German Research Foundation, Wilhelm Sander-Stiftung, Else Kr¨oner Fresenius Stiftung The remaining authors have no competing interests to disclose. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7593773","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":522480616,"identity":"e90fa82d-ecc4-4d5b-a119-ee2141b9bec5","order_by":0,"name":"Ayhan Can Erdur","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAtUlEQVRIiWNgGAWjYPACZgYG9gYog3gtPIdJ1iKRTKQW/vazBxg+7rFO3D7z/TGpGwzWcgS1SJzJS2Cc8Sw9cc7tZDbpHIZ0Y4JaDBhyDJh5DhxOnCEN1nI4sYGgFv43Bsx/QFokD4O11BPWIgG0hQGkRYIZrCWBsF9uvDE42HMg3XgGT7KxdY5BuiFBW/j7cwwf/DhgLTuD/eDD2zkV1vIEbQGBA0juJErDKBgFo2AUjAJCAAA44jYgaizKdwAAAABJRU5ErkJggg==","orcid":"","institution":"TUM University Hospital","correspondingAuthor":true,"prefix":"","firstName":"Ayhan","middleName":"Can","lastName":"Erdur","suffix":""},{"id":522480617,"identity":"3de29870-fc56-4561-91cc-2f7e3643db6a","order_by":1,"name":"Daniel Scholz","email":"","orcid":"","institution":"Technical University of Munich (TUM) and TUM University Hospital","correspondingAuthor":false,"prefix":"","firstName":"Daniel","middleName":"","lastName":"Scholz","suffix":""},{"id":522480618,"identity":"549cb803-9b59-481b-b9f0-a12522df11c1","order_by":2,"name":"Josef A. Buchner","email":"","orcid":"","institution":"TUM University Hospital","correspondingAuthor":false,"prefix":"","firstName":"Josef","middleName":"A.","lastName":"Buchner","suffix":""},{"id":522480619,"identity":"49e69f8d-db53-47f9-8eaa-c77e9a73c710","order_by":3,"name":"Denise Bernhardt","email":"","orcid":"","institution":"TUM University Hospital","correspondingAuthor":false,"prefix":"","firstName":"Denise","middleName":"","lastName":"Bernhardt","suffix":""},{"id":522480620,"identity":"e6ea6f9b-2a97-45b4-a74c-d0379034b91e","order_by":4,"name":"Stephanie E. Combs","email":"","orcid":"","institution":"TUM University Hospital","correspondingAuthor":false,"prefix":"","firstName":"Stephanie","middleName":"E.","lastName":"Combs","suffix":""},{"id":522480621,"identity":"818cc9c9-12be-44ae-b8f5-b3d257989664","order_by":5,"name":"Benedikt Wiestler","email":"","orcid":"","institution":"Technical University of Munich (TUM) and TUM University Hospital","correspondingAuthor":false,"prefix":"","firstName":"Benedikt","middleName":"","lastName":"Wiestler","suffix":""},{"id":522480622,"identity":"bc859474-d165-426a-9dde-c77582a520ef","order_by":6,"name":"Daniel Rueckert","email":"","orcid":"","institution":"Technical University of Munich (TUM) and TUM University Hospital","correspondingAuthor":false,"prefix":"","firstName":"Daniel","middleName":"","lastName":"Rueckert","suffix":""},{"id":522480623,"identity":"50e9f1c6-c52d-4597-99f1-242ea66a0190","order_by":7,"name":"Jan C. Peeken","email":"","orcid":"","institution":"TUM University Hospital","correspondingAuthor":false,"prefix":"","firstName":"Jan","middleName":"C.","lastName":"Peeken","suffix":""}],"badges":[],"createdAt":"2025-09-11 16:23:23","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7593773/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7593773/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":93262832,"identity":"ad2bc50a-a401-4d70-847f-81e7eeff0cc6","added_by":"auto","created_at":"2025-10-10 18:53:45","extension":"json","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":9908,"visible":true,"origin":"","legend":"","description":"","filename":"bcfdb70b53ff4c4cac22c724e9899c38.json","url":"https://assets-eu.researchsquare.com/files/rs-7593773/v1/4aa719385454111f8d4e7b58.json"},{"id":95313713,"identity":"d7a9047f-5e30-4ee3-87a5-1cbf868f453d","added_by":"auto","created_at":"2025-11-06 15:51:54","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":17248867,"visible":true,"origin":"","legend":"","description":"","filename":"submission.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7593773/v1_covered_e169fbac-ffa3-41a2-9a3f-f478210eea22.pdf"}],"financialInterests":"Competing interest reported. - SEC: Grants, Consulting fees and Honoraria: Roche, AstraZeneca, Medac, Dr.\nSennewald Medizintechnik, Elekta, Accuray, BMS, Brainlab, Daiichi Sankyo,\nIcotec AG, Carl Zeiss Meditec AG, HMG Systems Engineering, Janssen; Safety\nmonitoring/advisory board: CureVac DSMB Member; Leadership: NOA Board\nMember, DEGRO Board Member\n- DR: Grants: DFG, ERC, EPSRC, BMBF, Alexander von Humboldt Stiftung;\nConsulting fees: ERC\n- BW: Grants: DFG, NIH, Deutsche Krebshilfe, BMWi; Consulting fees and Stock:\nNeed; Honoraria: Philips, Novartis\n- JCP: Honoraria: AstraZeneca, Support for current manuscript: German Research\nFoundation, Wilhelm Sander-Stiftung, Else Kr¨oner Fresenius Stiftung\nThe remaining authors have no competing interests to disclose.","formattedTitle":"Independent Benchmarking of Prompt-Based Medical Segmentation Models\n","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Segmentation, Deep Learning, Vision-Language Models, Foundation Models 1","lastPublishedDoi":"10.21203/rs.3.rs-7593773/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7593773/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Medical image segmentation rapidly shifts toward vision(-language) foundation models that unify diverse modalities and tasks within a single framework. In this work, we systematically benchmark high-impact vision-language and segment-anything-based architectures across multiple clinically relevant CT and MRI tasks. We show that while these models achieve strong performance, each comes with specific (dis)advantages. Non-3D models are highly flexible but require substantial user guidance and are prone to over- or under-detection. 3D architectures offer overall more reliable volumetric consistency, but can still have detection problems.\nVision-language models appear sensitive to the coverage of training data, whereas click-prompted SAM-based models are more universal, with a, though limited, ability to address zero-shot targets. When tested with more complex text prompts, most vision-language models exhibit missing semantic language understanding. Overall, these models hold considerable promise but still express limitations. Our work highlights key areas where future research is needed to advance vision(-language) foundation models. ","manuscriptTitle":"Independent Benchmarking of Prompt-Based Medical Segmentation Models","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-10-10 18:53:22","doi":"10.21203/rs.3.rs-7593773/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"be62bc7d-852a-47ea-9051-4b7ea8932603","owner":[],"postedDate":"October 10th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":55534133,"name":"Biological sciences/Computational biology and bioinformatics"},{"id":55534134,"name":"Health sciences/Health care"},{"id":55534135,"name":"Physical sciences/Mathematics and computing"},{"id":55534136,"name":"Health sciences/Medical research"}],"tags":[],"updatedAt":"2025-11-06T01:08:30+00:00","versionOfRecord":[],"versionCreatedAt":"2025-10-10 18:53:22","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7593773","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7593773","identity":"rs-7593773","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.