Adapting General Representations of Pretrained Vision Foundation Models to Seismic Understanding via Geologically Informed Prompting and Lightweight Tuning | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Adapting General Representations of Pretrained Vision Foundation Models to Seismic Understanding via Geologically Informed Prompting and Lightweight Tuning Xinmin Wu, Zhengfa Bi, Nori Nakata, Nuo Chen This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7444252/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted You are reading this latest preprint version Abstract Interpreting seismic responses to subsurface structures is essential for reservoir characterization and energy exploration, yet remains challenging due to the spatial complexity and heterogeneity of geological formations. While deep learning has accelerated progress, most existing models depend on task-specific architectures trained on limited labeled data, which hinders their generalizability. Here, we introduce a transfer learning framework that repurposes vision foundation model for seismic interpretation through cross-domain adaptation. A lightweight bridge model maps seismic amplitudes into the representation space of vision backbones, and efficient adaptation through low-rank updates and prefix tuning balances flexibility with preservation of the foundation model’s representational priors. By embedding geological priors into the latent space, the model learns to internalize stratigraphic order and to deliver predictions that preserve structural consistency. A task-adaptive decoder further supports various objectives in seismic interpretation. Benchmark experiments across various geological settings demonstrate that our framework consistently surpasses baseline architectures, delivering superior accuracy with fewer trainable parameters. These results underscore the promise of foundation models not only as scalable backbones for seismic interpretation, but also as a basis for broader data-driven advances across geoscientific disciplines under supervision-limited conditions. Earth and environmental sciences/Solid Earth sciences Earth and environmental sciences/Environmental sciences Seismic interpretation Vision foundation model Parameter-efficient transfer learning Cross-domain adaptation Geological priors Full Text Additional Declarations There is NO Competing Interest. Cite Share Download PDF Status: Under Review Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7444252","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":510809655,"identity":"67ce2fce-b7e4-4b01-b65b-ef0aa5b4a7c0","order_by":0,"name":"Xinmin Wu","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAwElEQVRIiWNgGAWjYDACZjBpA+WxEa8lDaqaKC0QcJgELQbHeY9u5m07nzh/fvMDhg9lhxn4Zzfg1yLZzJd2m7ftduKGY2wGjDPOHWaQuHMAvxZ+Zh6z27kgLWw8DMy8bYcZDCQS8Gthg2g5lzi/DajlLzFaoLYcSGw4BtTCSIwWyWaglj/nko03HEszONhzLp1H4gYBLQbnz5jdnFFmJzu/+fDDBz/KrOX4ZxDQggIOADEPCepHwSgYBaNgFOACAPAlPpxsibOYAAAAAElFTkSuQmCC","orcid":"","institution":"University of Science and Technology of China","correspondingAuthor":true,"prefix":"","firstName":"Xinmin","middleName":"","lastName":"Wu","suffix":""},{"id":510809656,"identity":"79ba32b5-1918-45c5-8b59-fb4c36324f91","order_by":1,"name":"Zhengfa Bi","email":"","orcid":"","institution":"Lawrence Berkeley National Laboratory","correspondingAuthor":false,"prefix":"","firstName":"Zhengfa","middleName":"","lastName":"Bi","suffix":""},{"id":510809657,"identity":"4093fc7c-eca9-4d53-9f51-3f2e5a92ca06","order_by":2,"name":"Nori Nakata","email":"","orcid":"","institution":"","correspondingAuthor":false,"prefix":"","firstName":"Nori","middleName":"","lastName":"Nakata","suffix":""},{"id":510809658,"identity":"9b76033d-e0bd-41e3-90c4-cd5e740ac95f","order_by":3,"name":"Nuo Chen","email":"","orcid":"","institution":"","correspondingAuthor":false,"prefix":"","firstName":"Nuo","middleName":"","lastName":"Chen","suffix":""}],"badges":[],"createdAt":"2025-08-24 05:35:30","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7444252/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7444252/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":91218664,"identity":"4a5f9512-06a1-4d74-b8b4-9b642c9e7df6","added_by":"auto","created_at":"2025-09-12 20:33:00","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":3635307,"visible":true,"origin":"","legend":"Article File","description":"","filename":"TransferringLearningfromPretrainingPriorsofVisionFoundationModelsforSeismicInterpretation.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7444252/v1_covered_506b261d-4e42-4272-b9a2-22b521a1e304.pdf"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"Adapting General Representations of Pretrained Vision Foundation Models to Seismic Understanding via Geologically Informed Prompting and Lightweight Tuning","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"nature-portfolio","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"","title":"Nature Portfolio","twitterHandle":"","acdcEnabled":false,"dfaEnabled":false,"editorialSystem":"ejp","reportingPortfolio":"","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Seismic interpretation, Vision foundation model, Parameter-efficient, transfer learning, Cross-domain, adaptation, Geological priors","lastPublishedDoi":"10.21203/rs.3.rs-7444252/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7444252/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Interpreting seismic responses to subsurface structures is essential for reservoir characterization and energy exploration, yet remains challenging due to the spatial complexity and heterogeneity of geological formations. While deep learning has accelerated progress, most existing models depend on task-specific architectures trained on limited labeled data, which hinders their generalizability. Here, we introduce a transfer learning framework that repurposes vision foundation model for seismic interpretation through cross-domain adaptation. A lightweight bridge model maps seismic amplitudes into the representation space of vision backbones, and efficient adaptation through low-rank updates and prefix tuning balances flexibility with preservation of the foundation model’s representational priors. By embedding geological priors into the latent space, the model learns to internalize stratigraphic order and to deliver predictions that preserve structural consistency. A task-adaptive decoder further supports various objectives in seismic interpretation. Benchmark experiments across various geological settings demonstrate that our framework consistently surpasses baseline architectures, delivering superior accuracy with fewer trainable parameters. These results underscore the promise of foundation models not only as scalable backbones for seismic interpretation, but also as a basis for broader data-driven advances across geoscientific disciplines under supervision-limited conditions.","manuscriptTitle":"Adapting General Representations of Pretrained Vision Foundation Models to Seismic Understanding via Geologically Informed Prompting and Lightweight Tuning","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-09-12 20:24:52","doi":"10.21203/rs.3.rs-7444252/v1","editorialEvents":[],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"communications-engineering","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"commseng","sideBox":"Learn more about [Communications Engineering](http://link.springer.com/journal/44172)","snPcode":"44172","submissionUrl":"https://mts-commseng.nature.com/cgi-bin/main.plex","title":"Communications Engineering","twitterHandle":"@commseng","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"68eec145-2a56-45aa-8208-4ff79afe0184","owner":[],"postedDate":"September 12th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":54262585,"name":"Earth and environmental sciences/Solid Earth sciences"},{"id":54262586,"name":"Earth and environmental sciences/Environmental sciences"}],"tags":[],"updatedAt":"2026-04-17T13:27:13+00:00","versionOfRecord":[],"versionCreatedAt":"2025-09-12 20:24:52","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7444252","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7444252","identity":"rs-7444252","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.