Automatic extraction and structuring of cultural heritage analysis process documentation from audio and text files | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Automatic extraction and structuring of cultural heritage analysis process documentation from audio and text files Violette Abergel, Van Tuan Bui, Amine Berbagui This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8394677/v1 This work is licensed under a CC BY 4.0 License Status: Under Revision Version 1 posted 13 You are reading this latest preprint version Abstract Heritage science (HS) is an interdisciplinary field where collective knowledge emerges through an ongoing interplay between material objects and a wide range of research approaches that encompasses both the Humanities and experimental sciences. In this domain, the data management challenges are compounded by the strong heterogeneity of documentary sources, analytical data, and processes mobilized for condition reporting, analysis, monitoring or conservation purposes. Provenance metadata and paradata are essential for ensuring data reliability. Such documentation provides invaluable information on acquisition contexts and subsequent reuse possibilities. However, producing it rigorously is time-consuming, as the required information is diverse, context-dependent, and increasingly difficult to recover as time passes. In light of the massive daily data production in this field, developing methods to streamline data enrichment procedures is a clear priority. To address the risk of losing large amounts of undocumented data, the METAREVE project proposes a lightweight solution to help HS communities extract the key descriptive elements needed for minimal data understanding. Based on Automatic Speech Recognition (ASR) and Natural Language Understanding (NLU), it takes the form of a web application that automatically documents scientific activities related to cultural heritage, drawing from common outputs such as expert reports, research articles, or even audio recordings of in situ acquisition processes. This approach has been implemented within the digital ecosystem developed by the EquipEx + ESPADON project. Heritage sciences Knowledge extraction Data provenance Indexation NLP Speech recognition Documentation Knowledge graph Full Text Additional Declarations Competing interest reported. Funding : This work was supported by the Paris Seine Graduate School Humanities, Creation, Heritage, Investissement d'Avenir ANR-17-EURE-0021 – Fondation for cultural heritage sciences This work benefited from State aid managed by the Agence Nationale de la Recherche (French National Research Agency) under the Investment for the Future Program integrated into France 2030, bearing the reference ANR 21-ESRE-0050 EquipEx+ ESPADON. Cite Share Download PDF Status: Under Revision Version 1 posted Editorial decision: Revision requested 12 May, 2026 Reviews received at journal 18 Apr, 2026 Reviewers agreed at journal 08 Apr, 2026 Reviewers agreed at journal 07 Apr, 2026 Reviews received at journal 24 Mar, 2026 Reviewers agreed at journal 20 Mar, 2026 Reviews received at journal 21 Feb, 2026 Reviewers agreed at journal 15 Feb, 2026 Reviewers agreed at journal 11 Feb, 2026 Reviewers invited by journal 12 Jan, 2026 Editor assigned by journal 29 Dec, 2025 Submission checks completed at journal 29 Dec, 2025 First submitted to journal 18 Dec, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8394677","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":562339261,"identity":"25eb8619-7a87-4ff7-8035-bc0b3c25232b","order_by":0,"name":"Violette Abergel","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABJElEQVRIie3PMUvDQBTA8RcO4nJwa4KSfoVXAilC8bOcBNIl6qhDwQuFugTnQr+Ek/OFgzp1dalDQOgk0i6SSkCvqTUIia4O9x/uDo4f9w7AZPqPEbDk7kSr9YjVd1beQkACx29CXVHfYdtDPwnKPwg+ECnXRenBwTx7vnrvU/9RLXMYPl10BLFXTUTZPJtw9IGeh935bUSDRdRDmC2P7yQhkwbijigqyvFUQBy4Sao04YFjCYUITDUNVpGS47VgL71Nkn5Qfzp4q4gejDQRRjTR3+fgxIElCknxMN69ArKN6L+kkd8dO6+hm4iQOov4EvlMD6aaic1Ulhd9r8PYWbYW5YnHpoP7fDXUg92MGkltt4s1/jrz7fY72FfW3GQymUz7PgGYKF2QZabEpQAAAABJRU5ErkJggg==","orcid":"","institution":"French National Centre for Scientific Research","correspondingAuthor":true,"prefix":"","firstName":"Violette","middleName":"","lastName":"Abergel","suffix":""},{"id":562339262,"identity":"f4c685fb-ac96-42a9-95f2-e17d4a6117eb","order_by":1,"name":"Van Tuan Bui","email":"","orcid":"","institution":"Fondation des Sciences du Patrimoine","correspondingAuthor":false,"prefix":"","firstName":"Van","middleName":"Tuan","lastName":"Bui","suffix":""},{"id":562339263,"identity":"edae93f4-a09c-4e49-a6c8-3e9c353fb477","order_by":2,"name":"Amine Berbagui","email":"","orcid":"","institution":"Fondation des Sciences du Patrimoine","correspondingAuthor":false,"prefix":"","firstName":"Amine","middleName":"","lastName":"Berbagui","suffix":""}],"badges":[],"createdAt":"2025-12-18 11:08:28","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8394677/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8394677/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":103611752,"identity":"160bcbb2-0b03-4aa2-b35d-d63fb89f7a85","added_by":"auto","created_at":"2026-02-27 15:56:42","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1504431,"visible":true,"origin":"","legend":"","description":"","filename":"Metarevepaper231225.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8394677/v1_covered_603a9127-c5b9-4605-b9e6-f5eaeb73939c.pdf"}],"financialInterests":"Competing interest reported. Funding :\nThis work was supported by the Paris Seine Graduate School Humanities, Creation, Heritage, Investissement d'Avenir ANR-17-EURE-0021 – Fondation for cultural heritage sciences\n\nThis work benefited from State aid managed by the Agence Nationale de la Recherche (French National\nResearch Agency) under the Investment for the Future Program integrated into France 2030, bearing the reference ANR 21-ESRE-0050 EquipEx+ ESPADON.","formattedTitle":"Automatic extraction and structuring of cultural heritage analysis process documentation from audio and text files","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"discover-applied-sciences","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"Learn more about [Discover Applied Sciences](https://link.springer.com/journal/42452)","snPcode":"42452","submissionUrl":"https://submission.springernature.com/new-submission/42452/3","title":"Discover Applied Sciences","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Discover Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Heritage sciences, Knowledge extraction, Data provenance, Indexation, NLP, Speech recognition, Documentation, Knowledge graph","lastPublishedDoi":"10.21203/rs.3.rs-8394677/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8394677/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eHeritage science (HS) is an interdisciplinary field where collective knowledge emerges through an ongoing interplay between material objects and a wide range of research approaches that encompasses both the Humanities and experimental sciences. In this domain, the data management challenges are compounded by the strong heterogeneity of documentary sources, analytical data, and processes mobilized for condition reporting, analysis, monitoring or conservation purposes. Provenance metadata and paradata are essential for ensuring data reliability. Such documentation provides invaluable information on acquisition contexts and subsequent reuse possibilities. However, producing it rigorously is time-consuming, as the required information is diverse, context-dependent, and increasingly difficult to recover as time passes. In light of the massive daily data production in this field, developing methods to streamline data enrichment procedures is a clear priority. To address the risk of losing large amounts of undocumented data, the METAREVE project proposes a lightweight solution to help HS communities extract the key descriptive elements needed for minimal data understanding. Based on Automatic Speech Recognition (ASR) and Natural Language Understanding (NLU), it takes the form of a web application that automatically documents scientific activities related to cultural heritage, drawing from common outputs such as expert reports, research articles, or even audio recordings of in situ acquisition processes. This approach has been implemented within the digital ecosystem developed by the EquipEx\u0026thinsp;+\u0026thinsp;ESPADON project.\u003c/p\u003e","manuscriptTitle":"Automatic extraction and structuring of cultural heritage analysis process documentation from audio and text files","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-02-27 15:55:37","doi":"10.21203/rs.3.rs-8394677/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-05-12T06:31:57+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-18T09:00:29+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"135561466584701014285758042451382508217","date":"2026-04-08T14:54:31+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"296022072167787643031561677396635233115","date":"2026-04-07T10:09:49+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-03-24T21:41:48+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"311092012508161646416009580371663982880","date":"2026-03-20T09:50:58+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-02-21T11:19:52+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"173547769269213521176882476290450174491","date":"2026-02-15T11:58:10+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"136237111762830141247808283510104282998","date":"2026-02-11T09:46:57+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-01-12T13:37:13+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-12-29T09:41:51+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-12-29T09:40:08+00:00","index":"","fulltext":""},{"type":"submitted","content":"Discover Applied Sciences","date":"2025-12-18T10:46:33+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"discover-applied-sciences","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"Learn more about [Discover Applied Sciences](https://link.springer.com/journal/42452)","snPcode":"42452","submissionUrl":"https://submission.springernature.com/new-submission/42452/3","title":"Discover Applied Sciences","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Discover Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"121e17d4-1678-4192-a14e-9749a4c45f8c","owner":[],"postedDate":"February 27th, 2026","published":true,"recentEditorialEvents":[{"type":"decision","content":"Revision requested","date":"2026-05-12T06:31:57+00:00","index":"","fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"in-revision","subjectAreas":[],"tags":[],"updatedAt":"2026-05-12T06:41:44+00:00","versionOfRecord":[],"versionCreatedAt":"2026-02-27 15:55:37","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8394677","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8394677","identity":"rs-8394677","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.