A literature review of methods for assessment of reproducibility in science | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Systematic Review A literature review of methods for assessment of reproducibility in science Torbjörn Nordling, Tomas Melo Peralta This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-2267847/v5 This work is licensed under a CC BY 4.0 License Status: Posted Version 5 posted You are reading this latest preprint version Show more versions Abstract Introduction: In response to the US Congress petition, the National Academies of Sciences, Engineering, and Medicine investigated the status of reproducibility and replicability in science. A piece of work is reproducible if the same results can be obtained while following the methods under the same conditions and using the same data. Unavailable data, missing code, and unclear or incomplete method descriptions are common reasons for failure to reproduce results. Objectives: The motivation behind this review is to investigate the current methods for reproducibility assessment and analyze their strengths and weaknesses so that we can determine where there is room for improvement. Methods: We followed the PRISMA 2020 standard and conducted a literature review to find the current methods to assess the reproducibility of scientific articles. We made use of three databases for our search: Web of Science, Scopus, and Engineering Village. Our criteria to find relevant articles was to look for methods, algorithms, or techniques to evaluate, assess, or predict reproducibility in science. We discarded methods that were specific to a single study, or that could not be adapted to scientific articles in general. Results: We found ten articles describing methods to evaluate reproducibility, and classified them as either a prediction market, a survey, a machine learning algorithm, or a numerical method. A prediction market requires participants to bet on the reproducibility of a study. The surveys are simple and straightforward, but their performance has not been assessed rigorously. Two types of machine learning methods have been applied: handpicked features and natural language processing. Conclusion: While the machine learning methods are promising because they can be scaled to reduce time and cost for researchers, none of the models reviewed achieved an accuracy above 75%. Given the prominence of transformer models for state-of-the-art natural language processing (NLP) tasks, we believe a transformer model can achieve better accuracy. reproducibility replicability AI reproducibility assessment Full Text Cite Share Download PDF Status: Posted Version 5 posted You are reading this latest preprint version Show more versions Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-2267847","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Systematic Review","associatedPublications":[],"authors":[{"id":163274018,"identity":"d2e51302-1abf-4727-9acc-74af75383651","order_by":0,"name":"Torbjörn Nordling","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAy0lEQVRIiWNgGAWjYDAC9gYgwWbBw8DA3MBMlA4engMgLRJALYzEapFIAGthIF6LveTbg48ryiRkzPkXNn4uYLDJl3cgZIt0XrLhmXMSPJYzHjZLz2BIs9x4gKCWHDPJxjYJHoMbBxukeRgOGxg2ENIieQaupfk3cVokeKBazje2gW2RJ6CDgedMjrFhwzmQLYxt1jMM0gwMCGlhbz9j+LChzMbe4Pzhw7cLKmwM5Ak5DAHAEQS0wuAA0Vr4oUpJsGUUjIJRMApGCAAA5kw410p8MksAAAAASUVORK5CYII=","orcid":"","institution":"National Cheng Kung University","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Torbjörn","middleName":"","lastName":"Nordling","suffix":""},{"id":163274019,"identity":"c9e3b961-f370-4c7e-ae6c-6b88d51add88","order_by":1,"name":"Tomas Melo Peralta","email":"","orcid":"","institution":"National Cheng Kung University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Tomas","middleName":"Melo","lastName":"Peralta","suffix":""}],"badges":[],"createdAt":"2022-11-13 07:54:49","currentVersionCode":5,"declarations":{"humanSubjects":false,"vertebrateSubjects":false,"conflictsOfInterestStatement":true,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false,"coiExplicitlySet":false},"doi":"10.21203/rs.3.rs-2267847/v5","doiUrl":"https://doi.org/10.21203/rs.3.rs-2267847/v5","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":30899841,"identity":"0e9b8c68-33f2-45ea-9b80-60589cffaeb4","added_by":"auto","created_at":"2022-12-29 18:45:01","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":284413,"visible":true,"origin":"","legend":"","description":"","filename":"mainresearchsquareAPA.pdf","url":"https://assets-eu.researchsquare.com/files/rs-2267847/v5/7c7b6306aba08ef57d13f958.pdf"}],"financialInterests":"","formattedTitle":"\u003cp\u003eA literature review of methods for assessment of reproducibility in science\u003c/p\u003e","fulltext":[],"fulltextSource":"","fullText":"","funders":[{"identity":"5ce4ba09-2676-4be8-a40c-bf3ad861bceb","identifier":"10.13039/501100004663","name":"Ministry of Science and Technology, Taiwan","awardNumber":"110-2222-E-006-010","order_by":0},{"identity":"9461677d-d01e-42a6-85ba-5bf8e978e6a3","identifier":"10.13039/501100004663","name":"Ministry of Science and Technology, Taiwan","awardNumber":"111-2221-E-006-186","order_by":1}],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"National Cheng Kung University","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"reproducibility, replicability, AI reproducibility assessment","lastPublishedDoi":"10.21203/rs.3.rs-2267847/v5","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-2267847/v5","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eIntroduction: In response to the US Congress petition, the National Academies of Sciences, Engineering, and Medicine investigated the status of reproducibility and replicability in science. A piece of work is reproducible if the same results can be obtained while following the methods under the same conditions and using the same data. Unavailable data, missing code, and unclear or incomplete method descriptions are common reasons for failure to reproduce results. Objectives: The motivation behind this review is to investigate the current methods for reproducibility assessment and analyze their strengths and weaknesses so that we can determine where there is room for improvement. Methods: We followed the PRISMA 2020 standard and conducted a literature review to find the current methods to assess the reproducibility of scientific articles. We made use of three databases for our search: Web of Science, Scopus, and Engineering Village. Our criteria to find relevant articles was to look for methods, algorithms, or techniques to evaluate, assess, or predict reproducibility in science. We discarded methods that were specific to a single study, or that could not be adapted to scientific articles in general.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eResults: We found ten articles describing methods to evaluate reproducibility, and classified them as either a prediction market, a survey, a machine learning algorithm, or a numerical method. A prediction market requires participants to bet on the reproducibility of a study. The surveys are simple and straightforward, but their performance has not been assessed rigorously. Two types of machine learning methods have been applied: handpicked features and natural language processing.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eConclusion: While the machine learning methods are promising because they can be scaled to reduce time and cost for researchers, none of the models reviewed achieved an accuracy above 75%. Given the prominence of transformer models for state-of-the-art natural language processing (NLP) tasks, we believe a transformer model can achieve better accuracy.\u003c/p\u003e","manuscriptTitle":"A literature review of methods for assessment of reproducibility in science","msid":"","msnumber":"","nonDraftVersions":[{"code":5,"date":"2022-12-29 18:44:56","doi":"10.21203/rs.3.rs-2267847/v5","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}},{"code":4,"date":"2022-12-28 01:56:30","doi":"10.21203/rs.3.rs-2267847/v4","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}},{"code":3,"date":"2022-11-30 15:58:04","doi":"10.21203/rs.3.rs-2267847/v3","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}},{"code":2,"date":"2022-11-28 16:27:21","doi":"10.21203/rs.3.rs-2267847/v2","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}},{"code":1,"date":"2022-11-14 16:52:46","doi":"10.21203/rs.3.rs-2267847/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"0833a0e1-b08a-40c5-a8a1-9cdeed0c0ba5","owner":[],"postedDate":"December 29th, 2022","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2022-11-14T16:52:46+00:00","versionOfRecord":[],"versionCreatedAt":"2022-12-29 18:44:56","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v5","identity":"rs-2267847","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-2267847","identity":"rs-2267847","version":["v5"]},"buildId":"7rjqhiLT3MXkJMwkYKINL","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.