Evaluating Large Language Models in Scientific Discovery

preprint OA: closed CC-BY-4.0

Abstract

Abstract Large language models (LLMs) are increasingly applied to scientific research, yet prevailing science benchmarks probe decontextualized knowledge and overlook the iterative reasoning, hypothesis generation, and observation interpretation that drive scientific discovery. Here, we introduce a scenario-grounded benchmark that evaluates LLMs across biology, chemistry, material science, and physics. For this benchmark, domain experts define research projects of genuine value and interest, and decompose them into modular research scenarios from which vetted questions are sampled. The framework assesses models at two levels: (i) question-level accuracy on scenario-tied items and (ii) project-level performance, where models must propose testable hypotheses, execute simulations or experiments, and interpret results. Applying this two-phase scientific discovery evaluation (SDE) framework to state-of-the-art LLMs reveals a consistent performance gap relative to general science benchmarks, diminishing return of scaling up model sizes and reasoning, and systematic weaknesses shared across top-tier models from different providers. Large performance variation in research scenarios leads to changing choices of the best performing model on scientific discovery projects evaluated, suggesting all current LLMs are distant to general scientific “superintelligence”. Nevertheless, LLMs already demonstrate promise in a great variety of scientific discovery projects, including cases where constituent scenario scores are low, highlighting the role of guided exploration and serendipity in discovery. This SDE framework offers a reproducible benchmark for discovery- relevant evaluation of LLMs and charts practical paths to advance their development toward scientific innovation.
Full text 12,579 characters · extracted from preprint-html · click to expand
Evaluating Large Language Models in Scientific Discovery | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Resource Evaluating Large Language Models in Scientific Discovery Chenru Duan, Haojun Jia, Huan Sun, Seyed Mohamad Moosavi This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9045298/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted You are reading this latest preprint version Abstract Large language models (LLMs) are increasingly applied to scientific research, yet prevailing science benchmarks probe decontextualized knowledge and overlook the iterative reasoning, hypothesis generation, and observation interpretation that drive scientific discovery. Here, we introduce a scenario-grounded benchmark that evaluates LLMs across biology, chemistry, material science, and physics. For this benchmark, domain experts define research projects of genuine value and interest, and decompose them into modular research scenarios from which vetted questions are sampled. The framework assesses models at two levels: (i) question-level accuracy on scenario-tied items and (ii) project-level performance, where models must propose testable hypotheses, execute simulations or experiments, and interpret results. Applying this two-phase scientific discovery evaluation (SDE) framework to state-of-the-art LLMs reveals a consistent performance gap relative to general science benchmarks, diminishing return of scaling up model sizes and reasoning, and systematic weaknesses shared across top-tier models from different providers. Large performance variation in research scenarios leads to changing choices of the best performing model on scientific discovery projects evaluated, suggesting all current LLMs are distant to general scientific “superintelligence”. Nevertheless, LLMs already demonstrate promise in a great variety of scientific discovery projects, including cases where constituent scenario scores are low, highlighting the role of guided exploration and serendipity in discovery. This SDE framework offers a reproducible benchmark for discovery- relevant evaluation of LLMs and charts practical paths to advance their development toward scientific innovation. Physical sciences/Mathematics and computing/Computer science Physical sciences/Chemistry/Theoretical chemistry/Computational chemistry Physical sciences/Mathematics and computing/Scientific data Full Text Additional Declarations There is NO Competing Interest. Supplementary Files SDESI.pdf Supplementary Information PDF file data.zip Dataset 1 Cite Share Download PDF Status: Under Review Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9045298","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Resource","associatedPublications":[],"authors":[{"id":615505767,"identity":"ce22fea9-906b-4459-bf91-fbadc75c9ca4","order_by":0,"name":"Chenru Duan","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAoklEQVRIiWNgGAWjYHACAyC2YWADMXlI0JLGwMZGopbDDAxEa+GffXibxM8d56P55BsYH7xtI0KLxLm0MsneM7dz29gYmA3nEqPFgIfHTIK3DayFTZqXWC2Sf9vOgbSw/yZaC9DwA2BbmInSInGGrdhati0ZqCWxWXLOOSK08Pcwb7z5ts0ud37z4YMf3pQRoQUJMDaQpn4UjIJRMApGAW4AALuMLQ03XcnLAAAAAElFTkSuQmCC","orcid":"https://orcid.org/0000-0003-2592-4237","institution":"Deep Principle, Inc.","correspondingAuthor":true,"prefix":"","firstName":"Chenru","middleName":"","lastName":"Duan","suffix":""},{"id":615505768,"identity":"41e81d9e-8825-4d37-b89f-e56cf4c6363a","order_by":1,"name":"Haojun Jia","email":"","orcid":"","institution":"Deep Principle, Inc.","correspondingAuthor":false,"prefix":"","firstName":"Haojun","middleName":"","lastName":"Jia","suffix":""},{"id":615505769,"identity":"25461517-ff89-4b11-acad-a91007935756","order_by":2,"name":"Huan Sun","email":"","orcid":"","institution":"Ohio State University","correspondingAuthor":false,"prefix":"","firstName":"Huan","middleName":"","lastName":"Sun","suffix":""},{"id":615505770,"identity":"e3ce2e0e-22e7-4fc6-8a98-d601db01bee6","order_by":3,"name":"Seyed Mohamad Moosavi","email":"","orcid":"https://orcid.org/0000-0002-0357-5729","institution":"University of Toronto","correspondingAuthor":false,"prefix":"","firstName":"Seyed","middleName":"Mohamad","lastName":"Moosavi","suffix":""}],"badges":[],"createdAt":"2026-03-06 02:45:41","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9045298/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9045298/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":108182493,"identity":"9c54ced9-7888-4c8c-89ad-b2c1b0c2f844","added_by":"auto","created_at":"2026-04-30 08:59:24","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":5145073,"visible":true,"origin":"","legend":"Article File","description":"","filename":"SDEmain.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9045298/v1_covered_d6bc448a-2211-42ae-a792-489991151d6b.pdf"},{"id":108137151,"identity":"27dbc113-0278-43b2-b0db-cc12bcf4f1a4","added_by":"auto","created_at":"2026-04-29 17:54:35","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":5877709,"visible":true,"origin":"","legend":"Supplementary Information PDF file","description":"","filename":"SDESI.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9045298/v1/5e86fb2ca4d934191d72f7cc.pdf"},{"id":108137150,"identity":"553ee733-7172-418f-99a1-0c204cd69813","added_by":"auto","created_at":"2026-04-29 17:54:35","extension":"zip","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":9163196,"visible":true,"origin":"","legend":"Dataset 1","description":"","filename":"data.zip","url":"https://assets-eu.researchsquare.com/files/rs-9045298/v1/bba7017b0c20efa88f70d6b8.zip"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"Evaluating Large Language Models in Scientific Discovery","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"nature-portfolio","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"","title":"Nature Portfolio","twitterHandle":"","acdcEnabled":false,"dfaEnabled":false,"editorialSystem":"ejp","reportingPortfolio":"","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-9045298/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9045298/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Large language models (LLMs) are increasingly applied to scientific research, yet prevailing science benchmarks probe decontextualized knowledge and overlook the iterative reasoning, hypothesis generation, and observation interpretation that drive scientific discovery. Here, we introduce a scenario-grounded benchmark that evaluates LLMs across biology, chemistry, material science, and physics. For this benchmark, domain experts define research projects of genuine value and interest, and decompose them into modular research scenarios from which vetted questions are sampled. The framework assesses models at two levels: (i) question-level accuracy on scenario-tied items and (ii) project-level performance, where models must propose testable hypotheses, execute simulations or experiments, and interpret results. Applying this two-phase scientific discovery evaluation (SDE) framework to state-of-the-art LLMs reveals a consistent performance gap relative to general science benchmarks, diminishing return of scaling up model sizes and reasoning, and systematic weaknesses shared across top-tier models from different providers. Large performance variation in research scenarios leads to changing choices of the best performing model on scientific discovery projects evaluated, suggesting all current LLMs are distant to general scientific “superintelligence”. Nevertheless, LLMs already demonstrate promise in a great variety of scientific discovery projects, including cases where constituent scenario scores are low, highlighting the role of guided exploration and serendipity in discovery. This SDE framework offers a reproducible benchmark for discovery- relevant evaluation of LLMs and charts practical paths to advance their development toward scientific innovation.","manuscriptTitle":"Evaluating Large Language Models in Scientific Discovery","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-04-29 17:54:27","doi":"10.21203/rs.3.rs-9045298/v1","editorialEvents":[],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"nature-computational-science","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"natcomputsci","sideBox":"Learn more about [Nature Computational Science](http://www.nature.com/natcomputsci/)","snPcode":"","submissionUrl":"","title":"Nature Computational Science","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Nature Research","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"d23204e5-b729-49ee-9d27-5d3811c0206f","owner":[],"postedDate":"April 29th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":65494254,"name":"Physical sciences/Mathematics and computing/Computer science"},{"id":65494255,"name":"Physical sciences/Chemistry/Theoretical chemistry/Computational chemistry"},{"id":65494256,"name":"Physical sciences/Mathematics and computing/Scientific data"}],"tags":[],"updatedAt":"2026-04-29T17:54:27+00:00","versionOfRecord":[],"versionCreatedAt":"2026-04-29 17:54:27","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9045298","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9045298","identity":"rs-9045298","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-06-05T02:00:03.366016+00:00
License: CC-BY-4.0