Comprehensive prediction and analysis of human protein essentiality based on a pre-trained protein large language model

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Human essential genes and their protein products are indispensable for the viability and development of the individuals. Thus, it is quite important to decipher the essential proteins and up to now numerous computational methods have been developed for the above purpose. However, the current methods failed to comprehensively measure human protein essentiality at levels of humans, human cell lines, and mice orthologues. For doing so, here we developed Protein Importance Calculator (PIC), a sequence-based deep learning model, which was built by fine-tuning a pre-trained protein language model. As a result, PIC outperformed existing methods by increasing 5.13%-12.10% AUROC for predicting essential proteins at human cell-line level. In addition, it improved an average of 9.64% AUROC on 323 human cell lines compared to the only existing cell line-specific method, DeepCellEss. Moreover, we defined Protein Essential Score (PES) to quantify protein essentiality based on PIC and confirmed its power of measuring human protein essentiality and functional divergence across the above three levels. Finally, we successfully used PES to identify prognostic biomarkers of breast cancer and at the first time to quantify the essentiality of 617462 human microproteins.
Full text 17,556 characters · extracted from preprint-html · click to expand
Comprehensive prediction and analysis of human protein essentiality based on a pre-trained protein large language model | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Help Center Sign In Submit a Preprint Cite Share Download PDF Article Comprehensive prediction and analysis of human protein essentiality based on a pre-trained protein large language model Boming Kang, Rui Fan, Chunmei Cui, Qinghua Cui This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4246084/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 27 Nov, 2024 Read the published version in Nature Computational Science → Version 1 posted You are reading this latest preprint version Abstract Human essential genes and their protein products are indispensable for the viability and development of the individuals. Thus, it is quite important to decipher the essential proteins and up to now numerous computational methods have been developed for the above purpose. However, the current methods failed to comprehensively measure human protein essentiality at levels of humans, human cell lines, and mice orthologues. For doing so, here we developed Protein Importance Calculator (PIC), a sequence-based deep learning model, which was built by fine-tuning a pre-trained protein language model. As a result, PIC outperformed existing methods by increasing 5.13%-12.10% AUROC for predicting essential proteins at human cell-line level. In addition, it improved an average of 9.64% AUROC on 323 human cell lines compared to the only existing cell line-specific method, DeepCellEss. Moreover, we defined Protein Essential Score (PES) to quantify protein essentiality based on PIC and confirmed its power of measuring human protein essentiality and functional divergence across the above three levels. Finally, we successfully used PES to identify prognostic biomarkers of breast cancer and at the first time to quantify the essentiality of 617462 human microproteins. Biological sciences/Computational biology and bioinformatics/Machine learning Biological sciences/Computational biology and bioinformatics/Computational models Human protein essentiality protein language model deep learning Full Text Additional Declarations There is NO Competing Interest. Supplementary Files FigS1.pdf FigS2.pdf FigS3.pdf TableS1.csv TableS2.xls TableS3.xls TableS4.csv TableS5.xlsx TableS6.xlsx TableS7.csv TableS8.xlsx TableS9.xlsx TableS10.xlsx Cite Share Download PDF Status: Published Journal Publication published 27 Nov, 2024 Read the published version in Nature Computational Science → Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4246084","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":289807656,"identity":"891c3e19-fdf0-45c7-9638-b8ed2e9ab640","order_by":0,"name":"Boming Kang","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA1klEQVRIiWNgGAWjYBACCSBmZmCwSWBsAHHZiNeSRrqWwwkQLjFaJNvPHn5dmHM+j3naGQOGD2WHGfhnN+DXIs2Tl2Y9c9vtYsbZOQaMM84dZpC4cwC/FjmGHDNj3m23ExuBWph52w4zGEgkENDC/wak5RxEy19itEhL5Bg/5t12AKKFkRgtkjPemDHP3JYM9EtawcGec+k8EjcIaJE4n2P8uXCbXZ7h7OSND36UWcvxzyCgBQjYQHHDYNjAwHAASPMQVA8EzB9ApDwxSkfBKBgFo2BkAgAfPkLRkQwVhgAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0009-0008-8704-3949","institution":"Department of Biomedical Informatics, State Key Laboratory of Vascular Homeostasis and Remodeling, School of Basic Medical Sciences, Peking University","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Boming","middleName":"","lastName":"Kang","suffix":""},{"id":289807657,"identity":"80933bd2-543f-4b27-9e4a-32a12c1c4e94","order_by":1,"name":"Rui Fan","email":"","orcid":"","institution":"Department of Biomedical Informatics, State Key Laboratory of Vascular Homeostasis and Remodeling, School of Basic Medical Sciences, Peking University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Rui","middleName":"","lastName":"Fan","suffix":""},{"id":289807658,"identity":"d6596eef-e5c6-48cf-a9c8-01a690866196","order_by":2,"name":"Chunmei Cui","email":"","orcid":"","institution":"Department of Biomedical Informatics, State Key Laboratory of Vascular Homeostasis and Remodeling, School of Basic Medical Sciences, Peking University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Chunmei","middleName":"","lastName":"Cui","suffix":""},{"id":289807659,"identity":"20350d89-bcf3-4ad5-bd38-8c4252e14c19","order_by":3,"name":"Qinghua Cui","email":"","orcid":"","institution":"Department of Biomedical Informatics, State Key Laboratory of Vascular Homeostasis and Remodeling, School of Basic Medical Sciences, Peking University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Qinghua","middleName":"","lastName":"Cui","suffix":""}],"badges":[],"createdAt":"2024-04-10 08:21:05","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4246084/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4246084/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s43588-024-00733-1","type":"published","date":"2024-11-27T05:00:00+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":70085156,"identity":"b9f81324-0832-42ac-a4dc-dd77304787cb","added_by":"auto","created_at":"2024-11-28 08:06:52","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":9934684,"visible":true,"origin":"","legend":"","description":"","filename":"Fullissue.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4246084/v1_covered_dd0edfb1-d07e-4857-a9a6-2c58f75f71ff.pdf"},{"id":55064453,"identity":"98b222d1-fdc5-4e2b-a613-29c8e30dfc8b","added_by":"auto","created_at":"2024-04-22 03:36:00","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":2838453,"visible":true,"origin":"","legend":"","description":"","filename":"FigS1.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4246084/v1/773891860a75f6a00809dfad.pdf"},{"id":55064661,"identity":"fc3ed3af-1655-4e80-8d9a-1f8ea06875b5","added_by":"auto","created_at":"2024-04-22 03:44:00","extension":"pdf","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":23086,"visible":true,"origin":"","legend":"","description":"","filename":"FigS2.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4246084/v1/433d4bda1f38fa975a5582cb.pdf"},{"id":55064452,"identity":"046467a8-27c7-4721-a59b-ba1728a6397b","added_by":"auto","created_at":"2024-04-22 03:36:00","extension":"pdf","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":25167,"visible":true,"origin":"","legend":"","description":"","filename":"FigS3.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4246084/v1/56f1e8f8f7496379aed595fd.pdf"},{"id":55064451,"identity":"1e7ff4a3-8ec3-4636-aee0-0eaa76dbb7a9","added_by":"auto","created_at":"2024-04-22 03:36:00","extension":"csv","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":57551,"visible":true,"origin":"","legend":"","description":"","filename":"TableS1.csv","url":"https://assets-eu.researchsquare.com/files/rs-4246084/v1/8fc5923c784acaa68588181e.csv"},{"id":55064458,"identity":"fe182c56-f7ab-4a6c-8ce9-5595e8dfc9cd","added_by":"auto","created_at":"2024-04-22 03:36:01","extension":"xls","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":22528,"visible":true,"origin":"","legend":"","description":"","filename":"TableS2.xls","url":"https://assets-eu.researchsquare.com/files/rs-4246084/v1/3ad4be2b4f1af163bdbdab55.xls"},{"id":55064457,"identity":"95116212-736a-4193-aa82-81f9ce36db13","added_by":"auto","created_at":"2024-04-22 03:36:01","extension":"xls","order_by":6,"title":"","display":"","copyAsset":false,"role":"supplement","size":20992,"visible":true,"origin":"","legend":"","description":"","filename":"TableS3.xls","url":"https://assets-eu.researchsquare.com/files/rs-4246084/v1/ba206d709ee5b10f0e2961d5.xls"},{"id":55064459,"identity":"4e47cb58-f2ec-4be1-815b-0ccdcdf1e90c","added_by":"auto","created_at":"2024-04-22 03:36:01","extension":"csv","order_by":7,"title":"","display":"","copyAsset":false,"role":"supplement","size":18433,"visible":true,"origin":"","legend":"","description":"","filename":"TableS4.csv","url":"https://assets-eu.researchsquare.com/files/rs-4246084/v1/d21902417e62e89f6da4cc9c.csv"},{"id":55064882,"identity":"70f76d5c-a6d3-4e4f-ba67-ed8118eba345","added_by":"auto","created_at":"2024-04-22 03:52:01","extension":"xlsx","order_by":8,"title":"","display":"","copyAsset":false,"role":"supplement","size":18016,"visible":true,"origin":"","legend":"","description":"","filename":"TableS5.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-4246084/v1/339b0dcb1ab673f1dc902bca.xlsx"},{"id":55064449,"identity":"b3dfed54-8ca5-4f3d-ae03-6907b57560a3","added_by":"auto","created_at":"2024-04-22 03:35:59","extension":"xlsx","order_by":9,"title":"","display":"","copyAsset":false,"role":"supplement","size":18076,"visible":true,"origin":"","legend":"","description":"","filename":"TableS6.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-4246084/v1/e29f6db664f1a23035eb5245.xlsx"},{"id":55064663,"identity":"55ef630f-5147-42f6-90d8-eb823cc1acfb","added_by":"auto","created_at":"2024-04-22 03:44:01","extension":"csv","order_by":10,"title":"","display":"","copyAsset":false,"role":"supplement","size":34576,"visible":true,"origin":"","legend":"","description":"","filename":"TableS7.csv","url":"https://assets-eu.researchsquare.com/files/rs-4246084/v1/e84a63f2defba4ba540e4191.csv"},{"id":55064460,"identity":"9ccb2338-3d6d-4268-bad8-25b20dd43b2b","added_by":"auto","created_at":"2024-04-22 03:36:01","extension":"xlsx","order_by":11,"title":"","display":"","copyAsset":false,"role":"supplement","size":16911,"visible":true,"origin":"","legend":"","description":"","filename":"TableS8.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-4246084/v1/270a142764dd1de2dd30862a.xlsx"},{"id":55064454,"identity":"c3d7274d-b623-4b17-87f6-355b94fe2d28","added_by":"auto","created_at":"2024-04-22 03:36:01","extension":"xlsx","order_by":12,"title":"","display":"","copyAsset":false,"role":"supplement","size":17086,"visible":true,"origin":"","legend":"","description":"","filename":"TableS9.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-4246084/v1/ad3212414e1df8434a12bdc1.xlsx"},{"id":55064450,"identity":"272f182b-25be-49f5-a15d-28ba94f3f188","added_by":"auto","created_at":"2024-04-22 03:36:00","extension":"xlsx","order_by":13,"title":"","display":"","copyAsset":false,"role":"supplement","size":2770706,"visible":true,"origin":"","legend":"","description":"","filename":"TableS10.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-4246084/v1/431f23b9acad662399f43654.xlsx"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"Comprehensive prediction and analysis of human protein essentiality based on a pre-trained protein large language model","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"nature-portfolio","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"","title":"Nature Portfolio","twitterHandle":"","acdcEnabled":false,"dfaEnabled":false,"editorialSystem":"ejp","reportingPortfolio":"","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Human protein essentiality, protein language model, deep learning","lastPublishedDoi":"10.21203/rs.3.rs-4246084/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4246084/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Human essential genes and their protein products are indispensable for the viability and development of the individuals. Thus, it is quite important to decipher the essential proteins and up to now numerous computational methods have been developed for the above purpose. However, the current methods failed to comprehensively measure human protein essentiality at levels of humans, human cell lines, and mice orthologues. For doing so, here we developed Protein Importance Calculator (PIC), a sequence-based deep learning model, which was built by fine-tuning a pre-trained protein language model. As a result, PIC outperformed existing methods by increasing 5.13%-12.10% AUROC for predicting essential proteins at human cell-line level. In addition, it improved an average of 9.64% AUROC on 323 human cell lines compared to the only existing cell line-specific method, DeepCellEss. Moreover, we defined Protein Essential Score (PES) to quantify protein essentiality based on PIC and confirmed its power of measuring human protein essentiality and functional divergence across the above three levels. Finally, we successfully used PES to identify prognostic biomarkers of breast cancer and at the first time to quantify the essentiality of 617462 human microproteins.","manuscriptTitle":"Comprehensive prediction and analysis of human protein essentiality based on a pre-trained protein large language model","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-04-22 03:35:54","doi":"10.21203/rs.3.rs-4246084/v1","editorialEvents":[],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"nature-computational-science","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"natcomputsci","sideBox":"Learn more about [Nature Computational Science](http://www.nature.com/natcomputsci/)","snPcode":"","submissionUrl":"","title":"Nature Computational Science","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Nature Research","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"f04ae459-aa7d-4b81-94be-ca6c864ad73a","owner":[],"postedDate":"April 22nd, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":30520255,"name":"Biological sciences/Computational biology and bioinformatics/Machine learning"},{"id":30520256,"name":"Biological sciences/Computational biology and bioinformatics/Computational models"}],"tags":[],"updatedAt":"2024-11-28T08:06:43+00:00","versionOfRecord":{"articleIdentity":"rs-4246084","link":"https://doi.org/10.1038/s43588-024-00733-1","journal":{"identity":"nature-computational-science","isVorOnly":false,"title":"Nature Computational Science"},"publishedOn":"2024-11-27 05:00:00","publishedOnDateReadable":"November 27th, 2024"},"versionCreatedAt":"2024-04-22 03:35:54","video":"","vorDoi":"10.1038/s43588-024-00733-1","vorDoiUrl":"https://doi.org/10.1038/s43588-024-00733-1","workflowStages":[]},"version":"v1","identity":"rs-4246084","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4246084","identity":"rs-4246084","version":["v1"]},"buildId":"omnImTCwR2MFx8CMYfrG7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

⚙ Ask this paper AI returns verbatim quotes from the full text · source: preprint-html ⓘ

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00