The 2024 Report on the Human Proteome from the HUPO Human Proteome Project.

OA: closed CC-BY-NC-4.0
⚙ AI-generated summary by qwen3.7-flash, 2026-09-02 ⓘ

The 2024 Human Proteome Project report details the transition to UniProtKB and GENCODE, achieving protein expression detection for 93% of genes while reducing missing proteins to 1273.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

⚙ AI-generated deep summary by qwen3.7-flash, 2026-08-22 · read from full text ⓘ

The 2024 report from the Human Proteome Project outlines significant updates to its methodology, primarily transitioning the protein target list from neXtProt to a GENCODE-based system while maintaining UniProtKB as the central knowledge base for protein annotations. This shift resulted in the removal of approximately 978 entries deemed historically inaccurate or non-coding, such as immunoglobulin variable regions and endogenous retrovirus sequences, thereby refining the inventory of credible human proteins. The project reports that 93% of the current target list has been identified with high-confidence mass spectrometry evidence, leaving only 1,273 "missing proteins" yet to be credibly detected. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

The Human Proteome Project (HPP), the flagship initiative of the Human Proteome Organization (HUPO), has pursued two goals: (1) to credibly identify at least one isoform of every protein-coding gene and (2) to make proteomics an integral part of multiomics studies of human health and disease. The past year has seen major transitions for the HPP. neXtProt was retired as the official HPP knowledge base, UniProtKB became the reference proteome knowledge base, and Ensembl-GENCODE provides the reference protein target list. A function evidence FE1-5 scoring system has been developed for functional annotation of proteins, parallel to the PE1-5 UniProtKB/neXtProt scheme for evidence of protein expression. This report includes updates from neXtProt (version 2023-09) and UniProtKB release 2024_04, with protein expression detected (PE1) for 18138 of the 19411 GENCODE protein-coding genes (93%). The number of non-PE1 proteins ("missing proteins") is now 1273. The transition to GENCODE is a net reduction of 367 proteins (19,411 PE1-5 instead of 19,778 PE1-4 last year in neXtProt). We include reports from the Biology and Disease-driven HPP, the Human Protein Atlas, and the HPP Grand Challenge Project. We expect the new Functional Evidence FE1-5 scheme to energize the Grand Challenge Project for functional annotation of human proteins throughout the global proteomics community, including π-HuB in China.
Full text 48,091 characters · extracted from pmc-nxml · 10 sections · click to expand

The

The past year has seen several major developments in how the metrics of the HPP are computed and presented, including major changes in the protein target list, the reference genome, and the protein function score. The neXtProt knowledge base, started in 2011 by Amos Bairoch and Lydie Lane as an offshoot of UniProtKB/Swiss-Prot, and dedicated specifically to advancing our understanding of the human proteome, is sunsetting its operations. The final build and release of neXtProt, 2023–09, will remain available for the rest of 2024 but is becoming outdated as UniProtKB/Swiss-Prot continues to be updated. The neXtProt build process began each cycle with the most recent version of all proteins curated in Swiss-Prot and then added annotations to those proteins based on several additional data sources. 4 A major value-add process in the neXtProt build pipeline was to take the high-quality peptide lists produced by the PeptideAtlas and MassIVE-KB human mass-spectrometry-based proteomics resources and upgrade the protein PE (“protein evidence”) status to PE1 for all proteins that met the peptide uniqueness, length, and extent rules set forth in the HPP MS Data Interpretation Guidelines 3.0. 12 At least two high-quality uniquely-mapping peptides of at least nine aa in length covering at least 18 aa in total in at least one of the two resources. (Meeting the guidelines via a combination of single peptides from each of the resources is unacceptable due to amplification of false positives.) This process has now been implemented by UniProtKB in the same way, so that UniProtKB now reflects PE numbers as previously implemented in the neXtProt pipelines. Table 1 provides a summary of the transition from neXtProt to UniProtKB/Swiss-Prot as the knowledge base of protein information and computation of PE values. Column 1 lists various metrics for which values are presented in three different releases. Column 2 lists the metric values previously published in the 2023 HPP metrics from the neXtProt release 2023–04. Column 3 provides the metric values as contained in the UniProtKB/Swiss-Prot release 2024_03, which was prior to the integration of PeptideAtlas and MassIVE-KB mass spectrometry data (further described below). The last column provides the metric values for UniProtKB/Swiss-Prot release 2024_04 (exported 24 July 2024), after the integration of PeptideAtlas and MassIVE-KB mass spectrometry data. The numbers after this integration match closely to the final state of neXtProt. It is notable that the final release of neXtProt has substantially more PE2 proteins (insufficient protein-level evidence, but extant transcript-level evidence) and fewer PE3 and PE4 proteins (no transcript evidence) because neXtProt also integrated RNA-seq based data from the HPA to provide enhanced transcript-level evidence for PE2 classification. UniProtKB could consider a similar data approach. As discussed last year in the 2023 HPP metrics update, there are several hundred proteins that were added many years ago that are of dubious value in determining the HPP protein target list. These include immunoglobulin variable regions, T-cell receptor VDJ regions, endogenous retrovirus genes that are almost certainly not translated, and more. In each of these cases, the UniProtKB/Swiss-Prot entries represent an arbitrary subset of a much larger population of variable sequences not directly found in the reference genome or pseudogenes that are thought not to be translated. These entries are largely artifacts of history and do not reflect a modern set of protein-coding genes. Furthermore, UniProtKB/Swiss-Prot contains over 600 PE5 entries, described as dubious or uncertain genes, largely (but not entirely) thought to be pseudogenes and other potentially non-translated genes, retained for various reasons, including enabling the possibility of identification of protein expression by some of these genes by ensuring the inclusion of their sequence in mass spectrometry search engines. This list is under constant review. Starting in 2013, HPP excluded PE5 entries from the target list of proteins to be found. In order to update the target protein list, the HPP is transitioning its reference protein target list from a neXtProt-based system to an GENCODE-based system. 14 GENCODE’s mandate is to annotate all features of interest onto the reference human genome based on many lines of evidence, including evolutionary conservation derived from multiple genome alignments across many species, as well as data from RNA-seq, translatome sequencing information, proteomics, etc. Thus, GENCODE contains a comprehensive list of human protein-coding genes with precise links from genomic coordinates to transcripts to proteins according to the latest detailed understanding of the genome. This represents a far more appropriate HPP protein parts list than the subset of reviewed entries in UniProtKB. The HPP Executive Committee (EC) voted to make this switch. UniProtKB remains the definitive repository of information on protein expression, attributes, and function. HPP is aligning its protein parts list to GENCODE, which is itself linked to UniProtKB. Thus, the list now begins with GENCODE, but the extensive annotations, including PE levels, continue to come from UniProtKB/Swiss-Prot and UniProtKB/TrEMBL. It should be noted that there is an ongoing effort to more fully align GENCODE, UniProtKB, and RefSeq, but that will likely take several more years to complete. The effects of this change are minor at the overall summary level yet sweeping in detail. Table 2 provides the change in number of proteins per PE status from the final neXtProt-based list to the latest GENCODE version 46-based list. A total of 1254 protein entries have been dropped from the HPP protein parts list, but 276 were added for a net reduction of 978 in the total number of entries. In the 2024 list there are 19411 entries, including 92 PE5 entries and 23 entries for which there is currently no PE value because they cannot currently be mapped to UniProtKB (which will be resolved in the coming months). PE5 proteins are now included because the majority, 519, have been discarded, but the remainder, 92, have undergone some manual validation by GENCODE and may well be protein coding, and thus warrant further evaluation. Supplemental Table S1 provides the list of 1254 neXtProt entries that have been discarded because they are not found in the GENCODE protein-coding gene list. The entries are coded with a group name based on simple patterns in the protein description or gene symbol. The groups are “IgVar” for entries with “immunoglobulin … variable” in the description, “Ig” for entries with “immunoglobulin” without the “variable”, “HERV” for entries with “endogenous retrovirus” in the description, “TCRVar” with “T cell receptor … variable” in the description, “lncRNA” for entries with a gene symbol starting with “LINC”, “Humanin” for entries with “humanin” in the description, “OlfactoryRecep” for entries with “olfactory receptor” in the description, “PutUnchar” for entries with a description with “Putative uncharacterized” in the description, “Putative” for entries with a description with “Putative” but not “uncharacterized”, and finally “Other” if an entry cannot fit into one of the above categories. These groups are summarized in Figure 1 to show their relative numbers. Notably, 141 entries are immunoglobulin variable regions, 59 entries are human endogenous retrovirus sequences, 98 are T cell receptor variable regions, and 16 are lncRNAs. These sets of entries are all small incomplete subsets of a much larger population of similar entities that are not all in UniProtKB/Swiss-Prot due to historical happenstance; they are not representative of the inventory of distinct protein-coding genes. Supplementary Table S2 contains the entire list of 19411 entries in the GENCODE list (based on GENCODEv46) and the adopted HPP target list 2024. The table may be filtered and explored to examine important subsets of proteins. For example, of the 405 entries annotated as olfactory receptor, 27 are PE1 and just one (OR51E2) has canonical status in PeptideAtlas. (We note that OR51E2 is only detected in prostate tissues in PeptideAtlas, and UniProtKB offers an alternate name “Prostate-specific G-protein coupled receptor.”) Table 3 and Figure 2 show the progress achieved by the HPP over time in identifying the human proteins with protein-level evidence (PE1). Table 3 depicts the progress from 13975 PE1 proteins in neXtProt 2012–02 to 18397 in neXtProt 2023–04, and then 18138 for 2024 using the new GENCODE+UniProtKB-based target list. This represents 93% of the predicted proteins, with 1273 proteins not yet identified with PE1 evidence (“missing proteins”). Figure 2 shows the progress since 2016 when the more stringent Guidelines for Mass Spectrometry Data Interpretation v2.1 were implemented. The total length of the bars shows the sum of proteins in the HPP target list with the blue portions validated to PE1 via mass spectrometry evidence, the orange portions validated with various non-MS evidence, and the gray portions are the missing proteins. The current version 3.0 of the Guidelines was implemented starting for the year 2020 and provides a checklist at https://hupo.org/HPP-Data-Interpretation-Guidelines . 12 As noted above, the key requirements are two high-quality uniquely-mapping peptides of length nine or more residues covering at least 18 aa of the protein sequence, confirmed upon reanalysis by PeptideAtlas or MassIVE-KB. There are fewer PE1 proteins in the HPP target list in 2024 than in 2023, mainly due to the overall reduction in the number of proteins in the list. There were 410 PE1 proteins dropped from the target list that were largely immunoglobulin variable regions, T-cell receptor variable regions, and other similar sequences, many of which are indeed detectable with MS, but represent a subset of variations of otherwise included entities. The number of missing proteins has been reduced by 108, and PE1 proteins that have not yet been detected with MS by 213. PeptideAtlas added 214 new MS datasets in the 2024 build. Most of these datasets added few new proteins that had not been previously detected. However, of the 171 new canonical proteins in PeptideAtlas (having two or more uniquely-mapping peptides of length at least nine aa, covering a total of 18 aa or more), several datasets stand out. Figure 3 shows the contributions of the top five datasets of the new 214. PXD007985 from Bai et al., 18 based on deep whole-proteome and phosphoproteome profiling of human brain samples from different Alzheimer’s disease (AD) stages and non-AD neurodegenerative cases added by far the most, with 43 new canonical proteins. The next four highest contributing datasets are PXD006201, PXD006122, 19 PXD006939, 20 and PXD00519 , 21 providing 18, six, five and four new canonical proteins, respectively. All of these datasets are somewhat older but derive from uncommon sample types or enrichment that can probe further into the missing proteins. MassIVE-KB could not increase the number of datasets processed for 2024, but instituted more stringent thresholds, reducing the number of contributed peptides and proteins overall.

Launch

A chemical perturbation strategy is at the heart of the ChemBioFrance Grand Challenge project called “A Protein, a Ligand, a Function” ( https://chembiofrance.cn.cnrs.fr/en/ ). Individual experiments will utilize biologically-active small molecules from the National Chemical Library for exposure to the proteome of specific cell lines in order to permit assessment of the variation in the proteome in response to treatment. Of course, some of these biochemical effects could be off-target, making the analysis complex. Eight pilot projects have been selected to generate preliminary data, initiate and inform a database, stimulate hypotheses, and compete for major funding. Anonymized data from the first pilot are being used to establish a robust standardized workflow and analysis strategy. So far, all samples have been analyzed in DIA-PASEF mode on a large generation ion mobility mass spectrometer ( i.e., timsTOF Pro) with input from the HPP Mass Spectrometry Pillar for subsequent data analysis using SimpliFi (ProtiFi). The data generated are being mapped in Reactome ( https://reactome.org ) by superimposing the Function Evidence calculated for each protein on the pathways highlighted in Reactome . It should be feasible to focus on the fine function characterization of a selected protein carrying, e.g., the FE2 level label. Discussions are ongoing to replicate this approach in Spain and South Korea and Germany with their national chemical libraries. We hope to stimulate hundreds of biologists to participate in this Grand Challenge Project.

Protein

As the HPP draws near to achieving its first goal of completing confident detections of all entries in the human proteome target list (with the rate slowing down just before the finish line due to unexpressed genes and/or hard-to-detect proteins), the HPP has embarked on a perhaps more audacious goal of determining at least one molecular function for each protein. 4 There is always more to be learned about every protein, but the HPP Grand Challenge Project now sets as its goal a solid understanding of at least one function for each protein at the molecular level, ideally via at least two orthogonal methods. This goal has been under development since the C-HPP teams focused on 1260 functionally unannotated PE1 proteins (uPE1) in 2018, 22 which was reduced to 1134 in neXtProt 2023–03. 4 To track the progress toward this goal, a metric must be developed with which to measure it. The PE score developed by UniProtKB - decades ago and enhanced by the HPP Mass Spectrometry Data Interpretation Guidelines has served the HPP well in measuring progress across the community in detecting the expression of each predicted protein on the protein parts list. We have, therefore, developed an analogous FE (“function evidence”) score. As PE1 is the top tier goal for detection, FE1 is the top-tier goal for elucidating the function(s) for each protein. Of course, that does not imply that nothing more can be learned about an FE1 protein; in fact, many proteins have multiple functions, depending upon localization, binding partners, and intracellular conditions. FE5 represents the lowest category, indicating that essentially nothing is known. Table 4 lists a set of aspirational definitions for FE1 through FE5 as developed at the HPP Workshop at the HUPO World Congress in Cancún, Mexico, in December 2022. In practice, to be a useful metric, the FE score must be computable from a repository of protein function information; that repository should, of course, be UniProtKB. All functional information in UniProtKB can be downloaded as an XML file, parsed, and analyzed computationally. The FE working group of HPP has established a set of definitions for FE scores based on information available in UniProtKB, specifically based on the free-text functional description and associated publications, Gene Ontology terms, and EC (Enzyme Commission) numbers. Furthermore, all annotations have an “ECO code,” a term from the Evidence and Confidence Ontology describing how the particular information is known. 23 The best ECO codes describe annotations that are manual assertions based on direct experimental evidence. These stem from UniProtKB curators reading papers in the literature that describe experiments that provide functional evidence. The lowest confidence functional annotations are automatic assertions and predictions based on computational results, not individually validated by curators. Table 5 provides a set of practical definitions for FE1–5, based on information directly extractable from UniProtKB. This allows the FE scores to be computed quickly across all ~20,000 entries in the HPP target list and progress to be tracked over the coming years. These practical rules were drafted by the HPP FE Score Working Group by manually examining 12 proteins spanning the full range of FE scores, achieving a consensus FE score for them, and then adjusting the criteria so that they reflect the perceived quality of functional information present in the annotations found in UniProtKB. Table 6 provides a summary of the current FE scores for all 19,411 entries in the HPP target list as a function of their PE values. The scoring algorithm loops through all entries in the HPP target list and applies the rules in Table 5 sequentially to assign the FE score that first matches the UniProtKB annotations associated with the entry. There are now 5229 PE1 proteins that have FE1 scores, another 5416 have FE2 scores, and the remaining 7493 are in the lower FE3,4 and 5 categories. Note there are a small number of entries without PE or FE value, due to challenges in mapping the GENCODE-based target list to UniProtKB; these challenges will be mostly solved in the coming year, although since both are continually advancing databases, some minor discrepancies that need to be solved will continue to emerge. There are broadly two mechanisms for improving the FE scores for individual proteins. First, if the current UniProtKB annotations do not reflect what is currently known in the literature regarding a protein function, then the UniProtKB entry needs to be updated to reflect what is known in order to bring its FE score up to the appropriate level. We recommend that researchers write a brief summary about what is known in the literature with appropriate citations and submit the document to UniProtKB for curation and entry into UniProtKB/Swiss-Prot using the ‘Entry feedback’ link present on the web view of every UniProtKB entry. The second strategy is applicable when the UniProtKB annotations do accurately reflect that there is insufficient knowledge about a protein’s function in the literature. In this case, new experiments must be performed, and the results submitted as a peer-reviewed publication. Once the new peer-reviewed work is available in the literature, a brief report as described above should be submitted to UniProtKB to ensure timely inclusion in the knowledge base. Computational methods are also able to contribute findings into the FE scoring system. Algorithms and pipelines such as I-TASSER/COFACTOR predictions of Gene Ontology terms were already incorporated by neXtProt as functional annotation of the unannotated PE1 entries. 24 These predictions were subjected to blinded comparison with functional annotation by neXtProt and by CAFA. 25 Further development utilizes BLASTp and MMseqs2. 26 AlphaFold 2 and AlphaFold 3 surely can be applied for deeper functional annotation. With the retirement of neXtProt, which has served as the primary knowledge base of the HPP, we have set up an HPP portal that will focus on providing relevant lists of HPP target-list proteins by chromosome, PE level, FE level, and more. The Human Proteome Project Portal at hppportal.net enables the community to explore the latest metrics and protein lists of the human proteome. This portal hosts all 19411 proteins in the 2024 HPP target list, along with such attributes as their chromosomal coordinates, Ensembl identifiers, UniProtKB identifiers, PE scores, and FE scores. Users can explore summaries of proteins by PE score, FE score, and chromosome carrying the protein-coding gene. Users can download lists of proteins filtered by these attributes, such as the list of FE5 proteins (essentially nothing known about function) with a PE1 score (credibly observed at the protein level). In 2024, there are 1460 such proteins, previously called the uPE1 unannotated proteins 22 . For each of the 19411 proteins, there is a summary page providing general attributes, protein expression and localization information by tissue and organ, and RNA expression and localization information by tissue and organ. The portal will be updated yearly to reflect the latest HPP metrics. The new portal will become available by December 2024.

Antibody

The Antibody Technology Initiative aims to utilize antibody-based proteomics technologies for mapping proteins in health and disease. In close collaboration with the Human Protein Atlas (HPA) project, www.proteinatlas.org , the pillar contributes to the HPP Grand Challenge by providing a resource for spatial localization of proteins at body-wide, cellular, and subcellular levels, a first step towards understanding protein function. Since version 23 release of the open-access database, ongoing efforts have determined protein localization at a more detailed level, both by refining the antibody-based methods and by studying new tissues and samples. In the Tissue section, 13 multiplex immunofluorescence is integrated with single-cell transcriptomics analysis 125 to capture protein expression in subsets of cells not previously distinguishable by the human eye, including renal tubules, salivary glands, germ cells during spermatogenesis in testis, and subcellular compartments of motile cilia in various tissues. Rare expression patterns in ovarian follicle cells yielded evidence for several missing proteins. 126 Such cells are challenging to study with MS-based methods due to their low cell count; spatial proteomics methods add important insights. In the Subcellular section 127 a major effort examined proteins expressed in primary cilia or sperm at high resolution. The Disease section has extensive datasets utilizing the Olink Explore extension assay on blood plasma from patients with different diseases, including cancers. 128 In the upcoming version 24 of the HPA database, scheduled for launch in late 2024 at the Dresden HUPO World Congress, multiple major updates will be added to several sections, both in terms of internally generated data and importing information from other biological databases. HPA serves as an important knowledge resource gathering data on human proteins from different perspectives, identifying patterns of expression relevant for answering questions about presumed function or clinical significance.

Progress

As the HPP draws closer to completion of its first goal of achieving confident detection of expression of all the entries in the reference target protein parts list, the HPP has embarked on the Grand Challenge of determining and verifying at least one molecular function of each protein. For several years we have published a chromosome-by-chromosome table from neXtProt with a column for uPE1 protein numbers (PE1 lacking function annotation); for 2023 there were 1181 uPE1 entries. Several new annotations were stimulated by the C-HPP CP50 initiative. 22 As described above, we are establishing an FE1–5 scoring system for Function Evidence of the relevant data, in conjunction with UniProtKB. We are also preparing an HPP Portal in support of the Grand Challenge. We expect contributions from predictions based on sequence homology, protein:protein and protein:RNA interactions, and in silico predictions of detailed structures with AlphaFold2, 129 AlphaFold3, 130 and other algorithms. The 2022 HUPO Congress in Cancún, the 2023 HUPO Congress in Busan, Korea, and the planned 2024 HUPO Congress in Dresden, Germany, have highlighted the complementary goals of the investigators in the π-HuB Proteomics Navigator mega-project in China and the gradually emerging HPP Grand Challenge Project “A Function for Every Protein”.

Pathology

Trends noted in the 2023 HPP Metrics report 4 have continued with the transition from genomics to phenomics in personalized population health, addressing both wellness and disease. 101 An example of translation from research to clinical practice is a CLIA-compliant MS-based assay for thyroglobulin. 102 Emergence of more infectious variants of SARS-CoV-2 ( e.g., Omicron BA.4 and BA.5) has been investigated. High-throughput proteomic profiling of saliva samples identified new pathways of pathogenesis. 103 A multinational network of clinicians and informatics experts curated electronic health records data from patients hospitalized for COVID-19 with respiratory or neurologic complications to determine reasons for prolonged hospitalization and/or death. 104 , 105 An MS data portal (CoVProt) aims for comprehensive curation of COVID-19 clinical samples for deep proteomic investigations, data visualization, and facile data accessibility. 106 High-throughput sample preparation that is compatible with large-scale screening protocols remains a bottleneck. Refinement of a paramagnetic bead-based digestion protocol for automatic sample preparation was linked with an artificial neural network for different sample types. 107 The CLINSPECT-M consortium aims to develop a standardized start-to-finish, fit-for-purpose workflow for clinical specimens designed to benchmark and improve current best practices, as shown for plasma and cerebrospinal fluid specimens. 108 A flexible web-based application (MassSpecPreppy) uses a 96-sample liquid handling robot and Evotip trap columns. 109 Another bottleneck caused by the massive data sets being generated, often in different formats, combined with the associated electronic health records now being collected, is rapid analysis of the data. 110 Protocols include Propensity Score Matching, 111 AlphaPept 63 and TransDSI. 112 Etiologically different cholestasis was discriminated by modeling proteomics datasets. 113 Song et al used AlphaFold2 114 to predict the structures of 11000 human isoforms; multiple metrics were employed to identify splicing-induced structural alterations. 115 Serum continues to be the most commonly used biological matrix (see Cancers BD, above). A study of regulation of bile acids synthesis during the early stages of liver regeneration used serum from human liver donors. 116 Other sample types include tissue samples/biopsies, 87 , 117 organoids, 118 secretome, 119 and synaptosomes. 120 Biomarkers are valuable for estimating biological age and determining disease state, therapeutic response, and clinical outcomes. 121 – 124

Conclusion

The HUPO Human Proteome Project has entered a new phase, transitioning to UniProtKB and GENCODE for reference proteome and target list. In support of the HPP Grand Challenge Project, “A Function for Every Protein,” and the “π-HUB Protein Navigator for the Human Proteome”, we have developed a Function Evidence Annotation Scoring System. This system complements the well-established UniProtKB/SwissProt/neXtProt Protein Evidence scoring system for protein expression.

Highlights

These reports from several of the BD-HPP teams reflect specific work from the HPP team members, as well as some notable related publications identified by the BD team as significant in their domain. The HBPP ( https://hupo.org/brain ) comprises a diverse group of scientists promoting the connections of neuroproteomics and actively encouraging the inclusion of young investigators. HBPP workshops have connected the international neuroproteomics community. The 33rd HUPO Human Brain Proteome Project Workshop took place in May 2024 near Dublin, Ireland. Participants from eight countries addressed clinical, animal, and cellular findings for neurodegenerative disorders (Alzheimer’s diseases, Parkinson’s, multiple sclerosis) and neuropsychiatric conditions (schizophrenias, bipolar disorder, autism). While most studies employed mass spectrometry-based proteomics, large-scale protein arrays of different kinds and combinations of proteomics with other omics platforms were prominent. Multivariate machine learning and Mendelian randomization of GWAS findings were elegantly combined with proteomics findings, including pathway analyses and investigations of posttranslational modifications. Findings from the open-access proteomics resource profiling plasma samples from 50000 UK Biobank participants help elucidate biological mechanisms underlying proteogenomic discoveries. 27 Multiplex protein profiling of Alzheimer dementia significantly enhanced the utility of established CSF biomarkers of synaptic damage, such as GAP43, NRGN, or SNCB, when deployed with correlated brain-derived proteins such as PTPRN2, NCAN, or CHL1. 28 Bottom-up proteomic studies using CSF introduced an open-source quality control tool, MaCProQC , enabling rapid data comparison of raw data, identification, and quantitative features in the same run through pre-defined quality metrics. 29 Proteomics work in First Episode Psychosis showed that levels of complement proteins prior to treatment are associated with the subsequent response to treatment, as indexed by remission status and change in symptom severity. 30 To better understand the neurodevelopmental aspects of schizophrenias, another study investigated the mitochondrial and nuclear proteomes of neural stem cells (NSCs) and neurons derived from induced pluripotent stem cells from schizophrenia patients and mentally healthy individuals to assess possible alterations related to energy metabolism and mitochondrial function during neurodevelopment. Functional analysis of NSCs revealed alterations in mitochondrial oxygen consumption in schizophrenia-derived cells and a tendency of higher levels of intracellular reactive oxygen species (ROS). 31 The 34th HUPO HBPP workshop will take place in Toledo, Ohio, USA, in 2025. Notable human brain proteomics publications in the past year include these on Alzheimer disease, 32 – 35 Parkinson’s disease, 36 – 38 and aging 39 . The ultimate goals of proteomics-driven precision medicine are to achieve precise diagnosis and safe, effective treatment. A non-invasive early diagnostic panel (P4) for hepatocellular carcinoma (HCC) using an MS-based discovery-verification-validation proteomics workflow predicted the conversion of liver cirrhosis (LC) to HCC with a median lead time of 11.4 months earlier than imaging methods, with an accuracy of 90% 40 . This breakthrough provides a more accurate predictive tool for identifying occult HCC not visible through imaging. DeepRTAlign, a deep learning-based retention time alignment tool for large cohort LC-MS data analysis, served as a machine-learning-based HCC early recurrence prediction model that significantly enhanced diagnostic accuracy. 41 Another machine-learning-based early recurrence prediction model for HCC specifically targeted GSN and microvascular invasion; this model achieved an AUC of 0.803 (95% CI: 0.786 ~ 0.820), demonstrating robust performance even in AFP-negative patients. 42 DeepRTAlign is applicable to many large cohort LC-MS proteomic and metabolomic studies. 41 Multiomics analysis identified three proteomic molecular subtypes of HCC and elucidated significant differences among these subtypes in genetic alterations, microenvironment dysregulation, kinase substrate regulatory networks, and therapeutic responses. A precision treatment strategy using sorafenib was based on these proteomic subtypes. Machine-learning-based response prediction models enhanced the implementation of advanced personalized therapeutic strategies for HCC. 43 An Omics and Text driven Translational Medicine (OTTM) tool enabled large-scale drug target prediction on HCC proteomic data. This analysis identified forty FDA-approved or clinical trial drugs as potential agents for targeted intervention for HCC. 44 Lysozyme (LYZ), a secretory protein with antibacterial functions generally expressed in monocytes/macrophages, was identified as a prognostic marker for HCC. LYZ promoted HCC proliferation through activation of downstream signaling pathways via cell surface GRP78, making LYZ a potential new intervention target. 45 The HLPP has a pivotal role in advancing proteomics-driven precision medicine for HCC. The HGI over the past year focused on advancing its second community challenge, comparing and improving the bioinformatic tools for N- and O-glycopeptide identification and quantitation. This study is expected to involve more than 20 software developer teams from around the world. With the experimental design now in place, cross-lab data analysis and evaluation of software performance are scheduled for the second half of 2024 with completion in the first half of 2025. To bridge the first and the second HGI studies, the HGI organized a well-attended interactive session on state-of-the-art glycoproteomics software at the Glyco26 meeting in Taipei, Taiwan in August 2023. The HGI leadership team summarized the opportunities and future directions in glycoproteomics. 46 Finally, the HGI community and the Beilstein Institute have generated glycoproteomics guidelines with the minimum information required for a glycoproteomics experiment, expected to be released in late 2024. Significant advances have been achieved by the HIPP in enhancing peptide antigen acquisition and identification via LC-MS, advancing the definition of the non-canonical HLA-associated peptidome, and fostering clinical impact and translational research. New instrumentation and acquisition techniques have improved the sensitivity and scope of HLA peptide identifications, specifically methods utilizing Parallel Accumulation-Serial Fragmentation (PASEF) Mass Spectrometry. 47 Immunopeptidomics sequence mapping is more complex than enzyme-specific searches due to the larger theoretical search space, leading to higher false positive rates, particularly for less well-characterized protein sources, like non-human or non-canonical open reading frames. To mitigate false positive identifications, multiple tools have been employed to boost confidence in peptide identifications, including spectrum prediction and retention time prediction tools, 48 and de novo sequencing rescoring approaches which integrate HLA binding prediction. 49 These tools have been crucial in defining the non-canonical peptidome, including peptides derived from RNA genes. 50 Novel sources for HLA-peptide presentation include circRNA-derived HLA-peptides spanning fusion sections, 51 and pan-viral ORFs. 52 The field has further contributed to our understanding of tumor immune-recognition, and to characterization of the NSCLC microenvironment, providing evidence of immune-editing in inflamed tumors. 53 Clinical impact has been reported using an mRNA vaccination strategy incorporating MS-identified cancer antigens in pancreatic cancer. 54 Peptide vaccination in NSCLC overcame poor ICB response in mice. 55 Finally, immunopeptidomics identified physiologically relevant off-target antigens for TCR-like antibody drugs, 50 offering a future tool for drug screening and optimization approaches. Single cell analysis has revealed remarkable heterogeneity of tissues and decisively contributes to our understanding of complex biological mechanisms, such as organ development and disease progression. Global mass spectrometry (MS)-based proteomics of small sub-populations and single cell resolution aim to complement other omics approaches and to characterize cellular functionality. Despite advances in sample preparation, sensitivity of instrumentation, and data analysis, the field is still challenged to provide reproducible deep proteome coverage at high throughput. 56 – 58 Protocols dedicated to single cell proteomics encompass top-down and bottom-up sample preparation, label-free approaches, isobaric and non-isobaric multiplexing coupled with data-independent, data-dependent, and targeted acquisition methods for ultra-low sample input. 59 The introduction of MS instrumentation operating at up to 200 Hz with high-resolution ion mobility separation enables deep proteome coverage at short chromatographic separations of <15 minutes. 56 , 59 To increase the identification and quantification of such convoluted and sparse data sets, dedicated data analysis pipelines have been introduced. 60 – 63 The sensitivity improvements of dedicated SCP workflows and instrumentation enabled measurements of post-translational modifications from single cells to study protein turnover, protein transport, biological heterogeneity of oocytes, composite small lymphocytic and classical Hodgkin lymphoma, and mouse hematopoietic stem and progenitor classification. 60 – 65 An in-depth comparison of diverse setups from the identical sample batch acquired in multiple laboratories will facilitate informed decisions on experimental design to apply spatial or single-cell proteomics for diverse biological questions. Urine proteomics is a rapidly advancing field in the search for biomarkers that can aid in the diagnosis and monitoring of a wide variety of diseases. The sampling is non-invasive and sensitive to changes in organs throughout the body. MS-based urinary proteomics has expanded the catalog of proteins detected in urine, 66 but the urine proteome remains much less complex than those of serum or plasma. Articles last year identified biomarkers from urine for ovarian cancer, 66 pancreatic cancer, 67 acute pancreatitis, 68 and bladder cancer, 69 including use of machine learning models. 70 Additional biomarker studies addressed endometriosis, 71 heart failure, 72 diabetic kidney disease (with a panel of 8 proteins), 73 and necrotizing enterocolitis (NEC) in premature infants. 74 Monitoring a pregnant mother’s urine proteome to assess fetal development shows promise, as well. 75 Since its launch in 2002 and first major reports in 2005, 76 there have been periodic updates of the Human Plasma PeptideAtlas and Human Plasma Proteome. The latest is Geyer et al (under review at J Proteome Res ). Both intra-individual and inter-individual variation of circulating proteins is high. 77 Strategies and methods to analyze large cohorts with high precision include affinity-based platforms for pre-selected sets of proteins (Olink), 78 , 79 and NULISA (Alamar Biosciences) 80 . The challenges of comparing the detectability and concentrations of targeted proteins across these platforms and MS are discussed. The number of plasma proteins detectable by MS has been lower than with affinity methods but is significantly increased now with the Orbital Astral MS, 81 especially combined with enrichment. Superparamagnetic functionalized nanoparticles can compress the wide dynamic range of proteins in plasma; also, simple precipitation with perchloric acid has yielded up to 1300 proteins in plasma. 82 The 2023–04 PeptideAtlas build of Serum and Plasma combined has 113 datasets with 4608 canonical proteins meeting HPP guidelines (isoforms and immunoglobulins excluded), with FDR <0.04%. Extracellular vesicle proteins add 377 canonical proteins. There are many interesting developments with glycoproteins and phosphoproteins in circulation, and co-expression of proteins can feed into network analyses with clinical implications. Cancers continue to be a major thrust of the BD-HPP, as evidenced by the large number of HUPO-related publications in this reporting period, including kidney, 83 pancreas, 84 – 86 colorectal, 87 breast, 88 liver, 89 lung, 90 , 91 ovary, 92 prostate, 93 brain, 94 and melanoma. 95 The wide range of cancers studied indicates the general applicability of many of the approaches that are being used, including major recent pan-cancer proteogenomic characterizations. 96 , 97 CPTAC researchers have identified and validated potential therapeutic targets across 10 cancer types, utilizing overexpressed/hyperactivated proteins, tumor suppressor gene loss-associated dependencies, neoantigens and tumor-associated antigens; 98 kinases are especially attractive targets for medicinal chemistry. 99 The proteogenomic studies of various cancers by the CPTAC consortium have repeatedly demonstrated the necessity of direct analyses of the concentration, localization, post-translational modifications, protein-protein interactions, splicing, and functional features of proteins. 79 , 83 – 93 These features are not predictable from studies of RNA, DNA, methylation, or metabolomics and explain the generally low to moderate (0.2 to 0.5) correlation in expression levels for paired proteins and transcripts. The abundance of a protein also can be influenced by transcripts encoding its interacting partners, not just its own cognate transcript. Such knowledge can guide the development of drugs that target specific proteins or pathways involved in disease processes and capture inter-individual variation essential for personalized medicine. Cost-effective, reproducible mass spectrometry analytic workflows are crucial in drug discovery and biomarker development. Recent advances include high-throughput MS workflows that are designed to be more affordable and efficient. Examples are the Echo MS+ system that removes bottlenecks from sample preparation to data reporting and minimizes sample consumption, and the fast ion trap mass spectrometer Stellar combined with very short chromatographic gradients, yielding absolute peptide and protein quantification in biological matrices with intuitive data-independent acquisition (DIA). 100

Introduction

As the flagship initiative of the Human Proteome Organization (HUPO), 1 the global Human Proteome Project 2 (HPP) has pursued two goals: (1) to credibly identify at least one isoform of every protein-coding gene, primarily by mass spectrometry, and (2) to make proteomics an integral part of functional multiomics studies of health and disease 3 , 4 . The HPP is organized into 24 nuclear and one mitochondrial chromosome-based teams, 16 biology and disease categories, and four resource pillars ( www.hupo.org/human-proteome-project ). The HPP Grand Challenge Project aims to characterize the function(s) of each protein. The HPP is supported by several major knowledge bases and resources. Since 2011, neXtProt 5 has served as the primary knowledge base for the HPP. neXtProt was based on UniProtKB/Swiss-Prot, 6 which was started 38 years ago as the definitive reference knowledge base of protein-based information for all species. The mass spectrometry (MS) resources PeptideAtlas 7 , 8 and later MassIVE-KB 9 have provided comprehensive, stringent MS-based information obtained via large-scale reprocessing of data sets generated community-wide, shared via ProteomeXchange, 10 , 11 and meeting HPP Guidelines 12 . Large-scale information based on antibody assays and transcriptome is provided by the Human Protein Atlas. 13 Here we describe the progress in advancing our understanding of the human proteome since the 2023 update 4 and enhancements to the HPP. First, the HPP reliance on neXtProt for the protein parts list has been changed to the Ensembl-GENCODE (henceforth GENCODE)-based list of protein-coding genes 14 with UniProtKB becoming the knowledge base of protein information. We have created a preliminary new Function Evidence (FE) score that ranks the evidence for our current understanding of the molecular function(s) of each protein, tied to the Gene Ontology 15 , 16 annotations contained in UniProtKB. This is a key step in the pursuit of the HPP Grand Challenge Project, “A Function for Every Protein.” We also are developing an HPP Portal to provide up-to-date metrics and convenient lists of proteins in the HPP protein parts list.

Supplementary Material

Supplementary Table 1: The 1254 UniProtKB-Swiss-Prot entries that are not found in the Ensembl-GENCODE list of protein-coding genes and thus were dropped from the HPP target list. Supplementary Table 2: The 19,411 entries of the 2024 HPP target list based on the Ensembl-GENCODE list of protein-coding genes with additional information from UniProtKB-Swiss-Prot. Filtering the tabulation readily reveals protein entries by PE class or other criteria.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

⚙ Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml ⓘ

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-09-27T09:11:36.575535+00:00
unpaywall
last seen: 2026-09-18T06:25:56.777850+00:00
License: CC-BY-NC-4.0