Reference protein-coding transcripts of human genes annotated using long-read transcriptome datasets | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Reference protein-coding transcripts of human genes annotated using long-read transcriptome datasets Kuo-Feng Tung, Wen-chang Lin This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8974155/v1 This work is licensed under a CC BY 4.0 License Status: Under Revision Version 1 posted 12 You are reading this latest preprint version Abstract Accumulating NGS expression datasets suggest that protein-coding genes produce numerous alternatively spliced transcripts. However, this observation might be overestimated in short-read sequencing data, which often cannot accurately resolve distinct spliced isoforms and introduce ambiguity. Resolving tissue-specific expression profiles is crucial to identify bona fide translated peptide products. In this study, we identified the most highly expressed protein-coding transcripts by using long-read NGS datasets to better understand the biochemical and biological functions of human protein-coding genes. Using nanopore sequencing data from 30 normal human tissues in the GSE192955 dataset, we identified 18,094 dominantly expressed representative protein-coding transcripts (Ref-Tx) from 18,557 human genes. Comparison with MANE-select transcripts revealed that 14,546 Ref-Tx transcripts matched those in the MANE-select dataset. This result indicates improved agreement between long-read transcriptome data and MANE-select transcripts. A higher proportion of Rank1 transcripts were identified as Ref-Tx in the long-read dataset. Similar patterns were observed when Ref-Tx were compared with functional APPRIS annotations. Given the importance of tissue-specific expression profiles for protein-coding transcripts, we developed an expression visualization bioinformatic tool (eCPG). This webtool integrates the extensive expression information from 30 normal human tissues as well as from the GTEx project, which is designed to interrogate the dominant protein-coding transcripts. Biological sciences/Biological techniques Biological sciences/Biotechnology Biological sciences/Computational biology and bioinformatics Biological sciences/Genetics long-read RNA-seq dataset human protein-coding genes dominant protein-coding transcripts MANE-select transcripts wobble splicing transcripts Figures Figure 1 Figure 2 Figure 3 Introduction The completion of the Human Genome Project provided a major breakthrough and opportunity to comprehensively understand human gene composition and biological modulation. Despite considerable bioinformatic analyses of human genome sequences, full annotation of human protein-coding genes remains an ongoing and challenging task 1 – 4 . Gene annotation in eukaryotic genomes is challenging in part because of the presence of alternatively transcribed mRNA transcripts generated from multiple exon usages 5 . This challenge has been further amplified with the greatly expanded availability of next-generation sequencing (NGS) transcriptome datasets 2 , 6 . Massive short-read NGS datasets cannot completely resolve alternatively spliced mRNA transcript isoforms and may further generate unnecessary ambiguous transcripts, particularly in experiments with low-quality library preparations. Consequently, this problem has become more pronounced with increasing accumulation of short-read datasets. Current estimates indicate that more than 8.76 transcripts are expressed per human protein-coding gene 7 . For example, in the GENCODE v47 annotation, 170,270 mRNA transcripts are annotated for 19,433 human protein-coding genes, and more than 52% of the registered transcripts (89,832) are designated as protein-coding transcripts on the basis of the GENCODE biotype feature 7 . This finding implies that not all mRNA transcripts derived from protein-coding genes could be used for peptide translation. For protein-coding genes, understanding the expression profiles of alternative transcript isoforms is necessary to identify the peptides that are ultimately translated. This is crucial because the biochemical and biological functions of most protein-coding genes depend on their translated protein peptides. Although alternative splicing can produce multiple protein isoforms, most human protein-coding genes generate dominant protein-coding transcripts and peptides 8 – 10 . Thus, identifying the dominantly expressed protein-coding transcript in each protein-coding gene is crucial for subsequent analyses of biochemical function based on the final translated peptide product. The APPRIS database is a widely used bioinformatic resource for this purpose 11 . This database assigns a single reference sequence, termed the principal isoform, to each gene by integrating reliable protein structural and functional information as well as cross-species conservation data 12 . Recently, another major annotation reference for human protein-coding genes was established. The MANE project (Matched Annotation from the NCBI and EMBL-EBI) aims to provide consistent transcript annotations by designating a single, well-curated representative transcript for each human protein-coding gene 13 . The MANE-select transcript serves as the standardized transcript for every human protein-coding gene in both the RefSeq and Ensembl databases. On average, human protein-coding genes have eight transcript isoforms 7 . For each protein-coding gene, the MANE-select transcript is the only chosen transcript and may correspond to the dominant protein-coding transcript 13 . A study reported substantial agreement between APPRIS principal isoforms and MANE-select transcripts with major proteomic peptides 11 . By contrast, short-read NGS datasets cannot accurately reflect the true expression status of alternatively spliced transcript isoforms 14 . We previously demonstrated that long-read NGS datasets provide a better understanding of the expression profiles of transcript isoforms among various tissue types 15 . For some protein-coding genes, specific transcript isoforms are differentially expressed and used for specific biological functions during development 10 , 14 , 16 . For example, the biological functions of the VEGF gene are modulated by the balance between alternatively spliced proangiogenic and antiangiogenic variants during tissue development 17 . Thus, examining the tissue-specific expression of dominant protein-coding transcripts is crucial for functional studies of protein-coding genes. In this study, we extend our analysis of tissue expression profiles of major protein-coding transcripts by using additional long-read datasets. Gao et al generated useful long-read nanopore RNA sequencing (RNA-Seq) datasets from 30 normal human tissues, including two fetal brain samples. They produced more than one billion RNA-Seq reads that were analyzed using an improved bioinformatic pipeline 18 . This dataset also included some major tissue types that were not covered in the GTEx V9 long-read dataset used in our previous study 15 . In addition, visualization of alternative transcript isoform expression among different tissues can help understand patterns of functional peptide expression and modulation. In this study, we developed a web-based bioinformatic tool to identify dominant protein-coding transcripts by using updated MANE annotations and long-read transcriptome datasets that include additional tissue types. Methods MANE and APPRIS datasets The MANE human protein-coding gene dataset was retrieved from the NCBI website ( https://www.ncbi.nlm.nih.gov/refseq/MANE/ ). The file used in this study was MANE.GRCh38.v1.4.summary.txt. MANE-select transcript information was extracted using the MANE_status feature column and other annotation features, including gene symbol, gene name, chromosomal location, Ensembl gene ID, Ensembl transcript ID, Ensembl protein ID, NCBI gene ID, NCBI RefSeq NM ID, and RefSeq NP ID. The dataset contains 19,388 MANE-select transcript records for human protein-coding genes and 66 MANE-plus-clinical transcripts. In addition to MANE-select transcripts, MANE-plus-clinical transcripts are annotated to report clinically significant variants in specific protein-coding genes. Ensembl Gene ID was used as the primary key feature for all subsequent comparisons and analyses across datasets. The APPRIS database is another crucial resource that assigns a single coding sequence isoform as the principal isoform for each gene based on protein features, including evolutionary conservation 12 . Principal isoform scores range from 1 to 5, with 1 indicating the highest reliability 12 . We directly retrieved APPRIS score data from the APPRIS website ( https://appris.bioinfo.cnio.es/#/downloads ). Because the GSE192955 dataset used the GENCODE v34 annotation file, we retrieved the GENCODE34/Ensembl100 Principal Isoforms.txt file for this study. GSE192955 long-read datasets covering major human tissues ESPRESSO is a bioinformatic tool used for the quantification of transcript isoforms in long-read NGS datasets 18 . The GSE192955 dataset provides nanopore RNA-seq data from 30 human tissues. These include 10 brain subtypes, two fetal tissues, and 18 adult tissues. The brain subtypes comprise the caudate nucleus, cerebellum, cerebral cortex, corpus callosum, frontal lobe, hippocampus, medulla oblongata, pons, temporal lobe, and thalamus. The two fetal tissues are central nervous system tissues, namely the fetal brain and fetal spinal cord. The adult tissues include the bladder, blood, brain, colon, heart, kidney, liver, lung, ovary, pancreas, prostate, skeletal muscle, small intestine, spinal cord, spleen, stomach, testis, and thyroid. Inclusion of these tissues allows more comprehensive assessment of expression profiles for human protein-coding genes. When compared with tissues included in the Genotype-Tissue Expression (GTEx) V9 long-read dataset, only six major tissues overlap, namely the brain, heart, liver, lung, pancreas, and skeletal muscle. The GTEx project used the GENCODE v26 dataset as the reference annotation, whereas the GSE192955 dataset was annotated using GENCODE v34 data. Then, the gencode.v26.basic.annotation.gff3 and gencode.v34.basic.annotation.gff3 files were retrieved from the GENCODE project for the present study 19 . GTEx short-read (V8) and long-read (V9) datasets The GTEx Project is a major resource for studying genotypes and gene expression 20 and is supported by the Common Fund of the Office of the Director of the National Institutes of Health. All GTEx data used in the present study contain no participant data and comply with the NIH Genomic Data Sharing guideline. We directly downloaded both V8 and V9 normalized transcript expression datasets from the GTEx Portal ( https://www.gtexportal.org/home/downloads/adult-gtex ). The data files used were GTEx_Analysis_2017-06-05_v8_RSEMv1.3.0_transcript_tpm for the V8 dataset and quantification_gencode_tpm for the V9 dataset. The retrieved GTEx datasets were processed using Python scripts to separate transcript expression data by tissue subtype. Each tissue type was assigned a separate transcript expression file. For each tissue, average expression values were calculated for individual transcripts, and the expression data from all tissue types were then combined. The GTEx V8 dataset covered 54 tissue types from 948 donors, whereas the V9 dataset covered only 14 tissue types. The tissues included in the V9 dataset were the anterior cingulate cortex, caudate, cerebellar hemisphere, frontal cortex, putamen, adipose subcutaneous tissue, breast mammary tissue, heart atrial appendage, heart left ventricle, liver, lung, pancreas, skeletal muscle, and fibroblast cell lines. Ref-Tx: Topmost reference protein-coding transcripts We identified the dominantly expressed protein-coding transcript for each human protein-coding gene. For each gene, we first calculated the average expression values of all annotated transcripts and ranked them by expression level. Because not all transcripts are protein-coding, as indicated by the GENCODE biotype annotation, only transcripts annotated as protein-coding were considered. The highest-expressed protein-coding transcript for each gene was designated as the Ref-Tx. In the present study, only one Ref-Tx was assigned to each protein-coding gene. These Ref-Tx transcripts were then compared with MANE-select and GTEx transcripts. Construction of the eCPG web database An expression catalogue of protein-coding genes (eCPG) web database was hosted in a Docker-based web environment running on an Ubuntu Linux server. The eCPG database was implemented using PHP within an Apache web server framework and integrated with a MySQL database 15 , 21 . The JavaScript D3 library was used for interactive visualization of transcript expression levels. Transcript expression data for protein-coding genes from the GTEx and GSE192955 datasets were stored as flat files and then imported into the MySQL database for use by the eCPG web interface. The eCPG web database is freely accessible at https://ecpg.ibms.sinica.edu.tw/ . Figure illustration was done by using the GraphPad Prism (version 10) software package. Results Summary of protein-coding genes in the long-read NGS dataset To examine additional tissue expression profiles of human protein-coding genes, we analyzed the GSE192955 dataset, which includes 30 major tissue types that differ from those in the GTEx V9 dataset 18 . After the exclusion of nonprotein-coding genes based on GENCODE biotype annotations, the dataset contained 105,281 expressed transcripts corresponding to 18,557 protein-coding genes. On average, this long-read NGS dataset identified 5.67 expressed transcripts per human protein-coding gene. The number of expressed transcripts per protein-coding genes was lower than that observed in the GTEx V8 short-read dataset (7.43 transcripts per gene), likely indicating differences in sequencing depth and tissue coverage. We focused on these 105,281 transcripts. Using GENCODE v34 biotype annotations, we classified 60,090 transcripts (57.07%) as protein-coding (Supplementary Table 1). By comparison, the original GENCODE v34 annotation file used in the pipeline contained 19,959 protein-coding genes with a total of 153,435 transcripts, including 84,068 protein-coding transcripts. Thus, approximately 92.9% of GENCODE annotated protein-coding genes and 68.6% of transcripts associated with protein-coding genes were expressed in this long-read dataset. When only protein-coding transcripts were examined, approximately 70% exhibited detectable expression. However, these transcripts were not expressed equally. Most highly expressed protein-coding transcripts When we examined the distribution of expression among ranked transcripts in the long-read dataset, we observed a clear dominance of the top-ranked transcript in each gene, consistent with findings of other studies 10 , 22 . As presented in Table 1 , for the most highly expressed Rank1 transcripts (18,557 transcripts), the average transcript expression level in protein-coding genes was 79.81%. By contrast, the average transcript expression level of Rank2 transcripts (15,537 transcripts) was 17.87%. The average expression level of Rank1 transcripts in the short-read GTEx dataset was significantly lower (66.41%) (Supplementary Table 2). This difference supports the advantage of long-read NGS platforms in more accurately resolving dominant transcript expression patterns. In short-read datasets, protein-coding genes generally exhibit higher expression levels in Rank2 through Rank10 transcripts (Table 1 and Supplementary Table 2). Table 1 Numbers of transcripts and average expression levels of Rank1 and Rank10 expressed transcripts of protein-coding genes within GSE 192955 dataset. There is a total of 105,281 transcripts belonging to 18,557 protein-coding genes. Transcript Rank Numbers of Ranked transcripts Average expression percentage Numbers of protein-coding transcripts Protein-coding transcript / all Ranked transcripts (%) 1 18,557 79.81 17,282 93.12 2 15,537 17.87 9,613 61.87 3 13,168 6.91 6,867 52.14 4 11,061 3.42 5,311 48.01 5 9,139 2.01 4,182 45.75 6 7,471 1.26 3,413 45.68 7 6,015 0.84 2,687 44.67 8 4,807 0.62 2,074 43.14 9 3,848 0.47 1,688 43.86 10 3,040 0.38 1,326 43.61 The number of protein-coding transcripts was also the highest among Rank1 transcripts. In the long-read dataset, 17,282 Rank1 transcripts (93.12%) were annotated as protein-coding, whereas in the short-read GTEx dataset, 17,472 Rank1 transcripts (89.18%) were annotated as protein-coding. In other ranked transcripts (Rank2 to Rank10), the proportion of protein-coding transcripts was similar between the two datasets (Table 1 and Supplementary Table 2). These findings further support the conclusion that protein-coding genes typically produce a dominant peptide product, consistent with the findings of a previous study 22 . However, not all top-ranked transcripts were protein-coding because approximately 6.9% of Rank1 transcripts were noncoding. Ref-Tx defined by most highly expressed protein-coding transcripts We focused on the most highly expressed dominant protein-coding transcripts and annotated the top expressed protein-coding transcript for each gene as the Ref-Tx transcript. Similar to the MANE project, only one Ref-Tx was assigned to each protein-coding gene on the basis of transcript isoform expression percentage and transcript biotype information. Among the 18,557 protein-coding genes in the long-read dataset, we identified 18,094 Ref-Tx transcripts. This finding indicated that 463 genes did not have detectable protein-coding transcripts expression in the long-read dataset. This absence may result from limitations in sequencing depth or tissue-specific expression because some genes may be expressed only in tissue subtypes not included in this dataset. We compared the 18,094 Ref-Tx transcripts with the 19,338 MANE-select transcripts and identified 14,546 matched transcripts. Among these matched transcripts, 14,091 (96.8%) were Rank1 transcripts. By contrast, in the GTEx V8 short-read dataset, the matched gene proportion was only 67.6% (11,820 of 17,472). This comparison further supports the advantage of long-read sequencing for identifying dominant protein-coding transcripts. We then examined the 14,546 matched MANE-select transcripts. As expected, most Ref-Tx were Rank1 protein-coding transcripts (14,091 transcripts; Fig. 1 A). A large proportion of Ref-Tx transcript (9,290 transcripts) accounted for more than 80% of total gene expression levels (Fig. 1 B), indicating strong dominance. Notably, seven MANE-select transcripts had modified Ensembl gene IDs in the updated MANE annotations. After accounting for these updates, 14,539 MANE-select transcripts were matched with Ref-Tx in the GSE192955 dataset. Compared with Ref-Tx defined from the GTEx V8 and V9 datasets, the GSE192955 long-read dataset exhibited stronger concordance with MANE project genes (Fig. 2 ). Consistent with the findings of previous study 14 , long-read datasets (GSE192955 and GTEx V9) were determined to be more advantageous for protein-coding transcript isoform expression studies. 2,015 MANE-select transcripts differ from Ref-Tx transcripts Notably, the long-read dataset contained 16,561 matched MANE-select transcript ID records. Thus, 2,015 MANE-select transcripts present in the GSE dataset did not qualify as top-ranked protein-coding transcripts under our Ref-Tx selection criteria. We further examined these 2,015 MANE-select transcripts that did not match the assigned Ref-Tx transcripts. As expected, none of them were Rank1 transcripts. Some discrepancies likely resulted from annotation differences between datasets. In this study, gene and transcript IDs were used as the primary features for data comparison. Thus, modified gene or transcript annotations may lead to mismatches for certain MANE genes or MANE-select transcripts. Furthermore, some MANE-select transcripts were annotated as processed transcripts under the GENCODE v26 transcript biotype classification and were thus excluded from our Ref-Tx assignment pipeline. More than 60% of the 2,015 MANE-select transcripts (1,268 records) were classified as Rank2 transcripts on the basis of expression percentage criteria. In addition, there were 386 Rank3 transcripts, 167 Rank4 transcripts, and 61 Rank5 transcripts. As expected, these protein-coding genes exhibited complex alternative splicing patterns, with an average of 8.39 transcripts per gene. When Rank2 transcripts were excluded, the average number of alternative splicing isoforms increased to 10.38 per protein-coding gene, which is nearly double the overall mean transcript count in this long-read dataset (5.67 transcripts per gene). These findings indicate the challenges associated with annotating complex alternative splicing isoforms in protein-coding genes. Because of alternative splicing, some transcript isoforms differ only in their untranslated regions (UTRs) and produce identical protein peptide products. In such cases, we selected the transcript with higher expression as the Ref-Tx transcript. Among these 2,015 protein-coding genes, 561 genes had identical coding sequence lengths between the MANE-select transcript and the corresponding Ref-Tx transcript based on GENCODE annotation. Although these transcripts shared the same open reading frame, their 5′ or 3′ UTR regions differed because of alternative promoter usage or splicing events. Notably, some of these transcript pairs exhibited substantial differences in expression levels between the MANE-select transcripts and Ref-Tx transcripts. This observation may be relevant for studies of transcriptional regulation, such as tissue-specific or developmental stage–specific promoter usage, even though the translated protein products are the same. These 2,015 protein-coding genes are listed in the eCPG web tool to facilitate user exploration, particularly for examining tissue-specific expression patterns. Wobble splicing transcripts with subtle splicing site selection Besides the 561 genes, there were 222 Ref-Tx transcripts with longer CDS region length, whereas 1,232 MANE-select transcripts had longer CDS region length. In many cases, transcripts with longer coding regions are regarded as representative full-length products. However, the biological functions of individual alternatively spliced transcript isoforms require further investigation. Among the 1,232 and 222 protein-coding genes in which Ref-Tx and MANE-select transcripts differed in CDS region length, we observed features consistent with the wobble splicing phenomenon reported previously 23 , 24 . Wobble splicing is characterized by the use of alternative tandem splice sites, involving proximal donor or acceptor sites, rather than complete exon inclusion or deletion 25 . This mechanism often generates isoforms with single or small numbers of amino acid insertions or deletions 26 . Our laboratory previously demonstrated that one splicing isoform of the ING4 gene alters protein subnuclear localization and degradation 24 . When possible wobble spliced transcripts were restricted to those with CDS region differences within 20 base pairs between MANE-select and Ref-Tx transcripts, 142 protein-coding genes were identified with such small amino acid insertions or deletions events, including e.g. FGF14 and RNF19B genes. Not all protein-coding genes produce a single dominant protein peptide; therefore, variations in expression among different transcript isoforms must be examined. Some genes modulate their biological functions through the expression of alternatively spliced transcripts in specific tissue types or developmental stages. Accordingly, we developed the eCPG bioinformatic web tool to provide visualization functions that facilitate examination of such transcript isoforms among different tissue types ( https://eCPG.ibms.sinica.edu.tw ). These 2,015 protein-coding genes identified in this study are specifically listed for further exploration by users. Comparison of APPRIS annotations for 2,015 MANE-select transcripts APPRIS is a database that provides information on protein structures and evolutionary cross-species conservation for protein-coding genes 12 . APPRIS assigns annotations to alternatively spliced transcripts for many model organisms. An APPRIS Principal 1 (P:1) score indicates that a transcript is considered the representative isoform based on the core APPRIS computational modules. Among the 14,546 transcripts matched between Ref-Tx and MANE-select transcripts, a high proportion of transcripts were annotated as APPRIS P:1 (83.5%). We further examined APPRIS annotations for the 2,015 MANE-select transcripts that did not meet the Ref-Tx criteria. Among these 2,015 genes, 513 genes had both Ref-Tx and MANE-select transcripts annotated as APPRIS P:1. In total, 1,098 Ref-Tx transcript and 640 MANE-select transcripts within this group were annotated as APPRIS P:1. Score. These observations suggest that a large subset of Ref-Tx transcripts defined based on expression criteria may also have functional relevance and evolutionary conservation. eCPG web tool for expression visualization on protein-coding genes We previously developed a bioinformatic tool, TEx-MST, to display MANE-select transcript expression using GTEx datasets as the single data source 15 . In the present study, we primarily focused on the GSE192955 dataset and extended the previous webpage design to incorporate additional expression information from the GTEx V8, GTEx V9, and GSE192955 datasets. Users can examine transcript expression information for human protein-coding genes on a single gene-specific webpage (Fig. 3 ). Basic gene information derived from the MANE project is displayed at the top of the page, followed by detailed transcript-level information. Because the GTEx V8 dataset contains a larger number of expressed transcripts, GTEx V8 data were used as the default source for constructing the transcript data table. Users may reorder and customize the display by clicking the ascending or descending icons at the top of each column. Annotation information from both GENCODE v26 and GENCODE v34, corresponding to the GTEx and ESPRESSO analysis pipelines, respectively, is also included. Transcript expression profiles from 30 tissues are presented in two graphical panels, exhibiting either expression percentages or TPM values. Within these panels, users can turn on/off individual transcript expression profiles by selecting the corresponding “RankX” label in the legend. In addition, links to the previously developed TEx-MST and TREGT web databases are provided to support further exploration of gene expression data 15 , 27 . Expression information for the endothelin-converting enzyme 1 ( ECE1 ) gene is presented in Fig. 3 . The MANE-select transcript for ECE1 is ENST00000374893, which accounts for 14.36% of expression in the GSE dataset. The Ref-Tx transcript is ENST00000415912, accounting for 59.41%. The CDS length of ENST00000374893 is longer than that of ENST00000415912 (2,313 vs 2,265). APPRIS also annotated ENST00000415912 as the APPRIS P:1 transcript. This example highlights challenges in gene annotation because the transcript with the longest CDS is not necessarily the most highly expressed transcript for a given protein-coding gene. For ENST00000374893 and ENST00000415912, alternative exon usage occurs in the final exon, as demonstrated in the Ensembl database (Supplementary Fig. 1). The peptide products generated from these alternative transcripts and their functional consequences require further examination. A similar pattern was observed for the DLC1 Rho GTPase-activating protein gene (data not shown). This gene has 18 transcript isoforms in the GTEx V8 dataset. The MANE-select transcript for DLC1 is ENST00000276297, accounting for 17.7% of expression in the GSE dataset, whereas the Ref-Tx transcript is ENST00000358919, accounting for 54.8%. The CDS length of ENST00000276297 is longer than that of ENST00000358919 (4,587 vs 3,276). APPRIS also identified ENST00000358919 as the APPRIS P:1 transcript. Discussion In our previous TEx-MST database, we used GTEx short-read and long-read expression datasets to display MANE-select transcript expression in many human tissues 10 , 27 . We observed that MANE-select transcripts exhibited lower agreement with the GTEx V8 short-read dataset and higher agreement with the GTEx V9 long-read dataset. However, because the GTEx V9 dataset has limited tissue coverage 28 , we extended our expression analyses by using long-read NGS datasets for additional tissue types. The GSE192955 dataset provides long-read expression data from 30 human tissues 18 , allowing more comprehensive assessment of expression profiles for human protein-coding genes. In the present study, we observed improved agreement between MANE-select transcripts and dominantly expressed Ref-Tx transcripts in the GSE192955 dataset. However, some MANE-select transcripts were matched in the GTEx V8 or V9 datasets but not in the GSE192955 dataset. Although MANE-select transcripts serve as harmonized reference transcripts for human protein-coding genes, they are not necessarily the dominantly expressed protein-coding transcripts. Because many protein-coding genes exhibit tissue-specific modulation of alternative transcript isoform expression, the topmost expressed protein-coding transcript may differ across tissue types. We therefore consider it necessary to examine tissue-specific expression profiles of alternative transcript isoforms to better understand the biological functions of different protein-coding transcripts. The eCPG database was developed to provide additional tissue expression information for human protein-coding genes. Alternatively spliced transcript isoforms can have distinct functional roles in various cell types. Studies have reported that the BTF3a isoform is more effective in mediating transcriptional activity than the BTF3b isoform, despite the BTF3a isoform being expressed at lower levels 29 , 30 . These two isoforms differ in their N-terminal sequences and overall protein lengths. The CDS length encoded by ENST00000335895 (BTF3b) is 489 bps, whereas the CDS length by ENST00000380591 (BTF3a) is 621 bps. The MANE-select transcript corresponds to ENST00000380591, which accounts for approximately 8% of expression in the GSE long-read dataset. By contrast, the dominant Ref-Tx transcript is ENST00000335895- BTF3b, accounting for 88.4% of expression. The protein-coding length of the ENST00000335895 (BTF3b) transcript is shorter, and this transcript is also annotated as APPRIS P:1. By contrast, BTF3a is transcriptionally active, whereas BTF3b is transcriptionally inactive because it lacks the first 44 amino acids at the N terminus. Modulation of these two isoforms may therefore have biological significance. Accordingly, careful examination of tissue-specific expression profiles of alternatively transcript isoforms and their biological activities remains crucial. In addition to complete exon inclusion or exclusion events, we also observed previously reported wobble splicing isoforms among protein-coding transcripts. Wobble splicing isoforms are typically generated through tandem splice site selection and can result in single or a small number of amino acid insertions or deletions, which may exist after nonsense-mediated mRNA decay during transcript maturation. In the present study, we identified 142 potential wobble splicing gene transcripts among the 2,015 genes in which Ref-Tx and MANE-select transcripts were mismatched. Additional subtle wobble splicing events are likely present in the 14,546 matched transcript groups. For example, two protein-coding transcripts of the ZPBP gene (ENST00000046087 and ENST00000419417) differ by only a single amino acid. Thus, our eCPG web tool was developed to support detailed examination of tissue-specific expression profiles and dominantly expressed transcript isoforms of human protein-coding genes. Conclusion We used the GSE192955, GTEx V8, and GTEx V9 expression datasets to construct a bioinformatic web database for visualizing the expression of protein-coding transcripts in various human tissue types. The bioinformatic web tool we created is useful for analyzing the tissue-specific expression patterns of dominantly expressed protein-coding transcripts. Declarations Fundings This work was supported in part by fundings from Academia Sinica and the National Science and Technology Council, Taiwan (113-2311-B-001-019-MY3). Author Contributions K.-F. Tung retrieved and processed the GSE192955, MANE, GENCODE, and GTEx datasets and constructed the eCPG web tool. W.-c. Lin supervised the study and prepared the manuscript. All authors reviewed the manuscript. Data Availability Protein-coding gene and transcript expression information can be accessed without restriction at the following link: https://ecpg.ibms.sinica.edu.tw/. Competing Interests The authors declare no competing interests. References Collins, F. S. Genome research: the next generation. Cold Spring Harb. Symp. Quant. Biol . 68 , 49–54 (2003). Mudge, J. M. & Harrow, J. The state of play in higher eukaryote gene annotation. Nat. Rev. Genet. 17 , 758–772. https://doi.org/10.1038/nrg.2016.119 (2016). Nurk, S. et al. The complete sequence of a human genome. Science 376 , 44–53. https://doi.org/10.1126/science.abj6987 (2022). Salzberg, S. L. Next-generation genome annotation: we still struggle to get it right. Genome Biol. 20 , 92. https://doi.org/10.1186/s13059-019-1715-2 (2019). Deveson, I. W., Hardwick, S. A., Mercer, T. R. & Mattick, J. S. The Dimensions, Dynamics, and Relevance of the Mammalian Noncoding Transcriptome. Trends Genet. 33 , 464–478. https://doi.org/10.1016/j.tig.2017.04.004 (2017). Pertea, M. et al. CHESS: a new human gene catalog curated from thousands of large-scale RNA sequencing experiments reveals extensive transcriptional noise. Genome Biol. 19 , 208. https://doi.org/10.1186/s13059-018-1590-2 (2018). Mudge, J. M. et al. GENCODE 2025: reference gene annotation for human and mouse. Nucleic Acids Res. 53 , D966–D975. https://doi.org/10.1093/nar/gkae1078 (2025). Gonzalez-Porta, M., Frankish, A., Rung, J., Harrow, J. & Brazma, A. Transcriptome analysis of human tissues and cell lines reveals one dominant transcript per gene. Genome Biol. 14 , R70. https://doi.org/10.1186/gb-2013-14-7-r70 (2013). Rodriguez, J. M., Pozo, F., di Domenico, T., Vazquez, J. & Tress, M. L. An analysis of tissue-specific alternative splicing at the protein level. PLoS Comput. Biol. 16 , e1008287. https://doi.org/10.1371/journal.pcbi.1008287 (2020). Tung, K. F., Pan, C. Y. & Lin, W. C. Dominant transcript expression profiles of human protein-coding genes interrogated with GTEx dataset. Sci. Rep. 12 , 6969. https://doi.org/10.1038/s41598-022-10619-9 (2022). Pozo, F., Rodriguez, J. M., Gomez, M., Vazquez, L., Tress, M. & J. & L. APPRIS principal isoforms and MANE Select transcripts define reference splice variants. Bioinformatics 38 , ii89–ii94. https://doi.org/10.1093/bioinformatics/btac473 (2022). Rodriguez, J. M. et al. APPRIS: selecting functionally important isoforms. Nucleic Acids Res. 50 , D54–D59. https://doi.org/10.1093/nar/gkab1058 (2022). Morales, J. et al. A joint NCBI and EMBL-EBI transcript set for clinical genomics and research. Nature 604 , 310–315. https://doi.org/10.1038/s41586-022-04558-8 (2022). Miller, R. M. et al. Enhanced protein isoform characterization through long-read proteogenomics. Genome Biol. 23 , 69. https://doi.org/10.1186/s13059-022-02624-y (2022). Tung, K. F. & Lin, W. C. TEx-MST: tissue expression profiles of MANE select transcripts. Database (Oxford) (2022). (2022) https://doi.org/10.1093/database/baac089 Patowary, A. et al. Developmental isoform diversity in the human neocortex informs neuropsychiatric risk mechanisms. Science 384 , eadh7688. https://doi.org/10.1126/science.adh7688 (2024). Dehghanian, F., Hojati, Z. & Kay, M. New Insights into VEGF-A Alternative Splicing: Key Regulatory Switching in the Pathological Process. Avicenna J. Med. Biotechnol. 6 , 192–199 (2014). Gao, Y. et al. Robust discovery and quantification of transcript isoforms from error-prone long-read RNA-seq data. Sci. Adv. 9 , eabq5072. https://doi.org/10.1126/sciadv.abq5072 (2023). Cunningham, F. et al. Ensembl 2019. Nucleic Acids Res. 47 , D745–D751. https://doi.org/10.1093/nar/gky1113 (2019). Consortium, G. T. The Genotype-Tissue Expression (GTEx) project. Nat. Genet. 45 , 580–585. https://doi.org/10.1038/ng.2653 (2013). Chan, W. C. et al. MetaMirClust: discovery of miRNA cluster patterns using a data-mining approach. Genomics 100 , 141–148. https://doi.org/10.1016/j.ygeno.2012.06.007 (2012). Ezkurdia, I. et al. Most highly expressed protein-coding genes have a single dominant isoform. J. Proteome Res. 14 , 1880–1887. https://doi.org/10.1021/pr501286b (2015). Tsai, K. W., Tarn, W. Y. & Lin, W. C. Wobble splicing reveals the role of the branch point sequence-to-NAGNAG region in 3' tandem splice site selection. Mol. Cell. Biol. 27 , 5835–5848. https://doi.org/10.1128/MCB.00363-07 (2007). Tsai, K. W., Tseng, H. C. & Lin, W. C. Two wobble-splicing events affect ING4 protein subnuclear localization and degradation. Exp. Cell. Res. 314 , 3130–3141. https://doi.org/10.1016/j.yexcr.2008.08.002 (2008). Tsai, K. W. & Lin, W. C. Quantitative analysis of wobble splicing indicates that it is not tissue specific. Genomics 88 , 855–864. https://doi.org/10.1016/j.ygeno.2006.07.004 (2006). Tsai, K. W., Chan, W. C., Hsu, C. N. & Lin, W. C. Sequence features involved in the mechanism of 3' splice junction wobbling. BMC Mol. Biol. 11 , 34. https://doi.org/10.1186/1471-2199-11-34 (2010). Tung, K. F., Pan, C. Y., Chen, C. H. & Lin, W. C. Top-ranked expressed gene transcripts of human protein-coding genes investigated with GTEx dataset. Sci. Rep. 10 , 16245. https://doi.org/10.1038/s41598-020-73081-5 (2020). Glinos, D. A. et al. Transcriptome variation in human tissues revealed by long-read sequencing. Nature 608 , 353–359. https://doi.org/10.1038/s41586-022-05035-y (2022). Thakur, D. et al. Human beta casein fragment (54–59) modulates M. bovis BCG survival and basic transcription factor 3 (BTF3) expression in THP-1 cell line. PLoS One . 7 , e45905. https://doi.org/10.1371/journal.pone.0045905 (2012). Zhang, Y. et al. BTF3 confers oncogenic activity in prostate cancer through transcriptional upregulation of Replication Factor C. Cell. Death Dis. 12 , 12. https://doi.org/10.1038/s41419-020-03348-2 (2021). Additional Declarations No competing interests reported. Supplementary Files SuppTables.zip SupplementaryFigure1.png Supplementary Figure 1. Exon structure of major ECE1 transcripts. Basic exon distribution information of the ECE1 protein-coding gene was retrieved from the Ensembl Genome Browser website. The MANE-select transcript (ENST00000374893) is highlighted. Cite Share Download PDF Status: Under Revision Version 1 posted Editorial decision: Revision requested 02 Apr, 2026 Reviews received at journal 01 Apr, 2026 Reviews received at journal 27 Mar, 2026 Reviews received at journal 27 Mar, 2026 Reviewers agreed at journal 19 Mar, 2026 Reviewers agreed at journal 18 Mar, 2026 Reviewers agreed at journal 13 Mar, 2026 Reviewers invited by journal 13 Mar, 2026 Editor invited by journal 11 Mar, 2026 Editor assigned by journal 01 Mar, 2026 Submission checks completed at journal 01 Mar, 2026 First submitted to journal 26 Feb, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8974155","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":606354399,"identity":"09452db3-5de9-46ec-a6d8-96a70f84fa90","order_by":0,"name":"Kuo-Feng Tung","email":"","orcid":"","institution":"Academia Sinica","correspondingAuthor":false,"prefix":"","firstName":"Kuo-Feng","middleName":"","lastName":"Tung","suffix":""},{"id":606354401,"identity":"a6f0c225-9e1f-4a31-808f-e642985a0b63","order_by":1,"name":"Wen-chang Lin","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA/0lEQVRIiWNgGAWjYHACxgMJDDY8BmC2AQMPmOYhoAeoJQ1TiwReLQwMhxkMUITwaTE43nvgwMMd52XMxY4/YPxRcEeGv/0A44O3bQx1BgdwaDlzLuFA4pnbPJazcwyYeQye8UicSWA2nNvGIIFLi9mNHIMDiW23eQxu5zAwMxgcBnoqgU2aF6jFDJeW+29AWs4BtaQDHQbSwv+A/TdeLTd4QFoOALUkAIMLpEUigY0Znxb7M2CHJYP9chikReLGw2bJOeckJPfj0CLZfsbw4c82O3tz6fSHD3/8OWzP35988MObMht+yQbsWlAA1FhGkFq8MTkKRsEoGAWjgAAAAPgRW0a3vuTsAAAAAElFTkSuQmCC","orcid":"","institution":"Academia Sinica","correspondingAuthor":true,"prefix":"","firstName":"Wen-chang","middleName":"","lastName":"Lin","suffix":""}],"badges":[],"createdAt":"2026-02-26 06:40:12","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8974155/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8974155/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":104747559,"identity":"b9cc4245-f441-4de4-a214-c1ac96a0e45b","added_by":"auto","created_at":"2026-03-16 18:17:55","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":40765,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eRank and expression profiles of 14,546 MANE-select-matched Ref-Tx.\u003c/strong\u003e (A) Rank distribution of 14,546 Ref-Tx matched with MANE-select transcripts. Gene numbers are depicted at the top of each column. Most Ref-Tx are Rank1 transcripts. (B) Distribution of expression percentages for the same 14,546 Ref-Tx matched with MANE-select transcripts. Ref-Tx numbers are depicted at the top of each column.\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-8974155/v1/3b73b63a02556bab286e1618.png"},{"id":104783310,"identity":"db8f7078-bc8d-499a-8811-653f841e54db","added_by":"auto","created_at":"2026-03-17 07:58:36","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":43767,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eVenn diagram illustrating overlap between MANE-select transcripts and Ref-Tx defined from the GTEx V8, GTEx V9, and GSE192955 datasets.\u003c/strong\u003e Ref-Tx was defined as the most highly expressed protein-coding transcript within each protein-coding gene. A total of 12,979 MANE-select transcripts matched Ref-Tx in the GTEx V8 short-read dataset, 13,372 matched in the GTEx V9 long-read dataset, and 14,539 matched in the GSE192955 long-read dataset. Among these, 10,116 MANE-select transcripts were matched with Ref-Tx in all three datasets.\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-8974155/v1/227cd2091ffa792170aca6a2.png"},{"id":104747561,"identity":"4b96be08-06c1-4e63-95b6-cc017dafecee","added_by":"auto","created_at":"2026-03-16 18:17:55","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":256362,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eExample of \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eECE1\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003egene information age in eCPG web tool. \u003c/strong\u003eBasic gene and transcript information for the \u003cem\u003eECE1\u003c/em\u003eprotein-coding gene is displayed on a single webpage. The main transcript expression data table is displayed. The Ref-Tx identified from the GTEx and GSE192955 datasets is marked with a red star, and the MANE-select transcript is marked with a black star. The expression percentage and transcript per million plots are provided for Rank1 to Rank5 transcripts of the \u003cem\u003eECE1\u003c/em\u003e gene. Links to the \u003cem\u003eECE1\u003c/em\u003e gene in the TEx-MST and TREGT databases are also provided.\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-8974155/v1/74eb34497841a2358123738a.png"},{"id":104808454,"identity":"7185de6a-4721-440a-9e3e-cd531af27611","added_by":"auto","created_at":"2026-03-17 12:37:43","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1251399,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8974155/v1/02603e99-3c13-438c-95e6-6f5e9d0663e6.pdf"},{"id":104747560,"identity":"1c3485b6-1198-4974-8ef7-d6d3669d55c6","added_by":"auto","created_at":"2026-03-16 18:17:55","extension":"zip","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":46899,"visible":true,"origin":"","legend":"","description":"","filename":"SuppTables.zip","url":"https://assets-eu.researchsquare.com/files/rs-8974155/v1/4b1ebcc5acbaf7f2a0b11f5a.zip"},{"id":104747563,"identity":"db5cd9dd-1726-4265-bc56-8b598ccb4679","added_by":"auto","created_at":"2026-03-16 18:17:55","extension":"png","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":167454,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSupplementary Figure 1. Exon structure of major \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eECE1\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003e transcripts. \u003c/strong\u003eBasic exon distribution information of the \u003cem\u003eECE1\u003c/em\u003e protein-coding gene was retrieved from the Ensembl Genome Browser website. The MANE-select transcript (ENST00000374893) is highlighted.\u003c/p\u003e","description":"","filename":"SupplementaryFigure1.png","url":"https://assets-eu.researchsquare.com/files/rs-8974155/v1/84f553cfdb5acc278727ecce.png"}],"financialInterests":"No competing interests reported.","formattedTitle":"Reference protein-coding transcripts of human genes annotated using long-read transcriptome datasets","fulltext":[{"header":"Introduction","content":"\u003cp\u003eThe completion of the Human Genome Project provided a major breakthrough and opportunity to comprehensively understand human gene composition and biological modulation. Despite considerable bioinformatic analyses of human genome sequences, full annotation of human protein-coding genes remains an ongoing and challenging task \u003csup\u003e\u003cspan additionalcitationids=\"CR2 CR3\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e. Gene annotation in eukaryotic genomes is challenging in part because of the presence of alternatively transcribed mRNA transcripts generated from multiple exon usages \u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e. This challenge has been further amplified with the greatly expanded availability of next-generation sequencing (NGS) transcriptome datasets \u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e,\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e. Massive short-read NGS datasets cannot completely resolve alternatively spliced mRNA transcript isoforms and may further generate unnecessary ambiguous transcripts, particularly in experiments with low-quality library preparations. Consequently, this problem has become more pronounced with increasing accumulation of short-read datasets. Current estimates indicate that more than 8.76 transcripts are expressed per human protein-coding gene \u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e. For example, in the GENCODE v47 annotation, 170,270 mRNA transcripts are annotated for 19,433 human protein-coding genes, and more than 52% of the registered transcripts (89,832) are designated as protein-coding transcripts on the basis of the GENCODE biotype feature \u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e. This finding implies that not all mRNA transcripts derived from protein-coding genes could be used for peptide translation.\u003c/p\u003e \u003cp\u003eFor protein-coding genes, understanding the expression profiles of alternative transcript isoforms is necessary to identify the peptides that are ultimately translated. This is crucial because the biochemical and biological functions of most protein-coding genes depend on their translated protein peptides. Although alternative splicing can produce multiple protein isoforms, most human protein-coding genes generate dominant protein-coding transcripts and peptides \u003csup\u003e\u003cspan additionalcitationids=\"CR9\" citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e. Thus, identifying the dominantly expressed protein-coding transcript in each protein-coding gene is crucial for subsequent analyses of biochemical function based on the final translated peptide product. The APPRIS database is a widely used bioinformatic resource for this purpose \u003csup\u003e\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e. This database assigns a single reference sequence, termed the principal isoform, to each gene by integrating reliable protein structural and functional information as well as cross-species conservation data \u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eRecently, another major annotation reference for human protein-coding genes was established. The MANE project (Matched Annotation from the NCBI and EMBL-EBI) aims to provide consistent transcript annotations by designating a single, well-curated representative transcript for each human protein-coding gene \u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e. The MANE-select transcript serves as the standardized transcript for every human protein-coding gene in both the RefSeq and Ensembl databases. On average, human protein-coding genes have eight transcript isoforms \u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e. For each protein-coding gene, the MANE-select transcript is the only chosen transcript and may correspond to the dominant protein-coding transcript \u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e. A study reported substantial agreement between APPRIS principal isoforms and MANE-select transcripts with major proteomic peptides \u003csup\u003e\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e. By contrast, short-read NGS datasets cannot accurately reflect the true expression status of alternatively spliced transcript isoforms \u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eWe previously demonstrated that long-read NGS datasets provide a better understanding of the expression profiles of transcript isoforms among various tissue types \u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e. For some protein-coding genes, specific transcript isoforms are differentially expressed and used for specific biological functions during development \u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e,\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e,\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e. For example, the biological functions of the \u003cem\u003eVEGF\u003c/em\u003e gene are modulated by the balance between alternatively spliced proangiogenic and antiangiogenic variants during tissue development \u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e. Thus, examining the tissue-specific expression of dominant protein-coding transcripts is crucial for functional studies of protein-coding genes. In this study, we extend our analysis of tissue expression profiles of major protein-coding transcripts by using additional long-read datasets. Gao et al generated useful long-read nanopore RNA sequencing (RNA-Seq) datasets from 30 normal human tissues, including two fetal brain samples. They produced more than one billion RNA-Seq reads that were analyzed using an improved bioinformatic pipeline \u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e. This dataset also included some major tissue types that were not covered in the GTEx V9 long-read dataset used in our previous study \u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e. In addition, visualization of alternative transcript isoform expression among different tissues can help understand patterns of functional peptide expression and modulation. In this study, we developed a web-based bioinformatic tool to identify dominant protein-coding transcripts by using updated MANE annotations and long-read transcriptome datasets that include additional tissue types.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eMANE and APPRIS datasets\u003c/h2\u003e \u003cp\u003eThe MANE human protein-coding gene dataset was retrieved from the NCBI website (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/refseq/MANE/\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/refseq/MANE/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). The file used in this study was MANE.GRCh38.v1.4.summary.txt. MANE-select transcript information was extracted using the MANE_status feature column and other annotation features, including gene symbol, gene name, chromosomal location, Ensembl gene ID, Ensembl transcript ID, Ensembl protein ID, NCBI gene ID, NCBI RefSeq NM ID, and RefSeq NP ID. The dataset contains 19,388 MANE-select transcript records for human protein-coding genes and 66 MANE-plus-clinical transcripts. In addition to MANE-select transcripts, MANE-plus-clinical transcripts are annotated to report clinically significant variants in specific protein-coding genes. Ensembl Gene ID was used as the primary key feature for all subsequent comparisons and analyses across datasets.\u003c/p\u003e \u003cp\u003eThe APPRIS database is another crucial resource that assigns a single coding sequence isoform as the principal isoform for each gene based on protein features, including evolutionary conservation \u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e. Principal isoform scores range from 1 to 5, with 1 indicating the highest reliability \u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e. We directly retrieved APPRIS score data from the APPRIS website (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://appris.bioinfo.cnio.es/#/downloads\u003c/span\u003e\u003cspan address=\"https://appris.bioinfo.cnio.es/#/downloads\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). Because the GSE192955 dataset used the GENCODE v34 annotation file, we retrieved the GENCODE34/Ensembl100 Principal Isoforms.txt file for this study.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eGSE192955 long-read datasets covering major human tissues\u003c/h3\u003e\n\u003cp\u003eESPRESSO is a bioinformatic tool used for the quantification of transcript isoforms in long-read NGS datasets \u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e. The GSE192955 dataset provides nanopore RNA-seq data from 30 human tissues. These include 10 brain subtypes, two fetal tissues, and 18 adult tissues. The brain subtypes comprise the caudate nucleus, cerebellum, cerebral cortex, corpus callosum, frontal lobe, hippocampus, medulla oblongata, pons, temporal lobe, and thalamus. The two fetal tissues are central nervous system tissues, namely the fetal brain and fetal spinal cord. The adult tissues include the bladder, blood, brain, colon, heart, kidney, liver, lung, ovary, pancreas, prostate, skeletal muscle, small intestine, spinal cord, spleen, stomach, testis, and thyroid. Inclusion of these tissues allows more comprehensive assessment of expression profiles for human protein-coding genes. When compared with tissues included in the Genotype-Tissue Expression (GTEx) V9 long-read dataset, only six major tissues overlap, namely the brain, heart, liver, lung, pancreas, and skeletal muscle. The GTEx project used the GENCODE v26 dataset as the reference annotation, whereas the GSE192955 dataset was annotated using GENCODE v34 data. Then, the gencode.v26.basic.annotation.gff3 and gencode.v34.basic.annotation.gff3 files were retrieved from the GENCODE project for the present study \u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\n\u003ch3\u003eGTEx short-read (V8) and long-read (V9) datasets\u003c/h3\u003e\n\u003cp\u003eThe GTEx Project is a major resource for studying genotypes and gene expression \u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e and is supported by the Common Fund of the Office of the Director of the National Institutes of Health. All GTEx data used in the present study contain no participant data and comply with the NIH Genomic Data Sharing guideline. We directly downloaded both V8 and V9 normalized transcript expression datasets from the GTEx Portal (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.gtexportal.org/home/downloads/adult-gtex\u003c/span\u003e\u003cspan address=\"https://www.gtexportal.org/home/downloads/adult-gtex\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). The data files used were GTEx_Analysis_2017-06-05_v8_RSEMv1.3.0_transcript_tpm for the V8 dataset and quantification_gencode_tpm for the V9 dataset. The retrieved GTEx datasets were processed using Python scripts to separate transcript expression data by tissue subtype. Each tissue type was assigned a separate transcript expression file. For each tissue, average expression values were calculated for individual transcripts, and the expression data from all tissue types were then combined. The GTEx V8 dataset covered 54 tissue types from 948 donors, whereas the V9 dataset covered only 14 tissue types. The tissues included in the V9 dataset were the anterior cingulate cortex, caudate, cerebellar hemisphere, frontal cortex, putamen, adipose subcutaneous tissue, breast mammary tissue, heart atrial appendage, heart left ventricle, liver, lung, pancreas, skeletal muscle, and fibroblast cell lines.\u003c/p\u003e\n\u003ch3\u003eRef-Tx: Topmost reference protein-coding transcripts\u003c/h3\u003e\n\u003cp\u003eWe identified the dominantly expressed protein-coding transcript for each human protein-coding gene. For each gene, we first calculated the average expression values of all annotated transcripts and ranked them by expression level. Because not all transcripts are protein-coding, as indicated by the GENCODE biotype annotation, only transcripts annotated as protein-coding were considered. The highest-expressed protein-coding transcript for each gene was designated as the Ref-Tx. In the present study, only one Ref-Tx was assigned to each protein-coding gene. These Ref-Tx transcripts were then compared with MANE-select and GTEx transcripts.\u003c/p\u003e\n\u003ch3\u003eConstruction of the eCPG web database\u003c/h3\u003e\n\u003cp\u003eAn expression catalogue of protein-coding genes (eCPG) web database was hosted in a Docker-based web environment running on an Ubuntu Linux server. The eCPG database was implemented using PHP within an Apache web server framework and integrated with a MySQL database \u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e,\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e. The JavaScript D3 library was used for interactive visualization of transcript expression levels. Transcript expression data for protein-coding genes from the GTEx and GSE192955 datasets were stored as flat files and then imported into the MySQL database for use by the eCPG web interface. The eCPG web database is freely accessible at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ecpg.ibms.sinica.edu.tw/\u003c/span\u003e\u003cspan address=\"https://ecpg.ibms.sinica.edu.tw/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003cp\u003eFigure illustration was done by using the GraphPad Prism (version 10) software package.\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003eSummary of protein-coding genes in the long-read NGS dataset\u003c/h2\u003e \u003cp\u003eTo examine additional tissue expression profiles of human protein-coding genes, we analyzed the GSE192955 dataset, which includes 30 major tissue types that differ from those in the GTEx V9 dataset \u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e. After the exclusion of nonprotein-coding genes based on GENCODE biotype annotations, the dataset contained 105,281 expressed transcripts corresponding to 18,557 protein-coding genes. On average, this long-read NGS dataset identified 5.67 expressed transcripts per human protein-coding gene. The number of expressed transcripts per protein-coding genes was lower than that observed in the GTEx V8 short-read dataset (7.43 transcripts per gene), likely indicating differences in sequencing depth and tissue coverage. We focused on these 105,281 transcripts. Using GENCODE v34 biotype annotations, we classified 60,090 transcripts (57.07%) as protein-coding (Supplementary Table\u0026nbsp;1). By comparison, the original GENCODE v34 annotation file used in the pipeline contained 19,959 protein-coding genes with a total of 153,435 transcripts, including 84,068 protein-coding transcripts. Thus, approximately 92.9% of GENCODE annotated protein-coding genes and 68.6% of transcripts associated with protein-coding genes were expressed in this long-read dataset. When only protein-coding transcripts were examined, approximately 70% exhibited detectable expression. However, these transcripts were not expressed equally.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eMost highly expressed protein-coding transcripts\u003c/h3\u003e\n\u003cp\u003eWhen we examined the distribution of expression among ranked transcripts in the long-read dataset, we observed a clear dominance of the top-ranked transcript in each gene, consistent with findings of other studies \u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e,\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e. As presented in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, for the most highly expressed Rank1 transcripts (18,557 transcripts), the average transcript expression level in protein-coding genes was 79.81%. By contrast, the average transcript expression level of Rank2 transcripts (15,537 transcripts) was 17.87%. The average expression level of Rank1 transcripts in the short-read GTEx dataset was significantly lower (66.41%) (Supplementary Table\u0026nbsp;2). This difference supports the advantage of long-read NGS platforms in more accurately resolving dominant transcript expression patterns. In short-read datasets, protein-coding genes generally exhibit higher expression levels in Rank2 through Rank10 transcripts (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e and Supplementary Table\u0026nbsp;2).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eNumbers of transcripts and average expression levels of Rank1 and Rank10 expressed transcripts of protein-coding genes within GSE 192955 dataset. There is a total of 105,281 transcripts belonging to 18,557 protein-coding genes.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTranscript Rank\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNumbers of Ranked transcripts\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAverage expression percentage\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eNumbers of protein-coding transcripts\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eProtein-coding transcript / all Ranked transcripts (%)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003e1\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e18,557\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e79.81\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e17,282\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e93.12\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003e2\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e15,537\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e17.87\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e9,613\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e61.87\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003e3\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e13,168\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e6.91\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e6,867\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e52.14\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003e4\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e11,061\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.42\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e5,311\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e48.01\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003e5\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e9,139\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2.01\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e4,182\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e45.75\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003e6\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e7,471\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.26\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e3,413\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e45.68\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003e7\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6,015\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.84\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e2,687\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e44.67\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003e8\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e4,807\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.62\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e2,074\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e43.14\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003e9\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e3,848\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.47\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1,688\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e43.86\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003e10\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e3,040\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.38\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1,326\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e43.61\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe number of protein-coding transcripts was also the highest among Rank1 transcripts. In the long-read dataset, 17,282 Rank1 transcripts (93.12%) were annotated as protein-coding, whereas in the short-read GTEx dataset, 17,472 Rank1 transcripts (89.18%) were annotated as protein-coding. In other ranked transcripts (Rank2 to Rank10), the proportion of protein-coding transcripts was similar between the two datasets (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e and Supplementary Table\u0026nbsp;2). These findings further support the conclusion that protein-coding genes typically produce a dominant peptide product, consistent with the findings of a previous study \u003csup\u003e\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e. However, not all top-ranked transcripts were protein-coding because approximately 6.9% of Rank1 transcripts were noncoding.\u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eRef-Tx defined by most highly expressed protein-coding transcripts\u003c/h2\u003e \u003cp\u003eWe focused on the most highly expressed dominant protein-coding transcripts and annotated the top expressed protein-coding transcript for each gene as the Ref-Tx transcript. Similar to the MANE project, only one Ref-Tx was assigned to each protein-coding gene on the basis of transcript isoform expression percentage and transcript biotype information. Among the 18,557 protein-coding genes in the long-read dataset, we identified 18,094 Ref-Tx transcripts. This finding indicated that 463 genes did not have detectable protein-coding transcripts expression in the long-read dataset. This absence may result from limitations in sequencing depth or tissue-specific expression because some genes may be expressed only in tissue subtypes not included in this dataset.\u003c/p\u003e \u003cp\u003eWe compared the 18,094 Ref-Tx transcripts with the 19,338 MANE-select transcripts and identified 14,546 matched transcripts. Among these matched transcripts, 14,091 (96.8%) were Rank1 transcripts. By contrast, in the GTEx V8 short-read dataset, the matched gene proportion was only 67.6% (11,820 of 17,472). This comparison further supports the advantage of long-read sequencing for identifying dominant protein-coding transcripts.\u003c/p\u003e \u003cp\u003eWe then examined the 14,546 matched MANE-select transcripts. As expected, most Ref-Tx were Rank1 protein-coding transcripts (14,091 transcripts; Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e1\u003c/span\u003eA). A large proportion of Ref-Tx transcript (9,290 transcripts) accounted for more than 80% of total gene expression levels (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e1\u003c/span\u003eB), indicating strong dominance. Notably, seven MANE-select transcripts had modified Ensembl gene IDs in the updated MANE annotations. After accounting for these updates, 14,539 MANE-select transcripts were matched with Ref-Tx in the GSE192955 dataset. Compared with Ref-Tx defined from the GTEx V8 and V9 datasets, the GSE192955 long-read dataset exhibited stronger concordance with MANE project genes (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e2\u003c/span\u003e). Consistent with the findings of previous study \u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e, long-read datasets (GSE192955 and GTEx V9) were determined to be more advantageous for protein-coding transcript isoform expression studies.\u003c/p\u003e \u003cp\u003e \u003cb\u003e2,015 MANE-select transcripts differ from Ref-Tx transcripts\u003c/b\u003e \u003c/p\u003e \u003cp\u003eNotably, the long-read dataset contained 16,561 matched MANE-select transcript ID records. Thus, 2,015 MANE-select transcripts present in the GSE dataset did not qualify as top-ranked protein-coding transcripts under our Ref-Tx selection criteria. We further examined these 2,015 MANE-select transcripts that did not match the assigned Ref-Tx transcripts. As expected, none of them were Rank1 transcripts. Some discrepancies likely resulted from annotation differences between datasets. In this study, gene and transcript IDs were used as the primary features for data comparison. Thus, modified gene or transcript annotations may lead to mismatches for certain MANE genes or MANE-select transcripts. Furthermore, some MANE-select transcripts were annotated as processed transcripts under the GENCODE v26 transcript biotype classification and were thus excluded from our Ref-Tx assignment pipeline.\u003c/p\u003e \u003cp\u003eMore than 60% of the 2,015 MANE-select transcripts (1,268 records) were classified as Rank2 transcripts on the basis of expression percentage criteria. In addition, there were 386 Rank3 transcripts, 167 Rank4 transcripts, and 61 Rank5 transcripts. As expected, these protein-coding genes exhibited complex alternative splicing patterns, with an average of 8.39 transcripts per gene. When Rank2 transcripts were excluded, the average number of alternative splicing isoforms increased to 10.38 per protein-coding gene, which is nearly double the overall mean transcript count in this long-read dataset (5.67 transcripts per gene). These findings indicate the challenges associated with annotating complex alternative splicing isoforms in protein-coding genes.\u003c/p\u003e \u003cp\u003eBecause of alternative splicing, some transcript isoforms differ only in their untranslated regions (UTRs) and produce identical protein peptide products. In such cases, we selected the transcript with higher expression as the Ref-Tx transcript. Among these 2,015 protein-coding genes, 561 genes had identical coding sequence lengths between the MANE-select transcript and the corresponding Ref-Tx transcript based on GENCODE annotation. Although these transcripts shared the same open reading frame, their 5\u0026prime; or 3\u0026prime; UTR regions differed because of alternative promoter usage or splicing events. Notably, some of these transcript pairs exhibited substantial differences in expression levels between the MANE-select transcripts and Ref-Tx transcripts. This observation may be relevant for studies of transcriptional regulation, such as tissue-specific or developmental stage\u0026ndash;specific promoter usage, even though the translated protein products are the same. These 2,015 protein-coding genes are listed in the eCPG web tool to facilitate user exploration, particularly for examining tissue-specific expression patterns.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eWobble splicing transcripts with subtle splicing site selection\u003c/h2\u003e \u003cp\u003eBesides the 561 genes, there were 222 Ref-Tx transcripts with longer CDS region length, whereas 1,232 MANE-select transcripts had longer CDS region length. In many cases, transcripts with longer coding regions are regarded as representative full-length products. However, the biological functions of individual alternatively spliced transcript isoforms require further investigation. Among the 1,232 and 222 protein-coding genes in which Ref-Tx and MANE-select transcripts differed in CDS region length, we observed features consistent with the wobble splicing phenomenon reported previously \u003csup\u003e\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e,\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e. Wobble splicing is characterized by the use of alternative tandem splice sites, involving proximal donor or acceptor sites, rather than complete exon inclusion or deletion \u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e. This mechanism often generates isoforms with single or small numbers of amino acid insertions or deletions \u003csup\u003e\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/sup\u003e. Our laboratory previously demonstrated that one splicing isoform of the \u003cem\u003eING4\u003c/em\u003e gene alters protein subnuclear localization and degradation \u003csup\u003e\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e. When possible wobble spliced transcripts were restricted to those with CDS region differences within 20 base pairs between MANE-select and Ref-Tx transcripts, 142 protein-coding genes were identified with such small amino acid insertions or deletions events, including e.g. \u003cem\u003eFGF14\u003c/em\u003e and \u003cem\u003eRNF19B\u003c/em\u003e genes.\u003c/p\u003e \u003cp\u003eNot all protein-coding genes produce a single dominant protein peptide; therefore, variations in expression among different transcript isoforms must be examined. Some genes modulate their biological functions through the expression of alternatively spliced transcripts in specific tissue types or developmental stages. Accordingly, we developed the eCPG bioinformatic web tool to provide visualization functions that facilitate examination of such transcript isoforms among different tissue types (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://eCPG.ibms.sinica.edu.tw\u003c/span\u003e\u003cspan address=\"https://eCPG.ibms.sinica.edu.tw\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e ). These 2,015 protein-coding genes identified in this study are specifically listed for further exploration by users.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eComparison of APPRIS annotations for 2,015 MANE-select transcripts\u003c/h2\u003e \u003cp\u003eAPPRIS is a database that provides information on protein structures and evolutionary cross-species conservation for protein-coding genes \u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e. APPRIS assigns annotations to alternatively spliced transcripts for many model organisms. An APPRIS Principal 1 (P:1) score indicates that a transcript is considered the representative isoform based on the core APPRIS computational modules. Among the 14,546 transcripts matched between Ref-Tx and MANE-select transcripts, a high proportion of transcripts were annotated as APPRIS P:1 (83.5%). We further examined APPRIS annotations for the 2,015 MANE-select transcripts that did not meet the Ref-Tx criteria. Among these 2,015 genes, 513 genes had both Ref-Tx and MANE-select transcripts annotated as APPRIS P:1. In total, 1,098 Ref-Tx transcript and 640 MANE-select transcripts within this group were annotated as APPRIS P:1. Score. These observations suggest that a large subset of Ref-Tx transcripts defined based on expression criteria may also have functional relevance and evolutionary conservation.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eeCPG web tool for expression visualization on protein-coding genes\u003c/h2\u003e \u003cp\u003eWe previously developed a bioinformatic tool, TEx-MST, to display MANE-select transcript expression using GTEx datasets as the single data source \u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e. In the present study, we primarily focused on the GSE192955 dataset and extended the previous webpage design to incorporate additional expression information from the GTEx V8, GTEx V9, and GSE192955 datasets. Users can examine transcript expression information for human protein-coding genes on a single gene-specific webpage (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e3\u003c/span\u003e). Basic gene information derived from the MANE project is displayed at the top of the page, followed by detailed transcript-level information. Because the GTEx V8 dataset contains a larger number of expressed transcripts, GTEx V8 data were used as the default source for constructing the transcript data table. Users may reorder and customize the display by clicking the ascending or descending icons at the top of each column. Annotation information from both GENCODE v26 and GENCODE v34, corresponding to the GTEx and ESPRESSO analysis pipelines, respectively, is also included. Transcript expression profiles from 30 tissues are presented in two graphical panels, exhibiting either expression percentages or TPM values. Within these panels, users can turn on/off individual transcript expression profiles by selecting the corresponding \u0026ldquo;RankX\u0026rdquo; label in the legend. In addition, links to the previously developed TEx-MST and TREGT web databases are provided to support further exploration of gene expression data \u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e,\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eExpression information for the endothelin-converting enzyme 1 (\u003cem\u003eECE1\u003c/em\u003e) gene is presented in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e3\u003c/span\u003e. The MANE-select transcript for \u003cem\u003eECE1\u003c/em\u003e is ENST00000374893, which accounts for 14.36% of expression in the GSE dataset. The Ref-Tx transcript is ENST00000415912, accounting for 59.41%. The CDS length of ENST00000374893 is longer than that of ENST00000415912 (2,313 vs 2,265). APPRIS also annotated ENST00000415912 as the APPRIS P:1 transcript. This example highlights challenges in gene annotation because the transcript with the longest CDS is not necessarily the most highly expressed transcript for a given protein-coding gene. For ENST00000374893 and ENST00000415912, alternative exon usage occurs in the final exon, as demonstrated in the Ensembl database (Supplementary Fig.\u0026nbsp;1). The peptide products generated from these alternative transcripts and their functional consequences require further examination. A similar pattern was observed for the \u003cem\u003eDLC1\u003c/em\u003e Rho GTPase-activating protein gene (data not shown). This gene has 18 transcript isoforms in the GTEx V8 dataset. The MANE-select transcript for \u003cem\u003eDLC1\u003c/em\u003e is ENST00000276297, accounting for 17.7% of expression in the GSE dataset, whereas the Ref-Tx transcript is ENST00000358919, accounting for 54.8%. The CDS length of ENST00000276297 is longer than that of ENST00000358919 (4,587 vs 3,276). APPRIS also identified ENST00000358919 as the APPRIS P:1 transcript.\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eIn our previous TEx-MST database, we used GTEx short-read and long-read expression datasets to display MANE-select transcript expression in many human tissues \u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e,\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u003c/sup\u003e. We observed that MANE-select transcripts exhibited lower agreement with the GTEx V8 short-read dataset and higher agreement with the GTEx V9 long-read dataset. However, because the GTEx V9 dataset has limited tissue coverage \u003csup\u003e\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e, we extended our expression analyses by using long-read NGS datasets for additional tissue types. The GSE192955 dataset provides long-read expression data from 30 human tissues \u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e, allowing more comprehensive assessment of expression profiles for human protein-coding genes. In the present study, we observed improved agreement between MANE-select transcripts and dominantly expressed Ref-Tx transcripts in the GSE192955 dataset. However, some MANE-select transcripts were matched in the GTEx V8 or V9 datasets but not in the GSE192955 dataset. Although MANE-select transcripts serve as harmonized reference transcripts for human protein-coding genes, they are not necessarily the dominantly expressed protein-coding transcripts. Because many protein-coding genes exhibit tissue-specific modulation of alternative transcript isoform expression, the topmost expressed protein-coding transcript may differ across tissue types. We therefore consider it necessary to examine tissue-specific expression profiles of alternative transcript isoforms to better understand the biological functions of different protein-coding transcripts. The eCPG database was developed to provide additional tissue expression information for human protein-coding genes.\u003c/p\u003e \u003cp\u003eAlternatively spliced transcript isoforms can have distinct functional roles in various cell types. Studies have reported that the BTF3a isoform is more effective in mediating transcriptional activity than the BTF3b isoform, despite the BTF3a isoform being expressed at lower levels \u003csup\u003e\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e, \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e. These two isoforms differ in their N-terminal sequences and overall protein lengths. The CDS length encoded by ENST00000335895 (BTF3b) is 489 bps, whereas the CDS length by ENST00000380591 (BTF3a) is 621 bps. The MANE-select transcript corresponds to ENST00000380591, which accounts for approximately 8% of expression in the GSE long-read dataset. By contrast, the dominant Ref-Tx transcript is ENST00000335895- BTF3b, accounting for 88.4% of expression. The protein-coding length of the ENST00000335895 (BTF3b) transcript is shorter, and this transcript is also annotated as APPRIS P:1. By contrast, BTF3a is transcriptionally active, whereas BTF3b is transcriptionally inactive because it lacks the first 44 amino acids at the N terminus. Modulation of these two isoforms may therefore have biological significance. Accordingly, careful examination of tissue-specific expression profiles of alternatively transcript isoforms and their biological activities remains crucial.\u003c/p\u003e \u003cp\u003eIn addition to complete exon inclusion or exclusion events, we also observed previously reported wobble splicing isoforms among protein-coding transcripts. Wobble splicing isoforms are typically generated through tandem splice site selection and can result in single or a small number of amino acid insertions or deletions, which may exist after nonsense-mediated mRNA decay during transcript maturation. In the present study, we identified 142 potential wobble splicing gene transcripts among the 2,015 genes in which Ref-Tx and MANE-select transcripts were mismatched. Additional subtle wobble splicing events are likely present in the 14,546 matched transcript groups. For example, two protein-coding transcripts of the \u003cem\u003eZPBP\u003c/em\u003e gene (ENST00000046087 and ENST00000419417) differ by only a single amino acid. Thus, our eCPG web tool was developed to support detailed examination of tissue-specific expression profiles and dominantly expressed transcript isoforms of human protein-coding genes.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eWe used the GSE192955, GTEx V8, and GTEx V9 expression datasets to construct a bioinformatic web database for visualizing the expression of protein-coding transcripts in various human tissue types. The bioinformatic web tool we created is useful for analyzing the tissue-specific expression patterns of dominantly expressed protein-coding transcripts.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eFundings\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was supported in part by fundings from Academia Sinica and the National Science and Technology Council, Taiwan (113-2311-B-001-019-MY3).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eK.-F. Tung retrieved and processed the GSE192955, MANE, GENCODE, and GTEx datasets and constructed the eCPG web tool. W.-c. Lin supervised the study and prepared the manuscript. All authors reviewed the manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eProtein-coding gene and transcript expression information can be accessed without restriction at the following link: https://ecpg.ibms.sinica.edu.tw/.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting Interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eCollins, F. S. Genome research: the next generation. \u003cem\u003eCold Spring Harb. Symp. Quant. Biol\u003c/em\u003e. \u003cb\u003e68\u003c/b\u003e, 49\u0026ndash;54 (2003).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMudge, J. M. \u0026amp; Harrow, J. The state of play in higher eukaryote gene annotation. \u003cem\u003eNat. Rev. Genet.\u003c/em\u003e \u003cb\u003e17\u003c/b\u003e, 758\u0026ndash;772. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/nrg.2016.119\u003c/span\u003e\u003cspan address=\"10.1038/nrg.2016.119\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNurk, S. et al. The complete sequence of a human genome. \u003cem\u003eScience\u003c/em\u003e \u003cb\u003e376\u003c/b\u003e, 44\u0026ndash;53. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1126/science.abj6987\u003c/span\u003e\u003cspan address=\"10.1126/science.abj6987\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSalzberg, S. L. Next-generation genome annotation: we still struggle to get it right. \u003cem\u003eGenome Biol.\u003c/em\u003e \u003cb\u003e20\u003c/b\u003e, 92. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s13059-019-1715-2\u003c/span\u003e\u003cspan address=\"10.1186/s13059-019-1715-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDeveson, I. W., Hardwick, S. A., Mercer, T. R. \u0026amp; Mattick, J. S. The Dimensions, Dynamics, and Relevance of the Mammalian Noncoding Transcriptome. \u003cem\u003eTrends Genet.\u003c/em\u003e \u003cb\u003e33\u003c/b\u003e, 464\u0026ndash;478. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.tig.2017.04.004\u003c/span\u003e\u003cspan address=\"10.1016/j.tig.2017.04.004\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2017).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePertea, M. et al. CHESS: a new human gene catalog curated from thousands of large-scale RNA sequencing experiments reveals extensive transcriptional noise. \u003cem\u003eGenome Biol.\u003c/em\u003e \u003cb\u003e19\u003c/b\u003e, 208. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s13059-018-1590-2\u003c/span\u003e\u003cspan address=\"10.1186/s13059-018-1590-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMudge, J. M. et al. GENCODE 2025: reference gene annotation for human and mouse. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cb\u003e53\u003c/b\u003e, D966\u0026ndash;D975. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1093/nar/gkae1078\u003c/span\u003e\u003cspan address=\"10.1093/nar/gkae1078\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGonzalez-Porta, M., Frankish, A., Rung, J., Harrow, J. \u0026amp; Brazma, A. Transcriptome analysis of human tissues and cell lines reveals one dominant transcript per gene. \u003cem\u003eGenome Biol.\u003c/em\u003e \u003cb\u003e14\u003c/b\u003e, R70. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/gb-2013-14-7-r70\u003c/span\u003e\u003cspan address=\"10.1186/gb-2013-14-7-r70\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2013).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRodriguez, J. M., Pozo, F., di Domenico, T., Vazquez, J. \u0026amp; Tress, M. L. An analysis of tissue-specific alternative splicing at the protein level. \u003cem\u003ePLoS Comput. Biol.\u003c/em\u003e \u003cb\u003e16\u003c/b\u003e, e1008287. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1371/journal.pcbi.1008287\u003c/span\u003e\u003cspan address=\"10.1371/journal.pcbi.1008287\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTung, K. F., Pan, C. Y. \u0026amp; Lin, W. C. Dominant transcript expression profiles of human protein-coding genes interrogated with GTEx dataset. \u003cem\u003eSci. Rep.\u003c/em\u003e \u003cb\u003e12\u003c/b\u003e, 6969. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41598-022-10619-9\u003c/span\u003e\u003cspan address=\"10.1038/s41598-022-10619-9\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePozo, F., Rodriguez, J. M., Gomez, M., Vazquez, L., Tress, M. \u0026amp; J. \u0026amp; L. APPRIS principal isoforms and MANE Select transcripts define reference splice variants. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cb\u003e38\u003c/b\u003e, ii89\u0026ndash;ii94. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1093/bioinformatics/btac473\u003c/span\u003e\u003cspan address=\"10.1093/bioinformatics/btac473\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRodriguez, J. M. et al. APPRIS: selecting functionally important isoforms. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cb\u003e50\u003c/b\u003e, D54\u0026ndash;D59. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1093/nar/gkab1058\u003c/span\u003e\u003cspan address=\"10.1093/nar/gkab1058\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMorales, J. et al. A joint NCBI and EMBL-EBI transcript set for clinical genomics and research. \u003cem\u003eNature\u003c/em\u003e \u003cb\u003e604\u003c/b\u003e, 310\u0026ndash;315. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41586-022-04558-8\u003c/span\u003e\u003cspan address=\"10.1038/s41586-022-04558-8\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMiller, R. M. et al. Enhanced protein isoform characterization through long-read proteogenomics. \u003cem\u003eGenome Biol.\u003c/em\u003e \u003cb\u003e23\u003c/b\u003e, 69. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s13059-022-02624-y\u003c/span\u003e\u003cspan address=\"10.1186/s13059-022-02624-y\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTung, K. F. \u0026amp; Lin, W. C. TEx-MST: tissue expression profiles of MANE select transcripts. \u003cem\u003eDatabase (Oxford)\u003c/em\u003e (2022). (2022) \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1093/database/baac089\u003c/span\u003e\u003cspan address=\"10.1093/database/baac089\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePatowary, A. et al. Developmental isoform diversity in the human neocortex informs neuropsychiatric risk mechanisms. \u003cem\u003eScience\u003c/em\u003e \u003cb\u003e384\u003c/b\u003e, eadh7688. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1126/science.adh7688\u003c/span\u003e\u003cspan address=\"10.1126/science.adh7688\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDehghanian, F., Hojati, Z. \u0026amp; Kay, M. New Insights into VEGF-A Alternative Splicing: Key Regulatory Switching in the Pathological Process. \u003cem\u003eAvicenna J. Med. Biotechnol.\u003c/em\u003e \u003cb\u003e6\u003c/b\u003e, 192\u0026ndash;199 (2014).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGao, Y. et al. Robust discovery and quantification of transcript isoforms from error-prone long-read RNA-seq data. \u003cem\u003eSci. Adv.\u003c/em\u003e \u003cb\u003e9\u003c/b\u003e, eabq5072. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1126/sciadv.abq5072\u003c/span\u003e\u003cspan address=\"10.1126/sciadv.abq5072\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCunningham, F. et al. Ensembl 2019. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cb\u003e47\u003c/b\u003e, D745\u0026ndash;D751. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1093/nar/gky1113\u003c/span\u003e\u003cspan address=\"10.1093/nar/gky1113\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eConsortium, G. T. The Genotype-Tissue Expression (GTEx) project. \u003cem\u003eNat. Genet.\u003c/em\u003e \u003cb\u003e45\u003c/b\u003e, 580\u0026ndash;585. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/ng.2653\u003c/span\u003e\u003cspan address=\"10.1038/ng.2653\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2013).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChan, W. C. et al. MetaMirClust: discovery of miRNA cluster patterns using a data-mining approach. \u003cem\u003eGenomics\u003c/em\u003e \u003cb\u003e100\u003c/b\u003e, 141\u0026ndash;148. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.ygeno.2012.06.007\u003c/span\u003e\u003cspan address=\"10.1016/j.ygeno.2012.06.007\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2012).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEzkurdia, I. et al. Most highly expressed protein-coding genes have a single dominant isoform. \u003cem\u003eJ. Proteome Res.\u003c/em\u003e \u003cb\u003e14\u003c/b\u003e, 1880\u0026ndash;1887. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1021/pr501286b\u003c/span\u003e\u003cspan address=\"10.1021/pr501286b\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2015).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTsai, K. W., Tarn, W. Y. \u0026amp; Lin, W. C. Wobble splicing reveals the role of the branch point sequence-to-NAGNAG region in 3' tandem splice site selection. \u003cem\u003eMol. Cell. Biol.\u003c/em\u003e \u003cb\u003e27\u003c/b\u003e, 5835\u0026ndash;5848. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1128/MCB.00363-07\u003c/span\u003e\u003cspan address=\"10.1128/MCB.00363-07\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2007).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTsai, K. W., Tseng, H. C. \u0026amp; Lin, W. C. Two wobble-splicing events affect ING4 protein subnuclear localization and degradation. \u003cem\u003eExp. Cell. Res.\u003c/em\u003e \u003cb\u003e314\u003c/b\u003e, 3130\u0026ndash;3141. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.yexcr.2008.08.002\u003c/span\u003e\u003cspan address=\"10.1016/j.yexcr.2008.08.002\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2008).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTsai, K. W. \u0026amp; Lin, W. C. Quantitative analysis of wobble splicing indicates that it is not tissue specific. \u003cem\u003eGenomics\u003c/em\u003e \u003cb\u003e88\u003c/b\u003e, 855\u0026ndash;864. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.ygeno.2006.07.004\u003c/span\u003e\u003cspan address=\"10.1016/j.ygeno.2006.07.004\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2006).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTsai, K. W., Chan, W. C., Hsu, C. N. \u0026amp; Lin, W. C. Sequence features involved in the mechanism of 3' splice junction wobbling. \u003cem\u003eBMC Mol. Biol.\u003c/em\u003e \u003cb\u003e11\u003c/b\u003e, 34. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/1471-2199-11-34\u003c/span\u003e\u003cspan address=\"10.1186/1471-2199-11-34\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2010).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTung, K. F., Pan, C. Y., Chen, C. H. \u0026amp; Lin, W. C. Top-ranked expressed gene transcripts of human protein-coding genes investigated with GTEx dataset. \u003cem\u003eSci. Rep.\u003c/em\u003e \u003cb\u003e10\u003c/b\u003e, 16245. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41598-020-73081-5\u003c/span\u003e\u003cspan address=\"10.1038/s41598-020-73081-5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGlinos, D. A. et al. Transcriptome variation in human tissues revealed by long-read sequencing. \u003cem\u003eNature\u003c/em\u003e \u003cb\u003e608\u003c/b\u003e, 353\u0026ndash;359. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41586-022-05035-y\u003c/span\u003e\u003cspan address=\"10.1038/s41586-022-05035-y\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThakur, D. et al. Human beta casein fragment (54\u0026ndash;59) modulates M. bovis BCG survival and basic transcription factor 3 (BTF3) expression in THP-1 cell line. \u003cem\u003ePLoS One\u003c/em\u003e. \u003cb\u003e7\u003c/b\u003e, e45905. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1371/journal.pone.0045905\u003c/span\u003e\u003cspan address=\"10.1371/journal.pone.0045905\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2012).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang, Y. et al. BTF3 confers oncogenic activity in prostate cancer through transcriptional upregulation of Replication Factor C. \u003cem\u003eCell. Death Dis.\u003c/em\u003e \u003cb\u003e12\u003c/b\u003e, 12. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41419-020-03348-2\u003c/span\u003e\u003cspan address=\"10.1038/s41419-020-03348-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2021).\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"long-read RNA-seq dataset, human protein-coding genes, dominant protein-coding transcripts, MANE-select transcripts, wobble splicing transcripts","lastPublishedDoi":"10.21203/rs.3.rs-8974155/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8974155/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eAccumulating NGS expression datasets suggest that protein-coding genes produce numerous alternatively spliced transcripts. However, this observation might be overestimated in short-read sequencing data, which often cannot accurately resolve distinct spliced isoforms and introduce ambiguity. Resolving tissue-specific expression profiles is crucial to identify bona fide translated peptide products. In this study, we identified the most highly expressed protein-coding transcripts by using long-read NGS datasets to better understand the biochemical and biological functions of human protein-coding genes. Using nanopore sequencing data from 30 normal human tissues in the GSE192955 dataset, we identified 18,094 dominantly expressed representative protein-coding transcripts (Ref-Tx) from 18,557 human genes. Comparison with MANE-select transcripts revealed that 14,546 Ref-Tx transcripts matched those in the MANE-select dataset. This result indicates improved agreement between long-read transcriptome data and MANE-select transcripts. A higher proportion of Rank1 transcripts were identified as Ref-Tx in the long-read dataset. Similar patterns were observed when Ref-Tx were compared with functional APPRIS annotations. Given the importance of tissue-specific expression profiles for protein-coding transcripts, we developed an expression visualization bioinformatic tool (eCPG). This webtool integrates the extensive expression information from 30 normal human tissues as well as from the GTEx project, which is designed to interrogate the dominant protein-coding transcripts.\u003c/p\u003e","manuscriptTitle":"Reference protein-coding transcripts of human genes annotated using long-read transcriptome datasets","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-16 18:17:50","doi":"10.21203/rs.3.rs-8974155/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-04-02T05:13:34+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-01T11:39:46+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-03-28T01:24:15+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-03-27T16:01:13+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"73073108084799519075384824941671997016","date":"2026-03-19T14:11:07+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"299595455677077074193131870676248223549","date":"2026-03-18T19:20:03+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"167831647772837040646580750673084075495","date":"2026-03-13T15:51:21+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-03-13T06:57:47+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-03-11T06:28:17+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-03-02T04:14:41+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-03-02T04:14:15+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2026-02-26T06:32:38+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"8a5f839f-e3b0-4589-ae4f-1036539fd31b","owner":[],"postedDate":"March 16th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"in-revision","subjectAreas":[{"id":64526035,"name":"Biological sciences/Biological techniques"},{"id":64526036,"name":"Biological sciences/Biotechnology"},{"id":64526037,"name":"Biological sciences/Computational biology and bioinformatics"},{"id":64526038,"name":"Biological sciences/Genetics"}],"tags":[],"updatedAt":"2026-04-02T05:24:09+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-16 18:17:50","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8974155","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8974155","identity":"rs-8974155","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.