{"paper_id":"008376d4-d292-46ed-920a-889b1b7958f0","body_text":"bioRxiv, 2026, pp. 1–13\ndoi: DOI HERE\nAdvance Access Publication Date: Day Month Year\nPaper\nPAPER\nDeveloping SCL2205 : A Protein Sequence-based\nSpatial Modelling Dataset for the Protein Language\nModel Frontier\nDaniel Ouso 1,2,∗ and Gianluca Pollastri 1,∗\n1School of Computer Science, University College Dublin, Dublin 4, Belfield, Dublin, Ireland and 2Centre for Research Training in\nGenomics Data Science, University of Galway, University Rd, Galway, H91 TK33, Ireland\n∗Corresponding author. ousodaniel@gmail.com\nFOR PUBLISHER ONLY Received on Date Month Year; revised on Date Month Year; accepted on Date Month Year\nAbstract\nDeep learning (DL) has advanced computational genome annotation tasks such as protein sub-cellular localisation (SCL)\nprediction. Nonetheless, its potential remains underutilised, primarily because of the limited availability of high-quality\nreference data and suboptimal input preparation strategies. In this study, we develop and analyse a high-quality dataset\nderived from the latest release of the universal protein knowledgebase (UniProtKB), designed to address existing challenges\nand support robust DL-based SCL modelling. The dataset was constructed through extensive quality preprocessing\nto ensure reliability, manual label mapping to enhance the quantity and diversity of the training data, and stringent\npartitioning to minimise data leakage. We validated the dataset using independent test sets, achieving up to 10.8%\nperformance improvement, measured by the area under the precision–recall curve (PR-AUC), compared to the state-\nof-the-art (SoTA). Furthermore, we highlighted potential performance metric inflation in existing SoTA predictors by\ndemonstrating, for the first time, at least 4.8% training-to-testing data leakage (pre-sequence representation) when using\nonly 10% of the training set under homology augmentation (augmentation based on sequence similarity database searches;\ndetails in Sub-section 2.1 ), a commonly used data augmentation strategy in DL-based SCL prediction modelling.\nSCL2205 will efficiently support the development of robust, trustworthy, and generalisable DL-based SCL predictors,\nwhile minimising data leakage and promoting reproducibility. It is openly available under the Creative Commons Zero\n(CC0 1.0) licence on DRYAD and is conveniently deployed as a package on the Python Package Index – p-scldata.\nKey words: Subcellular localisation, Machine learning, Protein Language Models, Data leakage, Data augmentation,\nSequence modelling\nIntroduction\nAutomated sequence annotation is a critical component of\nfunctional genomics. It is mainly driven by widespread\naccess to massively parallel sequencing. The adoption of\nArtificial intelligence (AI) for protein cellular component\nlocation annotation remains an active research area currently\ndominated by sequence-based DL predictors. Owing to its\nability to model complex phenomena and directly utilise\nsequence information, DL has become more popular than\nclassical machine learning (ML). The latter approach relies on\npre-determined sequence characteristics as modelling features.\nHowever, sufficient quality training data is the main bottleneck\nto efficiently leveraging supervised DL predictors, which are\nvery data-intensive. Regardless of the data challenge, it is\nparamount to provide high-quality training input.\nAlthough UniProtKB is the primary source of sequences\nfor SCL DL predictors, researchers often prepare their data\ndifferently, introducing avoidable bias. These biases may\ncompromise the fair and accurate evaluation of different\npredictors. Moreover, suboptimal handling of intrinsic yet\nproblematic data characteristics, such as sequence homology,\ncan lead to costly consequences, such as training-to-testing\ndata overlap. In addition, some predictors continue to rely on\noutdated database versions, despite significant improvements\nin data quality and quantity. This is likely associated with the\nrigorous and time-consuming nature of data development.\nNotable advancements in AI have been driven by well-\ncurated datasets across domains, including computer vision\n[19], natural language processing [20], and protein structure\nprediction – for example, the Critical Assessment of Structure\nPrediction (CASP) datasets. In the SCL domain, the DeepLoc\ndataset [3] is well known, alongside others such as MultiLoc\n[15]. In DeepLoc, sequences were filtered to include only\neukaryotic, non-fragment, nucleus-encoded proteins, with a\nminimum length of 41 amino acids and confirmed experimental\nannotations [3, 29]. In contrast, MultiLoc [15, 5] applies filters\n© The Author 2026. Published by Oxford University Press. All rights reserved. For permissions, please e-mail:\njournals.permissions@oup.com\n1\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted March 10, 2026. ; https://doi.org/10.64898/2026.03.08.710388doi: bioRxiv preprint \n\n2 Ouso et al.\nbased on eukaryotic origin and selected keywords from sequence\ncomments and features. Given the processing disparities,\nwhich make performance comparisons problematic, it would be\ndesirable to standardise and consistently curate the input data\nacross all predictors.\nAs concerns emerge on sustainability and the environmental\ncosts of training DL-based AI, the need for high-quality training\ninput is critical to mitigate data noise that inflates training\ncosts. Consequently, recent research has demonstrated gains\nwith small yet well-curated training datasets in leading domains\nof AI such as large language models (LLMs) [21]. On the other\nhand, trustworthiness and the precision of the performance\nreport in AI research are common concerns, especially in critical\nsectors such as biology and health. Similarly, compliance\nefforts are proposed for biological AI modelling [31]. In\nlight of the above, we set out to develop a robust SCL\ndataset for DL predictor modelling, whilst underscoring current\nchallenges and positioning for emerging frontiers. Beyond\ncommon preprocessing practices like filtering, we make the\nfollowing major contributions with our dataset:\n1.We highlight the limitations of homology augmentation by\naffirming and, for the first time, quantifying its contribution\nto data leakage using metrics comparable to those used in\nhomology reduction, thereby underscoring trustworthiness in\npredictors.\n2.Unlike in other cases, we perform extensive and robust\nhomology reduction by minimising training-to-validation and\ntraining-to-testing sets overlap to ≤ 30% homology, while\npreserving the sequence-length distribution of the original\ndataset.\n3.We provide two streams of the dataset – SCL2205 – for\nmodel training: (i) a training-validation split and (ii) a five-\nfold cross-validation set. In addition, a held-out independent\ntesting set is provided for final evaluation.\n4.We employ domain knowledge to manually map labels to\nhigher-level cellular components commonly used with SCL\npredictors.\nWe propose a novel comprehensive dataset ( SCL2205) for\napplication in SCL modelling, unlike the minimally described\nones in most SCL models. For reproducibility, we discuss in\ndetail the processes and design choices used to generate the\ndataset. This work establishes a new benchmark dataset that\nwill empower researchers to build more sustainable, precise, and\ntrustworthy tools for the next frontier of AI in spatial genomic\ndiscovery.\nMethodology\nTerminology\nSequence homology According to Medical Subject Headings\n(MeSH), it is the degree of similarity between sequences.\nThe similarity is attributed to descent from a common\nancestor and can be based on percentage sequence identity\nand/or percentage positive substitutions. While similarity\nwithin diverse sets is important in modelling, class over-\nrepresentation can be problematic. Mitigation is needed to\navoid over-weighting closely related sequences. In database\nparlance, sequence homology is commonly synonymous with\nredundancy; therefore, redundancy reduction is associated\nwith homology mitigation.\nHomology reduction It is the process of minimising\nsequence homology across or within datasets. In this study,\nwe distinguish across and within dataset homology reduction\nby as overlap and redundancy, respectively.\nData leakage Refers to data overlap across independent\ndatasets. When in excess, it is undesirable in model training,\nwhere sufficient separation is needed across the training-\nvalidation and training-testing sets. The separation serves\ntwo major purposes: (i) enabling accurate model performance\nassessment and (ii) ensuring model inferential robustness to\nunseen data.\nData augmentation Data augmentation increases the size\nand/or enriches the quality of available data, aiming\nto enhance model performance. It helps mitigate the\nproblem of minimal training data, especially for supervised\nlearning. The details of data augmentation are elaborately\ncovered elsewhere [26, 27, 10]. From a broad perspective,\nit may be categorised into three approaches: (i) input-\nlevel augmentation, which involves transformations applied\nin the data space; (ii) embedding-level augmentation,\nwhich operates in the feature space and targets intrinsic\nrelationships and semantics; and (iii) hybrid augmentation,\nwhich combines elements of both input- and embedding-\nlevel strategies. Alternatively, from the perspective of\nthe augmentation source, methods may be classified as\neither internal or external. Internal augmentation refers\nto modifying the existing data directly, for example,\nreordering elements within a sequence. In contrast, external\naugmentation involves incorporating additional information,\nsuch as generating sequence profiles through sequence\nalignments. Although protein SCL prediction borrows much\nfrom the text domain of sequential learning, it has unique\nrequirements and/or characteristics. Therefore, most of the\nnatural language processing (NLP) augmentation techniques\nare not (yet) implemented in the SCL modelling domain.\nHowever, since the latter is affected by imbalanced and\nlimited labelled data, other augmentation techniques are\napplied, mainly, homology augmentation, yet not without\ninherent limitations, especially since it counteracts homology\nreduction. Homology augmentation can be categorised as\nembedding augmentation, while label mapping falls under\ninput augmentation.\nHomology augmentation Homology augmentation refers\nto sequence data enhancements involving database searches\nto retrieve additional sequences related to the original\ndataset. The additional sequences can be directly included\nin the original set or used to build sequence profiles from\nmultiple sequence alignment (MSA). The latter is common\nwith DL SCL predictors.\nData Source\nEthical considerations\nNo ethical pre-approval was required. The raw data were\nretrieved from an open, public protein sequence database–\nUniProtKB, which does not contain any personal information.\nThe Universal Protein Knowledgebase\nThe UniProtKB is the central repository for aggregating\nfunctional information on proteins characterised by consistent,\naccurate and rich annotations. The annotations include\nuniversally accepted cross-references, ontologies, classifications,\nand quality scores informed by experimental and computational\nevidence [7]. The knowledge-base consists of two sections: (i)\nUniProtKB/Swiss-Prot for manually-annotated records based\non extracted information from literature and expert-evaluated\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted March 10, 2026. ; https://doi.org/10.64898/2026.03.08.710388doi: bioRxiv preprint \n\nSCL DL Prediction and Data Augmentation 3\ncomputational analysis, and (ii) UniProtKB/TrEMBL for\ncomputationally analysed records awaiting full manual\nannotation [7]. We used the former source because of\nthe quality assurance of its records. More than 95% of\nproteins in UniProtKB are derived from translations of coding\nsequences (CDS) submitted to the International Nucleotide\nSequence Database Collaboration (INSDC) databases [7]. In\nUniProtKB/Swiss-Prot, all protein products encoded by a gene\nin a given species are represented in a single record; that is,\nidentical sequences are captured as a single record. However,\nif the sequences are from different species, they form different\nrecords [7].\nProtein subcellular location labels\nMost SCL predictors are trained with supervision. The\nSubcellular location subsection of UniProtKB contains the\nprotein cellular location annotations used to supervise training.\nThe UniProtKB provides location and topology information for\nthe mature protein in the cell using a structured hierarchy\nof controlled vocabulary, except for the Note, which is free\ntext [7]. Because proteins may be of mitochondrial, plastid or\nnucleic origin, it is important to note that the label annotations\nrepresent the mature protein, that is fully processed and\nfunctional, resulting from post-translation modification (PTM)\nand proteolytic processing. It is the ultimate version associated\nwith a biological function. Therefore, filtering sequences on\ngenomic origin (mitochondria/plastid/nuclear) [3, 29] may be\nunnecessary in general predictor training.\nRaw data\nProtein sequences and their annotation information were\nmanually retrieved in tabular format from the latest release\nof the Reviewed subset of UniProtKB (UniProtKB/Swiss-\nProt, Release 2022\n05; date: 2023-01-24). Each record includes\na stable and unique Entry identifier, which is essential for\nprovenance and reproducibility.\nData inclusion criteria\nA total of 469,935 sequence records were downloaded. The\nfollowing five filtering processes were applied; data field name:\nprocess description:\n1.Sub-cellular location [CC] : Records lacking SCL annotation\nremoved.\n2.Sub-cellular location [CC] : Annotations with an evidence\ncode ontology (ECO) code of ECO:0000269 for the evidence\n– annotation is experimentally determined – were retained.\nThese annotations are a credible basis for supervisory\nlearning.\n3.Taxonomic lineage : Retained records belonging to the\neukaryotic taxonomic group using Eukaryota (super-\nkingdom) as the match phrase in the description. The\neukaryotic cell is very similar across species; therefore, we\nminimise data fragmentation across taxonomic Kingdoms by\naggregating at the super-kingdom level.\n4.Annotation: Annotations with a quality score of at least\nthree were retained. Although often overlooked, the cut-\noff enforces stringent reference data quality to minimise\ncompounding errors.\n5.Length: Sequences with at least 30 and at most 5,000\namino acids. Evidence shows biological function in shorter\nsequences; however, most SCL predictors cap sequence length\nfrom 30 to 40 amino acids.\nMost current SCL datasets stop at some of the above\nminimal preprocessing before proceeding to the homology\nreduction step. However, we harness more from the data than in\nearlier research. The DeepLoc dataset [3, 29] incorporates label\nmapping for various locations. However, it is minimal, based on\nan older database release and may contain some avoidable noise.\nFor instance, while organellar-membrane-localising proteins are\nmapped to their respective organelles, membranes, relative to\nthe organelles, are structurally and functionally related across\norganelles, in a general sense.\nTherefore, it might be better to consider them as a location\nentity, at least for a general SCL predictor, rather than\norganelle-specific. The latter approach is adopted in some\nmembrane protein predictors: when predicting endomembrane\nsystem and secretory pathway proteins [18] or membrane and\nnon-membrane proteins [17].\nLabel mapping\nIn the next section, we exploit the biological controlled\nvocabulary SCL definitions to map sequences of rare sub-\ncompartment SCL labels, which would otherwise be discarded,\nto their higher-order compartments; the labels/locations we are\ninterested in predicting. The intuition is simple: for a sequence\nto belong in a sub-space, it must belong in a space. Because\nMembrane is a unique entity, sub-membranes (membranes\nwithin a compartment) were lump-summed under Membrane.\nTherefore, the granularity, or the predicted location of interest,\nmatters more and is partly determined by the amount of\nrepresentation for that location.\nMapping began by first extracting and restructuring label\ninformation from the structured annotation text using custom\ncode. The resulting labels are concise, sorted lists containing\nonly the unique SCL annotation(s) for a sequence. Afterwards,\nmanual label mapping using UniProtKB’s SCL ontology\ninformation [7] and the interactive cell map from SwissBioPics\nlibrary [22 ] as reference ensued. We mainly considered rare\nlabels with five or more sequences for mapping.\nFor completeness, single-location and multi-location\nproteins were included in the final dataset ( SCL2205).\nHowever, for simplicity, only the Single-location proteins\nwere used for experiments. Fig 1 illustrates the process of\nlabel mapping. Ultimately, the process delivered a substantial\nincrease in sequence count (see T able 1), which is valuable for\ntraining.\nHomology reduction\nEssentially, homology reduction is common in protein modelling\nfor two reasons: to mitigate data leakage and avoid learning\nbias; it mainly relies on sequence alignment. The CD-HIT\ntool [11], which is based on position-specific iterative BLAST\n(PSI-BLAST), is widely used by DL predictors to facilitate\nhomology reduction. CD-HIT prefers longer sequences, despite\nshort and long sequences having functional roles. However, at\nleast two issues arise: unnatural input data distribution and\nloss of functional domain (pattern) representation that could\nbe useful for learning SCL. Consequently, we implemented\na custom variant sequence similarity algorithm, which is\nstill based on basic local alignment sequence tool (BLAST),\nfactors sequence-pair length, but without preference for\nlonger sequences, and conveniently operates on training-\ntesting split dataset models for ML; avoiding multi-stage\nhomology reduction regardless of the homology threshold. The\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted March 10, 2026. ; https://doi.org/10.64898/2026.03.08.710388doi: bioRxiv preprint \n\n4 Ouso et al.\nFig. 1. Illustrating SCL label mapping using four proteins .\nThe left and right panels represent single-location and multi-location\nprotein mapping, respectively. Step 1: programmatically, extract labels\nfrom structured annotation text using custom code. Step 2 : manually\nmaps sub-location (sub-compartment) to location labels and harmonises\ninterchangeable labels.\nAlt text : Graphics illustrating the mapping processes using a pair of\nsingle-localisation and multi-localisation proteins; involving mining labels\nand manual mapping.\nT able 1. A summary of the sub-cellular location (SL)\ncomposition showing the mapping effect. The first and\nsecond columns are the before and after protein counts,\nrespectively. The fold increase is captured in the last column.\nOverall, the sample size grew by 71%. When considering only\nthe single-location proteins, the numbers improved by 80%.\nFor brevity,Centrosome; Cytoplasm; Cytoskeleton; Microtubule\norganising center is shortened as Cent; Cytop; Cytos; MTOC.\nLabel Original Mapped\nF old\nIncrease\nMembrane 290 5616 19.37\nNucleus 3621 4721 1.30\nSecreted 3067 3357 1.09\nCytoplasm 2209 2706 1.22\nCytoplasm; Nucleus* 2167 2618 1.21\nMitochondrion 838 1062 1.27\nPlastid 6 622 103.67\nCent; Cytop; Cytos; MTOC* 92 338 3.67\nCytoplasm; Membrane* 88 284 3.23\nCytoplasm; Cytoskeleton* 244 270 1.11\nER 150 204 1.36\nCell projection 6 194 32.33\nPeroxisome 144 160 1.11\nT otals 12922 22152 1.71\n* Multiple-localisation.\nsimilarity algorithm is detailed in the Methods section of\nSupplementary File S1 .\nBased on the reasons already provided, we had the following\ngoals for our homology reduction pipeline (Fig 2):\n1.Mitigating data leakage – similar samples stay in a bin;\ntraining or testing, not both.\n2.Avoiding learning bias; imbalance, by implication –\nmaintain a particular similarity threshold within a bin.\nWe then proceeded with a three-step homology reduction\nstrategy based on similarity thresholds and dataset partitions.\nIn the first step, we minimise overweighting associated with\nimbalanced classes and high sequence homology during training\nby using an 80% similarity threshold to redundancy-reduce\nthe preprocessed dataset ( Fig 2 , PrpS0). The second and\nthird steps use 30% similarity threshold while considering the\npartitioned data. The testing set (TstS0 ; subject; 30%) was\noverlap-reduced against the training set ( TrnS0; query; 70%)\nand the testing set (subject and query) redundancy-reduced,\nrespectively. Whereas the second step mitigates data leakage,\nthe third step mitigates evaluation bias, which is associated\nwith imbalanced classes and sample similarity. The overlapping\nand redundant testing set sequences were isolated (binned; Fig\n2 broken arrows) in the training set to align with the goals.\nMoreover, we used the resulting training set to generate two\ndataset tracks for model training, based on a common model\ndevelopment data partitioning approach: training, validation,\nand testing; and cross-validation and testing. Henceforth,\nwe refer to the two tracks as training-validation-testing\n(TVT) and cross-validation-testing (CVT), respectively, and\ncollectively call them SCL2205. Note that the held-out\ntesting set is shared because the original dataset is the same.\nTherefore, onwards we refer to the partitions associated with\nthe model training process as training-validation (TV) and\ncross-validation (CV). As before, the TV and CV sets were\nsubjected to the second and third steps of homology reduction.\nThis was across five folds for the CV set. In the TV set, the\ntraining (Fig 2, TrnS3) and validation (Fig 2, ValS0) sets\nwere split at 80% and 20% respectively.\nNote. T\no ensure reproducibility, all data partitioning\nwas performed using the train test split function\nfrom\nthe scikit-learn library, with a consistent random\nseed of 2023 maintained throughout the study. Other\nsoftware details are in Supplementary File S1 .\nExperiments\nManual mapping impact\nWe aimed to assess the contribution of manual label mapping\non model generalisation, compared to using the native labels.\nUsing SCL2205s as the reference, we generated SCL2205nomap\ns\nby copying newDatas and removing all sequences associated\nwith the label-mapping augmentation process. Since very few\nsamples in the Cell projection and Plastid categories remained\nwhen considering the original labelling, we omitted them from\nthe modelling process. The corresponding partitions of the sets\nwere used to train two convolutional neural network (CNN)-\nbased networks as later outlined in Sub-subsection 3.3.2.\nThe independent test sets were redundancy-reduced per the two\ndatasets and used for evaluation.\nImplication of data leakage from homologous\naugmentation\nHow does homology augmentation contribute to and impact\nthe overlap between training-testing datasets, especially for\npredictors based on pre-model-training homology reduction?\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted March 10, 2026. ; https://doi.org/10.64898/2026.03.08.710388doi: bioRxiv preprint \n\nSCL DL Prediction and Data Augmentation 5\nFig. 2. Summary illustration of the three-step homology\nreduction process based on global pairwise alignments . Solid-\nline arrows indicate data processing flows, broken-line arrows represent\ndata isolation, and bifurcating arrows are data splits (split ratios are\nincluded at the nodes). Magenta, dashed rectangles denote redundancy\nreduction, while the black ones correspond to overlap reduction. The\nshaded rectangles highlight the repeating motif within the homology\nreduction pipeline. The process takes the 22,152 preprocessed ( PrpS0)\ninstances as input. It then begins from a global redundancy reduction that\nresults in 19,074 instances (( PrpS1; 100%), resulting in 15,183 instances\nin the training set (TrnS5; 79.60%), 1,260 in the validation (aka discovery\nor development) set (VldS2; 6.60%), and 2,631 as the held-out testing set\n(TstS2; 13.79%).\nAlt text: Graphic illustrating the three-step homology reduction process.\nRedundancy reduction occurs within datasets, while overlap reduction\noccurs across dataset partitions.\nTo attempt answering the question, we designed an\nexperiment using the TVT dataset. Whilst maintaining\nclass proportions, we randomly selected 10% (n = 1280;\nrandom-training-seqs-A) sequences of the training set ( Fig 2;\nTrnS5 minus multi-localising sequences; 12,807) sequences and\nqueried the RefSeq (Version 5 of 202410260537) database for\nhomologues using BLAST global alignment – as is often the\ncase. The tweaked PSI-BLAST parameters were: evalue=0.001;\nnum\niterations=3; matrix=BLOSUM62, word size=3 and\nmax target seqs=500. We retrieved the full homologous\nsequence set (designated HmgS1) using their parsed sequence\naccessions and then performed overlap reduction (using the 30%\nthreshold applied at the initial homology reduction process)\nagainst the combined validation and testing set (CVTS); using\nHmgS1 as the query. Subsequently, we used the proportion\nof CVTS sequences overlapping with the HmgS1 as a proxy\nfor the extent of overlap. That is, we divided the number\nof overlapping sequences by the total number of sequences\nwithin CVTS. The rationale is that using overlapping sequences\nto generate profiles propagates the overlap (leakage) to the\nresulting sequence encoding.\nAblation of augmenting sequences from homology\nreduction in leakage calculation\nTo nullify the contribution of the overlapping sequences\nisolated within the training set to leakage calculation, we\ncreated random-training-seqs-B by removing the augmenting\nsequences resulting from the homology reduction process\n(shown as broken arrows in Fig 2 ) from random-training-\nseqs-A. A total of 384 sequences were removed, with 896\nremaining ( random-training-seqs-B), which constituted 10%\nof the corresponding training set (the single-localising subset\nof TrnS5 with all augmenting sequences removed). We then\npruned HmgS1 of all hits associated with the 384 sequences.\nSubsequently, we performed overlap reduction on the test set\nusing the remaining portion – HmgS2. This allowed us to\nobserve the impact of these specific sequences on the extent\nof overlap – net leakage.\nDataset benchmarking\nTo benchmark the reliability of our newly developed SCL\ndataset ( SCL2205; n = 19,074), we assessed its performance\nagainst a comparable SoTA: DeepLoc2, redundancy-reduced at\nan 80% threshold (SwissProt train-validation ( DEEP-TV ); n\n= 22,126). Finally, we evaluated these using two independent\nsets: SwissProt sorting signal (DEEP-SS ) and human protein\natlas (DEEP-HPA) [29].\nDataset pre-processing\nSCL2205 and DEEP-TV were redundancy-reduced similarly.\nFor training, SCL2205 and DEEP-TV were respectively\nsplit into training (n = 10,680; 12,390, 56%), validation\n(n = 2,671; 3,098 [14%]) and testing (n = 5,723; 6,638,\n30%) partitions. The splitting was random and stratified on\nclass labels. All DEEP-TV sequences localising to the Golgi\napparatus were excluded because they had fewer than n = 100\ncounts, which was the per-class count threshold used with\nSCL2205. For brevity, experiments were performed using only\nthe single-label classes, which were intuitively coded SCL2205s\nand DEEP-TVs, respectively. See (T able 2) for the class\ndistribution across datasets and partitions.\nSince the DEEP-SS independent dataset was initially\nlabelled by sorting signals, it was therefore relabelled. We\nretrieved subcellular location labels from UniProtKB using\npersistent sequence accessions. Only sequences with matching\nUniProtKB accessions were considered. The human protein\natlas (HPA) dataset had labels corresponding to our training\ndata.\nTest modelling\nInput representation Representation was either tokenizer-\nor protein language model (PLM)-based, depending on\nthe testing model architecture – CNN or PLM embedding\nnetwork, and a common fully connected neural network\n(FCNN) classification network.\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted March 10, 2026. ; https://doi.org/10.64898/2026.03.08.710388doi: bioRxiv preprint \n\n6 Ouso et al.\nT able 2. Label-wise distribution across datasets and\npartitions (T rain [T rn], V alidation [Vld], and T esting\n[Tst]) for single-location labels only. There is a notable difference\nfor the Membrane class, resulting from the divergent binning\napproach adopted in the datasets. DEEP-TV has a higher\nnumber of training examples than SCL2205 in 60% of the\nclasses. For brevity: Membrane (MEM), Nucleus (NUC), Secreted\n(SEC), Cytoplasm (CYT), Mitochondrion (MIT), Plastid (PLA),\nEndoplasmic reticulum (ER), Cell projection (CEP), Peroxisome\n(PER) and Lysosome/vacuole (LYS/VAC).\nLabel SCL2205s DEEP-TVs\nT rn Vld Tst T rn Vld Tst\nMEM 2642 661 1415 1117 599 279\nNUC 2340 585 1254 2581 1383 646\nSEC 1475 369 791 1307 701 327\nCYT 1380 345 740 2073 1111 518\nMIT 530 132 284 617 331 154\nPLA 329 82 176 332 178 83\nER 104 26 55 131 70 33\nCEP† 98 25 53 NA NA NA\nPER 76 19 41 75 40 19\nLYS/VAC† NA NA NA 106 57 26\nT otals 8974 2244 4809 8339 4470 2085\nCoverage 56% 14% 30% 56% 14% 30%\n† Missing for the corresponding dataset and omitted\nduring dataset-comparison model training.\nArchitecture While the ad hoc CNN architecture is briefly\ndescribed in Supplementary File S1 , for PLM embedding,\nwe used an independent PLM – Rostlab/prot t5 xl uniref50\n– by the Rost Lab [9], using the Transformers Application\nProgramming Interface (API) [32] from Hugging Face.\nT raining Training was run until the model overfitted. That\nis, a sustained increase in (training-validation) loss after\nan initial decrease, as observed in the loss plot. For\ncomputational robustness (training was on a managed\ncomputing cluster), each model training phase was run three\ntimes, followed by weight averaging to instantiate the final\nmodel used for testing. The model snapshot with the best\nvalidation accuracy just before the sustained loss increase\nwas saved for each phase.\nEvaluation The test evaluations were compared using\nmultiple strategies, in isolation or as combinations:\nmacro and per-class PR-AUC metrics for ranking and\npositive retrievals, McNemar’s error rates test, and\nstratified bootstrapping on metrics differences uncertainty\nquantification. Evaluation testing was performed using\nindependent test sets DEEP-SS and DEEP-HPA, see T able\n3.\nSCL2205s versus DEEP-TVs comparison\nTo compare datasets as holistic stand-alone entities, we used\nthe corresponding partitions of SCL2205s and DEEP-TVs to\ntrain ad hoc CNN-based model; we also isolated model effects\nby training an independent PLM-based model. The models\nresulting from training with SCL2205s and DEEP-TVs; which\nfor brevity we will, respectively, call Model A and Model\nB, were then evaluated. The independent test datasets were\noverlap-reduced with each corresponding training set, see\ndetails in Supplementary File S1 for the resulting ablation\ncounts.\nT able 3. Label-wise counts for independent test datasets.\nLabel DEEP-SS DEEP-HPA\nSecreted† 522 NA\nMembrane 170 161\nMitochondrion 111 175\nPeroxisome 77 7\nNucleus 63 664\nCytoplasm 10 263\nER 5 31\nPlastid† 1 NA\nT otals 958 1301\n† Missing for the corresponding dataset.\nDEEP-SS and DEEP-HPA independent test sets\nWe highlight distinct characteristics of these two external\nevaluation sets. DEEP-SS is drawn from universal protein\n(UniProt)-SwissProt, which is also the source of the training\ndatasets; therefore, it can be considered an in-distribution\nset. In contrast, DEEP-HPA is based on SCL annotations\nfrom the HPA, which employs a different annotation strategy.\nSCL in HPA is diversely investigated; using isolated or\nintegrated techniques comprising of immunocytochemistry-\nimmunofluorescence (ICC-IF), confocal microscopy, and\nstaining [28]. Moreover, the distinction is highlighted by the\nextent of overlap with the SCL2205 and DEEP-TV training\nsets, as captured in various overlap-reduction instances, see\nT ables 2 and 3 in Supplementary File S1 . Furthermore,\nwhereas the UniProt-SwissProt data covers multiple taxa, HPA\ncovers only humans. Therefore, we consider DEEP-HPA as an\nout-of-distribution (OOD) set.\nResults\nMapping improved generalisation\nWe employed manual label mapping to increase target diversity\nand boost the number of training samples, to improve model\ngeneralisability and performance. A representative outcome of\nthe mapping process usingPlastid – the most impacted location\n– is illustrated in Fig 3. For clarity, we refer to the model using\nlabel mapping as Model A and the model using native labelling\nas Model B.\nWe observed a substantial improvement of the macro PR-\nAUC for DEEP-SS when using label mapping (an increase of\n9.0% points); however, the improvement for DEEP-HPA was\nmarginal (1.5% points). Full details, including per-class PR-\nAUC with their respective prevalences, are provided in the\nResults section of Supplementary File S1 . Overall, label\nmapping enhanced model performance across most classes;\nhowever, we observed a slight deterioration in the Cytoplasm\nand Secreted categories for DEEP-SS, and the Nucleus\ncategory for DEEP-HPA.\nAn asymptotic McNemar’s test with continuity correction\nrevealed a significant performance difference between Model A\nand Model B on both test sets. This test evaluates error rates\nat a fixed threshold. For DEEP-SS, χ2(d f = 1) = 8.2, p <\n0.004, with Model A outperforming Model B (b = 108, c =\n69; Effect Size Z = 2.9). In contrast, for DEEP-HPA, χ2(d f=\n1) = 278.8, p < 0.001, with Model B outperforming Model A\n(b = 610, c = 149; Effect Size Z = −16.7).\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted March 10, 2026. ; https://doi.org/10.64898/2026.03.08.710388doi: bioRxiv preprint \n\nSCL DL Prediction and Data Augmentation 7\nFig. 3. Frequencies for the Plastid label mapping augmentation\ninvolving nine locations. We reported the highest fold increase for\nPlastid (103.67 fold).\nAlt text: Graphics of a mapping case study involving Plastid, illustrating\na fold change of over 100.\nThe significant performance difference between Model A\nand Model B was supported by the stratified paired bootstrap\nconfidence interval (CI) test. However, whereas Model A still,\nas in the McNemar‘s test, outperformedModel B on the DEEP-\nSS set, the reverse was observed for the DEEP-HPA set in\nthis test, where the label mapping of Model A proved superior.\nThe stratified paired bootstrap CI test evaluates the ranking\nand retrieval of positives. Fig 4 illustrates the uncertainty\nlevels based on the macro PR-AUC differences between the two\nmodels.\nData leakage from homology augmentation\nWe show and quantify that homology augmentation can\npartially undo prior training-testing dataset homology\nreduction. Interestingly, we found that the 10% training set\nsample resulted in 4.8% training-testing sequence overlap.\nHmgS1 had 664,719 sequences, while HmgS2 had 477,870,\na 28.1% ( n = 186, 849) reduction of the initial homology-\naugmenting sequence hits. Consequently, there was a 22.9%\ndecrease in the number of sequences overlapping the test\nset between using HmgS1 (n = 201) and HmgS2 (n =\n155) as query. These results underscore the need for\ncareful consideration when implementing homology search\naugmentation of training data, which may inflate performance\nmetrics and/or hamper generalisation. They also provide a\nbaseline quantification, arguably the first such indication,\nfor the amount of overlap resulting from database search\naugmentation, on the same similarity measurement scale.\nFig. 4. Histogram of the distribution of 1000 bootstrap PR-\nAUC-metric differences between models A and B as trained on\nSCL2205 map\ns and SCL2205 nomap\ns , respectively. The top panel ( A) was\ntested on DEEP-SS (A > B ), whereas the bottom panel ( B) was tested\non DEEP-HPA (A > B ).\nAlt text: Graphics of a histogram illustrating the effect of label mapping\nusing two independent test datasets. Panel A (top) represents the in-\ndistribution UniProtKB test data, while Panel B (bottom) represent the\nout-of-distribution human protein atlas test data.\nA benchmark dataset for subcellular localisation\nprediction\nWe developed SCL2205 from the latest UniProtKB release\n(UniProtKB/Swiss-Prot: Release 2022\n05; 20230124), with\nextensive quality preprocessing to ensure reliability, manual-\nmapping augmentation to increase training data and improve\ngeneralisation, and stringent data partitioning to minimise\ndata leakage. While we observed a 1.71-fold increase in\ndata during the development process, the inherent imbalance\nacross data classes persists, albeit reduced. Therefore, we\nopted for the macro PR-AUC metrics to evaluate performance\nrelative to other datasets. The precision–recall curve (PRC)\ncurve evaluates recall (sensitivity/true positive rate (TPR))\nversus precision (specificity/positive predictive value (PPV)) at\nmultiple thresholds to highlight the trade-off between them. On\nthe other hand, the PR-AUC score summarises the PRC curve\nto a number. Unlike in area under the ROC curve (ROC-AUC),\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted March 10, 2026. ; https://doi.org/10.64898/2026.03.08.710388doi: bioRxiv preprint \n\n8 Ouso et al.\nT able 4. SCL2205 s and DEEP-TV s performance comparisons:\na summary of the macro PR-AUC across model variants and test\ndatasets. Despite using vanilla models, all scores exceeded the\nexpected random guessing level. We observed the most significant\ngains in the PLM-based model – up to 10.8% points in favour of\nSCL2205 s.\nDataset &\nArchitecture\nBaseline\nPrevalence\nSCL2205s\n8-class\nDEEP-TV s\n8-class\nDEEP-SS CNN 0.143* 0.257 0 .334\nDEEP-HPACNN 0.167* 0.194 0 .184\nDEEP-SS PLM 0.143* 0.561 0 .453\nDEEP-HPAPLM 0.167* 0.345 0 .359\n* We computed the average prevalence across classes for each\ntest set.\nwhere the binary baseline is 0.5, positive-class prevalence is the\nbaseline in PR-AUC.\nNext, we consider the results for the two model architectures\nused for dataset comparisons. T able 4 provides the global\nPR-AUC scores, while the per-class scores are summarised in\nSupplementary File S1 .\nCNN-based architecture\nAn asymptotic McNemar’s test with continuity correction\nshowed a significant difference between the models trained\non SCL2205 s (Model A) and DEEP-TV s (Model B ) when\nevaluated on the two test sets. For DEEP-SS, χ2(d f= 1) =\n18.1, p < 0.001; (B > A [b = 90, c = 158]) with an Effect\nSize Z = −4.3. DEEP-HPA: χ2(d f = 1) = 143.0, p <\n0.001; (A > B [b = 145, c = 0]) with an Effect Size Z = 12.0.\nWe observed a similar pattern using the stratified bootstrap\non the PR-AUC metric difference test (∆ = A − B). For\nDEEP-SS, ¯x = −0.080, p < 0.001; (B > A [95% CI :\n−0.136 − −0.043]), see Fig 5, Panel A. However, although\nModel A exhibits a higher PR-AUC than Model B, the\ndifference was not significant forDEEP-HPA: ¯x = 0.012, p <\n0.094; (U nclear direction [95% CI : −0.001 − 0.031]), see\nFig 5, Panel B.\nPLM-based architecture\nSimilar to the above architecture, we implemented tests\non the PLM-based models. Interestingly here, we observed\nthe reverse trend for DEEP-SS: χ2(d f = 1) = 67 .0, p <\n0.001; (A > B [b = 90, c = 8]) with an Effect Size Z = 8.3,\nwhile the difference was not significant for DEEP-HPA:\nχ2(d f = 1) = 3.5, p < 0.060; (U nclear direction [b =\n109, c = 82]) with an Effect Size Z = 2.0.\nIn contrast, for the stratified bootstrap on the PR-AUC\nmetric difference test (∆ = A − B) for DEEP-SS: ¯x =\n0.105, p < 0.001; (A > B [ 95% CI : 0 .084 − 0.125]), see\nFig 6, Panel A. Nonetheless, despite Model B depicting\na higher PR-AUC than Model A , the difference was not\nsignificant, just as before; DEEP-HPA: ¯x = −0.014, p <\n0.07; ( U nclear direction [95% CI : −0.029 − 0.001]), see\nFig 6, Panel B.\nOverall, the choice of model architecture influenced the\nperformance of the training data. The PLM-based models had\nlower uncertainty compared to CNN-based models; Fig 6 and\nFig 5.\nThe final SCL2205 dataset\nWe propose SCL2205, a “leak-proof” dataset for SCL\nmodelling that incorporates expert manual curation. This\nFig. 5. Histogram of the distribution of 1,000 bootstrap PR-AUC-\nmetric differences between CNN-based modelsA and B as trained\non SCL2205 s and DEEP-TV s, respectively . The top panel ( A) was\ntested on DEEP-SS (B > A ), whereas the bottom panel ( B) was tested\non DEEP-HPA (N o signif icant dif f erence).\nAlt text: Graphics of a histogram illustrating the performance difference\nbetween our dataset and a SoTA based on a CNN model architecture.\nTwo independent test datasets were used, an in-distribution UniProtKB\ndataset (Panel A [top]) and an out-of-distribution human protein atlas\ndataset (Panel B [bottom]).\napproach ensures high-quality training input, thereby favouring\nefficient model development. The specific compositions for the\ntwo tracks of the SCL2205 dataset (n = 19074) are provided in\nthe Results section of Supplementary File S1 . Furthermore,\nwe provide a convenient, open-source interface ( p-scldata)\nto facilitate easy access to the dataset. This is available\nthrough the public Python Package Index (PyPI), ensuring\nthat SCL2205 is easily integrated into existing bioinformatics\nworkflows.\nDiscussion\nThe rapid advancement of AI in genomics has created an urgent\nneed for robust, standardised datasets. SCL2205 addresses this\ngap by providing a meticulously curated resource for sequence-\nbased spatial modelling. By implementing a rigorous homology-\nreduction procedure, we ensure that AI models are tested\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted March 10, 2026. ; https://doi.org/10.64898/2026.03.08.710388doi: bioRxiv preprint \n\nSCL DL Prediction and Data Augmentation 9\nFig. 6. Histogram of the distribution of 1,000 bootstrap PR-AUC-\nmetric differences between PLM-based models A and B as trained\non SCL2205 s and DEEP-TV s, respectively . The top panel ( A) was\ntested on DEEP-SS (A > B ), whereas the bottom panel ( B) was tested\non DEEP-HPA (N o signif icant dif f erence).\nAlt text: Graphics of a histogram illustrating the performance difference\nbetween our dataset and a SoTA based on a PLM model architecture.\nTwo independent test datasets were used, an in-distribution UniProtKB\ndataset (Panel A [top]) and an out-of-distribution human protein atlas\ndataset (Panel B [bottom]).\nagainst truly independent data, preventing the ‘overfitting‘\nthat often plagues genomic predictions. This work establishes a\nnew benchmark for the PLM era, offering a transparent and\nscalable foundation for characterising the protein landscape\nwith precision.\nInherently, DL modelling requires large volumes of\ntraining data. While true, studies have shown that smaller,\ncarefully curated datasets can improve model performance [21].\nAlthough pre-training data sources in NLP are comparatively\nless controlled than UniProtKB, the latter still contains poor-\nquality data that requires pruning.\nIn the life sciences, we are often confronted with life-\ncritical problems that demand more stringent performance\ncontrols. Therefore, we employed a variety of filters and,\nimportantly, careful manual curation, to ensure high-quality\npre-training SCL data. For instance, unlike in [3] and [29],\nwhere the UniProtKB annotation quality scores are ignored,\nprobably because they are associated with the overall sequence\nannotation (not SCL only), we include them in filtering.\nBy considering annotation quality scores, we emphasise the\nprotein‘s functional interrelationships as inseparable biological\nprocesses. Nonetheless, it may also be argued that protein\nfunctional annotation is progressive and that certain aspects\nof a protein may be better understood than others at any given\ntime. However, highly beneficial insights can be obtained from a\nbroadly understood basis, as reflected in the overall annotation\nquality score, thereby promoting discovery. Therefore, the\nannotation score metric is an important and readily accessible\nquality control component in pre-training data development.\nAnother key technical distinction of the SCL2205 dataset\nis our decision to retain a sequence cut-off of up to 5,000\namino acids, avoiding the common practice of aggressive\ntruncation. While many SoTA predictors truncate sequences to\n1,000 residues to reduce computational overhead, such choices\nrisk discarding critical biological information. Our approach is\npremised on signal positional agnosticism; we assume that SCL\nsignals – such as C-terminal ER-retention motifs (e.g., Lys-\nAsp-Glu-Leu (KDEL)) or internal nuclear localisation signals –\ncan occur anywhere within a protein‘s primary structure.\nBy preserving nearly the entire sequence length distribution\nof the original UniProtKB records, SCL2205 ensures that\nbidirectional architectures, including bidirectional long-short-\nterm memory (bi-LSTMs) and the latest PLMs, can leverage\nterminal information from both the N- and C-termini.\nTruncation at the 1,000-residue mark would effectively “blind”\nthe backward pass of a bi-LSTMs or the global attention\nmechanism of a Transformer to the C-terminal context of\nlarger proteins. Consequently, SCL2205 provides a more\nbiologically consistent representation, particularly for the\nnotable fraction of the eukaryotic proteome that exceeds\nstandard truncation limits, thereby mitigating the risk of false\nnegatives for locations defined by non-amino-terminal signals.\nSCL datasets are inherently imbalanced, which contributes\nto the omission of low-representation classes from modelling, as\nevidenced in DL predictors [17, 2]. With low target diversity,\ngeneralisation is bound to suffer. Therefore, although with\nsome precision trade-off, we alleviated the problem by manually\nmapping sequences to their higher-order subcellular locations as\nwas illustrated in ( Fig 1 ).\nNotably, from T able 1, the number of examples for certain\nlocations – such as the Membrane – drastically increased, a\ncategory like Plastid would be excluded from modelling by, say,\na mapping logic just considering the term “plastid”, which is\nbroadly defined in UniProtKB‘s SCL ontology.\nTo validate the mapping process, we use a bar chart (Fig 3)\nto illustrate the exact mapping for Plastid – the most enriched\nlocation – as a representative case. When considered alongside\nthe individual definitions from Gene Ontology and GO\nAnnotations, we can justify the observed logic. For example,\na Plastid is defined as: “Any member of a family of organelles\nfound in the cytoplasm of plants and some protists, which are\nmembrane-bounded and contain DNA.”\nHowever, nuanced cases sometimes required a “voting”\nstrategy. For instance, in the case ofChloroplast stroma;Chloroplast\nthylakoid membrane;Plastid, the majority of annotations\nfavoured a generalised mapping to Plastid instead of\nMembrane. In instances where a “draw” occurred between\ntwo locations of interest, such as Chloroplast thylakoid\nmembrane;Plastid, we argued that mapping to Membrane\nprovided a more definitive biological classification. While it\nis alternatively arguable that mapping to Plastid supports\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted March 10, 2026. ; https://doi.org/10.64898/2026.03.08.710388doi: bioRxiv preprint \n\n10 Ouso et al.\na broader generalisation, these rare edge cases represent the\ninherent trade-offs in mapping-based data augmentation.\nConsequently, manual intervention, though labour-intensive,\nnot only increased the number of training examples but\nalso improved class diversity. As our results demonstrate,\nthis directly benefits model performance. Although location\nimbalance persists, it has been notably reduced. Therefore,\nSCL2205 provides a robust foundation for further augmentation\nmechanisms, particularly during model development. We are\naddressing this in ongoing research and invite the community\nto explore these possibilities using the SCL2205 dataset.\nOur label-mapping results suggest a compelling enhancement\nto model ranking quality, especially for the PLM-based\nmodel, while underscoring the complexity of the quality-versus-\nquantity trade-off in spatial modelling. The two statistical tests\nemployed, McNemar‘s test and stratified bootstrapping CI on\nPR-AUC differences, address different notions of “superiority”.\nWhile the PR-AUC comparison determines which model better\nranks and retrieves positives across various thresholds (quality),\nMcNemar‘s test evaluates which model performs better at a\nfixed decision threshold (the decision rule).\nIn the case of DEEP-SS, the significant performance boost\nsuggests that broader, mapped labels helped the model identify\ngeneral biological “rules” for sorting signals. However, the\nresults for DEEP-HPA tell a conflicting story across the\ntests. Here, the native labelling (Model B ) outperformed the\nmapped version on the McNemar‘s test, whilst the opposite was\nobserved with the bootstrapping test, indicating that improved\nranking performance did not translate into superior hard\nclassification performance at the chosen operating point. It also\nsuggests that the high-precision human protein atlas data carry\nvital information that mapping or taxonomic heterogeneity may\ninadvertently “blur”.\nWe also evaluated a current SoTA augmentation approach,\nhomology augmentation (see Paragraph 2.1 for definition),\nto assess its impact on data partitioning. Previous studies\nhave demonstrated the benefits of incorporating evolutionary\ninformation into both structure prediction [4, 16] and SCL\nprediction [24, 1]. These benefits may stem from improved\ngeneralisation, as shown by Nair et al., [24], who observed that\nSCL could be transferred with up to 90% accuracy for proteins\nsharing as little as 50% sequence identity.\nInterestingly, their findings also indicated that SCL transfer\nremains feasible even at lower identity thresholds, similar\nto observations in structure prediction [23], particularly\nfor major cellular compartments, which are overrepresented\nin databases. Despite these advantages, however, certain\nissues are either perpetuated or newly introduced. While\nclass imbalance is inherent, data leakage remains a critical\nconcern, especially in SoTA scenarios that claim “stringent\ntrain–test data partitioning”. We conjectured residual leakage\ndue to homology augmentation, especially for classification,\nwhere sequence labels are inherent despite using unlabelled\nsequences for augmentation. While testing this hypothesis\nmay require a more complex study, primarily due to the\nintensive resource requirements for homology searches, we\ntested a simpler proxy. Homology augmentation is a standard\npractice in sequence-based DL SCL predictors. It is anchored\non conserving evolutionary function and has been shown to\nimprove prediction accuracy relative tosequence-feature-based-\nonly approaches [24, 12]. Although it can amplify signals using\nprofiles of homologous proteins, common implementations have\nbeen shown to bias model evaluation [30]. The bias results from\nthe overlap between the training and testing datasets.\nWe demonstrated that data leakage of at least 4.8%\noccurred, based on the same overlap-reduction (partitioning)\napproach applied prior to homology augmentation, when only\n10% of the examples in the training set were used for homology\naugmentation. In contrast, previous efforts to examine biases in\ncomparing profile- and non-profile-based SCL predictors, [30]\nused an ad hoc similarity proxy that differed from the\npartitioning methods of the models under consideration –\npotentially obscuring the true extent of overlap.\nTheoretically, our findings suggest that achieving 100% data\nleakage would require fewer examples than the full training\nset; however, this would depend on factors such as the\ndegree of homology in the original dataset and the extent\nof overlap between the training/testing datasets, and the\nhomology augmentation database. Regardless, we hypothesise\nthat, relative to common practices in homology augmentation,\nsuch as MSA and position-specific scoring matrices, the extent\nof data leakage would be even greater than that observed with\nour 10% subset of native sequence data. While our similarity\ncomparisons are conducted on a native-to-native, sequence-to-\nsequence basis, MSA- and PSSM-based approaches typically\ncompare an averaged (consensus) sequence against a native\nsequence. Because a consensus sequence is intrinsically more\nlikely to produce a broader range of hits, this implies greater\noverlap and, consequently, increased leakage, potentially even\nacross different subcellular locations.\nTherefore, at a minimum, predictors applying homology\nreduction followed by homology augmentation must assess\nand report any post-homology-augmentation training–testing\noverlap. We acknowledge that rigorously evaluating this\nis challenging, as the sequence representation following\naugmentation differs from that used during initial reduction\nunless equivalent consensus sequences are employed.\nIn light of the above, pursuing emerging alternative\naugmentation strategies, [14, 25, 9, 29], may be a superior\noption. Interestingly, we observed that SCL2205 better\ncomplements the new frontier of PLM by exhibiting improved\ngeneralisation over DEEP-TV for the in-distribution set\n(DEEP-SS), while exhibiting reduced performance on the\nOOD counter-set ( DEEP-HPA; Figure 6 ). To delineate true\neffect from noise, we performed a heterogeneity check using\nCochran‘s Q test. The significant heterogeneity observed across\nthe models suggests that the CNN and PLM architectures are\nlearning fundamentally different signals from the same data.\nConversely, we observed contrasting results between\nSCL2205 and DEEP-TV on both independent test sets when\nusing the CNN-based model (textbfFigure 5). These outcomes\ncan be attributed to several factors beyond architectural\ndesign. Foremost, DEEP-HPA is a unique OOD dataset due\nto its taxonomic bias; we therefore expect our multi-taxa\ndataset to notably diverge from it relative to DEEP-TV .\nSecondly, the training paradigm for PLMs aligns more naturally\nwith our label-mapping approach. Thirdly, the statistical\ntests are tailored for different notions of “superiority”.\nFinally, differences in model calibration may contribute to\noverconfidence, particularly in PLMs [13, 8, 6]. However, both\nmodels were tested in their vanilla forms, without calibration\nor optimisation, and reliability diagrams depicted very similar\nprofiles across the model variants.\nThese divergences highlight a critical trade-off in spatial\nmodelling: the delicate balance between breadth (e.g., through\nmapping and/or taxonomic diversity) and depth (the precision\nof native labels and/or taxonomic specificity).\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted March 10, 2026. ; https://doi.org/10.64898/2026.03.08.710388doi: bioRxiv preprint \n\nSCL DL Prediction and Data Augmentation 11\nOverall, SCL2205 alignsmore effectively than DEEP-\nTV with the increasing adoption of pre-trained PLMs in\ngenomics. This distinction is critical for the PLM frontier: if a\nresearcher requires a definitive “yes/no” prediction, the native\nlabelling of Model B is preferable. Conversely, for large-scale\ngenomic screening where ranking candidates is the priority, the\naugmented mapping of Model A provides greater utility.\nUltimately, we have proposed a rigorously developed dataset\nthat enhances the trustworthiness of AI models. It facilitates\nefficient and environmentally sustainable development by\nmitigating “noisy” overhead data, whilst enriching generalisation\nthrough diversity. By providing SCL2205 as an installable\npackage (p-scldata ), we adhere to open-science principles,\nensuring seamless integration with modern ML frameworks and\nproviding a compelling benchmark for future genomic discovery.\nConclusion\nThe computational annotation of proteins with SCL remains\na critical component of functional genome annotation. As\nthe field enters new frontiers, such as PLMs, new challenges\nemerge – most notably regarding the quality and diversity of\ninput data. While data are the primary driving force behind\nthe AI revolution, acquiring high-quality datasets for training\nbiological models remains a significant hurdle. Nevertheless,\ntechniques such as data augmentation and human-aided\ncuration have begun to alleviate these issues. In this study,\nwe present SCL2205: a new, high-quality dataset that\nhighlights the limitations of current state-of-the-art approaches\nto data development. Through independent validation, we\ndemonstrate that SCL2205 outperforms existing alternatives,\nparticularly for applications involving PLM-based modelling.\nWe acknowledge that certain persistent challenges, such as class\nimbalance, must still be addressed through further research\nbeyond data preparation. Nonetheless, SCL2205 serves as a\ntrustworthy and reliable benchmark for AI application in SCL\nprediction, laying the foundation for future discoveries in the\ngenomic landscape. Ultimately, by providing a more accurate\nspatial map of localisation within the cell, we move closer to\na future where we can rapidly identify the molecular drivers of\nrare diseases and accelerate the development of life-changing\ntargeted therapies.\nCompeting interests\nNo competing interest is declared.\nFunding\nThis work was supported by Research Ireland through the\nCentre for Research Training in Genomics Data Science\n[#18/CRT/6214 to G.P.]\nData availability\nData are available through the following open-source avenues:\n•DR Y AD: An archive (DOI) under the Creative Commons\nZero (CC0 1.0) licence.\n•PyPI: An installable Python data package (p-scldata) under\nthe MIT licence.\n•Data & Code: GitHub\nAuthor contributions statement\nDO and GP conceived the experiments, DO conducted the\nexperiments, DO analysed the results, and DO wrote the\nmanuscript. DO and GP reviewed and edited the manuscript.\nGP supervised the study. The project funding was through GP.\nAcknowledgments\nThe authors thank the anonymous reviewers for their valuable\nsuggestions. This research was funded by Research Ireland\nthrough the Centre for Research Training in Genomics Data\nScience under Grant number #18/CRT/6214.\nReferences\n1. Alessandro Adelfio, Viola Volpato, and Gianluca Pollastri.\nSclpredt: Ab initio and homology-based prediction\nof subcellular localization by n-to-1 neural networks.\nSpringerPlus, 2:1–11, 10 2013.\n2. Jose Juan Almagro Armenteros, Marco Salvatore, Olof\nEmanuelsson, Ole Winther, Gunnar Von Heijne, Arne\nElofsson, and Henrik Nielsen. Detecting sequence signals\nin targeting peptides using deep learning. Life Science\nAlliance, 2, 2019.\n3. Jose Juan Almagro Armenteros, Casper Kaae Sønderby,\nSøren Kaae Sønderby, Henrik Nielsen, and Ole Winther.\nDeeploc: prediction of protein subcellular localization using\ndeep learning. Bioinformatics, 33:4049–4049, 12 2017.\n4. Minkyung Baek, Frank DiMaio, Ivan Anishchenko,\nJustas Dauparas, Sergey Ovchinnikov, Gyu Rie Lee,\nJue Wang, Qian Cong, Lisa N. Kinch, R. Dustin\nSchaeffer, Claudia Mill´ an, Hahnbeom Park, Carson\nAdams, Caleb R. Glassman, Andy DeGiovanni, Jose H.\nPereira, Andria V. Rodrigues, Alberdina A. Van Dijk,\nAna C. Ebrecht, Diederik J. Opperman, Theo Sagmeister,\nChristoph Buhlheller, Tea Pavkov-Keller, Manoj K.\nRathinaswamy, Udit Dalwadi, Calvin K. Yip, John E.\nBurke, K. Christopher Garcia, Nick V. Grishin, Paul D.\nAdams, Randy J. Read, and David Baker. Accurate\nprediction of protein structures and interactions using a\nthree-track neural network. Science, 373:871–876, 8 2021.\n5. Torsten Blum, Sebastian Briesemeister, and Oliver\nKohlbacher. Multiloc2: integrating phylogeny and gene\nontology terms improves subcellular protein localization\nprediction. BMC bioinformatics, 10:274, 9 2009.\n6. Yangyi Chen, Lifan Yuan, Ganqu Cui, Zhiyuan Liu, and\nHeng Ji. A close look into the calibration of pre-trained\nlanguage models. In Anna Rogers, Jordan Boyd-Graber,\nand Naoaki Okazaki, editors, Proceedings of the 61st\nAnnual Meeting of the Association for Computational\nLinguistics (Volume 1: Long Papers) , pages 1343–1367,\nToronto, Canada, July 2023. Association for Computational\nLinguistics.\n7. The UniProt Consortium, Alex Bateman, Maria-Jesus\nMartin, Sandra Orchard, Michele Magrane, Shadab\nAhmad, Emanuele Alpi, Emily H Bowler-Barnett, Ramona\nBritto, Hema Bye-A-Jee, Austra Cukura, Paul Denny,\nTunca Dogan, ThankGod Ebenezer, Jun Fan, Penelope\nGarmiri, Leonardo Jose da Costa Gonzales, Emma Hatton-\nEllis, Abdulrahman Hussein, Alexandr Ignatchenko,\nGiuseppe Insana, Rizwan Ishtiaq, Vishal Joshi, Dushyanth\nJyothi, Swaathi Kandasaamy, Antonia Lock, Aurelien\nLuciani, Marija Lugaric, Jie Luo, Yvonne Lussi, Alistair\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted March 10, 2026. ; https://doi.org/10.64898/2026.03.08.710388doi: bioRxiv preprint \n\n12 Ouso et al.\nMacDougall, Fabio Madeira, Mahdi Mahmoudy, Alok\nMishra, Katie Moulang, Andrew Nightingale, Sangya\nPundir, Guoying Qi, Shriya Raj, Pedro Raposo, Daniel L\nRice, Rabie Saidi, Rafael Santos, Elena Speretta, James\nStephenson, Prabhat Totoo, Edward Turner, Nidhi Tyagi,\nPreethi Vasudev, Kate Warner, Xavier Watkins, Rossana\nZaru, Hermann Zellner, Alan J Bridge, Lucila Aimo,\nGhislaine Argoud-Puy, Andrea H Auchincloss, Kristian B\nAxelsen, Parit Bansal, Delphine Baratin, Teresa M Batista\nNeto, Marie-Claude Blatter, Jerven T Bolleman, Emmanuel\nBoutet, Lionel Breuza, Blanca Cabrera Gil, Cristina\nCasals-Casas, Kamal Chikh Echioukh, Elisabeth Coudert,\nBeatrice Cuche, Edouard de Castro, Anne Estreicher,\nMaria L Famiglietti, Marc Feuermann, Elisabeth Gasteiger,\nPascale Gaudet, Sebastien Gehant, Vivienne Gerritsen,\nArnaud Gos, Nadine Gruaz, Chantal Hulo, Nevila Hyka-\nNouspikel, Florence Jungo, Arnaud Kerhornou, Philippe Le\nMercier, Damien Lieberherr, Patrick Masson, Anne Morgat,\nVenkatesh Muthukrishnan, Salvo Paesano, Ivo Pedruzzi,\nSandrine Pilbout, Lucille Pourcel, Sylvain Poux, Monica\nPozzato, Manuela Pruess, Nicole Redaschi, Catherine\nRivoire, Christian J A Sigrist, Karin Sonesson, Shyamala\nSundaram, Cathy H Wu, Cecilia N Arighi, Leslie Arminski,\nChuming Chen, Yongxing Chen, Hongzhan Huang, Kati\nLaiho, Peter McGarvey, Darren A Natale, Karen Ross,\nC R Vinayaka, Qinghua Wang, Yuqi Wang, and Jian\nZhang. Uniprot: the universal protein knowledgebase in\n2023. Nucleic Acids Research, 51:D523–D531, 1 2023.\n8. Shrey Desai and Greg Durrett. Calibration of pre-trained\ntransformers. In Bonnie Webber, Trevor Cohn, Yulan\nHe, and Yang Liu, editors, Proceedings of the 2020\nConference on Empirical Methods in Natural Language\nProcessing (EMNLP), pages 295–302, Online, November\n2020. Association for Computational Linguistics.\n9. Ahmed Elnaggar, Michael Heinzinger, Christian Dallago,\nGhalia Rehawi, Yu Wang, Llion Jones, Tom Gibbs, Tamas\nFeher, Christoph Angerer, Martin Steinegger, Debsindhu\nBhowmik, and Burkhard Rost. Prottrans: Toward\nunderstanding the language of life through self-supervised\nlearning. IEEE Transactions on Pattern Analysis and\nMachine Intelligence, 44:7112–7127, 10 2022.\n10. Steven Y Feng, Varun Gangal, Jason Wei, Sarath\nChandar, Soroush Vosoughi, Teruko Mitamura, and\nEduard Hovy. A survey of data augmentation\napproaches for nlp. In Findings of the Association for\nComputational Linguistics: ACL-IJCNLP 2021, pages\n968–988. Association for Computational Linguistics, 8 2021.\n11. Limin Fu, Beifang Niu, Zhengwei Zhu, Sitao Wu, and\nWeizhong Li. Sequence analysis cd-hit: accelerated\nfor clustering the next-generation sequencing data.\nBioinformatics, 28:3150–3152, 12 2012.\n12. Maryam Gillani and Gianluca Pollastri. Impact\nof alignments on the accuracy of protein subcellular\nlocalization predictions. Proteins: Structure, Function and\nBioinformatics, 3 2024.\n13. Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q.\nWeinberger. On calibration of modern neural networks.\nIn Proceedings of the 34th International Conference on\nMachine Learning - Volume 70, ICML’17, page 1321–1330.\nJMLR.org, 2017.\n14. Michael Heinzinger, Ahmed Elnaggar, Yu Wang, Christian\nDallago, Dmitrii Nechaev, Florian Matthes, and Burkhard\nRost. Modeling aspects of the language of life through\ntransfer-learning protein sequences. BMC Bioinformatics,\n20, 12 2019.\n15. Annette H¨ oglund, Pierre D¨ onnes, Torsten Blum,\nHans Werner Adolph, and Oliver Kohlbacher. Multiloc:\nPrediction of protein subcellular localization using n-\nterminal targeting sequences, sequence motifs and amino\nacid composition. Bioinformatics, 22:1158–1165, 5 2006.\n16. John Jumper, Richard Evans, Alexander Pritzel, Tim\nGreen, Michael Figurnov, Olaf Ronneberger, Kathryn\nTunyasuvunakool, Russ Bates, Augustin ˇZ´ ıdek, Anna\nPotapenko, Alex Bridgland, Clemens Meyer, Simon A.A.\nKohl, Andrew J. Ballard, Andrew Cowie, Bernardino\nRomera-Paredes, Stanislav Nikolov, Rishub Jain, Jonas\nAdler, Trevor Back, Stig Petersen, David Reiman, Ellen\nClancy, Michal Zielinski, Martin Steinegger, Michalina\nPacholska, Tamas Berghammer, Sebastian Bodenstein,\nDavid Silver, Oriol Vinyals, Andrew W. Senior, Koray\nKavukcuoglu, Pushmeet Kohli, and Demis Hassabis.\nHighly accurate protein structure prediction with alphafold.\nNature 2021 596:7873, 596:583–589, 7 2021.\n17. Manaz Kaleel, Liam Ellinger, Clodagh Lalor, Gianluca\nPollastri, and Catherine Mooney. Sclpred-mem: Subcellular\nlocalization prediction of membrane proteins by deep n-\nto-1 convolutional neural networks. Proteins: Structure,\nFunction and Bioinformatics, 89:1233–1239, 10 2021.\n18. Manaz Kaleel, Yandan Zheng, Jialiang Chen, Xuanming\nFeng, Jeremy C Simpson, Gianluca Pollastri, and Catherine\nMooney. Sclpred-ems: subcellular localization prediction of\nendomembrane system and secretory pathway proteins by\ndeep n-to-1 convolutional neural networks. Bioinformatics,\n6 2020.\n19. Yann LeCun and Corinna Cortes. The mnist database of\nhandwritten digits, 1 2010.\n20. Andrew L. Maas, Raymond E. Daly, Peter T. Pham,\nDan Huang, Andrew Y. Ng, and Christopher Potts.\nLearning word vectors for sentiment analysis. In\nDekang Lin, Yuji Matsumoto, and Rada Mihalcea,\neditors, Proceedings of the 49th Annual Meeting of\nthe Association for Computational Linguistics: Human\nLanguage Technologies, pages 142–150, Portland, Oregon,\nUSA, jun 2011. Association for Computational Linguistics.\n21. Max Marion, Ahmet ¨Ust¨ un, Luiza Pozzobon, Alex Wang,\nMarzieh Fadaee, and Sara Hooker. When less is more:\nInvestigating data pruning for pretraining llms at scale,\n2023.\n22. Philippe Le Mercier, Jerven Bolleman, Edouard De Castro,\nElisabeth Gasteiger, Parit Bansal, Andrea H. Auchincloss,\nEmmanuel Boutet, Lionel Breuza, Cristina Casals-Casas,\nAnne Estreicher, Marc Feuermann, Damien Lieberherr,\nCatherine Rivoire, Ivo Pedruzzi, Nicole Redaschi, and\nAlan Bridge. Swissbiopics—an interactive library of\ncell images for the visualization of subcellular location\ndata. Database: The Journal of Biological Databases and\nCuration, 2022:1–5, 2022.\n23. Catherine Mooney and Gianluca Pollastri. Beyond\nthe twilight zone: Automated prediction of structural\nproperties of proteins by recursive neural networks and\nremote homology information. Proteins: Structure,\nFunction and Bioinformatics, 77:181–190, 10 2009.\n24. Rajesh Nair and Burkhard Rost. Sequence conserved for\nsubcellular localization. Protein science : a publication of\nthe Protein Society, 11:2836–2847, 4 2002.\n25. Alexander Rives, Joshua Meier, Tom Sercu, Siddharth\nGoyal, Zeming Lin, Jason Liu, Demi Guo, Myle Ott,\nC. Lawrence Zitnick, Jerry Ma, and Rob Fergus. Biological\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted March 10, 2026. ; https://doi.org/10.64898/2026.03.08.710388doi: bioRxiv preprint \n\nSCL DL Prediction and Data Augmentation 13\nstructure and function emerge from scaling unsupervised\nlearning to 250 million protein sequences. Proceedings of\nthe National Academy of Sciences of the United States of\nAmerica, 118:e2016239118, 4 2021.\n26. Connor Shorten and Taghi M. Khoshgoftaar. A survey on\nimage data augmentation for deep learning. Journal of Big\nData, 6:1–48, 12 2019.\n27. Connor Shorten, Taghi M Khoshgoftaar, and Borko Furht.\nText data augmentation for deep learning. Journal of Big\nData, 8, 7 2021.\n28. Peter J. Thul, Lovisa Akesson, Mikaela Wiking, Diana\nMahdessian, Aikaterini Geladaki, Hammou Ait Blal, Tove\nAlm, Anna Asplund, Lars Bj¨ ork, Lisa M. Breckels, Anna\nB¨ ackstr¨ om, Frida Danielsson, Linn Fagerberg, Jenny Fall,\nLaurent Gatto, Christian Gnann, Sophia Hober, Martin\nHjelmare, Fredric Johansson, Sunjae Lee, Cecilia Lindskog,\nJan Mulder, Claire M. Mulvey, Peter Nilsson, Per Oksvold,\nJohan Rockberg, Rutger Schutten, Jochen M. Schwenk,\nAsa Sivertsson, Evelina Sj¨ ostedt, Marie Skogs, Charlotte\nStadler, Devin P. Sullivan, Hanna Tegel, Casper Winsnes,\nCheng Zhang, Martin Zwahlen, Adil Mardinoglu, Fredrik\nPont´ en, Kalle Von Feilitzen, Kathryn S. Lilley, Mathias\nUhl´ en, and Emma Lundberg. A subcellular map of the\nhuman proteome. Science, 356, 5 2017.\n29. Vineet Thumuluri, Jos´ e Jos ´, Jos´ e Juan, Almagro\nArmenteros, Alexander Rosenberg Johansen, Henrik\nNielsen, and Ole Winther. Deeploc 2.0: multi-label\nsubcellular localization prediction using protein language\nmodels. Nucleic Acids Research, 50, 2022.\n30. Gregor Urban, Mirko Torrisi, Christophe N. Magnan,\nGianluca Pollastri, and Pierre Baldi. Protein profiles:\nBiases and protocols. Computational and Structural\nBiotechnology Journal, 18:2281–2289, 1 2020.\n31. Ian Walsh, Dmytro Fishman, Dario Garcia-Gasulla, Tiina\nTitma, Gianluca Pollastri, Emidio Capriotti, Rita Casadio,\nSalvador Capella-Gutierrez, Davide Cirillo, Alessio Del\nConte, Alexandros C. Dimopoulos, Victoria Dominguez Del\nAngel, Joaquin Dopazo, Piero Fariselli, Jos´ e Maria\nFern´ andez, Florian Huber, Anna Kreshuk, Tom Lenaerts,\nPier Luigi Martelli, Arcadi Navarro, Pilib Broin, Janet\nPi˜ nero, Damiano Piovesan, Martin Reczko, Francesco\nRonzano, Venkata Satagopam, Castrense Savojardo,\nVojtech Spiwok, Marco Antonio Tangaro, Giacomo Tartari,\nDavid Salgado, Alfonso Valencia, Federico Zambelli,\nJennifer Harrow, Fotis E. Psomopoulos, and Silvio C.E.\nTosatto. Dome: recommendations for supervised machine\nlearning validation in biology. Nature Methods 2021 18:10,\n18:1122–1127, 7 2021.\n32. Thomas Wolf, Lysandre Debut, Victor Sanh, Julien\nChaumond, Clement Delangue, Anthony Moi, Pierric\nCistac, Tim Rault, R´ emi Louf, Morgan Funtowicz,\nJoe Davison, Sam Shleifer, Patrick Von Platen, Clara\nMa, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le\nScao, Sylvain Gugger, Mariama Drame, Quentin Lhoest,\nand Alexander M Rush. Transformers: State-of-the-\nart natural language processing. In Proceedings of\nthe 2020 Conference on Empirical Methods in Natural\nLanguage Processing: System Demonstrations , pages 38–\n45. Association for Computational Linguistics, 10 2020.\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted March 10, 2026. ; https://doi.org/10.64898/2026.03.08.710388doi: bioRxiv preprint","source_license":"CC-BY-4.0","license_restricted":false}