{"paper_id":"08e94e9e-6314-4f1e-8a9c-779cfb058162","body_text":"1 \nRP3Net: a deep learning model for predicting recombinant \nprotein production in Escherichia coli \n \nEvgeny Tankhilevich1,2* †, Sergio Martinez Cuesta2, Ian Barrett2, Carolina Berg3, Lovisa Holmberg \nSchiavone3 & Andrew R Leach1,4 \n \n1European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton, Cam-\nbridge, CB10 1SD, United Kingdom. \n2Data Sciences and Quantitative Biology, Discovery Sciences, BioPharmaceuticals R&D, Astra-\nZeneca, Cambridge, UK \n3Protein Science, Structure and Biophysics, Discovery Sciences, BioPharmaceuticals R&D, \nAstraZeneca, Gothenburg, Sweden \n4current address: LifeArc, 7-12 Tavistock Square, London WC1H 9LT \n \nRecombinant protein expression can be a limiting step in the production of protein reagents \nfor drug discovery and other biotechnology applications. We introduce RP3Net (Recombinant \nProtein Production Prediction Network), an AI model of small-scale heterologous soluble pro-\ntein expression in Escherichia coli. RP3Net utilizes the most recent protein and genomic foun-\ndational models. A curated dataset of   internal experimental results from AstraZeneca  (AZ) \nand publicly  available data from the Structural Genomics Consortium (SGC) was used for \ntraining, validation and testing  of RP3Net. Set Transformer Pooling  (STP) aggregation and \nMeta Label Correction (MLC) with large scale purification data enabled RP3Net to improve \nArea Under Receiver Operator Curve (AUROC) by 0.15, compared to the baseline model. \nWhen experimentally validated on an independent, manually selected set of 97 constructs, \nRP3Net outperformed currently available models, with an AUROC of 0.83, delivering accu-\nrate predictions in 77% of the cases , and correctly identif ying successfully expressing con-\nstructs in 92% of cases. \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n 2 \nIntroduction \nMotivation\t\nThe production of protein reagents is an essential part of the research and development process in \nthe pharmaceutical and biotechnology industries. In drug discovery it is often a pre -requisite for \nscreening and hit identification 1,2. In living tissues,  the target protein may occur in very small \namounts alongside numerous other biomolecules.  To be used for high-throughput screening of drug \ncandidates, structural determination and the development of functional assays, the target protein \nneeds to be expressed in a cell culture and purified. The ability to express a protein depends on \nmultiple factors. First and foremost is the protein itself, but other factors include the cloning vector, \nthe species and strain of the host cells, the codon optimisation algorithm, the use of tags and fusion \nproteins, and other experimental conditions 3–21. The choice of these parameters is often influenced \nby the details  of the downstream experiments ,22 making protein production time-consuming and \nerror-prone, and often requiring multiple iterations and much trial and error. The purpose of this \nwork is to develop a deep learning model  to predict soluble protein expression in E. coli from the \nconstruct sequence, thus accelerating the timescales for protein production from months to weeks, \ncutting costs and reducing environmental impact. \nA recombinant protein production experimental pipeline involves several steps 23, including con-\nstruct design, cloning, small -scale expression screening, progression of expressing constructs to \nlarge-scale purification and quality control (QC). Small-scale soluble expression screening, shown \nin Fig. 1, is crucial for assessing whether to progress the construct to large -scale production. First, \ncells are transfected with vectors (e.g. plasmids) carrying the cloned DNA of the protein of interest \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n3 \n(step 1). Recombinant protein production is performed in deep well format (step 2). Cells are spun  \ndown and lysed (step 3). The lysate contains the total amount of protein produced. After an addi-\ntional centrifugation step the soluble protein is found in the supernatant and the insoluble material \nis discarded in the pellet (step 3). Soluble protein is captured in a one -step purification using the \nhistidine tag and immobilized metal affinity chromatography (IMAC24, step 4). The soluble protein \nyield and correct size are typically assessed by performing denaturing gel electrophoresis (SDS -\nPAGE) where yield and size are compared to a protein standard (step 5). The yield can be estimated \nby quantifying the amount of the target protein compared to the protein standard in the stained gel, \nFig. 1 The experimental workflow for small-scale recombinant soluble protein production with one-step purification. After cloning, \nplasmids with the genetic material of the protein of interest are transfected into E. coli cells (step 1). The cells are grow n for 24 \nhours (step 2). Harvesting involves two centrifugation stages: first to pellet down the cells, then, after lysis, to isolate the soluble \nprotein in the supernatant (step 3). This is followed by IMAC purification (step 4), yield estimation via densitometric analysis of \nSDS-PAGE gels (step 5) and, finally, data capture for further analysis and machine learning (ML, step 6). Image generated with \nBioRender.com. \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n4 \nusing densitometric analysis. At this stage, it is important to record both positive (produced) and \nnegative (failed to produce) experimental outcomes (step 6). Throughout the rest of this publication, \nunless stated explicitly, terms “protein production” and “protein expression” refer to this step of the \nexperimental pipeline. Constructs that pass this small-scale screening are then typically progressed \nto large-scale purification and further downstream applications.  \nProtein and DNA foundation models (FMs) have become ubiquitous tools for predicting structural \nand functional protein properties from amino acid and/or nucleic acid sequences  25–32. These FMs \nare typically trained on large corpora of sequences, such as UniRef33,34, GenBank35,36, MGnify37 or \nBFD38 (Table 1). During training a portion of the sequence is masked out, i.e. each residue is re-\nplaced with a special \"blank\" character. The training objective then becomes to reconstruct the \nmasked portion of the input, or \"fill in the blanks\". This technique, referred to as language modelling \ntask, originates from natural language processing39.  \nESM25,26,40, ProtBert27 and ProteinBert28 are examples of protein FMs; DNABert30,31 is a popular \nDNA FM. These models are all based on Transformer deep learning architecture 41 with different \nnumber of layers, feature dimensions and other details. HyenaDna 29 is another DNA FM that uses \na different architecture. \nThe intermediate layers of foundational models yield a residue -level sequence representation that \ncan be used to predict the protein property of interest, such as secondary or tertiary structure, binding \naffinity, fluorescence, thermodynamic stability, sol ubility, etc25,26,38,40,42–47. The experimental da-\ntasets that describe these properties typically contain orders of magnitude fewer entries when com-\npared to the sequence corpora. This scarcity of experimental datasets often makes it unfeasible to \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n5 \ntrain large foundational models from scratch for predicting protein properties. Such models are usu-\nally pre-trained with the language modelling objective on the large corpora first, and then further \ntrained to predict the property of interest using the smaller dataset 42,43,47. This final training step is \nreferred to as fine -tuning. RP3Net follows this architectural blueprint by encoding the biological \nsequence with a foundational model, feeding this encoding through an aggregation layer to obtain \na global representation for the entire construct, and then applying a fully connected classific ation \nhead to compute the predicted probability of recombinant expression in E. coli (Fig. 2A) as a binary \noutcome.  \nAlthough, in theory, a soluble protein production fine-tuning dataset could be designed and \nA \nFig. 2. A. Architecture diagram of RP3Net. The input biological sequence is encoded by the foundation model to obtain a sequence \nrepresentation, where each residue/codon/nucleotide is represented by a vector. The aggregation layer builds a global protein rep-\nresentation vector from the sequence representation. The predicted probability of successful recombinant expression of the protein \nin E. coli is computed by the fully connected classification head from the protein representation. B. Training with meta label correc-\ntion on a mixture of clean and noisy data. The standard training setup, where the model loss on clean inputs and labels is minimised \nwith gradient descent, is shown in the top row. A special “teacher” model is trained to predict the corrected labels from the noisy \ninput and labels. These corrected labels, along with noisy inputs, serve as inputs for training the “student” model. The latter model \nhas the same architecture and weights as the “clean” model. The bi-level optimisation algorithm that makes sure that the corrected \nlabels do not deviate from the (unknown) clean labels, relies on using the cross entropy (CE) loss. Model components with trainable \nweights are shown as blue boxes. Training data is shown as yellow boxes. Images generated with BioRender.com. \n \nB \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n6 \nexperimentally generated from scratch, in practice it would be too time -consuming and expensive. \nMoreover, there already exist publicly available datasets of protein expression that contain the re-\nsults of experiments worth millions of dollars and representing years of lab work48–50. For training \nand evaluating RP3Net, the internal AstraZeneca (AZ) small-scale expression screen data is com-\nbined with datasets from the Structural Genomics Consortium (SGC), specifically their sites in \nStockholm4,51 and Toronto23,50. The experimental pipeline for generating data from AZ and SGC \nStockholm has been already discussed above, see  Fig. 1, step 6. SGC Toronto captures the results \nof large-scale protein purification.  \n \nExisting\twork\t\nA number of models that predict soluble expression from construct sequence have been published \nin recent years. Most of these systems use datasets derived from the Protein Structure Initiative \n(PSI) compendium, also referred to as TargetTrac k 48,49. PSI was an experimental research effort \nrun across multiple laboratories in 2000-2017, with the objective of determining protein structures \nand depositing them in the Protein Data Bank (PDB)52,53 . This dataset records the pipeline position, \ni.e. the experimental stage where the work was terminated, for each target and construct. For exam-\nple, if a construct was selected and cloned, but could not be expressed, its pipeline position would \nbe recorded as \"cloned\". For another construct that has been selected, cloned, expressed and puri-\nfied, but could not be crystallised, the pipeline position would be \"purified\", e tc. One limitation of \nusing TargetTrack data in this work is that in this dataset a genuine inability to express the construct \nunder given experimental conditions can be confused with stopping to pursue the construct for other \nreasons (for example there being another well-behaving construct for the same target ). Different \nlabs that have provided data for TargetTack were using different experimental pipelines: sometimes \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n7 \nsmall-scale expression screening as shown in Fig. 1, but also large-scale purification results as SGC \nToronto, or the results of running SDS-PAGE on unpurified cell lysate. \nIt is important to make a distinction between solubility as a general physical property of the protein, \nwhich can be measured, for example, as peak concentration in the solution, and the ability to achieve \nsoluble expression of the protein under given experimental conditions, which is a binary outcome \nthat is modelled in this work. A protein that is generally soluble could still fail to express, for ex-\nample because it is toxic for the host cells, or because the chaperones that are required for forming \nthe correct structure are missing, or due to other reasons. Solubility is thus a necessary but insuffi-\ncient condition for soluble recombinant protein production. \nNetSolP44 uses PSI data to evaluate multiple transformer -based models available at the time for \npredicting soluble expression. NetSolP outputs two scores: solubility and \"usability\", the latter be-\ning a combined predictor of solubility and the ability of a protein to be expressed.  \nPLMC46 and SADeepCry45 also use data derived from TargetTrack and a Transformer architecture \nbut output the pipeline position given the construct sequence. PPCPred54, PredPPCrys55, Crysalis56  \nDCFCrystal57 are examples of older, simpler models that predict pipeline position, trained on vari-\nous subsets of TargetTrack. SoluProt58 uses a different PSI-based dataset with a Gradient Boosted \nMachine (GBM) model 59 and global features based on relative amino acid frequencies, predicted \nphysicochemical properties, similarity to E. coli proteome and output of various other bioinformat-\nics tools to predict soluble expression. \nCamSol60,61 is a well -established relative solubility prediction tool for libraries of similar protein \nsequences. There are many other solubility predictions that use deep neural networks, such as \nGPSFun62 and PLM_Sol63. A few methods exist for modelling expression and solubility of human \nantibodies, but their experimental protocols differ substantially from E. coli -based expression \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n8 \nanalysed in this work64,65. \n \n  \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n9 \nResults and discussion \nRP3Net\twith\tfixed\tfoundation\tmodel\tweights\toutperforms\tdecision\ttrees\twith\tglobal\t\nprotein\tfeatures\t\nThe RP3Net architecture with fixed foundation model weights and mean pooling  (Fig. 2A) was \nused for selecting the best performing foundation model. This architecture is denoted as Model A \n(Table 2). The models were trained and evaluated on SGC Stockholm dataset, with five-fold cross \nvalidation. SGC Stockholm was used because this dataset is of medium size, compared to AZ, which \nis much smaller, and SGC Toronto, which is much larger. This data set is also the only one of the \nthree that provides DNA sequences for all the constructs. \nA gradient boosted decision tree (XGBoost 59) with global protein features as inputs was used as a \nFig. 3. Performance of RP3Net with fixed foundation model (FM) weights and mean pooling (Model A) on SGC Stockholm, along \nwith FM parameter count. On the left y-axis each boxplot shows the area under Receiver Operator Curve (AUROC) of Model A with \na particular FM, evaluated on SGC Stockholm test data, with five-fold cross validation. On the right y-axis, black squares show the \nnumber of trainable parameters of the FM, in log scale. ESM2 (650M) and CaLM were selected for further analysis based on perfor-\nmance, consistency, parameter count and licensing restrictions. “Random embeddings” means using random residue embeddings \ninstead of a foundation model. The source data for all charts is available in the supplement. \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n10 \nbaseline model. As shown in Fig. 3, Model A with any protein FM outperforms the baseline model. \nOut of all the tested DNA and codon FMs, only CaLM shows better results than the baseline. A \nplausible explanation for this observation is due to the datasets used for pre-training the FMs: CaLM \nwas pre-trained on coding sequences from ENA, whereas other DNA FMs were pre -trained on a \nmixture of coding and non-coding sequences. \nRP3Net performance also varies depending on the training data subset that was used, sometimes \ndramatically. For example, for the more consistent FMs, such as ESM2 (650M) and CaLM, the \ndifference between the best and the worst runs is 0.03 and 0.01, respectively, whereas for Hye-\nnaDNA Medium the difference is 0.14. \nThe number of trainable foundation model parameters is used to indicate the resource requirements \nfor fine-tuning the FM (compute time, memory) . The foundation model for subsequent evaluation \nwas chosen based on the pragmatic trade-off between performance, training complexity and licens-\ning constraints (see Table 1). We selected ESM2 with 650 million parameters . The simple Model \nA training protocol, applied to this FM, achieves an average increase in AUROC of 0.03, compared \nto the baseline model. \nPerformance\ton\tdifferent\tdata\tsources\treveals\tdependency\ton\tdataset\tsize.\t\nModel A performance on SGC Stockholm dataset can be improved by replacing the mean pooling \naggregation layer with a more sophisticated set transformer pooling (STP 66,67). This configuration \nis denoted as Model B. The main difference between mean pooling and STP is that, whereas the \nformer just takes an average across the sequence, giving each residue the same weight, STP uses \ncontext-dependent weights for residue represe ntations, by computing multiheaded attention \n(MHA41, see Methods) between a special parameter, called the seed vector, and the output of the \nfoundation model. The seed vector is updated during training with gradient descent, along with the \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n11 \nrest of the model parameters.  \nModel B gives an AUROC improvement of 0.01% over Model A when trained and evaluated on \nSGC Stockholm (Table 3). The performance of Model B on other data sources varies with an AU-\nROC of 0.59 on the AZ dataset and of 0.84 on SGC Toronto. A plausible reason for this variation \nis dataset size. Training Model B on the combined AZ and SGC Stockholm data improves evalua-\ntion of AZ to 0.73, which is almost the same as evaluating the same model on SGC Stockholm. \nAdding SGC Toronto to the training data does not improve the evaluation results significantly for \nany data source.  \n \nMeta\tlabel\tcorrection\twith\tpurification\tdata\tyields\ta\t0.04\tincrease\tin\tAUROC\ton\tSGC\t\nStockholm\t\nBoth Model A and Model B are fine-tuned on soluble protein expression data with frozen parame-\nters of the foundation model. Unfreezing these parameters (Model C) and training on the full dataset \nleads to overfitting: perfect performance is quickly achieved on the training data set (AUROC»1.0), \nbut on the validation and test sets the AUROC remains below 0.75. Training Model C on individual \ndata sources also leads to overfitting, as expected. \nThis could be explained by the fact that the datasets contain the results of slightly different experi-\nments. The SGC Toronto dataset reports results of large-scale purification, whereas both AZ and \nSGC Stockholm report small-scale expression testing captured with one-step purification. Although \nthe exact experimental conditions, materials and methods used for purifications were not available \nduring model development, it is safe to assume that the SGC Toronto conditions are quite different \nfrom the small-scale expression testing. A natural question arises: given the construct sequence from \nSGC Toronto, and its binary purification result, what would be the result of small-scale expression \ntesting this construct under the conditions of SGC Stockholm or AZ?  \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n12 \nWe address this within the Meta Label Correction framework (MLC68,69), where a large, noisy data \nset is used to aid training the model on a small, clean set. Rather than adding the noisy data directly \nto the training set, a special model is trained to predict the corrected label from the noisy input and \nnoisy label. This is referred to as the \"teacher model\". The corrected labels are used to train the \n\"student model\", along with clean inputs and clean labels ( Fig. 2B). In our setup, SGC Toronto \nlarge scale purification data set is used to train the teacher model, and a union of SGC Stockholm \nand AZ small-scale expression data is used to train the student model.  \nUsing MLC with large scale purification data (Model D) achieves in AUROC of 0.74 on AZ dataset. \nThis is an improvement of 0.01 compared to the second-best result of Model B on AZ data. When \nevaluating Model D on SGC Stockholm, AU ROC reaches 0.77, which is an improvement of 0.04 \nover the next-best result. We have also observed that the MLC model is more robust across different \nsequence clusters – training, validation and testing – than other models, which tend to overfit the \ntraining data. The MLC framework thus allows utilising large scale purification data to improve \nmodelling of small-scale expression testing, whereas simple transfer learning (Model B or Model C \ntrained on all sources) fails to achieve that outcome. \n \nProspective\texperimental\tvalidation\tof\tthe\tmodel\tshows\tAUROC\tof\t0.83\t\nTo establish the utility of RP3Net for drug discovery projects, in addition to the normal train-vali-\ndate-test model development loop, we have conducted prospective model evaluation in a real-life \nscenario. A set of 46 proteins was curated from the human proteome to include viable drug targets, \nwhilst avoiding proteins with prior published evidence of successful expression. We started by gen-\nerating two full length constructs per target (with a 6-His affinity tag placed at the N- or C-Terminal) \nand running RP3Net on them. If both constructs were predicted not to express, we generated \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n13 \ntrimmed constructs, ran these through the model, and , if they were predicted to express , included \nthem in the dataset (see Methods). This resulted in a total of ninety-seven constructs for the exper-\nimental validation dataset, eight of which were generated by the trimming process  (see Methods). \nThe constructs were cloned and expressed in E. coli at the AZ protein production facility. 54% of \nthe constructs passed small -scale expression screening, including one-step purification by affinity \nchromatography. The remaining 46% were annotated as either “Ambiguous” or “Not Passed”.  \nThe performance of RP3Net models B and D was compared with the baseline model, and two third-\nparty predictors: SoluProt58 and NetSolP44 (Table 4). The highest AUROC of 0.83 is achieved by \nRP3Net model D. This is 0.06 better than the next-best result (RP3Net B trained on AZ and SGC \nStockholm), and 0.08 better than the best third -party predictor (NetSolP useability). RP3Net D \nshowed an accuracy of 0.77 when the score cut-off of 0.5 was used, and accuracy of 0.81 with the \ncut-off set to 0.79. \nFor the subset of eight trimmed constructs, RP3Net D shows an accuracy of 0.5 with score cut-off \nof 0.5, and accuracy of 0.62 with score cut -off of 0.79. This could be an artefact of the small eval-\nuation set, or that RP3Net does not consider if sequences will fold into stable protein domains. \nCuriously, the trimmed constructs that did result in soluble protein also contained degradation prod-\nucts (supplementary figure 1). With the score cut -off of 0.5 t he model predicts all trimmed con-\nstructs to express, whereas in fact only four out of eight were expressed successfully.  \nPerformance on trimmed constructs could thus be considered an area for improvement. Ho wever, \nconsidering the small number of trimmed constructs, and the model accuracy (0.77) and precision \n(0.73) on the larger experimental validation set, it could be argued that  an experimental scientist \nwould still find the modelling results helpful. An “overconfident”, high recall, model that predicts \ntoo many positives, which are then partly confirmed in the laboratory, is preferrable to a model that \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n14 \nmisses out constructs that would have expressed in the lab. Model precision could be increased, at \nthe expense of recall, by increasing the score cut-off threshold.  \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n15 \nConclusions \nThe recombinant production of proteins can require multiple experimental rounds of trial and error. \nTo improve the efficiency of such experiments, we have developed RP3Net, an AI model of heter-\nologous protein expression in E. coli. RP3Net predicts the results of protein expression as a binary \noutcome. It was built using the latest foundational models and was trained using a combination  of \ninternal experimental results from small -scale AZ expression screens, and publicly available data \nfrom the SGC. Using an STP aggregation layer and MLC with large scale purification data enables \nRP3Net to achieve state-of-the-art performance both on the take-out data from SGC Stockholm and \nAZ. RP3Net has been experimentally validated on a manually selected set of constructs for viable \nhuman drug targets and outperformed third party predictors on that set as well. Ablation studies \nshow that there is no single method that achieves a large performance inc rease, but rather many \nsmall incremental improvements.  \nThis work also underscores the need for large and well curated datasets of soluble protein expression \nand for the scientific community to agree on how the data should be captured following the \nFAIR23,70 principles, and to establish a protein production ontology. Unfortunately, in the field of \nprotein production there is not yet an equivalent of the PDB for structural biology. Significant time \nin this project was spent on data curation.  \nThe modelling results may also be further improved by making the model more aware of the exper-\nimental conditions, such as E. coli host strain, induction methods, time and temperature at which \nvarious experimental stages were performed, buffer formulations, etc. This information is largely \nmissing from the currently available data sets.  \nRP3Net is already deployed and used by the protein scientists at AZ. This publication and the \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n16 \naccompanying code repository at GitHub1 make the model available to the wider research commu-\nnity, both in industry and in academia. \n  \n \n1 www.github.com/RP3Net/RP3Net \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n17 \nAcknowledgements \nWe would like to thank Susanne Gräslund and Opher Gileadi from SGC Stockholm, and Matthieu \nSchapira and Peter Loppnau from SGC Toronto for sharing the ir respective datasets and helping \nwith the curation. We would like to acknowledge colleagues from the Quantitative Biology and \nProtein Science departments at AstraZeneca for constructive discussions during this project, and \nDavid Öling from BioPharmaceuticals R&D at AZ for overseeing the cloning of the experimental \nconstructs. We would like to acknowledge Mat thew Hall from the Industry Partnerships team at \nEMBL-EBI and Birgit Kerber and colleagues from EMBLEM for helping to organise the collabo-\nration; and the EMBL-EBI IT team for maintaining the computational facilities used to train the \nmodels. We also acknowledge the funding from the Member States of the European Molecular Bi-\nology Laboratory (ARL). \n \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n 18 \nTables \n \nTable 1. Foundational models. \nModel Type Architecture Training Data \nSources \nTraining \nData Size \nNumber of \nParameters \nComments \nESM2 \n650M26 \nProtein Transformer UniRef50, \nUniRef90 \n65M 650M  \nESM2 3B 26 Protein Transformer UniRef50, \nUniRef90 \n65M 2.8B  \nESM371 Protein Transformer Uniref, MGnify, \nJGI IMG/M, OAS \n2.78B 1.8B Licensing re-\nstrictions apply \nESMC \n60072 \nProtein Transformer Uniref, MGnify, \nJGI IMG/M \n2.3B 575M Licensing re-\nstrictions apply \nESMC \n30072  \nProtein Transformer Uniref, MGnify, \nJGI IMG/M \n2.3B 333M Licensing re-\nstrictions apply \nProtT5 XL27 Protein Transformer UniRef50 49M 1.2B Encoder only \nProtBert27 Protein Transformer UniRef100 217M 420M  \nProtein-\nBert28 \nProtein Transformer Uniref90 106M 16M  \nHyenaDNA \nMedium29 \nDNA Hyena Human Genome \nhg38 \n3.2B bp 24M Maximum se-\nquence length \n= 450K bases \nHyenaDNA \nLarge 29 \nDNA Hyena Human Genome \nhg38 \n3.2B bp 46M Maximum se-\nquence length \n= 1M bases \nDNABert31 DNA Transformer Genomes from \n136 species be-\nlonging to 6 clas-\nses \n32.49B \nbp \n117M  \nCaLM32 Codon Transformer European Nucle-\notide Archive, \ncoding se-\nquences \n8.7M 85M  \n  \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n19 \nTable 2. RP3Net training and architecture configurations. \nRP3Net Model Aggregation  FM weights Meta Label Correction \nA Mean Frozen No \nB STP Frozen No \nC STP Fine-tuned, LoRA No \nD STP Fine-tuned, LoRA Yes \n \n  \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n20 \nTable 3. Results of evaluating different RP3Net models trained on different data sources, versus the baseline model and third party \npredictors. \nModel Trained on Tested on AUROC Accuracy  Recall Precision \nNetSolP \nSolubility \n \nAZ 0.64 0.61 0.44 0.64 \nNetSolP \nUseability \n \nAZ 0.64 0.57 0.13 0.87 \nB AZ AZ 0.59 0.52 0.85 0.50 \nB SGC Stockholm, AZ AZ 0.73 0.66 0.72 0.63 \nB SGC Stockholm, \nSGC Toronto, AZ \nAZ 0.71 0.66 0.50 0.72 \nC SGC Stockholm, \nSGC Toronto, AZ \nAZ 0.72 0.66 0.52 0.71 \nD SGC Stockholm, AZ, \nSGC Toronto \nAZ 0.74 0.70 0.71 0.69 \nNetSolP \nSolubility \n SGC Stockholm 0.48 0.66 0.14 0.42 \nNetSolP \nUseability \n SGC Stockholm 0.39 0.68 0.00 0.00 \nBaseline SGC Stockholm SGC Stockholm 0.62 0.62 0.39 0.40 \nA SGC Stockholm SGC Stockholm 0.70 0.63 0.74 0.45 \nB SGC Stockholm SGC Stockholm 0.72 0.63 0.74 0.45 \nB SGC Stockholm, AZ SGC Stockholm 0.73 0.62 0.78 0.45 \nB SGC Stockholm, \nSGC Toronto, AZ \nSGC Stockholm 0.70 0.60 0.64 0.42 \nC SGC Stockholm, \nSGC Toronto, AZ \nSGC Stockholm 0.68 0.62 0.56 0.43 \nD SGC Stockholm, AZ, \nSGC Toronto \nSGC Stockholm 0.77 0.73 0.54 0.59 \nB SGC Stockholm, \nSGC Toronto, AZ \nSGC Toronto 0.75 0.87 0.12 0.27 \nC SGC Stockholm, \nSGC Toronto, AZ \nSGC Toronto 0.76 0.88 0.08 0.33 \n \n  \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n21 \nTable 4. Results of experimental validation of RP3Net, the baseline model, and third-party predictors. \nModel Trained on AUROC Accuracy Recall Precision \nSoluprot \n \n0.64 0.57 0.83 0.57 \nNetSolp  \nSolubility \n \n0.66 0.60 0.79 0.59 \nNetSolp  \nUseability \n \n0.75 0.57 0.23 0.86 \nBaseline SGC Stockholm 0.67 0.65 0.58 0.71 \nRP3Net B AZ 0.69 0.65 0.88 0.62 \nRP3Net B AS, SGC Stockholm,  0.77 0.71 0.94 0.66 \nRP3Net B AZ, SGC Stockholm, \nSGC Toronto \n0.76 0.69 0.65 0.74 \nRP3Net D AZ, SGC Stockholm, \nSGC Toronto MLC \n0.83 0.77 0.92 0.73 \n  \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n22 \nMethods \nThe\tdataset\t\nProtein production results from AZ, SGC Stockholm and SGC Toronto were used for training and \nevaluation of the models. AZ and SGC Stockholm report the results of small -scale protein \nexpression testing, after one purification step. In the AZ dataset the outcome is reported as an \nexpression yield category, manually estimated by the scientist who  has expressed the protein. In \naddition, for a subset of constructs, an estimate of the absolute concentration value in mg/L is \nprovided. The estimate is obtained by c omparing the size and intensity of the band on the SDS -\nPAGE gel for the protein of interest with the band for the reference protein of known concentration. \nThis comparison is performed by internal image analysis software. The amino acid sequence of the \nconstruct includes affinity and solubility tags; DNA sequences are available for a subset of \nconstructs. \nThe dataset from SGC Stockholm contains genetic sequences, with tags, annotated with categorical \noutcomes. \nFor the bulk of the SGC Toronto data, the outcome is reported  as a pipeline position, similarly to \nPSI/TargetTrack. Importantly, there is no dedicated stage for expression screening:  \"cloned\" is \nimmediately followed by \"purified\". Although it can generally be assumed that a protein has to be \nexpressed before it can be purified, it is sometimes the case that producing at larger scale (expression \nvolume) can rescue a construct that failed to yield soluble protein at small scale. For a small subset \nof SGC Toronto data, small-scale expression screening outcome is also provided as a  categorical \nvariable, similarly to SGC Stockholm. Genetic sequences are available for a subset of observations, \nand tags are included in the constructs.  \nGraphical overview of the datasets is shown in supplementary figure 2. There are a total of 67,055 \nunique sequences, covering 5,712 target proteins. Publicly available datasets are significantly larger \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n23 \nthat the internal AZ data set, SGC Toronto being the largest.  Datasets vary in terms of number of \nconstructs per target, availability of genetic sequences vs protein sequences and imbalance between \npositive and negative outcomes.  \nTo normalise the data across multiple sources and to  compute the outcome imbalance, the labels \nwere converted to binary form, with \"True\" indicating successful production, and \"False\" failed \nproduction. For AZ this binary outcome was computed based on the existing category annotation, \nestimate of the absolute concentration and manual  re-annotation. For SGC Stockholm the binary \noutcome was  derived directly from the existing category annotation. For  SGC Toronto, it was \nderived from the pipeline position.  \nHistorical AZ small -scale expression screening results were reported as concentration range \nestimates: \"0 to 1\", \"1 to 10\", \"10 to 20\", \"20 to 50\", and \"above 50\" mg/L. This data was converted \nto binary outcomes as follows. Results in \"0 to 1 mg/L\" concentration range were annotated as False \n(not produced). Results within \"10 to 20\", \"20 to 50\", and \"above 50\" mg/L were annotated as True \n(produced). Results in the \"1 to 10\" range were handled in a special manner. The experiments that \nhad an estimate of absolute concentration were annotated as either True or False by comparing this \nvalue with the threshold of 3.5 mg/L. The experiments where the absolute value had not been \nestimated were re-annotated manually, by re-examining the captured image of the SDS-PAGE gel. \nAZ data that were collected after April 2023 do not contain manual estimates of the concentration \nrange. Instead, each experiment outcome is manually classified by the scientist into three categories, \nbased on the SDS-PAGE gel: \"Passed\", \"Not passed\" and \"Ambiguous\". Outcomes belonging to the \n\"Passed\" category were annotated as True, and those belonging to \"Not passed\" and \"Ambiguous\" \ncategories as False. \nSGC Stockholm reports small -scale soluble expression screening outcomes with manual numeric \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n24 \nqualitative annotations by the lab scientist: 0 – no soluble expression, 1 – low soluble expression, 2 \n– medium, 3 – high and 4 – very high soluble expression. Outcomes from categories 0 and 1 were \nannotated as False, and with categories 2 and above – as True. Small-scale soluble expression results \nfor SGC Toronto were annotated in the identical manner. \nFor the SGC Toronto data with pipeline position outcome, results marked with \"cloned\" were \nannotated as False, and those with \"purified\" and beyond - as True. \n \nCross\tvalidation\t\nTo avoid  bias towards any particular protein sequence motifs , five -fold cross validation was \nperformed. All constructs were clustered using MMseqs2 v2568829073. The affinity and solubility \ntags were removed from the constructs, and the remaining \"target\" sequences were clustered.  \nSequence clusters that contain  the AZ results recorded after 1st of September 2023, as well as the \nconstructs used for the experimental validation, were grouped together to form the test set. The \nremaining constructs were divided into five cross validation subsets , such that  each cluster is \nentirely contained within a single subset.  Five-fold leave-one-out cross validation was performed \non Model A with SGC Stockholm data, and the worst performing data split was chosen for reporting \nand for subsequent model development. \n \nThe\tbaseline\tmodel\t\nA gradient -boosted decision tree (XGBoost v2.1.3 59) was used as a baseline model. The input \nfeatures for the tree were generated by analysing the sequences with ProtParam, as well as predicting \nglobal protein properties with Schrodinger API v2021-274, DisEMBL v2.0 75 and RaptorX76. The \nfull list of features and methods used to compute them is given in supplementary table 1. For tools \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n25 \nthat take protein sequence alignments as input, those were built against the Uniprot database \ndownloaded in February 2016, clustered at 20% cutoff (uniprot20_2016_0277). \n \nAggregation\t\nTwo types of aggregation layer were tested in this work: mean pooling and Set Transformer Pooling \n(STP66,67). For notation, assume that for a sequence of length 𝑁, the output of the foundational \nmodel for each residue 𝑖 is represented with a column vector 𝑥!\n\" from a 𝑑-dimensional space, 𝑥!\n\" ∈\nℝ#×%. The matrix representation for the entire protein, 𝑋\", is obtained by stacking these residue \nrepresentations along the sequence dimension: 𝑋\" ∈ ℝ#×&. In this notation, mean pooling, which \nis just averaging all these residue representations, can be written as  \n𝑋!\n\" = 1\n𝑁 % 𝑥#\n$\n%\n#&'\n. (1) \n \nThe advantage of mean pooling is that it is simple to interpret and fast to compute. The disadvantage \nis that it does not have any trainable parameters, or weights, so all training must happen upstream, \nin the foundational model, or downstream, in the clas sification head. STP is an example of an ag-\ngregation layer with trainable weights. Here, the global representation is the result of performing \nMultiheaded Attention (MHA41) with a seed vector, 𝑤' ∈ ℝ#, as a query, and residue representa-\ntions 𝑋\" \tas keys and values. Thus, \n𝑋'()\n* = MHA(𝑤', 𝑋\" , 𝑋\"). (2) \n \nMulti-headed\tattention\t\nComputing the MHA 41 involves weight matrices  𝑊+\n,, 𝑊+\n-, 𝑊+\n. and 𝑊/ for queries, keys, values \nand outputs, respectively, where ℎ = 1. . 𝐻, and 𝐻 is the number of heads: \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n26 \nMHA(𝑤', 𝑋\" , 𝑋\") = [concat(𝐴%, 𝐴0, … , 𝐴1)𝑊/]2,\n𝐴+ = softmax 8𝑄+𝐾+\n2\n;𝑑3\n< 𝑉+,\n𝑄+ = 𝑤'2𝑊+\n,,\n𝐾+ = 𝑋\"2𝑊+\n-,\n𝑉+ = 𝑋\"2𝑊+\n..\n(3) \nHere, the matrices 𝑊+\n, ∈ ℝ#×#!  are used to project 𝑤', into 𝑑3-dimensional space,  𝑊+\n- ∈ ℝ#×#!  \nand 𝑊+\n. ∈ ℝ#×#!  – to project 𝑋\"into 𝑑3-dimensional space, and the matrix  𝑊/ ∈ ℝ+#!×# – to \nproject the concatenated single head attention outputs back to the  𝑑-dimensional space. These ma-\ntrices, as well as the seed vector 𝑤',\tare updated during training. The number of heads, 𝐻, as well \nas the inputs and outputs dimension, 𝑑, and the attention dimension, 𝑑3, are hyperparameters, that \nare chosen to maximise model performance on the validation dataset.  Matrix transposition, denoted \nby ⊤, is required to keep the inputs and outputs in column form. \nMulti-headed attention is used for the STP aggregation layer in this work, as outlined above. It is \nalso an important part of the transformer architecture, that underpins most of the foundation models. \n \nRP3Net\timplementation\tand\ttraining\t\nRP3Net was implemented with PyTorch 78. Foundation models were downloaded from \nHuggingFace79. Training loop was implemented with PyTorch Lightning 80. For models C and D, \nwhen the foundation model weights were fine -tuned during training, low -rank adaptation \n(LoRA81,82) was used. Early stopping criterion was used, whe re training is terminated if AUROC \nfor the validation dataset does not improve for 10 epochs. Exact revisions of software packages and \nfoundation models, as well as training run configurations with hyperparameter values, are available \nin the RP3Net GitHub Repo.  \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n27 \n \nMeta\tlabel\tcorrection\twith\tpurification\tdata\t\nThe Meta Label Correction (MLC  68,69) framework utilises a larger, noisy, poor-quality dataset to \naugment the training process of the model that would normally use only a smaller, clean, high-\nquality dataset. A separate, “teacher” model is trained to predict the corrected soft label from the \nnoisy data and labels. These corrected labels, along with the clean inputs and labels, are used to \ntrain the original model, which in this setup is referred to as  the “student” model ( Fig. 2B in the \nmain text). \nFormally, we can denote the clean dataset as 𝐷 ≡ {𝑋, 𝑦}, where X is the input and y is the label. The \nnoisy dataset can be denoted as 𝐷E ≡ {𝑋F, 𝑦G}, and the corrected labels – as 𝑦c. The student model that \npredicts the probability of the clean label 𝑦, based on the clean input 𝑋, is  denoted as 𝑝4 (𝑋):  \n𝑃(𝑦|𝑋) ∼ 𝑝4(𝑋), where 𝑤 are the trainable parameters. In the normal deep learning framework \nthis model is trained by minimising the loss function ℒ(𝑤) between the true labels and the predicted \nlabels over the clean dataset: \n𝑤∗ = arg  min4ℒ(𝑤), (4) \nFor binary labels that take values of 0 and 1, and cross-entropy (CE) loss, we have \nℒ(𝑤) ≡ ℒ(𝑝4 (𝑋), 𝑦) ≡ 𝐶𝐸W𝑦, 𝑝4 (𝑋)X\n= 𝑦 × logW𝑝4(𝑋)X + (1 − 𝑦) × logW1 − 𝑝4(𝑋)X\t (5) \nSimple transfer learning would work by substituting 𝑋F and 𝑦G and for 𝑋 and 𝑦, respectively, in equa-\ntions (4) and (5). Instead, in the MLC framework, the noisy labels 𝑦G are replaced by the corrected \nlabels 𝑦c, modelled by the teacher model, based on the noisy sequences and the noisy labels: \n𝑃W𝑦c|𝑋F, 𝑦GX ∼ 𝑞6W𝑋F, 𝑦GX, with parameters 𝛼. The loss function ℒa between the corrected labels and \nthe noisy input is obtained by substituting the teacher model in place of 𝑦 in the equation (5): \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n28 \nℒa(𝑤, 𝛼) ≡ 𝐶𝐸 b𝑞6 W𝑋F, 𝑦GX, 𝑝4W𝑋FXc . (6) \nThe optimal parameters of the student model 𝑤∗ now depend on the parameters of the teacher model \n𝛼:\n𝑤∗(𝛼) = arg  min4ℒa(𝑤, 𝛼). (7) \nThe optimal value 𝛼∗ needs to be determined, such that the corrected labels 𝑦7 are indeed meaning-\nful in the context of the student model, or, in other words, that the student model trained on the \nnoisy data with corrected labels p erforms well on the clean data.  This can be done by substituting \nthe optimal value of 𝑤 defined by equation (7) in the equation (4): \n𝛼∗ = arg  min6ℒ(𝑤∗(𝛼)). (8) \n \nEquations (7) and (8) form the bi-level optimisation problem, that jointly determines the parameters \nof the teacher and the student models.  \nIn the context of this work, the clean data is a union of the AZ and SGC Stockholm data sets, and \nthe noisy data is SGC Toronto with pipeline position labels. The noisy dataset is thus several times \nlarger than the clean one. On each step of the algorithm several gradient steps through the noisy \ndata (Eqn. 7) are followed by a single step through the clean data (Eqn . 8). The number of noisy \nsteps per single clean step is a hyperparameter. Putting it all together, we get Algorithm 1 for com-\nputing 𝑤∗and 𝛼∗. \nThe teacher parameters 𝛼 at step 𝑡 are updated by computing the gradient 𝑔6\n(9) of the clean loss ℒ \nwith respect to (w.r.t) 𝛼. This gradient can be approximated by a formula involving the gradient of \nthe clean loss w.r.t student parameters 𝑤 at step 𝑡 + 1, 𝑔4\n(9;%), and the matrices of second derivatives \n(Hessian matrices) of the noisy loss w.r.t 𝑤 and 𝛼 at previous steps, 𝐻46\n(<) =\n=\"\n=4 =6 ℒaW𝑤(<), 𝛼(<)X: \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n29 \n𝑔6\n(>)2 ≈ −𝜂4𝑔4\n(9;%)2 k (1 − 𝜂4 )9?<\n9\n<@9?A;%\n𝐻46\n< . (9) \nThis assumes that gradients are represented as column vectors. For the special case of 𝑘 = 1, the  \nAlgorithm 1 Bi-level optimisation of teacher and student model parameters via stochastic gradient descent \nInput: Clean and noisy datasets 𝐷 and 𝐷E; number of training steps\tT; initial parameters  𝑤(0), 𝛼(0) ; \nlearning rates 𝜂+ , 𝜂,, number of noisy steps per clean step k.\t\nOutput: Optimised parameters 𝑤(T), 𝛼(T). \n1 for t\t=\t0,\t...,\tT\t–\t1 do  \n2  {𝑋, 𝑦} ← 𝑆𝑎𝑚𝑝𝑙𝑒(𝒟);\t{𝑋=, 𝑦>} ← 𝑆𝑎𝑚𝑝𝑙𝑒?𝒟@A  // sample the minibatches of clean and noisy data \n3  𝑤(./') ← 𝑤(.) − 𝜂+∇+ℒE?𝑤(.), 𝛼(.)A // update w by descending the noisy loss w.r.t. w \n4  if t mod k\t=\tk\t−\t1  then \n5   𝑔, = ∇,ℒ H𝑤(./')(𝛼)I // unroll 𝑤./'\tand approximate the gradient of the clean loss w.r.t. 𝛼 \n6   𝛼(./') ← 𝛼(.) − 𝜂,𝑔, // update 𝛼 by descending the clean loss w.r.t. 𝛼 \n7  else \n8   𝛼(./') ← 𝛼(.) \n9  end if \n10 end for \n  \nsum in equation (9) is reduced to just 𝐻46\n(9) ; 𝑔6\n(>)2 = −𝜂4𝑔4\n(9;%)2𝐻469  .  \nThe CE loss allows for efficient computation of the Hessian 𝐻46\t , by expressing it point-wise as a \nproduct of Jacobians, and averaging over the minibatch: \n𝐻46 = 1\n𝑁 k[𝐽4(𝑖)]2\n&\n!@%\n[𝐽6(𝑖)], (10) \nHere, 𝐽4 (𝑖) is the Jacobian (matrix of derivatives) of the student loss w.r.t 𝑤, and 𝐽6 (𝑖) – the Jaco-\nbian of the teacher loss w.r.t 𝛼 at input 𝑖, and 𝑁 is the size of the minibatch. \nTarget\tselection\tfor\texperimental\tvalidation\t\nThe target set for experimental validation of the model was curated to include viable human drug \ntargets and exclude proteins that are well known from literature to be successfully expressed. We \nmade sure that neither the protein itself, nor its close homologs, have been deposited in the PDB52,53. \nWe have also excluded the target from the validation set if it was referenced from ChEMBL 83. \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n30 \nOpenTargets84 was used to check for viability of a  drug target. Twenty thousand human protein s \nfrom UniProt34 were narrowed down to 454 viable targets. These targets were further curated man-\nually, to have distribution across different target classes and avoid too many DNA-binding proteins. \nIn the end, 46 targets were selected for experimental validation.  \nTwo full length constructs were created per target: one with a TEV -cleavable 6His tag and a GS \nlinker at the N-terminal (MHHHHHHENLYFQGS...), and another one with a GS linker and a 6His \ntag at the C -terminal (...GSHHHHHH). Soluble production of the full-length constructs was pre-\ndicted with RP3Net. For the targets where both full-length constructs were predicted to fail to be \nproduced, trimmed constructs were generated by iteratively removing residues from N - and C-ter-\nmini, with a minimum construct length of  50. Trimmed constructs that were predicted to express \nsuccessfully were included in the experimental validation set. 70% of the set comprised constructs \nthat were predicted to be produced, with the remaining 30% as negative controls. A total of 97 \nconstructs were available for expression testing after taking into account cost constraints and clon-\ning errors. \n \nExperimental\tprocedures\tfor\tconstruct\texpression.\t\t\nAll sequences were codon optimised for E. coli and synthesized as synthetic genes (Life Technolo-\ngies Europe BV) and cloned into backbone vector pET24a. One construct failed during the cloning \nprocess. Amino acid and nucleotide sequences of synthesized constructs are found in supplementary \ntable x. \nFor small-scale soluble expression screening, the plasmid DNA was transformed into competent \nphage resistant E. coli BL21(DE3) cells (New England Biolabs #C2527H) in 96 -well PCR plates. \nThe transformation mix was used to directly inoculate 3mL LB media supplemented with 100ug/mL \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n31 \nkanamycin in 24 deep-well plates and left shaking at 37°C overnight. Protein expression was auto \ninduced in rich ZYP-8012 media supplemented with 100ug/mL kanamycin in 24 deep-well plates, \nby inoculating 3mL with 50uL pre-culture and left shaking for 3 h at 37°C followed by 24 h at 18°C. \nAfter harvest (4000xg, 5min, 4°C), the pellets were lysed with 900uL lysis buffer (40mM HEPES, \n300mM NaCl, 5mM imidazole, 10% glycerol, 1mM TCEP, 0.1% DDM, 0.2mg/mL lysozyme, \nDNAse & protease inhibitors) and freeze -thawed once. The lysate was cleared by centrifugation \n(4000xg, 30min, 4°C) before subjecting to a one -step Nickel affinity purification using an auto-\nmated bead-based platform. The protein was captured on the magnetic beads for 30min at 4°C, \nfollowed by two wash st eps to wash off unbound proteins (40mM HEPES, 300mM NaCl, 5mM \nimidazole, 10% glycerol, 1mM TCEP) and eluted in 100uL elution buffer (40mM HEPES, 300mM \nNaCl, 300mM imidazole, 10% glycerol, 1mM TCEP). 10uL of the elution was loaded onto Nu-\nPAGE Bis- Tris gels (Invitrogen) together with Novex Pre- stained protein marker (Invitrogen) and \n5ug of an internal standard protein. The gels were stained in Der Blaue Jonas (GRP) and analysed \nusing the densitometry software Image Lab (BioRad). The protein yield (mg/L) was estimated from \nthe relative quantity. The “Passed”, “Not Passed” and “Ambiguous” outcome annotations were pro-\nvided manually by the lab scientist, based on the relative thickness and brightness of gel bands. \nAnnotations from two separate biological replicates are shown in supplementary Table X. Gels from \none of the two experiments are shown in Supplementary Figure X.  The maximum yield from the \ntwo experimental runs was used as the ground truth for the model evaluation.\n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n 32 \nReferences: \n1. Zanders, E. D. The Science and Business of Drug Discovery: Demystifying the Jargon: Second Edition. The Science \nand Business of Drug Discovery: Demystifying the Jargon: Second Edition (Springer International Publishing, \n2020). \n2. Singh, N. et al. Drug discovery and development: introduction to the general public and patient groups. Frontiers in \nDrug Discovery 3, 1201419 (2023). \n3. Structural Genomics Consortium et al. Protein production and purification. Nat Methods 5, 135–46 (2008). \n4. Gräslund, S. et al. The use of systematic N- and C-terminal deletions to promote production and structural studies of \nrecombinant proteins. Protein Expr Purif 58, 210–221 (2008). \n5. Burgess-Brown, N. A. et al. Codon optimization can improve expression of human genes in Escherichia coli: A \nmulti-gene study. Protein Expr Purif 59, 94–102 (2008). \n6. Hayashi, K. & Kojima, C. pCold-GST vector: A novel cold-shock vector containing GST tag for soluble protein \nproduction. Protein Expr Purif 62, 120–127 (2008). \n7. Gordon, E. et al. Effective high-throughput overproduction of membrane proteins in Escherichia coli. Protein Expr \nPurif 62, 1–8 (2008). \n8. Haacke, A., Fendrich, G., Ramage, P. & Geiser, M. Chaperone over-expression in Escherichia coli: Apparent in-\ncreased yields of soluble recombinant protein kinases are due mainly to soluble aggregates. Protein Expr Purif 64, \n185–193 (2009). \n9. Raymond, A. et al. Combined protein construct and synthetic gene engineering for heterologous protein expression \nand crystallization using Gene Composer. BMC Biotechnol 9, 1–15 (2009). \n10. Francis, D. M. & Page, R. Strategies to Optimize Protein Expression in E. coli. Curr Protoc Protein Sci 61, 5.24.1-\n5.24.29 (2010). \n11. Zhong, N. et al. Optimizing Production of Antigens and Fabs in the Context of Generating Recombinant Antibodies \nto Human Proteins. PLoS One 10, e0139695 (2015). \n12. Cooper, C. D. O. & Marsden, B. D. N- and C-Terminal Truncations to Enhance Protein Solubility and Crystalliza-\ntion: Predicting Protein Domain Boundaries with Bioinformatics Tools. in Methods in Molecular Biology vol. 1586 \n11–31 (Humana Press Inc., 2017). \n13. Söderberg, J. J., Grgic, M., Hjerde, E. & Haugen, P. Aliivibrio wodanis as a production host: Development of ge-\nnetic tools for expression of cold-active enzymes. Microb Cell Fact 18, 1–16 (2019). \n14. Strain-Damerell, C., Mahajan, P., Fernandez-Cid, A., Gileadi, O. & Burgess-Brown, N. A. Screening and Produc-\ntion of Recombinant Human Proteins: Ligation-Independent Cloning. Methods in Molecular Biology 2199, 23–43 \n(2021). \n15. Burgess-Brown, N. A. et al. Screening and Production of Recombinant Human Proteins: Protein Production in E. \ncoli. Methods in Molecular Biology 2199, 45–66 (2021). \n16. Chapple, S. D. & Dyson, M. R. High-Throughput Expression Screening in Mammalian Suspension Cells. Methods \nin Molecular Biology 2199, 117–125 (2021). \n17. Mahajan, P. et al. Screening and Production of Recombinant Human Proteins: Protein Production in Insect Cells. \nMethods in Molecular Biology 2199, 67–94 (2021). \n18. Mahajan, P. et al. Expression Screening of Human Integral Membrane Proteins Using BacMam. Methods in Molec-\nular Biology 2199, 95–115 (2021). \n19. Simm, D., Popova, B., Braus, G. H., Waack, S. & Kollmar, M. Design of typical genes for heterologous gene ex-\npression. Sci Rep 12, 1–14 (2022). \n20. Morão, L. G., Manzine, L. R., Clementino, L. O. D., Wrenger, C. & Nascimento, A. S. A scalable screening of E. \ncoli strains for recombinant protein expression. PLoS One 17, e0271403 (2022). \n21. Kurashiki, R. et al. Development of a thermophilic host–vector system for the production of recombinant proteins at \nelevated temperatures. Appl Microbiol Biotechnol 107, 7475–7488 (2023). \n22. Acton, T. B. et al. Preparation of Protein Samples for NMR Structure, Function, and Small-Molecule Screening \nStudies. Methods Enzymol 493, 21–60 (2011). \n23. Edfeldt, K. et al. A data science roadmap for open science organizations engaged in early-stage drug discovery. Na-\nture Communications 2024 15:1 15, 1–10 (2024). \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n33 \n24. Sulkowski, E. Purification of proteins by IMAC. Trends Biotechnol 3, 1–7 (1985). \n25. Lin, Z. et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science (1979) \n379, 1123–1130 (2023). \n26. Rives, A. et al. Biological structure and function emerge from scaling unsupervised learning to 250 million protein \nsequences. Proc Natl Acad Sci U S A 118, e2016239118 (2021). \n27. Elnaggar, A. et al. ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning. \nIEEE Trans Pattern Anal Mach Intell 44, 7112–7127 (2022). \n28. Brandes, N., Ofer, D., Peleg, Y., Rappoport, N. & Linial, M. ProteinBERT: a universal deep-learning model of pro-\ntein sequence and function. Bioinformatics 38, 2102–2110 (2022). \n29. Nguyen, E. et al. HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide Resolution. ArXiv \n(2023). \n30. Ji, Y., Zhou, Z., Liu, H. & Davuluri, R. V. DNABERT: pre-trained Bidirectional Encoder Representations from \nTransformers model for DNA-language in genome. Bioinformatics 37, 2112–2120 (2021). \n31. Zhou, Z. et al. DNABERT-2: Efficient Foundation Model and Benchmark For Multi-Species Genome. ArXiv \n(2023). \n32. Outeiral, C. & Deane, C. M. Codon language embeddings provide strong signals for use in protein engineering. Nat \nMach Intell 6, 170–179 (2024). \n33. Suzek, B. E., Wang, Y., Huang, H., McGarvey, P. B. & Wu, C. H. UniRef clusters: a comprehensive and scalable \nalternative for improving sequence similarity searches. Bioinformatics 31, 926 (2015). \n34. Bateman, A. et al. UniProt: the Universal Protein Knowledgebase in 2023. Nucleic Acids Res 51, D523–D531 \n(2023). \n35. Benson, D. A. et al. GenBank. Nucleic Acids Res 41, D36–D42 (2012). \n36. Sayers, E. W. et al. GenBank 2024 Update. Nucleic Acids Res 52, D134–D137 (2024). \n37. Richardson, L. et al. MGnify: the microbiome sequence data analysis resource in 2023. Nucleic Acids Res 51, \nD753–D759 (2023). \n38. Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021). \n39. Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for \nLanguage Understanding. Proceedings of the Conference of the North American Chapter of the Association for \nComputational Linguistics: Human Language Technologies 1, 4171–4186 (2018). \n40. Hayes, T. et al. Simulating 500 million years of evolution with a language model. bioRxiv (2024) \ndoi:10.1101/2024.07.01.600583. \n41. Vaswani, A. et al. Attention Is All You Need. Adv Neural Inf Process Syst 2017-December, 5999–6009 (2017). \n42. Rao, R. et al. Evaluating Protein Transfer Learning with TAPE. in Advances in Neural Information Processing Sys-\ntems 9686–9698 (2019). \n43. Dallago, C. et al. FLIP: Benchmark tasks in fitness landscape inference for proteins. in Thirty-fifth Conference on \nNeural Information Processing Systems Datasets and Benchmarks Track (Round 2) (2021). \n44. Thumuluri, V. et al. NetSolP: predicting protein solubility in Escherichia coli using language models. Bioinformat-\nics 38, 941–946 (2022). \n45. Wang, S. & Zhao, H. SADeepcry: a deep learning framework for protein crystallization propensity prediction using \nself-attention and auto-encoder networks. Brief Bioinform 23, (2022). \n46. Xiong, D., U, K., Sun, J. & Cribbs, A. P. PLMC: Language Model of Protein Sequences Enhances Protein Crystalli-\nzation Prediction. Interdiscip Sci 1–12 (2024) doi:10.1007/s12539-024-00639-6. \n47. Li, F.-Z., Amini, A. P., Yue, Y., Yang, K. K. & Lu, A. X. Feature Reuse and Scaling: Understanding Transfer \nLearning with Protein Language Models. in Proceedings of the 41st International Conference on Machine Learning \n(eds. Salakhutdinov, R. et al.) vol. 235 27351–27375 (PMLR, 2024). \n48. Berman, H. M. et al. Protein Structure Initiative - TargetTrack  2000-2017 - all data files. https://zenodo.org/rec-\nords/821654 https://doi.org/10.5281/zenodo.821654 (2017) doi:10.5281/zenodo.821654. \n49. Gabanyi, M. J. et al. The Structural Biology Knowledgebase: a portal to protein structures, sequences, functions, \nand methods. J Struct Funct Genomics 12, 45–54 (2011). \n50. Protein production data from the SGC. Preprint at https://www.ebi.ac.uk/biostudies/SGC/studies/S-BSST681 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n34 \n(2022). \n51. Savitsky, P. et al. High-throughput production of human proteins for crystallization: The SGC experience. J Struct \nBiol 172, 3–13 (2010). \n52. Berman, H. M. The Protein Data Bank. Nucleic Acids Res 28, 235–242 (2000). \n53. Burley, S. K. et al. RCSB Protein Data Bank (RCSB.org): delivery of experimentally-determined PDB structures \nalongside one million computed structure models of proteins from artificial intelligence/machine learning. Nucleic \nAcids Res 51, D488–D508 (2023). \n54. Mizianty, M. J. & Kurgan, L. Sequence-based prediction of protein crystallization, purification and production pro-\npensity. Bioinformatics 27, i24–i33 (2011). \n55. Wang, H. et al. PredPPCrys: Accurate Prediction of Sequence Cloning, Protein Production, Purification and Crys-\ntallization Propensity from Protein Sequences Using Multi-Step Heterogeneous Feature Fusion and Selection. PLoS \nOne 9, e105902 (2014). \n56. Wang, H. et al. Crysalis: an integrated server for computational analysis and design of protein crystallization. Sci \nRep 6, 21383 (2016). \n57. Zhu, Y.-H. et al. Accurate multistage prediction of protein crystallization propensity using deep-cascade forest with \nsequence-based features. Brief Bioinform 22, (2021). \n58. Hon, J. et al. SoluProt: prediction of soluble protein expression in \\textit{Escherichia coli}. Bioinformatics 37, 23–\n28 (2021). \n59. Friedman, J. H. Greedy function approximation: A gradient boosting machine. The Annals of Statistics 29, (2001). \n60. Sormanni, P., Aprile, F. A. & Vendruscolo, M. The CamSol Method of Rational Design of Protein Mutants with \nEnhanced Solubility. J Mol Biol 427, 478–490 (2015). \n61. Sormanni, P., Amery, L., Ekizoglou, S., Vendruscolo, M. & Popovic, B. Rapid and accurate in silico solubility \nscreening of a monoclonal antibody library. Scientific Reports 2017 7:1 7, 1–9 (2017). \n62. Yuan, Q. et al. GPSFun: geometry-aware protein sequence function predictions with language models. Nucleic Ac-\nids Res 52, W248–W255 (2024). \n63. Zhang, X. et al. PLM_Sol: predicting protein solubility by benchmarking multiple protein language models with the \nupdated Escherichia coli protein solubility dataset. Brief Bioinform 25, (2024). \n64. Basafa, M., Hashemi, A. & Behravan, A. Optimizing recombinant antibody fragment production: A comparison of \nartificial intelligence and statistical modeling. Biotechnol Appl Biochem 71, 1094–1104 (2024). \n65. Zhang, J. H., Shan, L. L., Liang, F., Du, C. Y. & Li, J. J. Strategies and Considerations for Improving Recombinant \nAntibody Production and Quality in Chinese Hamster Ovary Cells. Front Bioeng Biotechnol 10, 856049 (2022). \n66. Buterez, D., Janet, J. P., Kiddle, S. J., Oglic, D. & Liò, P. Graph Neural Networks with Adaptive Readouts. in Ad-\nvances in Neural Information Processing Systems (eds. Koyejo, S. et al.) vol. 35 19746–19758 (Curran Associates, \nInc., 2022). \n67. Lee, J. et al. Set Transformer: A Framework for Attention-based Permutation-Invariant Neural Networks. in Pro-\nceedings of the 36th International Conference on Machine Learning (eds. Chaudhuri, K. & Salakhutdinov, R.) vol. \n97 3744–3753 (PMLR, 2019). \n68. Taraday, M. K. & Baskin, C. Enhanced Meta Label Correction for Coping with Label Corruption. ArXiv (2023). \n69. Zheng, G., Awadallah, A. H. & Dumais, S. Meta Label Correction for Noisy Label Learning. in Proceedings of the \nAAAI Conference on Artificial Intelligence (AAAI) (2021). \n70. Wilkinson, M. D. et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data 3, \n160018 (2016). \n71. Hayes, T. et al. Simulating 500 million years of evolution with a language model. Science (1979) 387, 850–858 \n(2025). \n72. ESM Team. ESM Cambrian: Revealing the mysteries of proteins with unsupervised learning. Evolutionary Scale \nWebsite https://www.evolutionaryscale.ai/blog/esm-cambrian (2024). \n73. Steinegger, M. & Söding, J. MMseqs2 enables sensitive protein sequence searching for the analysis of massive data \nsets. Nat Biotechnol 35, 1026–1028 (2017). \n74. Sankar, K. et al. A Descriptor Set for Quantitative Structure‐property Relationship Prediction in Biologics. Mol In-\nform 41, (2022). \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint \n\n \n \n35 \n75. Linding, R. et al. Protein Disorder Prediction. Structure 11, 1453–1459 (2003). \n76. Wang, S., Li, W., Liu, S. & Xu, J. RaptorX-Property: a web server for protein structure property prediction. Nucleic \nAcids Res 44, W430–W435 (2016). \n77. Mirdita, M. et al. Uniclust databases of clustered and deeply annotated protein sequences and alignments. Nucleic \nAcids Res 45, D170–D176 (2016). \n78. Paszke, A. et al. PyTorch: An Imperative Style, High-Performance Deep Learning Library. in Advances in Neural \nInformation Processing Systems 32 (Curran Associates, Inc., 2019). \n79. Wolf, T. et al. Transformers: State-of-the-Art Natural Language Processing. in EMNLP 2020 - Conference on Em-\npirical Methods in Natural Language Processing, Proceedings of Systems Demonstrations 38–45 (Association for \nComputational Linguistics (ACL), 2020). doi:10.18653/V1/2020.EMNLP-DEMOS.6. \n80. Falcon, W. & The PyTorch Lightning team. PyTorch Lightning. Preprint at (2019). \n81. Yu, Y. et al. Low-Rank Adaptation of Large Language Model Rescoring for Parameter-Efficient Speech Recogni-\ntion. in 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) 1–8 (IEEE, 2023). \ndoi:10.1109/ASRU57964.2023.10389632. \n82. Mangrulkar, S. et al. PEFT: State-of-the-art Parameter-Efficient Fine-Tuning methods. Preprint at \nhttps://github.com/huggingface/peft (2022). \n83. Zdrazil, B. et al. The ChEMBL Database in 2023: a drug discovery platform spanning multiple bioactivity data \ntypes and time periods. Nucleic Acids Res 52, D1180–D1192 (2024). \n84. Ochoa, D. et al. The next-generation Open Targets Platform: reimagined, redesigned, rebuilt. Nucleic Acids Res 51, \nD1353–D1359 (2023). \n  \n \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2025.05.13.652824doi: bioRxiv preprint","source_license":"CC-BY-4.0","license_restricted":false}