{"paper_id":"10672abd-ae30-44f7-bc80-c2d19913dd46","body_text":"1 \nStructural bias in machine learning-guided peptide design \nVictor Daniel Aldas-Bulos2,4 and Fabien Plisson1,2,3,5,*. \n \n1 School of Chemistry, the University of Sydney, NSW 2006, Australia. \n2 Centro de Investigación y Estudios Avanzados del IPN (CINVESTAV), Unidad de \nGenómica Avanzada, Laboratorio Nacional de Genómica para la Biodiversidad \n(Langebio), Irapuato, Guanajuato 36824, Mexico. \n3 CINVESTAV, Unidad Irapuato, Departamento de Biotecnología y Bioquímica, \nIrapuato, Guanajuato 36824, Mexico.  \n4 Stowers Institute for Medical Research, Kansas City, MO, USA  \n5 Ingenie Bio, Sydney, NSW 2025, Australia. \n \n*Corresponding author \nE-mail address: fabien.plisson@sydney.edu.au \nV.D.A.-B.: https://orcid.org/0000-0002-7569-775X  \nF.P.: https://orcid.org/0000-0003-2246-9347 \n \n  \nABSTRACT (250 words) \nMachine learning continues to accelerate peptide and protein  design through the rapid \nprediction and generation of sequences with desired characteristics. Many applications focus \non predicting properties, functions, and structures, as well as generating point mutations and \nde novo designs. Nevertheless, many models prove less generalizable than initially claimed. \nMost predictors and generators are trained on sequential datasets, where imbalances can be \naddressed during preprocessing. In contrast,  structural bias, a subtype of algorithmic bias \narising from  uneven representation of structural classes in training datasets, and the \nlimitations of early protein structure predictors ha ve frequently remained undetected and \nuncorrected. The recent surge in powerful protein structure prediction tools, such as the \nAlphaFold and RosettaFold series and their variants, now presents opportunities to mitigate \nthis issue. We hypothesize that such structural sampling biases influence the downstream \nperformance of ML models . Using antimicrobial peptides as a case study,  we audited the \nstructural biases in 16 state-of-the-art predictors for antimicrobial activity and tested whether \nstructural information constrains their predicti ons. Our analysis revealed  that models \nexplicitly trained on s equential data still produce predictions biased by uneven fold  \nrepresentations and data leakage . These findings highlight the importance of  integrating \nbalanced structural data or implementing bias -mitigating strategies  to develop agnostic \nmodels that maximize bioactive protein discovery and multi-objective optimization. \n \nKeywords \nProtein design, machine learning, antimicrobial peptides, algorithmic bias, data leakage. \n  \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint \n\n2 \n1. INTRODUCTION \nFacial recognition systems that misidentify people with darker skin tones or women1, \nlending and credit scoring models  that disadvantage minority or low -income \ncommunities2, and image generators that reproduce stereotypes embedded in the \ntraining data3 illustrate the pervasive nature of algorithmic bias in our society. These \nexamples demonstrate  that bias can emerge from human decision -making and  the \ncomputational systems we design, often mirroring and exacerbating pre-existing \ninequalities. Algorithmic bias refers to systematic biases in the outputs of  \ncomputational models or decision -making systems, arising from the data on which \nthey are trained, the algorithms they use, or the contexts in which they operate. \nThe consequences of algorithmic bias in domains such as finance, law enforcement, \nand media production are increasingly recognized, prompting governments and \nprivate organizations to implement policies and guardrails that support the safe and \nresponsible use of artificial intelligence (AI) technologies4,5. By comparison, the effects \nof algorithmic bias in the natural sciences, including computational biology, have \nreceived less attention. Over the past decade, the field of protein biology has benefited \nfrom machine learning (ML) methods for engineering, structure prediction, functional \nannotation, and de novo  design6–13. However, because these models are trained on \nprotein sequence or structural databases , they inherit the historical, methodological, \nand experimental biases embedded in these resources, shaped by decades of research \nfunding priorities, experimental detection limits, taxonomic preferences, and domain-\nspecific data availability14.  \nRecognizing and addressing such biases early is essential to develop ing robust, \ngeneralizable, and reproducible applications in protein engineering, drug discovery, \nand synthetic biology15. It may also present significant biosecurity risks. In line with \nthese objectives, the Responsible AI x Biodesign initiative provides a global framework \nthat outlines  community values, guiding principles, and commitments to  ensure \nsound AI-guided protein design practices16. \nEarly evidence of algorithmic bias in protein biology has emerged across sequence-, \nstructure-, and function -prediction ML models. Sequence-based approaches rely on \ntraining datasets with phylogenetic and annotation biases that overrepresent certain \nmodel organisms, protein families, and conserved domains, thereby reducing model \nperformance on underrepresented sequences17–20. Structure-based models exhibit \nperformance variations due to imbalanced structure classes, fold-switching proteins, \nand intrinsically disordered regions in structural repositories 21–25. Functional \nannotation models propagat e historical misannotations from homology -based \nmethods, overlooking novel functions 26–28. These biases can distort predictions  and \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint \n\n3 \ngenerations, narrow the exploration of the sequence-function landscape to dominant \nprotein classes, and limit the discovery of rare folds29–31,24.  \nTo examine structural bias, a subtype of algorithmic bias, we recently applied a set of \nprotein structure predictors, including AlphaFold29, to map the structural landscape \nof medium-to-large-sized datasets using the GRAMPA repository of 5,980 \nantimicrobial peptides32. Our structural mapping revealed that most peptides (65.1%) \nadopted loose helices, whereas fewer formed random coils (17.8%)  or b-stranded or \nmixed structures ( 17.1%) – mimicking folds observed in X-ray crystallography and \naqueous NMR solutions . Because many state -of-the-art antimicrobial peptide  \npredictors are trained on GRAMPA or similar repositories1, we evaluated 16 models \nto assess whether the latent structural composition influences their predictions and to \nexplore how structural features shape predictive bias. \n \n2. MATERIALS AND METHODS \n2.1. Datasets.  \nGRAMPA. We obtained the peptide sequences from the GRAMPA repository (Giant \nRepository of AMP Activities), a robust database established in 2018 containing \nsequences of 5,980 peptides ranging from 5 to 50 residues.33 The GRAMPA repository \nand detailed information are available at  https://github.com/zswitten/Antimicrobial-\nPeptides.  \nNon-GRAMPA. We retrieved 3,385 peptide sequences ranging from 5 to 50 residues \nfrom the UniProt database 34. These sequences had no reported antimicrobial activity \nand originated from taxa known to produce antimicrobial peptides:  Arthropoda \n(n=1,867), Mammalia (n=501), Violaceae (n=114), Amphibia (n=322), Rubiaceae (n=49), \nCucurbitaceae (n=60), and Mollusca (n=472). The search excluded entries labeled as \npartial, putative, or predicted, and those annotated with keywords Antimicrobial [KW-\n0929], Antiviral, Anticancer, Antibiotic [KW-0044], or Fungicide [KW-0295], while \nrestricting results to reviewed records (reviewed: yes). \n \n2.2. Preprocessing.  \nStructural predictions. We used PEP2D (https://webs.iiitd.edu.in/raghava/pep2d)35 to \npredict the secondary structures of all peptides. Each peptide  is represented as a \nvector of  three-state values (H: helices, E: extended strand/β -sheets, and C: coils) \nexpressed as percentages. The collection of peptide structure predictions forms the \ncorresponding structural landscape, depicted in a ternary plot. \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint \n\n4 \nStructural classes. We defined seven structural regions (or classes) using the three states \n(H, E, C) and arbitrary borders at 80:0:20, 20:0:80, and 20:80:0. \nRemoval. We removed duplicated sequences , those containing non-canonical or \nunknown amino acids (represented as X), a nd sequences common to both the \nGRAMPA and non-GRAMPA datasets.  \nRedundancy. We used the CD-HIT36 web server (https://github.com/weizhongli/cdhit-\nweb-server) to remove highly redundant sequences from both datasets, grouping \nsequences with a 70% identity threshold and retaining only the most representative \nsequence from each group. \n \n2.3. Model selection, evaluation, and interpretation. \nAMP predictors . We selected 16 predictive models  for antimicrobial activity : AMP \nscanner v2 37, amPEPpy38, AMPlify variants39, CAMPr3 variants40, DBAASP41, IAMPE \nvariants42, iAMPpred43, PepNet44, and Sense the Moment 45 – see Table S1. \nEvaluating the model performances. Each model was evaluated as a binary classifier, with \nantimicrobial peptides (AMPs) as signed to  class 1 and non -antimicrobial peptides \n(non-AMPs) to class 0. Depending on the architecture, the model outputs were \ninterpreted as either discrete class labels, class probabilities (e.g., p(sequence X | AMP) \n= 0.67), or both. For models that produced only  class probabilities, sequences with \nvalues greater than 0.50 were labeled as AMPs. For each structural class, we computed \nthe confusion matrix, recording the number of True Positives (TP; correctly predicted \nactive AMP sequences), True Negatives (TN; correctly predicted non-AMP sequences \nas inactive), False Positives (FP; non-AMP sequences incorrectly predicted as AMPs), \nand False Negatives (FN; AMP sequences incorrectly predicted as non-AMPs). From \nthese, we derived standard performance  metrics, including accuracy, precision, \nspecificity, sensitivity, and Matthews Correlation coefficient (MCC). \n \n(1) Accuracy, the fraction of correct predictions obtained by the model, defined as:  \n  \t\n𝐴𝑐𝑐. = \t 𝑇𝑃 + 𝑇𝑁\n𝑇𝑃 + 𝑇𝑁 + 𝐹𝑃 + 𝐹𝑁 \n(2) Precision, the fraction of positive predictions correctly identified: \n𝑃𝑟𝑒𝑐. = \t 𝑇𝑃\n𝑇𝑃 + 𝐹𝑃 \n(3) Specificity, the fraction of negative predictions correctly identified: \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint \n\n5 \n𝑆𝑝𝑒𝑐. = \t 𝑇𝑁\n𝑇𝑁 + 𝐹𝑁 \n(4) Sensitivity, the fraction of true positives correctly identified:    \n𝑆𝑒𝑛𝑠. = \t 𝑇𝑃\n𝑇𝑃 + 𝐹𝑁 \n(5) Matthews Correlation coefficient integrates information from all four elements of \nthe confusion matrix, yielding a value between -1 and 1. Values of 1 and -1 indicate \nperfect and completely erroneous classification, respectively, whereas a value of 0 \nreflects random prediction. \n𝑀𝐶𝐶 = \t 𝑇𝑃 ∗ 𝑇𝑁 − 𝐹𝑃 ∗ 𝐹𝑁\n5(𝑇𝑃 + 𝐹𝑃) ∗ (𝑇𝑃 + 𝐹𝑁) ∗ (𝑇𝑁 + 𝐹𝑃) ∗ (𝑇𝑁 + 𝐹𝑁)\n \n \nEmbeddings and Dimensionality reduction of peptide datasets.  We encoded the structures \nof the training sets (amPEPpy, AMP scanner v2 , AMPlify, PepNet) and the external \nvalidation sets (GRAMPA and non -GRAMPA) using ProstT 46, a sequence -structure \nencoder that generates 3Di tokens as proxies for three -dimensional representations. \nThe resulting structural embeddings were projected onto a two-dimensional UMAP47 \nmanifold optimized to separate the structural classes – see Figure S6.  \n \n  \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint \n\n6 \n3. RESULTS \nMachine learning algorithms have become increasingly complex, gaining in \nperformance but often at the expense of  interpretability. Many existing models lack \ndetailed descriptions of the training data and algorithms, which limits their \nreproducibility and contributes to their perception as “black boxes” . This limited \ntransparency complicates the evaluation of  potential biases and the robustness of \nmodel predictions. To address these concerns, explainable AI techniques have gained \ntraction, and external (post-hoc) validation on independent datasets is a robust strategy \nfor uncovering inherent biases. \nIn our study, we assembled two curated datasets to critically assess the performance \nand potential biases of existing binary classifiers for predicting antimicrobial peptide \nactivity. The independent validation set comprised  5,980 antimicrobial peptide \nsequences from the GRAMPA repository 33 and 3,385 non-antimicrobial peptide \nsequences from the UniProt database (see Materials and Methods ). To explore the \nsequence-structure relationships with these sets, we predicted the secondary structure \nof all peptides using the PEP2D web server35. We then assigned each peptide to one of \nseven structural classes  based on defined thresholds  for the fractions of helix ( H), \nextended strand/β-sheet (E), and coil (C). \n \n3.1. Structural composition of the validation sets. \nThe GRAMPA repository is enriched in  α-helical and coiled structures, consistent \nwith PEP2D and AF2 predictions, which were experimentally validated, as previously \nreported32. To assess structural differences from non-antimicrobial peptides, we \npredicted the secondary structures of the  non-GRAMPA dataset. Initial analysis \nrevealed a limited presence of β-sheet conformations in non-GRAMPA sequences. To \nbetter capture this structural class , we supplemented the dataset with  53 sequences \nexplicitly annotated as β-sheet structures, bringing the total to 3,438.  \nWe subsequently removed duplicate sequences and those containing non-canonical \nor ambiguous residues, yielding 2,511 representative non-GRAMPA peptides. Both \nGRAMPA and non-GRAMPA datasets were further subjected to redundancy \nreduction via  CD-HIT36 clustering at 70% sequence identity . This process reduced \nGRAMPA to 1,731 sequences and non-GRAMPA to 1,185 sequences , thereby \neliminating highly homologous peptides while preserving sequence diversity.  \nImportantly, this reduction did not substantially alter the distribution of structural \nclasses across the datasets. The GRAMPA subset remained dominated by α -helices \nand coils, accounting for 64.6% of sequences, with 15.1% belonging to β -sheet/coiled \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint \n\n7 \nstructures and 14.1% to coils ( Table 1 ). Conversely, the non -GRAMPA subset \nexhibited a more balanced structural composition, comprising 35.9% helices, 32.7% β-\nsheets/coils, and 25.6% coils. The structural landscapes of both datasets are illustrated \nin Figure 1.  \n \nTable 1. Distributions of predicted structural classes (1)-(7) for the GRAMPA and non-\nGRAMPA datasets using protein secondary structure predictor PEP2D. \n Structural class GRAMPA non-GRAMPA \nN=5980  % N=1731   % N=2511   % N=1185   % \n1 Helices and coils 3892     65.1 1119     64.6 749       29.8 425       35.9 \n2 Mostly helices 88           1.5       23           1.4 5             0.2 5             0.2 \n3 Helices and strands 0             0.0 0             0.0 0             0.0 0             0.0 \n4 Mostly strands (β-sheets) 1             0.0 1             0.0 0             0.0 0             0.0 \n5 Strands and coils 754       12.6 261       15.1 767       30.6 387       32.7 \n6 Mostly coiled structures 1063     17.8 244       14.1 862       34.3 302       25.6 \n7 Mixed structures 182         3.0 83           4.8 128         5.1 66           5.6 \n \nFigure 1. Predicted structural landscapes of (A) GRAMPA and (B) non-GRAMPA. \n \n \n3.2. Performance of AMP predictors across structural classes. \nTo investigate whether the structural distributions of GRAMPA and non-GRAMPA \nsubsets influence antimicrobial peptide (AMP) prediction models, we  evaluated 16 \nstate-of-the-art predictors published between 2016 and 2024 (Table S1). These models \nvary in  machine learning architectures and training set sizes : AMP scanner v2 37, \namPEPpy38, AMPlify variants39, CAMPr3 variants40, DBAASP41, IAMPE variants42, \n20\n40\n60\n80\n100\n20\n40\n60\n80\n100\n20 40 60 80 100\n% Helix (H)\n% Strand (E)\n% Coil (C)\nDensity\n05 1 0 1 5\nA. GRAMPA \n5\n20\n40\n60\n80\n100\n20\n40\n60\n80\n100\n20 40 60 80 100\n% Helix (H)\n% Strand (E)\n% Coil (C)\nDensity\n05 1 0 1 5\nB. Non−GRAMPA \n1\n2\n3\n4\n7\n6\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint \n\n8 \niAMPpred43, PepNet44, and Sense the Moment 45. Predictions were stratified into the four \ndominant structural classes, as shown in Table 1: helices and coils (1), strands and \ncoils (5), mostly coils (6), and mixed structures (7). For each model and structural class, \nwe calculated accuracy, precision , sensitivity, and specificity from the confusion \nmatrices shown in Figures 2A-D and Figures S1-S4.  \nFor the helical and coil dataset (1, Figures 2A and S1), most classifiers achieved \naccuracy and precision consistently above 0.80, indicating robust classification of both \nAMPs and non-AMPs across all models. Here, the Sense the Moment model is a relative \nunderperformer with an accuracy dropping to nearly 0.70. In addition, nearly all \nmodels also showed high sensitivity (~0.80 -0.95) but lower specificity (~0.60 -0.80), \nsuggesting a bias toward predicting the positive (AMP) class. As such , the model \nIAMPE (SVM)  exhibited the largest gap between sensitivity (~0.95) and specificity \n(~0.42). The models AMPlify and AMPlify (imbalanced)  offer the best balance with \nrelatively high accuracy, precision, and sensitivity without low specificity.  \nThis sensitivity-specificity imbalance persists across other structural classes (5, 7) and \nclassifiers, suggesting a systemic bias in the benchmark and reflecting the composition \nof the training data. Most predictors in the strands and coils class (5; Figures 2B and \nS2) underperformed, with low accuracy, precision (~0.40), and specificity (~0.25), but \nexhibited the highest sensitivity (>0.90). Interestingly, three classifiers – AMPlify, \nAMPlify (imbalanced),  and PepNet – displayed all classification metrics above 0.75, \ndemonstrating strong predictive performance for this structural class.  We observed \nsimilar trends in the mixed structures class (7, Figures 2D and S4), where predictors \nperformed moderately well, with accuracy and precision in the 0.60 -0.70 range, \nspecificity averaging 0.30, and sensitivity above 0.90. Finally, several classifiers for the \ncoiled class (6, Figures 2C and S3) achieved near -equivalent performance across all \nmetrics, suggesting either a more tractable structural class or a smaller, less \ndiscriminating test set. The IAMPE (SVM)  and IAMPE (XGB)  exhibited the largest \nsensitivity-specificity gaps, indicating a bias toward coiled AMPs. The same three \nclassifiers – AMPlify, AMPlify (imbalanced),  and PepNet – outperformed the other \nmodels, suggesting they capture features relevant to stranded and coiled peptide \ndatasets that others do not.  \nDespite these observations, confusion-matrix metrics may overestimate performance \non imbalanced datasets48. To ensure fairer comparisons across structural classes, we \nalso evaluated performance using  the Matthews Correlation Coefficient (MCC )49, \nwhich balances predictions across both positive and negative classes (see Materials \nand Methods). As shown in Figure 2E, MCC values varied across structural classes. \nThe AMPlify models are the most robust classifiers across all four structural classes, \nwith AMPlify (imbalanced)  peaking at 0.82 on coiled structures. The two models \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint \n\n9 \nmaintained strong MCC  values even on the hardest classes (5: 0.74/0.79 and 7: \n0.66/0.73), where other models collapsed. The model PepNet is arguably the third most \nreliable classifier, with moderate MCC values  of 0.70 ( strands), 0.65 (coils), 0.63 \n(helices), and 0.59 (mixed structures). In contrast, the model Sense the Moment  (StM) \nexhibited the (near -)lowest MCC values in three of the four structural classes: 0.19 \n(strands), 0.12 (coils), 0.34 (helices), and 0.26 (mixed structures). Its moderate accuracy \n(Figure S2) on helical peptides may be due to class bias. All IAMPE variants and the \niAMPpred model exhibited near -random MCC values on stranded and mixed \nstructures (5 and 7), confirming the high-sensitivity/low-specificity pattern.  \nOverall, these results confirm that peptide structural features influence model \npredictions, with β-sheet-containing classes (5 and 7)  consistently performing worse \nthan helical and coiled classes (1 and 6). This pattern suggests an algorithmic bias that \nhinders the classification of AMPs versus non -AMPs within strand-rich structural \nspace. Structural class 1 (helices and coils) achieved the highest performance across \nevaluation approaches, confirming reliable classification across models. In contrast, \nstrand-rich classes (5 and 7) exhibited high sensitivity but low specifi city, and their \nlow MCC values revealed limited overall discriminatory reliability. This MCC -based \nevaluation aligns with the confusion matrix analyses ( Figures 2A-D and S1-S4), and \nMCC provides a more comprehensive representation of model efficacy than accuracy \nalone, particularly in skewed datasets where sensitivity -specificity trade -offs can \nobscure predictive power48.  \nAmong the 16 classifiers evaluated, AMPlify (both models) and PepNet showed the \nstrongest and most consistent performance across all structural classes. Notably, the \nAMPlify (imbalanced) variant achieved the highest single MCC value across the entire \nbenchmark (0.82 on coiled structures) and maintained strong discriminative power on \nthe helical and stranded classes alike, suggesting that accounting for structural class \nimbalance during model training contributes to its robustness. With the exception of \nthese models, our combined analyses indicate that most classifiers struggle to reliably \ndiscriminate AMP from non -AMP sequences containing β-sheet motifs, a limitation \nthat accuracy-based metrics alone would have underrepresented. \n \nFigure 2. Model performance varies across structural classes: (A) class 1 (helices/coils), \n(B) class 5 (strands/coils), (C) class 6 (predominantly coiled), and (D) class 7 (mixed). \nEvaluation is based on confusion-matrix metrics (specificity, sensitivity, accuracy, \nprecision), with global performance captured by Matthews correlation coefficients (E). \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint \n\n10 \n \nNext, we assessed per-sequence model performance by calculating the proportion of \nsequences correctly classified as GRAMPA or non -GRAMPA across different \nstructural classes (Figure 3). The figure displays the predictions – true positives (TP), \nfalse negatives ( FN), false positives ( FP), and true negatives ( TN) – for the 16 AMP \nclassifiers tested against the four structural classes (1, 5, 6, and 7). Most models achieve \nhigh true-positive rates for GRAMPA peptides (dominant dark blue bars on the left) \nbut also produce significant false positives among non -GRAMPA peptides (light red \nbars on the right). This further confirms the systemic bias in traini ng toward known \nAMPs. The right side of each panel shows that a few models, such as AMPlify and \nPepNet, consistently achieve high true -negative rates (dark red bars) acro ss all four \nstructural classes. \n0.74 0.82 0.49 0.65 0.120.620.44 0.21 0.35 0.46 0.33 0.350.36 0.38 0.460.37\n0.69 0.74 0.51 0.63 0.340.600.58 0.54 0.56 0.60 0.54 0.490.56 0.48 0.580.51\n0.74 0.79 0.43 0.70 0.190.260.10 0.17 0.12 0.22 0.16 0.080.06 0.04 0.060.12\n0.66 0.73 0.21 0.59 0.260.370.27 0.22 0.27 0.45 0.36 0.140.12 0.13 0.230.11Mixed structures (7)\t\nCoiled structures (6)\t\n\t\nStrands and coils (5)\t\n\t\nHelices and coils (1)\nAMP scanner v2\namPEPpy AMPlify\nAMPlify ( imbalanced)\nCAMPr 3  ( ANN)CAMPr 3  ( D AC)CAMPr 3  ( RF)CAMPr 3  ( SVM)\nDBAASPIAMPE (kNN)IAMPE (RF)IAMPE (SVM)IAMPE (XG B)iAMPpred PepNet\nSense t he Moment\nClassiﬁers\nStructural Class\nMetric\nE Matthews Correlation coefficients (MCC)\nA\nC Coiled structures (6) D\nB\n0.91\n0.88\n0.92\n0.77\n0.92\n0.90\n0.94\n0.77\n0.90\n0.77\n0.76\n0.79\n0.89\n0.85\n0.90\n0.72\n0.85\n0.69\n0.69\n0.69\n0.89\n0.84\n0.89\n0.70\n0.88\n0.83\n0.88\n0.69\n0.88\n0.81\n0.86\n0.69\n0.88\n0.82\n0.87\n0.69\n0.90\n0.83\n0.87\n0.74\n0.88\n0.81\n0.85\n0.71\n0.87\n0.78\n0.82\n0.69\n0.88\n0.83\n0.89\n0.67\n0.82\n0.81\n0.95\n0.44\n0.86\n0.84\n0.93\n0.60\n0.86\n0.81\n0.88\n0.62\nAMP scanner v2\namPEPpyAMPlify\nAMPlify ( imbalanced)\nCAMPr 3  ( ANN)CAMPr 3  ( D AC)CAMPr 3  ( RF)CAMPr 3  ( SVM)\nDBAASPIAMPE (kNN)IAMPE (RF)IAMPE (SVM)IAMPE (XG B)iAMPpredPepNet\nSense the Moment\n0.80\n0.87\n0.91\n0.85\n0.81\n0.89\n0.96\n0.85\n0.81\n0.73\n0.43\n0.93\n0.76\n0.85\n0.90\n0.81\n0.60\n0.63\n0.26\n0.88\n0.47\n0.55\n0.91\n0.32\n0.42\n0.44\n0.95\n0.10\n0.45\n0.52\n0.86\n0.28\n0.43\n0.49\n0.86\n0.24\n0.45\n0.52\n0.92\n0.26\n0.45\n0.51\n0.85\n0.28\n0.42\n0.45\n0.93\n0.12\n0.41\n0.44\n0.94\n0.09\n0.41\n0.42\n0.97\n0.04\n0.41\n0.43\n0.95\n0.08\n0.43\n0.47\n0.92\n0.16\nHelices and coils (1)\nSpecificity\nSensitivity\nStrands and coils (5)\n0.82\n0.83\n0.89\n0.76\n0.83\n0.87\n0.95\n0.76\n0.68\n0.59\n0.49\n0.71\n0.78\n0.80\n0.88\n0.70\n0.81\n0.57\n0.30\n0.91\n0.67\n0.69\n0.88\n0.45\n0.62\n0.64\n0.93\n0.27\n0.62\n0.62\n0.83\n0.36\n0.64\n0.64\n0.83\n0.41\n0.69\n0.72\n0.90\n0.50\n0.67\n0.68\n0.86\n0.47\n0.59\n0.59\n0.88\n0.23\n0.59\n0.58\n0.87\n0.23\n0.58\n0.58\n0.95\n0.12\n0.60\n0.62\n0.94\n0.21\n0.58\n0.58\n0.86\n0.23\n0.84\n0.87\n0.87\n0.87\n0.85\n0.91\n0.95\n0.87\n0.86\n0.74\n0.51\n0.93\n0.80\n0.83\n0.82\n0.83\n0.54\n0.58\n0.32\n0.78\n0.74\n0.81\n0.87\n0.75\n0.60\n0.69\n0.89\n0.53\n0.54\n0.60\n0.64\n0.56\n0.63\n0.68\n0.68\n0.68\n0.68\n0.73\n0.75\n0.72\n0.62\n0.66\n0.66\n0.67\n0.59\n0.66\n0.78\n0.57\n0.59\n0.66\n0.80\n0.56\n0.55\n0.63\n0.95\n0.37\n0.60\n0.69\n0.93\n0.50\n0.60\n0.67\n0.81\n0.56\nMixed structures (7)\nAccuracy\nPrecision\nSpecificity\nSensitivity\nAccuracy\nPrecision\nSpecificity\nSensitivity\nAccuracy\nPrecision\nSpecificity\nSensitivity\nAccuracy\nPrecision\nAMP scanner v2\namPEPpyAMPlify\nAMPlify ( imbalanced)\nCAMPr 3  ( ANN)CAMPr 3  ( D AC)CAMPr 3  ( RF)CAMPr 3  ( SVM)\nDBAASPIAMPE (kNN)IAMPE (RF)IAMPE (SVM)IAMPE (XG B)iAMPpredPepNet\nSense the Moment\nAMP scanner v2\namPEPpyAMPlify\nAMPlify ( imbalanced)\nCAMPr 3  ( ANN)CAMPr 3  ( D AC)CAMPr 3  ( RF)CAMPr 3  ( SVM)\nDBAASPIAMPE (kNN)IAMPE (RF)IAMPE (SVM)IAMPE (XG B)iAMPpredPepNet\nSense the Moment AMP scanner v2\namPEPpyAMPlify\nAMPlify ( imbalanced)\nCAMPr 3  ( ANN)CAMPr 3  ( D AC)CAMPr 3  ( RF)CAMPr 3  ( SVM)\nDBAASPIAMPE (kNN)IAMPE (RF)IAMPE (SVM)IAMPE (XG B)iAMPpredPepNet\nSense the Moment\n0.00 0.25 0.50 0.75 1.00\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint \n\n11 \nFigure 3  Mapping the proportions of predictions across classifiers and structural \nclasses - (A) 1: Helices and coils, (B) 5: Strands and coils, (C) 6: Mostly coiled structures, \nand (D) 7: Mixed structures. \n \nIn Figure 3A, models perform confidently overall against the helical structural class \n(1), as expected, since α-helical peptides are likely to dominate AMP training \ndatabases, such as GRAMPA 32. Most classifiers output high TP rates (about 75-95%) \nand moderate TN rates for non-GRAMPA peptides. IAMPE (RF) has the lowest TN \nrate; the model may struggle with helical non-AMPs. Sense the Moment (StM) performs \nwell on helices, consistent with its design to distinguish AMP sequences from \nrandomized variants that maintain the hydrophobic moment but lack sequence order \nand secondary structure. The authors of StM recommend using their tool as a \ncomplementary model, primarily for sequences prone to form helical motifs 45. The \nmodel performance declines in other structural classes.  \nFigure 3C  displays greater variation across models, with performance differences \nfalling between those of the other structure classes. Several classifiers show low TP \nrates (<60%) and high false-negative rates for GRAMPA peptides, especially DBAASP \nand StM, indicating that coiled structures are often misclassified as non-AMPs – likely \ndue to underrepresentation or structural ambiguity in the training sets.  Among the \nAMPscanner v2\namPEPpy\nAMPlify\nAMPlify (imbalanced)\nCAMPr3 (ANN)\nCAMPr3 (DAC)\nCAMPr3 (RF)\nCAMPr3 (SVM)\nDBAASP\nIAMPE (kNN)\nIAMPE (SVM)\nIAMPE (RF)\nIAMPE (XGB)\niAMPpred\nPepNet\nSense the Moment\n100 75 50 25 0 25 50 75 100\nProportion of predictions (%)\n100 75 50 25 0 25 50 75 100\nProportion of predictions (%)\n100 75 50 25 0 25 50 75 100\nProportion of predictions (%)\n100 75 50 25 0 25 50 75 100\nProportion of predictions (%)\nA Helices and coils (1) B Strands and coils (5)\nC Coiled structures (6) D Mixed structures (7)\nTrue Positive (GRAMPA) False Negative (GRAMPA) False Positive (Non −GRAMPA) True Negative (Non −GRAMPA)\nAMPscanner v2\namPEPpy\nAMPlify\nAMPlify (imbalanced)\nCAMPr3 (ANN)\nCAMPr3 (DAC)\nCAMPr3 (RF)\nCAMPr3 (SVM)\nDBAASP\nIAMPE (kNN)\nIAMPE (SVM)\nIAMPE (RF)\nIAMPE (XGB)\niAMPpred\nPepNet\nSense the Moment\nAMPscanner v2\namPEPpy\nAMPlify\nAMPlify (imbalanced)\nCAMPr3 (ANN)\nCAMPr3 (DAC)\nCAMPr3 (RF)\nCAMPr3 (SVM)\nDBAASP\nIAMPE (kNN)\nIAMPE (SVM)\nIAMPE (RF)\nIAMPE (XGB)\niAMPpred\nPepNet\nSense the Moment\nAMPscanner v2\namPEPpy\nAMPlify\nAMPlify (imbalanced)\nCAMPr3 (ANN)\nCAMPr3 (DAC)\nCAMPr3 (RF)\nCAMPr3 (SVM)\nDBAASP\nIAMPE (kNN)\nIAMPE (SVM)\nIAMPE (RF)\nIAMPE (XGB)\niAMPpred\nPepNet\nSense the Moment\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint \n\n12 \nIAMPE variants, differences in performance on coiled structures appear to be driven \nmore by model architecture than by training data, since all variants use the same \ntraining set. IAMPE (RF) achieves a high TP rate for GRAMPA coils but has the lowest \ntrue-negative rate among IAMPE variants , suggesting that the random forest \narchitecture tends to overfit AMPs for this structural class. The other IAMPE models \n(kNN, SVM, XGB) attain higher TN rates despite lower TP rates, indicating that these \narchitectures are more conservative in predicting coiled structures. A similar pattern \nis observed  among CAMPr3 variants, with the random forest model achieving the \nhighest TP and TN rates in the series. AMPlify, AMPlify (imbalanced), amPEPpy, and \nPepNet performed best on this structural class , achieving the highest combined true-\npositive and true-negative rates. \nIn Figures 3B and 3D, overall performance drops notably, reflecting that stranded \nAMPs (class 5) and mixed structures (class 7) are underrepresented in the training \ndata. Many models – including AMP scanner v2 , CAMPr3 variants, and IAMPE \nvariants – misclassify non-GRAMPA strands and mixed structures as AMPs, resulting \nin higher false -positive rates despite maintaining high TP rates for GRAMPA \nsequences. This simultaneous rise in both true-positive and false-positive rates \nsuggests that these models recognize broad physicochemical features common to \nhelical AMPs rather than structure-specific signatures. Conversely, DBAASP and StM \nshow the opposite pattern  – with higher false -negative and true -negative rates  –\nconsistent with a training bias toward helical AMPs and, in the case of StM, its original \nhelix-centric design. Amon g all tested models, AMPlify, AMPlify (imbalanced) , and \nPepNet perform best across these two structural classes, exhibiting relatively higher \nTP and TN rates, which aligns with their strong performance shown in Figure 3C. \nTo better understand model confidence, we examined class probability distributions \nfor nine out of the 16 classifiers that provided probability outputs (Figure S5). Overall, \nGRAMPA peptides (shown in blue) had median probabilities above 0.5 across all \nstructural subsets, consistent with their correct labe lling as AMPs. In contrast, non -\nGRAMPA sequences (depicted in red) exhibited wider distributions and greater \nuncertainty. Structural classes lacking stranded motifs (classes 1 and 6) typically fell \nbelow the 0.5 threshold, while strand-rich subsets (classes 5 and 7) often exceeded it—\nconsistent with the higher false positive rates seen in Figure 3. Model confidence also \nmirrored structural bias: probability distributions are tighter and higher for helical \nGRAMPA peptides, gradually broadening across coiled, stranded, and mixed \nstructural classes, reinforcing the training bias shown in Figure 3. Among the  nine \nmodels, AMPlify and AMPlify (imbalanced) produce the clearest separation between \nGRAMPA and non -GRAMPA probability distributions across all structural classes, \nwith tight distributions approaching 1.00 for GRAMPA sequences and 0.00 for non -\nGRAMPA sequences.  PepNet displayed similarly well -separated distributions for \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint \n\n13 \nmost structural classes , though notably more dispersed probabilities for  non-\nGRAMPA helices and mixed structures. In contrast, non -GRAMPA sequences with \nstranded or mixed structures were assigned intermediate class probabilities (0.25-0.75) \nacross most remaining models - including amPEPpy, AMP scanner v2 , CAMPr3 \nvariants, and iAMPpred - indicating that the ir decision boundaries are poorly \ncalibrated for non-helical structures.  \nIn summary, structural class strongly influences classifier performance, with helical \npeptides being the most consistently detected and β -sheet-rich subsets (5 and 7) \npresenting the greatest challenge. Most models exhibit a high false-positive rate across \nstructural classes, reflecting limited specificity for non -helical AMPs and suggesting \nthat current classifiers generalize poorly beyond the helical sequences that dominate \ntheir training data. Few models perform consistently across all four structural classes; \nAMPlify, AMPlify (imbalanced), and PepNet demonstrated the most balanced profiles \nin Figure 3, and the probability distributions in Figure S5 further corroborate their \nsuperior discrimination across all structural classes. \n \n3.3. Understanding structural bias in predictions. \nTo elucidate the origi ns of observed prediction biases , particularly the \nmisclassification of non-GRAMPA sequences containing β-sheet motifs as AMPs, we \ninvestigated two potential contributing factors: structural imbalance and data leakage \nin training sets.  \nWe first assessed whether structural diversity was adequately represented across both \nsequence classes. W e encoded the structures of both external GRAMPA and non -\nGRAMPA validation sets using ProstT546, a sequence-structure encoder that generates \n3Di tokens as computationally efficient proxies for three-dimensional representations. \nThe approach circumvented the need for explicit structure predictors such as PEP2D \nand AlphaFold2. The resulting structural embeddings were projected onto a two -\ndimensional UMAP 47 manifold fine -tuned to maximize the separation of the four \nstructural classes under consideration (folds 1, 5, 6, 7)  – see Figure S6. As shown in \nFigure 4 , the two datasets are distributed across  the structural space, forming  two \nclusters: a broad, continuous lower manifold (bottom center -right) and a compact \nupper cluster (top left). The first cluster contains  the four predicted folds, \npredominantly dominated by structural classes 1 and 6, whereas the second cluster  \nincludes only structural classes 5 and 7. Comparing the datasets fold by fold, \nGRAMPA helices assigned to fold 1 (light green, panel A) are abundant, whereas their \nnon-GRAMPA counterparts (yellow, panel B) remain underrepresented. Conversely, \nnon-GRAMPA sequences are enriched in strands and coils (5 and 6, red and orange) \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint \n\n14 \nrelative to their GRAMPA counterparts . Th ese asymmetries mirror the structural \nimbalance documented in Table 1 , namely , the overrepresentation of helical folds \namong AMPs and of strands and coils among non-AMPs – a bias that likely underlies \nthe misclassification of stranded and coiled non-GRAMPA sequences as AMPs.  \n \nFigure 4. Structural diversity of GRAMPA and non -GRAMPA external validation \nsets projected onto a UMAP structural space. UMAP projections of ProstT5-derived \n3Di structural embeddings for the GRAMPA (A) and non -GRAMPA (B) validation \nsets, colored by predicted structural fold (1, 5, 6, 7). The manifold was fine -tuned to \nmaximize the separation of the four structural classes. \n \nWe then compared the structural distributions of the two external validation sets with \nthose of the available training sets for four classifiers : GRAMPA against the positive \ntraining sets and non-GRAMPA against the negative training sets. The classifiers were \nAMP scanner v2  (panels A and B) , amPEPpy (panels C and D) , AMPlify, and PepNet \nsharing identical training sets (panels E and F). The results are illustrated in Figure 5.  \nAcross all four classifiers, a large fraction of  GRAMPA sequences overlap s with the \npositive training sets (dark blue, “shared”), indicating varying degrees of data leakage \n(defined here as sequence identity) into them. The shared sequences are not confined \nto a specific structural region but are distributed broadly across the UMAP manifold, \nsuggesting the leakage is uniform across structural classes. The proportion of shared \nsequences is notable : 626/1707 for AMP scanner v2 , 594/1707 for amPEPpy, and \n695/1707 for AMPlify/PepNet, representing roughly one third of GRAMPA sequences  \nare shared with the respective positive training sets.  Of note,  AMPlify (imbalanced)  \nshares the same positive training set as AMPlify and PepNet and therefore exhibits an \nidentical positive leakage (695/1707). \nA GRAMPA (N=1707) B Non-GRAMPA (N=1180)\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint \n\n15 \nOn the negative side, AMP scanner v2 (panel B) and amPEPpy (panel D) are effectively \nleakage-free; both show zero and 11 shared sequences, respectively, between the non-\nGRAMPA validation set and the negative training sets. AMPlify/PepNet (panel F) \nincludes 576 shared sequences out of 1180, representing nearly half of the non -\nGRAMPA validation set that overlaps with the negative training data.  AMPlify \n(imbalanced) retains this negative leakage and expands to over 102,000 peptide \nsequences (not depicted). Consequently, the fraction of non-GRAMPA sequences is at \nleast as high as in panel F, increasing the likelihood that additional non -GRAMPA \nsequences are covered.  \nThe models AMP scanner v2  and amPEPpy exhibit negligible negative leakage but \nperform poorly on non-GRAMPA strands and mixed structures (5 and 7 ; Figures 3B \nand 3D). The negative training sets (light red) are heavily concentrated in the lower \nbroad cluster, dominated by helices and coils (1 and 6). The non-GRAMPA validation \nsequences (dark red, “unique”), however, are distributed across both the lower \nmanifold and the upper cluster, which is occupied by folds 5 and 7 (strands and mixed \nstructures). AMP scanner v2  (panel B) contained no negative stranded and mixed \nexamples, whereas amPEPpy (panel D) spans a broader, more continuous structural \nspace, partially covering the upper cluster with fewer helical folds (1) - yet this partial \ncoverage remains insufficient to prevent misclassification of stranded and mixed \nstructures. This s tructural underrepresentation  in the negative training sets causes \nboth models to misclassify out-of-distribution non-GRAMPA sequences from folds 5 \nand 7 as AMPs by default. \nThe broader negative set  coverage visible in panel F  (Figure 5 ) offers a structural \nexplanation for the apparent superior performance of AMPlify/PepNet across \nstructural classes, including greater representation of stranded and mixed structures \nin the upper cluster. However, t he data leakage observed in panels E (~41% of \nGRAMPA) and F (~49% of non -GRAMPA) indicates that both validation sets are \ncontaminated with sequences from the AMPlify, AMPlify (imbalanced), PepNet training \ndatasets. All three models may simply be recognizing examples seen during training \n(memorization) rather than learning  structurally generalizable  discriminating \nfeatures. In the case of AMPlify (imbalanced), the disproportionate size of the negative \ntraining set increases the likelihood that non -GRAMPA validation sequences are \nrepresented in training. Their superior performance across all four structural classes \nmay therefore be inflated by this leakage, making it difficult to determine how much \nis attributable to learning versus memorization of identical training examples. \n  \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint \n\n16 \nFigure 5 Structural overlap between external validation sets and classifier training \nsets in UMAP space. UMAP projections of ProstT5 -derived 3Di structural \nembeddings compare the structural distributions of GRAMPA (Val1, blue shades) and \nnon-GRAMPA (Val2, red shades ) with the positive and negative training sets (light \nshades) for AMP scanner v2  (A, B), amPEPpy (C, D), and AMPlify/PepNet (E, F) , \nrespectively. Darker points denote the validation sequences shared with the \ncorresponding training set. Sample sizes are indicated in the legends. \n \nA AMP scanner v2 (+) B AMP scanner v2 (-)\nC amPEPpy (+) D amPEPpy (-)\nE F AMPlify (-) / PepNet (-)AMPlify (+) / PepNet (+)\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint \n\n17 \nTo disentangle data leakage from genuine pan -fold learning, we re -evaluated \nclassification performance on non -overlapping subsets of GRAMPA  (N=1,012) and \nnon-GRAMPA sequences (N=604). Sustained superior discrimination on these subsets \nwould support genuine model quality, whereas a performance drop would implicate \ndata leakage as a major confounding factor. We report prediction results – true \npositives (TP), false negatives (FN), false positives (FP), and true negatives (TN) – for \nthe AMPlify and PepNet models, both before ( Tables S2 and S4) and after removing \nduplicate sequences (Tables S3 and S5). These results are summarized in Figure 6.  \nThe removal of shared sequences between the training and validation sets resulted in \nsignificant losses for both GRAMPA and non -GRAMPA sequences across folds. \nGRAMPA lost 37.4% of helices (fold 1), half (52.5%) of stranded AMPs (fold 5), 41.8% \ncoils (fold 6) , and 44.6% mixed structures  (fold 7) . Similarly, non -GRAMPA \ndiminished 44% of helices, 46% of stranded non-AMPs, nearly two-thirds (62.6%) of \ncoils, and one-third of mixed structures.  \nAfter removing leakage, AMPlify consistently outperforms PepNet across all folds in \nboth sensitivity and specificity  (Figure 6 ). For AMPlify, FN counts are exactly \npreserved across all folds, and FP counts are nearly identical (fold 1: 96 to 93) , \nindicating that  the performance drops  are almost  entirely denominator -driven; \nconsistent with the near-complete memorization of the leaked sequences. For PepNet, \nboth FN and FP  counts decrease slightly  after de -leaking, indicating that a small \nnumber of leaked sequences were misclassified and that memorization was therefore \nless complete than in AMPlify. Sensitivity (DTP%) drops 5-11% for both models across \nfolds, while the false positive rate (DFP%) rises by 13-22% for AMPlify (Figure 6A) and \n13-28% for PepNet (Figure 6B). Specificity is more affected than sensitivity across all \nfolds and both models, suggesting that the non -AMP decision boundary has learned \nless from genuine sequence features and is more vulnerable to leakage removal. \nExamining performance by structural class reveals consistent patterns. Fold 1 (helices) \nretains the highest sensitivity for both models ( AMPlify 86.7%, PepNet 84.8%) and \nshows the smallest sensitivity drop ( DTP% = –5.0% for both), suggesting that helical \nAMPs carry more discriminative sequence features beyond the training set. Fold 5 \n(stranded AMPs – upper panels) exhibits greater sensitivity loss relative to baseline \n(DTP% = –10.1% and –10.6%), consistent with the largest proportion of leaked \nsequences (52.5%). Fold 7 (mixed structures) shows more modest declines in either \nsensitivity or specificity across both models, despite its small sample size. The largest \ndivergence between models occurs in fold 6 (coiled structures), where both models \nfall below 75% sensitivity. With PepNet, the proportion of true positives decreases \nfrom 81.6% to 70.4% after leakage removal ( Tables S4  and S5). AMPlify retains \napproximately 8% higher sensitivity and 9% lower FP rate than PepNet. Overall, these \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint \n\n18 \nresults indicate that both models had inflated performance metrics due to data \nleakage, with coiled AMPs representing a structurally distinct class that neither model \nhas learned to reliably distinguish from non-AMPs on novel sequences. \n \nFigure 6.  Classification proportions before and after data leakage removal for \nAMPlify (A) and PepNet (B). The positive set (GRAMPA/AMP) is shown in the upper \npanels as stacked proportions of TP (dark blue) and FN (light blue). The negative set \n(non-GRAMPA/non-AMP) is shown in the lower panels as stacked proportions of FP \n(light red) and TN (dark red). Solid ba rs represent performance before leakage \nremoval; hatched bars represent performance after removing shared sequences \nbetween training and validation sets. Delta values (DTP%, DFP%) indicate the change \nin proportion after leakage removal. Folds correspond to structural classes: 1 (helices), \n5 (strands), 6 (coils), and 7 (mixed structures). \n  \nA AMPlify B PepNet\n100\n60\n80\n20\n40\n0\n100\n60\n80\n20\n40\n0\n100\n60\n80\n20\n40\n0\n100\n60\n80\n20\n40\n0\nProportions (%)\nProportions (%)\nProportions (%)\nProportions (%)\nTrue Positive\n(GRAMPA)\nFalse Negative\n(GRAMPA)\nFalse Positive\n(Non-GRAMPA)\nTrue Negative\n(Non-GRAMPA)\nBefore removal After removal\nfold 1 fold 5 fold 6 fold 7\n-5.0 -10.6 -11.1 -9.7-5.0ΔTP (%)\nΔFP (%)\n-10.1 -9.1 -8.7\n+17.0 +13.5 +27.7 +12.9+16.5 +13.0 +22.1 +12.1\nfold 1 fold 5 fold 6 fold 7 fold 1 fold 5 fold 6 fold 7\nfold 1 fold 5 fold 6 fold 7\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint \n\n19 \n4. DISCUSSION \nOur study demonstrates that current antimicrobial peptide (AMP) predictive models \nexhibit structural bias  arising from the uneven representation of peptide structural \nclasses in training datasets. Most models achieve high predictive accuracy for \nsequences adopting helical folds and coils , but perform poorly on β-sheet motifs , \nthereby limiting their applicability across the broader AMP structural space . These \nobservations are consistent  with independent work by Dean and co -workers, who \nreported that computationally predicted and generated AMPs against E. coli, S. aureus, \nand P. aeruginosa predominantly folded into α-helices50. This mirrors the inherent bias \ntowards helical motifs in existing AMP datasets used for activity prediction and \ngeneration. Comparable biases have been reported across other protein classes51 and \ncomputational tasks, resulting in  skewed predictions in  functional annotation 52,53, \nprotein-protein interaction networks 27, protein stability54,55, protein structure32,56, and \nin sequence or structure generation57,58.  \nThis work contributes to a growing body of studies showing that biases in training \ndata inflate performance metrics in protein prediction and generation . Multiple \nstrategies have been proposed  to mitigate systematic bias in ML models, including  \nclass resampling, clustering-based methods , synthetic data augmentation via \ngenerative algorithms , feature augmentation, and optimized classification \nalgorithms55,59. In our recent work, subset selection  enabled the development of \nstructure-specific models with superior predictive performance, whereas targeted \ndata reduction led to information loss in structure-agnostic models60, underscoring the \nneed for structural diversity and rigorous curation in the next generation of  ML \nmodels56. \nThe structural biases documented here also carry underappreciated consequences for \nprospective experimental campaigns. Poor model recall of β-sheet AMPs leads to \nfewer stranded candidates being synthesized and tested, fewer representatives in \npublic databases, and continued training imbalance in future models – a self -\nreinforcing loop analogous to chemotype bias in drug discovery 61–63. Breaking this \ncycle requires treating structural diversity as an explicit experimental design criterion, \nprioritizing β-sheet and mixed -fold candidates for synthesis over relying solely on \npredicted activity scores. \nThe observed differences in classification performance across structural classes are \nattributable to training imbalance and variation in learned (physicochemical) features \nacross folds. α-Helical AMPs possess well-characterized cationic charge distribution, \nhydrophobicity, and amphipathicity that are captured by the hydrophobic moment, a \nvector sum of side -chain hydrophobicities projected along the helical axis  and the \nbasis of the StM model45,64. These properties are amenable to extraction by both \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint \n\n20 \ntraditional machine learning and neural network architectures. By comparison, \nstranded and mixed AMPs include structural components (e.g., hydrogen bonding, \nstrand orientation, disulfide bridges) that are non-local and poorly captured by \nsequence-only encodings. The superior sensitivity retained for helical AMPs after \nleakage removal (fold 1: AMPlify 86.7%, PepNet 84.8%) reflects the stronger \ndiscrimination in sequence space for this class and points to an intrinsic limitation of \nsequence-only models for stranded and coiled structures, motivating the development \nof structure-aware feature representations for pan-fold discrimination of AMPs. \nBeyond the structural imbalance, the de -leaking analysis further reveals that the two \ntop-performing models exhibited inflated performance metrics driven by  near-\ncomplete memorization of the leaked sequences rather than generalization. Some \nparallels can be drawn with large generative models, where training on insufficiently \ndiverse data  may lead to  memorization that masks poor  generalization. The \naccompanying sensitivity loss (5 -11%) and increase in false positive rate (13 -28%) \nhighlight that neither model has learned to reliably distinguish non-AMPs from novel \nsequences in the coiled structural space. A principled mitigation strategy is to develop \nselective classifiers that abstain from making predictions when an input falls outside \nthe structural training distribution 65, for which the ProstT5 -UMAP projections \nreported here provide a natural basis. \nFinally, our benchmarking analysis compared models of increasing algorithm ic \ncomplexity: early models ( pre-2022) primarily relied on SVM and Random Forest \nclassifiers, wh ile recent models incorporate CNNs or transformers. Greater \ncomplexity has improved predictive performance but has not resolved the underlying \nsampling bias. Joint sequence-structure embeddings (e.g., Prost T546, Saprot66) and co-\ngeneration approaches that pair sequence and structure data are promising avenues, \nthough they currently  risk biasing exploration to ward well-represented folds . As \nProstT5 was trained on AlphaFold2 predictions that were filtered for high structural \nconfidence, enriched in well -structured helical proteins, and excluded disordered \nsequences, the derived embeddings should be treated with caution. With sufficiently \ndiverse and balanced training data , these approaches hold the potential to expand \nexploration across broader protein structural space67.  \n \n \n \n \n \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint \n\n21 \n5. CONCLUSION \nOur study highlights structural bias in predictive models for identifying novel \nantimicrobial peptides. The combination of uneven structural representation in \ntraining data, data leakage, and the intrinsic learnability of helical versus non -helical \nfolds limits the generalization and reliability of current AMP classifiers across the full \nstructural space. While current efforts to integrate sequence, structure, and function \ninto machine learning -guided frameworks offer more controllable peptide and \nprotein d esign with improved success rates, they may also constrain novelty by \nreinforcing exploration of familiar structural territory. Addressing structural bias – \nthrough curated, structurally diverse training data, selective prediction strategies, and \njoint seq uence-structure modeling – will be essential to unlocking the ‘dark’ \npeptidome and proteome, creating opportunities for novel discoveries that could \ndrive the next generation of therapeutics, biopesticides, and biomaterials. \n  \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint \n\n22 \nData and code availability \nDatasets include p eptide sequences, ProsT5 embeddings, activity, and fold \npredictions. R and Python scripts to reproduce ternary plots and UMAP -ProsT5 \nprojections are all available at https://github.com/plissonf/AMP-structural-bias-audit. \n \nDeclaration of competing interest \nV.D.A.-B. declares no competing interests. F.P. is the founder and scientific director of \nIngenie Bio , a company that provides  consulting and services in computational \npeptide design.  \n \nAcknowledgments \nThe authors thank the Mexican Secretaría de Ciencias, Humanidades, Tecnología e \nInnovación (SECIHTI, formerly CONAHCYT), for financial support through grant A1-\nS-32579 (2019-2023) supporting this study . V.D.A. -B. acknowledges a national \npostgraduate scholarship awarded by CONAHCYT during the funded period. \n \nAuthor contributions \nF.P. conceptualized the investigation. V.D.A.-B. and F.P. carried out the investigation, \nincluding methodology, data curation, benchmarking, and bioinformatics analysis. \nBoth authors wrote, edited, and reviewed the manuscript. \n \n  \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint \n\n23 \nREFERENCES \n1. Buolamwini, J. & Gebru, T. Gender Shades: Intersectional Accuracy Disparities in \nCommercial Gender Classification. in Proceedings of Machine Learning Research vol. 81 1–15 \n(2018). \n2. Bartlett, R., Morse, A., Stanton, R. & Wallace, N. Consumer-lending discrimination in the \nFinTech Era. J. Financ. Econ. 143, 30–56 (2022). \n3. Bianchi, F. et al.  Easily Accessible Text -to-Image Generation Amplifies Demographic \nStereotypes at Large Scale. in 2023 ACM Conference on Fairness, Accountability, and \nTransparency 1493–1504 (ACM, Chicago IL USA, 2023). doi:10.1145/3593013.3594095. \n4. Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 \nlaying down harmonised rules on artificial intelligence and amending Regulations (EC) \nNo 300/2008, (EU) No 167/2013, (EU) No 168/2013, (EU) 2018/858, (EU) 2018/1139 and (EU) \n2019/2144 and Directives 2014/90/EU, (EU) 2016/797 and (EU) 2020/1828 (Artificial \nIntelligence Act). (2024). \n5. Artificial Intelligence Risk Management Framework (AI RMF 1.0). (2023). \n6. Yang, K. K., Wu, Z. & Arnold, F. H. Machine -learning-guided directed evolution for \nprotein engineering. Nat. Methods 16, 687–694 (2019). \n7. Alley, E. C., Khimulya, G., Biswas, S., AlQuraishi, M. & Church, G. M. Unified rational \nprotein engineering with sequence-based deep representation learning. Nat. Methods 16, \n1315–1322 (2019). \n8. Senior, A. W. et al.  Improved protein structure prediction using potentials from deep \nlearning. Nature 577, 706–710 (2020). \n9. Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, \n583–589 (2021). \n10. Bileschi, M. L. et al. Using deep learning to annotate the protein universe. Nat. Biotechnol. \n40, 932–937 (2022). \n11. Watson, J. L. et al. De novo design of protein structure and function with RFdiffusion. \nNature 620, 1089–1100 (2023). \n12. Abramson, J. et al.  Accurate structure prediction of biomolecular interactions with \nAlphaFold 3. Nature 630, 493–500 (2024). \n13. Ahern, W. et al.  Atom-level enzyme active site scaffolding using RFdiffusion2. Nat. \nMethods 23, 96–105 (2026). \n14. Albanese, K. I., Barbe, S., Tagami, S., Woolfson, D. N. & Schiex, T. Computational protein \ndesign. Nat. Rev. Methods Primer 5, 13 (2025). \n15. Eid, F.-E. et al. Systematic auditing is essential to debiasing machine learning in biology. \nCommun. Biol. 4, 183 (2021). \n16. Community Values, Guiding Principles, and Commitments for the Responsible \nDevelopment of AI for Protein Design. (2023). \n17. Schnoes, A. M., Ream, D. C., Thorman, A. W., Babbitt, P. C. & Friedberg, I. Biases in the \nExperimental Annotations of Protein Function and Their Effect on Our Understanding of \nProtein Function Space. PLoS Comput. Biol. 9, e1003063 (2013). \n18. Radivojac, P. et al. A large-scale evaluation of computational protein function prediction. \nNat. Methods 10, 221–227 (2013). \n19. Plisson, F., Ramírez -Sánchez, O. & Martínez -Hernández, C. Machine learning -guided \ndiscovery and design of non-hemolytic peptides. Sci. Rep. 10, 16581 (2020). \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint \n\n24 \n20. Rádai, Z., Kiss, J. & Nagy, N. A. Taxonomic bias in AMP prediction of invertebrate \npeptides. Sci. Rep. 11, 17924 (2021). \n21. Hiranuma, N. et al. Improved protein structure refinement guided by deep learning based \naccuracy estimation. Nat. Commun. 12, 1–11 (2021). \n22. Varadi, M. et al.  AlphaFold Protein Structure Database: massively expanding the \nstructural coverage of protein -sequence space with high -accuracy models. Nucleic Acids \nRes. 50, D439–D444 (2022). \n23. Omidi, A., Møller, M. H., Malhis, N., Bui, J. M. & Gsponer, J. AlphaFold -Multimer \naccurately captures interactions and dynamics of intrinsically disordered protein regions. \nProc. Natl. Acad. Sci. 121, e2406407121 (2024). \n24. Chakravarty, D., Lee, M. & Porter, L. L. Proteins with alternative folds reveal blind spots \nin AlphaFold-based protein structure prediction. Curr. Opin. Struct. Biol. 90, 102973 (2025). \n25. Graber, D. et al. Resolving data bias improves generalization in binding affinity prediction. \nNat. Mach. Intell. 7, 1713–1725 (2025). \n26. Tsishyn, M., Pucci, F. & Rooman, M. Quantification of biases in predictions of protein –\nprotein binding affinity changes upon mutations. Brief. Bioinform. 25, bbad491 (2023). \n27. Lannelongue, L. & Inouye, M. Pitfalls of machine learning models for protein –protein \ninteraction networks. Bioinformatics 40, btae012 (2024). \n28. Yılmaz, S., Yorgancioglu, K. & Koyutürk, M. Bias -aware training and evaluation of link \nprediction algorithms in network biology. Proc. Natl. Acad. Sci. 122, e2416646122 (2025). \n29. Biswas, S., Khimulya, G., Alley, E. C., Esvelt, K. M. & Church, G. M. Low -N protein \nengineering with data-efficient deep learning. Nat. Methods 18, 389–396 (2021). \n30. Batra, R. et al. Machine learning overcomes human bias in the discovery of self-assembling \npeptides. Nat. Chem. 14, 1427–1435 (2022). \n31. Wankowicz, S. A. Modeling Bias Toward Binding Sites in PDB Structural Models. Preprint \nat https://doi.org/10.1101/2024.12.14.628518 (2024). \n32. Aldas-Bulos, V. D. & Plisson, F. Benchmarking protein structure predictors to assist \nmachine learning-guided peptide discovery. Digit. Discov. 2, 981–993 (2023). \n33. Witten, J. & Witten, Z. Deep learning regression model for antimicrobial peptide design. \n692681 Preprint at https://doi.org/10.1101/692681 (2019). \n34. The UniProt Consortium et al.  UniProt: the Universal Protein Knowledgebase in 2025. \nNucleic Acids Res. 53, D609–D617 (2025). \n35. Singh, H., Singh, S. & Singh Raghava, G. P. Peptide Secondary Structure Prediction using \nEvolutionary Information. Preprint at https://doi.org/10.1101/558791 (2019). \n36. Huang, Y., Niu, B., Gao, Y., Fu, L. & Li, W. CD-HIT Suite: a web server for clustering and \ncomparing biological sequences. Bioinformatics 26, 680–682 (2010). \n37. Veltri, D., Kamath, U. & Shehu, A. Deep learning improves antimicrobial peptide \nrecognition. Bioinformatics 34, 2740–2747 (2018). \n38. Lawrence, T. J. et al.  amPEPpy 1.0: a portable and accurate antimicrobial peptide \nprediction tool. Bioinformatics 37, 2058–2060 (2021). \n39. Li, C. et al. AMPlify: attentive deep learning model for discovery of novel antimicrobial \npeptides effective against WHO priority pathogens. BMC Genomics 23, 77 (2022). \n40. Waghu, F. H., Barai, R. S., Gurung, P. & Idicula -Thomas, S. CAMPR3  : a database on \nsequences, structures and signatures of antimicrobial peptides. Nucleic Acids Res.  44, \nD1094–D1097 (2016). \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint \n\n25 \n41. Pirtskhalava, M. et al.  DBAASP v3: database of antimicrobial/cytotoxic activity and \nstructure of peptides as a resource for development of new therapeutics. Nucleic Acids Res. \n49, D288–D297 (2021). \n42. Kavousi, K. et al.  IAMPE: NMR -Assisted Computational Prediction of Antimicrobial \nPeptides. J. Chem. Inf. Model. 60, 4691–4701 (2020). \n43. Meher, P. K., Sahu, T. K., Saini, V. & Rao, A. R. Predicting antimicrobial peptides with \nimproved accuracy by incorporating the compositional, physico-chemical and structural \nfeatures into Chou’s general PseAAC. Sci. Rep. 7, 42362 (2017). \n44. Han, J., Kong, T. & Liu, J. PepNet: an interpretable neural network for anti-inflammatory \nand antimicrobial peptides prediction using a pre -trained protein language model. \nCommun. Biol. 7, 1198 (2024). \n45. Porto, W. F., Ferreira, K. C. V., Ribeiro, S. M. & Franco, O. L. Sense the moment: A highly \nsensitive antimicrobial activity predictor based on hydrophobic moment. Biochim. Biophys. \nActa BBA - Gen. Subj. 1866, 130070 (2022). \n46. Heinzinger, M. et al. Bilingual language model for protein sequence and structure. NAR \nGenomics Bioinforma. 6, lqae150 (2024). \n47. Healy, J. & McInnes, L. Uniform manifold approximation and projection. Nat. Rev. \nMethods Primer 4, 82 (2024). \n48. Chicco, D. Ten quick tips for machine learning in computational biology. BioData Min. 10, \n35 (2017). \n49. Chicco, D. & Jurman, G. The advantages of the Matthews correlation coefficient (MCC) \nover F1 score and accuracy in binary classification evaluation. BMC Genomics 21, 6 (2020). \n50. Dean, S. N., Alvarez, J. A. E., Zabetakis, D., Walper, S. A. & Malanoski, A. P. PepVAE: \nVariational Autoencoder Framework for Antimicrobial Peptide Generation and Activity \nPrediction. Front. Microbiol. 12, 725727 (2021). \n51. Ding, F. & Steinhardt, J. Protein language models are biased by unequal sequence \nsampling across the tree of life. Preprint at https://doi.org/10.1101/2024.03.07.584001 \n(2024). \n52. Liu, X. Deep Recurrent Neural Network for Protein Function Prediction from Sequence. \nhttps://doi.org/10.1101/103994 (2017) doi:10.1101/103994. \n53. Bonetta, R. & Valentino, G. Machine learning techniques for protein function prediction. \nProteins Struct. Funct. Bioinforma. 88, 397–413 (2020). \n54. Pucci, F., Bernaerts, K. V., Kwasigroch, J. M. & Rooman, M. Quantification of biases in \npredictions of protein stability changes upon mutations. Bioinformatics 34, 3659 –3665 \n(2018). \n55. Fang, J. The role of data imbalance bias in the prediction of protein stability change upon \nmutation. PLOS ONE 18, e0283727 (2023). \n56. Derry, A., Carpenter, K. A. & Altman, R. B. Training data composition affects performance \nof protein structure analysis algorithms. Pac. Symp. Biocomput. Pac. Symp. Biocomput.  27, \n10–21 (2022). \n57. Tang, X. et al. A survey of generative AI for de novo drug design: new frontiers in molecule \nand protein generation. Brief. Bioinform. 25, bbae338 (2024). \n58. Lu, T., Liu, M., Chen, Y., Kim, J. & Huang, P. -S. Assessing generative model coverage of \nprotein structures with SHAPES. Cell Syst. 16, 101347 (2025). \n59. Jiang, J. et al. A review of machine learning methods for imbalanced data challenges in \nchemistry. Chem. Sci. 16, 7637–7658 (2025). \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint \n\n26 \n60. Aguilera-Puga, M. D. C. & Plisson, F. Structure -aware machine learning strategies for \nantimicrobial peptide discovery. Sci. Rep. 14, 11995 (2024). \n61. Bickerton, G. R., Paolini, G. V., Besnard, J., Muresan, S. & Hopkins, A. L. Quantifying the \nchemical beauty of drugs. Nat. Chem. 4, 90–98 (2012). \n62. Saldívar-González, F. I., Aldas -Bulos, V. D., Medina -Franco, J. L. & Plisson, F. Natural \nproduct drug discovery in the artificial intelligence era. Chem. Sci. 13, 1526–1546 (2022). \n63. Van Den Broek, R. L., Patel, S., Van Westen, G. J. P., Jespers, W. & Sherman, W. In Search \nof Beautiful Molecules: A Perspective on Generative Modeling for Drug Design. J. Chem. \nInf. Model. 65, 9383–9397 (2025). \n64. Eisenberg, D., Weiss, R. M. & Terwilliger, T. C. The hydrophobic moment detects \nperiodicity in protein hydrophobicity. Proc. Natl. Acad. Sci. 81, 140–144 (1984). \n65. Rabanser, S. & Papernot, N. What Does It Take to Build a Performant Selective Classifier? \nPreprint at https://doi.org/10.48550/ARXIV.2510.20242 (2025). \n66. Su, J. et al. Democratizing protein language model training, sharing and collaboration. Nat. \nBiotechnol. https://doi.org/10.1038/s41587-025-02859-7 (2025) doi:10.1038/s41587 -025-\n02859-7. \n67. Wang, C., Alamdari, S., Domingo -Enrich, C., Amini, A. P. & Yang, K. K. Toward deep \nlearning sequence–structure co-generation for protein design. Curr. Opin. Struct. Biol. 91, \n103018 (2025). \n \n \n \n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted May 8, 2026. ; https://doi.org/10.64898/2026.05.06.721805doi: bioRxiv preprint","source_license":"CC-BY-4.0","license_restricted":false}