peleke-1: A Suite of Protein Language Models Fine-Tuned for Targeted Antibody Sequence Generation

preprint OA: closed
📄 Open PDF Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-16

Peleke-1 is a suite of protein language models fine-tuned on antibody-antigen complex data to generate targeted antibody Fv sequences for given antigen sequences at scale.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-16 · read from full text

The paper introduces peleke-1, a suite of protein language models fine-tuned on curated antibody–antigen complex data from SAbDab/PDB to generate targeted antibody Fv heavy and light chain sequences conditioned on an input antigen sequence. The authors collect 9,523 complete entries, use structural analysis (PandaProt) to identify epitope residues that mediate polar contacts, and encode these epitope residues in the model prompt while fine-tuning three LLMs derived from Phi-4 and Llama-3.1–8B-Instruct and Mistral-7B-Instruct-v0.2, with evaluation performed using benchmark antigen sequences such as EGFR. A major caveat stated by the authors is that their training/evaluation pipeline depends on the quality and completeness of available structural complex data and on defining epitope residues from contact analysis in curated PDB structures. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

The discovery of therapeutic antibodies is a traditionally arduous process. Today, the lab-based process of antibody discovery consists of several time-consuming steps that involve live animal immunization, B-cell harvesting, hybridoma creation, and then downstream engineering and evaluation. However, the use of artificial intelligence in drug design has previously been shown effective in the rapid generation of proteinspecific binders, small molecules, and even antibody therapeutics, thereby replacing some of the primary steps of the drug discovery process. Here we present peleke-1 , a suite of protein language models fine-tuned from state-of-the-art large language models using curated antibody-antigen complex data. These models generate targeted antibody Fv sequences for a given antigen sequence input at-scale. This suite of models provides a reliable, artificial intelligence-driven approach for in silico therapeutic antibody discovery along with an open-source framework for future antibody language model tuning.
Full text 36,103 characters · extracted from preprint-html · click to expand
peleke-1: A Suite of Protein Language Models Fine-Tuned for Targeted Antibody Sequence Generation | bioRxiv /* */ /* */ <!-- <!-- /*! * yepnope1.5.4 * (c) WTFPL, GPLv2 */ (function(a,b,c){function d(a){return"[object Function]"==o.call(a)}function e(a){return"string"==typeof a}function f(){}function g(a){return!a||"loaded"==a||"complete"==a||"uninitialized"==a}function h(){var a=p.shift();q=1,a?a.t?m(function(){("c"==a.t?B.injectCss:B.injectJs)(a.s,0,a.a,a.x,a.e,1)},0):(a(),h()):q=0}function i(a,c,d,e,f,i,j){function k(b){if(!o&&g(l.readyState)&&(u.r=o=1,!q&&h(),l.onload=l.onreadystatechange=null,b)){"img"!=a&&m(function(){t.removeChild(l)},50);for(var d in y[c])y[c].hasOwnProperty(d)&&y[c][d].onload()}}var j=j||B.errorTimeout,l=b.createElement(a),o=0,r=0,u={t:d,s:c,e:f,a:i,x:j};1===y[c]&&(r=1,y[c]=[]),"object"==a?l.data=c:(l.src=c,l.type=a),l.width=l.height="0",l.onerror=l.onload=l.onreadystatechange=function(){k.call(this,r)},p.splice(e,0,u),"img"!=a&&(r||2===y[c]?(t.insertBefore(l,s?null:n),m(k,j)):y[c].push(l))}function j(a,b,c,d,f){return q=0,b=b||"j",e(a)?i("c"==b?v:u,a,b,this.i++,c,d,f):(p.splice(this.i++,0,a),1==p.length&&h()),this}function k(){var a=B;return a.loader={load:j,i:0},a}var l=b.documentElement,m=a.setTimeout,n=b.getElementsByTagName("script")[0],o={}.toString,p=[],q=0,r="MozAppearance"in l.style,s=r&&!!b.createRange().compareNode,t=s?l:n.parentNode,l=a.opera&&"[object Opera]"==o.call(a.opera),l=!!b.attachEvent&&!l,u=r?"object":l?"script":"img",v=l?"script":u,w=Array.isArray||function(a){return"[object Array]"==o.call(a)},x=[],y={},z={timeout:function(a,b){return b.length&&(a.timeout=b[0]),a}},A,B;B=function(a){function b(a){var a=a.split("!"),b=x.length,c=a.pop(),d=a.length,c={url:c,origUrl:c,prefixes:a},e,f,g;for(f=0;f<d;f++)g=a[f].split("="),(e=z[g.shift()])&&(c=e(c,g));for(f=0;f<b;f++)c=x[f](c);return c}function g(a,e,f,g,h){var i=b(a),j=i.autoCallback;i.url.split(".").pop().split("?").shift(),i.bypass||(e&&(e=d(e)?e:e[a]||e[g]||e[a.split("/").pop().split("?")[0]]),i.instead?i.instead(a,e,f,g,h):(y[i.url]?i.noexec=!0:y[i.url]=1,f.load(i.url,i.forceCSS||!i.forceJS&&"css"==i.url.split(".").pop().split("?").shift()?"c":c,i.noexec,i.attrs,i.timeout),(d(e)||d(j))&&f.load(function(){k(),e&&e(i.origUrl,h,g),j&&j(i.origUrl,h,g),y[i.url]=2})))}function h(a,b){function c(a,c){if(a){if(e(a))c||(j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}),g(a,j,b,0,h);else if(Object(a)===a)for(n in m=function(){var b=0,c;for(c in a)a.hasOwnProperty(c)&&b++;return b}(),a)a.hasOwnProperty(n)&&(!c&&!--m&&(d(j)?j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}:j[n]=function(a){return function(){var b=[].slice.call(arguments);a&&a.apply(this,b),l()}}(k[n])),g(a[n],j,b,n,h))}else!c&&l()}var h=!!a.test,i=a.load||a.both,j=a.callback||f,k=j,l=a.complete||f,m,n;c(h?a.yep:a.nope,!!i),i&&c(i)}var i,j,l=this.yepnope.loader;if(e(a))g(a,0,l,0);else if(w(a))for(i=0;i (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0];var j=d.createElement(s);var dl=l!='dataLayer'?'&l='+l:'';j.src='//www.googletagmanager.com/gtm.js?id='+i+dl;j.type='text/javascript';j.async=true;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-M677548'); Skip to main content Home About Submit ALERTS / RSS Search for this keyword Advanced Search New Results peleke-1: A Suite of Protein Language Models Fine-Tuned for Targeted Antibody Sequence Generation View ORCID Profile Nicholas Santolla , View ORCID Profile Trey Pridgen , View ORCID Profile Prbhuv Nigam , View ORCID Profile Colby T. Ford doi: https://doi.org/10.1101/2025.10.16.682644 Nicholas Santolla 1 University of North Carolina at Charlotte, School of Data Science , Charlotte, NC, USA Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Nicholas Santolla Trey Pridgen 2 University of North Carolina at Charlotte, Center for Computational Intelligence to Predict Health and Environmental Risks (CIPHER) , Charlotte, NC, USA 3 University of North Carolina at Charlotte, Department of Bioinformatics and Genomics , Charlotte, NC, USA Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Trey Pridgen Prbhuv Nigam 4 North Carolina School of Science and Mathematics , Durham, NC, USA Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Prbhuv Nigam Colby T. Ford 1 University of North Carolina at Charlotte, School of Data Science , Charlotte, NC, USA 2 University of North Carolina at Charlotte, Center for Computational Intelligence to Predict Health and Environmental Risks (CIPHER) , Charlotte, NC, USA 3 University of North Carolina at Charlotte, Department of Bioinformatics and Genomics , Charlotte, NC, USA 5 Tuple LLC and Silico Biosciences , Charlotte, NC, USA Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Colby T. Ford For correspondence: colby{at}silico.bio Abstract Full Text Info/History Metrics Data/Code Preview PDF Abstract The discovery of therapeutic antibodies is a traditionally arduous process. Today, the lab-based process of antibody discovery consists of several time-consuming steps that involve live animal immunization, B-cell harvesting, hybridoma creation, and then downstream engineering and evaluation. However, the use of artificial intelligence in drug design has previously been shown effective in the rapid generation of proteinspecific binders, small molecules, and even antibody therapeutics, thereby replacing some of the primary steps of the drug discovery process. Here we present peleke-1 , a suite of protein language models fine-tuned from state-of-the-art large language models using curated antibody-antigen complex data. These models generate targeted antibody Fv sequences for a given antigen sequence input at-scale. This suite of models provides a reliable, artificial intelligence-driven approach for in silico therapeutic antibody discovery along with an open-source framework for future antibody language model tuning. Introduction Traditional antibody discovery is slow, expensive, and resource-intensive. This typically begins with antigen preparation, followed by animal immunization to elicit an immune response. Antibody-producing B-cells are then harvested and either fused with myeloma cells to create hybridomas or incorporated into phage display libraries to capture antibody diversity. Once sufficient antibody candidates are expressed, high-throughput screening is performed to evaluate binding affinity, specificity, and stability. Promising candidates are then subjected to engineering and optimization to improve pharmacokinetics, reduce immunogenicity, and enhance target binding. Finally, extensive in vitro and in vivo validation is required prior to clinical development for humans. Recent advances in computational biology, particularly the rise of protein language models (PLMs), offer a new paradigm. By learning rich sequence representations from massive protein corpora, PLMs can generalize structural and functional properties of proteins, enabling predictive and generative tasks once thought infeasible. Previously, PLMs such as ESM ( 1 , 2 ) and ProteinMPNN ( 3 ) have shown exceptional performance in generating realistic amino acid sequences from evolutionary-driven or structure-driven weights for general proteins. In the antibody domain, language models such as Ig-Bert ( 4 ), AbLang ( 5 ), AntiBERTy ( 6 ), and IgGM ( 7 ), have shown promise for accelerating candidate discovery, reducing dependence on animal immunization, and exploring immunoglobulin sequence space beyond what is accessible through natural immune repertoires or more general PLMs. However, a limitation of some of these models is that they are not trained on antibody-antigen complexes due to the lack of publicly available data. Here we introduce peleke-1 , a suite of protein language models fine-tuned for targeted antibody sequence generation. peleke-1 enables the rapid design of antibody candidates conditioned on an antigen input with desired epitopes, bridging the gap between large-scale pretraining and domain-specific fine-tuning for therapeutic discovery. This suite of models provides a reliable, artificial intelligence-driven approach for in silico therapeutic antibody discovery along with an opensource framework and training dataset for future antibody model tuning by the computational structural biology community. Methods The peleke-1 suite consists of multiple protein-language models (PLMs), fine-tuned from existing large language models (LLMs) that span varying architectures and parameter magnitudes. To perform the fine tuning, copious antibodyantigen sequence information was collected to form a curated training dataset. The overall workflow is shown in Figure 1 . Download figure Open in new tab Fig. 1. Model tuning and evaluation workflow. Data Curation Antibody-antigen complexes were collected from the Structural Antibody Database (SAbDab). The SAbDab data includes the PDB ID for the complex structure hosted on the Protein Data Bank website along with heavy and light chain and antigen chain identifiers. Also, this data includes other metadata about each antibody. This dataset included 18,971 antibodies across 16,263 target antigens. For each record, the PDB ID was used to retrieve the structure and sequence information from Protein Data Bank, filtering out any records where any of the antibody or antigen chains are missing. This resulted in a curated dataset of 9,523 complete entries that consisted of a heavy chain Fv sequence, light chain Fv sequence, and a target antigen sequence for each antibody. Example derived values are shown in Supplementary Table 2. Active Residue Determination Once the collection of antibody-antigen sequences was curated, the representative PDB structures were analyzed using PandaProt ( 8 ), which identifies epitope residues on the antigen that formed polar contacts with a CDR loop residue on the Fv structure. These interfacing residue numbers were collected and used to annotate the input sequences prompt, surrounding the interfacing residues with square brackets (“[]”). Prompt Example: Antigen: …FS[S][F][V]L[N]WY…\nAntibody: QVQL…|DIQM… Fine-Tuning Using the transformers package from Hugging Face, three LLMs were fine-tuned to the curated training dataset as listed in Table 1 . View this table: View inline View popup Download powerpoint Table 1. Percentage of generated antibodies that are predicted to be human per Humatch. First, the tokenizer for each model was updated to handle standard 20 amino acid characters and the Fv-chain delimiter “|”. Special tokens were added for denoting the epitope residues ( and ). The token embeddings of the model were then resized to accommodate these new tokens. Then, the training data are formatted and tokenized using the updated tokenizer. For example, KSF[E][D]AKCAA… becomes K,SF,, E,,,D,,AK,CAA,… after tokenization through the updated Phi-4 -based tokenizer. Stop tokens were also used at the end of the antigen sequence and antibody sequences to guide the LLM’s response. Phi-4 and Llama-3 . 1-8B-Instruct were tuned on a cloud-based NVIDIA H100 GPU in Microsoft Azure Machine Learning and Mistral-7B-Instruct-v0 . 2 was tuned on NVIDIA 5090 and 3090ti dual GPUs locally. Each LLM was fine-tuned to minimize loss over a few epochs, stopping when stable Fv-like sequences were consistently being generated for each chunk of steps. The models were then checkpointed for future use, evaluation, and inferencing. Evaluation Upon sufficient tuning of the PLMs, a test set of antigen sequences was selected as benchmarks for the antibody generation. The target benchmark antigens used were: Oncotargets - EGFR: Human epidermal growth factor receptor. (PDB: 8HGO) - PD-L1: Programmed death-ligand 1, an immune checkpoint inhibitor. (PDB: 4Z18) - PD-1: Programmed death protein 1, an immune check-point protein. (PDB: 5JXE) Infectious Disease Proteins - MBP: Maltose-binding protein from E. coli , a common protein expression tag. (PDB: 1NL5) - BHRF1: BCL-2 homolog from Epstein-Barr virus, an anti-apoptotic protein. (PDB: 2WH6) Others - IL-7Rα: Interleukin-7 receptor alpha chain, a cytokine receptor subunit. (PDB: 3DI3) - BBF-14: A synthetic 112-AA β-barrel protein. (PDB: 9HAC) This set encompasses 6 standard benchmark antigens to match AdapytvBio’s BenchBB service (BHRF1, EGFR, IL-7Ra, MBP, PD-L1, and BBF-14) ( 12 ) and 1 additional antigen (PD-1). For each antigen, 50 novel Fv antibody sequences were generated. Each paired set of Fv sequences were co-folded with their respective antigen using the esm3-medium-multimer-2024-09 model from Evolutionary Scale ( 2 ). Each Fv sequence was also quality-checked using multiple methods. This included a test to number the sequences through the Chothia numbering system using the ANARCI library ( 13 ) (to ensure Fv-like sequences), an analysis for evaluating humanness with Humatch ( 14 ), and stability predictions with FoldX ( 15 ). This closely follows the evaluation logic as published in Santolla and Ford, 2025. HADDOCK3’s haddock3-score method was used to predict the binding affinity of the predicted Fv structure to the target antigen ( 17 , 18 ). Then, these affinity metrics were collected to assess the generated Fv structure’s binding to the desired antigen. The metrics are as follows: Van der Waals intermolecular energy ( vdw ) in kcal/mol Electrostatic intermolecular energy ( elec ) in kcal/mol Desolvation energy ( desolv ) in kcal/mol Buried surface area ( bsa ) in Å 2 Total energy ( total ): 1.0 vdw + 1.0 elec in kcal/mol HADDOCK score: 1.0 vdw + 0.2 elec + 1.0 desolv + 0.1 air Evaluation outputs were then summarized by base LLM and antigen to assess overall model performance. Results Across the 7 benchmark antigens, 50 Fv sequences were generated for each of the 3 fine-tuned models, a total of 1,050 test complexes. The tuned models were inferenced locally on NVIDIA 5090 and 3090ti GPUs. Inferencing took <1 minute per Fv generation. Note that each model was inferenced 350 times (7 benchmark antigens × 50 runs) to generate pairs of heavy and light chain Fv sequences in each run. Antibody Realness and Stability Each model consistently produced Fv heavy and light chain sequences, 100% of which were successfully numbered using the Chothia numbering scheme (with only 4 eliciting duplicate residue numbering warnings). Furthermore, the models consistently generate human-like sequences, with 82% and 80% of Fv sequences predicted to be human for heavy and light chains, respectively. See Supplementary Table 1 for a full breakdown by model and antigen. Regarding protein stability, the generated antibody-antigen complexes averaged a total energy of −215.59 kcal/mol, polar solvation of −715.03 kcal/mol, and hydrophic solvation of 1,246.67 kcal/- mol across all models and antigens. These indicate predicted stable binding and solubility. While overall stability was similar across models, antigen-level metrics varied. EGFR had the most stable interactions and BBF-14 had the least. See Supplementary Figure 1 for the stability metric distributions by model and antigen. Sequence Diversity Due to the nature of language models, output sequence lengths varied. The generated heavy chain sequences were between 103 and 145 amino acids in length and light chains were between 98 and 116 amino acids. Conversely to our previous work in Santolla and Ford, 2025 that used a Microsoft EvoDiff diffusion model ( 19 ), where the conditionally-diffused sequences exhibited relatively low sequence diversity, the peleke-1 language models generated sequences with significant diversity. In the generated heavy chain sequences, various standard antibody patterns were common (e.g. sequences starting with EVQL or QVQL and ending with TVSS), though there was significant diversity in the other regions. In the generated light chains, lower diversity was experienced outside the CDR loop regions with standard DIQM or DIVL beginning patterns and ending with TKVEIK. See Supplementary Figure 2 for full WebLogos. Binding Performance As shown in Table 2 , the models were better overall at generating antibodies for some target antigens compared to others. For example, antibodies against targets MBP, PD-1, and PD-L1 had much better (lower) affinity scores and van der Waals energies as compared to those generated for BBF-14, BHRF1, and EGFR. View this table: View inline View popup Download powerpoint Table 2. Curated antibody-antigen complex training data from SAbDab. The fine-tuned models also performed similarly overall for a given antigen. As shown in Figure 2 , PD-1 and PD-L1 had the best overall distributions of values, indicating stable performance of the antibody generation for those antigens specifically. Conversely, in BBF-14 and BHRF1, have the largest disparity in overall generated antibody binding. Download figure Open in new tab Fig. 2. Boxplots depicting the predicted binding affinity score distribution by antigen and model. More negative values indicate better overall binding. Note that the y-axis is showing scores between [-250, 250]. Structural Examples Across the 1,050 generated sequences in this study, many have epitopes in the resulting predicted structures that closely match the desired epitope residues defined in the model input prompt. For example, in the BHRF1 antigen, a normal target protein is the BCL-2-like protein, shown in yellow in Figure 3B . This interaction enables apoptosis inhibition by the Epstein-Barr virus ( 20 ). Using the peleke-phi-4 model with an input prompt highlighting the active residues between BHRF1 and the BCL-2-like protein from PDB 2WH6, we are able to generate the antibody shown in pink in Figure 3A , which is predicted to block the interaction between BHRF1 and the BCL-2 homolog protein. Download figure Open in new tab Fig. 3. Comparison of (A) peleke-phi-4_BHRF1_0a , a novel generated anti-BHRF1 Fv structure, in pink; and (B) a BCL-2-like protein 11 (in yellow) bound to BHRF1 (in white, PDB: 2WH6). One of the best predicted binders across the set of 1,050 generated sequences was peleke-phi-4_MBP_20 with an overall binding score of −215.63 and a van der Waals energy of −102.422 kcal/mol. The antigen prompt for MBP (maltose binding protein) highlighted the residues surrounding the maltose ligand with the goal of blocking the binding pocket with the antibody. As shown in Figure 4A , this generated candidate is predicted to be a strong binder on the opposite side of the protein, missing the target epitope. Download figure Open in new tab Fig. 4. Comparison of (A) a novel anti-MBP candidate peleke-phi-4_MBP_20 (in pink); and (B) the maltose binding protein in its closed conformation (in white) with maltose bound inside (in purple, PDB: 1NL5). Interestingly, the side where the antibody bound is a conformationally-important site known as the “balancing interface” that controls the ligand-binding cleft to allow/block the maltose lig- and ( 21 ). Thus, this novel antibody, shown in pink in Figure 4A , could play a role in inhibiting this conformational change. Lastly, PD-1, a checkpoint inhibitor for which there are multiple approved antibody therapeutics, is a common oncotarget in various cancers. Ranging from metastatic melanoma to metastatic nonsmall cell lung cancer (NSCLC) to Hodgkin lymphoma, anti-PD-1 antibodies are used as a combination therapy with other chemotherapy drugs. As previously reported in Ford, 2024, pembrolizumab (a leading approved anti-PD-1 antibody, marketed as Keytruda ® by Merck & Co., Inc.) had a predicted van der Waals energy of −62.22 kcal/mol ( 22 ). Of the 150 candidates generated by the peleke-1 models, 83 (55.3%) of the candidates have better affinity scores than pembrolizumab and other anti-PD-1 antibodies. Shown in Figure 5 are the best candidates per model, all with predicted van der Waals energies <-100 kcal/mol. Download figure Open in new tab Fig. 5. Comparison of 3 novel anti-PD-1 candidates (A) peleke-phi-4_PD-1_18 (in pink); (B) peleke-mistral-7b-instruct-v0 . 2_PD-1_12 (in orange); and (C) pelekellama-3 . 1-8b-instruct_PD-1_30 (in blue); with (D) pembrolizumab (in yellow) bound to PD-1 (in white, PDB: 5JXE). As shown in these aforementioned examples, across various targets, the peleke-1 series of models often generates novel, targeted antibodies with strong binding potential. While additional testing across more target antigens is needed, the set of 7 benchmark antigens shown here proved useful in highlighting the learned behavior of the tuned models and where future improvements may increase performance, binding affinity, and epitope specificity. Discussion The use of AI in antibody discovery poses a great benefit to the therapeutic development process. As shown in this study, specialized PLMs such as peleke-1 provide a cost-effective method to computationally generate novel antibody sequences that reliably fold into desired structures and bind to the targets of interest. By fine-tuning pretrained language models on antibody-specific corpora, peleke-1 bridges the gap between large-scale protein representation learning and targeted therapeutic design. Compared to traditional workflows, which rely on animal immunization and hybridoma production, peleke-1 can significantly reduce the experimental lab workload. In contrast to prior PLMs that have been applied broadly across proteins, our suite of models explicitly encodes antibody sequence and antigen biases, enabling more targeted generation. This domain specialization highlights the value of task-specific fine-tuning for biologics development. Nevertheless, several limitations remain. As always, computational generation does not eliminate the need for experimental validation. Thus, candidate antibodies must still undergo biochemical characterization and functional assays. Second, while peleke-1 demonstrates strong performance in generating binders to predefined epitopes, generalization across diverse antigen classes remains an open challenge. Finally, like other deep learning systems, model interpretability and robustness require further investigation. The peleke-1 project was designed for ease-of-use and flexibility for antibody discovery. By utilizing the transformers framework from Hugging Face, we have enabled easy inferencing for users and simplified the tuning process for future developers wanting to extend or adapt this work. Looking forward, integrating peleke-1 with high-throughput screening pipelines and other structure-prediction frameworks could accelerate end-to-end antibody discovery. Moreover, additional tuning time and automated evaluation, along with incorporating feedback from experimental assays into iterative model design, may further improve binding specificity and developability. Ultimately, by combining computational generation with empirical validation, approaches like peleke-1 have the potential to transform antibody discovery into a faster, more scalable and accessible process with a higher likelihood of therapeutic success. Key Contributions This work showcases the following contributions: Domain-specialized PLMs for antibodies: We introduce peleke-1 , a suite of protein language models fine-tuned on antibody-specific corpora to enable targeted sequence generation. Epitope-conditioned design: We demonstrate that peleke-1 can generate antibody sequences conditioned on desired epitopes, enabling controllable design beyond general-purpose protein models. Bridging AI and molecular generation: By treating antibodies as structured sequences of amino acid tokens, we highlight how task-specific fine-tuning of PLMs will accelerate therapeutic discovery. Open-sourced for the community: We release model weights, curated training data, tuning code, and generation pipelines to promote reproducibility and enable further research in machine learning for antibody engineering. Contributors Authors NS and CTF developed the peleke fine-tuning pipeline. TP and CTF developed the training data preparation logic and evaluation pipeline. PN assisted in the research into candidate LLMs. CTF performed the evaluation of the models. All authors wrote and reviewed this manuscript. Declaration of Interests Author CTF is the owner of Tuple, LLC, a biotechnology consulting firm, and its subsidiary, Silico Biosciences. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest. Data Sharing Statement All code, data, results, and additional analyses are openly available on GitHub at: https://github.com/silicobio/peleke . This repository includes the open-source logic for tuning additional peleke models. Model weights are hosted on Hugging Face at https://huggingface.co/collections/silicobio/peleke-1-686c7e91c90f7f65a748698d . This collection of models includes the updated tokenizers and safetensors for each of the adapter models tuned and sample inferencing code. Funding Statement Personnel funding was provided in part by the NCBiotech Industrial Internship Program (Grant #: 2025-IIP-0021). Funding for cloud computational resources was provided by the Microsoft Most Valuable Professionals program. Supplementary Materials Download figure Open in new tab Fig. 1. Density plots of (A) total energy, (B) polar solvation, and (C) hydrophobic solvation. These stability metrics are shown by model and benchmark antigen. Download figure Open in new tab Fig. 2. WebLogos of the generated Fv sequences ( 23 ). View this table: View inline View popup Download powerpoint Table 1. Base models and LoRA configurations used in this study. View this table: View inline View popup Download powerpoint Table 2. Percentage of complexes with van der Waals energies <-25 kcal/mol (median value shown in parentheses) by fine-tuned model and antigen. Acknowledgments We acknowledge the following entities at the University of North Carolina at Charlotte: the Center for Computational Intelligence to Predict Health and Environmental Risks (CIPHER), the Department of Bioinformatics and Genomics, and the School of Data Science. We would like to think Rafael Jaimes III from MIT Lincoln Lab for his help in reviewing this manuscript. Funder Information Declared North Carolina Biotechnology Center, https://ror.org/03hj4jr80 , 2025-IIP-0021 Footnotes https://github.com/silicobio/peleke http://silico.bio/peleke Bibliography 1. ↵ Zeming Lin , Halil Akin , Roshan Rao , Brian Hie , Zhongkai Zhu , Wenting Lu , Nikita Smetanin , Robert Verkuil , Ori Kabeli , Yaniv Shmueli , Allan dos Santos Costa , Maryam Fazel-Zarandi , Tom Sercu , Salvatore Candido , and Alexander Rives . Evolutionary-scale prediction of atomic-level protein structure with a language model . Science , 379 ( 6637 ): 1123 – 1130 , 2023 . doi: 10.1126/science.ade2574 . OpenUrl CrossRef PubMed 2. ↵ Thomas Hayes , Roshan Rao , Halil Akin , Nicholas J. Sofroniew , Deniz Oktay , Zeming Lin , Robert Verkuil , Vincent Q. Tran , Jonathan Deaton , Marius Wiggert , Rohil Badkundri , Irhum Shafkat , Jun Gong , Alexander Derry , Raul S. Molina , Neil Thomas , Yousuf A. Khan , Chetan Mishra , Carolyn Kim , Liam J. Bartie , Matthew Nemeth , Patrick D. Hsu , Tom Sercu , Salvatore Candido , and Alexander Rives . Simulating 500 million years of evolution with a language model . Science , 387 ( 6736 ): 850 – 858 , 2025 . doi: 10.1126/science.ads0018 . OpenUrl CrossRef PubMed 3. ↵ J. Dauparas , I. Anishchenko , N. Bennett , H. Bai , R. J. Ragotte , L. F. Milles , B. I. M. Wicky , A. Courbet , R. J. de Haas , N. Bethel , P. J. Y. Leung , T. F. Huddy , S. Pellock , D. Tischer , F. Chan , B. Koepnick , H. Nguyen , A. Kang , B. Sankaran , A. K. Bera , N. P. King , and D. Baker . Robust deep learning–based protein sequence design using ProteinMPNN . Science , 378 ( 6615 ): 49 – 56 , 2022 . doi: 10.1126/science.add2187 . OpenUrl CrossRef PubMed 4. ↵ Henry Kenlay , Frédéric A. Dreyer , Aleksandr Kovaltsuk , Dom Miketa , Douglas Pires , and Charlotte M. Deane . Large scale paired antibody language models , 2024 . 5. ↵ Tobias H Olsen , Iain H Moal , and Charlotte M Deane . AbLang: an antibody language model for completing antibody sequences . Bioinformatics Advances , 2 ( 1 ): vbac046 , 06 2022 .ISSN 2635-0041 . doi: 10.1093/bioadv/vbac046 . OpenUrl CrossRef 6. ↵ Jeffrey A. Ruffolo , Jeffrey J. Gray , and Jeremias Sulam . Deciphering antibody affinity maturation with language models and weakly supervised learning , 2021 . 7. ↵ Rubo Wang , Fandi Wu , Jiale Shi , Yidong Song , Yu Kong , Jian Ma , Bing He , Qihong Yan , Tianlei Ying , Peilin Zhao , Xingyu Gao , and Jianhua Yao . A Generative Foundation Model for Antibody Design . bioRxiv , 2025 . doi: 10.1101/2025.09.12.675771 . OpenUrl Abstract / FREE Full Text 8. ↵ Pritam Kumar Panda . A comprehensive python application for mapping and visualizing molecular interactions at protein/nucleic acid interfaces and tool for mapping protein-protein, protein-nucleic acid, and antigen-antibody interactions . https://github.com/pritampanda15/PandaProt , 2025 . GitHub repository . 9. Marah Abdin , Jyoti Aneja , Harkirat Behl , Sébastien Bubeck , Ronen Eldan , Suriya Gunasekar , Michael Harrison , Russell J. Hewett , Mojan Javaheripi , Piero Kauffmann , James R. Lee , Yin Tat Lee , Yuanzhi Li , Weishung Liu , Caio C. T. Mendes , Anh Nguyen , Eric Price , Gustavo de Rosa , Olli Saarikivi , Adil Salim , Shital Shah , Xin Wang , Rachel Ward , Yue Wu , Dingli Yu , Cyril Zhang , and Yi Zhang . Phi-4 technical report , 2024 . 10. Aaron Grattafiori et al. The Llama 3 Herd of Models , 2024 . 11. Albert Q. Jiang , Alexandre Sablayrolles , Arthur Mensch , Chris Bamford , Devendra Singh Chaplot , Diego de las Casas , Florian Bressand , Gianna Lengyel , Guillaume Lample , Lucile Saulnier , Lélio Renard Lavaud , Marie-Anne Lachaux , Pierre Stock , Teven Le Scao , Thibaut Lavril , Thomas Wang , Timothée Lacroix , and William El Sayed . Mistral 7b , 2023 . 12. ↵ Tudor-Stefan Cotet , Igor Krawczuk , Filippo Stocco , Noelia Ferruz , Anthony Gitter , Yoichi Kurumida , Lucas de Almeida Machado , Francesco Paesani , Cianna N. Calia , Chance A. Challacombe , Nikhil Haas , Ahmad Qamar , Bruno E. Correia , Martin Pacesa , Lennart Nickel , Kartic Subr , Leonardo V. Castorina , Maxwell J. Campbell , Constance Ferragu , Patrick Kidger , Logan Hallee , Christopher W. Wood , Michael J. Stam , Tadas Kluonis , Süleyman Mert ünal , Elian Belot , Alexander Naka , and Adaptyv Competition Organizers . Crowdsourced Protein Design: Lessons From the Adaptyv EGFR Binder Competition . bioRxiv , 2025 . doi: 10.1101/2025.04.17.648362 . OpenUrl Abstract / FREE Full Text 13. ↵ James Dunbar and Charlotte M. Deane . Anarci: antigen receptor numbering and receptor classification . Bioinformatics , 32 ( 2 ): 298 – 300 , 09 2015 .ISSN 1367-4803 . doi: 10.1093/bioinformatics/btv552 . OpenUrl CrossRef PubMed 14. ↵ Lewis Chinery , Jeliazko R. Jeliazkov , and Charlotte M. Deane . Humatch - fast, gene-specific joint humanisation of antibody heavy and light chains . bioRxiv , 2024 . doi: 10.1101/2024.09.16.613210 . OpenUrl Abstract / FREE Full Text 15. ↵ Javier Delgado , Leandro G Radusky , Damiano Cianferoni , and Luis Serrano . FoldX 5.0: working with RNA, small molecules and a new graphical interface . Bioinformatics , 35 ( 20 ): 4168 – 4169 , 03 2019 .ISSN 1367-4803 . doi: 10.1093/bioinformatics/btz184 . OpenUrl CrossRef PubMed 16. Nicholas Santolla and Colby T. Ford . Ai-based antibody design targeting recent h5n1 avian influenza strains . Computational and Structural Biotechnology Journal , 27 : 2915 – 2923 , Jan 2025 .ISSN 2001-0370 . doi: 10.1016/j.csbj.2025.06.026 . OpenUrl CrossRef 17. ↵ João M.C. Teixeira , Rodrigo Vargas Honorato , Marco Giulini , Alexandre Bonvin , SarahAlidoost , Victor Reys , Brian Jimenez , Douwe Schulte , Charlotte van Noort , Stefan Verhoeven , Barbara Vreede , SSchott , and Regen Tsai . haddocking/haddock3: v3.0.0-beta.5 . Zenodo , January 2024 . doi: 10.5281/zenodo.10527751 . OpenUrl CrossRef 18. ↵ Bonvin Lab . HADDOCK3 Antibody-Antigen Tutorial , 2024 . Accessed: 2024-06-21 . 19. ↵ Sarah Alamdari , Nitya Thakkar , Rianne van den Berg , Neil Tenenholtz , Robert Strome , Alan M. Moses , Alex X. Lu , Nicolò Fusi , Ava P. Amini , and Kevin K. Yang . Protein generation with evolutionary diffusion: sequence is all you need . bioRxiv , 2024 . doi: 10.1101/2023.09.11.556673 . OpenUrl Abstract / FREE Full Text 20. ↵ Marc Kvansakul , Andrew H. Wei , Jamie I. Fletcher , Simon N. Willis , Lin Chen , Andrew W. Roberts , David C. S. Huang , and Peter M. Colman . Structural Basis for Apoptosis Inhibition by Epstein-Barr Virus BHRF1 . PLOS Pathogens , 6 ( 12 ): 1 – 10 , 12 2010 . doi: 10.1371/journal.ppat.1001236 . OpenUrl CrossRef PubMed 21. ↵ Patrick G. Telmer and Brian H. Shilton . Insights into the Conformational Equilibria of Maltose-binding Protein by Analysis of High Affinity Mutants . Journal of Biological Chemistry , 278 ( 36 ): 34555 – 34567 , Sep 2003 .ISSN 0021-9258 . doi: 10.1074/jbc.M301004200 . OpenUrl Abstract / FREE Full Text 22. ↵ Colby T. Ford . PD-1 Targeted Antibody Discovery Using AI Protein Diffusion . Technology in Cancer Research & Treatment , 23 : 15330338241275947 , 2024 . doi: 10.1177/15330338241275947 . PMID: 39228166 . OpenUrl CrossRef PubMed 23. ↵ Gavin E Crooks , Gary Hon , John-Marc Chandonia , and Steven E Brenner . WebLogo: a sequence logo generator . Genome Res ., 14 ( 6 ): 1188 – 1190 , June 2004 . OpenUrl Abstract / FREE Full Text View the discussion thread. Back to top Previous Next Posted October 16, 2025. Download PDF Data/Code Email Thank you for your interest in spreading the word about bioRxiv. NOTE: Your email address is requested solely to identify you as the sender of this article. Your Email * Your Name * Send To * Enter multiple addresses on separate lines or separate them with commas. You are going to email the following peleke-1: A Suite of Protein Language Models Fine-Tuned for Targeted Antibody Sequence Generation Message Subject (Your Name) has forwarded a page to you from bioRxiv Message Body (Your Name) thought you would like to see this page from the bioRxiv website. Your Personal Message CAPTCHA This question is for testing whether or not you are a human visitor and to prevent automated spam submissions. Share peleke-1: A Suite of Protein Language Models Fine-Tuned for Targeted Antibody Sequence Generation Nicholas Santolla , Trey Pridgen , Prbhuv Nigam , Colby T. Ford bioRxiv 2025.10.16.682644; doi: https://doi.org/10.1101/2025.10.16.682644 Share This Article: Copy Citation Tools peleke-1: A Suite of Protein Language Models Fine-Tuned for Targeted Antibody Sequence Generation Nicholas Santolla , Trey Pridgen , Prbhuv Nigam , Colby T. Ford bioRxiv 2025.10.16.682644; doi: https://doi.org/10.1101/2025.10.16.682644 Citation Manager Formats BibTeX Bookends EasyBib EndNote (tagged) EndNote 8 (xml) Medlars Mendeley Papers RefWorks Tagged Ref Manager RIS Zotero Tweet Widget Facebook Like Google Plus One Subject Area Immunology Subject Areas All Articles Animal Behavior and Cognition (7618) Biochemistry (17636) Bioengineering (13860) Bioinformatics (41847) Biophysics (21401) Cancer Biology (18536) Cell Biology (25424) Clinical Trials (138) Developmental Biology (13353) Ecology (19860) Epidemiology (2067) Evolutionary Biology (24287) Genetics (15583) Genomics (22463) Immunology (17701) Microbiology (40300) Molecular Biology (17141) Neuroscience (88434) Paleontology (666) Pathology (2825) Pharmacology and Toxicology (4813) Physiology (7633) Plant Biology (15107) Scientific Communication and Education (2042) Synthetic Biology (4285) Systems Biology (9808) Zoology (2268)

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00