Reinforcement Learning-Augmented ProteinMPNN Improve the Binding Affinity of TNFR1-Targeting Minibinders

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Minibinders are compact protein molecules that hold great promise as therapeutic agents due to their target-binding specificity, stability, and potential for oral delivery. However, the primary objective ofthe widely used minibinder sequence design method, ProteinMPNN, is structural fidelity rather than the optimization of specific functional properties such as binding affinity. Consequently, the de novo design of high-affinity minibinders still largely relies on extensive wet-lab screening of numerous candidates. Directly designing high-affinity minibinders in silico thus remains a critical challenge for numerous therapeutic applications. Here, we present a computational framework that integrates a reinforcement learning (RL) framework with ProteinMPNN network to directly generate minibinder sequences with improved functional features. We demonstrate the power of this framework by optimizing the activities of the minibinders targeting Tumor Necrosis Factor Receptor 1 (TNFR1). Experimental validation showed that the optimized minibinders OPT1 and OPT7 exhibited 3-fold and 7-fold higher binding affinity, respectively, and 6-fold and 4-fold greater neutralizing activity in cells compared to the original minibinder S1B2 . Our work establishes the framework as a promising tool that augments AI-driven protein design for the de novo development of high-affinity therapeutic minibinders.
Full text 126,921 characters · extracted from preprint-html · click to expand
Reinforcement Learning-Augmented ProteinMPNN Improve the Binding Affinity of TNFR1-Targeting Minibinders | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Reinforcement Learning-Augmented ProteinMPNN Improve the Binding Affinity of TNFR1-Targeting Minibinders Zigong Wei, Lin Wei, Zhiyong Wu, Yang Hu, Yihe Fang, Miaomiao Geng, and 4 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8645836/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Minibinders are compact protein molecules that hold great promise as therapeutic agents due to their target-binding specificity, stability, and potential for oral delivery. However, the primary objective ofthe widely used minibinder sequence design method, ProteinMPNN, is structural fidelity rather than the optimization of specific functional properties such as binding affinity. Consequently, the de novo design of high-affinity minibinders still largely relies on extensive wet-lab screening of numerous candidates. Directly designing high-affinity minibinders in silico thus remains a critical challenge for numerous therapeutic applications. Here, we present a computational framework that integrates a reinforcement learning (RL) framework with ProteinMPNN network to directly generate minibinder sequences with improved functional features. We demonstrate the power of this framework by optimizing the activities of the minibinders targeting Tumor Necrosis Factor Receptor 1 (TNFR1). Experimental validation showed that the optimized minibinders OPT1 and OPT7 exhibited 3-fold and 7-fold higher binding affinity, respectively, and 6-fold and 4-fold greater neutralizing activity in cells compared to the original minibinder S1B2 . Our work establishes the framework as a promising tool that augments AI-driven protein design for the de novo development of high-affinity therapeutic minibinders. Biological sciences/Computational biology and bioinformatics/Computational models Health sciences/Medical research/Drug development Minibinder Protein design Reinforcement learning ProteinMPNN Tumor Necrosis Factor Receptor 1 (TNFR1) Figures Figure 1 Figure 2 Figure 3 Introduction Minibinders, as conceptualized by Baker and colleagues 1 , are compact protein molecules (typically 4–12 kDa) composed of natural amino acids that exhibit specific target-binding capabilities. To enable data-driven de novo minibinder design, Baker et al. constructed a set of small proteins comprising 20–50 amino acid peptides, which fold into stable conformations and demonstrate resistance to chymotrypsin and trypsin 2 . Leveraging this database, minibinders targeting to influenza virus hemagglutinin A (HA) and botulinum neurotoxin B were successfully designed 1 . These molecules exhibit stable folded conformations, protease resistance, low immunogenicity, and enhanced tissue permeability compared to antibodies 1 – 3 . Berger et al. 4 further validated oral delivery efficacy of an IL-23R-targeting minibinder in murine colitis models, achieving therapeutic outcomes comparable to injected antibodies. Designing conformationally stable minibinders remains challenging. Baker et al., pioneered computational de novo minibinders design by Rosetta, a physics-based approach, which treats sequence design as an energy optimization problem for a given input structure. Based on Rosetta, minibinders targeting SARS-CoV-2 spike protein 3 , IL-7Rα, VirB8, TrkA 5 , TNFR1 6 , PRV gD 7 , IL-6R, GP130, and IL-1R1 receptor subunits 8 has been successfully designed. However, due to incomplete sampling of conformational space with physics-based Rosetta, experimental success rates were low, often requiring the screening of thousands of designs 9 . Moreover, reliance on pre-defined protein scaffolds inherently limits the diversity and shape complementarity of potential solutions 9 . AI-driven strategies can overcome these limitations by leveraging large protein datasets. Deep-learning approaches can efficiently generate sequences for monomeric backbones without rotamer enumeration 10 – 17 . Crucially, ProteinMPNN 18 enhances the sequences folding fidelity of the backbones designed by Rosetta, as predicted by AlphaFold2(AF2) 19 , increasing the success rate from 2.7% to 54.1% and improving utility for functional sites 18 . The autoregressive MPNN utilizes backbone geometry (distances, orientations, dihedrals) to achieve 50.5% native sequence recovery, outperforming predecessors (52.4% vs. 32.9%). 18 A backbones design strategy, RFdiffusion 17 , employs denoising diffusion probabilistic models (DDPMs) 20 with RoseTTAFold (RF) 21 frame representations (Cα coordinates + N-Cα-C orientations) to create diverse structures from random initialization. 20 Baker et al. accurately reproduced the crystal structure of the binder LCB1 targeting the SARS-CoV-2 receptor binding domain (RBD), originally reported by the lab in 2020 18 , using RFDiffusion (skeleton design) combined with ProteinMPNN (sequence design), demonstrating the reliability of this AI approach 9 . Similarly, in our previous work 22 , 23 , with the same protocol, minibinders targeting M-like protein (SzM), Glyceraldehyde-3-phosphate dehydrogenase (GAPDH) of Streptococcus equi ssp. Zooepidemicus (SeZ) has been successfully designed, improved the survival rate (80%, 60%, respectively) of mice infected with a 100-fold lethal dose of SEZ and significantly reduced the organ bacterial load. While ProteinMPNN significantly improves sequence recovery rates compared to Rosetta (~ 52% vs. ~33%) 18 , its primary objective remains structural fidelity rather than optimizing specific functional properties like binding affinity. Consequently, the de novo design of high-affinity minibinders still largely relies on extensive wet-lab screening of numerous candidates. Therefore, beyond ensuring high sequence design success, directly designing high-affinity minibinders in silico represents a critical challenge for numerous therapeutic applications 24 . Reinforcement learning(RL) 25 is a machine learning paradigm well-suited to address such optimization problems. An RL agent learns to make decisions by interacting with an environment to maximize cumulative rewards. RL-based models like AlphaGo 26 and DeepSeek-R1 27 have achieved significant success. RL concepts are widely applied to drug design 28 – 30 and protein design 31 , often directing generative models by maximizing a multi-parameter optimization (MPO) scoring function reflecting key features. 29 , 32 , 33 REINVENT 28 , REINVENT2 29 , Reinvent4 30 tackle the inverse design problem through reinforcement learning using RNNs or transformers as deep learning architectures to generate molecules with desirable properties by learning an augmented episodic likelihood combining prior likelihood 34 and a user-defined scoring function. Baker et al. reported a "top-down" RL approach using Monte Carlo tree search to sample conformers within an architecture under functional constraints 31 , successfully designing disk-shaped nanopores and ultracompact icosahedra verified by cryo-EM. In this work, we introduce a computational platform that integrates a RL framework with ProteinMPNN network to directly generate minibinder sequences for high-affinity binding. We demonstrate the efficacy of the framework by optimizing a previously reported TNF receptor type 1 (TNFR1)-targeting minibinder, S1B2 6 . Tumor necrosis factor alpha (TNFα) is a pleiotropic cytokine with dual roles: it provides beneficial defense against infections and cancer, yet also exerts deleterious pro-inflammatory and cytotoxic effects during inflammation 35 , 36 . TNFα signals via two distinct cell surface receptors, TNFR1 and type 2 (TNFR2) 37 . Broadly, TNFR1 primarily mediates TNFα’s inflammatory and apoptotic effects. 38 Targeting TNFR1 specifically disrupts the TNFα-TNFR1 interaction while preserving TNFα's beneficial homeostatic immune functions. In the previous work 6 , binder S1B2 exhibited strong binding activity (K D = 890 nM). Therefore, it is necessary to optimize the small minibinder like S1B2 , to enhance its activity. In this work, the generated minibinders by the RL-augmented ProteinMPNN framework shown “Prefect” (minibinders with predicted negative logarithm dissociation constant to base 10, pK D >7) rate increasing with 28%. And OPT7 were designed from S1B2 resulting in up to 7-fold enhanced activity. Results and Discussion Reinforcement learning framework. The RL-augmented framework with ProteinMPNN network presents an integrated computational pipeline for designing target-binding minibinders (Fig. 1 ). The process begins with generated backbone, which was created to bind specific target sites using tools like RFdiffusion or Rosetta. The framework is based on policy-based reinforcement learning (RL) framework. Within this framework, an “Agent” acts as the sequence designer. A pre-trained general ProteinMPNN model 18 serves as a “Prior” policy, which initially shares its network weights with the “Agent”. During training, the “Agent” generates protein sequences through ProteinMPNN with an augmented training model. The environment is represented by a multi-parameter objective (MPO) scoring function, which evaluates these sequences and provides a reward. A key step involves forming an augmented training target by combining the scaled MPO reward with the evaluation of “Prior” sequence likelihood (negative log-likelihood, NLL). The “Agent” training loss is then defined as the squared difference between its own predicted likelihood and this augmented target. This objective guides the “Agent” to learn a policy that generates sequences with improved functional scores while remaining grounded in plausible sequence-structure relationships. Throughout the RL process, only the Agent's parameters are updated via this loss, iteratively refining its policy. Finally, optimized sequences are sampled using the trained Agent model and are structurally validated with tools like AF2-multimer 14 . More details see Materials and Methods. Designed TNFR1 minibinders by RL-augmented ProteinMPNN. In RL training, The NLL of the augmented likelihood and “Agent” is shown in Fig. S1 , the NLL of the augmented likelihood stabilized at 2.25, while the NLL of “Agent” increased from 1.9 to 2.0 in the first 800 steps and then showed a slight downward trend. Following RL training, 1000 minibinders were generated from “Agents” model at every 100 steps. The negative logarithm dissociation constant to base 10 (pK D ) of all candidates were predicted by ProAffinity-GNN 39 , whose applicability for predicting the binding affinity of de novo designed minibinders were validated here with the data from Berger’s work 4 (see Supplementary information Part II ). Although Fig. 2 A shows that the binding affinities of the samples from the 'Agent' models are primarily distributed between 5 and 6.5, similar to the samples from original ProteinMPNN network model (“Prior”), a distinct “peak” emerged around pK D =7 in the “Agents” distributions, indicating an increased probability of generating high-affinity candidates. This “peak” intensified up to 800 training steps, after which it gradually diminished. When examining the pK D values of candidates sampled by each “Agent” model, we found that the top 100 candidates and numbers of candidates with pK D >7 showed similar trends. The average pKa raised from 7.35 in the “Prior” to 7.40 at 800 steps with no obvious trending, then decreased to 7.28 at 1000 steps (Fig. 2 B), while numbers of candidates with pK D >7 (here we treat that as “Prefect” minibinders) rose from 110 in the “Prior” to 141 at 800 steps, showing a 28% increase in the "Perfect" rate, although it then decreasing to 96 at 1000 steps (Fig. 2 C). The above data indicate that during the first 800 training steps, the prediction activity of the sampling from “Agent” model gradually increased, reflecting the “Agent” model in RL process learned to explore sequences with higher binding affinities. This confirms the power of the design of high-activity minibinders was enhanced through RL beginning with original ProteinMPNN model. Since we are generating protein sequences through reinforcement learning-augmented ProteinMPNN rather than the backbone, and no structural properties were used as a reward function, naturally, the structural rationality is not expected to undergo significant changes. This is supported by AF3Score 40 analysis of the Agent's sampling results. For minibinder Cα atom predicted local distance difference test (pLDDT), predicted TM score (pTM) and predicted aligned error (pAE), although a slight worsening tendency is observed, the distribution ranges remain consistent, suggesting that the changes stay within a controllable range (Fig. 2 D-F). There was no significant difference in Cα atom pLDDT, pAE, pTM or interface predicted TM score(iPTM) for the complexes between “Prior” and “Agents” (Fig. 2 G-J). More importantly, the distribution of the binder-target interface metric, pAE_interaction, shows no significant difference (Fig. 2 K). All of this indicates that, in this RL process, the structural rationality was not significantly compromised. To further evaluate the activity of the candidate structures sampling from the “Agent” models, we performed molecular dynamics simulation (MD) optimization on all candidates from the above “Agents” after the condition-based filtering (pAE_interaction < 10), and calculated the binding affinity using MM/GBSA. For the structures with predicted binding affinity higher than S1B2 , we conducted manual inspection to select minibinders with a higher likelihood of binding. Subsequently, we selected 7 sequences for expression and purification. The obtained OPT1-OPT7 reduced the predicted binding free energy (ΔG) by 36.1–1.93 kcal/mol (Table 1 ), indicating extremely high potential activity and demonstrating the feasibility and advancement of the RL-Augmented ProteinMPNN framework. Table 1 The top binding affinities predicted by MM/GBSA for RL-augmented ProteinMPNN designed mini-binder targeting TNFR1. ID Sequence Weight (kDa) ΔG(MM/GBSA) (kcal/mol) S1B2 DAEEQLRIQERLIELAFRLGDPETAERIARAAEILLKEDGDPEAVEIIRELLERLR 6.51 -53.29 ± 6.52 OPT1 SLEENLKFALNEIKRAFRQGKPEEARRTAEAAISWFKEQGDEESVKEIEELLKKLE 6.56 -89.36 ± 6.46 OPT2 SMKENLELALTQIRLAVRRGEPEEAKRIADAAVSVFREFGDEESVKKIEELLKELL 6.40 -65.30 ± 6.19 OPT3 SLKENLKFTLNQIRLAFRRGNPEEAKAIAEAAISVFKEFGDEESVEELKKLLEELL 6.41 -65.18 ± 3.71 OPT4 SSKENTEFAINQIKLAFRQGNPEEARQIAKAAVSVLTEMGDPEGVERVKAVLEELE 6.17 -62.13 ± 4.20 OPT5 SSKKNLEFALNQIRLAFRQGKPEEAKRIAKAAVKVLLEQGDEEGAKKVEELLKELE 6.32 -56.70 ± 5.78 OPT6 MAKENKEFALNQIRLAVSRGEPEEAEKIARASVKVLKEMGDPESAEEVEKLLKELL 6.30 -56.67 ± 3.85 OPT7 SAEENLEFALTQIRLAFRQGKPEEAKRIAEAAIKVLSEMGYPEGVEKIKKLLEELE 6.35 -55.22 ± 3.46 Expression, purification for the 7 minibinders. Each minibinder was expressed in Escherichia coli BL21(DE3) using a bacterial expression system, taking advantage of the well-established suitability of prokaryotic hosts for producing small, stable, single-domain protein 41 . Except for OPT3 , all others were successfully expressed. Following expression, the minibinders were purified by immobilized metal affinity chromatography (IMAC, Ni-NTA) followed by size-exclusion chromatography (FPLC). SDS-PAGE analysis confirmed successful expression with high purity, and molecular weights between 6–7 kDa, ( Fig. S2 ), consistent with theoretical calculations (Table 1 ). BLI and cell assays of designd minibinders The affinities of purified minibinders to human TNFR1 were assessed by Biolayer Interferometry (BLI). Here, S1B2 showed an affinity with K D = 1.47 µM ( Fig. S3 ), consistent with previous work (K D = 0.88 µM 6 ). After optimization, the binding affinities (K D ) of OPT1 , OPT2 , OPT4 , OPT5 , and OPT7 (Fig. 3 A-E) were determined to be 0.48 µM, 1.59 µM, 20.30 µM, 54.88 µM, and 0.23 µM, respectively. (Fig. 3 F-J) However, OPT6 did not show high binding affinity ( Fig. S4 ). Compared to S1B2 , OPT1 and OPT7 increased 3-fold and 7-fold, respectively, while OPT2 maintained its activity. The mouse fibroblast L929 cell line is sensitive to TNFα when treated with actinomycin D. Therefore, as in our previous work 6 , we applied an L929 cell assay to assess the blocking effect of antagonist molecules on TNFR1. S1B2 exhibited dose-dependent TNFα-blocking activity with IC 50 of 7.75 nM (Fig. 3 K), consistent with it in our previous work (4.32 nM 6 ). OPT1 , OPT2 , OPT4 , OPT5 , and OPT7 presented higher potency to TNFR1 in L929 cell assay, The IC 50 s were 1.25 nM, 12.92 nM, 22.39 nM, 13.04 nM and 1.94 nM, respectively (Fig. 3 K). Compared to S1B2 , the neutralizing activity of OPT1 and OPT7 increased by 6-fold and 4-fold, respectively, while the IC 50 of OPT2 is comparable. The consistency between K D and IC 50 suggests that these binders indeed target TNFR1. Analysis binding energy of the binder-TNFR1 complexes. We analyzed the energy contribution of each residue predicted by MM/GBSA. The results showed that for OPT1 , the energy contribution between each residue range from − 13.30 to 0.54 kcal/mol, with 75% being positive. The residues R15, E38, and R18 made prominent contributions through strong electrostatic interactions with E149, R77, and E149 of TNFR1, respectively. Residues N5, E12, and E24 of OPT1 also contributed significantly to binding affinity through hydrogen bond interactions with W107, Y103, and Q102 of TNFR1, respectively (Fig. 3 L). For OPT7 , the energy contribution between each residue ranged from − 13.16 to 0.21 kcal/mol, shown nearly no negative contribution. The strong electrostatic interaction between R18 and E149 (TNFR1) had a prominent contribution to affinity, while the electrostatic interactions between E38, K34 and R77, E79 of TNFR1, as well as interactions involving N5, Q12, and E24 with W107, Y103, N116 (TNFR1), also played important roles (Fig. 3 M). In contrast, in S1B2 , residues at positions 12, 15, and 34, which are LEU, LEU, and ILE respectively, cannot form electrostatic or hydrogen bond interactions. We speculate that the probability of incorporating residues at these positions capable of forming favorable interactions was increased. Discussion In this work, by employing affinity prediction as the reward function in a reinforcement learning framework, we successfully introduced multiple key interactions onto the S1B2 backbone (Fig. 3 L-M). The minibinder OPT7 , designed with the RL-augmented ProteinMPNN model based on S1B2 backbone, achieved a 7-fold increase in activity without compromising structural rationality. While this study was primarily intended as a methodological proof‑of‑concept rather than the discovery of a clinical candidate, and the achieved affinity does not yet match that of some reported high-affinity binders(K D =4 pM in Baker’s work 42 ), but the immunogenicity risk and challenges associated with oral delivery for the ~ 60 amino acid minibinders like OPT7 have been evaluated. Thus, our work successfully demonstrates that RL can effectively enhance the functional performance of the previous pre-trained general 18 ProteinMPNN model, steering protein design towards desired functional properties. Conclusion In this work, we successfully develop an innovative AI framework that integrates Reinforcement Learning (RL) with the protein sequence design network ProteinMPNN to advance the de novo design of therapeutic minibinders. excel at generating sequences that fold into stable structures, they are not inherently optimized for functional properties like binding affinity. Our Reinforcement Learning-Augmented ProteinMPNN framework addresses this gap by employing an RL agent to steer the sequence generation process, directly optimizing for higher affinity binding against a target of interest. Using TNFR1 as a proof-of-concept target, we demonstrated that the framework efficiently generates minibinders with superior predicted binding affinity. Experimental validation confirmed that the designed candidates exhibit significantly improved binding affinity and neutralizing potency in cellular assays. Structural and energetic analyses further elucidated the origin of these gains, linking them to specific, optimized interactions at the protein–protein interface. This work highlights a pivotal step in AI-driven protein design, shifting the focus from mere structural recapitulation toward the direct, goal-oriented optimization of biological function. Extending this approach, the Reinforcement Learning-Augmented ProteinMPNN framework is generalizable, offering a powerful and efficient strategy, can be trained with alternative reward functions targeting other traits, such as thermal stability and solubility, to directly generate proteins with the desired properties. This strategy paves the way for the future development of improved, fit‑for‑purpose minibinders. Materials and Methods Backbone Generation. Before run RL training, we must design a backbone. RFdiffusion 17 was employed to directly generate target-bound minibinder backbone structures. For therapeutic goals such as blocking protein–protein interaction, binding specific target sites ("interface hotspots"). Rosetta can also use to generated backbone, as our previous work. 6 , 7 Cost Function for Reinforcement Learning (RL). Figure 1 illustrates the components of the reinforcement learning (RL) within RL-Augmented ProteinMPNN framework. Here, we frame the problem of designing a protein sequence with specified properties as a Reinforcement Learning (RL) task. In policy-based 25 , 43 reinforcement learning (RL), an agent interacts with an environment by selecting actions based on its current state. Here, the state space is denoted as \(\:S\) , and for a given state \(\:s\in\:S\) , the set of available actions is A(s). The agent's behavior is defined by a stochastic policy \(\:\pi\:\left(a\right|s)\) , which outputs a probability distribution over actions for each state. When the agent generates a sequence of actions \(\:A=({a}_{1},...,{a}_{i},...,{a}_{T})\) given corresponding states \(\:({s}_{1},...,{s}_{T})\) , the probability of this trajectory under the policy is: $$\:P\left(A\right)={\prod\:}_{t=1}^{T}\pi\:\left({a}_{t}\right|{s}_{t})$$ 1 Each state-action pair yields an immediate reward \(\:r\left(a\right|s)\) , and the long-term return from time \(\:t\:\) is the cumulative reward \(\:G({a}_{t},{s}_{t})={\sum\:}_{t}^{T}{r}_{t}\) . The goal of RL is to improve the policy to maximize the expected return over trajectories. In the standard algorithm, REINFORCE 43 . with a sampled trajectory, the policy parameters \(\:\theta\:\) are updated via: \(\:{\theta\:}_{t+1}={\theta\:}_{t}-\alpha\:{\nabla\:}_{\theta\:}\sum\:_{t=0}^{T}\text{log}{\pi\:}_{\theta\:}\left({a}_{t}\right|{s}_{t}\left)\right(G\left({a}_{t}|{s}_{t}\right)-b)\) (2) 44 where \(\:\alpha\:\) is the step size and \(\:b\) is a baseline reward (often set to zero). Setting \(\:b=0\) simplifies the cost function to maximizing: \(\:J\left(\theta\:\right)=\sum\:_{t=0}^{T}\text{log}{\pi\:}_{\theta\:}\left({a}_{t}\right|{s}_{t})G\left({a}_{t}|{s}_{t}\right)\) (3) 44 If we consider a sparse reward setting, where reward is zero at all steps except the final one, and the final reward equals the total return \(\:G\left(A\right)\) , the objective reduces to: $$\:J\left(\theta\:\right)=\sum\:_{t=0}^{T}\text{log}{\pi\:}_{\theta\:}\left({a}_{t}|{s}_{t}\right)G\left({a}_{t}|{s}_{t}\right)=G\left(A\right)\sum\:_{t=0}^{T}\text{log}{\pi\:}_{\theta\:}\left({a}_{t}|{s}_{t}\right)=G\left(A\right)\text{l}\text{o}\text{g}\sum\:_{t=0}^{T}{\pi\:}_{\theta\:}\left({a}_{t}|{s}_{t}\right)=G\left(A\right)\text{log}P\left(A\right)$$ 4 Here, \(\:\text{log}P\left(A\right)\) represents the log-likelihood of the trajectory under the agent’s policy. We can reinterpret this objective by introducing an augmented likelihood, which combines a prior model’s assessment of the sequence with a measure of its desirability. Let \(\:\text{log}{P\left(A\right)}_{Prior}\) be the log-likelihood under a fixed prior policy, we then define an augmented log-likelihood as: \(\:\text{log}{P\left(A\right)}_{Augmented}=\text{log}{P\left(A\right)}_{Prior}+\sigma\:\:\times\:\:\text{S}\left(\text{A}\right)\) (5) 28 where \(\:\text{S}\left(\text{A}\right)\) is an external scoring function that quantifies sequence quality, and \(\:\sigma\:>0\) is a scaling factor. If we choose the reward for the final step of the sequence A to: (6) \(\:G\left(A\right)\:=\:{\left[\text{log}{P\left(A\right)}_{Agent}-\text{log}{P\left(A\right)}_{Augmented}\right]}^{2}/\text{log}{P\left(A\right)}_{Agent}\) (6) 28 then the agent’s objective becomes: $$\:J\left(\theta\:\right)=\frac{{\left[\text{log}{P\left(A\right)}_{Agent}-\text{log}{P\left(A\right)}_{Augmented}\right]}^{2}}{\text{log}{P\left(A\right)}_{Agent}}\text{l}\text{o}\text{g}{\text{P}\left(A\right)}_{Agent}={\left[\text{log}{P\left(A\right)}_{Agent}-\text{log}{P\left(A\right)}_{Augmented}\right]}^{2}$$ 7 This formulation encourages the agent’s policy to produce log-likelihoods that align with the augmented target, which itself balances prior likelihood and an external score. The Agent Network. In our specific setting, the task is to generate protein sequences that fold into a given backbone structure \(\:X\) , The prior policy is a pre-trained ProteinMPNN network model, which defines the conditional distribution: $$\:{p}_{Prior}\left(Y|X\right)={p}_{\text{P}\text{r}\text{o}\text{t}\text{e}\text{i}\text{n}\text{M}\text{P}\text{N}\text{N}}\left(Y|X\right)=P\left({y}_{t}\right|{y}_{t-1},\cdots\:,{y}_{1},\varvec{X})$$ 8 Its negative log-likelihood (NLL) for a sequence \(\:Y\) is: $$\:{NLL\left(Y\right|X)}_{Prior}\:=-{\sum\:}_{n}\text{log}{p}_{Prior}\left(Y\right|X)$$ 9 The agent is another ProteinMPNN network model with the same architecture but trainable parameters. Sequences generated by the agent are evaluated using a multiparameter objective (MPO) score, which measures desired biochemical or structural properties. According to Eq. 5, the MPO score is scaled by \(\:\sigma\:\) and combined with the prior NLL to form the augmented target: $$\:{NLL\left(\varvec{Y}\right|\varvec{X})}_{Augmented}={NLL\left(\varvec{Y}\right|\varvec{X})}_{Prior}+\sigma\:\:\times\:\:MPO\left(\:\varvec{Y}\right|\varvec{X})$$ 10 According to Eq. 7 , the agent is trained to minimize the squared difference between its own predicted NLL and this augmented target: $$\:loss={[{NLL\left(\varvec{Y}\right|\varvec{X})}_{Augmented}\:-\:{NLL\left(\varvec{Y}\right|\varvec{X})}_{Agent}]}^{2}$$ 11 Only the agent’s parameters are updated via backpropagation; the prior remains fixed. This approach anchors the agent to the prior distribution while steering it toward sequences with higher MPO scores, effectively integrating generative modeling and goal-directed optimization into a single differentiable framework. This method provides a principled alternative to direct reward maximization with REINFORCE 43 . By framing the update as matching an augmented likelihood, which blends prior knowledge and task-specific desirability, we maintain proximity to the prior distribution, avoid high-variance gradient estimates, and enable stable fine-tuning. The result is an agent that generates sequences which are both plausible under the structural prior and optimized for the specified scoring function. Score function. The framework provides a flexible scoring framework defined by a general formulation (Eq. 12 ). This framework allows users to combine individual score components p either through a weighted sum or a weighted product. 45 For sequence \(\:x\) , Each score, \(\:{score}_{i}\) , is assigned a strictly positive weight coefficient \(\:w(w\in\:\left(0,\:+\infty\:\right))\) , reflecting its relative contribution to the final score. Individual component scores span the non-negative range [0, +∞): $$\:MPO\left(x\right)=\frac{{\sum\:}_{i}{w}_{i}\times\:{score}_{i}\left(x\right)}{{\sum\:}_{i}{w}_{i}}$$ 12 The formulation is provided for user convenience and flexibility. Generating New Sequence Samples. During RL training, the amino acid sequences were sampled by ProteinMPNN with the conditional probability from RL training, and the coordinates were then generated with the backbone by Rosetta and refined by FastRelax. TNFR1 study case for RL-Augmented ProteinMPNN framework. In this study, we focus here on optimizing S1B2 to enhance its functional activity. the inputs files for this framework are available at GitHub( https://github.com/weilin199204/SuperBinder ), and σ is set to 0.05, while batch_size set to 1, and num_seq_per_target (numbers of sequence at per sample step) set to 20. Other parameters are default. Here MPO score function is as Eq. 13: \(\:p\left(x\right)={pK}_{D}(x\) ) (13) Here, the negative logarithm dissociation constant to base 10 ( \(\:{pK}_{D}\) ) (Eq. 14 ), was predicted by ProAffinity-GNN 39 , a novel approach to structure-based protein–protein binding affinity prediction model, $$\:{pK}_{D}=-\text{log}{K}_{D}$$ 14 Molecule dynamics (MD) and binding free energy calculation with MM/GBSA. Same as our previous work 6 , to evaluate the binding capability of designed minibinders to TNFR1, molecular dynamics (MD) simulations of each minibinder complexed with TNFR1 alone were performed for 300 ns by using the AMBER force field of GROMACS 2019.6 software. Three hundred snapshot structures for each minibinder-TNFR1 complex were extracted from the smooth MD trajectory at equal intervals for binding free energy calculation using the MM/GBSA method. Expression, purification of binders and swine TNFR1. The genes sequences for the minibinders were cloned into a pET-22b vector, incorporating a C-terminal 6×His tag. The resulting plasmids were transformed into Escherichia coli BL21(DE3) competent cells. Protein expression was induced by the addition of IPTG, followed by incubation at 18 ℃ for 20 hours. Cells were harvested via centrifugation at 7,000 rpm for 10 minutes and lysed by sonication. To prevent proteolytic degradation, PMSF was added to a final concentration of 100 mM. The clarified lysate was subjected to purification by Ni-NTA affinity chromatography, which included a wash step with buffer containing 30 mM imidazole and subsequent elution using 300 mM imidazole. Further purification was achieved by size-exclusion chromatography on an FPLC system equipped with a Superdex 75 Increase 10/300 column (Cytiva). Protein concentration was quantified using a BCA assay kit (Biosharp, BL521A), and sample purity was confirmed by SDS-PAGE. The Avi-tag extracellular domain of human TNFR1 was expression, purification some as our previous work 6 . Biolayer interferometry . The binding affinity of the purified minibinders for TNFR1 was quantitatively evaluated by biolayer interferometry (BLI) using an Octet RED96 instrument (ForteBio), following an established procedure 6 . Kinetics constants (K on , K off ) were determined by globally fitting the association and dissociation curves to a 1:1 binding model. The affinity constant (K D ) for each minibinder was calculated following the equation: K D = K off / K on . Data were processed using Octet Analysis Studio v.13.0.1.35. Neutralizing potency measured in L929 cell-based assay. The neutralizing potency of the minibinders was assessed using a TNFα-sensitive murine L929 cell-based assay, as described previously 6 . Authenticated L929 cells were cultured in complete DMEM under standard conditions. For the assay, cells were seeded in 96-well plates (5,000 cells/well) and treated with serially diluted minibinders (0.78–400 nM) for 30 min, followed by co-treatment with TNFα (10 pM) and actinomycin D (1 µg/mL). After 24 h incubation, cell viability was measured using a CCK-8 kit. Dose-response curves were analyzed with GraphPad Prism to determine IC 50 values. Declarations Data and Code Availability Statement The code for reinforcement learning-augmented ProteinMPNN is available at GitHub (https://github.com/weilin199204/SuperBinder). All data generated or analyzed during the current study are included in this article and Supplementary information . Acknowledgments This work was supported by Key Research and Development Progarm of Wuhan (2025020602030105), Hubei Provincial Natural Science Foundation of China (2024AFB465), Major Special Projects for the Development of Agricultural Microbial Industry in Hubei Province (NYWSWZX2025-2027-03, NYWSWZX2025-2027-10), State Key Laboratory Open Project (SKLBEE2022007) from State Key Laboratory of Biocatalysis and Enzyme Engineering. We thank HPC center at State Key Laboratory of Biocatalysis and Enzyme Engineering for cloud computing support. Author information contributions L. W.: formal analysis, funding acquisition, methodology, resources, software, visualization, writing – original draft, writing – review & editing; Z. Wu, Y. H.: formal analysis, methodology, validation, visualization, writing – original draft, writing – review & editing; Y. F.: formal analysis, validation, writing – original draft, writing – review & editing; M. G., B. X.: validation, writing – review & editing; J. W.: funding acquisition, methodology, supervision, writing – original draft, writing – review & editing; S. L.: formal analysis, software, writing – original draft, writing – review & editing; K. M.: funding acquisition, Writing – review & editing; Z. W.: Funding acquisition, Supervision, Writing – review & editing. Ethics declarations Competing interests The authors declare no competing interests. Supplementary information. More computational and experimental methods, details of the results can be found online. References Chevalier, A., et al.: Massively parallel de novo protein design for targeted therapeutics. Nature. 550 , 74–79 (2017) Rocklin, G.J., et al.: Global analysis of protein folding using massively parallel design, synthesis, and testing. Science. 357 , 168–175 (2017) Cao, L., et al.: De novo design of picomolar SARS-CoV-2 miniprotein inhibitors. Science. 370 , 426–431 (2020) Berger, S., et al.: Preclinical proof of principle for orally delivered Th17 antagonist miniproteins. Cell. 187 , 1–13 (2024) Cao, L., et al.: Design of protein-binding proteins from the target structure alone. Nature. 605 , 551–560 (2022) Weng, J., et al.: Design of minibinder proteins specific to TNFR1. Int. J. Biol. Macromol. 293 , 139403 (2025) Wei, L., et al.: De novo design mini-binder proteins targeting the glycoproteins D to inhibit PRV replication in PK15 cells. Int. J. Biol. Macromol., 144403 (2025) Huang, B., et al.: De novo design of miniprotein antagonists of cytokine storm inducers. Nat. Commun. 15 , 7064 (2024) Bennett, N.R., et al.: Improving de novo protein binder design with deep learning. Nat. Commun. 14 , 2625 (2023) Ingraham, J., Garg, V.K., Barzilay, R., Jaakkola, T.: Generative models for graph-based protein design. NeurIPS 2019 (2019) Anand, N., et al.: Protein sequence design with a learned potential. Nat. Commun. 13 , 746 (2022) Jing, B., Eismann, S., Suriana, P., Townshend, R.J.L., Dror, R.: Learning from Protein Structure with Geometric Vector Perceptrons. arXiv:2009.01411 (2020) Strokach, A., Becerra, D., Corbi-Verge, C., Perez-Riba, A., Kim, P.M.: Fast and Flexible Protein Design Using Deep Graph Neural Networks. Cell. Syst. 11 , 402–411e404 (2020) Hsu, C., et al.: Learning inverse folding from millions of predicted structures. bioRxiv:2022.04.10.487779v2 (2022) Zhang, Y., et al.: ProDCoNN: Protein design using a convolutional neural network. Proteins. 88 , 819–829 (2020) Qi, Y., Zhang, J.Z.H., DenseCPD: Improving the Accuracy of Neural-Network-Based Computational Protein Sequence Design with DenseNet. J. Chem. Inf. Model. 60 , 145–1252 (2020) Watson, J.L., et al.: De novo design of protein structure and function with RFdiffusion. Nature. 620 , 1089–1100 (2023) Dauparas, J., et al.: Robust deep learning–based protein sequence design using ProteinMPNN. Science. 378 , 49–56 (2022) Jumper, J., et al.: Highly accurate protein structure prediction with AlphaFold. Nature. 596 , 583–589 (2021) Ho, J., Jain, A., Abbeel, P.: Denoising Diffusion Probabilistic Models. arXiv:2006.11239 (2020) Baek, M., et al.: Accurate prediction of protein structures and interactions using a three-track neural network. Science. 373 , 871–876 (2021) Liu, Z., et al.: De novo designed mini-binders targeting glyceraldehyde-3-phosphate dehydrogenase of Streptococcus equi ssp. zooepidemicus provided partial protection in mice model of infection. Int. J. Biol. Macromol. 307 , 142293 (2025) Ming, K., et al.: Mini-binders targeting Streptococcus equi ssp. zooepidemicus M-like protein inhibit the bacterial adhesion and exert protective effects in vivo. Int. J. Biol. Macromol. 304 , 140803 (2025) Quijano-Rubioa, A., Ulge, U.Y., Walkey, C.D., Silva, D.: The advent of de novo proteins for cancer immunotherapy. Curr. Opin. Chem. Biol. 56 , 119–128 (2020) Sutton, R.S., Barto, A.G.: Reinforcement Learning: An Introduction. The MIT Press (2015) Silver, D., et al.: Mastering the game of Go without human knowledge. Nature. 550 , 354–359 (2017) Guo, D., et al.: DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature. 645 , 633–638 (2025) Olivecrona, M., Blaschke, T., Engkvist, O., Chen, H.: Molecular de-novo design through deep reinforcement learning. J. Cheminform. 9 , 48 (2017) Blaschke, T., et al.: REINVENT 2.0: An AI Tool for De Novo Drug Design. J. Chem. Inf. Model. 60 , 5918–5922 (2020) Loeffler, H.H., et al.: Reinvent 4: Modern AI–driven generative molecule design. J. Cheminform. 16 , 20 (2024) Lutz, I.D., et al.: Top-down design of protein architectures with reinforcement learning. Science. 380 , 266–273 (2023) Popova, M., Isayev, O., Tropsha, A.: Deep reinforcement learning for de novo drug design. Sci. Adv. 4 , eaap7885 (2018) Ivanenkov, Y.A., et al.: Chemistry42: An AI-Driven Platform for Molecular Design and Optimization. J. Chem. Inf. Model. 63 , 695–701 (2023) Jaques, N., Gu, S., Turner, R.E., Eck, D.: Tuning Recurrent Neural Networks with Reinforcement Learning. arXiv:1611.02796v1 (2017) van Loo, G., Bertrand, M.J.M.: Death by TNF: a road to inflammation. Nat. Rev. Immunol. 23 , 289–303 (2022) Leone, G.M., Mangano, K., Petralia, M.C., Nicoletti, F., Fagone, P., Past: Present and (Foreseeable) Future of Biological Anti-TNF Alpha Therapy. J. Clin. Med. 12 , 1630 (2023) Aggarwal, B.B.: Signalling pathways of the TNF superfamily: a double-edged sword. Nat. Rev. Immunol. 3 , 745–756 (2003) Siegmund, D., Wajant, H.: TNF and TNF receptors as therapeutic targets for rheumatic diseases and beyond. Nat. Rev. Rheumatol. 19 , 576–591 (2023) Zhou, Z., et al.: ProAffinity-GNN: A Novel Approach to Structure-Based Protein–Protein Binding Affinity Prediction via a Curated Data Set and Graph Neural Networks. J. Chem. Inf. Model. 64 , 8796–8808 (2024) Liu, Y., Yu, Q., Wang, D., Chen, M.: AF3Score: A Score-Only Adaptation of AlphaFold3 for Biomolecular Structure Evaluation. J. Chem. Inf. Model. 65 , 8207–8214 (2025) Cao, L., et al.: De novo design of picomolar SARS-CoV-2 miniprotein inhibitors. Science. 370 , 426–431 (2020) Glögl, M., et al.: Target-conditioned diffusion generates potent TNFR superfamily antagonists and agonists. Science. 386 , 1154–1161 (2024) J., W.R. Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning. Mach. Learn. 8 , 229–256 (1992) Peters, J.: Policy gradient methods. Scholarpedia. 5 , 3698 (2010) Cummins, D.J., Bell, M.A., Integrating Everything: The Molecule Selection Toolkit, a System for Compound Prioritization in Drug Discovery. J. Med. Chem. 59 , 6999–7010 (2016) Additional Declarations There is NO Competing Interest. Supplementary Files SIZIP.zip Supplementary information SuperbinderSI5CB.docx Reinforcement Learning-Augmented ProteinMPNN Improve the Binding Affinity of TNFR1-Targeting Minibinders Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8645836","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":577374180,"identity":"597d4291-4edb-4544-ad8d-bceae7c14dbe","order_by":0,"name":"Zigong Wei","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAz0lEQVRIiWNgGAWjYFACxgaGBBDN3tj48ANpWngONxtLkGabRHqbAA8xCg2ONzd/ePDLJk8+8mEbgwSDnZxuAyEtZw42GCT2pRUb3k5se1DAkGxsdoCAFrMbiQ0JiT2HEzfOTmw3kGA4kLiNoJb7DxsOgLXMPNgmwUOUlhuMjQ0JPw4nzpdgJFKL/ZnEZobEhrTEDTyJwEA2IMIvku3HH3/88ccmcX778YcPP1TYyRHUAgaMbcCgA6s0IEY5GPxhYJBvIFr1KBgFo2AUjDQAACfbTBADkSFQAAAAAElFTkSuQmCC","orcid":"https://orcid.org/0000-0002-2445-6171","institution":"Hubei University","correspondingAuthor":true,"prefix":"","firstName":"Zigong","middleName":"","lastName":"Wei","suffix":""},{"id":577374181,"identity":"a422924f-c418-4ed2-bde3-2376b2c063a9","order_by":1,"name":"Lin Wei","email":"","orcid":"https://orcid.org/0000-0002-1977-2700","institution":"Hubei Univeristy","correspondingAuthor":false,"prefix":"","firstName":"Lin","middleName":"","lastName":"Wei","suffix":""},{"id":577374182,"identity":"f3e4c778-15d5-4c66-804d-c7657b68d254","order_by":2,"name":"Zhiyong Wu","email":"","orcid":"","institution":"Hubei University","correspondingAuthor":false,"prefix":"","firstName":"Zhiyong","middleName":"","lastName":"Wu","suffix":""},{"id":577374183,"identity":"cda666b8-685a-435c-bfa2-33f3d0aafd22","order_by":3,"name":"Yang Hu","email":"","orcid":"","institution":"Hubei University","correspondingAuthor":false,"prefix":"","firstName":"Yang","middleName":"","lastName":"Hu","suffix":""},{"id":577374184,"identity":"5f0a338e-8373-40f9-befa-34371f46f740","order_by":4,"name":"Yihe Fang","email":"","orcid":"","institution":"Hubei University","correspondingAuthor":false,"prefix":"","firstName":"Yihe","middleName":"","lastName":"Fang","suffix":""},{"id":577374185,"identity":"ea3dc309-1944-4207-9d69-026465e195c9","order_by":5,"name":"Miaomiao Geng","email":"","orcid":"","institution":"Hubei University","correspondingAuthor":false,"prefix":"","firstName":"Miaomiao","middleName":"","lastName":"Geng","suffix":""},{"id":577374186,"identity":"bcd191d5-7ae2-471d-9f47-37468f6b4012","order_by":6,"name":"Banbin Xing","email":"","orcid":"","institution":"Hubei Univeristy","correspondingAuthor":false,"prefix":"","firstName":"Banbin","middleName":"","lastName":"Xing","suffix":""},{"id":577374188,"identity":"66fad8bd-33d2-4618-9038-ab89c99c8720","order_by":7,"name":"Jun Weng","email":"","orcid":"https://orcid.org/0000-0002-8295-3459","institution":"Huazhong Univeristy of Science and Technology","correspondingAuthor":false,"prefix":"","firstName":"Jun","middleName":"","lastName":"Weng","suffix":""},{"id":577374189,"identity":"9d6d9b79-5d22-49de-8a51-4a3bfde7790f","order_by":8,"name":"Song Liu","email":"","orcid":"","institution":"Hubei University","correspondingAuthor":false,"prefix":"","firstName":"Song","middleName":"","lastName":"Liu","suffix":""},{"id":577374190,"identity":"59b1ebc0-b3d1-4a80-bd0b-8ddc3ecdf1cb","order_by":9,"name":"Ke Ming","email":"","orcid":"","institution":"Hubei Univeristy","correspondingAuthor":false,"prefix":"","firstName":"Ke","middleName":"","lastName":"Ming","suffix":""}],"badges":[],"createdAt":"2026-01-20 06:51:43","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8645836/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8645836/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":100784413,"identity":"17dfdb1b-cb20-41ff-88bf-55c2eeaf41f7","added_by":"auto","created_at":"2026-01-21 11:53:19","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1463347,"visible":true,"origin":"","legend":"","description":"","filename":"Superbinder13CB.docx","url":"https://assets-eu.researchsquare.com/files/rs-8645836/v1/576e9417199423e59ee3866f.docx"},{"id":100784503,"identity":"1fd8bc72-bb85-42c0-a5c0-20847ca667ea","added_by":"auto","created_at":"2026-01-21 11:53:29","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":10505,"visible":true,"origin":"","legend":"","description":"","filename":"COMMSBIO260709.json","url":"https://assets-eu.researchsquare.com/files/rs-8645836/v1/ae0a1f6ca90c3ec44758d961.json"},{"id":100784311,"identity":"2cc80e76-71ac-427c-9533-23e98a3f6b5e","added_by":"auto","created_at":"2026-01-21 11:52:59","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":496663,"visible":true,"origin":"","legend":"","description":"","filename":"SuperbinderSI5CB.docx","url":"https://assets-eu.researchsquare.com/files/rs-8645836/v1/8d1d2548e08af3107a894233.docx"},{"id":100784414,"identity":"ab940318-4b66-42b9-bac4-3d8d13b2979f","added_by":"auto","created_at":"2026-01-21 11:53:19","extension":"xml","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":104887,"visible":true,"origin":"","legend":"","description":"","filename":"COMMSBIO2607090enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-8645836/v1/3664c166438909835ded4fac.xml"},{"id":100784302,"identity":"64970104-32f7-473b-ac2a-09d84743d9ae","added_by":"auto","created_at":"2026-01-21 11:52:58","extension":"png","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":78684,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8645836/v1/c79759a4eae8528296057eb2.png"},{"id":100784306,"identity":"bccbd8ad-9986-4f79-baef-459a9e575269","added_by":"auto","created_at":"2026-01-21 11:52:58","extension":"png","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":117932,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8645836/v1/432473ac1318f04c3bf1ff16.png"},{"id":100784696,"identity":"634f46dd-6cce-4b5a-9fb0-9dbf412ca2d0","added_by":"auto","created_at":"2026-01-21 11:54:03","extension":"png","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":171742,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8645836/v1/e663fe3302e4aea5273cdfae.png"},{"id":100784402,"identity":"dc3060c6-ab28-435f-82f5-528703ada699","added_by":"auto","created_at":"2026-01-21 11:53:16","extension":"xml","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":103635,"visible":true,"origin":"","legend":"","description":"","filename":"COMMSBIO2607090structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-8645836/v1/6d0642a4bd569facdc8f1c81.xml"},{"id":100784437,"identity":"6c753061-973f-4921-9b64-38e52bf46e1e","added_by":"auto","created_at":"2026-01-21 11:53:25","extension":"html","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":116183,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-8645836/v1/eb41e4ad199cdf3a6afe9aa7.html"},{"id":100784593,"identity":"9477fdca-5092-42a1-b707-085251a7ca88","added_by":"auto","created_at":"2026-01-21 11:53:43","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":422052,"visible":true,"origin":"","legend":"\u003cp\u003eThe framework for reinforcement learning-augmented ProteinMPNN.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8645836/v1/ea17a3dd6264b35f7e034cdb.png"},{"id":100784257,"identity":"0c9408b4-b586-4595-be08-b84a49472e1d","added_by":"auto","created_at":"2026-01-21 11:52:51","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":381915,"visible":true,"origin":"","legend":"\u003cp\u003e(A) The distribution for pK\u003csub\u003eD\u003c/sub\u003e predicted by ProAffinity-GNN for the samples from each “Agent” models. (B) The distribution of the top 100 pK\u003csub\u003eD\u003c/sub\u003e samples for each “Agent” models. (C) Number of candidates with pK\u003csub\u003eD\u003c/sub\u003e \u0026gt;7 for each \"Agent\" model\u003cstrong\u003e.\u003c/strong\u003e (D-F) The distribution for pLDDT of minibinder Cα atom, pTM and pAE predicted by AF3Score for the samples from each “Agent” model. (G-J) The distribution for pLDDT of complex Cα atom, pTM, pAE and ipTM predicted by AF3Score for the samples from each “Agent”. (K) The distribution for pAE_interaction predicted by AF3Score for the samples from each “Agent” model.\u003c/p\u003e","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8645836/v1/458f550ca12891dfcf296e77.jpeg"},{"id":100784742,"identity":"a4bc30ec-afb6-4217-94d1-d5bb9b5ba90a","added_by":"auto","created_at":"2026-01-21 11:54:14","extension":"jpeg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":535161,"visible":true,"origin":"","legend":"\u003cp\u003e(A-E) The AF2-predicted binding modes for \u003cstrong\u003eOPT1\u003c/strong\u003e, \u003cstrong\u003eOPT2\u003c/strong\u003e,\u003cstrong\u003e OPT4\u003c/strong\u003e, \u003cstrong\u003eOPT5\u003c/strong\u003e, \u003cstrong\u003eOPT7\u003c/strong\u003e. (F-J) The binding affinity of \u003cstrong\u003eOPT1\u003c/strong\u003e, \u003cstrong\u003eOPT2\u003c/strong\u003e, \u003cstrong\u003eOPT4\u003c/strong\u003e, \u003cstrong\u003eOPT5\u003c/strong\u003e, \u003cstrong\u003eOPT7\u003c/strong\u003e to swine TNFR1 was determined using biolayer interferometry. K\u003csub\u003eon\u003c/sub\u003e and K\u003csub\u003eoff\u003c/sub\u003e were determined by globally fitting the association and dissociation curves respectively and K\u003csub\u003eD\u003c/sub\u003e for each minibinder was calculated following the equation: K\u003csub\u003eD\u003c/sub\u003e = K\u003csub\u003eoff \u003c/sub\u003e/ K\u003csub\u003eon\u003c/sub\u003e. (K) The blocking effects on TNFα-mediated cell death of L929 cell line for \u003cstrong\u003eOPT1\u003c/strong\u003e, \u003cstrong\u003eOPT2\u003c/strong\u003e, \u003cstrong\u003eOPT4\u003c/strong\u003e, \u003cstrong\u003eOPT5\u003c/strong\u003e, \u003cstrong\u003eOPT7\u003c/strong\u003e and\u003cstrong\u003e S1B2\u003c/strong\u003e. (L-M) The heatmaps of binding affinities of each pair of residues in the binding surface of \u003cstrong\u003eOPT1 \u003c/strong\u003eand \u003cstrong\u003eOPT7\u003c/strong\u003e.\u003c/p\u003e","description":"","filename":"floatimage3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8645836/v1/7027a77a6c75e3117b6bef4b.jpeg"},{"id":101753481,"identity":"9181fa09-cbfe-474f-9276-88ea2ca3d948","added_by":"auto","created_at":"2026-02-03 10:40:09","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2168982,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8645836/v1/390653cb-2c3a-455a-af55-1a4f010fbec5.pdf"},{"id":100784720,"identity":"be87e19f-51e8-4fc4-86b2-d2d15be56367","added_by":"auto","created_at":"2026-01-21 11:54:06","extension":"zip","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":11725979,"visible":true,"origin":"","legend":"Supplementary information","description":"","filename":"SIZIP.zip","url":"https://assets-eu.researchsquare.com/files/rs-8645836/v1/3f847e25b20c2ed1157164d8.zip"},{"id":100784590,"identity":"ef454227-2605-4fea-9370-91fcc7ce9d80","added_by":"auto","created_at":"2026-01-21 11:53:42","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":496663,"visible":true,"origin":"","legend":"Reinforcement Learning-Augmented ProteinMPNN Improve the Binding Affinity of TNFR1-Targeting Minibinders","description":"","filename":"SuperbinderSI5CB.docx","url":"https://assets-eu.researchsquare.com/files/rs-8645836/v1/3ec0866747171e1fd6a4d7b1.docx"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"Reinforcement Learning-Augmented ProteinMPNN Improve the Binding Affinity of TNFR1-Targeting Minibinders","fulltext":[{"header":"Introduction","content":"\u003cp\u003eMinibinders, as conceptualized by Baker and colleagues\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e, are compact protein molecules (typically 4\u0026ndash;12 kDa) composed of natural amino acids that exhibit specific target-binding capabilities. To enable data-driven \u003cem\u003ede novo\u003c/em\u003e minibinder design, Baker et al. constructed a set of small proteins comprising 20\u0026ndash;50 amino acid peptides, which fold into stable conformations and demonstrate resistance to chymotrypsin and trypsin\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e. Leveraging this database, minibinders targeting to influenza virus hemagglutinin A (HA) and botulinum neurotoxin B were successfully designed\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e. These molecules exhibit stable folded conformations, protease resistance, low immunogenicity, and enhanced tissue permeability compared to antibodies\u003csup\u003e\u003cspan additionalcitationids=\"CR2\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. Berger et al.\u003csup\u003e4\u003c/sup\u003e further validated oral delivery efficacy of an IL-23R-targeting minibinder in murine colitis models, achieving therapeutic outcomes comparable to injected antibodies.\u003c/p\u003e \u003cp\u003eDesigning conformationally stable minibinders remains challenging. Baker et al., pioneered computational \u003cem\u003ede novo\u003c/em\u003e minibinders design by Rosetta, a physics-based approach, which treats sequence design as an energy optimization problem for a given input structure. Based on Rosetta, minibinders targeting SARS-CoV-2 spike protein\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e, IL-7Rα, VirB8, TrkA\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e, TNFR1\u003csup\u003e6\u003c/sup\u003e, PRV gD\u003csup\u003e7\u003c/sup\u003e, IL-6R, GP130, and IL-1R1 receptor subunits\u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e has been successfully designed. However, due to incomplete sampling of conformational space with physics-based Rosetta, experimental success rates were low, often requiring the screening of thousands of designs\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e. Moreover, reliance on pre-defined protein scaffolds inherently limits the diversity and shape complementarity of potential solutions\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eAI-driven strategies can overcome these limitations by leveraging large protein datasets. Deep-learning approaches can efficiently generate sequences for monomeric backbones without rotamer enumeration\u003csup\u003e\u003cspan additionalcitationids=\"CR11 CR12 CR13 CR14 CR15 CR16\" citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e. Crucially, ProteinMPNN\u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e enhances the sequences folding fidelity of the backbones designed by Rosetta, as predicted by AlphaFold2(AF2)\u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e, increasing the success rate from 2.7% to 54.1% and improving utility for functional sites\u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e. The autoregressive MPNN utilizes backbone geometry (distances, orientations, dihedrals) to achieve 50.5% native sequence recovery, outperforming predecessors (52.4% vs. 32.9%).\u003csup\u003e18\u003c/sup\u003e A backbones design strategy, RFdiffusion\u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e, employs denoising diffusion probabilistic models (DDPMs)\u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e with RoseTTAFold (RF)\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e frame representations (Cα coordinates\u0026thinsp;+\u0026thinsp;N-Cα-C orientations) to create diverse structures from random initialization.\u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e Baker et al. accurately reproduced the crystal structure of the binder LCB1 targeting the SARS-CoV-2 receptor binding domain (RBD), originally reported by the lab in 2020\u003csup\u003e18\u003c/sup\u003e, using RFDiffusion (skeleton design) combined with ProteinMPNN (sequence design), demonstrating the reliability of this AI approach\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e. Similarly, in our previous work\u003csup\u003e\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e, \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u003c/sup\u003e, with the same protocol, minibinders targeting M-like protein (SzM), Glyceraldehyde-3-phosphate dehydrogenase (GAPDH) of Streptococcus equi ssp. Zooepidemicus (SeZ) has been successfully designed, improved the survival rate (80%, 60%, respectively) of mice infected with a 100-fold lethal dose of SEZ and significantly reduced the organ bacterial load.\u003c/p\u003e \u003cp\u003eWhile ProteinMPNN significantly improves sequence recovery rates compared to Rosetta (~\u0026thinsp;52% vs. ~33%)\u003csup\u003e18\u003c/sup\u003e, its primary objective remains structural fidelity rather than optimizing specific functional properties like binding affinity. Consequently, the \u003cem\u003ede novo\u003c/em\u003e design of high-affinity minibinders still largely relies on extensive wet-lab screening of numerous candidates. Therefore, beyond ensuring high sequence design success, directly designing high-affinity minibinders \u003cem\u003ein silico\u003c/em\u003e represents a critical challenge for numerous therapeutic applications\u003csup\u003e\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eReinforcement learning(RL)\u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e is a machine learning paradigm well-suited to address such optimization problems. An RL agent learns to make decisions by interacting with an environment to maximize cumulative rewards. RL-based models like AlphaGo\u003csup\u003e\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/sup\u003e and DeepSeek-R1\u003csup\u003e27\u003c/sup\u003e have achieved significant success. RL concepts are widely applied to drug design\u003csup\u003e\u003cspan additionalcitationids=\"CR29\" citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e and protein design\u003csup\u003e\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e, often directing generative models by maximizing a multi-parameter optimization (MPO) scoring function reflecting key features.\u003csup\u003e\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e, \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e, \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e\u003c/sup\u003e REINVENT\u003csup\u003e\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e, REINVENT2\u003csup\u003e29\u003c/sup\u003e, Reinvent4\u003csup\u003e30\u003c/sup\u003e tackle the inverse design problem through reinforcement learning using RNNs or transformers as deep learning architectures to generate molecules with desirable properties by learning an augmented episodic likelihood combining prior likelihood\u003csup\u003e\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e\u003c/sup\u003e and a user-defined scoring function. Baker et al. reported a \"top-down\" RL approach using Monte Carlo tree search to sample conformers within an architecture under functional constraints\u003csup\u003e\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e, successfully designing disk-shaped nanopores and ultracompact icosahedra verified by cryo-EM.\u003c/p\u003e \u003cp\u003eIn this work, we introduce a computational platform that integrates a RL framework with ProteinMPNN network to directly generate minibinder sequences for high-affinity binding. We demonstrate the efficacy of the framework by optimizing a previously reported TNF receptor type 1 (TNFR1)-targeting minibinder, \u003cb\u003eS1B2\u003c/b\u003e\u003csup\u003e6\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eTumor necrosis factor alpha (TNFα) is a pleiotropic cytokine with dual roles: it provides beneficial defense against infections and cancer, yet also exerts deleterious pro-inflammatory and cytotoxic effects during inflammation\u003csup\u003e\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e, \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e\u003c/sup\u003e. TNFα signals via two distinct cell surface receptors, TNFR1 and type 2 (TNFR2)\u003csup\u003e\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e\u003c/sup\u003e. Broadly, TNFR1 primarily mediates TNFα\u0026rsquo;s inflammatory and apoptotic effects.\u003csup\u003e\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e\u003c/sup\u003e Targeting TNFR1 specifically disrupts the TNFα-TNFR1 interaction while preserving TNFα's beneficial homeostatic immune functions. In the previous work\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e, binder \u003cb\u003eS1B2\u003c/b\u003e exhibited strong binding activity (K\u003csub\u003eD\u003c/sub\u003e = 890 nM). Therefore, it is necessary to optimize the small minibinder like \u003cb\u003eS1B2\u003c/b\u003e, to enhance its activity. In this work, the generated minibinders by the RL-augmented ProteinMPNN framework shown \u0026ldquo;Prefect\u0026rdquo; (minibinders with predicted negative logarithm dissociation constant to base 10, pK\u003csub\u003eD\u003c/sub\u003e \u0026gt;7) rate increasing with 28%. And \u003cb\u003eOPT7\u003c/b\u003e were designed from \u003cb\u003eS1B2\u003c/b\u003e resulting in up to 7-fold enhanced activity.\u003c/p\u003e"},{"header":"Results and Discussion","content":"\u003cp\u003e \u003cem\u003eReinforcement learning framework.\u003c/em\u003e \u003c/p\u003e \u003cp\u003eThe RL-augmented framework with ProteinMPNN network presents an integrated computational pipeline for designing target-binding minibinders (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). The process begins with generated backbone, which was created to bind specific target sites using tools like RFdiffusion or Rosetta.\u003c/p\u003e \u003cp\u003eThe framework is based on policy-based reinforcement learning (RL) framework. Within this framework, an \u0026ldquo;Agent\u0026rdquo; acts as the sequence designer. A pre-trained general ProteinMPNN model\u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e serves as a \u0026ldquo;Prior\u0026rdquo; policy, which initially shares its network weights with the \u0026ldquo;Agent\u0026rdquo;. During training, the \u0026ldquo;Agent\u0026rdquo; generates protein sequences through ProteinMPNN with an augmented training model. The environment is represented by a multi-parameter objective (MPO) scoring function, which evaluates these sequences and provides a reward.\u003c/p\u003e \u003cp\u003eA key step involves forming an augmented training target by combining the scaled MPO reward with the evaluation of \u0026ldquo;Prior\u0026rdquo; sequence likelihood (negative log-likelihood, NLL). The \u0026ldquo;Agent\u0026rdquo; training loss is then defined as the squared difference between its own predicted likelihood and this augmented target. This objective guides the \u0026ldquo;Agent\u0026rdquo; to learn a policy that generates sequences with improved functional scores while remaining grounded in plausible sequence-structure relationships.\u003c/p\u003e \u003cp\u003eThroughout the RL process, only the Agent's parameters are updated via this loss, iteratively refining its policy. Finally, optimized sequences are sampled using the trained Agent model and are structurally validated with tools like AF2-multimer\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e. More details see \u003cb\u003eMaterials and Methods.\u003c/b\u003e\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eDesigned TNFR1 minibinders by RL-augmented ProteinMPNN.\u003c/em\u003e \u003c/p\u003e \u003cp\u003eIn RL training, The NLL of the augmented likelihood and \u0026ldquo;Agent\u0026rdquo; is shown in \u003cb\u003eFig. \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u003c/b\u003e, the NLL of the augmented likelihood stabilized at 2.25, while the NLL of \u0026ldquo;Agent\u0026rdquo; increased from 1.9 to 2.0 in the first 800 steps and then showed a slight downward trend. Following RL training, 1000 minibinders were generated from \u0026ldquo;Agents\u0026rdquo; model at every 100 steps. The negative logarithm dissociation constant to base 10 (pK\u003csub\u003eD\u003c/sub\u003e) of all candidates were predicted by ProAffinity-GNN\u003csup\u003e\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e\u003c/sup\u003e, whose applicability for predicting the binding affinity of \u003cem\u003ede novo\u003c/em\u003e designed minibinders were validated here with the data from Berger\u0026rsquo;s work\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e(see \u003cb\u003eSupplementary information Part II\u003c/b\u003e). Although Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA shows that the binding affinities of the samples from the 'Agent' models are primarily distributed between 5 and 6.5, similar to the samples from original ProteinMPNN network model (\u0026ldquo;Prior\u0026rdquo;), a distinct \u0026ldquo;peak\u0026rdquo; emerged around pK\u003csub\u003eD\u003c/sub\u003e =7 in the \u0026ldquo;Agents\u0026rdquo; distributions, indicating an increased probability of generating high-affinity candidates. This \u0026ldquo;peak\u0026rdquo; intensified up to 800 training steps, after which it gradually diminished. When examining the pK\u003csub\u003eD\u003c/sub\u003e values of candidates sampled by each \u0026ldquo;Agent\u0026rdquo; model, we found that the top 100 candidates and numbers of candidates with pK\u003csub\u003eD\u003c/sub\u003e \u0026gt;7 showed similar trends. The average pKa raised from 7.35 in the \u0026ldquo;Prior\u0026rdquo; to 7.40 at 800 steps with no obvious trending, then decreased to 7.28 at 1000 steps (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB), while numbers of candidates with pK\u003csub\u003eD\u003c/sub\u003e \u0026gt;7 (here we treat that as \u0026ldquo;Prefect\u0026rdquo; minibinders) rose from 110 in the \u0026ldquo;Prior\u0026rdquo; to 141 at 800 steps, showing a 28% increase in the \"Perfect\" rate, although it then decreasing to 96 at 1000 steps (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eC). The above data indicate that during the first 800 training steps, the prediction activity of the sampling from \u0026ldquo;Agent\u0026rdquo; model gradually increased, reflecting the \u0026ldquo;Agent\u0026rdquo; model in RL process learned to explore sequences with higher binding affinities. This confirms the power of the design of high-activity minibinders was enhanced through RL beginning with original ProteinMPNN model.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eSince we are generating protein sequences through reinforcement learning-augmented ProteinMPNN rather than the backbone, and no structural properties were used as a reward function, naturally, the structural rationality is not expected to undergo significant changes. This is supported by AF3Score\u003csup\u003e\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u003c/sup\u003e analysis of the Agent's sampling results. For minibinder Cα atom predicted local distance difference test (pLDDT), predicted TM score (pTM) and predicted aligned error (pAE), although a slight worsening tendency is observed, the distribution ranges remain consistent, suggesting that the changes stay within a controllable range (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eD-F). There was no significant difference in Cα atom pLDDT, pAE, pTM or interface predicted TM score(iPTM) for the complexes between \u0026ldquo;Prior\u0026rdquo; and \u0026ldquo;Agents\u0026rdquo; (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eG-J). More importantly, the distribution of the binder-target interface metric, pAE_interaction, shows no significant difference (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eK). All of this indicates that, in this RL process, the structural rationality was not significantly compromised.\u003c/p\u003e \u003cp\u003eTo further evaluate the activity of the candidate structures sampling from the \u0026ldquo;Agent\u0026rdquo; models, we performed molecular dynamics simulation (MD) optimization on all candidates from the above \u0026ldquo;Agents\u0026rdquo; after the condition-based filtering (pAE_interaction\u0026thinsp;\u0026lt;\u0026thinsp;10), and calculated the binding affinity using MM/GBSA. For the structures with predicted binding affinity higher than \u003cb\u003eS1B2\u003c/b\u003e, we conducted manual inspection to select minibinders with a higher likelihood of binding. Subsequently, we selected 7 sequences for expression and purification. The obtained \u003cb\u003eOPT1-OPT7\u003c/b\u003e reduced the predicted binding free energy (ΔG) by 36.1\u0026ndash;1.93 kcal/mol (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e), indicating extremely high potential activity and demonstrating the feasibility and advancement of the RL-Augmented ProteinMPNN framework.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eThe top binding affinities predicted by MM/GBSA for RL-augmented ProteinMPNN designed mini-binder targeting TNFR1.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\"\u0026plusmn;\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eID\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSequence\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eWeight\u003c/p\u003e \u003cp\u003e(kDa)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eΔG(MM/GBSA) (kcal/mol)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eS1B2\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDAEEQLRIQERLIELAFRLGDPETAERIARAAEILLKEDGDPEAVEIIRELLERLR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e6.51\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e-53.29\u0026thinsp;\u0026plusmn;\u0026thinsp;6.52\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eOPT1\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSLEENLKFALNEIKRAFRQGKPEEARRTAEAAISWFKEQGDEESVKEIEELLKKLE\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e6.56\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e-89.36\u0026thinsp;\u0026plusmn;\u0026thinsp;6.46\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eOPT2\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSMKENLELALTQIRLAVRRGEPEEAKRIADAAVSVFREFGDEESVKKIEELLKELL\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e6.40\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e-65.30\u0026thinsp;\u0026plusmn;\u0026thinsp;6.19\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eOPT3\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSLKENLKFTLNQIRLAFRRGNPEEAKAIAEAAISVFKEFGDEESVEELKKLLEELL\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e6.41\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e-65.18\u0026thinsp;\u0026plusmn;\u0026thinsp;3.71\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eOPT4\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSSKENTEFAINQIKLAFRQGNPEEARQIAKAAVSVLTEMGDPEGVERVKAVLEELE\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e6.17\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e-62.13\u0026thinsp;\u0026plusmn;\u0026thinsp;4.20\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eOPT5\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSSKKNLEFALNQIRLAFRQGKPEEAKRIAKAAVKVLLEQGDEEGAKKVEELLKELE\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e6.32\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e-56.70\u0026thinsp;\u0026plusmn;\u0026thinsp;5.78\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eOPT6\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMAKENKEFALNQIRLAVSRGEPEEAEKIARASVKVLKEMGDPESAEEVEKLLKELL\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e6.30\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e-56.67\u0026thinsp;\u0026plusmn;\u0026thinsp;3.85\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eOPT7\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSAEENLEFALTQIRLAFRQGKPEEAKRIAEAAIKVLSEMGYPEGVEKIKKLLEELE\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e6.35\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e-55.22\u0026thinsp;\u0026plusmn;\u0026thinsp;3.46\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eExpression, purification for the 7 minibinders.\u003c/em\u003e \u003c/p\u003e \u003cp\u003eEach minibinder was expressed in \u003cem\u003eEscherichia coli\u003c/em\u003e BL21(DE3) using a bacterial expression system, taking advantage of the well-established suitability of prokaryotic hosts for producing small, stable, single-domain protein\u003csup\u003e\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e\u003c/sup\u003e. Except for \u003cb\u003eOPT3\u003c/b\u003e, all others were successfully expressed. Following expression, the minibinders were purified by immobilized metal affinity chromatography (IMAC, Ni-NTA) followed by size-exclusion chromatography (FPLC). SDS-PAGE analysis confirmed successful expression with high purity, and molecular weights between 6\u0026ndash;7 kDa, (\u003cb\u003eFig. S2\u003c/b\u003e), consistent with theoretical calculations (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eBLI and cell assays of designd minibinders\u003c/h2\u003e \u003cp\u003eThe affinities of purified minibinders to human TNFR1 were assessed by Biolayer Interferometry (BLI). Here, \u003cb\u003eS1B2\u003c/b\u003e showed an affinity with K\u003csub\u003eD\u003c/sub\u003e = 1.47 \u0026micro;M (\u003cb\u003eFig. S3\u003c/b\u003e), consistent with previous work (K\u003csub\u003eD\u003c/sub\u003e = 0.88 \u0026micro;M\u003csup\u003e6\u003c/sup\u003e). After optimization, the binding affinities (K\u003csub\u003eD\u003c/sub\u003e) of \u003cb\u003eOPT1\u003c/b\u003e, \u003cb\u003eOPT2\u003c/b\u003e, \u003cb\u003eOPT4\u003c/b\u003e, \u003cb\u003eOPT5\u003c/b\u003e, and \u003cb\u003eOPT7\u003c/b\u003e (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA-E) were determined to be 0.48 \u0026micro;M, 1.59 \u0026micro;M, 20.30 \u0026micro;M, 54.88 \u0026micro;M, and 0.23 \u0026micro;M, respectively. (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eF-J) However, \u003cb\u003eOPT6\u003c/b\u003e did not show high binding affinity (\u003cb\u003eFig. S4\u003c/b\u003e). Compared to \u003cb\u003eS1B2\u003c/b\u003e, \u003cb\u003eOPT1\u003c/b\u003e and \u003cb\u003eOPT7\u003c/b\u003e increased 3-fold and 7-fold, respectively, while \u003cb\u003eOPT2\u003c/b\u003e maintained its activity.\u003c/p\u003e \u003cp\u003eThe mouse fibroblast L929 cell line is sensitive to TNFα when treated with actinomycin D. Therefore, as in our previous work\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e, we applied an L929 cell assay to assess the blocking effect of antagonist molecules on TNFR1. \u003cb\u003eS1B2\u003c/b\u003e exhibited dose-dependent TNFα-blocking activity with IC\u003csub\u003e50\u003c/sub\u003e of 7.75 nM (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eK), consistent with it in our previous work (4.32 nM\u003csup\u003e6\u003c/sup\u003e). \u003cb\u003eOPT1\u003c/b\u003e, \u003cb\u003eOPT2\u003c/b\u003e, \u003cb\u003eOPT4\u003c/b\u003e, \u003cb\u003eOPT5\u003c/b\u003e, and \u003cb\u003eOPT7\u003c/b\u003e presented higher potency to TNFR1 in L929 cell assay, The IC\u003csub\u003e50\u003c/sub\u003es were 1.25 nM, 12.92 nM, 22.39 nM, 13.04 nM and 1.94 nM, respectively (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eK). Compared to \u003cb\u003eS1B2\u003c/b\u003e, the neutralizing activity of \u003cb\u003eOPT1\u003c/b\u003e and \u003cb\u003eOPT7\u003c/b\u003e increased by 6-fold and 4-fold, respectively, while the IC\u003csub\u003e50\u003c/sub\u003e of \u003cb\u003eOPT2\u003c/b\u003e is comparable. The consistency between K\u003csub\u003eD\u003c/sub\u003e and IC\u003csub\u003e50\u003c/sub\u003e suggests that these binders indeed target TNFR1.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eAnalysis binding energy of the binder-TNFR1 complexes.\u003c/em\u003e \u003c/p\u003e \u003cp\u003eWe analyzed the energy contribution of each residue predicted by MM/GBSA. The results showed that for \u003cb\u003eOPT1\u003c/b\u003e, the energy contribution between each residue range from \u0026minus;\u0026thinsp;13.30 to 0.54 kcal/mol, with 75% being positive. The residues R15, E38, and R18 made prominent contributions through strong electrostatic interactions with E149, R77, and E149 of TNFR1, respectively. Residues N5, E12, and E24 of \u003cb\u003eOPT1\u003c/b\u003e also contributed significantly to binding affinity through hydrogen bond interactions with W107, Y103, and Q102 of TNFR1, respectively (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eL). For \u003cb\u003eOPT7\u003c/b\u003e, the energy contribution between each residue ranged from \u0026minus;\u0026thinsp;13.16 to 0.21 kcal/mol, shown nearly no negative contribution. The strong electrostatic interaction between R18 and E149 (TNFR1) had a prominent contribution to affinity, while the electrostatic interactions between E38, K34 and R77, E79 of TNFR1, as well as interactions involving N5, Q12, and E24 with W107, Y103, N116 (TNFR1), also played important roles (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eM). In contrast, in \u003cb\u003eS1B2\u003c/b\u003e, residues at positions 12, 15, and 34, which are LEU, LEU, and ILE respectively, cannot form electrostatic or hydrogen bond interactions. We speculate that the probability of incorporating residues at these positions capable of forming favorable interactions was increased.\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eIn this work, by employing affinity prediction as the reward function in a reinforcement learning framework, we successfully introduced multiple key interactions onto the \u003cb\u003eS1B2\u003c/b\u003e backbone (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eL-M). The minibinder \u003cb\u003eOPT7\u003c/b\u003e, designed with the RL-augmented ProteinMPNN model based on \u003cb\u003eS1B2\u003c/b\u003e backbone, achieved a 7-fold increase in activity without compromising structural rationality. While this study was primarily intended as a methodological proof‑of‑concept rather than the discovery of a clinical candidate, and the achieved affinity does not yet match that of some reported high-affinity binders(K\u003csub\u003eD\u003c/sub\u003e =4 pM in Baker\u0026rsquo;s work\u003csup\u003e\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e\u003c/sup\u003e), but the immunogenicity risk and challenges associated with oral delivery for the ~\u0026thinsp;60 amino acid minibinders like \u003cb\u003eOPT7\u003c/b\u003e have been evaluated. Thus, our work successfully demonstrates that RL can effectively enhance the functional performance of the previous pre-trained general \u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e ProteinMPNN model, steering protein design towards desired functional properties.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eIn this work, we successfully develop an innovative AI framework that integrates Reinforcement Learning (RL) with the protein sequence design network ProteinMPNN to advance the \u003cem\u003ede novo\u003c/em\u003e design of therapeutic minibinders. excel at generating sequences that fold into stable structures, they are not inherently optimized for functional properties like binding affinity. Our Reinforcement Learning-Augmented ProteinMPNN framework addresses this gap by employing an RL agent to steer the sequence generation process, directly optimizing for higher affinity binding against a target of interest.\u003c/p\u003e \u003cp\u003eUsing TNFR1 as a proof-of-concept target, we demonstrated that the framework efficiently generates minibinders with superior predicted binding affinity. Experimental validation confirmed that the designed candidates exhibit significantly improved binding affinity and neutralizing potency in cellular assays. Structural and energetic analyses further elucidated the origin of these gains, linking them to specific, optimized interactions at the protein\u0026ndash;protein interface.\u003c/p\u003e \u003cp\u003eThis work highlights a pivotal step in AI-driven protein design, shifting the focus from mere structural recapitulation toward the direct, goal-oriented optimization of biological function. Extending this approach, the Reinforcement Learning-Augmented ProteinMPNN framework is generalizable, offering a powerful and efficient strategy, can be trained with alternative reward functions targeting other traits, such as thermal stability and solubility, to directly generate proteins with the desired properties. This strategy paves the way for the future development of improved, fit‑for‑purpose minibinders.\u003c/p\u003e"},{"header":"Materials and Methods","content":"\u003cp\u003e \u003cem\u003eBackbone Generation.\u003c/em\u003e \u003c/p\u003e \u003cp\u003eBefore run RL training, we must design a backbone. RFdiffusion\u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e was employed to directly generate target-bound minibinder backbone structures. For therapeutic goals such as blocking protein\u0026ndash;protein interaction, binding specific target sites (\"interface hotspots\"). Rosetta can also use to generated backbone, as our previous work.\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003e \u003cem\u003eCost Function for Reinforcement Learning (RL).\u003c/em\u003e \u003c/p\u003e \u003cp\u003eFigure\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e illustrates the components of the reinforcement learning (RL) within RL-Augmented ProteinMPNN framework. Here, we frame the problem of designing a protein sequence with specified properties as a Reinforcement Learning (RL) task. In policy-based\u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e, \u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e\u003c/sup\u003e reinforcement learning (RL), an agent interacts with an environment by selecting actions based on its current state. Here, the state space is denoted as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:S\\)\u003c/span\u003e\u003c/span\u003e, and for a given state \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:s\\in\\:S\\)\u003c/span\u003e\u003c/span\u003e, the set of available actions is A(s). The agent's behavior is defined by a stochastic policy \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\pi\\:\\left(a\\right|s)\\)\u003c/span\u003e\u003c/span\u003e, which outputs a probability distribution over actions for each state. When the agent generates a sequence of actions \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:A=({a}_{1},...,{a}_{i},...,{a}_{T})\\)\u003c/span\u003e\u003c/span\u003e given corresponding states \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:({s}_{1},...,{s}_{T})\\)\u003c/span\u003e\u003c/span\u003e, the probability of this trajectory under the policy is:\u003cdiv id=\"Equ1\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ1\" name=\"EquationSource\"\u003e\n$$\\:P\\left(A\\right)={\\prod\\:}_{t=1}^{T}\\pi\\:\\left({a}_{t}\\right|{s}_{t})$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e1\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eEach state-action pair yields an immediate reward \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:r\\left(a\\right|s)\\)\u003c/span\u003e\u003c/span\u003e, and the long-term return from time \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:t\\:\\)\u003c/span\u003e\u003c/span\u003eis the cumulative reward \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:G({a}_{t},{s}_{t})={\\sum\\:}_{t}^{T}{r}_{t}\\)\u003c/span\u003e\u003c/span\u003e. The goal of RL is to improve the policy to maximize the expected return over trajectories.\u003c/p\u003e \u003cp\u003eIn the standard algorithm, REINFORCE\u003csup\u003e\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e\u003c/sup\u003e. with a sampled trajectory, the policy parameters \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\theta\\:\\)\u003c/span\u003e\u003c/span\u003e are updated via:\u003c/p\u003e \u003cp\u003e \u003cspan class=\"InlineEquation\"\u003e \u003cspan class=\"mathinline\"\u003e\\(\\:{\\theta\\:}_{t+1}={\\theta\\:}_{t}-\\alpha\\:{\\nabla\\:}_{\\theta\\:}\\sum\\:_{t=0}^{T}\\text{log}{\\pi\\:}_{\\theta\\:}\\left({a}_{t}\\right|{s}_{t}\\left)\\right(G\\left({a}_{t}|{s}_{t}\\right)-b)\\)\u003c/span\u003e \u003c/span\u003e (2)\u003csup\u003e44\u003c/sup\u003e\u003c/p\u003e \u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\alpha\\:\\)\u003c/span\u003e\u003c/span\u003e is the step size and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:b\\)\u003c/span\u003e\u003c/span\u003e is a baseline reward (often set to zero). Setting \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:b=0\\)\u003c/span\u003e\u003c/span\u003e simplifies the cost function to maximizing:\u003c/p\u003e \u003cp\u003e \u003cspan class=\"InlineEquation\"\u003e \u003cspan class=\"mathinline\"\u003e\\(\\:J\\left(\\theta\\:\\right)=\\sum\\:_{t=0}^{T}\\text{log}{\\pi\\:}_{\\theta\\:}\\left({a}_{t}\\right|{s}_{t})G\\left({a}_{t}|{s}_{t}\\right)\\)\u003c/span\u003e \u003c/span\u003e (3)\u003csup\u003e44\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eIf we consider a sparse reward setting, where reward is zero at all steps except the final one, and the final reward equals the total return \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:G\\left(A\\right)\\)\u003c/span\u003e\u003c/span\u003e, the objective reduces to:\u003cdiv id=\"Equ2\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ2\" name=\"EquationSource\"\u003e\n$$\\:J\\left(\\theta\\:\\right)=\\sum\\:_{t=0}^{T}\\text{log}{\\pi\\:}_{\\theta\\:}\\left({a}_{t}|{s}_{t}\\right)G\\left({a}_{t}|{s}_{t}\\right)=G\\left(A\\right)\\sum\\:_{t=0}^{T}\\text{log}{\\pi\\:}_{\\theta\\:}\\left({a}_{t}|{s}_{t}\\right)=G\\left(A\\right)\\text{l}\\text{o}\\text{g}\\sum\\:_{t=0}^{T}{\\pi\\:}_{\\theta\\:}\\left({a}_{t}|{s}_{t}\\right)=G\\left(A\\right)\\text{log}P\\left(A\\right)$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e4\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eHere, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\text{log}P\\left(A\\right)\\)\u003c/span\u003e\u003c/span\u003e represents the log-likelihood of the trajectory under the agent\u0026rsquo;s policy.\u003c/p\u003e \u003cp\u003eWe can reinterpret this objective by introducing an augmented likelihood, which combines a prior model\u0026rsquo;s assessment of the sequence with a measure of its desirability. Let \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\text{log}{P\\left(A\\right)}_{Prior}\\)\u003c/span\u003e\u003c/span\u003e be the log-likelihood under a fixed prior policy, we then define an augmented log-likelihood as:\u003c/p\u003e \u003cp\u003e \u003cspan class=\"InlineEquation\"\u003e \u003cspan class=\"mathinline\"\u003e\\(\\:\\text{log}{P\\left(A\\right)}_{Augmented}=\\text{log}{P\\left(A\\right)}_{Prior}+\\sigma\\:\\:\\times\\:\\:\\text{S}\\left(\\text{A}\\right)\\)\u003c/span\u003e \u003c/span\u003e (5) \u003csup\u003e28\u003c/sup\u003e\u003c/p\u003e \u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\text{S}\\left(\\text{A}\\right)\\)\u003c/span\u003e\u003c/span\u003e is an external scoring function that quantifies sequence quality, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\sigma\\:\u0026gt;0\\)\u003c/span\u003e\u003c/span\u003e is a scaling factor.\u003c/p\u003e \u003cp\u003eIf we choose the reward for the final step of the sequence A to:\u003c/p\u003e\n\u003ch3\u003e (6)\u003c/h3\u003e\n\u003cdiv class=\"Heading\"\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:G\\left(A\\right)\\:=\\:{\\left[\\text{log}{P\\left(A\\right)}_{Agent}-\\text{log}{P\\left(A\\right)}_{Augmented}\\right]}^{2}/\\text{log}{P\\left(A\\right)}_{Agent}\\)\u003c/span\u003e\u003c/span\u003e (6)\u003csup\u003e28\u003c/sup\u003e\u003c/div\u003e \u003cp\u003ethen the agent\u0026rsquo;s objective becomes:\u003cdiv id=\"Equ3\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ3\" name=\"EquationSource\"\u003e\n$$\\:J\\left(\\theta\\:\\right)=\\frac{{\\left[\\text{log}{P\\left(A\\right)}_{Agent}-\\text{log}{P\\left(A\\right)}_{Augmented}\\right]}^{2}}{\\text{log}{P\\left(A\\right)}_{Agent}}\\text{l}\\text{o}\\text{g}{\\text{P}\\left(A\\right)}_{Agent}={\\left[\\text{log}{P\\left(A\\right)}_{Agent}-\\text{log}{P\\left(A\\right)}_{Augmented}\\right]}^{2}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e7\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eThis formulation encourages the agent\u0026rsquo;s policy to produce log-likelihoods that align with the augmented target, which itself balances prior likelihood and an external score.\u003c/p\u003e \u003cp\u003e \u003cem\u003eThe Agent Network.\u003c/em\u003e \u003c/p\u003e \u003cp\u003eIn our specific setting, the task is to generate protein sequences that fold into a given backbone structure \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:X\\)\u003c/span\u003e\u003c/span\u003e, The prior policy is a pre-trained ProteinMPNN network model, which defines the conditional distribution:\u003cdiv id=\"Equ4\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ4\" name=\"EquationSource\"\u003e\n$$\\:{p}_{Prior}\\left(Y|X\\right)={p}_{\\text{P}\\text{r}\\text{o}\\text{t}\\text{e}\\text{i}\\text{n}\\text{M}\\text{P}\\text{N}\\text{N}}\\left(Y|X\\right)=P\\left({y}_{t}\\right|{y}_{t-1},\\cdots\\:,{y}_{1},\\varvec{X})$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e8\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eIts negative log-likelihood (NLL) for a sequence \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:Y\\)\u003c/span\u003e\u003c/span\u003e is:\u003cdiv id=\"Equ5\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ5\" name=\"EquationSource\"\u003e\n$$\\:{NLL\\left(Y\\right|X)}_{Prior}\\:=-{\\sum\\:}_{n}\\text{log}{p}_{Prior}\\left(Y\\right|X)$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e9\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eThe agent is another ProteinMPNN network model with the same architecture but trainable parameters. Sequences generated by the agent are evaluated using a multiparameter objective (MPO) score, which measures desired biochemical or structural properties. According to Eq.\u0026nbsp;5, the MPO score is scaled by \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\sigma\\:\\)\u003c/span\u003e\u003c/span\u003e and combined with the prior NLL to form the augmented target:\u003cdiv id=\"Equ6\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ6\" name=\"EquationSource\"\u003e\n$$\\:{NLL\\left(\\varvec{Y}\\right|\\varvec{X})}_{Augmented}={NLL\\left(\\varvec{Y}\\right|\\varvec{X})}_{Prior}+\\sigma\\:\\:\\times\\:\\:MPO\\left(\\:\\varvec{Y}\\right|\\varvec{X})$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e10\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eAccording to Eq.\u0026nbsp;\u003cspan refid=\"Equ3\" class=\"InternalRef\"\u003e7\u003c/span\u003e, the agent is trained to minimize the squared difference between its own predicted NLL and this augmented target:\u003cdiv id=\"Equ7\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ7\" name=\"EquationSource\"\u003e\n$$\\:loss={[{NLL\\left(\\varvec{Y}\\right|\\varvec{X})}_{Augmented}\\:-\\:{NLL\\left(\\varvec{Y}\\right|\\varvec{X})}_{Agent}]}^{2}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e11\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eOnly the agent\u0026rsquo;s parameters are updated via backpropagation; the prior remains fixed. This approach anchors the agent to the prior distribution while steering it toward sequences with higher MPO scores, effectively integrating generative modeling and goal-directed optimization into a single differentiable framework.\u003c/p\u003e \u003cp\u003eThis method provides a principled alternative to direct reward maximization with REINFORCE\u003csup\u003e\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e\u003c/sup\u003e. By framing the update as matching an augmented likelihood, which blends prior knowledge and task-specific desirability, we maintain proximity to the prior distribution, avoid high-variance gradient estimates, and enable stable fine-tuning. The result is an agent that generates sequences which are both plausible under the structural prior and optimized for the specified scoring function.\u003c/p\u003e \u003cp\u003e \u003cem\u003eScore function.\u003c/em\u003e \u003c/p\u003e \u003cp\u003eThe framework provides a flexible scoring framework defined by a general formulation (Eq.\u0026nbsp;\u003cspan refid=\"Equ8\" class=\"InternalRef\"\u003e12\u003c/span\u003e). This framework allows users to combine individual score components p either through a weighted sum or a weighted product.\u003csup\u003e\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e\u003c/sup\u003e For sequence \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:x\\)\u003c/span\u003e\u003c/span\u003e, Each score, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{score}_{i}\\)\u003c/span\u003e\u003c/span\u003e, is assigned a strictly positive weight coefficient \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:w(w\\in\\:\\left(0,\\:+\\infty\\:\\right))\\)\u003c/span\u003e\u003c/span\u003e, reflecting its relative contribution to the final score. Individual component scores span the non-negative range [0, +\u0026infin;):\u003cdiv id=\"Equ8\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ8\" name=\"EquationSource\"\u003e\n$$\\:MPO\\left(x\\right)=\\frac{{\\sum\\:}_{i}{w}_{i}\\times\\:{score}_{i}\\left(x\\right)}{{\\sum\\:}_{i}{w}_{i}}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e12\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eThe formulation is provided for user convenience and flexibility.\u003c/p\u003e \u003cp\u003e \u003cem\u003eGenerating New Sequence Samples.\u003c/em\u003e \u003c/p\u003e \u003cp\u003eDuring RL training, the amino acid sequences were sampled by ProteinMPNN with the conditional probability from RL training, and the coordinates were then generated with the backbone by Rosetta and refined by FastRelax.\u003c/p\u003e \u003cp\u003e \u003cem\u003eTNFR1 study case for RL-Augmented ProteinMPNN framework.\u003c/em\u003e \u003c/p\u003e \u003cp\u003eIn this study, we focus here on optimizing \u003cb\u003eS1B2\u003c/b\u003e to enhance its functional activity. the inputs files for this framework are available at GitHub(\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/weilin199204/SuperBinder\u003c/span\u003e\u003cspan address=\"https://github.com/weilin199204/SuperBinder\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), and σ is set to 0.05, while batch_size set to 1, and num_seq_per_target (numbers of sequence at per sample step) set to 20. Other parameters are default.\u003c/p\u003e \u003cp\u003eHere MPO score function is as Eq.\u0026nbsp;13:\u003c/p\u003e \u003cp\u003e \u003cspan class=\"InlineEquation\"\u003e \u003cspan class=\"mathinline\"\u003e\\(\\:p\\left(x\\right)={pK}_{D}(x\\)\u003c/span\u003e \u003c/span\u003e) (13)\u003c/p\u003e \u003cp\u003eHere, the negative logarithm dissociation constant to base 10 (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{pK}_{D}\\)\u003c/span\u003e\u003c/span\u003e) (Eq.\u0026nbsp;\u003cspan refid=\"Equ9\" class=\"InternalRef\"\u003e14\u003c/span\u003e), was predicted by ProAffinity-GNN\u003csup\u003e\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e\u003c/sup\u003e, a novel approach to structure-based protein\u0026ndash;protein binding affinity prediction model,\u003cdiv id=\"Equ9\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ9\" name=\"EquationSource\"\u003e\n$$\\:{pK}_{D}=-\\text{log}{K}_{D}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e14\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003e \u003cem\u003eMolecule dynamics (MD) and binding free energy calculation with MM/GBSA.\u003c/em\u003e \u003c/p\u003e \u003cp\u003eSame as our previous work\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e, to evaluate the binding capability of designed minibinders to TNFR1, molecular dynamics (MD) simulations of each minibinder complexed with TNFR1 alone were performed for 300 ns by using the AMBER force field of GROMACS 2019.6 software. Three hundred snapshot structures for each minibinder-TNFR1 complex were extracted from the smooth MD trajectory at equal intervals for binding free energy calculation using the MM/GBSA method.\u003c/p\u003e \u003cp\u003e \u003cem\u003eExpression, purification of binders and swine TNFR1.\u003c/em\u003e \u003c/p\u003e \u003cp\u003eThe genes sequences for the minibinders were cloned into a pET-22b vector, incorporating a C-terminal 6\u0026times;His tag. The resulting plasmids were transformed into \u003cem\u003eEscherichia coli\u003c/em\u003e BL21(DE3) competent cells. Protein expression was induced by the addition of IPTG, followed by incubation at 18 ℃ for 20 hours. Cells were harvested via centrifugation at 7,000 rpm for 10 minutes and lysed by sonication. To prevent proteolytic degradation, PMSF was added to a final concentration of 100 mM. The clarified lysate was subjected to purification by Ni-NTA affinity chromatography, which included a wash step with buffer containing 30 mM imidazole and subsequent elution using 300 mM imidazole. Further purification was achieved by size-exclusion chromatography on an FPLC system equipped with a Superdex 75 Increase 10/300 column (Cytiva). Protein concentration was quantified using a BCA assay kit (Biosharp, BL521A), and sample purity was confirmed by SDS-PAGE. The Avi-tag extracellular domain of human TNFR1 was expression, purification some as our previous work\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003e \u003cem\u003eBiolayer interferometry\u003c/em\u003e.\u003c/p\u003e \u003cp\u003eThe binding affinity of the purified minibinders for TNFR1 was quantitatively evaluated by biolayer interferometry (BLI) using an Octet RED96 instrument (ForteBio), following an established procedure\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e. Kinetics constants (K\u003csub\u003eon\u003c/sub\u003e, K\u003csub\u003eoff\u003c/sub\u003e) were determined by globally fitting the association and dissociation curves to a 1:1 binding model. The affinity constant (K\u003csub\u003eD\u003c/sub\u003e) for each minibinder was calculated following the equation: K\u003csub\u003eD\u003c/sub\u003e = K\u003csub\u003eoff\u003c/sub\u003e / K\u003csub\u003eon\u003c/sub\u003e. Data were processed using Octet Analysis Studio v.13.0.1.35.\u003c/p\u003e \u003cp\u003e \u003cem\u003eNeutralizing potency measured in L929 cell-based assay.\u003c/em\u003e \u003c/p\u003e \u003cp\u003eThe neutralizing potency of the minibinders was assessed using a TNFα-sensitive murine L929 cell-based assay, as described previously\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e. Authenticated L929 cells were cultured in complete DMEM under standard conditions. For the assay, cells were seeded in 96-well plates (5,000 cells/well) and treated with serially diluted minibinders (0.78\u0026ndash;400 nM) for 30 min, followed by co-treatment with TNFα (10 pM) and actinomycin D (1 \u0026micro;g/mL). After 24 h incubation, cell viability was measured using a CCK-8 kit. Dose-response curves were analyzed with GraphPad Prism to determine IC\u003csub\u003e50\u003c/sub\u003e values.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eData\u0026nbsp;\u003c/strong\u003e\u003cstrong\u003eand Code Availability Statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe code for reinforcement learning-augmented ProteinMPNN is available at GitHub (https://github.com/weilin199204/SuperBinder). All data generated or analyzed during the current study are included in this article and \u003cstrong\u003eSupplementary information\u003c/strong\u003e.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgments\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was supported by Key Research and Development Progarm of Wuhan (2025020602030105), Hubei Provincial Natural Science Foundation of China (2024AFB465), Major Special Projects for the Development of Agricultural Microbial Industry in Hubei Province (NYWSWZX2025-2027-03, NYWSWZX2025-2027-10), State Key Laboratory Open Project (SKLBEE2022007) from State Key Laboratory of Biocatalysis and Enzyme Engineering. We thank HPC center at State Key Laboratory of Biocatalysis and Enzyme Engineering for cloud computing support.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor information\u0026nbsp;\u003c/strong\u003e\u003cstrong\u003econtributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eL. W.: formal analysis, funding acquisition, methodology, resources, software, visualization, writing \u0026ndash; original draft, writing \u0026ndash; review \u0026amp; editing; Z. Wu, Y. H.: formal analysis, methodology, validation, visualization, writing \u0026ndash; original draft, writing \u0026ndash; review \u0026amp; editing; Y. F.: formal analysis, validation, writing \u0026ndash; original draft, writing \u0026ndash; review \u0026amp; editing; M. G., B. X.: validation, writing \u0026ndash; review \u0026amp; editing; J. W.: funding acquisition, methodology, supervision, writing \u0026ndash; original draft, writing \u0026ndash; review \u0026amp; editing; S. L.: formal analysis, software, writing \u0026ndash; original draft, writing \u0026ndash; review \u0026amp; editing; K. M.: funding acquisition, Writing \u0026ndash; review \u0026amp; editing; Z. W.: Funding acquisition, Supervision, Writing \u0026ndash; review \u0026amp; editing.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics declarations\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSupplementary information.\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eMore computational and experimental methods, details of the results can be found online.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eChevalier, A., et al.: Massively parallel de novo protein design for targeted therapeutics. Nature. \u003cb\u003e550\u003c/b\u003e, 74\u0026ndash;79 (2017)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRocklin, G.J., et al.: Global analysis of protein folding using massively parallel design, synthesis, and testing. Science. \u003cb\u003e357\u003c/b\u003e, 168\u0026ndash;175 (2017)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCao, L., et al.: De novo design of picomolar SARS-CoV-2 miniprotein inhibitors. Science. \u003cb\u003e370\u003c/b\u003e, 426\u0026ndash;431 (2020)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBerger, S., et al.: Preclinical proof of principle for orally delivered Th17 antagonist miniproteins. Cell. \u003cb\u003e187\u003c/b\u003e, 1\u0026ndash;13 (2024)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCao, L., et al.: Design of protein-binding proteins from the target structure alone. Nature. \u003cb\u003e605\u003c/b\u003e, 551\u0026ndash;560 (2022)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWeng, J., et al.: Design of minibinder proteins specific to TNFR1. Int. J. Biol. Macromol. \u003cb\u003e293\u003c/b\u003e, 139403 (2025)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWei, L., et al.: De novo design mini-binder proteins targeting the glycoproteins D to inhibit PRV replication in PK15 cells. Int. J. Biol. Macromol., 144403 (2025)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuang, B., et al.: De novo design of miniprotein antagonists of cytokine storm inducers. Nat. Commun. \u003cb\u003e15\u003c/b\u003e, 7064 (2024)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBennett, N.R., et al.: Improving de novo protein binder design with deep learning. Nat. Commun. \u003cb\u003e14\u003c/b\u003e, 2625 (2023)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIngraham, J., Garg, V.K., Barzilay, R., Jaakkola, T.: Generative models for graph-based protein design. \u003cem\u003eNeurIPS 2019\u003c/em\u003e (2019)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAnand, N., et al.: Protein sequence design with a learned potential. Nat. Commun. \u003cb\u003e13\u003c/b\u003e, 746 (2022)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJing, B., Eismann, S., Suriana, P., Townshend, R.J.L., Dror, R.: Learning from Protein Structure with Geometric Vector Perceptrons. \u003cem\u003earXiv:2009.01411\u003c/em\u003e (2020)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStrokach, A., Becerra, D., Corbi-Verge, C., Perez-Riba, A., Kim, P.M.: Fast and Flexible Protein Design Using Deep Graph Neural Networks. Cell. Syst. \u003cb\u003e11\u003c/b\u003e, 402\u0026ndash;411e404 (2020)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHsu, C., et al.: Learning inverse folding from millions of predicted structures. \u003cem\u003ebioRxiv:2022.04.10.487779v2\u003c/em\u003e (2022)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang, Y., et al.: ProDCoNN: Protein design using a convolutional neural network. Proteins. \u003cb\u003e88\u003c/b\u003e, 819\u0026ndash;829 (2020)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eQi, Y., Zhang, J.Z.H., DenseCPD: Improving the Accuracy of Neural-Network-Based Computational Protein Sequence Design with DenseNet. J. Chem. Inf. Model. \u003cb\u003e60\u003c/b\u003e, 145\u0026ndash;1252 (2020)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWatson, J.L., et al.: De novo design of protein structure and function with RFdiffusion. Nature. \u003cb\u003e620\u003c/b\u003e, 1089\u0026ndash;1100 (2023)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDauparas, J., et al.: Robust deep learning\u0026ndash;based protein sequence design using ProteinMPNN. Science. \u003cb\u003e378\u003c/b\u003e, 49\u0026ndash;56 (2022)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJumper, J., et al.: Highly accurate protein structure prediction with AlphaFold. Nature. \u003cb\u003e596\u003c/b\u003e, 583\u0026ndash;589 (2021)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHo, J., Jain, A., Abbeel, P.: Denoising Diffusion Probabilistic Models. \u003cem\u003earXiv:2006.11239\u003c/em\u003e (2020)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBaek, M., et al.: Accurate prediction of protein structures and interactions using a three-track neural network. Science. \u003cb\u003e373\u003c/b\u003e, 871\u0026ndash;876 (2021)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu, Z., et al.: De novo designed mini-binders targeting glyceraldehyde-3-phosphate dehydrogenase of Streptococcus equi ssp. zooepidemicus provided partial protection in mice model of infection. Int. J. Biol. Macromol. \u003cb\u003e307\u003c/b\u003e, 142293 (2025)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMing, K., et al.: Mini-binders targeting Streptococcus equi ssp. zooepidemicus M-like protein inhibit the bacterial adhesion and exert protective effects in vivo. Int. J. Biol. Macromol. \u003cb\u003e304\u003c/b\u003e, 140803 (2025)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eQuijano-Rubioa, A., Ulge, U.Y., Walkey, C.D., Silva, D.: The advent of de novo proteins for cancer immunotherapy. Curr. Opin. Chem. Biol. \u003cb\u003e56\u003c/b\u003e, 119\u0026ndash;128 (2020)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSutton, R.S., Barto, A.G.: Reinforcement Learning: An Introduction. The MIT Press (2015)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSilver, D., et al.: Mastering the game of Go without human knowledge. Nature. \u003cb\u003e550\u003c/b\u003e, 354\u0026ndash;359 (2017)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuo, D., et al.: DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature. \u003cb\u003e645\u003c/b\u003e, 633\u0026ndash;638 (2025)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOlivecrona, M., Blaschke, T., Engkvist, O., Chen, H.: Molecular de-novo design through deep reinforcement learning. J. Cheminform. \u003cb\u003e9\u003c/b\u003e, 48 (2017)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBlaschke, T., et al.: REINVENT 2.0: An AI Tool for De Novo Drug Design. J. Chem. Inf. Model. \u003cb\u003e60\u003c/b\u003e, 5918\u0026ndash;5922 (2020)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLoeffler, H.H., et al.: Reinvent 4: Modern AI\u0026ndash;driven generative molecule design. J. Cheminform. \u003cb\u003e16\u003c/b\u003e, 20 (2024)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLutz, I.D., et al.: Top-down design of protein architectures with reinforcement learning. Science. \u003cb\u003e380\u003c/b\u003e, 266\u0026ndash;273 (2023)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePopova, M., Isayev, O., Tropsha, A.: Deep reinforcement learning for de novo drug design. Sci. Adv. \u003cb\u003e4\u003c/b\u003e, eaap7885 (2018)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIvanenkov, Y.A., et al.: Chemistry42: An AI-Driven Platform for Molecular Design and Optimization. J. Chem. Inf. Model. \u003cb\u003e63\u003c/b\u003e, 695\u0026ndash;701 (2023)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJaques, N., Gu, S., Turner, R.E., Eck, D.: Tuning Recurrent Neural Networks with Reinforcement Learning. \u003cem\u003earXiv:1611.02796v1\u003c/em\u003e (2017)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003evan Loo, G., Bertrand, M.J.M.: Death by TNF: a road to inflammation. Nat. Rev. Immunol. \u003cb\u003e23\u003c/b\u003e, 289\u0026ndash;303 (2022)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLeone, G.M., Mangano, K., Petralia, M.C., Nicoletti, F., Fagone, P., Past: Present and (Foreseeable) Future of Biological Anti-TNF Alpha Therapy. J. Clin. Med. \u003cb\u003e12\u003c/b\u003e, 1630 (2023)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAggarwal, B.B.: Signalling pathways of the TNF superfamily: a double-edged sword. Nat. Rev. Immunol. \u003cb\u003e3\u003c/b\u003e, 745\u0026ndash;756 (2003)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSiegmund, D., Wajant, H.: TNF and TNF receptors as therapeutic targets for rheumatic diseases and beyond. Nat. Rev. Rheumatol. \u003cb\u003e19\u003c/b\u003e, 576\u0026ndash;591 (2023)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhou, Z., et al.: ProAffinity-GNN: A Novel Approach to Structure-Based Protein\u0026ndash;Protein Binding Affinity Prediction via a Curated Data Set and Graph Neural Networks. J. Chem. Inf. Model. \u003cb\u003e64\u003c/b\u003e, 8796\u0026ndash;8808 (2024)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu, Y., Yu, Q., Wang, D., Chen, M.: AF3Score: A Score-Only Adaptation of AlphaFold3 for Biomolecular Structure Evaluation. J. Chem. Inf. Model. \u003cb\u003e65\u003c/b\u003e, 8207\u0026ndash;8214 (2025)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCao, L., et al.: De novo design of picomolar SARS-CoV-2 miniprotein inhibitors. Science. \u003cb\u003e370\u003c/b\u003e, 426\u0026ndash;431 (2020)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGl\u0026ouml;gl, M., et al.: Target-conditioned diffusion generates potent TNFR superfamily antagonists and agonists. Science. \u003cb\u003e386\u003c/b\u003e, 1154\u0026ndash;1161 (2024)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJ., W.R. Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning. Mach. Learn. \u003cb\u003e8\u003c/b\u003e, 229\u0026ndash;256 (1992)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePeters, J.: Policy gradient methods. Scholarpedia. \u003cb\u003e5\u003c/b\u003e, 3698 (2010)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCummins, D.J., Bell, M.A., Integrating Everything: The Molecule Selection Toolkit, a System for Compound Prioritization in Drug Discovery. J. Med. Chem. \u003cb\u003e59\u003c/b\u003e, 6999\u0026ndash;7010 (2016)\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Minibinder, Protein design, Reinforcement learning, ProteinMPNN, Tumor Necrosis Factor Receptor 1 (TNFR1)","lastPublishedDoi":"10.21203/rs.3.rs-8645836/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8645836/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eMinibinders are compact protein molecules that hold great promise as therapeutic agents due to their target-binding specificity, stability, and potential for oral delivery. However, the primary objective ofthe widely used minibinder sequence design method, ProteinMPNN, is structural fidelity rather than the optimization of specific functional properties such as binding affinity. Consequently, the \u003cem\u003ede novo\u003c/em\u003e design of high-affinity minibinders still largely relies on extensive wet-lab screening of numerous candidates. Directly designing high-affinity minibinders \u003cem\u003ein silico\u003c/em\u003e thus remains a critical challenge for numerous therapeutic applications. Here, we present a computational framework that integrates a reinforcement learning (RL) framework with ProteinMPNN network to directly generate minibinder sequences with improved functional features. We demonstrate the power of this framework by optimizing the activities of the minibinders targeting Tumor Necrosis Factor Receptor 1 (TNFR1). Experimental validation showed that the optimized minibinders \u003cb\u003eOPT1\u003c/b\u003e and \u003cb\u003eOPT7\u003c/b\u003e exhibited 3-fold and 7-fold higher binding affinity, respectively, and 6-fold and 4-fold greater neutralizing activity in cells compared to the original minibinder \u003cb\u003eS1B2\u003c/b\u003e. Our work establishes the framework as a promising tool that augments AI-driven protein design for the \u003cem\u003ede novo\u003c/em\u003e development of high-affinity therapeutic minibinders.\u003c/p\u003e","manuscriptTitle":"Reinforcement Learning-Augmented ProteinMPNN Improve the Binding Affinity of TNFR1-Targeting Minibinders","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-01-21 10:53:58","doi":"10.21203/rs.3.rs-8645836/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"c3e83974-31ec-4844-aa56-6d6feffa8e2b","owner":[],"postedDate":"January 21st, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":61428056,"name":"Biological sciences/Computational biology and bioinformatics/Computational models"},{"id":61428057,"name":"Health sciences/Medical research/Drug development"}],"tags":[],"updatedAt":"2026-02-02T06:16:15+00:00","versionOfRecord":[],"versionCreatedAt":"2026-01-21 10:53:58","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8645836","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8645836","identity":"rs-8645836","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-29T02:00:03.542394+00:00
License: CC-BY-4.0