IRIS Integrates Sparse Sequence, Experimental, and AI-Predicted Structures for Protein-RNA Affinity Prediction and Motif Discovery

preprint OA: closed CC-BY-NC-ND-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Protein–RNA interactions are fundamental to numerous cellular processes, yet quantitatively characterizing their binding specificity remains a major challenge. We present IRIS (Integrative RNA–protein interaction prediction Informed by Structure and sequence), a biophysical framework that integrates residue-level sequence and structural features without relying on large-scale affinity data to predict binding affinities and identify binding motifs. Applied across different protein–RNA systems, IRIS predicts relative binding free energies (ΔΔ G ) with consistent correlations and competitive error metrics, and its performance is further improved by incorporating additional high-affinity sequences into the training set. By leveraging predicted structural complexes, IRIS reveals alternative binding modes not observed in experimental structures, extends applicability to systems lacking experimental protein–RNA complexes, and generates a library of favorable RNA-binding motifs at protein–RNA interfaces. Collectively, these results establish IRIS as a versatile framework that leverages increasingly accurate structural predictions to enable quantitative modeling and rational engineering of protein–RNA interactions.
Full text 100,072 characters · extracted from oa-pdf · 2 sections · click to expand

Materials

and Methods355 Protein–RNA Model Architecture and Training356 IRIS uses a residue-level biophysical framework that integrates detailed physicochemical357 modeling with structural data, either experimentally determined or predicted by structure-358 prediction tools, 39–42 to quantify sequence-specific protein-RNA binding a !nities (Figure 1).359 The model estimates relative binding free energies, ranks binding strengths, and converts360 these free energies into dissociation constants via the thermodynamic relation: 90361 ##Gbinding = RT ln ( #Kd ) . (1) This approach integrates structural and sequence signatures of the protein-RNA binding362 interface into an energy model to predict protein-RNA interactions with high precision.363 For training, IRIS identifies all residue pairs within the protein-RNA structural inter-364 face using C ε-P atom pairs that are closer than 0.95 nm in the training structures. The365 training sets include high-a!nity “native” binders and 10 , 000 randomized RNA sequence366 “decoy” binders. The integration of structural and sequence information of given protein-367 RNA complexes maximizes the usage of information encoded in the evolutionarily favorable368 21 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint protein-RNA interfaces while reducing reliance on large-scale experimental binding sequence369 data.370 Each complex is represented by a 20-by-4 amino-acid-nucleotide interaction matrix that371 captures pairwise contact frequencies between each amino acid ( ai)a n dn u c l e o t i d e(nj )w i t h i n372 the protein-RNA interface, computed using the equation:373 ϑ(a, n)= ∑ i↑protein ∑ j↑RNA $i,j (ri,j )( 2) where $i,j (rij )c a nb ec a l c u l a t e dv i a :374 $i,j (ri,j )= 1 2 [ tanh ( ϖ(r ↑ rmin) ) · tanh ( ϖ(rmax ↑ r) )] + 1 2 (3) Here, rmin = ↑ 9.5 ˚A,r max =9 .5 ˚ A,ϖ =0 .7. This defines two residues to be “in contact” 375 if they are separated less than 9.5 ˚ A. The parameter ϖ modulates the steepness of the 376 hyperbolic tangent function. i and j are sequence indices of the amino acid and nucleotide.377 The solvent-averaged binding energy between the protein and RNA is computed as the378 summed individual amino-acid-nucleotide pair energies at the protein-RNA interface:379 Ebinding = ∑ a↑protein,n↑RNA ϱ(a, n)ϑ(a, n)( 4) Here, ϱ(a, n)i sa2 0 - b y - 4e n e r g ym a t r i xe n c o d i n gr e s i d u e - t y p es p e c i fi ci n t e r a c t i o ns t r e n g t h380 between amino acid ( a) and nucleotide ( n). By defining binding free energy in this manner,381 our model captures sequence-specific protein-RNA binding while considering their structural382 context.383 The model optimizes ϱ(a, n)e n e r g ym a t r i xb ym a x i m i z i n gt h ee n e r g yg a pb e t w e e nt h e384 given strong binders and their corresponding decoy binders. This approach draws inspi-385 ration from methods used in protein folding, protein-protein and protein-DNA interaction386 studies.47,91–95 In practice, IRIS computes the average energy gap between the strong and387 decoy binders using ςE = ↔Edecoy↗↑↔Estrong↗,a n dt h es t a n d a r dd e v i a t i o no fd e c o yb i n d i n g388 22 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint energies are computed by #E =S t d (Edecoy), where Std stands for the standard deviation.389 Using the interaction matrix (Eq. 2), we have390 A(a, n)= ↔ϑdecoy↗↑↔ϑstrong↗↘ R1↓300 (5) 391 B(a, n)= ⟨ ϑdecoy ϑT decoy ⟩ ↑↔ϑdecoy↗↔ϑT decoy ↗↘ R300↓300 . (6) and IRIS maximizes the separation of strong and decoy binding energies using:392 ςE #E = Aϱ√ ϱT Bϱ (7) by optimizing393 R(ϱ)= AT ϱ ↑ φ √ ϱT Bϱ . (8) whose solution is ϱ ≃ B↔1 A.T o e n s u r e r o b u s t n e s s , w e r e g u l a r i z e t h e s o l u t i o n b y :( 1 ) d i a g o -394 nalizing B = P %P ↔1 ,( 2 )r e t a i n i n gt h et o p2 5e i g e n m o d e s ,( 3 )r e p l a c i n gs m a l l e re i g e n v a l u e s395 with the 25th largest value, and (4) reconstructing B↔1 from these filtered components.396 The optimized interaction energy matrix ϱ can be visualized in the 20 ⇐ 4 form (Figure 1),397 revealing the physicochemical preference at amino acid–nucleotide resolution.398 To predict the binding a !nity of a target protein-RNA sequence, we substitute it into399 the native high-a !nity protein-RNA complex structure and recalculate the target ϑ matrix400 using Eq. 2.T h i s i n t e r a c t i o n m a t r i x w i l l b e c o m b i n e d w i t h t h e t r a i n e d i n t e r a c t i o n e n e r g yϱ401 matrix to compute the target protein-RNA binding free energies using:402 Etarget = ϑT target ϱ. (9) Since we focused our predictions on various RNA sequences binding to the same RNA-403 binding protein, we hypothesized that all of them share the same conformational entropy404 and approximated the di ”erence in binding free energy ##Gtarget to be the di” erence in405 23 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint computed solvent-average binding free energy Etarget.T h i s h y p o t h e s i s w i l l n e e d t o b e r e v i s e d406 when considering plasticity in RNA structural ensemble, 55 especially for those RNAs with407 more drastic sequence mutations. Additionally, due to the presence of an undetermined408 scaling factor when optimizing the ϱ energy matrix, the predicted free energies are presented409 in reduced units. Despite this, the relative free energies can be used to accurately rank410 binding a!nities of target protein-RNA pairs, and the relative free energies can be converted411 to #Kd using Eq. 1. Additionally, the predicted binding free energies can be used to identify412 preferred RNA binding motifs.413 Predicting Binding A!nities of Mutant RNA Sequences414 As described in Section Protein–RNA Model Architecture and Training,m u t a n tb i n d i n g415 a!nities were predicted by substituting the native RNA sequence in the protein–RNA struc-416 ture with the target sequences. Binding free energies were then computed by combining the417 updated ϑ interaction matrix with the learned ϱ energy matrix using Eq. 4.T o a s s e s s p r e d i c -418 tive performance, mutant sequences were grouped by the number of nucleotide substitutions419 and by their proximity to the protein interface, defined using distance cuto”s ranging from420 0.3 to 0.95 nm.421 All MS2 testing sequences were obtained from Reference 24.O n l y s e q u e n c e s w i t h i n422 the measurable range of their RNA array were included, excluding those with maximum423 (##G =6 .661693623 kB T )o rm i n i m u m(# #G = ↑ 0.885494 kB T )v a l u e s . T h ec r y s t a l424 structure used in this study (PDB ID: 2C4Q) lacks nucleotides at both the 5 → and 3→ termini.425 For testing sequences carrying mutations at these missing p ositions (-15 at the 5 → end or426 the terminal 3 → nucleotide), the corresponding bases were omitted when constructing the427 ϑ contact matrix. This allowed such sequences to be included in Figure 2A, but did not428 influence a!nity predictions because these terminal nucleotides are absent in the 2C4Q429 training structure and therefore indistinguishable to the model.430 In addition to the native crystal structure, we also used the deep learning–based tool431 24 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint AlphaFold3 (AF3)39 to generate protein–RNA complex structures for the mutant sequences.432 These predicted structures were then used to calculate the ϑ matrices and the corresponding433 binding free energies using Eqs. 2 and 4.434 Correlation and Linear Regression Analysis of Binding Free Ener-435 gies436 To assess predictive accuracy, we compared IRIS-predicted and exp erimentally measured437 relative binding free energies (## G) for MS2 protein binding on mutated RNA sequences.438 Sequences were grouped into two categories: those with 0–3 substitutions and those with four439 substitutions (4). Pearson’s r and Spearman’s ω were calculated to quantify linear agreement440 and rank consistency, respectively. For the 4-substitution group, correlations were computed441 using only quadruple mutants, with high-a!nity outliers retained in the analysis but omitted442 from the Sub 4 legend for clarity.443 A linear regression constrained through the origin (intercept = 0) was fitted to the 0–3444 substitution group to obtain a best-fit slope capturing the relationship between predicted445 and experimental binding energies (Figures 2Aa n d 3C). This regression line was applied446 as a common reference across both the 0–3 and the 4-substitution panels, enabling direct447 visual comparison between the two groups. High-a !nity outliers are indicated with purple448 triangles.449 All plots were generated with standardized axis limits and consistent color and shape450 encodings, facilitating visual comparison across groups. This design highlights both the451 overall predictive strength of the model and its limitations when applied to highly divergent452 sequences.453 25 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint Binding Interface Cuto” Modulates Predictive Accuracy454 To evaluate how the definition of the binding interface influences predictive p erformance, we455 calculated Pearson correlations between predicted and experimental binding energies across456 residue distance cuto”s ranging from 0.3 to 0.95 nm. This analysis was performed on five pro-457 tein–RNA complexes with experimentally resolved structures: 2C4Q, 2ERR, 1IM8, 1HQ1,458 and 1URN (Figures 2Ba n d4 B). For each complex, predicted binding free energies were com-459 puted at five interface distance thresholds, and correlations with experimental values were460 determined. For MS2 (2C4Q), results were further stratified by the number of nucleotide461 substitutions (1–4) at each cuto”. The data were visualized as faceted bar plots, including a462 zero-correlation reference line to highlight thresholds where predictive accuracy diminishes463 or reverses. This analysis underscores the model’s sensitivity to interface definitions and464 helps identify optimal distance thresholds for accurate binding a !nity prediction.465 Binding Interface Motifs Across Distance Cuto”s466 To examine how the definition of the binding interface a ”ects conserved RNA-binding motifs,467 we analyzed the base composition of interface residues at di ”erent distance cuto”s (0.3, 0.4,468 and 0.5 nm) using sequence logo representations. This analysis was performed on the MS2469 protein–RNA complex, focusing on sequences with 1 and 4 substitutions (Figure 3D), as well470 as on test sequences from four additional protein–RNA complexes: 2ERR, 1IM8, 1HQ1, and471 1URN (Figure 4A).472 For each cuto”, RNA nucleotide identities were extracted using residue-level annotations473 and mapped to sequence positions. To avoid visualization redundancy and highlight the474 unique contributions of each cuto” , only residues within the shortest cuto” were retained.475 Nucleotide frequencies (A, C, G, U) were then computed to generate nucleotide distributions476 at each distance threshold.477 Stacked sequence logos were generated from these frequencies using the ggseqlogo package478 in R. 96 These visualizations illustrate how nucleotide enrichment patterns shift across dis-479 26 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint tance cuto”s, revealing both shared and unique sequence preferences within the protein–RNA480 interface. This analysis provides insight into which distance thresholds capture meaningful481 RNA motifs, guiding the selection of high-a !nity binding motifs for target RNA-binding482 proteins (Figure 6).483 Total Residual Error Across Mutation Positions and Nucleotide484 Identities485 To investigate how specific nucleotide substitutions impact the predictive accuracy of MS2486 coat protein–RNA interactions, we computed the Total Residual Error stratified by both487 mutation position and nucleotide identity. All mutational variants and the wild-type RNA488 sequences used in Figure 1 were obtained and processed as described in Section Predicting489 Binding A”nities of Mutant RNA Sequences,w h e r eb i n d i n gf r e ee n e r g i e sw e r ep r e d i c t e d490 by substituting the native RNA sequence and computing ##G values from the ϑ and ϱ491 matrices.492 We leveraged the linear regression mo del describ ed in Section Correlation and Linear Re-493 gression Analysis of Binding Free Energies, which was fitted through the origin on sequences494 containing three or fewer substitutions ( Substitution ⇒ 3). This regression of predicted495 versus experimentally measured ##G values provides the expected predicted binding energy496 for a given experimental measurement under near-native mutational load, with the best-fit497 line shown in Figures 2Aa n d 3C.498 For each sequence, the residual error was calculated as the absolute di ”erence between499 the predicted binding energy and the value projected by the regression model:500 Residual Error = |##Gpredicted ↑ ##Gbest-fit|, (10) where ##Gbest-fit = m · ##Gexperimental and m is the slope of the fitted linear regression.501 Residual errors were then aggregated across all sequences, grouped by mutation position502 27 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint and substituted nucleotide (A, C, G, or U). This yields a total residual error for each nu-503 cleotide at each position. The results are visualized as a stacked bar plot (Fig. 3B), with504 bars colored by nucleotide identity and a fixed y-axis for consistent scaling. This representa-505 tion highlights positions and substitutions that disproportionately contribute to prediction506 errors, revealing model blind spots and guiding potential refinement.507 Jaccard Similarity Heatmap of RNA Sequences Against Known508 MS2 Motifs509 To compare structural similarity between input RNA sequences and known RNA-binding510 motifs, we generated a Jaccard similarity heatmap using a k-mer–based approach. For each511 pairwise comparison between a test RNA sequence and a motif, the Jaccard similarity was512 computed using all unique k-mers of length 3 ( k =3 ) ,d e fi n e dm a t h e m a t i c a l l ya s :513 J(A, B)= |K(A) ⇑ K(B)| |K(A) ⇓ K(B)| , (11) where K(A)a n d K(B)a r et h es e t so fk - m e r si ns e q u e n c e sA and B, respectively.514 To compare similarity among RNA sequences, we selected 10 RNA sequences from Ref-515 erence 24,i n c l u d i n gt h en a t i v es e q u e n c ei nt h ec r y s t a ls t r u c t u r e( P D BI D :2 C 4 Q ) ,7s t r o n g -516 binding 4-mutation RNA sequences (B1–B7), and two additional RNA mutants (R1 and517 R2). Experimentally identified MS2-RNA binding motifs from prior CLIP-seq and CryoEM518 studies56–58 (Motifs I–XIV) were also included. Each sequence was converted into a set of519 3-mers, and pairwise Jaccard similarities were calculated between the two groups.520 The resulting similarity matrix was visualized as a heatmap, with experimentally iden-521 tified motifs on the y-axis and RNA sequences from Reference 24 on the x-axis (Fig. 5A).522 This k-mer–level analysis enabled rapid structural comparison of short RNA sequences and523 facilitated identification of which motifs each test sequence most closely resembled. Because524 this approach is alignment-free, it is robust to minor shifts or indels, making it particularly525 28 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint well-suited for comparing short, structured RNAs.526 K-mer Network Construction and Clustering of High-A!nity RNA527 Fra gments528 To identify and visualize structure-informed patterns of RNA fragments associated with high529 protein-binding a!nity, we extracted contiguous runs of residue IDs and their corresponding530 nucleotide identities from all AF3-predicted RNA-binding structures. Blocks of protein-531 interacting RNA residues were defined under varying interaction distance cuto ”s, and for532 each residue, we recorded the lowest cuto” at which it remained in contact with the protein.533 Contiguous segments with more than four residues were further divided into overlapping534 sub-k-mers (e.g., 4-mers, 5-mers), with each k-mer annotated with its sequence, length,535 per-position cuto ” values, and the experimentally measured ##Gexp of the parent RNA536 sequence. To standardize comparisons across the dataset, ##Gexp values were converted537 into Z-scores:538 Z = ##Gexp ↑ µ!!Gexp ↼!!Gexp , where µ!!Gexp and ↼!!Gexp are the mean and standard deviation of all ##Gexp val-539 ues for the target RNA-binding protein, calculated across the full dataset prior to k-mer540 segmentation to avoid bias from fragment counts.541 After collecting all k-mers, Z-scores were aggregated for each unique sequence. We542 then computed a pairwise string similarity matrix using the Levenshtein distance via the543 stringdist R package. 97 This matrix served as input for agglomerative hierarchical clus-544 tering with average linkage (implemented using hclust in R). The resulting dendrogram545 (Fig. 6) highlights relationships among k-mers based on sequence similarity, revealing dis-546 tinct sequence families. This approach uncovers recurring structural and sequence features547 of RNA that strongly bind proteins and provides guidance for designing high-a !nity RNA548 29 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint sequences.549 AF3-Predicted HnRNPK–B Motif Structures and IRIS A!nity550 Prediction551 In the absence of experimentally determined hnRNPK-B motif complexes, we used AF3 39552 to generate predicted structures for IRIS model training and binding-a !nity prediction.553 The hnRNPK protein sequence was obtained from Uniprot 98 (https://www.uniprot.org/554 uniprotkb/P61978/entry#sequences), and the B motif RNA GCAGCCCCAGCCCCAGC-555 CCCUACCCCUGCCCCUGCCCCUGC was selected for its stable secondary structure and556 cytosine-rich patches. 68 We performed twenty independent AF3 runs using di ”erent random557 seeds and observed modest overall confidence (highest-confident mode: ipTM = 0.5 for the558 complex; average pLDDT scores = 66.22 for protein and 43.05 for the RNA). Particularly, the559 intrinsically disordered RG/RGG region (average pLDDT score = 40.71) and the KH3–RNA560 binding interface (average pLDDT scores = 81.99 for KH3 and 46.77 for the contacting RNA561 region) have low model confidence (Fig. 7).562 To focus on regions with more confident protein-RNA structural predictions, we trained563 the IRIS model using the KH1+KH2-B motif interface from the top-ranked predicted struc-564 ture. The resulting ϱ energy matrix was applied to three prediction tasks: binding a !nities565 for ten biologically relevant RNAs, for B-motif variants containing two cytosine patches, and566 for variants containing three cytosine patches.567 When target RNAs have a di ”erent length from the training sequence (such as biolog-568 ically relevant RNAs), we adopted a sliding-window approach using the training sequence.569 Specifically, when the target sequence is longer than the training sequence, we segmented570 the target sequence into a series of 41-bp segments from the 5 → to 3 → end in 1-nucleotide571 increments, and calculated the binding free energy for each of the target segments. If the572 target sequence is shorter than the training sequence, we scan the target sequence to replace573 a portion of the training sequence, incrementing one nucleotide at a time, thereby constitut-574 30 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint ing a set of new target sequences for binding free energy evaluation. For example, the 43 bp575 Sirloin motif yields 43 ↑ 41 + 1 = 3 overlapping 41 bp fragments, and the shorter 36 bp Ucp2576 sequence yields 41 ↑ 36 + 1 = 6 fragments within the same 41 bp framework. IRIS predicts577 a binding free energy for each fragment, and we report the lowest (i.e., strongest binding)578 value among them as the final predicted ##G for the full-length target RNA sequence. This579 approach ensures that IRIS identifies the highest-a!nity subsequence for each target.580 We also evaluated the model trained on the lower-confidence full-length hnRNPK-B-RNA581 complex and the KH1+KH2+RG/RGG-RNA complex. The results are reported in Figure582 S3-S5.583 Acknowledgment584 We are grateful to the Reformer authors for their assistance in using the pretrained hnRNPK585 model.586 Funding587 This work was supported by startup funding from North Carolina State University. Ad-588 ditional support was provided by the NC State Genetics and Genomics Academy and the589 Comparative Medicine Institute. E.C.R. acknowledges support by the GAANN Fellowship590 in Molecular Biotechnology at NC State University.591 Author Contributions592 Conceptualization: E.C.R, Y.Z., X.L.593 Methodology: E.C.R., Y.Z., X.L.594 Investigation: E.C.R., Y.Z., X.L.595 Visualization: E.C.R, Y.Z., X.L.596 31 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint Supervision: X.L.597 Writing—original draft: E.C.R., Y.Z., X.L.598 Writing—review & editing: E.C.R., Y.Z., X.L.599 Competing Interests600 The authors declare no competing interests.601 Data and Materials Availability602 The implementation of the IRIS model, along with training and prediction examples, is603 available at our GitHub repository . All training and testing datasets used in this study604 were collected from previously published articles and are publicly available from open-source605 repositories.606 32 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint References607 (1) Lee, Y.; Rio, D. C. Mechanisms and Regulation of Alternative Pre-mRNA Splicing.608 Annu Rev Biochem 2015, 84,2 9 1 – 3 2 3 .609 (2) Corley, M.; Burns, M. C.; Yeo, G. W. How RNA-Binding Proteins Interact with RNA:610 Molecules and Mechanisms. Molecular Cell 2020, 78,9 – 2 9 .611 (3) Ban, K.-Y.; Na, Y.-w.; Song, J.; Kim, J.-S.; Kim, J. Protein-RNA interaction dynamics612 reveal key regulators of oncogenic KRAS-driven cancers. Sci Rep 2024, 14,2 7 1 1 9 ,613 Publisher: Nature Publishing Group.614 (4) Rinn, J. L.; Chang, H. Y. Genome Regulation by Long Noncoding RNAs. Annu. Rev.615 Biochem. 2012, 81,1 4 5 – 1 6 6 .616 (5) Hentze, M. W.; Castello, A.; Schwarzl, T.; Preiss, T. A brave new world of RNA-617 binding proteins. Nat Rev Mol Cell Biol 2018, 19,3 2 7 – 3 4 1 .618 (6) Langdon, E. M.; Qiu, Y.; Ghanbari Niaki, A.; McLaughlin, G. A.; Weidmann, C. A.;619 Gerbich, T. M.; Smith, J. A.; Crutchley, J. M.; Termini, C. M.; Weeks, K. M.; My-620 ong, S.; Gladfelter, A. S. mRNA structure determines specificity of a polyQ-driven621 phase separation. Science 2018, 360,9 2 2 – 9 2 7 .622 (7) Zhang, V.; Gladfelter, A. S.; Roden, C. A. Biomolecular condensates: It was RNA all623 along! Molecular Cell 2025, 85, 461–463, Publisher: Elsevier BV.624 (8) Wiedner, H. J.; Giudice, J. It’s not just a phase: function and characteristics of625 RNA-binding proteins in phase separation. Nat Struct Mol Biol 2021, 28,4 6 5 – 4 7 3 ,626 Publisher: Nature Publishing Group.627 (9) Sanders, D. W. et al. Competing Protein-RNA Interaction Networks Control Multi-628 phase Intracellular Organization. Cell 2020, 181,3 0 6 – 3 2 4 . e 2 8 .629 33 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint (10) Ramanathan, M.; Porter, D. F.; Khavari, P. A. Methods to study RNA–protein inter-630 actions. Nat Methods 2019, 16, 225–234, Publisher: Nature Publishing Group.631 (11) Germer, K.; Leonard, M.; Zhang, X. RNA aptamers and their therapeutic and diag-632 nostic applications. Int J Biochem Mol Biol 2013, 4,2 7 – 4 0 .633 (12) Keefe, A. D.; Pai, S.; Ellington, A. Aptamers as therapeutics. Nat Rev Drug Discov634 2010, 9, 537–550, Publisher: Nature Publishing Group.635 (13) Zhu, Y.; Zhu, L.; Wang, X.; Jin, H. RNA-based therapeutics: an overview and prospec-636 tus. Cell Death Dis 2022, 13, 644, Publisher: Nature Publishing Group.637 (14) Gebauer, F.; Schwarzl, T.; Valc´ arcel, J.; Hentze, M. W. RNA-binding proteins in hu-638 man genetic disease. Nat Rev Genet 2021, 22, 185–198, Publisher: Nature Publishing639 Group.640 (15) Licatalosi, D. D.; Ye, X.; Jankowsky, E. Approaches for measuring the dynamics of641 RNA-protein interactions. Wiley Interdiscip Rev RNA 2020, 11,e 1 5 6 5 .642 (16) Mattay, J. Current Technical Approaches to Study RNA–Protein Interactions in mR-643 NAs and Long Non-Coding RNAs. BioChem 2023, 3, 1–14, Number: 1 Publisher:644 Multidisciplinary Digital Publishing Institute.645 (17) Lambert, N.; Robertson, A.; Jangi, M.; McGeary, S.; Sharp, P. A.; Burge, C. B.646 RNA Bind-n-Seq: Quantitative Assessment of the Sequence and Structural Binding647 Specificity of RNA Binding Proteins. Molecular Cell 2014, 54,8 8 7 – 9 0 0 ,P u b l i s h e r :648 Elsevier.649 (18) Jouravleva, K.; Vega-Badillo, J.; Zamore, P. D. Principles and pitfalls of high-650 throughput analysis of microRNA-binding thermodynamics and kinetics by RNA651 Bind-n-Seq. Cell Reports Methods 2022, 2,P u b l i s h e r : E l s e v i e r .652 34 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint (19) Ray, D.; Ha, K. C. H.; Nie, K.; Zheng, H.; Hughes, T. R.; Morris, Q. D. RNAcompete653 methodology and application to determine sequence preferences of unconventional654 RNA-binding proteins. Methods 2017, 118-119,3 – 1 5 .655 (20) Van Nostrand, E. L.; Pratt, G. A.; Shishkin, A. A.; Gelboin-Burkhart, C.; Fang, M. Y.;656 Sundararaman, B.; Blue, S. M.; Nguyen, T. B.; Surka, C.; Elkins, K.; Stanton, R.;657 Rigo, F.; Guttman, M.; Yeo, G. W. Robust transcriptome-wide discovery of RNA-658 binding protein binding sites with enhanced CLIP (eCLIP). Nat Methods 2016, 13,659 508–514, Publisher: Nature Publishing Group.660 (21) Huppertz, I.; Attig, J.; D’Ambrogio, A.; Easton, L. E.; Sibley, C. R.; Sugimoto, Y.;661 Tajnik, M.; K¨ onig, J.; Ule, J. iCLIP: protein-RNA interactions at nucleotide resolu-662 tion. Methods 2014, 65,2 7 4 – 2 8 7 .663 (22) Cordiner, R. A.; Dou, Y.; Thomsen, R.; Bugai, A.; Granneman, S.; Heick Jensen, T.664 Temporal-iCLIP captures co-transcriptional RNA-protein interactions. Nat Commun665 2023, 14, 696, Publisher: Nature Publishing Group.666 (23) Danan, C.; Manickavel, S.; Hafner, M. In Post-Transcriptional Gene Regulation ;667 Dassi, E., Ed.; Springer: New York, NY, 2016; pp 153–173.668 (24) Buenrostro, J. D.; Araya, C. L.; Chircus, L. M.; Layton, C. J.; Chang, H. Y.; Sny-669 der, M. P.; Greenleaf, W. J. Quantitative analysis of RNA-protein interactions on a670 massively parallel array reveals biophysical and evolutionary landscapes. Nat Biotech-671 nol 2014, 32,5 6 2 – 5 6 8 .672 (25) Martin, L.; Meier, M.; Lyons, S. M.; Sit, R. V.; Marzlu ”, W. F.; Quake, S. R.;673 Chang, H. Y. Systematic reconstruction of RNA functional motifs with high-674 throughput microfluidics. Nat Methods 2012, 9, 1192–1194, Publisher: Nature Pub-675 lishing Group.676 35 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint (26) Ridgeway, W. K.; Seitaridou, E.; Phillips, R.; Williamson, J. R. RNA–protein binding677 kinetics in an automated microfluidic reactor. Nucleic Acids Research 2009, 37,e 1 4 2 .678 (27) Hellman, L. M.; Fried, M. G. Electrophoretic mobility shift assay (EMSA) for detecting679 protein-nucleic acid interactions. Nat Protoc 2007, 2,1 8 4 9 – 1 8 6 1 .680 (28) Yang, Y.; Wang, Q.; Guo, D. A Novel Strategy for Analyzing RNA-Protein Interac-681 tions by Surface Plasmon Resonance Biosensor. Mol Biotechnol 2008, 40,8 7 – 9 3 .682 (29) Feig, A. L. Methods in Enzymology ;B i o p h y s i c a l ,C h e m i c a l ,a n dF u n c t i o n a lP r o b e so f683 RNA Structure, Interactions and Folding: Part A; Academic Press, 2009; Vol. 468; pp684 409–422.685 (30) Moon, M. H.; Hilimire, T. A.; Sanders, A. M.; Schneekloth, J. S. J. Measuring686 RNA–Ligand Interactions with Microscale Thermophoresis. Biochemistry 2018, 57,687 4638–4643, Publisher: American Chemical Society.688 (31) Rube, H. T.; Rastogi, C.; Feng, S.; Kribelbauer, J. F.; Li, A.; Becerra, B.; Melo, L.689 A. N.; Do, B. V.; Li, X.; Adam, H. H.; Shah, N. H.; Mann, R. S.; Bussemaker, H. J.690 Prediction of protein–ligand binding a! nity from sequencing data with interpretable691 machine learning. Nat Biotechnol 2022, 40,1 5 2 0 – 1 5 2 7 .692 (32) Feng, H.; Bao, S.; Rahman, M. A.; Weyn-Vanhentenryck, S. M.; Khan, A.; Wong, J.;693 Shah, A.; Flynn, E. D.; Krainer, A. R.; Zhang, C. Modeling RNA-Binding Protein694 Specificity In Vivo by Precisely Registering Protein-RNA Crosslink Sites. Molecular695 Cell 2019, 74,1 1 8 9 – 1 2 0 4 . e 6 .696 (33) Sugimoto, Y.; K¨ onig, J.; Hussain, S.; Zupan, B.; Curk, T.; Frye, M.; Ule, J. Anal-697 ysis of CLIP and iCLIP methods for nucleotide-resolution studies of protein-RNA698 interactions. Genome Biology 2012, 13,R 6 7 .699 36 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint (34) Hafner, M.; Landthaler, M.; Burger, L.; Khorshid, M.; Hausser, J.; Berninger, P.;700 Rothballer, A.; Ascano, M.; Jungkamp, A.-C.; Munschauer, M.; Ulrich, A.; War-701 dle, G. S.; Dewell, S.; Zavolan, M.; Tuschl, T. Transcriptome-wide Identification of702 RNA-Binding Protein and MicroRNA Target Sites by PAR-CLIP. Cell 2010, 141,703 129–141, Publisher: Elsevier.704 (35) Alipanahi, B.; Delong, A.; Weirauch, M. T.; Frey, B. J. Predicting the sequence speci-705 ficities of DNA- and RNA-binding proteins by deep learning. Nat Biotechnol 2015,706 33,8 3 1 – 8 3 8 .707 (36) Shen, X.; Hou, Y.; Wang, X.; Zhang, C.; Liu, J.; Shen, H.; Wang, W.; Yang, Y.;708 Yang, M.; Li, Y.; Zhang, J.; Sun, Y.; Chen, K.; Shi, L.; Li, X. A deep learning model709 for characterizing protein-RNA interactions from sequences at single-base resolution.710 Patterns 2025, 6,1 0 1 1 5 0 .711 (37) Mitra, R.; Cohen, A. S.; Sagendorf, J. M.; Berman, H. M.; Rohs, R. DNAproDB: an712 updated database for the automated and interactive analysis of protein-DNA com-713 plexes. Nucleic Acids Res 2025, 53,D 3 9 6 – D 4 0 2 .714 (38) Burley, S. et al. Updated resources for exploring experimentally-determined PDB715 structures and Computed Structure Models at the RCSB Protein Data Bank. Nu-716 cleic Acids Research 2025, 53,D 5 6 4 – D 5 7 4 .717 (39) Abramson, J. et al. Accurate structure prediction of biomolecular interactions with718 AlphaFold 3. Nature 2024,1 – 3 .719 (40) Baek, M.; McHugh, R.; Anishchenko, I.; Jiang, H.; Baker, D.; DiMaio, F. Accurate720 prediction of protein–nucleic acid complexes using RoseTTAFoldNA. Nat Methods721 2024, 21,1 1 7 – 1 2 1 .722 (41) Chai Discovery; Boitreaud, J.; Dent, J.; McPartlon, M.; Meier, J.; Reis, V.; Ro-723 37 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint gozhnikov, A.; Wu, K. Chai-1: Decoding the molecular interactions of life. 2024;724 http://biorxiv.org/lookup/doi/10.1101/2024.10.10.615955.725 (42) Wohlwend, J.; Corso, G.; Passaro, S.; Getz, N.; Reveiz, M.; Leidal, K.; Swiderski, W.;726 Atkinson, L.; Portnoi, T.; Chinn, I.; Silterra, J.; Jaakkola, T.; Barzilay, R. Boltz-1 De-727 mocratizing Biomolecular Interaction Modeling. 2024; http://biorxiv.org/lookup/728 doi/10.1101/2024.11.19.624167.729 (43) Passaro, S.; Corso, G.; Wohlwend, J.; Reveiz, M.; Thaler, S.; Somnath, V. R.; Getz, N.;730 Portnoi, T.; Roy, J.; Stark, H.; Kwabi-Addo, D.; Beaini, D.; Jaakkola, T.; Barzilay, R.731 Boltz-2: Towards Accurate and E! cient Binding A!nity Prediction. 2025; http:732 //biorxiv.org/lookup/doi/10.1101/2025.06.14.659707.733 (44) Xu, Y.; Zhu, J.; Huang, W.; Xu, K.; Yang, R.; Zhang, Q.; Sun, L. PrismNet: predicting734 protein–RNA interaction using in vivo RNA structural information. Nucleic Acids735 Research 2023, 51, W468–W477.736 (45) Zhu, H.; Yang, Y.; Wang, Y.; Wang, F.; Huang, Y.; Chang, Y.; Wong, K.-c.; Li, X. Dy-737 namic characterization and interpretation for protein-RNA interactions across diverse738 cellular conditions using HDRNet. Nat Commun 2023, 14, 6824, Publisher: Nature739 Publishing Group.740 (46) Kappel, K.; Jarmoskaite, I.; Vaidyanathan, P. P.; Greenleaf, W. J.; Herschlag, D.;741 Das, R. Blind tests of RNA-protein binding a !nity prediction. Proc Natl Acad Sci U742 SA 2019, 116,8 3 3 6 – 8 3 4 1 .743 (47) Zhang, Y.; Silvernail, I.; Lin, Z.; Lin, X. Interpretable Protein-DNA Interactions744 Captured by Structure-Sequence Optimization. 2025; https://elifesciences.org/745 reviewed-preprints/105565v2.746 (48) Berman, H. M. The Protein Data Bank. Nucleic Acids Research 2000, 28,2 3 5 – 2 4 2 .747 38 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint (49) Grahn, E.; Moss, T.; Helgstrand, C.; Fridborg, K.; Sundaram, M.; Tars, K.; Lago, H.;748 Stonehouse, N. J.; Davis, D. R.; Stockley, P. G.; Liljas, L. Structural basis of pyrimi-749 dine specificity in the MS2 RNA hairpin-coat-protein complex. RNA 2001, 7,1 6 1 6 –750 1627.751 (50) Auweter, S. D.; Fasan, R.; Reymond, L.; Underwood, J. G.; Black, D. L.; Pitsch, S.;752 Allain, F. H.-T. Molecular basis of RNA recognition by the human alternative splicing753 factor Fox-1. EMBO J 2006, 25,1 6 3 – 1 7 3 .754 (51) Wang, X.; McLachlan, J.; Zamore, P. D.; Hall, T. M. Modular Recognition of RNA755 by a Human Pumilio-Homology Domain. Cell 2002, 110,5 0 1 – 5 1 2 .756 (52) Batey, R. T.; Sagar, M.; Doudna, J. A. Structural and energetic analysis of RNA757 recognition by a universally conserved protein from the signal recognition particle.758 Journal of Molecular Biology 2001, 307,2 2 9 – 2 4 6 .759 (53) Oubridge, C.; Ito, N.; Evans, P. R.; Teo, C.-H.; Nagai, K. Crystal structure at 1.92760 ˚A resolution of the RNA-binding domain of the U1A spliceosomal protein complexed761 with an RNA hairpin. Nature 1994, 372,4 3 2 – 4 3 8 .762 (54) Humphrey, W.; Dalke, A.; Schulten, K. VMD – Visual Molecular Dynamics. Journal763 of Molecular Graphics 1996, 14,3 3 – 3 8 .764 (55) Ganser, L. R.; Kelly, M. L.; Herschlag, D.; Al-Hashimi, H. M. The roles of structural765 dynamics in the cellular functions of RNAs. Nat Rev Mol Cell Biol 2019, 20,4 7 4 – 4 8 9 .766 (56) ´Ottar Rolfsson; Middleton, S.; Manfield, I. W.; White, S. J.; Fan, B.; Vaughan, R.;767 Ranson, N. A.; Dykeman, E.; Twarock, R.; Ford, J.; Cheng Kao, C.; Stockley, P. G. Di-768 rect Evidence for Packaging Signal-Mediated Assembly of Bacteriophage MS2. Journal769 of Molecular Biology 2016, 428,4 3 1 – 4 4 8 .770 39 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint (57) Dai, X.; Li, Z.; Lai, M.; Shu, S.; Du, Y.; Zhou, Z. H.; Sun, R. In situ structures of771 the genome and genome-delivery apparatus in a single-stranded RNA virus. Nature772 2017, 541,1 1 2 – 1 1 6 .773 (58) Bukina, V.; Boˇ ziˇ c, A. Context-dependent structure formation of hairpin motifs in774 bacteriophage MS2 genomic RNA. Biophys J 2024, 123,3 3 9 7 – 3 4 0 7 .775 (59) Dominguez, D.; Freese, P.; Alexis, M. S.; Su, A.; Hochman, M.; Palden, T.; Bazile, C.;776 Lambert, N. J.; Van Nostrand, E. L.; Pratt, G. A.; Yeo, G. W.; Graveley, B. R.;777 Burge, C. B. Sequence, Structure, and Context Preferences of Human RNA Binding778 Proteins. Mol Cell 2018, 70,8 5 4 – 8 6 7 . e 9 .779 (60) Harris, S. E.; Alexis, M. S.; Giri, G.; Cavazos, F. F.; Hu, Y.; Murn, J.; Aleman, M. M.;780 Burge, C. B.; Dominguez, D. Understanding species-specific and conserved RNA-781 protein interactions in vivo and in vitro. Nat Commun 2024, 15,8 4 0 0 .782 (61) Harris, S. E.; Hu, Y.; Bridges, K.; Cavazos, F. F.; Martyr, J. G.; Guzm´ an, B. B.;783 Murn, J.; Aleman, M. M.; Dominguez, D. Dissecting RNA selectivity mediated by784 tandem RNA-binding domains. J Biol Chem 2025, 301,1 0 8 4 3 5 .785 (62) Hinkle, E. R.; Wiedner, H. J.; Torres, E. V.; Jackson, M.; Black, A. J.; Blue, R. E.;786 Harris, S. E.; Guzman, B. B.; Gentile, G. M.; Lee, E. Y.; Tsai, Y.-H.; Parker, J.;787 Dominguez, D.; Giudice, J. Alternative splicing regulation of membrane tra !cking788 genes during myogenesis. RNA 2022, 28,5 2 3 – 5 4 0 .789 (63) Gentile, G. M.; Blue, R. E.; Goda, G. A.; Guzman, B. B.; Szymanski, R. A.; Lee, E. Y.;790 Engels, N. M.; Hinkle, E. R.; Wiedner, H. J.; Bishop, A. N.; Harrison, J. T.; Zhang, H.;791 Wehrens, X. H. T.; Dominguez, D.; Giudice, J. Alternative splicing of the Snap23792 microexon is regulated by MBNL, QKI, and RBFOX2 in a tissue-specific manner and793 is altered in striated muscle diseases. RNA Biol 2025, 22,1 – 2 0 .794 40 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint (64) Ostareck, D. H.; Ostareck-Lederer, A.; Wilm, M.; Thiele, B. J.; Mann, M.;795 Hentze, M. W. mRNA silencing in erythroid di ”erentiation: hnRNP K and hnRNP796 E1 regulate 15-lipoxygenase translation from the 3’ end. Cell 1997, 89,5 9 7 – 6 0 6 .797 (65) Almeida, M.; Pintacuda, G.; Masui, O.; Koseki, Y.; Gdula, M.; Cerase, A.; Brown, D.;798 Mould, A.; Innocent, C.; Nakayama, M.; Schermelleh, L.; Nesterova, T. B.; Koseki, H.;799 Brockdor”, N. PCGF3/5-PRC1 initiates Polycomb recruitment in X chromosome in-800 activation. Science 2017, 356,1 0 8 1 – 1 0 8 4 .801 (66) Pintacuda, G.; Wei, G.; Roustan, C.; Kirmizitas, B. A.; Solcan, N.; Cerase, A.;802 Castello, A.; Mohammed, S.; Moindrot, B.; Nesterova, T. B.; Brockdor ”, N. hnRNPK803 Recruits PCGF3/5-PRC1 to the Xist RNA B-Repeat to Establish Polycomb-Mediated804 Chromosomal Silencing. Mol Cell 2017, 68,9 5 5 – 9 6 9 . e 1 0 .805 (67) Bomsztyk, K.; Denisenko, O.; Ostrowski, J. hnRNP K: one protein multiple processes.806 Bioessays 2004, 26,6 2 9 – 6 3 8 .807 (68) Nakamoto, M. Y.; Lammer, N. C.; Batey, R. T.; Wuttke, D. S. hnRNPK recognition808 of the B motif of Xist and other biological RNAs. Nucleic Acids Res 2020, 48,9 3 2 0 –809 9335.810 (69) Paziewska, A.; Wyrwicz, L. S.; Bujnicki, J. M.; Bomsztyk, K.; Ostrowski, J. Coopera-811 tive binding of the hnRNP K three KH domains to mRNA targets. FEBS Lett 2004,812 577,1 3 4 – 1 4 0 .813 (70) Klimek-Tomczak, K.; Wyrwicz, L. S.; Jain, S.; Bomsztyk, K.; Ostrowski, J. Character-814 ization of hnRNP K Protein–RNA Interactions. Journal of Molecular Biology 2004,815 342,1 1 3 1 – 1 1 4 1 .816 (71) Ozdilek, B. A.; Thompson, V. F.; Ahmed, N. S.; White, C. I.; Batey, R. T.;817 Schwartz, J. C. Intrinsically disordered RGG/RG domains mediate degenerate speci-818 ficity in RNA binding. Nucleic Acids Res 2017, 45,7 9 8 4 – 7 9 9 6 .819 41 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint (72) Rastogi, C.; Rube, H. T.; Kribelbauer, J. F.; Crocker, J.; Loker, R. E.; Martini, G. D.;820 Laptenko, O.; Freed-Pastor, W. A.; Prives, C.; Stern, D. L.; Mann, R. S.; Busse-821 maker, H. J. Accurate and sensitive quantification of protein-DNA binding a !nity.822 Proc. Natl. Acad. Sci. U.S.A. 2018, 115 .823 (73) Li, X.; Melo, L. A. N.; Bussemaker, H. J. Benchmarking and building DNA binding824 a!nity models using allele-specific and allele-agnostic transcription factor binding825 data. Genome Biol 2024, 25,2 8 4 .826 (74) Chu, W.-T.; Yan, Z.; Chu, X.; Zheng, X.; Liu, Z.; Xu, L.; Zhang, K.; Wang, J. Physics827 of biomolecular recognition and conformational dynamics. Rep. Prog. Phys. 2021, 84,828 126601.829 (75) Mitra, R.; Li, J.; Sagendorf, J. M.; Jiang, Y.; Cohen, A. S.; Chiu, T.-P.; Glass-830 cock, C. J.; Rohs, R. Geometric deep learning of protein–DNA binding specificity. Nat831

Methods

2024, 21,1 6 7 4 – 1 6 8 3 .832 (76) Liu, S.; Gomez-Alcala, P.; Leemans, C.; Glassford, W. J.; Melo, L. A. N.; Lu, X.-J.;833 Mann, R. S.; Bussemaker, H. J. Predicting the DNA binding specificity of transcription834 factor mutants using family-level biophysically interpretable machine learning. bioRxiv835 2025,2 0 2 4 . 0 1 . 2 4 . 5 7 7 1 1 5 .836 (77) Wetzel, J. L.; Zhang, K.; Singh, M. Learning probabilistic protein-DNA recognition837 codes from DNA-binding specificities using structural mappings. Genome Res 2022,838 32,1 7 7 6 – 1 7 8 6 .839 (78) Kretsch, R. C.; Hummer, A. M.; He, S.; Yuan, R.; Zhang, J.; Karagianes, T.; Cong, Q.;840 Kryshtafovych, A.; Das, R. Assessment of nucleic acid structure prediction in CASP16.841 2025; http://biorxiv.org/lookup/doi/10.1101/2025.05.06.652459.842 (79) Kaczmarek, J. C.; Kowalski, P. S.; Anderson, D. G. Advances in the delivery of RNA843 therapeutics: from concept to clinical reality. Genome Medicine 2017, 9,6 0 .844 42 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint (80) Ruscio, A. D.; Franciscis, V. d. Minding the gap: Unlocking the therapeutic potential845 of aptamers and making up for lost time. Molecular Therapy Nucleic Acids 2022, 29,846 384–386, Publisher: Elsevier.847 (81) Dias, R.; Kolazckowski, B. Di” erent combinations of atomic interactions predict848 protein-small molecule and protein-DNA/RNA a !nities with similar accuracy. Pro-849 teins 2015, 83,2 1 0 0 – 2 1 1 4 .850 (82) Leaver-Fay, A. et al. ROSETTA3: an object-oriented software suite for the simulation851 and design of macromolecules. Methods Enzymol 2011, 487,5 4 5 – 5 7 4 .852 (83) Shchepachev, V.; Bresson, S.; Spanos, C.; Petfalski, E.; Fischer, L.; Rappsilber, J.;853 Tollervey, D. Defining the RNA interactome by total RNA-associated protein purifi-854 cation. Molecular Systems Biology 2019, 15, e8689, Publisher: John Wiley & Sons,855 Ltd.856 (84) Kang, J. et al. RNAInter v4.0: RNA interactome repository with redefined confidence857 scoring system and improved accessibility. Nucleic Acids Research 2022, 50,D 3 2 6 –858 D332.859 (85) Caudron-Herger, M. Uncovering the complexity of RNA–protein interactions. Nat Rev860 Mol Cell Biol 2025, 26, 499–499, Publisher: Nature Publishing Group.861 (86) Cui, L.; Ma, R.; Cai, J.; Guo, C.; Chen, Z.; Yao, L.; Wang, Y.; Fan, R.; Wang, X.;862 Shi, Y. RNA modifications: importance in immune cell biology and related diseases.863 Sig Transduct Target Ther 2022, 7, 334, Publisher: Nature Publishing Group.864 (87) Delaunay, S.; Helm, M.; Frye, M. RNA modifications in physiology and disease: to-865 wards clinical applications. Nat Rev Genet 2024, 25, 104–122, Publisher: Nature866 Publishing Group.867 43 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint (88) Kar, M.; Dar, F.; Welsh, T. J.; Vogel, L. T.; K¨ uhnemuth, R.; Majumdar, A.;868 Krainer, G.; Franzmann, T. M.; Alberti, S.; Seidel, C. A. M.; Knowles, T. P. J.;869 Hyman, A. A.; Pappu, R. V. Phase-separating RNA-binding proteins form heteroge-870 neous distributions of clusters in subsaturated solutions. Proceedings of the National871 Academy of Sciences 2022, 119, e2202222119, Publisher: Proceedings of the National872 Academy of Sciences.873 (89) Phase-separation behaviour of RNAs. Nat. Chem. 2023, 15,1 6 6 0 – 1 6 6 1 ,P u b l i s h e r :874 Nature Publishing Group.875 (90) Stormo, G. D.; Zhao, Y. Determining the specificity of protein–DNA interactions. Nat876 Rev Genet 2010, 11,7 5 1 – 7 6 0 .877 (91) Bryngelson, J. D.; Wolynes, P. G. Spin glasses and the statistical mechanics of protein878 folding. Proc. Natl. Acad. Sci. U.S.A. 1987, 84,7 5 2 4 – 7 5 2 8 .879 (92) Davtyan, A.; Schafer, N. P.; Zheng, W.; Clementi, C.; Wolynes, P. G.; Papoian, G. A.880 A WSEM-MD: Protein Structure Prediction Using Coarse-Grained Physical Potentials881 and Bioinformatically Based Local Structure Biasing. J. Phys. Chem. B 2012, 116,882 8494–8503.883 (93) Schafer, N. P.; Kim, B. L.; Zheng, W.; Wolynes, P. G. Learning to fold proteins using884 energy landscape theory. Israel journal of chemistry 2014, 54,1 3 1 1 – 1 3 3 7 .885 (94) Lin, X.; George, J. T.; Schafer, N. P.; Ng Chau, K.; Birnbaum, M. E.; Clementi, C.;886 Onuchic, J. N.; Levine, H. Rapid assessment of T-cell receptor specificity of the im-887 mune repertoire. Nat Comput Sci 2021, 1,3 6 2 – 3 7 3 .888 (95) Wang, A.; Lin, X.; Chau, K. N.; Onuchic, J. N.; Levine, H.; George, J. T. RACER-m889 leverages structural features for sparse T cell specificity prediction. Sci. Adv. 2024,890 10,e a d l 0 1 6 1 .891 44 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint (96) Wagih, O. ggseqlogo: a versatile R package for drawing sequence logos. Bioinformatics892 2017, 33,3 6 4 5 – 3 6 4 7 .893 (97) van der Loo, M. P. The stringdist Package for Approximate String Matching. The R894 Journal 2014, 6,1 1 1 – 1 2 2 .895 (98) The UniProt Consortium et al. UniProt: the Universal Protein Knowledgebase in896 2025. Nucleic Acids Research 2025, 53,D 6 0 9 – D 6 1 7 .897 (99) Lorenz, R.; Bernhart, S. H.; H¨ oner zu Siederdissen, C.; Tafer, H.; Flamm, C.;898 Stadler, P. F.; Hofacker, I. L. ViennaRNA Package 2.0. Algorithms for Molecular899 Biology 2011, 6,2 6 .900 (100) Lyskov, S. et al. Serverification of Molecular Modeling Applications: The Rosetta901 Online Server That Includes Everyone (ROSIE). PLoS ONE 2013, 8,e 6 3 9 0 6 .902 45 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-pdf

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-26T02:00:01.498150+00:00
License: CC-BY-NC-ND-4.0