Materials
and Methods355
Protein–RNA Model Architecture and Training356
IRIS uses a residue-level biophysical framework that integrates detailed physicochemical357
modeling with structural data, either experimentally determined or predicted by structure-358
prediction tools, 39–42 to quantify sequence-specific protein-RNA binding a !nities (Figure 1).359
The model estimates relative binding free energies, ranks binding strengths, and converts360
these free energies into dissociation constants via the thermodynamic relation: 90361
##Gbinding = RT ln
(
#Kd
)
. (1)
This approach integrates structural and sequence signatures of the protein-RNA binding362
interface into an energy model to predict protein-RNA interactions with high precision.363
For training, IRIS identifies all residue pairs within the protein-RNA structural inter-364
face using C ε-P atom pairs that are closer than 0.95 nm in the training structures. The365
training sets include high-a!nity “native” binders and 10 , 000 randomized RNA sequence366
“decoy” binders. The integration of structural and sequence information of given protein-367
RNA complexes maximizes the usage of information encoded in the evolutionarily favorable368
21
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint
protein-RNA interfaces while reducing reliance on large-scale experimental binding sequence369
data.370
Each complex is represented by a 20-by-4 amino-acid-nucleotide interaction matrix that371
captures pairwise contact frequencies between each amino acid ( ai)a n dn u c l e o t i d e(nj )w i t h i n372
the protein-RNA interface, computed using the equation:373
ϑ(a, n)=
∑
i↑protein
∑
j↑RNA
$i,j (ri,j )( 2)
where $i,j (rij )c a nb ec a l c u l a t e dv i a :374
$i,j (ri,j )= 1
2
[
tanh
(
ϖ(r ↑ rmin)
)
· tanh
(
ϖ(rmax ↑ r)
)]
+ 1
2 (3)
Here, rmin = ↑ 9.5 ˚A,r max =9 .5 ˚
A,ϖ =0 .7. This defines two residues to be “in contact”
375
if they are separated less than 9.5 ˚
A. The parameter ϖ modulates the steepness of the
376
hyperbolic tangent function. i and j are sequence indices of the amino acid and nucleotide.377
The solvent-averaged binding energy between the protein and RNA is computed as the378
summed individual amino-acid-nucleotide pair energies at the protein-RNA interface:379
Ebinding =
∑
a↑protein,n↑RNA
ϱ(a, n)ϑ(a, n)( 4)
Here, ϱ(a, n)i sa2 0 - b y - 4e n e r g ym a t r i xe n c o d i n gr e s i d u e - t y p es p e c i fi ci n t e r a c t i o ns t r e n g t h380
between amino acid ( a) and nucleotide ( n). By defining binding free energy in this manner,381
our model captures sequence-specific protein-RNA binding while considering their structural382
context.383
The model optimizes ϱ(a, n)e n e r g ym a t r i xb ym a x i m i z i n gt h ee n e r g yg a pb e t w e e nt h e384
given strong binders and their corresponding decoy binders. This approach draws inspi-385
ration from methods used in protein folding, protein-protein and protein-DNA interaction386
studies.47,91–95 In practice, IRIS computes the average energy gap between the strong and387
decoy binders using ςE = ↔Edecoy↗↑↔Estrong↗,a n dt h es t a n d a r dd e v i a t i o no fd e c o yb i n d i n g388
22
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint
energies are computed by #E =S t d (Edecoy), where Std stands for the standard deviation.389
Using the interaction matrix (Eq. 2), we have390
A(a, n)= ↔ϑdecoy↗↑↔ϑstrong↗↘ R1↓300 (5)
391
B(a, n)=
⟨
ϑdecoy ϑT
decoy
⟩
↑↔ϑdecoy↗↔ϑT
decoy
↗↘ R300↓300 . (6)
and IRIS maximizes the separation of strong and decoy binding energies using:392
ςE
#E = Aϱ√
ϱT Bϱ
(7)
by optimizing393
R(ϱ)= AT ϱ ↑ φ
√
ϱT Bϱ . (8)
whose solution is ϱ ≃ B↔1 A.T o e n s u r e r o b u s t n e s s , w e r e g u l a r i z e t h e s o l u t i o n b y :( 1 ) d i a g o -394
nalizing B = P %P ↔1 ,( 2 )r e t a i n i n gt h et o p2 5e i g e n m o d e s ,( 3 )r e p l a c i n gs m a l l e re i g e n v a l u e s395
with the 25th largest value, and (4) reconstructing B↔1 from these filtered components.396
The optimized interaction energy matrix ϱ can be visualized in the 20 ⇐ 4 form (Figure 1),397
revealing the physicochemical preference at amino acid–nucleotide resolution.398
To predict the binding a !nity of a target protein-RNA sequence, we substitute it into399
the native high-a !nity protein-RNA complex structure and recalculate the target ϑ matrix400
using Eq. 2.T h i s i n t e r a c t i o n m a t r i x w i l l b e c o m b i n e d w i t h t h e t r a i n e d i n t e r a c t i o n e n e r g yϱ401
matrix to compute the target protein-RNA binding free energies using:402
Etarget = ϑT
target
ϱ. (9)
Since we focused our predictions on various RNA sequences binding to the same RNA-403
binding protein, we hypothesized that all of them share the same conformational entropy404
and approximated the di ”erence in binding free energy ##Gtarget to be the di” erence in405
23
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint
computed solvent-average binding free energy Etarget.T h i s h y p o t h e s i s w i l l n e e d t o b e r e v i s e d406
when considering plasticity in RNA structural ensemble, 55 especially for those RNAs with407
more drastic sequence mutations. Additionally, due to the presence of an undetermined408
scaling factor when optimizing the ϱ energy matrix, the predicted free energies are presented409
in reduced units. Despite this, the relative free energies can be used to accurately rank410
binding a!nities of target protein-RNA pairs, and the relative free energies can be converted411
to #Kd using Eq. 1. Additionally, the predicted binding free energies can be used to identify412
preferred RNA binding motifs.413
Predicting Binding A!nities of Mutant RNA Sequences414
As described in Section Protein–RNA Model Architecture and Training,m u t a n tb i n d i n g415
a!nities were predicted by substituting the native RNA sequence in the protein–RNA struc-416
ture with the target sequences. Binding free energies were then computed by combining the417
updated ϑ interaction matrix with the learned ϱ energy matrix using Eq. 4.T o a s s e s s p r e d i c -418
tive performance, mutant sequences were grouped by the number of nucleotide substitutions419
and by their proximity to the protein interface, defined using distance cuto”s ranging from420
0.3 to 0.95 nm.421
All MS2 testing sequences were obtained from Reference 24.O n l y s e q u e n c e s w i t h i n422
the measurable range of their RNA array were included, excluding those with maximum423
(##G =6 .661693623 kB T )o rm i n i m u m(# #G = ↑ 0.885494 kB T )v a l u e s . T h ec r y s t a l424
structure used in this study (PDB ID: 2C4Q) lacks nucleotides at both the 5 → and 3→ termini.425
For testing sequences carrying mutations at these missing p ositions (-15 at the 5 → end or426
the terminal 3 → nucleotide), the corresponding bases were omitted when constructing the427
ϑ contact matrix. This allowed such sequences to be included in Figure 2A, but did not428
influence a!nity predictions because these terminal nucleotides are absent in the 2C4Q429
training structure and therefore indistinguishable to the model.430
In addition to the native crystal structure, we also used the deep learning–based tool431
24
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint
AlphaFold3 (AF3)39 to generate protein–RNA complex structures for the mutant sequences.432
These predicted structures were then used to calculate the ϑ matrices and the corresponding433
binding free energies using Eqs. 2 and 4.434
Correlation and Linear Regression Analysis of Binding Free Ener-435
gies436
To assess predictive accuracy, we compared IRIS-predicted and exp erimentally measured437
relative binding free energies (## G) for MS2 protein binding on mutated RNA sequences.438
Sequences were grouped into two categories: those with 0–3 substitutions and those with four439
substitutions (4). Pearson’s r and Spearman’s ω were calculated to quantify linear agreement440
and rank consistency, respectively. For the 4-substitution group, correlations were computed441
using only quadruple mutants, with high-a!nity outliers retained in the analysis but omitted442
from the Sub 4 legend for clarity.443
A linear regression constrained through the origin (intercept = 0) was fitted to the 0–3444
substitution group to obtain a best-fit slope capturing the relationship between predicted445
and experimental binding energies (Figures 2Aa n d 3C). This regression line was applied446
as a common reference across both the 0–3 and the 4-substitution panels, enabling direct447
visual comparison between the two groups. High-a !nity outliers are indicated with purple448
triangles.449
All plots were generated with standardized axis limits and consistent color and shape450
encodings, facilitating visual comparison across groups. This design highlights both the451
overall predictive strength of the model and its limitations when applied to highly divergent452
sequences.453
25
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint
Binding Interface Cuto” Modulates Predictive Accuracy454
To evaluate how the definition of the binding interface influences predictive p erformance, we455
calculated Pearson correlations between predicted and experimental binding energies across456
residue distance cuto”s ranging from 0.3 to 0.95 nm. This analysis was performed on five pro-457
tein–RNA complexes with experimentally resolved structures: 2C4Q, 2ERR, 1IM8, 1HQ1,458
and 1URN (Figures 2Ba n d4 B). For each complex, predicted binding free energies were com-459
puted at five interface distance thresholds, and correlations with experimental values were460
determined. For MS2 (2C4Q), results were further stratified by the number of nucleotide461
substitutions (1–4) at each cuto”. The data were visualized as faceted bar plots, including a462
zero-correlation reference line to highlight thresholds where predictive accuracy diminishes463
or reverses. This analysis underscores the model’s sensitivity to interface definitions and464
helps identify optimal distance thresholds for accurate binding a !nity prediction.465
Binding Interface Motifs Across Distance Cuto”s466
To examine how the definition of the binding interface a ”ects conserved RNA-binding motifs,467
we analyzed the base composition of interface residues at di ”erent distance cuto”s (0.3, 0.4,468
and 0.5 nm) using sequence logo representations. This analysis was performed on the MS2469
protein–RNA complex, focusing on sequences with 1 and 4 substitutions (Figure 3D), as well470
as on test sequences from four additional protein–RNA complexes: 2ERR, 1IM8, 1HQ1, and471
1URN (Figure 4A).472
For each cuto”, RNA nucleotide identities were extracted using residue-level annotations473
and mapped to sequence positions. To avoid visualization redundancy and highlight the474
unique contributions of each cuto” , only residues within the shortest cuto” were retained.475
Nucleotide frequencies (A, C, G, U) were then computed to generate nucleotide distributions476
at each distance threshold.477
Stacked sequence logos were generated from these frequencies using the ggseqlogo package478
in R. 96 These visualizations illustrate how nucleotide enrichment patterns shift across dis-479
26
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint
tance cuto”s, revealing both shared and unique sequence preferences within the protein–RNA480
interface. This analysis provides insight into which distance thresholds capture meaningful481
RNA motifs, guiding the selection of high-a !nity binding motifs for target RNA-binding482
proteins (Figure 6).483
Total Residual Error Across Mutation Positions and Nucleotide484
Identities485
To investigate how specific nucleotide substitutions impact the predictive accuracy of MS2486
coat protein–RNA interactions, we computed the Total Residual Error stratified by both487
mutation position and nucleotide identity. All mutational variants and the wild-type RNA488
sequences used in Figure 1 were obtained and processed as described in Section Predicting489
Binding A”nities of Mutant RNA Sequences,w h e r eb i n d i n gf r e ee n e r g i e sw e r ep r e d i c t e d490
by substituting the native RNA sequence and computing ##G values from the ϑ and ϱ491
matrices.492
We leveraged the linear regression mo del describ ed in Section Correlation and Linear Re-493
gression Analysis of Binding Free Energies, which was fitted through the origin on sequences494
containing three or fewer substitutions ( Substitution ⇒ 3). This regression of predicted495
versus experimentally measured ##G values provides the expected predicted binding energy496
for a given experimental measurement under near-native mutational load, with the best-fit497
line shown in Figures 2Aa n d 3C.498
For each sequence, the residual error was calculated as the absolute di ”erence between499
the predicted binding energy and the value projected by the regression model:500
Residual Error = |##Gpredicted ↑ ##Gbest-fit|, (10)
where ##Gbest-fit = m · ##Gexperimental and m is the slope of the fitted linear regression.501
Residual errors were then aggregated across all sequences, grouped by mutation position502
27
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint
and substituted nucleotide (A, C, G, or U). This yields a total residual error for each nu-503
cleotide at each position. The results are visualized as a stacked bar plot (Fig. 3B), with504
bars colored by nucleotide identity and a fixed y-axis for consistent scaling. This representa-505
tion highlights positions and substitutions that disproportionately contribute to prediction506
errors, revealing model blind spots and guiding potential refinement.507
Jaccard Similarity Heatmap of RNA Sequences Against Known508
MS2 Motifs509
To compare structural similarity between input RNA sequences and known RNA-binding510
motifs, we generated a Jaccard similarity heatmap using a k-mer–based approach. For each511
pairwise comparison between a test RNA sequence and a motif, the Jaccard similarity was512
computed using all unique k-mers of length 3 ( k =3 ) ,d e fi n e dm a t h e m a t i c a l l ya s :513
J(A, B)= |K(A) ⇑ K(B)|
|K(A) ⇓ K(B)| , (11)
where K(A)a n d K(B)a r et h es e t so fk - m e r si ns e q u e n c e sA and B, respectively.514
To compare similarity among RNA sequences, we selected 10 RNA sequences from Ref-515
erence 24,i n c l u d i n gt h en a t i v es e q u e n c ei nt h ec r y s t a ls t r u c t u r e( P D BI D :2 C 4 Q ) ,7s t r o n g -516
binding 4-mutation RNA sequences (B1–B7), and two additional RNA mutants (R1 and517
R2). Experimentally identified MS2-RNA binding motifs from prior CLIP-seq and CryoEM518
studies56–58 (Motifs I–XIV) were also included. Each sequence was converted into a set of519
3-mers, and pairwise Jaccard similarities were calculated between the two groups.520
The resulting similarity matrix was visualized as a heatmap, with experimentally iden-521
tified motifs on the y-axis and RNA sequences from Reference 24 on the x-axis (Fig. 5A).522
This k-mer–level analysis enabled rapid structural comparison of short RNA sequences and523
facilitated identification of which motifs each test sequence most closely resembled. Because524
this approach is alignment-free, it is robust to minor shifts or indels, making it particularly525
28
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint
well-suited for comparing short, structured RNAs.526
K-mer Network Construction and Clustering of High-A!nity RNA527
Fra gments528
To identify and visualize structure-informed patterns of RNA fragments associated with high529
protein-binding a!nity, we extracted contiguous runs of residue IDs and their corresponding530
nucleotide identities from all AF3-predicted RNA-binding structures. Blocks of protein-531
interacting RNA residues were defined under varying interaction distance cuto ”s, and for532
each residue, we recorded the lowest cuto” at which it remained in contact with the protein.533
Contiguous segments with more than four residues were further divided into overlapping534
sub-k-mers (e.g., 4-mers, 5-mers), with each k-mer annotated with its sequence, length,535
per-position cuto ” values, and the experimentally measured ##Gexp of the parent RNA536
sequence. To standardize comparisons across the dataset, ##Gexp values were converted537
into Z-scores:538
Z = ##Gexp ↑ µ!!Gexp
↼!!Gexp
,
where µ!!Gexp and ↼!!Gexp are the mean and standard deviation of all ##Gexp val-539
ues for the target RNA-binding protein, calculated across the full dataset prior to k-mer540
segmentation to avoid bias from fragment counts.541
After collecting all k-mers, Z-scores were aggregated for each unique sequence. We542
then computed a pairwise string similarity matrix using the Levenshtein distance via the543
stringdist R package. 97 This matrix served as input for agglomerative hierarchical clus-544
tering with average linkage (implemented using hclust in R). The resulting dendrogram545
(Fig. 6) highlights relationships among k-mers based on sequence similarity, revealing dis-546
tinct sequence families. This approach uncovers recurring structural and sequence features547
of RNA that strongly bind proteins and provides guidance for designing high-a !nity RNA548
29
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint
sequences.549
AF3-Predicted HnRNPK–B Motif Structures and IRIS A!nity550
Prediction551
In the absence of experimentally determined hnRNPK-B motif complexes, we used AF3 39552
to generate predicted structures for IRIS model training and binding-a !nity prediction.553
The hnRNPK protein sequence was obtained from Uniprot 98 (https://www.uniprot.org/554
uniprotkb/P61978/entry#sequences), and the B motif RNA GCAGCCCCAGCCCCAGC-555
CCCUACCCCUGCCCCUGCCCCUGC was selected for its stable secondary structure and556
cytosine-rich patches. 68 We performed twenty independent AF3 runs using di ”erent random557
seeds and observed modest overall confidence (highest-confident mode: ipTM = 0.5 for the558
complex; average pLDDT scores = 66.22 for protein and 43.05 for the RNA). Particularly, the559
intrinsically disordered RG/RGG region (average pLDDT score = 40.71) and the KH3–RNA560
binding interface (average pLDDT scores = 81.99 for KH3 and 46.77 for the contacting RNA561
region) have low model confidence (Fig. 7).562
To focus on regions with more confident protein-RNA structural predictions, we trained563
the IRIS model using the KH1+KH2-B motif interface from the top-ranked predicted struc-564
ture. The resulting ϱ energy matrix was applied to three prediction tasks: binding a !nities565
for ten biologically relevant RNAs, for B-motif variants containing two cytosine patches, and566
for variants containing three cytosine patches.567
When target RNAs have a di ”erent length from the training sequence (such as biolog-568
ically relevant RNAs), we adopted a sliding-window approach using the training sequence.569
Specifically, when the target sequence is longer than the training sequence, we segmented570
the target sequence into a series of 41-bp segments from the 5 → to 3 → end in 1-nucleotide571
increments, and calculated the binding free energy for each of the target segments. If the572
target sequence is shorter than the training sequence, we scan the target sequence to replace573
a portion of the training sequence, incrementing one nucleotide at a time, thereby constitut-574
30
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint
ing a set of new target sequences for binding free energy evaluation. For example, the 43 bp575
Sirloin motif yields 43 ↑ 41 + 1 = 3 overlapping 41 bp fragments, and the shorter 36 bp Ucp2576
sequence yields 41 ↑ 36 + 1 = 6 fragments within the same 41 bp framework. IRIS predicts577
a binding free energy for each fragment, and we report the lowest (i.e., strongest binding)578
value among them as the final predicted ##G for the full-length target RNA sequence. This579
approach ensures that IRIS identifies the highest-a!nity subsequence for each target.580
We also evaluated the model trained on the lower-confidence full-length hnRNPK-B-RNA581
complex and the KH1+KH2+RG/RGG-RNA complex. The results are reported in Figure582
S3-S5.583
Acknowledgment584
We are grateful to the Reformer authors for their assistance in using the pretrained hnRNPK585
model.586
Funding587
This work was supported by startup funding from North Carolina State University. Ad-588
ditional support was provided by the NC State Genetics and Genomics Academy and the589
Comparative Medicine Institute. E.C.R. acknowledges support by the GAANN Fellowship590
in Molecular Biotechnology at NC State University.591
Author Contributions592
Conceptualization: E.C.R, Y.Z., X.L.593
Methodology: E.C.R., Y.Z., X.L.594
Investigation: E.C.R., Y.Z., X.L.595
Visualization: E.C.R, Y.Z., X.L.596
31
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint
Supervision: X.L.597
Writing—original draft: E.C.R., Y.Z., X.L.598
Writing—review & editing: E.C.R., Y.Z., X.L.599
Competing Interests600
The authors declare no competing interests.601
Data and Materials Availability602
The implementation of the IRIS model, along with training and prediction examples, is603
available at our GitHub repository . All training and testing datasets used in this study604
were collected from previously published articles and are publicly available from open-source605
repositories.606
32
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint
References607
(1) Lee, Y.; Rio, D. C. Mechanisms and Regulation of Alternative Pre-mRNA Splicing.608
Annu Rev Biochem 2015, 84,2 9 1 – 3 2 3 .609
(2) Corley, M.; Burns, M. C.; Yeo, G. W. How RNA-Binding Proteins Interact with RNA:610
Molecules and Mechanisms. Molecular Cell 2020, 78,9 – 2 9 .611
(3) Ban, K.-Y.; Na, Y.-w.; Song, J.; Kim, J.-S.; Kim, J. Protein-RNA interaction dynamics612
reveal key regulators of oncogenic KRAS-driven cancers. Sci Rep 2024, 14,2 7 1 1 9 ,613
Publisher: Nature Publishing Group.614
(4) Rinn, J. L.; Chang, H. Y. Genome Regulation by Long Noncoding RNAs. Annu. Rev.615
Biochem. 2012, 81,1 4 5 – 1 6 6 .616
(5) Hentze, M. W.; Castello, A.; Schwarzl, T.; Preiss, T. A brave new world of RNA-617
binding proteins. Nat Rev Mol Cell Biol 2018, 19,3 2 7 – 3 4 1 .618
(6) Langdon, E. M.; Qiu, Y.; Ghanbari Niaki, A.; McLaughlin, G. A.; Weidmann, C. A.;619
Gerbich, T. M.; Smith, J. A.; Crutchley, J. M.; Termini, C. M.; Weeks, K. M.; My-620
ong, S.; Gladfelter, A. S. mRNA structure determines specificity of a polyQ-driven621
phase separation. Science 2018, 360,9 2 2 – 9 2 7 .622
(7) Zhang, V.; Gladfelter, A. S.; Roden, C. A. Biomolecular condensates: It was RNA all623
along! Molecular Cell 2025, 85, 461–463, Publisher: Elsevier BV.624
(8) Wiedner, H. J.; Giudice, J. It’s not just a phase: function and characteristics of625
RNA-binding proteins in phase separation. Nat Struct Mol Biol 2021, 28,4 6 5 – 4 7 3 ,626
Publisher: Nature Publishing Group.627
(9) Sanders, D. W. et al. Competing Protein-RNA Interaction Networks Control Multi-628
phase Intracellular Organization. Cell 2020, 181,3 0 6 – 3 2 4 . e 2 8 .629
33
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint
(10) Ramanathan, M.; Porter, D. F.; Khavari, P. A. Methods to study RNA–protein inter-630
actions. Nat Methods 2019, 16, 225–234, Publisher: Nature Publishing Group.631
(11) Germer, K.; Leonard, M.; Zhang, X. RNA aptamers and their therapeutic and diag-632
nostic applications. Int J Biochem Mol Biol 2013, 4,2 7 – 4 0 .633
(12) Keefe, A. D.; Pai, S.; Ellington, A. Aptamers as therapeutics. Nat Rev Drug Discov634
2010, 9, 537–550, Publisher: Nature Publishing Group.635
(13) Zhu, Y.; Zhu, L.; Wang, X.; Jin, H. RNA-based therapeutics: an overview and prospec-636
tus. Cell Death Dis 2022, 13, 644, Publisher: Nature Publishing Group.637
(14) Gebauer, F.; Schwarzl, T.; Valc´ arcel, J.; Hentze, M. W. RNA-binding proteins in hu-638
man genetic disease. Nat Rev Genet 2021, 22, 185–198, Publisher: Nature Publishing639
Group.640
(15) Licatalosi, D. D.; Ye, X.; Jankowsky, E. Approaches for measuring the dynamics of641
RNA-protein interactions. Wiley Interdiscip Rev RNA 2020, 11,e 1 5 6 5 .642
(16) Mattay, J. Current Technical Approaches to Study RNA–Protein Interactions in mR-643
NAs and Long Non-Coding RNAs. BioChem 2023, 3, 1–14, Number: 1 Publisher:644
Multidisciplinary Digital Publishing Institute.645
(17) Lambert, N.; Robertson, A.; Jangi, M.; McGeary, S.; Sharp, P. A.; Burge, C. B.646
RNA Bind-n-Seq: Quantitative Assessment of the Sequence and Structural Binding647
Specificity of RNA Binding Proteins. Molecular Cell 2014, 54,8 8 7 – 9 0 0 ,P u b l i s h e r :648
Elsevier.649
(18) Jouravleva, K.; Vega-Badillo, J.; Zamore, P. D. Principles and pitfalls of high-650
throughput analysis of microRNA-binding thermodynamics and kinetics by RNA651
Bind-n-Seq. Cell Reports Methods 2022, 2,P u b l i s h e r : E l s e v i e r .652
34
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint
(19) Ray, D.; Ha, K. C. H.; Nie, K.; Zheng, H.; Hughes, T. R.; Morris, Q. D. RNAcompete653
methodology and application to determine sequence preferences of unconventional654
RNA-binding proteins. Methods 2017, 118-119,3 – 1 5 .655
(20) Van Nostrand, E. L.; Pratt, G. A.; Shishkin, A. A.; Gelboin-Burkhart, C.; Fang, M. Y.;656
Sundararaman, B.; Blue, S. M.; Nguyen, T. B.; Surka, C.; Elkins, K.; Stanton, R.;657
Rigo, F.; Guttman, M.; Yeo, G. W. Robust transcriptome-wide discovery of RNA-658
binding protein binding sites with enhanced CLIP (eCLIP). Nat Methods 2016, 13,659
508–514, Publisher: Nature Publishing Group.660
(21) Huppertz, I.; Attig, J.; D’Ambrogio, A.; Easton, L. E.; Sibley, C. R.; Sugimoto, Y.;661
Tajnik, M.; K¨ onig, J.; Ule, J. iCLIP: protein-RNA interactions at nucleotide resolu-662
tion. Methods 2014, 65,2 7 4 – 2 8 7 .663
(22) Cordiner, R. A.; Dou, Y.; Thomsen, R.; Bugai, A.; Granneman, S.; Heick Jensen, T.664
Temporal-iCLIP captures co-transcriptional RNA-protein interactions. Nat Commun665
2023, 14, 696, Publisher: Nature Publishing Group.666
(23) Danan, C.; Manickavel, S.; Hafner, M. In Post-Transcriptional Gene Regulation ;667
Dassi, E., Ed.; Springer: New York, NY, 2016; pp 153–173.668
(24) Buenrostro, J. D.; Araya, C. L.; Chircus, L. M.; Layton, C. J.; Chang, H. Y.; Sny-669
der, M. P.; Greenleaf, W. J. Quantitative analysis of RNA-protein interactions on a670
massively parallel array reveals biophysical and evolutionary landscapes. Nat Biotech-671
nol 2014, 32,5 6 2 – 5 6 8 .672
(25) Martin, L.; Meier, M.; Lyons, S. M.; Sit, R. V.; Marzlu ”, W. F.; Quake, S. R.;673
Chang, H. Y. Systematic reconstruction of RNA functional motifs with high-674
throughput microfluidics. Nat Methods 2012, 9, 1192–1194, Publisher: Nature Pub-675
lishing Group.676
35
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint
(26) Ridgeway, W. K.; Seitaridou, E.; Phillips, R.; Williamson, J. R. RNA–protein binding677
kinetics in an automated microfluidic reactor. Nucleic Acids Research 2009, 37,e 1 4 2 .678
(27) Hellman, L. M.; Fried, M. G. Electrophoretic mobility shift assay (EMSA) for detecting679
protein-nucleic acid interactions. Nat Protoc 2007, 2,1 8 4 9 – 1 8 6 1 .680
(28) Yang, Y.; Wang, Q.; Guo, D. A Novel Strategy for Analyzing RNA-Protein Interac-681
tions by Surface Plasmon Resonance Biosensor. Mol Biotechnol 2008, 40,8 7 – 9 3 .682
(29) Feig, A. L. Methods in Enzymology ;B i o p h y s i c a l ,C h e m i c a l ,a n dF u n c t i o n a lP r o b e so f683
RNA Structure, Interactions and Folding: Part A; Academic Press, 2009; Vol. 468; pp684
409–422.685
(30) Moon, M. H.; Hilimire, T. A.; Sanders, A. M.; Schneekloth, J. S. J. Measuring686
RNA–Ligand Interactions with Microscale Thermophoresis. Biochemistry 2018, 57,687
4638–4643, Publisher: American Chemical Society.688
(31) Rube, H. T.; Rastogi, C.; Feng, S.; Kribelbauer, J. F.; Li, A.; Becerra, B.; Melo, L.689
A. N.; Do, B. V.; Li, X.; Adam, H. H.; Shah, N. H.; Mann, R. S.; Bussemaker, H. J.690
Prediction of protein–ligand binding a! nity from sequencing data with interpretable691
machine learning. Nat Biotechnol 2022, 40,1 5 2 0 – 1 5 2 7 .692
(32) Feng, H.; Bao, S.; Rahman, M. A.; Weyn-Vanhentenryck, S. M.; Khan, A.; Wong, J.;693
Shah, A.; Flynn, E. D.; Krainer, A. R.; Zhang, C. Modeling RNA-Binding Protein694
Specificity In Vivo by Precisely Registering Protein-RNA Crosslink Sites. Molecular695
Cell 2019, 74,1 1 8 9 – 1 2 0 4 . e 6 .696
(33) Sugimoto, Y.; K¨ onig, J.; Hussain, S.; Zupan, B.; Curk, T.; Frye, M.; Ule, J. Anal-697
ysis of CLIP and iCLIP methods for nucleotide-resolution studies of protein-RNA698
interactions. Genome Biology 2012, 13,R 6 7 .699
36
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint
(34) Hafner, M.; Landthaler, M.; Burger, L.; Khorshid, M.; Hausser, J.; Berninger, P.;700
Rothballer, A.; Ascano, M.; Jungkamp, A.-C.; Munschauer, M.; Ulrich, A.; War-701
dle, G. S.; Dewell, S.; Zavolan, M.; Tuschl, T. Transcriptome-wide Identification of702
RNA-Binding Protein and MicroRNA Target Sites by PAR-CLIP. Cell 2010, 141,703
129–141, Publisher: Elsevier.704
(35) Alipanahi, B.; Delong, A.; Weirauch, M. T.; Frey, B. J. Predicting the sequence speci-705
ficities of DNA- and RNA-binding proteins by deep learning. Nat Biotechnol 2015,706
33,8 3 1 – 8 3 8 .707
(36) Shen, X.; Hou, Y.; Wang, X.; Zhang, C.; Liu, J.; Shen, H.; Wang, W.; Yang, Y.;708
Yang, M.; Li, Y.; Zhang, J.; Sun, Y.; Chen, K.; Shi, L.; Li, X. A deep learning model709
for characterizing protein-RNA interactions from sequences at single-base resolution.710
Patterns 2025, 6,1 0 1 1 5 0 .711
(37) Mitra, R.; Cohen, A. S.; Sagendorf, J. M.; Berman, H. M.; Rohs, R. DNAproDB: an712
updated database for the automated and interactive analysis of protein-DNA com-713
plexes. Nucleic Acids Res 2025, 53,D 3 9 6 – D 4 0 2 .714
(38) Burley, S. et al. Updated resources for exploring experimentally-determined PDB715
structures and Computed Structure Models at the RCSB Protein Data Bank. Nu-716
cleic Acids Research 2025, 53,D 5 6 4 – D 5 7 4 .717
(39) Abramson, J. et al. Accurate structure prediction of biomolecular interactions with718
AlphaFold 3. Nature 2024,1 – 3 .719
(40) Baek, M.; McHugh, R.; Anishchenko, I.; Jiang, H.; Baker, D.; DiMaio, F. Accurate720
prediction of protein–nucleic acid complexes using RoseTTAFoldNA. Nat Methods721
2024, 21,1 1 7 – 1 2 1 .722
(41) Chai Discovery; Boitreaud, J.; Dent, J.; McPartlon, M.; Meier, J.; Reis, V.; Ro-723
37
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint
gozhnikov, A.; Wu, K. Chai-1: Decoding the molecular interactions of life. 2024;724
http://biorxiv.org/lookup/doi/10.1101/2024.10.10.615955.725
(42) Wohlwend, J.; Corso, G.; Passaro, S.; Getz, N.; Reveiz, M.; Leidal, K.; Swiderski, W.;726
Atkinson, L.; Portnoi, T.; Chinn, I.; Silterra, J.; Jaakkola, T.; Barzilay, R. Boltz-1 De-727
mocratizing Biomolecular Interaction Modeling. 2024; http://biorxiv.org/lookup/728
doi/10.1101/2024.11.19.624167.729
(43) Passaro, S.; Corso, G.; Wohlwend, J.; Reveiz, M.; Thaler, S.; Somnath, V. R.; Getz, N.;730
Portnoi, T.; Roy, J.; Stark, H.; Kwabi-Addo, D.; Beaini, D.; Jaakkola, T.; Barzilay, R.731
Boltz-2: Towards Accurate and E! cient Binding A!nity Prediction. 2025; http:732
//biorxiv.org/lookup/doi/10.1101/2025.06.14.659707.733
(44) Xu, Y.; Zhu, J.; Huang, W.; Xu, K.; Yang, R.; Zhang, Q.; Sun, L. PrismNet: predicting734
protein–RNA interaction using in vivo RNA structural information. Nucleic Acids735
Research 2023, 51, W468–W477.736
(45) Zhu, H.; Yang, Y.; Wang, Y.; Wang, F.; Huang, Y.; Chang, Y.; Wong, K.-c.; Li, X. Dy-737
namic characterization and interpretation for protein-RNA interactions across diverse738
cellular conditions using HDRNet. Nat Commun 2023, 14, 6824, Publisher: Nature739
Publishing Group.740
(46) Kappel, K.; Jarmoskaite, I.; Vaidyanathan, P. P.; Greenleaf, W. J.; Herschlag, D.;741
Das, R. Blind tests of RNA-protein binding a !nity prediction. Proc Natl Acad Sci U742
SA 2019, 116,8 3 3 6 – 8 3 4 1 .743
(47) Zhang, Y.; Silvernail, I.; Lin, Z.; Lin, X. Interpretable Protein-DNA Interactions744
Captured by Structure-Sequence Optimization. 2025; https://elifesciences.org/745
reviewed-preprints/105565v2.746
(48) Berman, H. M. The Protein Data Bank. Nucleic Acids Research 2000, 28,2 3 5 – 2 4 2 .747
38
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint
(49) Grahn, E.; Moss, T.; Helgstrand, C.; Fridborg, K.; Sundaram, M.; Tars, K.; Lago, H.;748
Stonehouse, N. J.; Davis, D. R.; Stockley, P. G.; Liljas, L. Structural basis of pyrimi-749
dine specificity in the MS2 RNA hairpin-coat-protein complex. RNA 2001, 7,1 6 1 6 –750
1627.751
(50) Auweter, S. D.; Fasan, R.; Reymond, L.; Underwood, J. G.; Black, D. L.; Pitsch, S.;752
Allain, F. H.-T. Molecular basis of RNA recognition by the human alternative splicing753
factor Fox-1. EMBO J 2006, 25,1 6 3 – 1 7 3 .754
(51) Wang, X.; McLachlan, J.; Zamore, P. D.; Hall, T. M. Modular Recognition of RNA755
by a Human Pumilio-Homology Domain. Cell 2002, 110,5 0 1 – 5 1 2 .756
(52) Batey, R. T.; Sagar, M.; Doudna, J. A. Structural and energetic analysis of RNA757
recognition by a universally conserved protein from the signal recognition particle.758
Journal of Molecular Biology 2001, 307,2 2 9 – 2 4 6 .759
(53) Oubridge, C.; Ito, N.; Evans, P. R.; Teo, C.-H.; Nagai, K. Crystal structure at 1.92760
˚A resolution of the RNA-binding domain of the U1A spliceosomal protein complexed761
with an RNA hairpin. Nature 1994, 372,4 3 2 – 4 3 8 .762
(54) Humphrey, W.; Dalke, A.; Schulten, K. VMD – Visual Molecular Dynamics. Journal763
of Molecular Graphics 1996, 14,3 3 – 3 8 .764
(55) Ganser, L. R.; Kelly, M. L.; Herschlag, D.; Al-Hashimi, H. M. The roles of structural765
dynamics in the cellular functions of RNAs. Nat Rev Mol Cell Biol 2019, 20,4 7 4 – 4 8 9 .766
(56) ´Ottar Rolfsson; Middleton, S.; Manfield, I. W.; White, S. J.; Fan, B.; Vaughan, R.;767
Ranson, N. A.; Dykeman, E.; Twarock, R.; Ford, J.; Cheng Kao, C.; Stockley, P. G. Di-768
rect Evidence for Packaging Signal-Mediated Assembly of Bacteriophage MS2. Journal769
of Molecular Biology 2016, 428,4 3 1 – 4 4 8 .770
39
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint
(57) Dai, X.; Li, Z.; Lai, M.; Shu, S.; Du, Y.; Zhou, Z. H.; Sun, R. In situ structures of771
the genome and genome-delivery apparatus in a single-stranded RNA virus. Nature772
2017, 541,1 1 2 – 1 1 6 .773
(58) Bukina, V.; Boˇ ziˇ c, A. Context-dependent structure formation of hairpin motifs in774
bacteriophage MS2 genomic RNA. Biophys J 2024, 123,3 3 9 7 – 3 4 0 7 .775
(59) Dominguez, D.; Freese, P.; Alexis, M. S.; Su, A.; Hochman, M.; Palden, T.; Bazile, C.;776
Lambert, N. J.; Van Nostrand, E. L.; Pratt, G. A.; Yeo, G. W.; Graveley, B. R.;777
Burge, C. B. Sequence, Structure, and Context Preferences of Human RNA Binding778
Proteins. Mol Cell 2018, 70,8 5 4 – 8 6 7 . e 9 .779
(60) Harris, S. E.; Alexis, M. S.; Giri, G.; Cavazos, F. F.; Hu, Y.; Murn, J.; Aleman, M. M.;780
Burge, C. B.; Dominguez, D. Understanding species-specific and conserved RNA-781
protein interactions in vivo and in vitro. Nat Commun 2024, 15,8 4 0 0 .782
(61) Harris, S. E.; Hu, Y.; Bridges, K.; Cavazos, F. F.; Martyr, J. G.; Guzm´ an, B. B.;783
Murn, J.; Aleman, M. M.; Dominguez, D. Dissecting RNA selectivity mediated by784
tandem RNA-binding domains. J Biol Chem 2025, 301,1 0 8 4 3 5 .785
(62) Hinkle, E. R.; Wiedner, H. J.; Torres, E. V.; Jackson, M.; Black, A. J.; Blue, R. E.;786
Harris, S. E.; Guzman, B. B.; Gentile, G. M.; Lee, E. Y.; Tsai, Y.-H.; Parker, J.;787
Dominguez, D.; Giudice, J. Alternative splicing regulation of membrane tra !cking788
genes during myogenesis. RNA 2022, 28,5 2 3 – 5 4 0 .789
(63) Gentile, G. M.; Blue, R. E.; Goda, G. A.; Guzman, B. B.; Szymanski, R. A.; Lee, E. Y.;790
Engels, N. M.; Hinkle, E. R.; Wiedner, H. J.; Bishop, A. N.; Harrison, J. T.; Zhang, H.;791
Wehrens, X. H. T.; Dominguez, D.; Giudice, J. Alternative splicing of the Snap23792
microexon is regulated by MBNL, QKI, and RBFOX2 in a tissue-specific manner and793
is altered in striated muscle diseases. RNA Biol 2025, 22,1 – 2 0 .794
40
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint
(64) Ostareck, D. H.; Ostareck-Lederer, A.; Wilm, M.; Thiele, B. J.; Mann, M.;795
Hentze, M. W. mRNA silencing in erythroid di ”erentiation: hnRNP K and hnRNP796
E1 regulate 15-lipoxygenase translation from the 3’ end. Cell 1997, 89,5 9 7 – 6 0 6 .797
(65) Almeida, M.; Pintacuda, G.; Masui, O.; Koseki, Y.; Gdula, M.; Cerase, A.; Brown, D.;798
Mould, A.; Innocent, C.; Nakayama, M.; Schermelleh, L.; Nesterova, T. B.; Koseki, H.;799
Brockdor”, N. PCGF3/5-PRC1 initiates Polycomb recruitment in X chromosome in-800
activation. Science 2017, 356,1 0 8 1 – 1 0 8 4 .801
(66) Pintacuda, G.; Wei, G.; Roustan, C.; Kirmizitas, B. A.; Solcan, N.; Cerase, A.;802
Castello, A.; Mohammed, S.; Moindrot, B.; Nesterova, T. B.; Brockdor ”, N. hnRNPK803
Recruits PCGF3/5-PRC1 to the Xist RNA B-Repeat to Establish Polycomb-Mediated804
Chromosomal Silencing. Mol Cell 2017, 68,9 5 5 – 9 6 9 . e 1 0 .805
(67) Bomsztyk, K.; Denisenko, O.; Ostrowski, J. hnRNP K: one protein multiple processes.806
Bioessays 2004, 26,6 2 9 – 6 3 8 .807
(68) Nakamoto, M. Y.; Lammer, N. C.; Batey, R. T.; Wuttke, D. S. hnRNPK recognition808
of the B motif of Xist and other biological RNAs. Nucleic Acids Res 2020, 48,9 3 2 0 –809
9335.810
(69) Paziewska, A.; Wyrwicz, L. S.; Bujnicki, J. M.; Bomsztyk, K.; Ostrowski, J. Coopera-811
tive binding of the hnRNP K three KH domains to mRNA targets. FEBS Lett 2004,812
577,1 3 4 – 1 4 0 .813
(70) Klimek-Tomczak, K.; Wyrwicz, L. S.; Jain, S.; Bomsztyk, K.; Ostrowski, J. Character-814
ization of hnRNP K Protein–RNA Interactions. Journal of Molecular Biology 2004,815
342,1 1 3 1 – 1 1 4 1 .816
(71) Ozdilek, B. A.; Thompson, V. F.; Ahmed, N. S.; White, C. I.; Batey, R. T.;817
Schwartz, J. C. Intrinsically disordered RGG/RG domains mediate degenerate speci-818
ficity in RNA binding. Nucleic Acids Res 2017, 45,7 9 8 4 – 7 9 9 6 .819
41
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint
(72) Rastogi, C.; Rube, H. T.; Kribelbauer, J. F.; Crocker, J.; Loker, R. E.; Martini, G. D.;820
Laptenko, O.; Freed-Pastor, W. A.; Prives, C.; Stern, D. L.; Mann, R. S.; Busse-821
maker, H. J. Accurate and sensitive quantification of protein-DNA binding a !nity.822
Proc. Natl. Acad. Sci. U.S.A. 2018, 115 .823
(73) Li, X.; Melo, L. A. N.; Bussemaker, H. J. Benchmarking and building DNA binding824
a!nity models using allele-specific and allele-agnostic transcription factor binding825
data. Genome Biol 2024, 25,2 8 4 .826
(74) Chu, W.-T.; Yan, Z.; Chu, X.; Zheng, X.; Liu, Z.; Xu, L.; Zhang, K.; Wang, J. Physics827
of biomolecular recognition and conformational dynamics. Rep. Prog. Phys. 2021, 84,828
126601.829
(75) Mitra, R.; Li, J.; Sagendorf, J. M.; Jiang, Y.; Cohen, A. S.; Chiu, T.-P.; Glass-830
cock, C. J.; Rohs, R. Geometric deep learning of protein–DNA binding specificity. Nat831
Methods
2024, 21,1 6 7 4 – 1 6 8 3 .832
(76) Liu, S.; Gomez-Alcala, P.; Leemans, C.; Glassford, W. J.; Melo, L. A. N.; Lu, X.-J.;833
Mann, R. S.; Bussemaker, H. J. Predicting the DNA binding specificity of transcription834
factor mutants using family-level biophysically interpretable machine learning. bioRxiv835
2025,2 0 2 4 . 0 1 . 2 4 . 5 7 7 1 1 5 .836
(77) Wetzel, J. L.; Zhang, K.; Singh, M. Learning probabilistic protein-DNA recognition837
codes from DNA-binding specificities using structural mappings. Genome Res 2022,838
32,1 7 7 6 – 1 7 8 6 .839
(78) Kretsch, R. C.; Hummer, A. M.; He, S.; Yuan, R.; Zhang, J.; Karagianes, T.; Cong, Q.;840
Kryshtafovych, A.; Das, R. Assessment of nucleic acid structure prediction in CASP16.841
2025; http://biorxiv.org/lookup/doi/10.1101/2025.05.06.652459.842
(79) Kaczmarek, J. C.; Kowalski, P. S.; Anderson, D. G. Advances in the delivery of RNA843
therapeutics: from concept to clinical reality. Genome Medicine 2017, 9,6 0 .844
42
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint
(80) Ruscio, A. D.; Franciscis, V. d. Minding the gap: Unlocking the therapeutic potential845
of aptamers and making up for lost time. Molecular Therapy Nucleic Acids 2022, 29,846
384–386, Publisher: Elsevier.847
(81) Dias, R.; Kolazckowski, B. Di” erent combinations of atomic interactions predict848
protein-small molecule and protein-DNA/RNA a !nities with similar accuracy. Pro-849
teins 2015, 83,2 1 0 0 – 2 1 1 4 .850
(82) Leaver-Fay, A. et al. ROSETTA3: an object-oriented software suite for the simulation851
and design of macromolecules. Methods Enzymol 2011, 487,5 4 5 – 5 7 4 .852
(83) Shchepachev, V.; Bresson, S.; Spanos, C.; Petfalski, E.; Fischer, L.; Rappsilber, J.;853
Tollervey, D. Defining the RNA interactome by total RNA-associated protein purifi-854
cation. Molecular Systems Biology 2019, 15, e8689, Publisher: John Wiley & Sons,855
Ltd.856
(84) Kang, J. et al. RNAInter v4.0: RNA interactome repository with redefined confidence857
scoring system and improved accessibility. Nucleic Acids Research 2022, 50,D 3 2 6 –858
D332.859
(85) Caudron-Herger, M. Uncovering the complexity of RNA–protein interactions. Nat Rev860
Mol Cell Biol 2025, 26, 499–499, Publisher: Nature Publishing Group.861
(86) Cui, L.; Ma, R.; Cai, J.; Guo, C.; Chen, Z.; Yao, L.; Wang, Y.; Fan, R.; Wang, X.;862
Shi, Y. RNA modifications: importance in immune cell biology and related diseases.863
Sig Transduct Target Ther 2022, 7, 334, Publisher: Nature Publishing Group.864
(87) Delaunay, S.; Helm, M.; Frye, M. RNA modifications in physiology and disease: to-865
wards clinical applications. Nat Rev Genet 2024, 25, 104–122, Publisher: Nature866
Publishing Group.867
43
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint
(88) Kar, M.; Dar, F.; Welsh, T. J.; Vogel, L. T.; K¨ uhnemuth, R.; Majumdar, A.;868
Krainer, G.; Franzmann, T. M.; Alberti, S.; Seidel, C. A. M.; Knowles, T. P. J.;869
Hyman, A. A.; Pappu, R. V. Phase-separating RNA-binding proteins form heteroge-870
neous distributions of clusters in subsaturated solutions. Proceedings of the National871
Academy of Sciences 2022, 119, e2202222119, Publisher: Proceedings of the National872
Academy of Sciences.873
(89) Phase-separation behaviour of RNAs. Nat. Chem. 2023, 15,1 6 6 0 – 1 6 6 1 ,P u b l i s h e r :874
Nature Publishing Group.875
(90) Stormo, G. D.; Zhao, Y. Determining the specificity of protein–DNA interactions. Nat876
Rev Genet 2010, 11,7 5 1 – 7 6 0 .877
(91) Bryngelson, J. D.; Wolynes, P. G. Spin glasses and the statistical mechanics of protein878
folding. Proc. Natl. Acad. Sci. U.S.A. 1987, 84,7 5 2 4 – 7 5 2 8 .879
(92) Davtyan, A.; Schafer, N. P.; Zheng, W.; Clementi, C.; Wolynes, P. G.; Papoian, G. A.880
A WSEM-MD: Protein Structure Prediction Using Coarse-Grained Physical Potentials881
and Bioinformatically Based Local Structure Biasing. J. Phys. Chem. B 2012, 116,882
8494–8503.883
(93) Schafer, N. P.; Kim, B. L.; Zheng, W.; Wolynes, P. G. Learning to fold proteins using884
energy landscape theory. Israel journal of chemistry 2014, 54,1 3 1 1 – 1 3 3 7 .885
(94) Lin, X.; George, J. T.; Schafer, N. P.; Ng Chau, K.; Birnbaum, M. E.; Clementi, C.;886
Onuchic, J. N.; Levine, H. Rapid assessment of T-cell receptor specificity of the im-887
mune repertoire. Nat Comput Sci 2021, 1,3 6 2 – 3 7 3 .888
(95) Wang, A.; Lin, X.; Chau, K. N.; Onuchic, J. N.; Levine, H.; George, J. T. RACER-m889
leverages structural features for sparse T cell specificity prediction. Sci. Adv. 2024,890
10,e a d l 0 1 6 1 .891
44
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint
(96) Wagih, O. ggseqlogo: a versatile R package for drawing sequence logos. Bioinformatics892
2017, 33,3 6 4 5 – 3 6 4 7 .893
(97) van der Loo, M. P. The stringdist Package for Approximate String Matching. The R894
Journal 2014, 6,1 1 1 – 1 2 2 .895
(98) The UniProt Consortium et al. UniProt: the Universal Protein Knowledgebase in896
2025. Nucleic Acids Research 2025, 53,D 6 0 9 – D 6 1 7 .897
(99) Lorenz, R.; Bernhart, S. H.; H¨ oner zu Siederdissen, C.; Tafer, H.; Flamm, C.;898
Stadler, P. F.; Hofacker, I. L. ViennaRNA Package 2.0. Algorithms for Molecular899
Biology 2011, 6,2 6 .900
(100) Lyskov, S. et al. Serverification of Molecular Modeling Applications: The Rosetta901
Online Server That Includes Everyone (ROSIE). PLoS ONE 2013, 8,e 6 3 9 0 6 .902
45
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted September 16, 2025. ; https://doi.org/10.1101/2025.09.10.675247doi: bioRxiv preprint