Application of vision transformers to protein-ligand affinity prediction | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Application of vision transformers to protein-ligand affinity prediction Jakub Poziemski, Pawel Siedlecki This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9539600/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 5 You are reading this latest preprint version Abstract Predicting protein-ligand binding affinity from three-dimensional (3D) structural data is a central task in structure-based drug discovery, yet it remains challenging due to limited data availability, structural complexity, and the sparse nature of 3D molecular representations. In this study, we investigate the application of vision transformers (ViTs) to the problem of affinity prediction. Unlike other neural networks used in this problem, the ViT framework can capture global, long-range interactions across the entire protein-ligand complex via self-attention, without relying on local receptive fields or predefined interaction cutoffs. We evaluate this advantage in representation of spatial information across two benchmark datasets, demonstrating competitive performance, and in some cases surpassing, state-of-the-art models. We study the models behavior using explainable AI (XAI) techniques, revealing that spatially proximal patches with similar attention scores cluster around biologically relevant regions, confirming the model’s ability to capture key interaction features. Furthermore, we show that data augmentation strategies can yield performance improvements, highlighting the potential for further enhancement. Despite challenges related to data sparsity and conformational variability, ViTs show strong performance and high robustness in structure-based affinity prediction tasks. Our findings underscore their effectiveness in learning spatial patterns and suggest broader applicability to related tasks, such as protein-protein or protein-nucleic acid interaction modeling. Biological sciences/Computational biology and bioinformatics Biological sciences/Drug discovery Physical sciences/Mathematics and computing Affinity protein-ligand complex machine learning deep learning data augmentation data representation Figures Figure 1 Figure 2 Figure 3 Introduction Extracting high-level information from 3D image data poses significant challenges, particularly when only a limited number of examples are available. In structure-based drug discovery, 3D molecular structures are commonly represented as cloud of points (atoms) with Cartesian coordinates, derived either from experimental techniques or in silico predictions [ 1 , 2 ]. Protein-ligand complexes are fundamental components of cellular processes, playing key roles in disease progression and therapeutic outcomes. Understanding structural characteristics and dynamic behavior of these complexes is essential for utilizing the mechanisms of molecular interactions and the forces driving them. Given the complexity and high cost of experimental methods, there is a growing need for computational approaches capable of accurately predicting protein-ligand interactions. In particular, the ability to estimate binding affinity from the 3D structure of a protein-ligand complex is useful in the continuous search for novel molecules and in many subsequent stages leading to the development of therapeutic agents or tool compounds [ 3 ]. The problem of predicting the protein-ligand complex affinity has been approached from many angles, including the application of machine learning and deep learning methods [ 4 , 5 ] and with the use of different types of data [ 6 – 8 ]. Convolutional neural networks (CNN) gained popularity especially combined with 3D molecular structures, due to their ability to identify localized patterns and extract features from raw data. Briefly, the convolutional layers employ various filters to detect patterns, while pooling layers are used to reduce dimensionality and enhance robustness to variations. However, CNN emphasizes local features therefore their inductive bias can lead to a loss of more spatially distant correlations in the protein-ligand complex [ 9 ]. In contrast, graph-based approaches including graph neural networks (GNN) offer a different approach for structure-based protein-ligand affinity prediction, particularly due to their natural ability to encode structural and interaction features [ 10 ]. GNNs can describe molecules and and their interactions as graphs, where nodes can represent atoms or residues and edges represent bonds, spatial proximity or non-covalent interactions. While GNNs align closely with the structural and chemical nature of molecules, they may not fully capture 3D spatial relationships. The geometric GNNs (e.g., incorporating distance or angular information) can mitigate this problem [ 11 ]. Challenges remain as GNNs heavily depends on how the graph is constructed (e.g., cutoff distances, types of edges) and as the information is propagated across many layers, node features can become indistinguishable (over-smoothing), reducing the model’s ability to capture local structural details [ 10 , 12 ]. Transformer architectures, introduced by Vaswani et al. [ 13 ], have found wide application in many areas of computer science. Transformers use the attention mechanism which has been shown to better capture long-range relationships between data. This mechanism has proven highly versatile, extending beyond original NLP applications into various domains, including cheminformatics and drug discovery. In these fields, transformers have enhanced molecular property prediction [ 14 ], molecule generation [ 15 , 16 ] and enabling predictions of molecular interactions [ 17 ]. The Vision Transformers [ 18 ], originally developed for image processing tasks, typically outperform traditional convolutional architectures [ 19 , 20 ]. ViTs offer new opportunities for analyzing complex data structures such as those coded in 3D structures of protein-ligand complexes, however they also require more computational resources to train than e.g. CNNs. In short, the main idea is to represent an image as a sequence of small patches for which the attention mechanism can be employed. The patches are transformed into embedding vectors to which the positional information is added. Embeddings provide a rich data representation scheme by mapping discrete data into continuous vector spaces, allowing for the exploration of more complex patterns. The resulting vectors, called tokens, are given as input to the transformer. The architecture of the Vision transformers (ViT) consist of many encoder layers, where each layer is composed of two main components; multi-head attention and a feedforward neural network. The attention mechanism allows the model to weigh the importance of the different image parts and to capture the long-range dependencies between them [ 13 ]. Importantly, ViTs assume minimal prior knowledge, making them significantly dependent on large datasets. A direct comparison with other transformer-based approaches or related 3D ViT implementations, is not straightforward. In particular, current transformer-based models are designed for tasks such as binding pose prediction, virtual screening, or classification of active versus inactive compounds, rather than quantitative affinity prediction. Guo and Wang [ 21 ] showed that ViT can benefit from pre-training on large datasets and then fine-tuning on specific tasks; their model was able to recognize the near native pose among a pool of decoy poses, showing superior docking power compared to CNN-based models and other SFs benchmarked against CASF_2016. Later, the same team presented an approach utilizing vision transformers in screening [ 22 ], demonstrating ViT can potentially be used to accurately identify near-native poses. In this work we present the application of Vision Transformers to the problem of predicting affinity of protein-ligand complexes using 3D structures as input data. Unlike existing methods, we introduce a tailored ViT-inspired architecture adapted to molecular data, with dedicated design choices for encoding spatial interactions within complexes. We extensively evaluate the model's hyperparameters on diverse tasks connected to affinity prediction, including the influence of data augmentation techniques, generalizability and explainability research. We show performance estimates compared to a number of current neural networks architectures and machine learning models, drawing conclusions on the future employment of vision transformers for the structure based drug discovery field. Results Model overview The ViT_for_affinity model is based on the Vision Transformer (ViT) architecture. Inspired by the work of [ 23 ] we have adapted its architecture specifically for the 3D protein-ligand complex data. Transformers excel when supplied with rich, well‑ordered tokens. To take advantage of this, we process the spatial information using Shifted Patch Tokenization (SPT) concept, i.e. we shift the 21×21×21 Å protein-ligand grid by 3Å along every axis and axis pair, generating twelve shifted replicas in addition to the reference orientation (see Methods for details). By augmenting every complex with twelve shifts the model can explore spatial motifs from multiple local coordinate frames, subsequently forming a concatenated, consensus representation. Another modification introduced in ViT_for_affinity is the Locality Self‑Attention (LSA) mechanism. This modification involved two key changes: (1) introducing a learned temperature parameter to scale attention weights, and (2) applying a diagonal masking strategy to prevent each token from attending to itself. These changes encourage the model to focus more effectively on relationships between distinct spatial regions in the molecular structure, which are often overlooked by other types of neural networks, particularly by CNNs. This geometry‑aware design translates into tangible performance gains. On the PDBBind CoreSet_2013 benchmark, our model achieved the best RMSE versus the strongest published baselines (see Table 1 ). On the CASF_2016 benchmark, the model performs slightly worse, however still achieves very strong results. More detailed studies on performance, ablation, explainability and generalizability of the ViT_for_affinity model are presented in the following sections below. A graphical representation of the model architecture is presented on Fig. 1 . Performance assessment To evaluate the predictive capabilities of the ViT_for_affinity model, we benchmarked its performance against a range of state-of-the-art scoring functions using two widely adopted test sets: CASF_2016 and CoreSet_2013. These datasets differ in difficulty and distributional characteristics, providing a framework for assessing model generalization and accuracy. Table 1 Predictive performance of the ViT_for_affinity model compared to other state-of-the-art models. Sorted by mean performance on both test datasets. PCC: Pearson Correlation Coefficient, RMSE: Root Mean Square Error. Model Architecture RMSE PCC RMSE PCC RMSE PCC CASF_2016 CoreSet_2013 Mean OnionNet-2 CNN [ 24 ] 1.164 0.864 1.357 0.821 1.260 0.842 SS-GNN GNN [ 25 ] 1.181 0.853 1.347 0.816 1.264 0.834 ViT_for_ Affinity Vision transformer [ 23 ] 1.233 0.829 1.332 0.811 1.282 0.820 PerSpect ML Persistent spectral-based ML [ 26 ] 1.724 0.840 1.956 0.793 1.840 0.816 AGL Algebraic Graph Learning Score [ 27 ] 1.773 0.833 1.973 0.792 1.873 0.812 CAPLA Cross-Attention [ 28 ] 1.200 0.843 1.446 0.770 1.323 0.806 OnionNet Contact based CNN [ 29 ] 1.278 0.816 1.503 0.782 1.391 0.799 RF-Score v3 Random forest [ 30 ] 1.390 0.800 1.510 0.740 1.45 0.770 Pafnucy 3D grid based CNN [ 31 ] 1.418 0.775 1.620 0.700 1.519 0.737 DeepDTAF non 3D CNN [ 32 ] 1.355 0.789 2.103 0.608 1.729 0.698 The ViT_for_affinity model, employing a grid-based representation, achieves top 3 overall performance on the combined CoreSet_2013 and CASF_2016 test sets (Table 1 ). A more detailed performance assessment, where all individual predictions are shown, is presented on Fig. 2 . ViT outperforms other grid-based models like Pafnucy and DeepDTAF, both of which had significantly lower PCC (Pearson Correlation Coefficient) and RMSE (Root Mean Square Error) values. However, graph neural network models such as SS-GNN and interaction based convolution neural network model OnionNet-2 show an advantage in predictive performance. Interestingly, it is evident that the CoreSet_2013 is a more difficult test set than CASF_2016; all tested models achieved worse performance (PCC and RMSE) irrespectively of the methods or architecture used. Importantly, on the CoreSet_2013 dataset ViT_for_affinity model achieves the lowest RMSE of 1.332, outperforming all tested methods, including both SS-GNN and OnionNet-2. What is more, the difference in performance as measured by PCC between CASF_2016 and CoreSet_2013 is the lowest of all tested methods, which may indicate ViT_for_affinity high generalization potential. Position encoding impact Positional encoding in transformers is a technique used to inject information about the order of tokens in a sequence, addressing the model's lack of inherent positional awareness due to its self-attention mechanism. The literature describes several approaches to positional encoding, in this study we evaluate four most widely used, namely: 1) no positional encoding: in this configuration the model does not receive any information about the position of individual elements in the sequence, i.e data processing relies solely on the self-attention mechanism; 2) 1D sinusoidal encoding: employs sine and cosine functions to encode positions in a one-dimensional space; 3) 3D sinusoidal encoding: an extension of the sinusoidal encoding, where the position of each element is represented using independent sinusoidal components for all three spatial dimensions; and 4) learnable positional encoding: in this method, position vectors are optimized along with the model parameters, allowing the positional representation to be tailored to the specific task and data structure. This approach eliminates the limitations of manually defined encoding schemes, potentially enhancing the model's flexibility and effectiveness in 3D data analysis. The analysis of ViT_for_affinity performance, depending on four different encoding types are presented in Table 2 . Table 2 Performance of the ViT_for_affinity model using different positional encoding strategies. Positional encodings CASF_2016 CoreSet_2013 PCC RMSE PCC RMSE no positional encoding 0.794 1.334 0.753 1.476 1D sinusoidal encoding 0.801 1.300 0.792 1.370 3D sinusoidal encoding 0.779 1.387 0.749 1.513 learnable positional encoding 0.829 1.233 0.811 1.283 Across both datasets, models employing learnable positional encodings achieved the highest predictive performance. These models consistently outperformed those using fixed encodings (1D and 3D sinusoidal) as well as configurations without positional encoding. Notably, the absence of positional encoding resulted in a substantial decline in performance, highlighting the importance of explicitly incorporating spatial information (i.e. patch position). Together, these findings indicate that model performance improves with increasing flexibility of the positional encoding scheme. More sophisticated encoding strategies enable the model to adapt spatial representations to the underlying data distribution and task, resulting in more accurate predictions. Augmentation impact ViT_for_affinity utilizes voxelized 3D grids which are not inherently rotation-invariant. Without augmentation, the model may learn orientation-specific patterns which can harm its generalization potential. To examine how rotational augmentation influences error and performance of the model, we applied a series of 3D rotations to each input along the three axes, rotating a certain number of times in clockwise or anticlockwise directions (Table 3 ). The results show that increasing the number of rotated training examples improved both the error and performance metrics of the model, on both test sets. In fact ViT_for_affinity required strong augmentation, which might be due to the intrinsic nature of the 3D protein-ligand complexes, or to a relatively small training dataset for the ViT architecture. The results also show that by further increasing the number of applied transformations there still might be room to improve the performance of the model. Table 3 Impact of rotational data augmentation on the performance of the ViT_for_affinity model. no. of transformations CASF_2016 CoreSet_2013 PCC RMSE PCC RMSE 0 0.623 1.810 0.596 1.956 8 0.709 1.618 0.669 1.742 16 0.686 1.614 0.622 1.800 32 0.801 1.305 0.734 1.523 48 0.811 1.284 0.770 1.454 64 0.829 1.233 0.811 1.283 Ablation studies To assess the contribution of different molecular and protein descriptors to the performance of the ViT_for_affinity model, we conducted a series of ablation studies. In each experiment, a specific group of features was removed from the input representation, and the model was re-evaluated on both the CASF_2016 and CoreSet_2013 benchmarks. This approach provides insights into the individual and combined impacts of representation descriptors. The results are summarized in Table 4 . Table 4 Ablation studies for the ViT_for_affinity model. Removed features CASF_2016 CoreSet_2013 PCC RMSE PCC RMSE all features 0.829 1.233 0.811 1.283 protein secondary structures 0.793 1.367 0.777 1.459 pharmacophore features 0.778 1.383 0.746 1.494 amino acid types 0.799 1.320 0.788 1.405 atom types 0.795 1.372 0.753 1.401 Molecule type 0.803 1.310 0.774 1.430 protein secondary structures + amino acid types 0.791 1.321 0.754 1.470 protein secondary structures + amino acid types + pharmacophore features 0.784 1.355 0.761 1.456 Overall, removing any individual descriptor group resulted in a moderate decline in predictive performance, as measured by both PCC and RMSE. Among all tested features, the removal of pharmacophore descriptors had the most pronounced negative effect, highlighting their importance in capturing structure-activity relationships crucial for accurate affinity prediction. Interestingly, the exclusion of protein-related features, specifically secondary structure and amino acid type encodings, led to only a modest decrease in performance. Consistently, the model performed well when trained without high-level protein context. These suggest that basic atom-type encodings alone provide a sufficiently informative representation for the model, which can maintain reasonable accuracy even when relying primarily on atomic-level features. The trends observed in both test sets were consistent, with performance degradation, and followed a similar pattern across CASF_2016 and CoreSet_2013. Influence of complex conformation on model performance Affinity prediction methods are often applied to sub-optimal structures generated through different molecular modeling efforts. To test the ViT_for_affinity model in such real-world applications and assess its robustness to conformational changes, we conducted experiments using molecular dynamics (MD) data from the MDD dataset [ 33 ]. Specifically, we collected 15 complexes from the CASF_2016 benchmark, for which MD simulations were available. Next, from each simulation, 10 representative structures were selected using hierarchical agglomerative clustering (dissimilarity defined as C-α and C-β RMSD after superposition, as implemented in [ 34 ]), resulting in a total of 150 MD-derived structures. The results for the ViT_for_affinity model, in comparison with other affinity prediction methods, are shown in Table 5 . Table 5 Model robustness to conformational changes introduced by MD simulation. Based on molecular dynamics simulations of 15 unique complexes derived from CASF_2016. Model MAE RMSE PCC ViT_for_affinity 1.364 1.600 0.678 OnionNet-2 1.423 1.931 0.634 Pafnucy 1.732 1.983 0.416 Vina 1.920 2.328 0.626 PLEC_NN 1.842 2.428 0.566 Among all tested methods, ViT_for_affinity achieved the best overall results, yielding the lowest MAE (1.364) and RMSE (1.600), as well as the highest Pearson correlation (0.678). OnionNet-2 showed performance comparable to ViT_for_affinity in terms of Pearson correlation of 0.634, but it also exhibited higher error values. Pafnucy produced moderate MAE and RMSE values (1.732 and 1.983, respectively), however, its substantially lower Pearson correlation (0.416) suggests limited prediction quality. The more traditional approaches, Vina and PLEC_NN, were characterized by higher prediction errors. See Supplementary Figures S4 and S5 for detailed results of all tested model predictions. Overall, the results suggest that the transformer architecture can provide a favorable balance between prediction accuracy and robustness when evaluated on MD-derived sub-optimal poses, highlighting its potential applicability to affinity prediction in dynamically varying protein-ligand systems. Explainable AI (XAI) To better understand how the ViT_for_affinity model processes 3D protein-ligand complexes, we performed a set of diagnostic and explainability analyses (Fig. 3 ). We first examined the distribution of attention scores for patches containing ligand atoms versus non-ligand patches (Fig. 3 A). The two distributions differ substantially, with ligand-containing patches showing a clear shift toward higher attention values. This indicates that the model selectively focuses on regions directly involved in molecular interactions rather than distributing attention uniformly. Such behavior is consistent with established biochemical principles, as binding affinity is primarily determined by interactions within and near the binding site. Importantly, this suggests that the model captures meaningful signals rather than relying on spurious correlations or artifacts, a known issue in classical scoring functions [ 36 ]. To assess how spatial relationships are encoded, we analyzed the similarity of the learned positional embeddings (Fig. 3 B). The resulting heatmap reveals a structured and symmetric organization, showing that the positional embedding space is non-uniform and reflects spatial dependencies between patches. Patches that are close in 3D space generally display stronger similarity (note the chessboard like pattern is a result flattening the 3D space into a 1D). The most similar clusters are located within approximately 5 Å of the ligand’s geometric center, while more distant regions exhibit weaker correlations. These observations suggest that the model preferentially emphasizes central ligand-proximal regions and encodes recurring local spatial relationships instead of assigning equivalent representations to all positions. Taken together, this indicates that the learned positional embeddings capture meaningful aspects of molecular geometry that are likely relevant for protein-ligand interactions. Finally, we analyzed the attention weights directed to the [CLS] token, which acts as the global readout of the model (Fig. 3 C). A high peak is observed at patch index 171, corresponding to the central patch, aligned with the ligand position. Beside the central patch, there are 26 patches adjacent to it, which also show higher influence on the [CLS] token, compared to all other patches. This broader attention profile mirrors the spatial patterns observed in the positional embedding analysis, further emphasizing the importance of spatial proximity in the model’s reasoning. Together, these results show that the ViT_for_affinity model learns structured and interpretable representations that are consistent with domain knowledge. The alignment between attention patterns and spatial organization supports the reliability and transparency of the approach. Generalizability assessment The ability of predictive models to generalize to new, unknown data is a key factor determining their practical usefulness in drug design. In the context of predicting protein-ligand affinity, it is particularly important how models deal with complexes in which ligands or protein targets differ structurally from those found in the training set. Performance analysis in such cases allows for the assessment of models' resistance to distribution shifts and their ability to generalize [ 35 , 36 ]. To assess the models' ability to generalize on out-of-distribution data, we compiled a set of complexes with most dissimilar ligands and protein targets to the training set complexes (see Methods section for details, and Supplementary Table S1 ). Briefly, 20 complexes were extracted from the CASF_2016 dataset, which had the lowest structurally similar ligands to the training set (based on ECFP4 1024 bits fingerprint; average Tanimoto similarity Ts = 0.32, highest Ts = 0.38). Although not without drawbacks, this methodology allows to select the most difficult cases to test how well the models perform in predicting affinity for the mostly unseen chemical structures. The results obtained for the ligand generalizability experiment are presented in Table 6 . It can be seen that the ViT_for_affinity and OnionNet-2 models achieve similar performance results to those obtained on the full test set, which indicates their high generalization ability even for ligands that deviate from the training data. In contrast, the CAPLA model shows a decrease in prediction accuracy, suggesting its lower ability to generalize to unknown ligand structures. These results highlight that the effectiveness of prediction methods depends also on the degree of similarity between the test data and the training set. Table 6 Model performance on least similar ligands and targets. Least similar ligands Model PCC MAE RMSE CAPLA 0.777 1.075 1.344 Pafnucy 0.770 1.173 1.406 OnionNet-2 0.869 0.950 1.159 ViT_for_affinity 0.857 0.957 1.129 Least similar targets Model PCC MAE RMSE CAPLA 0.732 1.479 1.689 Pafnucy 0.481 1.864 2.094 OnionNet-2 0.784 1.396 1.704 ViT_for_affinity 0.756 1.364 1.597 In the case of protein target generalizability potential we followed a similar methodology as above. For each Coreset_2013 and CASF_2016 complex we calculated a structural similarity measure with TM-Score [ 37 ] relative to all other protein structures in the training set. On this basis, we selected a subset of 18 complexes in which each protein had a maximum of three structurally similar examples (TM-score > 0.75) in the training set (see Methods and Supplementary Table S2). It is worth noting that in the original test sets there were no cases in which a given protein had zero or only one similar counterpart in the training data. This finding highlights the limitations of existing benchmarks in terms of reliably assessing the structural generalization of models. The results of the models on the selected subset are presented in Table 6 . All the tested models showed a significant decrease in performance (measured by PCC, MAE, and RMSE) compared to the results on the full test sets. This indicates that the effectiveness of all tested affinity prediction methods is strongly dependent on protein structural similarity, much more so than on ligand similarity. However the ViT_for_affinity model stood out from the others, achieving second best results in PCC along with lowest MAE and RMSE values, confirming its high generalization potential. Discussion In this study, we explored the feasibility of using Vision Transformer (ViT) architecture to predict protein-ligand affinity from three-dimensional structural data. This problem is inherently challenging due to factors such as limited and noisy datasets, the presence of activity clefts, and the added complexity of representing molecular systems in a sparse three-dimensional grid. Despite these obstacles, our results show that the developed ViT_for_affinity model can achieve performance on par with, and in some cases surpass state-of-the-art methods reported in the literature. From a methodological perspective our findings highlight that positional encoding and patch representation are critical to capturing spatial relationships within the protein-ligand complex. The explainability (XAI) experiments indicate that the transformer architecture focuses its attention on relevant patterns in the data, allowing it to learn useful spatial context rather than focusing uniformly on all elements of the 3D information. Our results show that areas with similar attentional values tend to cluster in contextually similar regions such as the binding site, and that patches containing ligand atoms exert significantly more influence on the model's prediction. Such behavior is expected and intuitively aligns with analyses commonly conducted by humans. While the overall predictive performance of ViT_for_affinity is comparable to state-of-the-art CNN- and GNN-based approaches, its primary advantage and uniqueness lies in the way spatial information is represented and integrated. Unlike convolutional architectures, which rely on local receptive fields, or graph-based models, which depend on predefined interaction cutoffs, the ViT framework enables global, long-range interactions between all spatial regions of the protein–ligand complex through self-attention. We demonstrate that current neural network-based models exhibit high sensitivity to conformational changes, highlighting the importance of evaluating model robustness with respect to structural variability. The transformer architecture used in the ViT_for_affinity model shows high robustness to sub-optimal poses derived from MD simulation. Given that in real-world applications 3D complexes often originate from molecular modeling and docking, the demonstrated resilience is particularly valuable for improving the practical applicability of affinity prediction. The ViT_for_affinity model shows strong generalization to structurally diverse ligands, maintaining its high performance even for the least similar compounds. In the case of diverse protein targets its generalization potential is slightly weaker, however this trend is consistent across essentially all evaluated models. These results highlight the difficulty of transferring binding affinity patterns across heterogeneous target families. In conclusion, despite inherent challenges such as limited datasets, data sparsity, and the complex nature of molecular interactions, the ViT model demonstrated notable effectiveness in predicting protein-ligand affinities, successfully learning meaningful spatial patterns from structural data. What is more, the model demonstrates high generalizability and robustness when dealing with novel chemistry and sub-optimal data; features which are crucial in many real-life scenarios in early drug discovery. Our work positions the ViT-based model as a promising alternative for problems where global context, spatial flexibility, generalizability to novel chemistry are critical, even when raw predictive metrics are similar to existing approaches. Methods Model development We customised the architecture presented in [ 23 ] to operate on three-dimensional protein-ligand complexes. To capture the spatial context of this type of data we used the Shifted Patch Tokenization methodology, as originally introduced in the said study. In our 3D adaptation, the input grid was augmented using a series of spatial shifts: a total of 12 transformations were generated: 6 shifting the grid by 3 Å in both directions along each of the three Cartesian axes (i.e. two shifts per axis) and 6 shifting the grid by 3 Å in both directions along each combination of two axes: (X,Y),(X, Z),(Y, Z). The original input grid was then concatenated with these 12 shifted versions to increase spatial diversity and contextual richness. The reference and shifted grids (13 in total) were subdivided into non-overlapping 3×3×3 Å patches producing reshaped 7x7x7 grids. Next all grids were concatenated by the feature axis, producing a 7×7×7 grid with a feature dimension of 11883, which was then flattened and normalized. A linear projection layer reduced the feature dimension of each patch from 11883 to 256, resulting in a token sequence of size 343×256 (343 patches, each represented as a 256-dimensional vector). We then applied a learned positional encoding to the sequence of visual tokens and prepended a special [CLS] token, which served to capture the global context of the entire input. The final representation of the [CLS] token was used to predict the binding affinity value, consistent with common transformer practices in vision and NLP tasks. The resulting token sequence is passed to the transformer encoder, which includes a modified version of the scaled dot-product attention mechanism called Locality Self-Attention (LSA). We modify the typical attention mechanism by adding a learnable temperature term for rescaling weights and by zeroing the diagonal of the attention matrix, to force each token to attend exclusively to other regions. A graphical representation of the architecture used is presented on Fig. 1 . We used a single transformer block with 32 attention heads. The model was trained using the Adam optimizer with a learning rate of 0.0001, weight decay of 0.001, and a batch size of 32. A dropout rate of 0.2 was applied to reduce overfitting. The training objective was to minimize mean squared error (MSE). The model was trained for 40 epochs, and after each epoch, the version with the lowest validation loss was saved for downstream evaluation. Training, validation and testing The training, validation, and test sets were derived from the main PDBbind 2020 collection, comprising 19443 protein-ligand complexes with accompanying binding affinity measurements reported as -log Ki, -log Kd, or -log IC₅₀ values. The training, test and validation sets were derived from the main PDBBind 2020 collection (19443 complexes, each with accompanying affinity values. All complexes with peptides and complexes raising errors in feature generation procedure were discarded. CASF_2016 (285 complexes) and CoreSet_2013 (195 complexes) were used as test sets therefore these sets were removed from the main collection. The remaining 13599 complexes were divided into the training set (90%, 12239) and validation set (10%, 1357). Additional comparison analyses of the above datasets is available in the Supplementary Materials section. Representation We represent the protein-ligand 3D complex with a 21x21x21A grid of 1Å (9261 grid points), centered on the ligands’ geometric center. The position of atoms are adjusted to fit the nearest grid point, similar to [ 31 ]. Atoms are represented by a vector of 33 diverse features listed in Table 7 . Protein secondary structures features were obtained with DSSP [ 38 ], pharmacophore features by RDKit [ 39 ], the rest of the features either by self written scripts or MDAnalysis [ 40 ]. Table 7 List of features used for learnable representation. Atom Type (8) Protein Secondary structures (9) Amino Acid type (4) Pharmacophore features (9) Molecule type (3) - C_alpha - N_back - O_back - C - N - O - S - Halogen − 310 - a_helix - pi_helix - b_strand - b-bridge - b_turn - b_end - coil - other - hydrophobic - polar - charged - aromatic - heavy valence - aromatic - ring - valence - charge - donor - acceptor - hydrophobic - hybridization - is protein - is ligand - distance to ligand A similar voxel-based representation with a spatial resolution of 1 Å is commonly employed in the literature [ 31 , 41 , 42 ]. Interestingly, data represented in this manner exhibit a high degree of sparsity, approximately 96% of grid nodes are empty, containing no heavy atoms. This observation suggests that, when working with such representations, it is essential to adopt methods specifically designed to address the challenges posed by sparse data. Vision transformers seem to be a suitable method for this type of data [ 43 ]. Model architecture and training We implemented a Vision Transformer (ViT) architecture inspired by the work of [ 23 ], adapting it specifically for 3D protein-ligand complex data. To process spatial information, we employed Shifted Patch Tokenization (SPT), as originally introduced in [ 23 ]. In our 3D adaptation, the input grid was augmented using a series of spatial shifts: a total of 12 transformations were generated: 6 shifting the grid by 3 Å in both directions along each of the three Cartesian axes (i.e. two shifts per axis) and 6 shifting the grid by 3 Å in both directions along each combination of two axes: (X,Y),(X, Z),(Y, Z). The original input grid was then concatenated with these 12 shifted versions to increase spatial diversity and contextual richness. The reference and shifted grids (13 in total) were subdivided into non-overlapping 3×3×3 Å patches producing reshaped 7x7x7 grids. Next all grids were concatenated by the feature axis, producing a 7×7×7 grid with a feature dimension of 11883, which was then flattened and normalized. A linear projection layer reduced the feature dimension of each patch from 11883 to 256, resulting in a token sequence of size 343×256 (343 patches, each represented as a 256-dimensional vector). We then applied a learned positional encoding to the sequence of visual tokens and prepended a special [CLS] token, which served to capture the global context of the entire input. The final representation of the [CLS] token was used to predict the binding affinity value, consistent with common transformer practices in vision and NLP tasks. The resulting token sequence was passed to the transformer encoder, which included a modified version of the scaled dot-product attention mechanism called Locality Self-Attention (LSA) [ 23 ]. This modification involved two key changes: (1) introducing a learned temperature parameter to scale attention weights, and (2) applying a diagonal masking strategy to prevent each token from attending to itself. These changes encourage the model to focus more effectively on relationships between distinct spatial regions in the molecular structure. A graphical representation of the architecture used is presented on Fig. 1 . We used a single transformer block with 32 attention heads. The model was trained using the Adam optimizer with a learning rate of 0.0001, weight decay of 0.001, and a batch size of 32. A dropout rate of 0.2 was applied to reduce overfitting. The training objective was to minimize mean squared error (MSE). The model was trained for 40 epochs, and after each epoch, the version with the lowest validation loss was saved for downstream evaluation. Declarations Competing Interests Statement The authors declare they have no competing interests Funding This work was sponsored by grant 2020/39/B/ST4/02747 obtained from the Polish National Science Center. Computational resources were provided partially by POL-OPENSCREEN HE ERIC project. Author Contribution PS and JP designed the study. JP implemented the ViT architecture and conducted the calculations, JP & PS analysed the results, PS wrote the main manuscript text. All authors reviewed the manuscript. Data Availability Source code of the final model, including pretrained weights and the script used for generating molecular representation is available in the GitHub repository: https://github.com/JPoziemski/VIT_for_affinity References Berman, H. M., Kleywegt, G. J., Nakamura, H. & Markley, J. L. The Protein Data Bank archive as an open data resource. J. Comput. Aided Mol. Des. 28 , 1009–1014 (2014). Abramson, J. et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630 , 493–500 (2024). Sadybekov, A. V. & Katritch, V. Computational approaches streamlining drug discovery. Nature 616 , 673–685 (2023). Wójcikowski, M., Ballester, P. J. & Siedlecki, P. Performance of machine-learning scoring functions in structure-based virtual screening. Sci. Rep. 7 , 46710 (2017). Özçelik, R., van Tilborg, D., Jiménez-Luna, J. & Grisoni, F. Structure-based drug discovery with deep learning. Chembiochem 24 , e202200776 (2023). Michel, J. & Essex, J. W. Prediction of protein-ligand binding affinity by free energy simulations: assumptions, pitfalls and expectations. J. Comput. Aided Mol. Des. 24 , 639–658 (2010). Liu, Q., Kwoh, C. K. & Li, J. Binding affinity prediction for protein-ligand complexes based on β contacts and B factor. J. Chem. Inf. Model. 53 , 3076–3085 (2013). Boyles, F., Deane, C. M. & Morris, G. M. Learning from the ligand: using ligand-based features to improve binding affinity prediction. Bioinformatics 36 , 758–764 (2020). Gawehn, E., Hiss, J. A. & Schneider, G. Deep learning in drug discovery. Mol. Inf. 35 , 3–14 (2016). Zhang, Z. et al. Graph neural network approaches for drug-target interactions. Curr. Opin. Struct. Biol. 73 , 102327 (2022). Han, J. et al. A survey of geometric Graph Neural Networks: Data structures, models and applications. ArXiv (2024). ;abs/2403.00485: 0 . Zhou, J. et al. Graph neural networks: A review of methods and applications. AI Open. 1 , 57–81 (2020). Vaswani, A. et al. Attention is all you need. arXiv [cs.CL]. (2017). Available: http://arxiv.org/abs/1706.03762 Sultan, A., Sieg, J., Mathea, M. & Volkamer, A. Transformers for molecular Property Prediction: Lessons learned from the past five years. J. Chem. Inf. Model. 64 , 6259–6280 (2024). Bagal, V., Aggarwal, R., Vinod, P. K. & Priyakumar, U. D. MolGPT: Molecular generation using a transformer-decoder model. J. Chem. Inf. Model. 62 , 2064–2076 (2022). Mazuz, E., Shtar, G., Shapira, B. & Rokach, L. Molecule generation using transformers and policy gradient reinforcement learning. Sci. Rep. 13 , 8799 (2023). Huang, K., Xiao, C., Glass, L. M. & Sun, J. MolTrans: Molecular Interaction Transformer for drug-target interaction prediction. Bioinformatics 37 , 830–836 (2021). Dosovitskiy, A. et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv [cs.CV]. (2020). Available: http://arxiv.org/abs/2010.11929 Maurício, J., Domingues, I. & Bernardino, J. Comparing Vision Transformers and Convolutional Neural Networks for image classification: A literature review. Appl. Sci. (Basel) . 13 , 5521 (2023). Rodrigo, M., Cuevas, C. & García, N. Comprehensive comparison between vision transformers and convolutional neural networks for face recognition tasks. Sci. Rep. 14 , 21392 (2024). Guo, L. & Wang, J. ViTRMSE: a three-dimensional RMSE scoring method for protein-ligand docking models based on Vision Transformer. 2022 IEEE International Conference on Bioinformatics and Biomedicine (BIBM). IEEE; pp. 328–333. (2022). Guo, L., Qiu, T., Wang, J. & ViTScore: A Novel Three-Dimensional Vision Transformer Method for Accurate Prediction of Protein-Ligand Docking Poses. IEEE Trans. Nanobiosci. 22 , 734–743 (2023). Lee, S. H., Lee, S. & Song, B. C. Vision Transformer for Small-Size Datasets. arXiv [cs.CV]. 2021. Available: http://arxiv.org/abs/2112.13492 Wang, Z. et al. OnionNet-2: A Convolutional Neural Network Model for Predicting Protein-Ligand Binding Affinity Based on Residue-Atom Contacting Shells. Front. Chem. 9 , 753002 (2021). Zhang, S. et al. SS-GNN: A Simple-Structured Graph Neural Network for Affinity Prediction. ACS Omega . 8 , 22496–22507 (2023). Meng, Z. & Xia, K. Persistent spectral-based machine learning (PerSpect ML) for protein-ligand binding affinity prediction. Sci. Adv. 7 10.1126/sciadv.abc5329 (2021). Nguyen, D. D. & Wei, G-W. AGL-score: Algebraic graph learning score for protein-ligand binding scoring, ranking, docking, and screening. J. Chem. Inf. Model. 59 , 3291–3304 (2019). Jin, Z. et al. CAPLA: improved prediction of protein-ligand binding affinity by a deep learning approach based on a cross-attention mechanism. Bioinformatics 39 10.1093/bioinformatics/btad049 (2023). Zheng, L., Fan, J. & Mu, Y. OnionNet: a multiple-layer inter-molecular contact based convolutional neural network for protein-ligand binding affinity prediction. arXiv [physics.bio-ph]. (2019). Available: http://arxiv.org/abs/1906.02418 Ballester, P. J. & Mitchell, J. B. O. A machine learning approach to predicting protein-ligand binding affinity with applications to molecular docking. Bioinformatics 26 , 1169–1175 (2010). Stepniewska-Dziubinska, M. M., Zielenkiewicz, P. & Siedlecki, P. Development and evaluation of a deep learning model for protein-ligand binding affinity prediction. Bioinformatics 10.1093/bioinformatics/bty374 (2018). Wang, K., Zhou, R., Li, Y. & Li, M. DeepDTAF: a deep learning method to predict protein-ligand binding affinity. Brief. Bioinform . 22 10.1093/bib/bbab072 (2021). Poziemski, J., Yurkevych, A. & Siedlecki, P. Assessment of molecular dynamics time series descriptors in protein–ligand affinity prediction. Digit. Discov . 10.1039/d5dd00452g (2026). Case, D. A. et al. AmberTools J. Chem. Inf. Model. ; 63 : 6183–6191. (2023). Chen, L. et al. Hidden bias in the DUD-E dataset leads to misleading performance of deep learning in structure-based virtual screening. PLoS One . 14 , e0220113 (2019). Wang, M. et al. Deep learning approaches for de novo drug design: An overview. Curr. Opin. Struct. Biol. 72 , 135–144 (2022). Xu, J. & Zhang, Y. How significant is a protein structure similarity with TM-score = 0.5? Bioinformatics 26 , 889–895 (2010). Touw, W. G. et al. A series of PDB-related databanks for everyday needs. Nucleic Acids Res. 43 , D364–D368 (2015). RDKit. Open-source cheminformatics. 10.5281/zenodo.591637 Michaud-Agrawal, N., Denning, E. J., Woolf, T. B. & Beckstein, O. MDAnalysis: a toolkit for the analysis of molecular dynamics simulations. J. Comput. Chem. 32 , 2319–2327 (2011). Jiménez, J., Škalič, M., Martínez-Rosell, G. & De Fabritiis, G. K. D. E. E. P. Protein-Ligand Absolute Binding Affinity Prediction via 3D-Convolutional Neural Networks. J. Chem. Inf. Model. 58 , 287–296 (2018). Aggarwal, R., Gupta, A., Chelur, V., Jawahar, C. V. & Priyakumar, U. D. DeepPocket: Ligand binding site detection and segmentation using 3D convolutional neural networks. J. Chem. Inf. Model. 62 , 5069–5079 (2022). Farina, M. et al. Sparsity in transformers: A systematic literature review. Neurocomputing 582 , 127468 (2024). Additional Declarations No competing interests reported. Supplementary Files SUPViTforaffinity.docx Cite Share Download PDF Status: Under Review Version 1 posted Reviewers invited by journal 05 May, 2026 Editor invited by journal 04 May, 2026 Editor assigned by journal 28 Apr, 2026 Submission checks completed at journal 28 Apr, 2026 First submitted to journal 27 Apr, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9539600","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":635037671,"identity":"955d62fa-a6db-43e4-af86-753cb9d5489c","order_by":0,"name":"Jakub Poziemski","email":"","orcid":"","institution":"Institute of Biochemistry and Biophysics, Polish Academy of Sciences","correspondingAuthor":false,"prefix":"","firstName":"Jakub","middleName":"","lastName":"Poziemski","suffix":""},{"id":635037672,"identity":"5396617a-9bd5-4ff8-80d2-737e03ea64a5","order_by":1,"name":"Pawel Siedlecki","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAuElEQVRIiWNgGAWjYLCCBxUSCQZQtgxxWhLOILTwEKclsY2BBC3m7GcMHyTOs8gzZ+Ax/MzDcIewFsueHGODxG0SxZYNPMbSPAzPCGsxOJC7TQKoJXHDAd4N0jkMh4nQcv7t9h+Jc8BaNv8mTsuN3G0MiQ1gLduIs8VyxvvPEgnHgH5p5v9m/ceACL+Y86clfvhQU5dnzt6WfHNGxR05wg6Ds5jB3AMEdSBpgQAitIyCUTAKRsGIAwBomDutiNxl7gAAAABJRU5ErkJggg==","orcid":"","institution":"Institute of Biochemistry and Biophysics, Polish Academy of Sciences","correspondingAuthor":true,"prefix":"","firstName":"Pawel","middleName":"","lastName":"Siedlecki","suffix":""}],"badges":[],"createdAt":"2026-04-27 09:40:24","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9539600/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9539600/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":109216131,"identity":"7d5210fa-5bdc-4806-af2c-bd230c93117c","added_by":"auto","created_at":"2026-05-13 18:00:00","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":568249,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eArchitecture of vision transformer for 3D images (ViT_for_affinity model).\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-9539600/v1/6c7edbc27eb539acdc041c46.jpeg"},{"id":109405087,"identity":"6b02b546-544d-4d4f-a63a-82bffb09af39","added_by":"auto","created_at":"2026-05-17 12:54:49","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":109124,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eViT_for_affinity performance details. \u003c/strong\u003ePrediction results on CASF_2016 (blue panel) and CorSet_2013 (green panel) test sets.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-9539600/v1/1d01db5688c0196e60e3031d.png"},{"id":109216134,"identity":"4140e9e1-41cb-4dd4-baef-fccd157c92e2","added_by":"auto","created_at":"2026-05-13 18:00:00","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":950235,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eVision Transformer (ViT) explainability analysis\u003c/strong\u003e. A) Histogram of attention scores for the [CLS] token, comparing attention distribution between patches containing ligand atoms versus non-ligand regions. B) Heatmap of dot-product similarity matrix of positional embeddings. Darker areas indicate higher dot-product similarity between patches, highlighting spatial correlations learned during training. The \"checkerboard\" pattern is an artifact of flattening a 3D grid into a 1D sequence. C) [CLS] token attention scores for all 343 patches, indicating which regions of the input contribute most strongly to the model’s final prediction (red color).\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-9539600/v1/b3de83ed8f9868b4e0a87604.png"},{"id":109405908,"identity":"abe19dff-0fa5-4585-ac8f-9ddc03e28419","added_by":"auto","created_at":"2026-05-17 13:22:14","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1742318,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9539600/v1/23bc7dd5-5fdd-461e-95f1-7f8a727dff3a.pdf"},{"id":109298501,"identity":"88d1b124-9a3a-499b-b1f2-a2bac8cd150b","added_by":"auto","created_at":"2026-05-15 09:14:14","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":1236573,"visible":true,"origin":"","legend":"","description":"","filename":"SUPViTforaffinity.docx","url":"https://assets-eu.researchsquare.com/files/rs-9539600/v1/55d1b2098c592e07f3402547.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Application of vision transformers to protein-ligand affinity prediction","fulltext":[{"header":"Introduction","content":"\u003cp\u003eExtracting high-level information from 3D image data poses significant challenges, particularly when only a limited number of examples are available. In structure-based drug discovery, 3D molecular structures are commonly represented as cloud of points (atoms) with Cartesian coordinates, derived either from experimental techniques or \u003cem\u003ein silico\u003c/em\u003e predictions [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. Protein-ligand complexes are fundamental components of cellular processes, playing key roles in disease progression and therapeutic outcomes. Understanding structural characteristics and dynamic behavior of these complexes is essential for utilizing the mechanisms of molecular interactions and the forces driving them. Given the complexity and high cost of experimental methods, there is a growing need for computational approaches capable of accurately predicting protein-ligand interactions. In particular, the ability to estimate binding affinity from the 3D structure of a protein-ligand complex is useful in the continuous search for novel molecules and in many subsequent stages leading to the development of therapeutic agents or tool compounds [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe problem of predicting the protein-ligand complex affinity has been approached from many angles, including the application of machine learning and deep learning methods [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e] and with the use of different types of data [\u003cspan additionalcitationids=\"CR7\" citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. Convolutional neural networks (CNN) gained popularity especially combined with 3D molecular structures, due to their ability to identify localized patterns and extract features from raw data. Briefly, the convolutional layers employ various filters to detect patterns, while pooling layers are used to reduce dimensionality and enhance robustness to variations. However, CNN emphasizes local features therefore their inductive bias can lead to a loss of more spatially distant correlations in the protein-ligand complex [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. In contrast, graph-based approaches including graph neural networks (GNN) offer a different approach for structure-based protein-ligand affinity prediction, particularly due to their natural ability to encode structural and interaction features [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. GNNs can describe molecules and and their interactions as graphs, where nodes can represent atoms or residues and edges represent bonds, spatial proximity or non-covalent interactions. While GNNs align closely with the structural and chemical nature of molecules, they may not fully capture 3D spatial relationships. The geometric GNNs (e.g., incorporating distance or angular information) can mitigate this problem [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. Challenges remain as GNNs heavily depends on how the graph is constructed (e.g., cutoff distances, types of edges) and as the information is propagated across many layers, node features can become indistinguishable (over-smoothing), reducing the model\u0026rsquo;s ability to capture local structural details [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eTransformer architectures, introduced by Vaswani \u003cem\u003eet al.\u003c/em\u003e [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e], have found wide application in many areas of computer science. Transformers use the attention mechanism which has been shown to better capture long-range relationships between data. This mechanism has proven highly versatile, extending beyond original NLP applications into various domains, including cheminformatics and drug discovery. In these fields, transformers have enhanced molecular property prediction [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e], molecule generation [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e, \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e] and enabling predictions of molecular interactions [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. The Vision Transformers [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e], originally developed for image processing tasks, typically outperform traditional convolutional architectures [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e, \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. ViTs offer new opportunities for analyzing complex data structures such as those coded in 3D structures of protein-ligand complexes, however they also require more computational resources to train than e.g. CNNs. In short, the main idea is to represent an image as a sequence of small patches for which the attention mechanism can be employed. The patches are transformed into embedding vectors to which the positional information is added. Embeddings provide a rich data representation scheme by mapping discrete data into continuous vector spaces, allowing for the exploration of more complex patterns. The resulting vectors, called tokens, are given as input to the transformer. The architecture of the Vision transformers (ViT) consist of many encoder layers, where each layer is composed of two main components; multi-head attention and a feedforward neural network. The attention mechanism allows the model to weigh the importance of the different image parts and to capture the long-range dependencies between them [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. Importantly, ViTs assume minimal prior knowledge, making them significantly dependent on large datasets. A direct comparison with other transformer-based approaches or related 3D ViT implementations, is not straightforward. In particular, current transformer-based models are designed for tasks such as binding pose prediction, virtual screening, or classification of active versus inactive compounds, rather than quantitative affinity prediction. Guo and Wang [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e] showed that ViT can benefit from pre-training on large datasets and then fine-tuning on specific tasks; their model was able to recognize the near native pose among a pool of decoy poses, showing superior docking power compared to CNN-based models and other SFs benchmarked against CASF_2016. Later, the same team presented an approach utilizing vision transformers in screening [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e], demonstrating ViT can potentially be used to accurately identify near-native poses.\u003c/p\u003e \u003cp\u003eIn this work we present the application of Vision Transformers to the problem of predicting affinity of protein-ligand complexes using 3D structures as input data. Unlike existing methods, we introduce a tailored ViT-inspired architecture adapted to molecular data, with dedicated design choices for encoding spatial interactions within complexes. We extensively evaluate the model's hyperparameters on diverse tasks connected to affinity prediction, including the influence of data augmentation techniques, generalizability and explainability research. We show performance estimates compared to a number of current neural networks architectures and machine learning models, drawing conclusions on the future employment of vision transformers for the structure based drug discovery field.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eModel overview\u003c/p\u003e \u003cp\u003eThe ViT_for_affinity model is based on the Vision Transformer (ViT) architecture. Inspired by the work of [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e] we have adapted its architecture specifically for the 3D protein-ligand complex data. Transformers excel when supplied with rich, well‑ordered tokens. To take advantage of this, we process the spatial information using Shifted Patch Tokenization (SPT) concept, i.e. we shift the 21\u0026times;21\u0026times;21 \u0026Aring; protein-ligand grid by 3\u0026Aring; along every axis and axis pair, generating twelve shifted replicas in addition to the reference orientation (see Methods for details). By augmenting every complex with twelve shifts the model can explore spatial motifs from multiple local coordinate frames, subsequently forming a concatenated, consensus representation. Another modification introduced in ViT_for_affinity is the Locality Self‑Attention (LSA) mechanism. This modification involved two key changes: (1) introducing a learned temperature parameter to scale attention weights, and (2) applying a diagonal masking strategy to prevent each token from attending to itself. These changes encourage the model to focus more effectively on relationships between distinct spatial regions in the molecular structure, which are often overlooked by other types of neural networks, particularly by CNNs. This geometry‑aware design translates into tangible performance gains. On the PDBBind CoreSet_2013 benchmark, our model achieved the best RMSE versus the strongest published baselines (see Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). On the CASF_2016 benchmark, the model performs slightly worse, however still achieves very strong results. More detailed studies on performance, ablation, explainability and generalizability of the ViT_for_affinity model are presented in the following sections below. A graphical representation of the model architecture is presented on Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003ePerformance assessment\u003c/p\u003e \u003cp\u003eTo evaluate the predictive capabilities of the ViT_for_affinity model, we benchmarked its performance against a range of state-of-the-art scoring functions using two widely adopted test sets: CASF_2016 and CoreSet_2013. These datasets differ in difficulty and distributional characteristics, providing a framework for assessing model generalization and accuracy.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003e\u003cb\u003ePredictive performance of the ViT_for_affinity model compared to other state-of-the-art models.\u003c/b\u003e Sorted by mean performance on both test datasets. PCC: Pearson Correlation Coefficient, RMSE: Root Mean Square Error.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"8\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eArchitecture\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRMSE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePCC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eRMSE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003ePCC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eRMSE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003ePCC\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c4\" namest=\"c3\"\u003e \u003cp\u003eCASF_2016\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c6\" namest=\"c5\"\u003e \u003cp\u003eCoreSet_2013\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c8\" namest=\"c7\"\u003e \u003cp\u003eMean\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOnionNet-2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCNN [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.164\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.864\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.357\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.821\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1.260\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.842\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSS-GNN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGNN [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.181\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.853\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.347\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.816\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1.264\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.834\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eViT_for_\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003eAffinity\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eVision\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003etransformer\u003c/b\u003e [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e1.233\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e0.829\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e1.332\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e0.811\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e\u003cb\u003e1.282\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e\u003cb\u003e0.820\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePerSpect ML\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePersistent spectral-based ML [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.724\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.840\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.956\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.793\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1.840\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.816\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAGL\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAlgebraic Graph Learning Score [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.773\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.833\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.973\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.792\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1.873\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.812\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCAPLA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCross-Attention [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.200\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.843\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.446\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.770\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1.323\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.806\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOnionNet\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eContact based CNN [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.278\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.816\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.503\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.782\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1.391\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.799\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRF-Score v3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRandom forest [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.390\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.800\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.510\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.740\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1.45\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.770\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePafnucy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e3D grid based CNN [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.418\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.775\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.620\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.700\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1.519\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.737\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDeepDTAF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003enon 3D CNN [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.355\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.789\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e2.103\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.608\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1.729\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.698\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe ViT_for_affinity model, employing a grid-based representation, achieves top 3 overall performance on the combined CoreSet_2013 and CASF_2016 test sets (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). A more detailed performance assessment, where all individual predictions are shown, is presented on Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e. ViT outperforms other grid-based models like Pafnucy and DeepDTAF, both of which had significantly lower PCC (Pearson Correlation Coefficient) and RMSE (Root Mean Square Error) values. However, graph neural network models such as SS-GNN and interaction based convolution neural network model OnionNet-2 show an advantage in predictive performance. Interestingly, it is evident that the CoreSet_2013 is a more difficult test set than CASF_2016; all tested models achieved worse performance (PCC and RMSE) irrespectively of the methods or architecture used. Importantly, on the CoreSet_2013 dataset ViT_for_affinity model achieves the lowest RMSE of 1.332, outperforming all tested methods, including both SS-GNN and OnionNet-2. What is more, the difference in performance as measured by PCC between CASF_2016 and CoreSet_2013 is the lowest of all tested methods, which may indicate ViT_for_affinity high generalization potential.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003ePosition encoding impact\u003c/p\u003e \u003cp\u003ePositional encoding in transformers is a technique used to inject information about the order of tokens in a sequence, addressing the model's lack of inherent positional awareness due to its self-attention mechanism. The literature describes several approaches to positional encoding, in this study we evaluate four most widely used, namely: 1) no positional encoding: in this configuration the model does not receive any information about the position of individual elements in the sequence, i.e data processing relies solely on the self-attention mechanism; 2) 1D sinusoidal encoding: employs sine and cosine functions to encode positions in a one-dimensional space; 3) 3D sinusoidal encoding: an extension of the sinusoidal encoding, where the position of each element is represented using independent sinusoidal components for all three spatial dimensions; and 4) learnable positional encoding: in this method, position vectors are optimized along with the model parameters, allowing the positional representation to be tailored to the specific task and data structure. This approach eliminates the limitations of manually defined encoding schemes, potentially enhancing the model's flexibility and effectiveness in 3D data analysis. The analysis of ViT_for_affinity performance, depending on four different encoding types are presented in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePerformance of the ViT_for_affinity model using different positional encoding strategies.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003ePositional encodings\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eCASF_2016\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e \u003cp\u003eCoreSet_2013\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePCC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRMSE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePCC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eRMSE\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eno positional encoding\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.794\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.334\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.753\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.476\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e1D sinusoidal encoding\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.801\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.300\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.792\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.370\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e3D sinusoidal encoding\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.779\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.387\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.749\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.513\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003elearnable positional encoding\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e0.829\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e1.233\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e0.811\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e1.283\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eAcross both datasets, models employing learnable positional encodings achieved the highest predictive performance. These models consistently outperformed those using fixed encodings (1D and 3D sinusoidal) as well as configurations without positional encoding. Notably, the absence of positional encoding resulted in a substantial decline in performance, highlighting the importance of explicitly incorporating spatial information (i.e. patch position). Together, these findings indicate that model performance improves with increasing flexibility of the positional encoding scheme. More sophisticated encoding strategies enable the model to adapt spatial representations to the underlying data distribution and task, resulting in more accurate predictions.\u003c/p\u003e \u003cp\u003eAugmentation impact\u003c/p\u003e \u003cp\u003eViT_for_affinity utilizes voxelized 3D grids which are not inherently rotation-invariant. Without augmentation, the model may learn orientation-specific patterns which can harm its generalization potential. To examine how rotational augmentation influences error and performance of the model, we applied a series of 3D rotations to each input along the three axes, rotating a certain number of times in clockwise or anticlockwise directions (Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). The results show that increasing the number of rotated training examples improved both the error and performance metrics of the model, on both test sets. In fact ViT_for_affinity required strong augmentation, which might be due to the intrinsic nature of the 3D protein-ligand complexes, or to a relatively small training dataset for the ViT architecture. The results also show that by further increasing the number of applied transformations there still might be room to improve the performance of the model.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eImpact of rotational data augmentation on the performance of the ViT_for_affinity model.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eno. of transformations\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eCASF_2016\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e \u003cp\u003eCoreSet_2013\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePCC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRMSE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePCC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eRMSE\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.623\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.810\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.596\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.956\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.709\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.618\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.669\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.742\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e16\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.686\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.614\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.622\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.800\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e32\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.801\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.305\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.734\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.523\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e48\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.811\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.284\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.770\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.454\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e64\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.829\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.233\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.811\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.283\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eAblation studies\u003c/p\u003e \u003cp\u003eTo assess the contribution of different molecular and protein descriptors to the performance of the ViT_for_affinity model, we conducted a series of ablation studies. In each experiment, a specific group of features was removed from the input representation, and the model was re-evaluated on both the CASF_2016 and CoreSet_2013 benchmarks. This approach provides insights into the individual and combined impacts of representation descriptors. The results are summarized in Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eAblation studies for the ViT_for_affinity model.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eRemoved features\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eCASF_2016\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e \u003cp\u003eCoreSet_2013\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePCC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRMSE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePCC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eRMSE\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eall features\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.829\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.233\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.811\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.283\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eprotein secondary structures\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.793\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.367\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.777\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.459\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003epharmacophore features\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.778\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.383\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.746\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.494\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eamino acid types\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.799\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.320\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.788\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.405\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eatom types\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.795\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.372\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.753\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.401\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMolecule type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.803\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.310\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.774\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.430\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eprotein secondary structures\u0026thinsp;+\u0026thinsp;amino acid types\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.791\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.321\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.754\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1.470\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eprotein secondary structures\u0026thinsp;+\u0026thinsp;amino acid types\u003c/p\u003e \u003cp\u003e+ pharmacophore features\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.784\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.355\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.761\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1.456\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eOverall, removing any individual descriptor group resulted in a moderate decline in predictive performance, as measured by both PCC and RMSE. Among all tested features, the removal of pharmacophore descriptors had the most pronounced negative effect, highlighting their importance in capturing structure-activity relationships crucial for accurate affinity prediction.\u003c/p\u003e \u003cp\u003eInterestingly, the exclusion of protein-related features, specifically secondary structure and amino acid type encodings, led to only a modest decrease in performance. Consistently, the model performed well when trained without high-level protein context. These suggest that basic atom-type encodings alone provide a sufficiently informative representation for the model, which can maintain reasonable accuracy even when relying primarily on atomic-level features. The trends observed in both test sets were consistent, with performance degradation, and followed a similar pattern across CASF_2016 and CoreSet_2013.\u003c/p\u003e \u003cp\u003eInfluence of complex conformation on model performance\u003c/p\u003e \u003cp\u003eAffinity prediction methods are often applied to sub-optimal structures generated through different molecular modeling efforts. To test the ViT_for_affinity model in such real-world applications and assess its robustness to conformational changes, we conducted experiments using molecular dynamics (MD) data from the MDD dataset [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. Specifically, we collected 15 complexes from the CASF_2016 benchmark, for which MD simulations were available. Next, from each simulation, 10 representative structures were selected using hierarchical agglomerative clustering (dissimilarity defined as C-α and C-β RMSD after superposition, as implemented in [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]), resulting in a total of 150 MD-derived structures. The results for the ViT_for_affinity model, in comparison with other affinity prediction methods, are shown in Table\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003e\u003cb\u003eModel robustness to conformational changes introduced by MD simulation.\u003c/b\u003e Based on molecular dynamics simulations of 15 unique complexes derived from CASF_2016.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMAE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRMSE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePCC\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eViT_for_affinity\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e1.364\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e1.600\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e0.678\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOnionNet-2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1.423\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.931\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.634\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePafnucy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1.732\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.983\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.416\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eVina\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1.920\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2.328\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.626\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePLEC_NN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1.842\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2.428\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.566\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eAmong all tested methods, ViT_for_affinity achieved the best overall results, yielding the lowest MAE (1.364) and RMSE (1.600), as well as the highest Pearson correlation (0.678). OnionNet-2 showed performance comparable to ViT_for_affinity in terms of Pearson correlation of 0.634, but it also exhibited higher error values. Pafnucy produced moderate MAE and RMSE values (1.732 and 1.983, respectively), however, its substantially lower Pearson correlation (0.416) suggests limited prediction quality. The more traditional approaches, Vina and PLEC_NN, were characterized by higher prediction errors. See Supplementary Figures S4 and S5 for detailed results of all tested model predictions.\u003c/p\u003e \u003cp\u003eOverall, the results suggest that the transformer architecture can provide a favorable balance between prediction accuracy and robustness when evaluated on MD-derived sub-optimal poses, highlighting its potential applicability to affinity prediction in dynamically varying protein-ligand systems.\u003c/p\u003e \u003cp\u003eExplainable AI (XAI)\u003c/p\u003e \u003cp\u003eTo better understand how the ViT_for_affinity model processes 3D protein-ligand complexes, we performed a set of diagnostic and explainability analyses (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). We first examined the distribution of attention scores for patches containing ligand atoms versus non-ligand patches (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA). The two distributions differ substantially, with ligand-containing patches showing a clear shift toward higher attention values. This indicates that the model selectively focuses on regions directly involved in molecular interactions rather than distributing attention uniformly. Such behavior is consistent with established biochemical principles, as binding affinity is primarily determined by interactions within and near the binding site. Importantly, this suggests that the model captures meaningful signals rather than relying on spurious correlations or artifacts, a known issue in classical scoring functions [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eTo assess how spatial relationships are encoded, we analyzed the similarity of the learned positional embeddings (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB). The resulting heatmap reveals a structured and symmetric organization, showing that the positional embedding space is non-uniform and reflects spatial dependencies between patches. Patches that are close in 3D space generally display stronger similarity (note the chessboard like pattern is a result flattening the 3D space into a 1D). The most similar clusters are located within approximately 5 \u0026Aring; of the ligand\u0026rsquo;s geometric center, while more distant regions exhibit weaker correlations. These observations suggest that the model preferentially emphasizes central ligand-proximal regions and encodes recurring local spatial relationships instead of assigning equivalent representations to all positions. Taken together, this indicates that the learned positional embeddings capture meaningful aspects of molecular geometry that are likely relevant for protein-ligand interactions.\u003c/p\u003e \u003cp\u003eFinally, we analyzed the attention weights directed to the [CLS] token, which acts as the global readout of the model (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eC). A high peak is observed at patch index 171, corresponding to the central patch, aligned with the ligand position. Beside the central patch, there are 26 patches adjacent to it, which also show higher influence on the [CLS] token, compared to all other patches. This broader attention profile mirrors the spatial patterns observed in the positional embedding analysis, further emphasizing the importance of spatial proximity in the model\u0026rsquo;s reasoning.\u003c/p\u003e \u003cp\u003eTogether, these results show that the ViT_for_affinity model learns structured and interpretable representations that are consistent with domain knowledge. The alignment between attention patterns and spatial organization supports the reliability and transparency of the approach.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eGeneralizability assessment\u003c/p\u003e \u003cp\u003eThe ability of predictive models to generalize to new, unknown data is a key factor determining their practical usefulness in drug design. In the context of predicting protein-ligand affinity, it is particularly important how models deal with complexes in which ligands or protein targets differ structurally from those found in the training set. Performance analysis in such cases allows for the assessment of models' resistance to distribution shifts and their ability to generalize [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e, \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eTo assess the models' ability to generalize on out-of-distribution data, we compiled a set of complexes with most dissimilar ligands and protein targets to the training set complexes (see Methods section for details, and Supplementary Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e). Briefly, 20 complexes were extracted from the CASF_2016 dataset, which had the lowest structurally similar ligands to the training set (based on ECFP4 1024 bits fingerprint; average Tanimoto similarity Ts\u0026thinsp;=\u0026thinsp;0.32, highest Ts\u0026thinsp;=\u0026thinsp;0.38). Although not without drawbacks, this methodology allows to select the most difficult cases to test how well the models perform in predicting affinity for the mostly unseen chemical structures. The results obtained for the ligand generalizability experiment are presented in Table\u0026nbsp;\u003cspan refid=\"Tab6\" class=\"InternalRef\"\u003e6\u003c/span\u003e. It can be seen that the ViT_for_affinity and OnionNet-2 models achieve similar performance results to those obtained on the full test set, which indicates their high generalization ability even for ligands that deviate from the training data. In contrast, the CAPLA model shows a decrease in prediction accuracy, suggesting its lower ability to generalize to unknown ligand structures. These results highlight that the effectiveness of prediction methods depends also on the degree of similarity between the test data and the training set.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab6\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 6\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eModel performance on least similar ligands and targets.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"4\" nameend=\"c4\" namest=\"c1\"\u003e \u003cp\u003eLeast similar ligands\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePCC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMAE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eRMSE\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCAPLA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.777\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.075\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.344\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePafnucy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.770\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.173\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.406\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOnionNet-2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e0.869\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e0.950\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.159\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eViT_for_affinity\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.857\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.957\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e1.129\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"4\" nameend=\"c4\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003eLeast similar targets\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eModel\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003ePCC\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003eMAE\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003eRMSE\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCAPLA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.732\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.479\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.689\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePafnucy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.481\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.864\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e2.094\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOnionNet-2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e0.784\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.396\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.704\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eViT_for_affinity\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.756\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e1.364\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e1.597\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eIn the case of protein target generalizability potential we followed a similar methodology as above. For each Coreset_2013 and CASF_2016 complex we calculated a structural similarity measure with TM-Score [\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e] relative to all other protein structures in the training set. On this basis, we selected a subset of 18 complexes in which each protein had a maximum of three structurally similar examples (TM-score\u0026thinsp;\u0026gt;\u0026thinsp;0.75) in the training set (see Methods and Supplementary Table S2). It is worth noting that in the original test sets there were no cases in which a given protein had zero or only one similar counterpart in the training data. This finding highlights the limitations of existing benchmarks in terms of reliably assessing the structural generalization of models.\u003c/p\u003e \u003cp\u003eThe results of the models on the selected subset are presented in Table\u0026nbsp;\u003cspan refid=\"Tab6\" class=\"InternalRef\"\u003e6\u003c/span\u003e. All the tested models showed a significant decrease in performance (measured by PCC, MAE, and RMSE) compared to the results on the full test sets. This indicates that the effectiveness of all tested affinity prediction methods is strongly dependent on protein structural similarity, much more so than on ligand similarity. However the ViT_for_affinity model stood out from the others, achieving second best results in PCC along with lowest MAE and RMSE values, confirming its high generalization potential.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eIn this study, we explored the feasibility of using Vision Transformer (ViT) architecture to predict protein-ligand affinity from three-dimensional structural data. This problem is inherently challenging due to factors such as limited and noisy datasets, the presence of activity clefts, and the added complexity of representing molecular systems in a sparse three-dimensional grid. Despite these obstacles, our results show that the developed ViT_for_affinity model can achieve performance on par with, and in some cases surpass state-of-the-art methods reported in the literature.\u003c/p\u003e \u003cp\u003eFrom a methodological perspective our findings highlight that positional encoding and patch representation are critical to capturing spatial relationships within the protein-ligand complex. The explainability (XAI) experiments indicate that the transformer architecture focuses its attention on relevant patterns in the data, allowing it to learn useful spatial context rather than focusing uniformly on all elements of the 3D information. Our results show that areas with similar attentional values tend to cluster in contextually similar regions such as the binding site, and that patches containing ligand atoms exert significantly more influence on the model's prediction. Such behavior is expected and intuitively aligns with analyses commonly conducted by humans.\u003c/p\u003e \u003cp\u003eWhile the overall predictive performance of ViT_for_affinity is comparable to state-of-the-art CNN- and GNN-based approaches, its primary advantage and uniqueness lies in the way spatial information is represented and integrated. Unlike convolutional architectures, which rely on local receptive fields, or graph-based models, which depend on predefined interaction cutoffs, the ViT framework enables global, long-range interactions between all spatial regions of the protein–ligand complex through self-attention.\u003c/p\u003e \u003cp\u003eWe demonstrate that current neural network-based models exhibit high sensitivity to conformational changes, highlighting the importance of evaluating model robustness with respect to structural variability. The transformer architecture used in the ViT_for_affinity model shows high robustness to sub-optimal poses derived from MD simulation. Given that in real-world applications 3D complexes often originate from molecular modeling and docking, the demonstrated resilience is particularly valuable for improving the practical applicability of affinity prediction.\u003c/p\u003e \u003cp\u003eThe ViT_for_affinity model shows strong generalization to structurally diverse ligands, maintaining its high performance even for the least similar compounds. In the case of diverse protein targets its generalization potential is slightly weaker, however this trend is consistent across essentially all evaluated models. These results highlight the difficulty of transferring binding affinity patterns across heterogeneous target families.\u003c/p\u003e \u003cp\u003eIn conclusion, despite inherent challenges such as limited datasets, data sparsity, and the complex nature of molecular interactions, the ViT model demonstrated notable effectiveness in predicting protein-ligand affinities, successfully learning meaningful spatial patterns from structural data. What is more, the model demonstrates high generalizability and robustness when dealing with novel chemistry and sub-optimal data; features which are crucial in many real-life scenarios in early drug discovery. Our work positions the ViT-based model as a promising alternative for problems where global context, spatial flexibility, generalizability to novel chemistry are critical, even when raw predictive metrics are similar to existing approaches.\u003c/p\u003e "},{"header":"Methods","content":"\u003cp\u003eModel development\u003c/p\u003e\u003cp\u003eWe customised the architecture presented in [\u003cspan class=\"CitationRef\"\u003e23\u003c/span\u003e] to operate on three-dimensional protein-ligand complexes. To capture the spatial context of this type of data we used the Shifted Patch Tokenization methodology, as originally introduced in the said study. In our 3D adaptation, the input grid was augmented using a series of spatial shifts: a total of 12 transformations were generated: 6 shifting the grid by 3 Å in both directions along each of the three Cartesian axes (i.e. two shifts per axis) and 6 shifting the grid by 3 Å in both directions along each combination of two axes: (X,Y),(X, Z),(Y, Z). The original input grid was then concatenated with these 12 shifted versions to increase spatial diversity and contextual richness.\u003c/p\u003e\u003cp\u003eThe reference and shifted grids (13 in total) were subdivided into non-overlapping 3×3×3 Å patches producing reshaped 7x7x7 grids. Next all grids were concatenated by the feature axis, producing a 7×7×7 grid with a feature dimension of 11883, which was then flattened and normalized. A linear projection layer reduced the feature dimension of each patch from 11883 to 256, resulting in a token sequence of size 343×256 (343 patches, each represented as a 256-dimensional vector).\u003c/p\u003e\u003cp\u003eWe then applied a learned positional encoding to the sequence of visual tokens and prepended a special [CLS] token, which served to capture the global context of the entire input. The final representation of the [CLS] token was used to predict the binding affinity value, consistent with common transformer practices in vision and NLP tasks.\u003c/p\u003e\u003cp\u003eThe resulting token sequence is passed to the transformer encoder, which includes a modified version of the scaled dot-product attention mechanism called Locality Self-Attention (LSA). We modify the typical attention mechanism by adding a learnable temperature term for rescaling weights and by zeroing the diagonal of the attention matrix, to force each token to attend exclusively to other regions. A graphical representation of the architecture used is presented on Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e\u003cp\u003eWe used a single transformer block with 32 attention heads. The model was trained using the Adam optimizer with a learning rate of 0.0001, weight decay of 0.001, and a batch size of 32. A dropout rate of 0.2 was applied to reduce overfitting. The training objective was to minimize mean squared error (MSE). The model was trained for 40 epochs, and after each epoch, the version with the lowest validation loss was saved for downstream evaluation.\u003c/p\u003e\u003cp\u003eTraining, validation and testing\u003c/p\u003e\u003cp\u003eThe training, validation, and test sets were derived from the main PDBbind 2020 collection, comprising 19443 protein-ligand complexes with accompanying binding affinity measurements reported as -log Ki, -log Kd, or -log IC₅₀ values. The training, test and validation sets were derived from the main PDBBind 2020 collection (19443 complexes, each with accompanying affinity values. All complexes with peptides and complexes raising errors in feature generation procedure were discarded. CASF_2016 (285 complexes) and CoreSet_2013 (195 complexes) were used as test sets therefore these sets were removed from the main collection. The remaining 13599 complexes were divided into the training set (90%, 12239) and validation set (10%, 1357). Additional comparison analyses of the above datasets is available in the Supplementary Materials section.\u003c/p\u003e\u003cp\u003eRepresentation\u003c/p\u003e\u003cp\u003eWe represent the protein-ligand 3D complex with a 21x21x21A grid of 1Å (9261 grid points), centered on the ligands’ geometric center. The position of atoms are adjusted to fit the nearest grid point, similar to [\u003cspan class=\"CitationRef\"\u003e31\u003c/span\u003e]. Atoms are represented by a vector of 33 diverse features listed in Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e7\u003c/span\u003e. Protein secondary structures features were obtained with DSSP [\u003cspan class=\"CitationRef\"\u003e38\u003c/span\u003e], pharmacophore features by RDKit [\u003cspan class=\"CitationRef\"\u003e39\u003c/span\u003e], the rest of the features either by self written scripts or MDAnalysis [\u003cspan class=\"CitationRef\"\u003e40\u003c/span\u003e].\u003c/p\u003e\u003cdiv class=\"gridtable\"\u003e\u003cdiv align=\"left\" class=\"colspec\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\"\u003e\u003c/div\u003e\u003ctable id=\"Tab7\" border=\"1\"\u003e \u003ccaption\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 7\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eList of features used for learnable representation.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003c/colgroup\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\"\u003e \u003cp\u003eAtom Type (8)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\"\u003e \u003cp\u003eProtein Secondary structures (9)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\"\u003e \u003cp\u003eAmino Acid type (4)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\"\u003e \u003cp\u003ePharmacophore features (9)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\"\u003e \u003cp\u003eMolecule type (3)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e- C_alpha\u003c/p\u003e \u003cp\u003e- N_back\u003c/p\u003e \u003cp\u003e- O_back\u003c/p\u003e\u003cp\u003e- C\u003c/p\u003e\u003cp\u003e- N\u003c/p\u003e\u003cp\u003e- O\u003c/p\u003e\u003cp\u003e- S\u003c/p\u003e\u003cp\u003e- Halogen\u003c/p\u003e\u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e− 310\u003c/p\u003e \u003cp\u003e- a_helix\u003c/p\u003e \u003cp\u003e- pi_helix\u003c/p\u003e\u003cp\u003e- b_strand\u003c/p\u003e\u003cp\u003e- b-bridge\u003c/p\u003e\u003cp\u003e- b_turn\u003c/p\u003e\u003cp\u003e- b_end\u003c/p\u003e\u003cp\u003e- coil\u003c/p\u003e\u003cp\u003e- other\u003c/p\u003e\u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e- hydrophobic\u003c/p\u003e \u003cp\u003e- polar\u003c/p\u003e \u003cp\u003e- charged\u003c/p\u003e \u003cp\u003e- aromatic\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e- heavy valence\u003c/p\u003e \u003cp\u003e- aromatic\u003c/p\u003e \u003cp\u003e- ring\u003c/p\u003e \u003cp\u003e- valence\u003c/p\u003e \u003cp\u003e- charge\u003c/p\u003e \u003cp\u003e- donor\u003c/p\u003e \u003cp\u003e- acceptor\u003c/p\u003e \u003cp\u003e- hydrophobic\u003c/p\u003e \u003cp\u003e- hybridization\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\"\u003e \u003cp\u003e- is protein\u003c/p\u003e \u003cp\u003e- is ligand\u003c/p\u003e \u003cp\u003e- distance to ligand\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/table\u003e\u003c/div\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eA similar voxel-based representation with a spatial resolution of 1 Å is commonly employed in the literature [\u003cspan class=\"CitationRef\"\u003e31\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e41\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e42\u003c/span\u003e]. Interestingly, data represented in this manner exhibit a high degree of sparsity, approximately 96% of grid nodes are empty, containing no heavy atoms. This observation suggests that, when working with such representations, it is essential to adopt methods specifically designed to address the challenges posed by sparse data. Vision transformers seem to be a suitable method for this type of data [\u003cspan class=\"CitationRef\"\u003e43\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eModel architecture and training\u003c/p\u003e\u003cp\u003eWe implemented a Vision Transformer (ViT) architecture inspired by the work of [\u003cspan class=\"CitationRef\"\u003e23\u003c/span\u003e], adapting it specifically for 3D protein-ligand complex data. To process spatial information, we employed Shifted Patch Tokenization (SPT), as originally introduced in [\u003cspan class=\"CitationRef\"\u003e23\u003c/span\u003e]. In our 3D adaptation, the input grid was augmented using a series of spatial shifts: a total of 12 transformations were generated: 6 shifting the grid by 3 Å in both directions along each of the three Cartesian axes (i.e. two shifts per axis) and 6 shifting the grid by 3 Å in both directions along each combination of two axes: (X,Y),(X, Z),(Y, Z). The original input grid was then concatenated with these 12 shifted versions to increase spatial diversity and contextual richness.\u003c/p\u003e\u003cp\u003eThe reference and shifted grids (13 in total) were subdivided into non-overlapping 3×3×3 Å patches producing reshaped 7x7x7 grids. Next all grids were concatenated by the feature axis, producing a 7×7×7 grid with a feature dimension of 11883, which was then flattened and normalized. A linear projection layer reduced the feature dimension of each patch from 11883 to 256, resulting in a token sequence of size 343×256 (343 patches, each represented as a 256-dimensional vector).\u003c/p\u003e\u003cp\u003eWe then applied a learned positional encoding to the sequence of visual tokens and prepended a special [CLS] token, which served to capture the global context of the entire input. The final representation of the [CLS] token was used to predict the binding affinity value, consistent with common transformer practices in vision and NLP tasks.\u003c/p\u003e\u003cp\u003eThe resulting token sequence was passed to the transformer encoder, which included a modified version of the scaled dot-product attention mechanism called Locality Self-Attention (LSA) [\u003cspan class=\"CitationRef\"\u003e23\u003c/span\u003e]. This modification involved two key changes: (1) introducing a learned temperature parameter to scale attention weights, and (2) applying a diagonal masking strategy to prevent each token from attending to itself. These changes encourage the model to focus more effectively on relationships between distinct spatial regions in the molecular structure. A graphical representation of the architecture used is presented on Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e\u003cp\u003eWe used a single transformer block with 32 attention heads. The model was trained using the Adam optimizer with a learning rate of 0.0001, weight decay of 0.001, and a batch size of 32. A dropout rate of 0.2 was applied to reduce overfitting. The training objective was to minimize mean squared error (MSE). The model was trained for 40 epochs, and after each epoch, the version with the lowest validation loss was saved for downstream evaluation.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e \u003ch2\u003eCompeting Interests Statement\u003c/h2\u003e \u003cp\u003eThe authors declare they have no competing interests\u003c/p\u003e \u003c/p\u003e\u003ch2\u003eFunding\u003c/h2\u003e \u003cp\u003eThis work was sponsored by grant 2020/39/B/ST4/02747 obtained from the Polish National Science Center. Computational resources were provided partially by POL-OPENSCREEN HE ERIC project.\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003ePS and JP designed the study. JP implemented the ViT architecture and conducted the calculations, JP \u0026amp; PS analysed the results, PS wrote the main manuscript text. All authors reviewed the manuscript.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eSource code of the final model, including pretrained weights and the script used for generating molecular representation is available in the GitHub repository: https://github.com/JPoziemski/VIT_for_affinity\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eBerman, H. M., Kleywegt, G. J., Nakamura, H. \u0026amp; Markley, J. L. The Protein Data Bank archive as an open data resource. \u003cem\u003eJ. Comput. Aided Mol. Des.\u003c/em\u003e \u003cb\u003e28\u003c/b\u003e, 1009\u0026ndash;1014 (2014).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAbramson, J. et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. \u003cem\u003eNature\u003c/em\u003e \u003cb\u003e630\u003c/b\u003e, 493\u0026ndash;500 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSadybekov, A. V. \u0026amp; Katritch, V. Computational approaches streamlining drug discovery. \u003cem\u003eNature\u003c/em\u003e \u003cb\u003e616\u003c/b\u003e, 673\u0026ndash;685 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eW\u0026oacute;jcikowski, M., Ballester, P. J. \u0026amp; Siedlecki, P. Performance of machine-learning scoring functions in structure-based virtual screening. \u003cem\u003eSci. Rep.\u003c/em\u003e \u003cb\u003e7\u003c/b\u003e, 46710 (2017).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e\u0026Ouml;z\u0026ccedil;elik, R., van Tilborg, D., Jim\u0026eacute;nez-Luna, J. \u0026amp; Grisoni, F. Structure-based drug discovery with deep learning. \u003cem\u003eChembiochem\u003c/em\u003e \u003cb\u003e24\u003c/b\u003e, e202200776 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMichel, J. \u0026amp; Essex, J. W. Prediction of protein-ligand binding affinity by free energy simulations: assumptions, pitfalls and expectations. \u003cem\u003eJ. Comput. Aided Mol. Des.\u003c/em\u003e \u003cb\u003e24\u003c/b\u003e, 639\u0026ndash;658 (2010).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu, Q., Kwoh, C. K. \u0026amp; Li, J. Binding affinity prediction for protein-ligand complexes based on β contacts and B factor. \u003cem\u003eJ. Chem. Inf. Model.\u003c/em\u003e \u003cb\u003e53\u003c/b\u003e, 3076\u0026ndash;3085 (2013).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBoyles, F., Deane, C. M. \u0026amp; Morris, G. M. Learning from the ligand: using ligand-based features to improve binding affinity prediction. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cb\u003e36\u003c/b\u003e, 758\u0026ndash;764 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGawehn, E., Hiss, J. A. \u0026amp; Schneider, G. Deep learning in drug discovery. \u003cem\u003eMol. Inf.\u003c/em\u003e \u003cb\u003e35\u003c/b\u003e, 3\u0026ndash;14 (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang, Z. et al. Graph neural network approaches for drug-target interactions. \u003cem\u003eCurr. Opin. Struct. Biol.\u003c/em\u003e \u003cb\u003e73\u003c/b\u003e, 102327 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHan, J. et al. A survey of geometric Graph Neural Networks: Data structures, models and applications. \u003cem\u003eArXiv\u003c/em\u003e (2024). ;abs/2403.00485: 0 .\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhou, J. et al. Graph neural networks: A review of methods and applications. \u003cem\u003eAI Open.\u003c/em\u003e \u003cb\u003e1\u003c/b\u003e, 57\u0026ndash;81 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVaswani, A. et al. Attention is all you need. arXiv [cs.CL]. (2017). Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/1706.03762\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/1706.03762\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSultan, A., Sieg, J., Mathea, M. \u0026amp; Volkamer, A. Transformers for molecular Property Prediction: Lessons learned from the past five years. \u003cem\u003eJ. Chem. Inf. Model.\u003c/em\u003e \u003cb\u003e64\u003c/b\u003e, 6259\u0026ndash;6280 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBagal, V., Aggarwal, R., Vinod, P. K. \u0026amp; Priyakumar, U. D. MolGPT: Molecular generation using a transformer-decoder model. \u003cem\u003eJ. Chem. Inf. Model.\u003c/em\u003e \u003cb\u003e62\u003c/b\u003e, 2064\u0026ndash;2076 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMazuz, E., Shtar, G., Shapira, B. \u0026amp; Rokach, L. Molecule generation using transformers and policy gradient reinforcement learning. \u003cem\u003eSci. Rep.\u003c/em\u003e \u003cb\u003e13\u003c/b\u003e, 8799 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuang, K., Xiao, C., Glass, L. M. \u0026amp; Sun, J. MolTrans: Molecular Interaction Transformer for drug-target interaction prediction. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cb\u003e37\u003c/b\u003e, 830\u0026ndash;836 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDosovitskiy, A. et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv [cs.CV]. (2020). Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/2010.11929\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/2010.11929\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMaur\u0026iacute;cio, J., Domingues, I. \u0026amp; Bernardino, J. Comparing Vision Transformers and Convolutional Neural Networks for image classification: A literature review. \u003cem\u003eAppl. Sci. (Basel)\u003c/em\u003e. \u003cb\u003e13\u003c/b\u003e, 5521 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRodrigo, M., Cuevas, C. \u0026amp; Garc\u0026iacute;a, N. Comprehensive comparison between vision transformers and convolutional neural networks for face recognition tasks. \u003cem\u003eSci. Rep.\u003c/em\u003e \u003cb\u003e14\u003c/b\u003e, 21392 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuo, L. \u0026amp; Wang, J. ViTRMSE: a three-dimensional RMSE scoring method for protein-ligand docking models based on Vision Transformer. 2022 IEEE International Conference on Bioinformatics and Biomedicine (BIBM). IEEE; pp. 328\u0026ndash;333. (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuo, L., Qiu, T., Wang, J. \u0026amp; ViTScore: A Novel Three-Dimensional Vision Transformer Method for Accurate Prediction of Protein-Ligand Docking Poses. \u003cem\u003eIEEE Trans. Nanobiosci.\u003c/em\u003e \u003cb\u003e22\u003c/b\u003e, 734\u0026ndash;743 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLee, S. H., Lee, S. \u0026amp; Song, B. C. Vision Transformer for Small-Size Datasets. arXiv [cs.CV]. 2021. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/2112.13492\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/2112.13492\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, Z. et al. OnionNet-2: A Convolutional Neural Network Model for Predicting Protein-Ligand Binding Affinity Based on Residue-Atom Contacting Shells. \u003cem\u003eFront. Chem.\u003c/em\u003e \u003cb\u003e9\u003c/b\u003e, 753002 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang, S. et al. SS-GNN: A Simple-Structured Graph Neural Network for Affinity Prediction. \u003cem\u003eACS Omega\u003c/em\u003e. \u003cb\u003e8\u003c/b\u003e, 22496\u0026ndash;22507 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMeng, Z. \u0026amp; Xia, K. Persistent spectral-based machine learning (PerSpect ML) for protein-ligand binding affinity prediction. \u003cem\u003eSci. Adv.\u003c/em\u003e \u003cb\u003e7\u003c/b\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1126/sciadv.abc5329\u003c/span\u003e\u003cspan address=\"10.1126/sciadv.abc5329\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNguyen, D. D. \u0026amp; Wei, G-W. AGL-score: Algebraic graph learning score for protein-ligand binding scoring, ranking, docking, and screening. \u003cem\u003eJ. Chem. Inf. Model.\u003c/em\u003e \u003cb\u003e59\u003c/b\u003e, 3291\u0026ndash;3304 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJin, Z. et al. CAPLA: improved prediction of protein-ligand binding affinity by a deep learning approach based on a cross-attention mechanism. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cb\u003e39\u003c/b\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/bioinformatics/btad049\u003c/span\u003e\u003cspan address=\"10.1093/bioinformatics/btad049\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZheng, L., Fan, J. \u0026amp; Mu, Y. OnionNet: a multiple-layer inter-molecular contact based convolutional neural network for protein-ligand binding affinity prediction. arXiv [physics.bio-ph]. (2019). Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/1906.02418\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/1906.02418\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBallester, P. J. \u0026amp; Mitchell, J. B. O. A machine learning approach to predicting protein-ligand binding affinity with applications to molecular docking. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cb\u003e26\u003c/b\u003e, 1169\u0026ndash;1175 (2010).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStepniewska-Dziubinska, M. M., Zielenkiewicz, P. \u0026amp; Siedlecki, P. Development and evaluation of a deep learning model for protein-ligand binding affinity prediction. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/bioinformatics/bty374\u003c/span\u003e\u003cspan address=\"10.1093/bioinformatics/bty374\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, K., Zhou, R., Li, Y. \u0026amp; Li, M. DeepDTAF: a deep learning method to predict protein-ligand binding affinity. \u003cem\u003eBrief. Bioinform\u003c/em\u003e. \u003cb\u003e22\u003c/b\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/bib/bbab072\u003c/span\u003e\u003cspan address=\"10.1093/bib/bbab072\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePoziemski, J., Yurkevych, A. \u0026amp; Siedlecki, P. Assessment of molecular dynamics time series descriptors in protein\u0026ndash;ligand affinity prediction. \u003cem\u003eDigit. Discov\u003c/em\u003e. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1039/d5dd00452g\u003c/span\u003e\u003cspan address=\"10.1039/d5dd00452g\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2026).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCase, D. A. et al. \u003cem\u003eAmberTools J. Chem. Inf. Model.\u003c/em\u003e ;\u003cb\u003e63\u003c/b\u003e: 6183\u0026ndash;6191. (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen, L. et al. Hidden bias in the DUD-E dataset leads to misleading performance of deep learning in structure-based virtual screening. \u003cem\u003ePLoS One\u003c/em\u003e. \u003cb\u003e14\u003c/b\u003e, e0220113 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, M. et al. Deep learning approaches for de novo drug design: An overview. \u003cem\u003eCurr. Opin. Struct. Biol.\u003c/em\u003e \u003cb\u003e72\u003c/b\u003e, 135\u0026ndash;144 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXu, J. \u0026amp; Zhang, Y. How significant is a protein structure similarity with TM-score\u0026thinsp;=\u0026thinsp;0.5? \u003cem\u003eBioinformatics\u003c/em\u003e \u003cb\u003e26\u003c/b\u003e, 889\u0026ndash;895 (2010).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTouw, W. G. et al. A series of PDB-related databanks for everyday needs. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cb\u003e43\u003c/b\u003e, D364\u0026ndash;D368 (2015).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRDKit. Open-source cheminformatics. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.5281/zenodo.591637\u003c/span\u003e\u003cspan address=\"10.5281/zenodo.591637\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMichaud-Agrawal, N., Denning, E. J., Woolf, T. B. \u0026amp; Beckstein, O. MDAnalysis: a toolkit for the analysis of molecular dynamics simulations. \u003cem\u003eJ. Comput. Chem.\u003c/em\u003e \u003cb\u003e32\u003c/b\u003e, 2319\u0026ndash;2327 (2011).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJim\u0026eacute;nez, J., Škalič, M., Mart\u0026iacute;nez-Rosell, G. \u0026amp; De Fabritiis, G. K. D. E. E. P. Protein-Ligand Absolute Binding Affinity Prediction via 3D-Convolutional Neural Networks. \u003cem\u003eJ. Chem. Inf. Model.\u003c/em\u003e \u003cb\u003e58\u003c/b\u003e, 287\u0026ndash;296 (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAggarwal, R., Gupta, A., Chelur, V., Jawahar, C. V. \u0026amp; Priyakumar, U. D. DeepPocket: Ligand binding site detection and segmentation using 3D convolutional neural networks. \u003cem\u003eJ. Chem. Inf. Model.\u003c/em\u003e \u003cb\u003e62\u003c/b\u003e, 5069\u0026ndash;5079 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFarina, M. et al. Sparsity in transformers: A systematic literature review. \u003cem\u003eNeurocomputing\u003c/em\u003e \u003cb\u003e582\u003c/b\u003e, 127468 (2024).\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Affinity, protein-ligand complex, machine learning, deep learning, data augmentation, data representation","lastPublishedDoi":"10.21203/rs.3.rs-9539600/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9539600/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003ePredicting protein-ligand binding affinity from three-dimensional (3D) structural data is a central task in structure-based drug discovery, yet it remains challenging due to limited data availability, structural complexity, and the sparse nature of 3D molecular representations. In this study, we investigate the application of vision transformers (ViTs) to the problem of affinity prediction. Unlike other neural networks used in this problem, the ViT framework can capture global, long-range interactions across the entire protein-ligand complex via self-attention, without relying on local receptive fields or predefined interaction cutoffs. We evaluate this advantage in representation of spatial information across two benchmark datasets, demonstrating competitive performance, and in some cases surpassing, state-of-the-art models. We study the models behavior using explainable AI (XAI) techniques, revealing that spatially proximal patches with similar attention scores cluster around biologically relevant regions, confirming the model\u0026rsquo;s ability to capture key interaction features. Furthermore, we show that data augmentation strategies can yield performance improvements, highlighting the potential for further enhancement. Despite challenges related to data sparsity and conformational variability, ViTs show strong performance and high robustness in structure-based affinity prediction tasks. Our findings underscore their effectiveness in learning spatial patterns and suggest broader applicability to related tasks, such as protein-protein or protein-nucleic acid interaction modeling.\u003c/p\u003e","manuscriptTitle":"Application of vision transformers to protein-ligand affinity prediction","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-05-13 17:59:56","doi":"10.21203/rs.3.rs-9539600/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"reviewersInvited","content":"","date":"2026-05-05T20:42:35+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-05-04T12:23:04+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-04-28T04:18:31+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-04-28T04:17:33+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2026-04-27T09:22:51+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"3cf86c77-2100-4009-be95-050645a845a9","owner":[],"postedDate":"May 13th, 2026","published":true,"recentEditorialEvents":[{"type":"reviewersInvited","content":"6","date":"2026-05-05T20:42:35+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-05-04T12:23:04+00:00","index":"","fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":67587323,"name":"Biological sciences/Computational biology and bioinformatics"},{"id":67587324,"name":"Biological sciences/Drug discovery"},{"id":67587325,"name":"Physical sciences/Mathematics and computing"}],"tags":[],"updatedAt":"2026-05-13T17:59:56+00:00","versionOfRecord":[],"versionCreatedAt":"2026-05-13 17:59:56","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9539600","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9539600","identity":"rs-9539600","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.