Optimal Division of Molecules Into Training And Test Sets With A New Tool To Predict Pharmacophore In 3D-QSAR

preprint OA: closed
View at publisher

Abstract

According to the descriptors in the pharmacophore model, dividing molecules into training and test sets serves to create a good model. It is difficult to track the Local Reactive Descriptor (LRD) effect of the pharmacophore at each interaction point in the 3D metric system. A subset of clusters of atoms can correspond to all or part of the pharmacophore structure. In this study, the multidimensional system of the subset was reduced to a one-dimensional index and the Vector Fingerprint Functions (VFF) of the molecules were created. Models were established by dividing molecules with close and similar VFFs into training and test sets. Sub-clusters were examined for all molecules by applying the Genetic Algorithm (GA). The model was predicted using the Leave One Out-Cross Validation (LOO-CV) method and verified with an external test set. The statistical results of the model obtained according to the division in the new method we developed (Q 2 = 0.604 and R 2 = 0.760 for training-80 and external test-20 sets, respectively) were compared with random and manual division results.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00