References
Altschul, S. F., Madden, T. L., Sch¨ affer, A. A., Zhang, J., Zhang, Z.,
Miller, W., and Lipman, D. J. (1997). Gapped BLAST and PSI-
BLAST: a new generation of protein database search programs.
Nucleic Acids Research, 25(17), 3389–3402.
Ashburner, M., Ball, C. A., Blake, J. A., Botstein, D., Butler, H.,
Cherry, M. J., Davis, A. P., Dolinski, K., Dwight, S. S., Eppig,
J. T., Harris, M. A., Hill, D. P., Tarver, L. I., Kasarskis, A.,
Lewis, S., Matese, J. C., Richardson, J. E., Ringwald, M., Rubin,
G. M., and Sherlock, G. (2000). Gene ontology: tool for the
unification of biology. Nature Genetics, 25(1), 25–29.
Bekker, J. and Davis, J. (2020). Learning from positive and
unlabeled data: a survey. Machine Learning, 109(4), 719–760.
Buchfink, B., Xie, C., and Huson, D. H. (2014). Fast and sensitive
protein alignment using diamond. Nature Methods, 12, 59 EP –.
[PubMed:25402007] [doi:10.1038/nmeth.3176].
Cao, Y. and Shen, Y. (2021a). TALE: Transformer-based protein
function Annotation with joint sequence–Label Embedding.
Bioinformatics, 37(18), 2825–2833.
Cao, Y. and Shen, Y. (2021b). TALE: Transformer-based protein
function Annotation with joint sequence–Label Embedding.
Bioinformatics, 37(18), 2825–2833.
Chen, Y., Li, Z., Wang, X., Feng, J., and Hu, X. (2010). Predicting
gene function using few positive examples and unlabeled ones.
BMC Genomics , 11(Suppl 2), S11.
Consortium, T. U. (2022). UniProt: the Universal Protein
Knowledgebase in 2023. Nucleic Acids Research , 51(D1),
D523–D531.
Cortes, C. and Vapnik, V. (1995). Support-vector networks.
Machine Learning, 20(3), 273–297.
du Plessis, M. C., Niu, G., and Sugiyama, M. (2014). Analysis of
learning from positive and unlabeled data. In Z. Ghahramani,
M. Welling, C. Cortes, N. Lawrence, and K. Weinberger, editors,
.CC-BY 4.0 International licenseperpetuity. It is made available under a
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
The copyright holder for thisthis version posted January 31, 2024. ; https://doi.org/10.1101/2024.01.28.577662doi: bioRxiv preprint
8 Zhapa-Camacho et al.
Advances in Neural Information Processing Systems, volume 27.
Curran Associates, Inc.
du Plessis, M. C., Niu, G., and Sugiyama, M. (2016). Class-
prior estimation for learning from positive and unlabeled data.
Machine Learning, 106(4), 463–492.
Eisenberg, D., Marcotte, E. M., Xenarios, I., and Yeates, T. O.
(2000). Protein function in the post-genomic era. Nature,
405(6788), 823–826.
Elkan, C. and Noto, K. (2008). Learning classifiers from only
positive and unlabeled data. In Proceedings of the 14th ACM
SIGKDD International Conference on Knowledge Discovery
and Data Mining , KDD ’08, page 213–220, New York, NY, USA.
Association for Computing Machinery.
Elnaggar, A., Heinzinger, M., Dallago, C., Rehawi, G., Wang,
Y., Jones, L., Gibbs, T., Feher, T., Angerer, C., Steinegger,
M., Bhowmik, D., and Rost, B. (2022). Prottrans: Toward
understanding the language of life through self-supervised
learning. IEEE Transactions on Pattern Analysis and Machine
Intelligence, 44(10), 7112–7127.
Hsieh, Y.-G., Niu, G., and Sugiyama, M. (2019). Classification from
positive, unlabeled and biased negative data. In K. Chaudhuri and
R. Salakhutdinov, editors, Proceedings of the 36th International
Conference on Machine Learning , volume 97 of Proceedings of
Machine Learning Research, pages 2820–2829. PMLR.
Kenneth Lange, D. R. H. and Yang, I. (2000). Optimization transfer
using surrogate objective functions. Journal of Computational
and Graphical Statistics , 9(1), 1–20.
Kingma, D. P. and Ba, J. (2015). Adam: A method for
stochastic optimization. In Y. Bengio and Y. LeCun, editors, 3rd
International Conference on Learning Representations, ICLR
2015, San Diego, CA, USA, May 7-9, 2015, Conference Track
Proceedings.
Kiryo, R., Niu, G., du Plessis, M. C., and Sugiyama, M. (2017).
Positive-unlabeled learning with non-negative risk estimator. In
I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus,
S. Vishwanathan, and R. Garnett, editors, Advances in Neural
Information Processing Systems, volume 30. Curran Associates,
Inc.
Kulmanov, M. and Hoehndorf, R. (2019). DeepGOPlus: improved
protein function prediction from sequence. Bioinformatics.
[PubMed:31350877] [doi:10.1093/bioinformatics/btz595].
Kulmanov, M. and Hoehndorf, R. (2022). DeepGOZero: improving
protein function prediction from sequence and zero-shot learning
based on ontology axioms. Bioinformatics, 38(Supplement
1),
i238–i245.
Kulmanov, M., Khan, M. A., and Hoehndorf, R. (2017). DeepGO:
predicting protein functions from sequence and interactions
using a deep ontology-aware classifier. Bioinformatics, 34(4),
660–668. [PubMed:29028931] [PubMed Central:PMC5860606]
[doi:10.1093/bioinformatics/btx624].
Kulmanov, M., Liu-Wei, W., Yan, Y., and Hoehndorf, R.
(2019). El embeddings: Geometric construction of models
for the description logic el++. In Proceedings of the
Twenty-Eighth International Joint Conference on Artificial
Intelligence, IJCAI-19 , pages 6103–6109. International Joint
Conferences on Artificial Intelligence Organization.
Lan, W., Wang, J., Li, M., Liu, J., Li, Y., Wu, F.-X., and Pan,
Y. (2016). Predicting drug–target interaction using positive-
unlabeled learning. Neurocomputing, 206, 50–57. SI:DMSB.
Li, F., Dong, S., Leier, A., Han, M., Guo, X., Xu, J., Wang,
X., Pan, S., Jia, C., Zhang, Y., Webb, G. I., Coin, L.
J. M., Li, C., and Song, J. (2021). Positive-unlabeled learning
in bioinformatics and computational biology: a brief review.
Briefings in Bioinformatics , 23(1).
Lin, Z., Akin, H., Rao, R., Hie, B., Zhu, Z., Lu, W., Smetanin, N.,
dos Santos Costa, A., Fazel-Zarandi, M., Sercu, T., Candido, S.,
et al. (2022). Language models of protein sequences at the scale
of evolution enable accurate structure prediction. bioRxiv.
Liu, W., Wu, A., Pellegrini, M., and Wang, X. (2015). Integrative
analysis of human protein, function and disease networks.
Scientific Reports, 5(1).
Peng, L., Zhu, W., Liao, B., Duan, Y., Chen, M., Chen, Y., and
Yang, J. (2017). Screening drug-target interactions with positive-
unlabeled learning. Scientific Reports, 7(1).
Plessis, M. D., Niu, G., and Sugiyama, M. (2015). Convex
formulation for learning from positive and unlabeled data.
In F. Bach and D. Blei, editors, Proceedings of the 32nd
International Conference on Machine Learning , volume 37 of
Proceedings of Machine Learning Research , pages 1386–1394,
Lille, France. PMLR.
Radivojac, P., Clark, W. T., Oron, T. R., Schnoes, A. M., Wittkop,
T., Sokolov, A., Graim, K., Funk, C., Verspoor, K., Ben-Hur,
A., Pandey, G., Yunes, J. M., Talwalkar, A. S., Repo, S., Souza,
M. L., Piovesan, D., Casadio, R., Wang, Z., Cheng, J., Fang, H.,
Gough, J., Koskinen, P., Toronen, P., Nokso-Koivisto, J., Holm,
L., Cozzetto, D., Buchan, D. W. A., Bryson, K., Jones, D. T.,
Limaye, B., Inamdar, H., Datta, A., Manjari, S. K., Joshi, R.,
Chitale, M., Kihara, D., Lisewski, A. M., Erdin, S., Venner, E.,
Lichtarge, O., Rentzsch, R., Yang, H., Romero, A. E., Bhat, P.,
Paccanaro, A., Hamp, T., Kaszner, R., Seemayer, S., Vicedo,
E., Schaefer, C., Achten, D., Auer, F., Boehm, A., Braun, T.,
Hecht, M., Heron, M., Honigschmid, P., Hopf, T. A., Kaufmann,
S., Kiening, M., Krompass, D., Landerer, C., Mahlich, Y., Roos,
M., Bjorne, J., Salakoski, T., Wong, A., Shatkay, H., Gatzmann,
F., Sommer, I., Wass, M. N., Sternberg, M. J. E., Skunca, N.,
Supek, F., Bosnjak, M., Panov, P., Dzeroski, S., Smuc, T.,
Kourmpetis, Y. A. I., van Dijk, A. D. J., Braak, C. J. F. t.,
Zhou, Y., Gong, Q., Dong, X., Tian, W., Falda, M., Fontana, P.,
Lavezzo, E., Di Camillo, B., Toppo, S., Lan, L., Djuric, N., Guo,
Y., Vucetic, S., Bairoch, A., Linial, M., Babbitt, P. C., Brenner,
S. E., Orengo, C., Rost, B., Mooney, S. D., and Friedberg, I.
(2013). A large-scale evaluation of computational protein function
prediction. Nat Meth , 10(3), 221–227. [PubMed:23353650]
[PubMed Central:PMC3584181] [doi:10.1038/nmeth.2340].
Rasmussen, C. E. and Williams, C. K. I. (2005). Gaussian processes
for machine learning . Adaptive Computation and Machine
Learning series. MIT Press, London, England.
Schenone, M., Danˇ c´ ık, V., Wagner, B. K., and Clemons, P. A.
(2013). Target identification and mechanism of action in chemical
biology and drug discovery. Nature Chemical Biology , 9(4),
232–240.
Shahriari, B., Swersky, K., Wang, Z., Adams, R. P., and de Freitas,
N. (2016). Taking the human out of the loop: A review of bayesian
optimization. Proceedings of the IEEE , 104(1), 148–175.
Smith, L. N. (2017). Cyclical learning rates for training neural
networks. In 2017 IEEE Winter Conference on Applications
of Computer Vision (WACV) , pages 464–472.
Song, H., Bremer, B. J., Hinds, E. C., Raskutti, G., and Romero,
P. A. (2021). Inferring protein sequence-function relationships
with large-scale positive-unlabeled learning. Cell Systems , 12(1),
92–101.e8.
Stolfi, P., Mastropietro, A., Pasculli, G., Tieri, P., and Vergni, D.
(2023). NIAPU: network-informed adaptive positive-unlabeled
learning for disease gene identification. Bioinformatics, 39(2),
btac848.
Tang, Z., Pei, S., Zhang, Z., Zhu, Y., Zhuang, F., Hoehndorf,
R., and Zhang, X. (2022). Positive-unlabeled learning with
adversarial data augmentation for knowledge graph completion.
In L. D. Raedt, editor, Proceedings of the Thirty-First
International Joint Conference on Artificial Intelligence,
IJCAI-22, pages 2248–2254. International Joint Conferences on
Artificial Intelligence Organization. Main Track.
Vasighizaker, A. and Jalili, S. (2018). C-pugp: A cluster-based
positive unlabeled learning method for disease gene prediction
and prioritization. Computational Biology and Chemistry , 76,
23–31.
Wang, S., You, R., Liu, Y., Xiong, Y., and Zhu, S. (2023).
Netgo 3.0: Protein language model improves large-scale functional
annotations. Genomics, Proteomics & Bioinformatics , 21(2),
349–358.
Yang, P., Li, X.-L., Mei, J.-P., Kwoh, C.-K., and Ng, S.-K.
(2012). Positive-unlabeled learning for disease gene identification.
Bioinformatics, 28(20), 2640–2647.
Youngs, N., Penfold-Brown, D., Drew, K., Shasha, D., and
Bonneau, R. (2013). Parametric bayesian priors and better
choice of negative examples improve protein function prediction.
Bioinformatics, 29(9), 1190–1198.
.CC-BY 4.0 International licenseperpetuity. It is made available under a
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
The copyright holder for thisthis version posted January 31, 2024. ; https://doi.org/10.1101/2024.01.28.577662doi: bioRxiv preprint
9
Yuan, Q., Xie, J., Xie, J., Zhao, H., and Yang, Y. (2023a). Fast
and accurate protein function prediction from sequence through
pretrained language model and homology-based label diffusion.
Briefings in Bioinformatics , 24(3), bbad117.
Yuan, Q., Xie, J., Xie, J., Zhao, H., and Yang, Y. (2023b). Fast
and accurate protein function prediction from sequence through
pretrained language model and homology-based label diffusion.
Briefings in Bioinformatics .
Zhao, X.-M., Wang, Y., Chen, L., and Aihara, K. (2008). Gene
function prediction using labeled and unlabeled data. BMC
Bioinformatics, 9(1).
.CC-BY 4.0 International licenseperpetuity. It is made available under a
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
The copyright holder for thisthis version posted January 31, 2024. ; https://doi.org/10.1101/2024.01.28.577662doi: bioRxiv preprint