References
Ivan A Adzhubei, Steffen Schmidt, Leonid Peshkin, Vasily E Ramensky, Anna Gerasimova, Peer
Bork, Alexey S Kondrashov, and Shamil R Sunyaev. A method and server for predicting
damaging missense mutations. Nature Methods, 7(4):248–249, April 2010. ISSN 1548-7091,
1548-7105. doi: 10.1038/nmeth0410-248. URL https://www.nature.com/articles/
nmeth0410-248.
Ethan Ahler, Ames C. Register, Sujata Chakraborty, Linglan Fang, Emily M. Dieter, Kather-
ine A. Sitko, Rama Subba Rao Vidadala, Bridget M. Trevillian, Martin Golkowski, Hannah Gel-
man, Jason J. Stephany, Alan F. Rubin, Ethan A. Merritt, Douglas M. Fowler, and Dustin J.
Maly. A Combined Approach Reveals a Regulatory Mechanism Coupling Src’s Kinase Activ-
ity, Localization, and Phosphotransferase-Independent Functions. Molecular Cell, 74(2):393–
408.e20, April 2019. ISSN 10972765. doi: 10.1016/j.molcel.2019.02.003. URL https:
//linkinghub.elsevier.com/retrieve/pii/S1097276519300930.
Joshua D. Backman, Alexander H. Li, Anthony Marcketta, Dylan Sun, Joelle Mbatchou, Michael D.
Kessler, Christian Benner, Daren Liu, Adam E. Locke, Suganthi Balasubramanian, Ashish Ya-
dav, Nilanjana Banerjee, Christopher E. Gillies, Amy Damask, Simon Liu, Xiaodong Bai, Alicia
Hawes, Evan Maxwell, Lauren Gurski, Kyoko Watanabe, Jack A. Kosmicki, Veera Rajagopal, Ja-
son Mighty, Regeneron Genetics Center, DiscovEHR, Marcus Jones, Lyndon Mitnaul, Eli Stahl,
Giovanni Coppola, Eric Jorgenson, Lukas Habegger, William J. Salerno, Alan R. Shuldiner,
Luca A. Lotta, John D. Overton, Michael N. Cantor, Jeffrey G. Reid, George Yancopoulos,
Hyun M. Kang, Jonathan Marchini, Aris Baras, Gonc ¸alo R. Abecasis, and Manuel A. R. Fer-
reira. Exome sequencing and analysis of 454,787 UK Biobank participants. Nature, 599(7886):
628–634, November 2021. ISSN 0028-0836, 1476-4687. doi: 10.1038/s41586-021-04103-z.
URL https://www.nature.com/articles/s41586-021-04103-z.
Dan Biderman, Jacob Portes, Jose Javier Gonzalez Ortiz, Mansheej Paul, Philip Greengard, Connor
Jennings, Daniel King, Sam Havens, Vitaliy Chiley, Jonathan Frankle, Cody Blakeney, and John P.
Cunningham. LoRA Learns Less and Forgets Less, 2024. URL https://arxiv.org/abs/
2405.09673. Version Number: 2.
B. Boeckmann. The SWISS-PROT protein knowledgebase and its supplement TrEMBL in 2003.
Nucleic Acids Research, 31(1):365–370, January 2003. ISSN 13624962. doi: 10.1093/
nar/gkg095. URL https://academic.oup.com/nar/article-lookup/doi/10.
1093/nar/gkg095.
Nadav Brandes, Grant Goldman, Charlotte H. Wang, Chun Jimmie Ye, and Vasilis Ntra-
nos. Genome-wide prediction of disease variant effects with a deep protein language
model. Nature Genetics, 55(9):1512–1522, September 2023. ISSN 1061-4036, 1546-
1718. doi: 10.1038/s41588-023-01465-0. URL https://www.nature.com/articles/
s41588-023-01465-0.
Anton Bushuiev, Roman Bushuiev, Nikola Zadorozhny, Raman Samusevich, Hannes St¨ark, Jiri Sed-
lar, Tom´aˇs Pluskal, and Josef Sivic. Training on test proteins improves fitness, structure, and func-
tion prediction, 2024. URL https://arxiv.org/abs/2411.02109. Version Number: 1.
Jun Cheng, Guido Novati, Joshua Pan, Clare Bycroft, Akvil ˙e ˇZemgulyt˙e, Taylor Applebaum,
Alexander Pritzel, Lai Hong Wong, Michal Zielinski, Tobias Sargeant, Rosalia G. Schneider,
Andrew W. Senior, John Jumper, Demis Hassabis, Pushmeet Kohli, and ˇZiga Avsec. Accu-
rate proteome-wide missense variant effect prediction with AlphaMissense. Science, 381(6664):
eadg7492, September 2023. ISSN 0036-8075, 1095-9203. doi: 10.1126/science.adg7492. URL
https://www.science.org/doi/10.1126/science.adg7492.
Xingyi Cheng, Bo Chen, Pan Li, Jing Gong, Jie Tang, and Le Song. Training compute-optimal
protein language models. bioRxiv, 2024. doi: 10.1101/2024.06.06.597716. URL https://
www.biorxiv.org/content/early/2024/06/09/2024.06.06.597716.
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher R ´e. Flashattention:
Fast and memory-efficient exact attention with io-awareness. In S. Koyejo, S. Mo-
hamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (eds.), Advances in Neural In-
formation Processing Systems, volume 35, pp. 16344–16359. Curran Associates, Inc.,
11
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted February 7, 2025. ; https://doi.org/10.1101/2025.02.03.635741doi: bioRxiv preprint
2022. URL https://proceedings.neurips.cc/paper_files/paper/2022/
file/67d57c32e20fd0a7a302cb81d36e40d5-Paper-Conference.pdf.
Houssemeddine Derbel, Zhongming Zhao, and Qian Liu. Accurate prediction of functional effect
of single amino acid variants with deep learning. Computational and Structural Biotechnology
Journal, 21:5776–5784, 2023. ISSN 20010370. doi: 10.1016/j.csbj.2023.11.017. URL https:
//linkinghub.elsevier.com/retrieve/pii/S2001037023004312.
Alistair S Dunham and Pedro Beltrao. Exploring amino acid functions in a deep mutational land-
scape. Molecular Systems Biology, 17(7):e10305, July 2021. ISSN 1744-4292, 1744-4292.
doi: 10.15252/msb.202110305. URL https://www.embopress.org/doi/10.15252/
msb.202110305.
Daniel Esposito, Jochen Weile, Jay Shendure, Lea M. Starita, Anthony T. Papenfuss,
Frederick P. Roth, Douglas M. Fowler, and Alan F. Rubin. MaveDB: an open-
source platform to distribute and interpret data from multiplexed assays of variant ef-
fect. Genome Biology, 20(1):223, December 2019. ISSN 1474-760X. doi: 10.
1186/s13059-019-1845-6. URL https://genomebiology.biomedcentral.com/
articles/10.1186/s13059-019-1845-6.
Exome Aggregation Consortium, Monkol Lek, Konrad J. Karczewski, Eric V . Minikel, Kaitlin E.
Samocha, Eric Banks, Timothy Fennell, Anne H. O’Donnell-Luria, James S. Ware, Andrew J.
Hill, Beryl B. Cummings, Taru Tukiainen, Daniel P. Birnbaum, Jack A. Kosmicki, Laramie E.
Duncan, Karol Estrada, Fengmei Zhao, James Zou, Emma Pierce-Hoffman, Joanne Berghout,
David N. Cooper, Nicole Deflaux, Mark DePristo, Ron Do, Jason Flannick, Menachem Fromer,
Laura Gauthier, Jackie Goldstein, Namrata Gupta, Daniel Howrigan, Adam Kiezun, Mitja I.
Kurki, Ami Levy Moonshine, Pradeep Natarajan, Lorena Orozco, Gina M. Peloso, Ryan Poplin,
Manuel A. Rivas, Valentin Ruano-Rubio, Samuel A. Rose, Douglas M. Ruderfer, Khalid Shakir,
Peter D. Stenson, Christine Stevens, Brett P. Thomas, Grace Tiao, Maria T. Tusie-Luna, Ben Weis-
burd, Hong-Hee Won, Dongmei Yu, David M. Altshuler, Diego Ardissino, Michael Boehnke,
John Danesh, Stacey Donnelly, Roberto Elosua, Jose C. Florez, Stacey B. Gabriel, Gad Getz,
Stephen J. Glatt, Christina M. Hultman, Sekar Kathiresan, Markku Laakso, Steven McCarroll,
Mark I. McCarthy, Dermot McGovern, Ruth McPherson, Benjamin M. Neale, Aarno Palotie,
Shaun M. Purcell, Danish Saleheen, Jeremiah M. Scharf, Pamela Sklar, Patrick F. Sullivan,
Jaakko Tuomilehto, Ming T. Tsuang, Hugh C. Watkins, James G. Wilson, Mark J. Daly, and
Daniel G. MacArthur. Analysis of protein-coding genetic variation in 60,706 humans. Nature,
536(7616):285–291, August 2016. ISSN 0028-0836, 1476-4687. doi: 10.1038/nature19057.
URL https://www.nature.com/articles/nature19057.
Jonathan Frazer, Pascal Notin, Mafalda Dias, Aidan Gomez, Joseph K. Min, Kelly Brock,
Yarin Gal, and Debora S. Marks. Disease variant prediction with deep generative models
of evolutionary data. Nature, 599(7883):91–95, November 2021. ISSN 0028-0836, 1476-
4687. doi: 10.1038/s41586-021-04043-8. URL https://www.nature.com/articles/
s41586-021-04043-8.
Hong Gao, Tobias Hamp, Jeffrey Ede, Joshua G. Schraiber, Jeremy McRae, Moriel Singer-Berk,
Yanshen Yang, Anastasia S. D. Dietrich, Petko P. Fiziev, Lukas F. K. Kuderna, Laksshman
Sundaram, Yibing Wu, Aashish Adhikari, Yair Field, Chen Chen, Serafim Batzoglou, Francois
Aguet, Gabrielle Lemire, Rebecca Reimers, Daniel Balick, Mareike C. Janiak, Martin Kuhlwilm,
Joseph D. Orkin, Shivakumara Manu, Alejandro Valenzuela, Juraj Bergman, Marjolaine Rous-
selle, Felipe Ennes Silva, Lidia Agueda, Julie Blanc, Marta Gut, Dorien De Vries, Ian Good-
head, R. Alan Harris, Muthuswamy Raveendran, Axel Jensen, Idriss S. Chuma, Julie E. Hor-
vath, Christina Hvilsom, David Juan, Peter Frandsen, Fabiano R. De Melo, Fabr ´ıcio Bertuol,
Hazel Byrne, Iracilda Sampaio, Izeni Farias, Jo˜ao Valsecchi Do Amaral, Mariluce Messias, Maria
N. F. Da Silva, Mihir Trivedi, Rogerio Rossi, Tomas Hrbek, Nicole Andriaholinirina, Cl´ement J.
Rabarivola, Alphonse Zaramody, Clifford J. Jolly, Jane Phillips-Conroy, Gregory Wilkerson,
Christian Abee, Joe H. Simmons, Eduardo Fernandez-Duque, Sree Kanthaswamy, Fekadu
Shiferaw, Dongdong Wu, Long Zhou, Yong Shao, Guojie Zhang, Julius D. Keyyu, Sascha Knauf,
Minh D. Le, Esther Lizano, Stefan Merker, Arcadi Navarro, Thomas Bataillon, Tilo Nadler,
Chiea Chuen Khor, Jessica Lee, Patrick Tan, Weng Khong Lim, Andrew C. Kitchener, Diet-
mar Zinner, Ivo Gut, Amanda Melin, Katerina Guschanski, Mikkel Heide Schierup, Robin M. D.
12
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted February 7, 2025. ; https://doi.org/10.1101/2025.02.03.635741doi: bioRxiv preprint
Beck, Govindhaswamy Umapathy, Christian Roos, Jean P. Boubli, Monkol Lek, Shamil Sun-
yaev, Anne O’Donnell-Luria, Heidi L. Rehm, Jinbo Xu, Jeffrey Rogers, Tomas Marques-Bonet,
and Kyle Kai-How Farh. The landscape of tolerated genetic variation in humans and primates.
Science, 380(6648):eabn8153, June 2023. ISSN 0036-8075, 1095-9203. doi: 10.1126/science.
abn8197. URL https://www.science.org/doi/10.1126/science.abn8197.
Nilah M. Ioannidis, Joseph H. Rothstein, Vikas Pejaver, Sumit Middha, Shannon K. McDon-
nell, Saurabh Baheti, Anthony Musolf, Qing Li, Emily Holzinger, Danielle Karyadi, Lisa A.
Cannon-Albright, Craig C. Teerlink, Janet L. Stanford, William B. Isaacs, Jianfeng Xu, Kath-
leen A. Cooney, Ethan M. Lange, Johanna Schleutker, John D. Carpten, Isaac J. Powell, Olivier
Cussenot, Geraldine Cancel-Tassin, Graham G. Giles, Robert J. MacInnis, Christiane Maier,
Chih-Lin Hsieh, Fredrik Wiklund, William J. Catalona, William D. Foulkes, Diptasri Mandal,
Rosalind A. Eeles, Zsofia Kote-Jarai, Carlos D. Bustamante, Daniel J. Schaid, Trevor Hastie,
Elaine A. Ostrander, Joan E. Bailey-Wilson, Predrag Radivojac, Stephen N. Thibodeau, Alice S.
Whittemore, and Weiva Sieh. REVEL: An Ensemble Method for Predicting the Pathogenicity
of Rare Missense Variants. The American Journal of Human Genetics, 99(4):877–885, Octo-
ber 2016. ISSN 00029297. doi: 10.1016/j.ajhg.2016.08.016. URL https://linkinghub.
elsevier.com/retrieve/pii/S0002929716303706.
Sudarshan R Iyer, Kevin Nusser, Kristen Jones, Pushkar Shinde, Clare Keddy, Catherine Z Beach,
Erin Aguero, Jeremy Force, Ujwal Shinde, and Monika A Davare. Discovery of oncogenic ROS1
missense mutations with sensitivity to tyrosine kinase inhibitors. EMBO Molecular Medicine,
15(10):e17367, October 2023. ISSN 1757-4676, 1757-4684. doi: 10.15252/emmm.202217367.
URL https://www.embopress.org/doi/10.15252/emmm.202217367.
Milind Jagota, Chengzhong Ye, Carlos Albors, Ruchir Rastogi, Antoine Koehl, Nilah Ioanni-
dis, and Yun S. Song. Cross-protein transfer learning substantially improves disease vari-
ant prediction. Genome Biology, 24(1):182, August 2023. ISSN 1474-760X. doi: 10.
1186/s13059-023-03024-6. URL https://genomebiology.biomedcentral.com/
articles/10.1186/s13059-023-03024-6.
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ron-
neberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin ˇZ´ıdek, Anna Potapenko, Alex Bridg-
land, Clemens Meyer, Simon A. A. Kohl, Andrew J. Ballard, Andrew Cowie, Bernardino
Romera-Paredes, Stanislav Nikolov, Rishub Jain, Jonas Adler, Trevor Back, Stig Petersen,
David Reiman, Ellen Clancy, Michal Zielinski, Martin Steinegger, Michalina Pacholska, Tamas
Berghammer, Sebastian Bodenstein, David Silver, Oriol Vinyals, Andrew W. Senior, Koray
Kavukcuoglu, Pushmeet Kohli, and Demis Hassabis. Highly accurate protein structure pre-
diction with AlphaFold. Nature, 596(7873):583–589, August 2021. ISSN 0028-0836, 1476-
4687. doi: 10.1038/s41586-021-03819-2. URL https://www.nature.com/articles/
s41586-021-03819-2.
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child,
Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling Laws for Neural Language
Models, 2020. URL https://arxiv.org/abs/2001.08361. Version Number: 1.
Konrad J. Karczewski, Laurent C. Francioli, Grace Tiao, Beryl B. Cummings, Jessica Alf ¨oldi,
Qingbo Wang, Ryan L. Collins, Kristen M. Laricchia, Andrea Ganna, Daniel P. Birnbaum,
Laura D. Gauthier, Harrison Brand, Matthew Solomonson, Nicholas A. Watts, Daniel Rhodes,
Moriel Singer-Berk, Eleina M. England, Eleanor G. Seaby, Jack A. Kosmicki, Raymond K. Wal-
ters, Katherine Tashman, Yossi Farjoun, Eric Banks, Timothy Poterba, Arcturus Wang, Cotton
Seed, Nicola Whiffin, Jessica X. Chong, Kaitlin E. Samocha, Emma Pierce-Hoffman, Zachary
Zappala, Anne H. O’Donnell-Luria, Eric Vallabh Minikel, Ben Weisburd, Monkol Lek, James S.
Ware, Christopher Vittal, Irina M. Armean, Louis Bergelson, Kristian Cibulskis, Kristen M.
Connolly, Miguel Covarrubias, Stacey Donnelly, Steven Ferriera, Stacey Gabriel, Jeff Gentry,
Namrata Gupta, Thibault Jeandet, Diane Kaplan, Christopher Llanwarne, Ruchi Munshi, Sam
Novod, Nikelle Petrillo, David Roazen, Valentin Ruano-Rubio, Andrea Saltzman, Molly Schle-
icher, Jose Soto, Kathleen Tibbetts, Charlotte Tolonen, Gordon Wade, Michael E. Talkowski,
Genome Aggregation Database Consortium, Carlos A. Aguilar Salinas, Tariq Ahmad, Chris-
tine M. Albert, Diego Ardissino, Gil Atzmon, John Barnard, Laurent Beaugerie, Emelia J.
Benjamin, Michael Boehnke, Lori L. Bonnycastle, Erwin P. Bottinger, Donald W. Bowden,
13
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted February 7, 2025. ; https://doi.org/10.1101/2025.02.03.635741doi: bioRxiv preprint
Matthew J. Bown, John C. Chambers, Juliana C. Chan, Daniel Chasman, Judy Cho, Mina K.
Chung, Bruce Cohen, Adolfo Correa, Dana Dabelea, Mark J. Daly, Dawood Darbar, Ravin-
dranath Duggirala, Jos ´ee Dupuis, Patrick T. Ellinor, Roberto Elosua, Jeanette Erdmann, T ˜onu
Esko, Martti F¨arkkil¨a, Jose Florez, Andre Franke, Gad Getz, Benjamin Glaser, Stephen J. Glatt,
David Goldstein, Clicerio Gonzalez, Leif Groop, Christopher Haiman, Craig Hanis, Matthew
Harms, Mikko Hiltunen, Matti M. Holi, Christina M. Hultman, Mikko Kallela, Jaakko Kaprio,
Sekar Kathiresan, Bong-Jo Kim, Young Jin Kim, George Kirov, Jaspal Kooner, Seppo Koskinen,
Harlan M. Krumholz, Subra Kugathasan, Soo Heon Kwak, Markku Laakso, Terho Lehtim ¨aki,
Ruth J. F. Loos, Steven A. Lubitz, Ronald C. W. Ma, Daniel G. MacArthur, Jaume Marru-
gat, Kari M. Mattila, Steven McCarroll, Mark I. McCarthy, Dermot McGovern, Ruth McPher-
son, James B. Meigs, Olle Melander, Andres Metspalu, Benjamin M. Neale, Peter M. Nilsson,
Michael C. O’Donovan, Dost Ongur, Lorena Orozco, Michael J. Owen, Colin N. A. Palmer,
Aarno Palotie, Kyong Soo Park, Carlos Pato, Ann E. Pulver, Nazneen Rahman, Anne M. Remes,
John D. Rioux, Samuli Ripatti, Dan M. Roden, Danish Saleheen, Veikko Salomaa, Nilesh J.
Samani, Jeremiah Scharf, Heribert Schunkert, Moore B. Shoemaker, Pamela Sklar, Hilkka Soini-
nen, Harry Sokol, Tim Spector, Patrick F. Sullivan, Jaana Suvisaari, E. Shyong Tai, Yik Ying
Teo, Tuomi Tiinamaija, Ming Tsuang, Dan Turner, Teresa Tusie-Luna, Erkki Vartiainen, Mar-
quis P. Vawter, James S. Ware, Hugh Watkins, Rinse K. Weersma, Maija Wessman, James G.
Wilson, Ramnik J. Xavier, Benjamin M. Neale, Mark J. Daly, and Daniel G. MacArthur. The
mutational constraint spectrum quantified from variation in 141,456 humans. Nature, 581(7809):
434–443, May 2020. ISSN 0028-0836, 1476-4687. doi: 10.1038/s41586-020-2308-7. URL
https://www.nature.com/articles/s41586-020-2308-7.
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. ImageNet classification with deep con-
volutional neural networks. Communications of the ACM, 60(6):84–90, May 2017. ISSN 0001-
0782, 1557-7317. doi: 10.1145/3065386. URL https://dl.acm.org/doi/10.1145/
3065386.
Dimitri M Kullmann and Michael G Hanna. Neurological disorders caused by inherited ion-
channel mutations. The Lancet Neurology, 1(3):157–166, July 2002. ISSN 14744422.
doi: 10.1016/S1474-4422(02)00071-6. URL https://linkinghub.elsevier.com/
retrieve/pii/S1474442202000716.
Aleix Lafita, Ferran Gonzalez, Mahmoud Hossam, Paul Smyth, Jacob Deasy, Ari Allyn-Feuer,
Daniel Seaton, and Stephen Young. Fine-tuning Protein Language Models with Deep Muta-
tional Scanning improves Variant Effect Prediction, 2024. URLhttps://arxiv.org/abs/
2405.06729. Version Number: 1.
Melissa J Landrum, Jennifer M Lee, Mark Benson, Garth R Brown, Chen Chao, Shanmuga Chi-
tipiralla, Baoshan Gu, Jennifer Hart, Douglas Hoffman, Wonhee Jang, Karen Karapetyan, Ken-
neth Katz, Chunlei Liu, Zenith Maddipatla, Adriana Malheiro, Kurt McDaniel, Michael Ovetsky,
George Riley, George Zhou, J Bradley Holmes, Brandi L Kattman, and Donna R Maglott. Clin-
Var: improving access to variant interpretations and supporting evidence.Nucleic Acids Research,
46(D1):D1062–D1067, January 2018. ISSN 0305-1048, 1362-4962. doi: 10.1093/nar/gkx1153.
URL http://academic.oup.com/nar/article/46/D1/D1062/4641904.
Jonas Leichsenring, Peter Horak, Simon Kreutzfeldt, Christoph Heining, Petros Christopou-
los, Anna-Lena V olckmar, Olaf Neumann, Martina Kirchner, Carolin Ploeger, Jan Budczies,
Christoph E. Heilig, Barbara Hutter, Martina Fr ¨ohlich, Sebastian Uhrig, Daniel Kazdal, Michael
Allg¨auer, Alexander Harms, Eugen Rempel, Ulrich Lehmann, Michael Thomas, Nicole Pfarr,
Ninel Azoitei, Irina Bonzheim, Ralf Marienfeld, Peter M ¨oller, Martin Werner, Falko Fend,
Melanie Boerries, Nikolas V on Bubnoff, Silke Lassmann, Thomas Longerich, Michael Bitzer,
Thomas Seufferlein, Nisar Malek, Wilko Weichert, Peter Schirmacher, Roland Penzel, V olker
Endris, Benedikt Brors, Frederick Klauschen, Hanno Glimm, Stefan Fr ¨ohling, and Albrecht
Stenzinger. Variant classification in precision oncology. International Journal of Cancer, 145
(11):2996–3010, December 2019. ISSN 0020-7136, 1097-0215. doi: 10.1002/ijc.32358. URL
https://onlinelibrary.wiley.com/doi/10.1002/ijc.32358.
Francesca-Zhoufan Li, Ava P. Amini, Yisong Yue, Kevin K. Yang, and Alex X. Lu. Feature Reuse
and Scaling: Understanding Transfer Learning with Protein Language Models, February 2024.
URL http://biorxiv.org/lookup/doi/10.1101/2024.02.05.578959.
14
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted February 7, 2025. ; https://doi.org/10.1101/2025.02.03.635741doi: bioRxiv preprint
Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Nikita Smetanin,
Robert Verkuil, Ori Kabeli, Yaniv Shmueli, Allan Dos Santos Costa, Maryam Fazel-Zarandi,
Tom Sercu, Salvatore Candido, and Alexander Rives. Evolutionary-scale prediction of atomic-
level protein structure with a language model.Science, 379(6637):1123–1130, March 2023. ISSN
0036-8075, 1095-9203. doi: 10.1126/science.ade2574. URL https://www.science.org/
doi/10.1126/science.ade2574.
Benjamin J Livesey and Joseph A Marsh. Updated benchmarking of variant effect predictors using
deep mutational scanning. Molecular Systems Biology, 19(8):e11474, August 2023. ISSN 1744-
4292, 1744-4292. doi: 10.15252/msb.202211474. URL https://www.embopress.org/
doi/10.15252/msb.202211474.
C´eline Marquet, Michael Heinzinger, Tobias Olenyi, Christian Dallago, Kyra Erckert, Michael
Bernhofer, Dmitrii Nechaev, and Burkhard Rost. Embeddings from protein language models
predict conservation and variant effects. Human Genetics , 141(10):1629–1647, October 2022.
ISSN 0340-6717, 1432-1203. doi: 10.1007/s00439-021-02411-y. URL https://link.
springer.com/10.1007/s00439-021-02411-y.
Joaquin Mateo, Lotte Steuten, Philippe Aftimos, Fabrice Andr ´e, Mark Davies, Elena Garralda, Jan
Geissler, Don Husereau, Iciar Martinez-Lopez, Nicola Normanno, Jorge S. Reis-Filho, Stephen
Stefani, David M. Thomas, C. Benedikt Westphalen, and Emile V oest. Delivering precision oncol-
ogy to patients with cancer.Nature Medicine, 28(4):658–665, April 2022. ISSN 1078-8956, 1546-
170X. doi: 10.1038/s41591-022-01717-2. URL https://www.nature.com/articles/
s41591-022-01717-2.
Joshua Meier, Roshan Rao, Robert Verkuil, Jason Liu, Tom Sercu, and Alexander Rives. Language
models enable zero-shot prediction of the effects of mutations on protein function, July 2021.
URL http://biorxiv.org/lookup/doi/10.1101/2021.07.09.450648.
Marco Necci, Damiano Piovesan, Zsuzsanna Doszt ´anyi, and Silvio C.E Tosatto. MobiDB-lite:
fast and highly specific consensus prediction of intrinsic disorder in proteins. Bioinformat-
ics, 33(9):1402–1404, May 2017. ISSN 1367-4803, 1367-4811. doi: 10.1093/bioinformatics/
btx015. URL https://academic.oup.com/bioinformatics/article/33/9/
1402/2908909.
Pascal Notin, Aaron W. Kollasch, Daniel Ritter, Lood Van Niekerk, Steffanie Paul, Hansen Spinner,
Nathan Rollins, Ada Shaw, Ruben Weitzman, Jonathan Frazer, Mafalda Dias, Dinko Franceschi,
Rose Orenbuch, Yarin Gal, and Debora S. Marks. ProteinGym: Large-Scale Benchmarks for Pro-
tein Design and Fitness Prediction, December 2023. URLhttp://biorxiv.org/lookup/
doi/10.1101/2023.12.07.570727.
Cl´emence T. B. Pasmans, Bastiaan B. J. Tops, Elisabeth M. P. Steeghs, Veerle M. H. Coup´e, Katrien
Gr¨unberg, Eiko K De Jong, Ed M. D. Schuuring, Stefan M. Willems, Marjolijn J. L. Ligten-
berg, Valesca P. Ret`el, Hans Van Snellenberg, Ewart De Bruijn, Edwin Cuppen, and Geert W. J.
Frederix. Micro-costing diagnostics in oncology: from single-gene testing to whole- genome
sequencing. Expert Review of Pharmacoeconomics & Outcomes Research, 21(3):413–414, May
2021. ISSN 1473-7167, 1744-8379. doi: 10.1080/14737167.2021.1917385. URL https:
//www.tandfonline.com/doi/full/10.1080/14737167.2021.1917385.
Roshan Rao, Joshua Meier, Tom Sercu, Sergey Ovchinnikov, and Alexander Rives. Transformer
protein language models are unsupervised structure learners, December 2020. URL http://
biorxiv.org/lookup/doi/10.1101/2020.12.15.422761.
Alexander Rives, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Demi Guo,
Myle Ott, C. Lawrence Zitnick, Jerry Ma, and Rob Fergus. Biological structure and function
emerge from scaling unsupervised learning to 250 million protein sequences. Proceedings of the
National Academy of Sciences, 118(15):e2016239118, April 2021. ISSN 0027-8424, 1091-6490.
doi: 10.1073/pnas.2016239118. URL https://pnas.org/doi/full/10.1073/pnas.
2016239118.
Alan F Rubin, Joseph K Min, Nathan J Rollins, Estelle Y Da, Daniel Esposito, Matthew Har-
rington, Jeremy Stone, Aisha Haley Bianchi, Mafalda Dias, Jonathan Frazer, Yunfan Fu, Molly
15
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted February 7, 2025. ; https://doi.org/10.1101/2025.02.03.635741doi: bioRxiv preprint
Gallaher, Iris Li, Olivia Moscatelli, Jesslyn Yl Ong, Joshua E Rollins, Matthew J Wakefield,
Shenyi “Sunny” Ye, Amy Tam, Abbye E McEwen, Lea M Starita, Vanessa L Bryant, Deb-
ora S Marks, and Douglas M Fowler. MaveDB v2: a curated community database with
over three million variant effects from multiplexed functional assays, November 2021. URL
http://biorxiv.org/lookup/doi/10.1101/2021.11.29.470445.
Ali Saadat and Jacques Fellay. Fine-tuning the ESM2 protein language model to understand the
functional impact of missense variants, 2024. URL https://arxiv.org/abs/2410.
10919. Version Number: 1.
Robert Schmirler, Michael Heinzinger, and Burkhard Rost. Fine-tuning protein language mod-
els boosts predictions across diverse tasks. Nature Communications, 15(1):7407, August 2024.
ISSN 2041-1723. doi: 10.1038/s41467-024-51844-2. URL https://www.nature.com/
articles/s41467-024-51844-2.
Amelie Schreiber. Esmbind and qbind: Lora, qlora, and esm-2 for predicting binding sites and
post translational modification. bioRxiv, 2023. doi: 10.1101/2023.11.13.566930. URL https:
//www.biorxiv.org/content/early/2023/11/14/2023.11.13.566930.
David Stein, Meltem Ece Kars, Yiming Wu, C ¸ i ˘gdem Sevim Bayrak, Peter D. Stenson,
David N. Cooper, Avner Schlessinger, and Yuval Itan. Genome-wide prediction of
pathogenic gain- and loss-of-function variants from ensemble learning of a diverse fea-
ture set. Genome Medicine, 15(1):103, November 2023. ISSN 1756-994X. doi: 10.
1186/s13073-023-01261-9. URL https://genomemedicine.biomedcentral.com/
articles/10.1186/s13073-023-01261-9.
Baris E. Suzek, Hongzhan Huang, Peter McGarvey, Raja Mazumder, and Cathy H. Wu.
UniRef: comprehensive and non-redundant UniProt reference clusters. Bioinformatics, 23
(10):1282–1288, May 2007. ISSN 1367-4811, 1367-4803. doi: 10.1093/bioinformatics/
btm098. URL https://academic.oup.com/bioinformatics/article/23/10/
1282/197795.
The UniProt Consortium, Alex Bateman, Maria-Jesus Martin, Sandra Orchard, Michele Magrane,
Shadab Ahmad, Emanuele Alpi, Emily H Bowler-Barnett, Ramona Britto, Hema Bye-A-Jee,
Austra Cukura, Paul Denny, Tunca Dogan, ThankGod Ebenezer, Jun Fan, Penelope Garmiri,
Leonardo Jose Da Costa Gonzales, Emma Hatton-Ellis, Abdulrahman Hussein, Alexandr Ig-
natchenko, Giuseppe Insana, Rizwan Ishtiaq, Vishal Joshi, Dushyanth Jyothi, Swaathi Kan-
dasaamy, Antonia Lock, Aurelien Luciani, Marija Lugaric, Jie Luo, Yvonne Lussi, Alistair Mac-
Dougall, Fabio Madeira, Mahdi Mahmoudy, Alok Mishra, Katie Moulang, Andrew Nightingale,
Sangya Pundir, Guoying Qi, Shriya Raj, Pedro Raposo, Daniel L Rice, Rabie Saidi, Rafael San-
tos, Elena Speretta, James Stephenson, Prabhat Totoo, Edward Turner, Nidhi Tyagi, Preethi Va-
sudev, Kate Warner, Xavier Watkins, Rossana Zaru, Hermann Zellner, Alan J Bridge, Lucila
Aimo, Ghislaine Argoud-Puy, Andrea H Auchincloss, Kristian B Axelsen, Parit Bansal, Del-
phine Baratin, Teresa M Batista Neto, Marie-Claude Blatter, Jerven T Bolleman, Emmanuel
Boutet, Lionel Breuza, Blanca Cabrera Gil, Cristina Casals-Casas, Kamal Chikh Echioukh, Elis-
abeth Coudert, Beatrice Cuche, Edouard De Castro, Anne Estreicher, Maria L Famiglietti, Marc
Feuermann, Elisabeth Gasteiger, Pascale Gaudet, Sebastien Gehant, Vivienne Gerritsen, Arnaud
Gos, Nadine Gruaz, Chantal Hulo, Nevila Hyka-Nouspikel, Florence Jungo, Arnaud Kerhornou,
Philippe Le Mercier, Damien Lieberherr, Patrick Masson, Anne Morgat, Venkatesh Muthukrish-
nan, Salvo Paesano, Ivo Pedruzzi, Sandrine Pilbout, Lucille Pourcel, Sylvain Poux, Monica Poz-
zato, Manuela Pruess, Nicole Redaschi, Catherine Rivoire, Christian J A Sigrist, Karin Sonesson,
Shyamala Sundaram, Cathy H Wu, Cecilia N Arighi, Leslie Arminski, Chuming Chen, Yongx-
ing Chen, Hongzhan Huang, Kati Laiho, Peter McGarvey, Darren A Natale, Karen Ross, C R
Vinayaka, Qinghua Wang, Yuqi Wang, and Jian Zhang. UniProt: the Universal Protein Knowl-
edgebase in 2023. Nucleic Acids Research, 51(D1):D523–D531, January 2023. ISSN 0305-
1048, 1362-4962. doi: 10.1093/nar/gkac1052. URL https://academic.oup.com/nar/
article/51/D1/D523/6835362.
Jesse Vig, Ali Madani, Lav R. Varshney, Caiming Xiong, Richard Socher, and Nazneen Fatema
Rajani. BERTology Meets Biology: Interpreting Attention in Protein Language Models, 2020.
URL https://arxiv.org/abs/2006.15222. Version Number: 3.
16
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted February 7, 2025. ; https://doi.org/10.1101/2025.02.03.635741doi: bioRxiv preprint
Guojie Zhong and Yufeng Shen. Representation of missense variants for predicting modes of action.
Machine Learning for Structural Biology Workshop, NeurIPS, 2022.
Guojie Zhong, Yige Zhao, Demi Zhuang, Wendy K Chung, and Yufeng Shen. PreMode predicts
mode-of-action of missense variants by deep graph representation learning of protein sequence
and structural context, February 2024. URL http://biorxiv.org/lookup/doi/10.
1101/2024.02.20.581321.
17
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted February 7, 2025. ; https://doi.org/10.1101/2025.02.03.635741doi: bioRxiv preprint
7 A PPENDIX
7.1 T HE LANDSCAPE OF PATHOGENICITY PREDICTORS
With the advent of artificial intelligence, advanced deep learning models (Krizhevsky et al., 2017)
join the traditionally machine-learning-dominated landscape of mutation prediction models (Ioan-
nidis et al., 2016; Adzhubei et al., 2010). The current landscape is characterized by two axes (cf.
Figure 7):
• (a) whether the mutation effect is a unidirectional pathogenicity score or a bidirectional
functional effect (i.e., increasing or decreasing a specific property or activity) and
• (b) whether the model performs classification or regression.
Regression: pathogenicity score (0-1)
Regression: DMS measurement
Pathogenicity/Fitness
Classification:
Gain-of-Function vs Loss-of-Function
Functional Effect
Protein-specific DMS Fine-Tuning:
ESM1b (Brandes et al. Nat. Gen.)
-no fine-tuning, just likelihood ratio of
mutation according to MLM pretraining
-outperforms EVE
AlphaMissense (Chen et al.
Science)
-AlphaFold & MLM pre-training
-supervised fine-tuning on binary
labels according to variant
frequency
-outperforms EVE & ESM1b
EVE (Frazer et al. Nature)
-unsupervised training on MSAs
-outperforms previous SOTA
(PolyPhen, REVEL etc.)
CPT-1 (Jagota et al. Genome Biol.)
-EVE, ESM-1v, MSA, AF2 etc. features with linear
regression model
-pretrained on 5 DMS -> transfer to other DMS in inference
-outperforms ESM-1v, EVE and partly ESM-NLR (cf. table)
VariPred (Lin et al. Scientific
Reports)
-outperforms EVE & ESM-1v
on ClinVar
-PLM wildtype & mutant
mutation position
embeddings -> shallow DNN
-> binary classification
LoGoFunc (Stein et al. Genom. Med.)
-classifier on GoF , LoF samples
collated from databases with AF2,
physicochemical, biological features
-train data with unrealistic
distribution: LoF and GoF are
overrepresented, neutral lacking
PreMode (Zhong et al.)
-fine-tuning: protein-specific with DMS
-outperforms augmented EVE and
ESM1b
PreMode (Zhong et al.)
-supervised pathogenicity
prediction pretraining, GAT model
with AF2 structure, MSA & PLM
embeddings as features
-slightly outperforms
AlphaMissense, EVE etc. on
ClinVar
ESM-1v-NLR (Lafita et al. ICLR)
-fine-tuning on 25 Mutation Scans with param-free
normalization head
-overall moderate performance gain compared to ESM-1v
base
Other approaches:
-static ESM-1b embeddings + NN
-FT ESM2 for residue annotations
Classification: binary labels
(pathogenic vs benign)
outperformed by
pan-protein prediction
outperformed
by
New SOTA: Our ESM-Effect
Framework
-FT ESM2 model + custom regression
head with inductive biases
-outperforms multi-modal PreMode
-SOTA for Protein-specific DMS Fine-
Tuning
Generalist vs. Specialist
Models Tradeoff
Generalists
-predictions not limited to one
protein -> pan-protein use
-predictions are less accurate
-Application to high-
throughput assessment of
variants BUT limited clinical
usefulness
Specialists
-predictions limited to training
protein (perhaps similar
proteins)
-predictions have higher
accuracy/correlation & reflect
DMS measurement range
-Application to mutations in
training protein, but outside
of mutations covered by DMS
-> somewhat clinically useful
CPT-
1
NLR
-
ESM
ESM-
1vEVE
Alpha
Misse
nse
BRCA10.560.440.430.520.56
RASH0.450.490.40.480.46
MSH20.420.420.40.390.42
AI struggles to capture
the full complexity of
biological mechanisms
Abbreviations
-AF2 AlphaFold2
-MSA Multiple-Sequence-Alignment
-PLM Protein Language model
-DNN Dense neural network
-SOTA state-of-the-art
-FT fine-tuned
-GoF , Lof Gain/Loss-of-Function
-DMS Deep Mutational Scan
-MLM Masked Language Modeling
Figure 7: Survey of existing methods illustrating the trade-off between broadly applicable but less
precise models and highly precise models limited to their training protein. Notably, the latter can
produce high-quality predictions only for mutations within the same protein as the training DMS.
Despite this limitation, such models remain valuable, as DMS datasets typically focus on specific
protein domains and often contain incomplete data due to failed mutagenesis experiments.
Most existing models focus on pathogenicity prediction (i.e., how physiological or wildtype-similar
a mutation is) and use regression-based approaches. These models adopt a generalist strategy, scor-
ing all possible variants across the (human) proteome (Cheng et al., 2023). However, pathogenicity
predictors — whether trained on multiple DMS datasets, ClinVar annotations or physiological se-
quences — struggle to accurately predict the multi-faceted functional effects of specific mutations,
such as rare gain-of-function enzyme mutations (Figure 8). This limitation arises from the bio-
logical complexity and specificity required for such tasks, which cannot be reliably captured by
large-scale pre-training and the current architectures (Livesey & Marsh, 2023). However, clinical
decision-making often depends on understanding the precise functional effect of mutations (i.e.,
increase/decrease of a specific protein property) (Iyer et al., 2023). Thus, functional effect predic-
tors have been developed which are fine-tuned on DMS datasets achieving higher protein-specifity
18
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted February 7, 2025. ; https://doi.org/10.1101/2025.02.03.635741doi: bioRxiv preprint
and accuracy yet are limited to unseen mutations and regions in their training protein. Current ap-
proaches rely on non-fine-tuned embeddings or intricate multi-modal features (Marquet et al., 2022;
Zhong et al., 2024).
7.2 P ROTEIN MODELING AND PATHOGENICITY PREDICTION
Models like AlphaFold2 (AF2) predict protein structures from MSAs, capturing evolutionary in-
formation about residue interactions (Jumper et al., 2021) and Transformer-based Protein Language
Models (PLMs), like ESM-1b and ESM2, learn protein semantics by predicting masked amino acids
from evolutionary sequences (Rives et al., 2021; Lin et al., 2023; Rao et al., 2020). Because these
models learn sequence and structure physiology they are directly applied to predict the lack thereof
in form of the likelihood ratio of a mutant and wildtype residue (e.g., AlphaMissense, EVE building
on MSAs (Cheng et al., 2023; Frazer et al., 2021) and pre-trained PLMs like ESM-1v (Meier et al.,
2021; Brandes et al., 2023)). Some methods refine their training data with DMSs, which offer suf-
ficient signal for pathogenicity despite heterogeneous properties across different DMSs. Examples
include fine-tuning ESM-1v on 25 DMSs with a Normalized Log-odds Ratio (NLR) head (Lafita
et al., 2024) and combining EVE, ESM-1v, and AF2 features in a regression model (Jagota et al.,
2023). However, these methods struggle with multi-directional functional effects, particularly for
Gain-of-Function mutations in DMSs like SNCA (Livesey & Marsh, 2023). In summary, while
pathogenicity models effectively distinguish benign and pathogenic mutations, they fall short in
predicting multi-dimensional functional effects as demonstrated in Figure 8.
2
1
0 1 2 3 4 5
DMS score
(Has LoF, Neutral & GoF mutations)
0.0
0.5
1.0Density
SRC (3372 mutations)
4
3
2
1
0 1
DMS score
(Only has LoF and Neutral mutations
-> pathogenicity concept applicable)
0.0
0.5
1.0Density
BRCA1 (2086 mutations)
8
6
4
2
0 2 4 6
DMS score
(Has all three types of mutations)
0.0
0.1
0.2
0.3Density
MSH2 (16749 mutations)
1.5
1.0
0.5
0.0 0.5
DMS score
(Has rare GoF mutations)
0
1
2Density
HRAS (3135 mutations)
2
0 2 4
Score
0.0
0.2
0.4
0.6
0.8
1.0AM Pathogenicity Score
Spearman r = -0.58
R² = 0.09
MSE = 0.05
BME = 1.58
5
4
3
2
1
0 1 2
Score
0.00
0.25
0.50
0.75
1.00
1.25
1.50AM Pathogenicity Score
Spearman r = -0.56
R² = 0.38
MSE = 0.07
BME = 2.19
8
6
4
2
0 2 4 6
Score
0.0
0.2
0.4
0.6
0.8
1.0
1.2AM Pathogenicity Score
Spearman r = 0.42
R² = 0.18
MSE = 0.09
BME = 3.30
1.5
1.0
0.5
0.0 0.5
Score
0.0
0.2
0.4
0.6
0.8
1.0
1.2AM Pathogenicity Score
Spearman r = -0.46
R² = 0.09
MSE = 0.06
BME = 1.36
low AM score indicates neutral mutation,
but high AM score spans all functional effects
(incl. neutral)
low AM score indicates neutral mutation,
AND high AM score (roughly) indicates LOF mutation
-> pathogenicity prediction aligns with functional effect,
because there is not GoF
note that all metrics fail to reflect this decent performance
No useful predictions:
high & low AM score can indicate any effect
low score indicates neutral variant,
but high score spans all possible effects
0.0 0.2 0.4 0.6 0.8 1.0
AM pathogenicity
0
50Density
0.0 0.2 0.4 0.6 0.8 1.0
AM pathogenicity
0.0
2.5Density
0.0 0.2 0.4 0.6 0.8 1.0
AM pathogenicity
0
20Density
0.2 0.4 0.6 0.8 1.0
AM pathogenicity
0
50Density
AlphaMissense (AM) Pathogenicity vs. Functional Score from DMS:
Pathogenicity Predictors struggle to discern the bi-directional Functional Effect
Figure 8: SOTA pathogenicity predictor AlphaMissense on DMS data. Note that the DMSs some-
times not cover the entire protein sequence.
7.3 A BLATION AND MODEL ANALYSIS
Layer Probing. To investigate how the number of trainable layers affects performance, we retrained
optimized ESM-Effect with a descending number of layers frozen: the results show that the number
of frozen layers has no impact on test performance, as long as at least one layer remains unfrozen,
allowing the model to adapt to the specific task (cf. Figure 9). Given that a single unfrozen layer
can suffice for fine-tuning, we further explored whether its position within the network affects per-
formance: the test performance remains consistent regardless of the unfrozen layer’s position. Even
when only the first layer (immediately after the embedding layer) is unfrozen, it can still influence
the subsequent layers, enabling the model to produce informative embeddings for the regression
head at the final layer.
19
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted February 7, 2025. ; https://doi.org/10.1101/2025.02.03.635741doi: bioRxiv preprint
Figure 9: Ablation study of ESM-Effect: Fine-Tuning and Layer probing. Ablating Transformer
parts of optimized ESM-Effect on 3 SNCA seeds.
Transformer Parts Ablation. To investigate which components of the Transformer architecture
contribute most to performance, multiple models were trained with specific parts of the last two lay-
ers unfrozen. These include feed-forward layers, attention mechanisms, and individual components
of the attention module (i.e., key, query, value and output projection layers). Performance (cf. Fig-
ure 9) increases progressively, starting from the embedding layer, followed by key, query, value and
output projections, then the feed-forward and attention layers, and finally, the full last two layers.
This analysis suggests that ESM2 does not encode mutation-specific knowledge in individual layers,
as it does for structural features such as contacts and binding sites (Vig et al., 2020). Fine-tuning
performance is largely invariant to the position or number of fine-tuned layers, indicating that adap-
tation likely arises from task-specific tuning of the overall embeddings rather than mutation-specific
mechanisms. Notably, the differences observed across Transformer components demonstrate the
parameter efficiency of multi-head self-attention, which achieves competitive performance with ap-
proximately half the parameters of the feed-forward layers.
7.4 E XPERIMENTS WITH TEST-TIME -TRAINING (TTT)
As Bushuiev et al. (2024) showed, fine-tuning a pre-trained PLM backbone on a specific protein
sequence that is used for a given inference task improves performance (Bushuiev et al., 2024). For
instance, unsupervised mutation pathogenicity prediction from PLMs without a regression head
benefited from TTT. Here, we sought to apply this technique to ESM-Effect using a similar approach
for supevised functional effect prediction: first we customize (i.e., fine-tune) the ESM2 backbone on
the protein sequence of the DMS. Then we train the backbone with the ESM-Effect head on top on
a DMS. To customise the 35M ESM2 model, we the hyperparamter settings and training duration
recommended by Bushuiev et al. (2024) for 35M ESM2. However, this led to rapid overfitting
to the DMS sequence in this case: for the target DMS sequence and another (non-DMS protein
family) sequence, we monitored the percentage of correctly predicted tokens and their probability
when predicting each token in the sequence individually (with a mask for that token). We used this
strategy to adjust the learning rate to maintain accuracy of the non-related sequence while achieving
increased accuracy on the TTT/DMS sequence (cf. Figure 10). Based on the results we selected 1e-5
as optimal learning rate, customized the ESM2 backbone and trained ESM-Effect on three seeds of
the SNCA DMS.
Experiments with SNCA (seeds 0–2) reveal only minor performance differences between the non-
TTT and TTT models, depending on the metric used. Consequently, no significant benefit from TTT
is observed for the SNCA protein.
20
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted February 7, 2025. ; https://doi.org/10.1101/2025.02.03.635741doi: bioRxiv preprint
Test-Time-Training (TTT) Analysis: TTT-Sequence vs. General
Performance Tradeoff - Optimizing the learning rate
Figure 10: Top: Customizing ESM2 backbone on SNCA sequence while maintaining general knowl-
edge to prevent catastrophic forgetting. Bottom: Mixed results from TTT fine-tuning on SNCA
DMS.
7.5 G ENERALIZATION TEST ON DISTINCT DMS
To examine whether ESM-Effect learns transferable features from one member of a protein family to
another, we trained the model on the Glucokinase (GCK) DMS (with a 20% test split) and evaluated
its performance on both the held-out GCK test set and an independent DMS from the SRC tyrosine
kinase (Ahler et al., 2019).
Figure 11: Catalytic score distributions for the
GCK training and the SRC kinase testing DMS.
The distinct distributions highlight the violation of
the I.I.D. (independent and identically distributed)
assumption. SRC DMS scores were min-max
scaled to align with the range of GCK DMS val-
ues.
Note that the distribution of the catalytic activ-
ity scores for both DMS is distinct, thus the
i.i.d. (independent and identically distributed)
assumption does not hold true (cf. Histogram
Figure 11).
As illustrated in Figure 11, the catalytic activity
score distributions for the GCK and SRC DMS
datasets are distinct, violating the I.I.D. as-
sumption fundamental to many machine learn-
ing models. This is biologically expected:
despite both proteins being kinases, differ-
ences in their binding pockets and catalytic do-
mains reflect their specialization for different
substrates. Consequently, generalization from
GCK to SRC is challenging.
While ESM-Effect achieves reasonable perfor-
mance on the GCK test split (Spearman ρ =
0.67), it performs poorly on the SRC DMS
(Spearman ρ = 0.03), underscoring the limited
transferability of learned features across struc-
turally and functionally divergent kinases.
21
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted February 7, 2025. ; https://doi.org/10.1101/2025.02.03.635741doi: bioRxiv preprint
7.6 D ISSECTING THE NOTION OF A
WILDTYPE SEQUENCE
Over the course of ongoing evolution many different variants of sequences evolve and are selected
for fitness. Thus, one fixed, unique ”wild-type” sequence does not exist. Only different versions
of sequences exist which have different properties. The term ”mutation” and ”variant” build on the
arguable existence of one unique, static ”wild-type” sequence in which one amino-acid is substituted
forming the mutant sequence. Nonetheless, a physiological, natural sequence space exists compris-
ing many functionally and fitness-regarding equivalent ”wild-type” sequences which are curated in
databases like UniProt (The UniProt Consortium et al., 2023), UniRef or SwissProt (Suzek et al.,
2007; Boeckmann, 2003). These databases typically list one fixed, reference/”wild-type” sequence
but also other isoforms. And different amino acid alterations in these physiological sequences may
be viewed as mutations in contexts like precision medicine, where the wildtype sequence (space) for
a given oncogene is established. In this light, the task of variant pathogenicity prediction equates to
carving out the edges of the physiological sequence space. So the notion of one unique wild-type
sequence is less applicable to variant pathogenicity prediction models, since the models learn a no-
tion of physiological sequence spaces to which they compare a given sequence at inference. Yet they
require a reference sequence (one version of the physiological wildtype) to compare the likelihood
of the variant amino acid to: There is no effect without a reference to compare the effect to. The
same logic applies to supervised, specialists models trained on DMSs. While we train models that
only take the mutated sequence as input to predict the DMS score, the DMS score itself is being
calculated by comparing the property of the cell expressing the mutant sequence to cells expressing
the reference sequence. In general, variant prediction is not possible without a reference sequence
(as part of the physiological sequence space).
7.7 E MBEDDING ANALYSIS
Seeking to understand how fine-tuned ESM2 embeddings compare to baseline ESM2 embeddings
- the reason ESM-Effect outperforms PreMode - we obtained the embeddings for 100 GCK DMS
test mutations from both models and analyzed them using the UMAP dimensionality reduction tech-
nique. However, there are no clusters and coloring the data points according to their catalytic activity
does not show any relationship either. This might be due to the regression head’s importance in ex-
tracting meaningful features from the embedding (since it is trained with an order-of-magnitude
higher learning rate).
7.8 T RAINING MEAN EMBEDDING MODELS
Fine-tuning ESM2 models with a regression head using the mean sequence embedding presents
unique challenges that do not arise when using the mutation position head. Notably, these issues are
specific to the PTEN stability and enzyme activity DMS datasets and are not observed for SNCA
or NUDT15. Training with the mean embedding often exhibits instability, characterized by spiking
losses and abrupt fluctuations in performance. Additionally, convergence is slow, requiring more
than 20–30 epochs, because the mean embedding condenses fine-grained information into a less
informative representation, making it harder for the model to capture fine-grained mutation effects.
Furthermore, the gradients from the regression head propagate less directly through the mean oper-
ation to the ESM2 model, compared to using the mutation position embedding, where the gradients
flow directly from the head to the model parameters. This instability mainly applies to fine-tuning the
full ESM2 model on the PTEN enzyme activity DMS compared to frozen or LoRA-based models.
Therefore the PTEN enzyme activity comparison in Figure 2 is lacking given our limited compute
resources in order to train enough models for enough epochs in the fully unfrozen setups.
7.9 M ETHODS
7.9.1 T RAINING
A split the learning rate approach was used for ESM2 backbone and the regression head: ESM2 pa-
rameters are updated with a learning rate of 1e-5 and prediction head parameters by 1e-4 multiplied
by the batch size (8), respectively. Dropout rate is set to 20% and we train for 10 epochs with a one
cycle learning rate scheduler. AdamW was used with β1 = 0.9, β2 = 0.999, ϵ = 1e−8 and weight
22
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted February 7, 2025. ; https://doi.org/10.1101/2025.02.03.635741doi: bioRxiv preprint
decay coefficient = 0.01. Training time for a DMS with 3k mutations and sequence length of 400
(e.g., PTEN) for 10 epochs is roughly 3
4 hour on a NVIDIA T4 GPU depending on evaluation and
monitoring.
7.9.2 E STIMATING TRAINING TIME AND RESOURCES OF PREMODE AND ESM-E FFECT
While we were not able to run PreMode due to limited resources, the authors noted that the average
DMS fine-tuning time for PreMode is 30 minutes on an NVIDIA A40 GPU which has 74.8 TFLOPS.
ESM-Effect trains between 20 minutes and one hour on an NVIDIA T4 GPU with 8.1 TFLOPS.
Correcting for the performance difference, PreMode trains for 4.5 T4 hours, while ESM-Effect trains
for 0.67 hours (average). The ESM-Effect Speed Implementation caches the embeddings from the
10th layer in the first epoch and only performs forward-passes (and back-propagations) on the last
two layers. Thus the average training time is about 5 minutes - an 8.4x speedup over ESM-Effect
optimized and a 56.3x speedup over PreMode. Note that ESM-Effect can be further sped-up with
FlashAttention (Dao et al., 2022) but requires NVIDIA Ampere GPUs or newer and does not work
on Tesla GPUs.
While ESM-Effect only requires the model weights (ca. 200 MB), PreMode requires pre-computed
ESM2 650M embeddings, AF2 structures and MSAs which amounts to roughly 60 gigabytes of
compressed tar files. However, the memory requirements of ESM-Effect can increase during training
depending on batch size and sequence length.
7.9.3 D ATA
We used the same DMSs as in PreMode to compare performance: the same 20 percent test split
with five (or three) different seeds was used for random splitting. Note that when using data from
the PreMode repository the same .csv file contains scores for all properties of the DMS if there
are multiple. As the score column names are not indicative of the measurement, and the same
measurement type has different score column indices for different datasets we specify them here:
Protein column name for enzyme activity
SNCA score.1
CYP2C9 score.1
NUDT15 score.2
ASPA score.2
GCK score.1
PTEN score.2
Table 2: Mapping of proteins to column names containing enzyme activity scores.
To evaluate generalization from training on GCK we use a DMS of the SRC kinase from MaveDB
containing 3372 mutations (Ahler et al., 2019; Rubin et al., 2021). To adjust the scale of the score
measurement from the SRC DMS to GCK DMS we use min-max scaling.
23
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted February 7, 2025. ; https://doi.org/10.1101/2025.02.03.635741doi: bioRxiv preprint