⚙
AI-generated deep summary
by claude@2026-06, 2026-06-24
· read from full text
ⓘ
The paper developed a hybrid experimental–computational framework combining high-throughput molecular assays with active learning to study intrinsically disordered fungal transcriptional activators and quantify the strength of their activation domains across evolutionary space. Using ADhunter, a high-capacity regression model, the authors analyzed 7.8 million proteins from 2,400 fungal genomes and functionally characterized 9,836 activation domains from 1,071 genomes, expanding representation by 15.5-fold and producing functional annotations for thousands of proteins from non-model fungi. A key stated limitation is that existing datasets are biased toward model organisms, motivating their sampling strategy, and the authors emphasize model generalizability improvements from non-model genomes. This paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.
Abstract
Biological discovery and design are increasingly being guided by predictive models in place of costly experimentation. However, existing datasets are often biased by overrepresentation from model organisms, leading to failures in evolutionary studies of non-model species. We present a hybrid framework that leverages high-throughput molecular assays and active learning to quantify biological properties across evolutionary space. We focus on transcriptional activators, which contain activation domains (ADs) that promote gene expression. ADs are intrinsically disordered and poorly conserved, which limits their study using comparative genomics. Here, we developed ADhunter, a high-capacity regression model that outperforms state-of-theart algorithms in identifying and quantifying the strength of transcriptional activators. Model uncertainty was used to guide evolutionary sampling across 7.8 million proteins from 2,400 fungal genomes. We functionally characterized 9,836 ADs from 1,071 fungal genomes, providing a 15.5-fold expansion in genome representation compared to existing datasets. Comprehensive sampling from non-model genomes improved model generalizability and provides the first functional annotation for 3,416 proteins from 670 non-model fungi. Model interpretability analysis aligns with the biophysical model of AD function and reveals novel, underrepresented protein codes, highlighting the importance of sampling from non-model organisms to build evolutionarily robust models for predicting biological properties.
Full text
2,014 characters
· extracted from
oa-doi-fallback
· click to expand
Abstract
Biological discovery and design are increasingly being guided by predictive models in place of costly experimentation. However, existing datasets are often biased by overrepresentation from model organisms, leading to failures in evolutionary studies of non-model species. We present a hybrid framework that leverages high-throughput molecular assays and active learning to quantify biological properties across evolutionary space. We focus on transcriptional activators, which contain activation domains (ADs) that promote gene expression. ADs are intrinsically disordered and poorly conserved, which limits their study using comparative genomics. Here, we developed ADhunter, a high-capacity regression model that outperforms state-of-theart algorithms in identifying and quantifying the strength of transcriptional activators. Model uncertainty was used to guide evolutionary sampling across 7.8 million proteins from 2,400 fungal genomes. We functionally characterized 9,836 ADs from 1,071 fungal genomes, providing a 15.5-fold expansion in genome representation compared to existing datasets. Comprehensive sampling from non-model genomes improved model generalizability and provides the first functional annotation for 3,416 proteins from 670 non-model fungi. Model interpretability analysis aligns with the biophysical model of AD function and reveals novel, underrepresented protein codes, highlighting the importance of sampling from non-model organisms to build evolutionarily robust models for predicting biological properties.
Competing Interest Statement
J.D.K. has financial interests in Amyris, Ansa Biotechnologies, Apertor Pharma, Berkeley Yeast, BioMia, Demetrix, Lygos, Napigen, ResVita Bio, and Zero Acre Farms. P.M.S. has financial interests in BasidioBio.
Footnotes
The title has been updated; author affiliations have been updated for completeness; the font has been changed from serif to sans serif; an error where the taxonomical labels were not shown in Figure 2C has been fixed.
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.