Introduction
Fungi play essential roles in ecosystems as decomposers, pathogens and symbionts,
and many taxa exhibit flexibility across these trophic modes ( Berbee et al. 2017 ). Some
species occupy multiple ecological strategies depending on host identity or
environmental context, while others remain specialised to a single mode ( Martin and Tan
2025). Understanding this trophic versatility is critical for biodiversity assessments,
ecosystem modelling and agricultural applications where a fungus’ ecological role affects
plant performance.
Trait databases, such as FUNGuild ( Nguyen et al. 2016 ) and FungalTraits ( Põlme et al.
2021), have advanced fungal functional annotation at scale, but manual literature
extraction remains a bottleneck. Human-curated classifications are time-intensive,
subjective and limited by the accessibility of trait-relevant language in publications.
Moreover, trait databases are valuable but can be limited in applicability, as annotations
are often performed at the genus or family level despite often substantial interspecific
variability ( Violle et al. 2015 ).
Natural language processing (NLP) offers a scalable way to extract trait-relevant
information directly from text. Transformer-based models, such as BioBERT, pretrained
on large biomedical corpora, excel at contextual understanding and have achieved state-
of-the-art results in various text classification tasks ( Lee et al. 2019 ). However, their use in
fungal ecology and trait data integration remains largely unexplored.
This pilot study tests the feasibility of fine-tuning BioBERT to classify fungal trophic modes
from abstracts. The workflow is designed for transparency and future scaling, providing a
reproducible pipeline that can complement or benchmark existing trait databases. By
linking automated text classification with open trait resources, this work demonstrates a
path towards more consistent, interoperable fungal functional data.
Related Work
Recent advances in domain-specific NLP illustrate the potential for scaling and refining
workflows like this one. BiodivBERT ( Abdelmageed et al. 2023 ) represents the first
pretrained language model tailored specifically for biodiversity research, achieving
significant gains in named entity recognition and relation extraction. Its development from
life-sciences corpora highlights how domain adaptation can substantially improve the
precision and recall of ecological information retrieval. Similarly, Cornelius et al. (2025)
demonstrated a machine-learning framework for extracting arthropod organismal traits,
translating unstructured text into a machine-actionable database (ArTraDB). Their
approach shows how targeted trait extraction can efficiently transform literature into
structured ecological data.
2
Author-formatted, not peer-reviewed document posted on 03/11/2025. DOI:
https://doi.org/10.3897/arphapreprints.e176591
Parallel work in plant functional ecology supports the broader utility of transformer-based
extraction for trait data. Domazetoski et al. (2025) developed a natural language pipeline
that automatically identifies both categorical and numerical plant traits from unstructured
descriptions with high precision and recall. This scalability across morphological, life
history and functional trait types underscores the potential to adapt such methods to
fungal traits, where similar data gaps persist.
Beyond ecological trait extraction, innovations in model architecture and pretraining
further extend applicability. BioT5 ( Pei et al. 2023 ) introduced cross-modal pretraining to
connect textual and molecular data using chemically informed representations, an
approach that could eventually allow integration of genomic or metabolomic predictors
into ecological models. ModernBERT ( Warner et al. 2025 ) provides a computationally
efficient encoder capable of handling long-context inputs, which is particularly relevant
for large-scale biodiversity corpora. Together, these efforts suggest practical pathways to
scale fungal trait text mining beyond small datasets, while maintaining interpretability and
reproducibility. Foundational work by Gu et al. (2021) also reinforces the importance of
domain-specific pretraining, demonstrating that models trained entirely within a target
domain outperform general-domain models adapted later.
References
• Abdelmageed N, Löffler F , König-Ries B (2023) BiodivBERT : a Pre-Trained Language
Model for the Biodiversity Domain. In: Yamaguchi A, et al. (Ed.) SWAT4HCLS. 62-71 pp.
URL: https://ceur-ws.org/Vol-3415/paper-7.pdf
• Berbee M, James T , Strullu-Derrien C (2017) Early Diverging Fungi: Diversity and Impact
at the Dawn of T errestrial Life. Annual Review of Microbiology 71 (1): 41‑60. https://
doi.org/10.1146/annurev-micro-030117-020324
• Bock B (2026) beabock/biobert_dualsolo: Reproducible BioBERT & BERT model
comparison (4 models, 5-fold CV). Zenodo https://doi.org/10.5281/zenodo.17343492
• Cornelius J, Detering H, Lithgow-Serrano O, Agosti D, Rinaldi F , Waterhouse R (2025)
From literature to biodiversity data: mining arthropod organismal traits with machine
learning. Biodiversity Data Journal 13 https://doi.org/10.3897/bdj.13.e153070
• Devlin J, Chang M, Lee K, T outanova K (2018) BERT : Pre-training of Deep Bidirectional
Transformers for Language Understanding. CoRR URL: http://arxiv.org/abs/1810.04805
• Domazetoski V , Kreft H, Bestova H, Wieder P , Koynov R, Zarei A, Weigelt P (2025)
Using large language models to extract plant functional traits from unstructured text.
Applications in Plant Sciences 13 (3). https://doi.org/10.1002/aps3.70011
• Gu Y , Tinn R, Cheng H, Lucas M, Usuyama N, Liu X, Naumann T , Gao J, Poon H (2021)
Domain-Specific Language Model Pretraining for Biomedical Natural Language
Processing. ACM Transactions on Computing for Healthcare 3 (1): 1‑23. https://doi.org/
10.1145/3458754
• Lee J, Y oon W, Kim S, Kim D, Kim S, So CH, Kang J (2019) BioBERT : a pre-trained
biomedical language representation model for biomedical text mining. Bioinformatics 36
(4): 1234‑1240. https://doi.org/10.1093/bioinformatics/btz682
• Martin F , T an H (2025) Saprotrophy-to-symbiosis continuum in fungi. Current Biology 35
(11). https://doi.org/10.1016/j.cub.2025.01.032
• Nguyen N, Song Z, Bates S, Branco S, T edersoo L, Menke J, Schilling J, Kennedy P
(2016) FUNGuild: An open annotation tool for parsing fungal community datasets by
ecological guild. Fungal Ecology 20: 241‑248. https://doi.org/10.1016/j.funeco.2015.06.006
• Paszke A, Gross S, Massa F , Lerer G, Bradbury J, Chanan G, Killeen T , Lin Z,
Gimelshein N, Chintala S, et al. (2019) PyT orch: An Imperative Style, High-Performance
Deep Learning Library. Advances in Neural Information Processing Systems 32:
8024‑8035.
8
Author-formatted, not peer-reviewed document posted on 03/11/2025. DOI:
https://doi.org/10.3897/arphapreprints.e176591
• Pedregosa F , Varoquaux G, Gramfort A, Michel V , Thirion B, Grisel O, Blondel M,
Prettenhofer P , Weiss R, Dubourg V , Vanderplas J, Passos A, Cournapeau D, Brucher M,
Perrot M, Duchesnay E (2011) Scikit-learn: Machine Learning in {P}ython. Journal of
Machine Learning Research 12: 2825‑2830.
• Pei Q, Zhang W, Zhu J, Wu K, Gao K, Wu L, Xia Y , Yan R (2023) BioT5: Enriching
Cross-modal Integration in Biology with Chemical Knowledge and Natural Language
Associations. Proceedings of the 2023 Conference on Empirical Methods in Natural
Language Processing https://doi.org/10.18653/v1/2023.emnlp-main.70
• Põlme S, Abarenkov K, Henrik Nilsson R, Lindahl B, Clemmensen KE, Kauserud H,
Nguyen N, Kjøller R, Bates S, Baldrian P , Frøslev TG, Adojaan K, Vizzini A, Suija A,
Pfister D, Baral H, Järv H, Madrid H, Nordén J, Liu J, Pawlowska J, Põldmaa K, Pärtel K,
Runnel K, Hansen K, Larsson K, Hyde KD, Sandoval-Denis M, Smith M, T oome-Heller M,
Wijayawardene N, Menolli N, Reynolds N, Drenkhan R, Maharachchikumbura SN,
Gibertoni T , Læssøe T , Davis W, T okarev Y , Corrales A, Soares AM, Agan A, Machado
AR, Argüelles-Moyao A, Detheridge A, de Meiras-Ottoni A, Verbeken A, Dutta AK, Cui B,
Pradeep CK, Marín C, Stanton D, Gohar D, Wanasinghe D, Otsing E, Aslani F , Griffith G,
Lumbsch T , Grossart H, Masigol H, Timling I, Hiiesalu I, Oja J, Kupagme J, Geml J,
Alvarez-Manjarrez J, Ilves K, Loit K, Adamson K, Nara K, Küngas K, Rojas-Jimenez K,
Bitenieks K, Irinyi L, Nagy L, Soonvald L, Zhou L, Wagner L, Aime MC, Öpik M, Mujica
MI, Metsoja M, Ryberg M, Vasar M, Murata M, Nelsen M, Cleary M, Samarakoon M,
Doilom M, Bahram M, Hagh-Doust N, Dulya O, Johnston P , Kohout P , Chen Q, Tian Q,
Nandi R, Amiri R, Perera RH, dos Santos Chikowski R, Mendes-Alvarenga R, Garibay-
Orijel R, Gielen R, Phookamsak R, Jayawardena R, Rahimlou S, Karunarathna S,
Tibpromma S, Brown S, Sepp S, Mundra S, Luo Z, Bose T , Vahter T , Netherway T , Yang
T , May T , Varga T , Li W, Coimbra VRM, de Oliveira VRT , de Lima VX, Mikryukov V , Lu Y ,
Matsuda Y , Miyamoto Y , Kõljalg U, T edersoo L (2021) FungalTraits: a user-friendly traits
database of fungi and fungus-like stramenopiles. Fungal Diversity 105 (1): 1‑16. https://
doi.org/10.1007/s13225-020-00466-2
• Python Software Foundation (2020) Python 3.9.0. URL: https://docs.python.org/3.9
• Violle C, Borgy B, Choler P (2015) Trait databases: misuses and precautions. Journal of
Vegetation Science 26 (5): 826‑827. https://doi.org/10.1111/jvs.12325
• Warner B, Chaffin A, Clavié B, Weller O, Hallström O, T aghadouini S, Gallagher A,
Biswas R, Ladhak F , Aarsen T , Adams GT , Howard J, Poli I (2025) Smarter, Better,
Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long
Context Finetuning and Inference. Proceedings of the 63rd Annual Meeting of the
Association for Computational Linguistics (Volume 1: Long Papers)2526‑2547. https://
doi.org/10.18653/v1/2025.acl-long.127
• Wolf T , et al. (2020) Transformers: State-of-the-Art Natural Language Processing.
Proceedings of the 2020 Conference on Empirical Methods in Natural Language
Processing: System Demonstrations. Online. Association for Computational Linguistics
URL: https://www.aclweb.org/anthology/2020.emnlp-demos.6
9
Author-formatted, not peer-reviewed document posted on 03/11/2025. DOI:
https://doi.org/10.3897/arphapreprints.e176591
Figure 1.
Workflow diagram of the classification pipeline. The diagram summarises the full end-to-end
process: literature search (Web of Science queries), manual curation (56 labelled abstracts),
preprocessing (text cleaning, tokenisation with truncation at 512 tokens and token-length QC),
stratified 5-fold cross-validation, model fine-tuning across four models (BERT-base-uncased,
BERT-base-cased, BioBERT v.1.1, BiodivBERT) with standardised hyperparameters,
evaluation (metrics, confusion matrices, learning curves) and outputs (fine-tuned models,
predictions).
10
Author-formatted, not peer-reviewed document posted on 03/11/2025. DOI:
https://doi.org/10.3897/arphapreprints.e176591
Figure 2.
Comparative model performance across four transformer-based architectures. Bar charts
show mean classification metrics (accuracy, precision, recall, F1-score) ± standard deviation
from stratified 5-fold cross-validation (n = 56 abstracts). BioBERT (biomedical domain-
adapted) and BERT-base-cased achieved statistically equivalent performance (~ 89%
accuracy), substantially outperforming BERT-base-uncased (~ 75%) and BiodivBERT (~ 77%).
Case sensitivity proved critical, with cased models outperforming uncased by ~ 15 percentage
points. Metrics are calculated as macro averages (unweighted mean across dual and solo
classes).
11
Author-formatted, not peer-reviewed document posted on 03/11/2025. DOI:
https://doi.org/10.3897/arphapreprints.e176591
Figure 3.
Aggregated confusion matrices for all four models (BioBERT , BERT-base-cased, BERT-base-
uncased, BiodivBERT) across all five folds (total 56 predictions per model). Each matrix shows
true labels (Solo = single trophic mode, Dual = multiple trophic modes) versus predicted labels,
allowing direct comparison of error patterns and class balance for each model. BioBERT and
BERT-base-cased show balanced performance, while uncased and BiodivBERT models
display more misclassifications. Colour intensity indicates prediction frequency; diagonal cells
represent correct predictions.
12
Author-formatted, not peer-reviewed document posted on 03/11/2025. DOI:
https://doi.org/10.3897/arphapreprints.e176591
Figure 4.
Training time comparison across models. BioBERT and BERT-base-cased completed 5-fold
cross-validation training in ~ 10-11 minutes, while BiodivBERT and BERT-base-uncased
required ~ 35 minutes. Differences likely reflect tokenisation efficiency and convergence
patterns rather than model size (all models have ~ 110M parameters). Faster convergence in
cased models correlates with higher classification accuracy, suggesting that case-preserving
tokenisation provides stronger learning signals for this taxonomic text classification task.
Training performed on NAU Monsoon HPC cluster (T esla K80 GPU, CUDA 11.4).
13
Author-formatted, not peer-reviewed document posted on 03/11/2025. DOI:
https://doi.org/10.3897/arphapreprints.e176591
Metric Value Note
Accuracy 89.4% ± 11.6% Fraction of correctly predicted labels
Precision 89.9% ± 11.5% Positive predictive value (macro average)
Recall 88.8% ± 12.4% True positive rate (macro average)
F1-Score 89.2% ± 12.0% Harmonic mean of precision and recall (macro average)
Table 1.
BioBERT classification performance on fungal trophic modes (5-fold CV). Precision, recall and F1-
score are reported as macro averages (unweighted mean across both classes), which is
appropriate for balanced binary classification and treats both classes equally regardless of support.
14
Author-formatted, not peer-reviewed document posted on 03/11/2025. DOI:
https://doi.org/10.3897/arphapreprints.e176591