TaxonMatch: taxonomic integration and tree construction from heterogeneous biological databases

preprint OA: closed
Full text JSON View at publisher
AI-generated deep summary by claude@2026-07, 2026-07-05 · read from full text

This paper describes TaxonMatch, a tool for integrating taxonomic information across heterogeneous biological databases that use non-standard or evolving names, including resolving synonymy and correcting typographical and structural inconsistencies. Using examples, the authors show how it can align names to construct a shared backbone arthropod taxonomy across NCBI, GBIF, and iNaturalist, identify molecular data closest to a fossil, and match IUCN endangered species to available molecular data. A key limitation implied by the task focus is that the approach depends on taxonomic name matching and the quality/coverage of the participating databases rather than experimentally validating biological relationships. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Integrating taxonomic data across heterogeneous biological databases remains a major challenge in biodiversity research due to non-standardized nomenclature, incomplete synonym annotation, and inconsistencies in taxonomic hierarchies. These issues limit interoperability between key resources such as the Global Biodiversity Information Facility (GBIF), the National Center for Biotechnology Information (NCBI), and citizen science platforms such as iNaturalist. Here, we present TaxonMatch, a scalable and reproducible framework for taxonomic reconciliation and cross-database integration. The workflow combines string-based candidate generation using TF–IDF vectorization, supervised machine learning for match classification, and lineage-aware synonym resolution to align taxonomic entities across multiple sources. By integrating both declared and implicit equivalences, TaxonMatch resolves typographical variation, synonymy, and structural inconsistencies in taxonomic data. The framework produces a unified taxonomic structure in which equivalent entities are reconciled while preserving source-specific identifiers, provenance information, and hierarchical relationships. We evaluate its robustness across multiple classifiers and demonstrate its effectiveness in resolving ambiguous taxonomic cases that are not handled by traditional matching approaches. We illustrate the applicability of TaxonMatch through three use cases: the construction of a unified arthropod taxonomy integrating GBIF, NCBI, and iNaturalist data; the identification of closest extant relatives of fossil taxa with molecular information; and the integration of genomic resources with conservation data from the IUCN Red List. These applications highlight the ability of the workflow to support the integration of ecological, genomic, and paleontological datasets. TaxonMatch provides a flexible and generalizable solution for taxonomic data integration, enabling the construction of coherent and interoperable biodiversity datasets for downstream analyses in ecology, evolution, and conservation biology.
Full text 1,175 characters · extracted from oa-doi-fallback · click to expand
Abstract Integrating taxonomic data from various sources presents a significant challenge in the study of biodiversity research, due to non-standardized nomenclature and evolving species classifications. Discrepancies between major repositories like the Global Biodiversity Information Facility (GBIF) and the National Center for Biotechnology Information (NCBI), as well as citizen science platforms such as iNaturalist, lead to fragmented and sometimes inaccurate biological data. We present TaxonMatch, a tool designed to address these challenges. TaxonMatch aligns taxonomic names, resolves synonymy, and corrects typographical and structural inconsistencies across databases. We show how it can be used to build a common backbone arthropod taxonomy over NCBI, GBIF and iNaturalist, to find the closest molecular data to a given fossil, and to identify IUCN endangered species with molecular data. TaxonMatch provides a cohesive taxonomic framework and a consistent taxonomic backbone, and can be applied to any taxonomic source. The tool is available at https://github.com/MoultDB/TaxonMatch. Competing Interest Statement The authors have declared no competing interest.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00