Abstract
Eukaryotic genome annotation is currently bottlenecked by limitations in the generality, scalability and accuracy of computational methods. Deep learning approaches have recently achieved large improvements in ab initio gene prediction accuracy. We extend the deep learning-based ab initio gene predictor Tiberius beyond mammals by training lineage-specific models for Mesangiospermae, Fungi, Vertebrata, Insecta, Chlorophyta and Bacillariophyta. Across a benchmark of 33 species, Tiberius consistently achieves higher accuracy than the other evaluated ab initio methods, Helixer and ANNEVO, while also having the fastest runtimes overall. Compared with BRAKER3, which incorporates RNA-Seq and protein evidence, Tiberius approaches state-of-the-art accuracy in Mesangiospermae, Fungi, Bacillariophyta and Chlorophyta, while being on average 80 times faster when using a GPU. Availability and implementation https://github.com/Gaius-Augustus/Tiberius
Full text
1,036 characters
· extracted from
oa-doi-fallback
· click to expand
Abstract
Eukaryotic genome annotation is currently bottlenecked by limitations in the generality, scalability and accuracy of computational methods. Deep learning approaches have recently achieved large improvements in ab initio gene prediction accuracy. We extend the deep learning-based ab initio gene predictor Tiberius beyond mammals by training lineage-specific models for Mesangiospermae, Fungi, Vertebrata, Insecta, Chlorophyta and Bacillariophyta. Across a benchmark of 33 species, Tiberius consistently achieves higher accuracy than the other evaluated ab initio methods, Helixer and ANNEVO, while also having the fastest runtimes overall. Compared with BRAKER3, which incorporates RNA-Seq and protein evidence, Tiberius approaches state-of-the-art accuracy in Mesangiospermae, Fungi, Bacillariophyta and Chlorophyta, while being on average 80 times faster when using a GPU.
Availability and implementation https://github.com/Gaius-Augustus/Tiberius
Competing Interest Statement
The authors have declared no competing interest.
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.