EasyPseudogene: an easy-to-use and multithreaded pipeline for pseudogene detection

preprint OA: closed CC-BY-NC-ND-4.0

Abstract

Pseudogenes are recognized as essential components for reconstructing adaptive evolutionary trajectories and understanding genomic remodeling. However, identifying these sequences in large eukaryotic genomes remains technically challenging due to fragmented workflows, complex manual configurations, and the lack of high-performance, parallelized tools capable of processing rapidly growing data volumes. We present EasyPseudogene, an automated and multithreaded pipeline designed for the end-to-end identification of pseudogenes across diverse eukaryotic lineages. Unlike traditional self-mapping tools that often fail to detect unitary pseudogenes when functional counterparts are absent, EasyPseudogene introduces an inter-species reference-driven paradigm that utilizes high-quality proteomes as probes to scan target genomes for evolutionary relics. The pipeline employs a modular “hierarchical screening and precision detection” architecture, integrating high-speed homology searching via MMseqs2 and spliced alignments via miniprot with high-fidelity, three-frame alignments using GeneWise. Performance benchmarking on cetacean genomes demonstrates that EasyPseudogene can replicate known gene loss events, such as the functional decay of the ADRB3 gene, with 100% consistency relative to established manual workflows. By encapsulating complex comparative genomics logic into a standardized framework with interactive HTML visualization for mutation auditing at single-base resolution, EasyPseudogene provides a versatile and reproducible solution for marine ecology and evolutionary research.
Full text 1,687 characters · extracted from oa-html · click to expand
Abstract Pseudogenes are recognized as essential components for reconstructing adaptive evolutionary trajectories and understanding genomic remodeling. However, identifying these sequences in large eukaryotic genomes remains technically challenging due to fragmented workflows, complex manual configurations, and the lack of high-performance, parallelized tools capable of processing rapidly growing data volumes. We present EasyPseudogene, an automated and multithreaded pipeline designed for the end-to-end identification of pseudogenes across diverse eukaryotic lineages. Unlike traditional self-mapping tools that often fail to detect unitary pseudogenes when functional counterparts are absent, EasyPseudogene introduces an inter-species reference-driven paradigm that utilizes high-quality proteomes as probes to scan target genomes for evolutionary relics. The pipeline employs a modular “hierarchical screening and precision detection” architecture, integrating high-speed homology searching via MMseqs2 and spliced alignments via miniprot with high-fidelity, three-frame alignments using GeneWise. Performance benchmarking on cetacean genomes demonstrates that EasyPseudogene can replicate known gene loss events, such as the functional decay of the ADRB3 gene, with 100% consistency relative to established manual workflows. By encapsulating complex comparative genomics logic into a standardized framework with interactive HTML visualization for mutation auditing at single-base resolution, EasyPseudogene provides a versatile and reproducible solution for marine ecology and evolutionary research. Competing Interest Statement The authors have declared no competing interest.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-NC-ND-4.0