Abstract
The escalating antimicrobial resistance crisis has intensified the urgent need for alternative antibacterial agents. Phage lysins exhibit potent bactericidal activity with low resistance potential, yet their large-scale discovery from rapidly expanding genomic resources remains limited, as existing computational methods often rely on sequence homology and few reproducible, independently validated, and practically deployable prediction frameworks are currently available. We present LysinFusion, a reproducible deep learning framework that integrates heterogeneous sequence features via a hybrid CNN–Transformer architecture for accurate lysin identification. Trained on a curated, de-redundant corpus (PHROG + inphared) and validated on 148 experimentally confirmed UniProt proteins, LysinFusion achieves an accuracy of 0.81, an AUC of 0.89, and an MCC of 0.62, outperforming the state-of-the-art DeepMineLys by up to 52.3% across metrics while reducing false positives by 64% (12 vs. 33). Ablation studies confirm that both CNN and Transformer modules are individually essential and that serial processing outperforms parallel alternatives. Model interpretability via occlusion and LIME aligns with established lysin biology, with attention concentrating on the N-terminal catalytic domain and decision boundaries reflecting characteristic charge-composition thresholds documented in Gram-negative lysin C-terminal regions. The framework, including all source code, curated datasets, and an accompanying web server, is freely available at https://github.com/sinuo560/LysinFusion .
Full text
1,941 characters
· extracted from
oa-doi-fallback
· click to expand
Abstract
The escalating antimicrobial resistance crisis has intensified the urgent need for alternative antibacterial agents. Phage lysins exhibit potent bactericidal activity with low resistance potential, yet their large-scale discovery from rapidly expanding genomic resources remains limited, as existing computational methods often rely on sequence homology and few reproducible, independently validated, and practically deployable prediction frameworks are currently available. We present LysinFusion, a reproducible deep learning framework that integrates heterogeneous sequence features via a hybrid CNN–Transformer architecture for accurate lysin identification. Trained on a curated, de-redundant corpus (PHROG + inphared) and validated on 148 experimentally confirmed UniProt proteins, LysinFusion achieves an accuracy of 0.81, an AUC of 0.89, and an MCC of 0.62, outperforming the state-of-the-art DeepMineLys by up to 52.3% across metrics while reducing false positives by 64% (12 vs. 33). Ablation studies confirm that both CNN and Transformer modules are individually essential and that serial processing outperforms parallel alternatives. Model interpretability via occlusion and LIME aligns with established lysin biology, with attention concentrating on the N-terminal catalytic domain and decision boundaries reflecting characteristic charge-composition thresholds documented in Gram-negative lysin C-terminal regions. The framework, including all source code, curated datasets, and an accompanying web server, is freely available at https://github.com/sinuo560/LysinFusion.
Competing Interest Statement
The authors have declared no competing interest.
Footnotes
In this revised version (v2), we have refined and condensed the Abstract and Introduction. We also added an online prediction tool, which is described in the new Section 3.6. Consequently, the contributor who built this tool has been added to the author list.
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.