LysinFusion: Integrating Multi-Feature Encoding and Hybrid CNN-Transformer Architecture for Phage Lysin Prediction

preprint OA: closed
Full text JSON View at publisher

Abstract

The escalating antimicrobial resistance crisis has intensified the urgent need for alternative antibacterial agents. Phage lysins exhibit potent bactericidal activity with low resistance potential, yet their large-scale discovery from rapidly expanding genomic resources remains limited, as existing computational methods often rely on sequence homology and few reproducible, independently validated, and practically deployable prediction frameworks are currently available. We present LysinFusion, a reproducible deep learning framework that integrates heterogeneous sequence features via a hybrid CNN–Transformer architecture for accurate lysin identification. Trained on a curated, de-redundant corpus (PHROG + inphared) and validated on 148 experimentally confirmed UniProt proteins, LysinFusion achieves an accuracy of 0.81, an AUC of 0.89, and an MCC of 0.62, outperforming the state-of-the-art DeepMineLys by up to 52.3% across metrics while reducing false positives by 64% (12 vs. 33). Ablation studies confirm that both CNN and Transformer modules are individually essential and that serial processing outperforms parallel alternatives. Model interpretability via occlusion and LIME aligns with established lysin biology, with attention concentrating on the N-terminal catalytic domain and decision boundaries reflecting characteristic charge-composition thresholds documented in Gram-negative lysin C-terminal regions. The framework, including all source code, curated datasets, and an accompanying web server, is freely available at https://github.com/sinuo560/LysinFusion .
Full text 1,941 characters · extracted from oa-doi-fallback · click to expand
Abstract The escalating antimicrobial resistance crisis has intensified the urgent need for alternative antibacterial agents. Phage lysins exhibit potent bactericidal activity with low resistance potential, yet their large-scale discovery from rapidly expanding genomic resources remains limited, as existing computational methods often rely on sequence homology and few reproducible, independently validated, and practically deployable prediction frameworks are currently available. We present LysinFusion, a reproducible deep learning framework that integrates heterogeneous sequence features via a hybrid CNN–Transformer architecture for accurate lysin identification. Trained on a curated, de-redundant corpus (PHROG + inphared) and validated on 148 experimentally confirmed UniProt proteins, LysinFusion achieves an accuracy of 0.81, an AUC of 0.89, and an MCC of 0.62, outperforming the state-of-the-art DeepMineLys by up to 52.3% across metrics while reducing false positives by 64% (12 vs. 33). Ablation studies confirm that both CNN and Transformer modules are individually essential and that serial processing outperforms parallel alternatives. Model interpretability via occlusion and LIME aligns with established lysin biology, with attention concentrating on the N-terminal catalytic domain and decision boundaries reflecting characteristic charge-composition thresholds documented in Gram-negative lysin C-terminal regions. The framework, including all source code, curated datasets, and an accompanying web server, is freely available at https://github.com/sinuo560/LysinFusion. Competing Interest Statement The authors have declared no competing interest. Footnotes In this revised version (v2), we have refined and condensed the Abstract and Introduction. We also added an online prediction tool, which is described in the new Section 3.6. Consequently, the contributor who built this tool has been added to the author list.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00