Functional Profiling of Thousands of Sequence-Diverse Protease Homologs with GROQ-seq
preprint
OA: closed
AI-generated summary
Growth-based Quantitative Sequencing (GROQ-seq) was used to measure function across an 11,722 member protease library, identifying sequence-divergent homologs that preserve native TEV protease activity.
One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works
Abstract
High-quality datasets that span broad sequence diversity are essential for understanding protein sequence–function relationships beyond local mutational landscapes. Here, we applied Growth-based Quantitative Sequencing (GROQ-seq) to measure function across an 11,722 member protease library, comprised of natural homologs and AI-shrunken variants. This library spans vast sequence diversity, with Levenshtein distances of up to 245 and a mean pairwise sequence identity of 41% to TEV protease S219V. We identified sequence-divergent TEV protease homologs that preserve function against the native TEV protease substrate. These findings reveal the robustness of protease activity across highly diverse sequences. Here, we demonstrate the aptitude of the GROQ-seq assay for screening large, diverse protein libraries for function, enabling efficient data generation at scale for training machine learning models across broad sequence landscapes.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.
Source provenance
- europepmc
- last seen: 2026-05-20T01:45:00.602351+00:00