GROQ-seq Datasets Across Transcription Factors (LacI, RamR, VanR), T7 RNA Polymerase and TEV Protease
preprint
OA: closed
CC-BY-ND-4.0
Abstract
Predicting any protein’s function from its sequence alone would be a significant breakthrough in molecular biology. Although machine learning approaches have sought to tackle this, their limited generalizability reflects the absence of sufficiently large, open, diverse, and unified datasets. To address this data gap, we developed a high-throughput experimental platform called GROQ-seq ( Gro wth-based Q uantitative Seq uencing). In GROQ-seq, a protein’s function can be linked to a sequencing-based readout that enables scalable characterization of large variant libraries in Escherichia coli . Here, we present pilot datasets demonstrating its performance across three distinct protein function classes: transcription factors, polymerases, and proteases. The objective of this report is to present the datasets and to provide users with a clear and transparent characterization of their properties, including both the strengths and limitations.
My notes (saved in your browser only)
Citation neighborhood (sparse)
Too few in-corpus citations on either side for a chart; here are the lists.
Cites (4)
- GROQ-seq Enables Cross-site Reproducibility for High-Throughput Measurement of Protein Function 2026
- Natural protein structures have evolved exceptional robustness to mutations 2025
- Using enantioselective biosensors to evolve asymmetric biocatalysts 2025
- FLIGHTED: Inferring Fitness Landscapes from Noisy High-Throughput Experimental Data 2024
References (39)
- FLIGHTED: Inferring Fitness Landscapes from Noisy High-Throughput Experimental Data via crossref
- GROQ-seq Enables Cross-site Reproducibility for High-Throughput Measurement of Protein Function via crossref
- Natural protein structures have evolved exceptional robustness to mutations via crossref
- Using enantioselective biosensors to evolve asymmetric biocatalysts via crossref
- doi:10.1016/j.jmb.2011.03.057 via crossref
- doi:10.1016/j.bpc.2011.06.005 via crossref
- doi:10.7717/peerj.498 via crossref
- doi:10.1016/j.cels.2018.02.004 via crossref
- doi:10.1002/pro.3695 via crossref
- doi:10.1126/science.271.5253.1247 via crossref
- doi:10.1038/ncomms3078 via crossref
- doi:10.1093/nar/gkad1011 via crossref
- doi:10.1006/jmbi.1996.0479 via crossref
- doi:10.15252/msb.202010179 via crossref
- doi:10.1016/j.cels.2024.03.002 via crossref
- doi:10.1016/j.jmb.2015.09.015 via crossref
- doi:10.1038/s41589-018-0168-3 via crossref
- doi:10.1021/acssynbio.9b00071 via crossref
- doi:10.1128/aem.00939-22 via crossref
- doi:10.1038/19999 via crossref
- doi:10.3389/fmolb.2024.1504876 via crossref
- doi:10.1016/j.jmb.2007.05.070 via crossref
- doi:10.1038/nature01129 via crossref
- doi:10.1016/s0092-8674(04)00059-5 via crossref
- doi:10.1038/s41598-017-06586-1 via crossref
- doi:10.1016/j.bpj.2017.04.039 via crossref
- doi:10.1006/jmbi.1994.1205 via crossref
- doi:10.1016/s0014-5793(98)01484-7 via crossref
- doi:10.1021/bi962397i via crossref
- doi:10.1016/j.jbiotec.2005.08.006 via crossref
- doi:10.1038/s41592-019-0665-7 via crossref
- doi:10.1093/protein/14.12.993 via crossref
- doi:10.1002/prot.26693 via crossref
- doi:10.7554/elife.32472 via crossref
- doi:10.1074/jbc.m207224200 via crossref
- doi:10.1093/bioinformatics/btx655 via crossref
- doi:10.5281/zenodo.19451958 via crossref
- doi:10.5281/zenodo.19551443 via crossref
- doi:10.5281/zenodo.19551807 via crossref
Source provenance
- crossref
- last seen: 2026-05-25T01:00:13.309563+00:00
- europepmc
- last seen: 2026-05-20T01:45:00.602351+00:00
- unpaywall
- last seen: 2026-05-21T05:10:58.409756+00:00
License: CC-BY-ND-4.0