uSort-M: Scalable isolation of user-defined sequences from diverse pooled libraries
preprint
OA: closed
Abstract
Advances in high-throughput sequencing and computational protein design are expanding the catalog of known protein sequences far more rapidly than they can be functionally characterized. Functional characterization through biochemical and biophysical assays often requires isolating variants to profile them individually, a process which is often laborious and expensive. To address this, we developed u ser-defined Sort ed M utants (uSort-M), which can rapidly isolate and identify individual variants from widely available pools of genes by leveraging automated cell sorting and long-read sequencing technologies. To develop the uSort-M pipeline, we first parsed a 328-member scanning mutagenesis library of a 300-bp gene. Direct comparison between short-read and long-read sequencing demonstrated that both methods resolve variants with high fidelity, enabling recovery of 96% of desired library members by sorting eight 384-well plates at fivefold lower cost than traditional synthesis. After optimizing for long-read sequencing, we demonstrated uSort-M’s generalizability to complex libraries by parsing a sequence- and length-diverse 500-member library, recovering 88% of variants from shallow oversampling (<3-fold). Library recoveries could be accurately predicted by simulating library uniformity, transformation number, sorting efficiency, and per-base error rates, and we extrapolated these simulations to predict the sampling depth needed to process large libraries containing thousands of members. To facilitate adoption, these simulations are packaged alongside data analysis and workflow management tools in an open-source Python toolkit with an interactive dashboard. By implementing standard instrumentation in a generalizable workflow, uSort-M provides an efficient and cost-effective solution for large library generation, thereby removing a key barrier to large-scale protein functional characterization.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.
Source provenance
- europepmc
- last seen: 2026-05-20T01:45:00.602351+00:00