Quantifying the ∼75-95% of Peptides in DIA-MS Datasets that were not Previously Quantified

preprint OA: closed CC-BY-ND-4.0
📄 Open PDF View at publisher

Abstract

We demonstrate an algorithm termed GoldenHaystack (GH) that, compared to the leading DIA-MS algorithm, (a) quantifies and identifies with better FDR accuracy the peptides found in FASTA search spaces (∼5-25% of analytes in DIA-MS datasets), (b) quantifies the remaining ∼75-95% of analytes that were previously unquantified, and (c) runs ∼40-200x faster (or ∼1-10x faster than the LC-MS). Specifically, without a FASTA or spectral library, GH can deconvolute and accurately quantify chimeric LC-MS spectra. The central idea that enables this claim is: for sufficiently sized projects (e.g., ≥ ∼50 LC-MS files), pairs of peptides that co-elute in one subset of LC-MS files do not exactly co-elute in a different subset of files. GH thus analyzes a project holistically: it uses multi -partite matching to match fragment ions across all samples, separates and regroups the fragment ions into unique analyte signatures, reduces stochastic noise, and then quantifies those unique analyte signatures.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-23T02:00:01.238055+00:00
License: CC-BY-ND-4.0