In-Depth Benchmarking of DIA-type Proteomics Data Analysis Strategies Using a Large-Scale Benchmark Dataset Comprising Inter-Patient Heterogeneity

preprint OA: closed CC-BY-4.0
📄 Open PDF View at publisher

Abstract

Abstract An overwhelming number of proteomics software tools and algorithms have been published for different steps of Data Independent Acquisition analysis of clinical samples. Nonetheless, there is still a lack of comprehensive benchmark studies evaluating which combinations of those isolated components perform best.Here, we used 92 lymph nodes from distinct patients to create a unique benchmark dataset representing real-world inter-individual heterogeneity. The publicly available dataset comprises 118 LC-MS/MS runs with > 12 million MS2 spectra and allowed us to objectively evaluate how well different combinations of spectral libraries, DIA software, sparsity reduction, normalization and statistical tests can detect differentially abundant proteins, while also taking sample size into account.Evaluation of 2 million data analysis workflows showed that a gas phase fractionation refined spectral library in combination with DIA-NN and Significance Analysis of Microarrays reliably detected differentially abundant proteins. Furthermore, DIA-NN and Spectronaut robustly avoided the false detection of truly absent proteins.*KF and EB share first authorship. CK and OS share last authorship.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0