DIA-BERT: pre-trained end-to-end transformer models for enhanced DIA proteomics data analysis

preprint OA: closed
📄 Open PDF View at publisher

Abstract

Data-independent acquisition mass spectrometry (DIA-MS) plays an increasingly important role in quantitative proteomics. Here, we introduce DIA-BERT, a software tool that leverages a transformer-based pre-trained artificial intelligence (AI) model for the analysis of DIA proteomics data. Over 276 million of high-quality peptide precursors extracted from existing DIA-MS files were used for training the identification model, while 34 million peptide precursors from synthetic DIA-MS files were used for training the quantification model. Compared to DIA-NN, DIA-BERT led to on average 54% more protein identifications and 37% more peptide precursors in five different human cancer (cervical cancer, pancreatic adenocarcinoma, myosarcoma, gallbladder cancer, and gastric carcinoma) sample sets with a high degree of quantitative accuracy. This study highlights the potential of utilizing pre-trained models and synthetic datasets to advance DIA proteomics analysis.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00