Empirical-Bayes and Bayesian Hierarchical Modelling for Missingness and Differential Expression in Proteomics

preprint OA: closed
Full text JSON View at publisher

Abstract

Mass spectrometry-based label-free proteomics data often suffer from missing values, especially for low-abundance proteins or when a protein is completely absent in one condition. This makes it challenging to estimate fold changes reliably and perform downstream analyses. Traditional imputation methods often show inconsistent performance across datasets and they typically treat imputed values as fixed rather than uncertain. This can lead to an underestimation of variability in downstream analyses. To address those problems, we present a hierarchical model that accounts for both observed protein intensities and patterns of missing data. Missing values are modelled as left-censored observations below protein-specific detection limits, reflecting the limited sensitivity of the instrument, or being missing with the probability of an intensity dependent manner. Our proposed model captures structure at multiple levels: intensity-level measurements, group-level effects (e.g., experimental conditions), and protein-level variation. To estimate model parameters, we employ an empirical Bayes framework to infer hyperparameters across proteins and use Markov Chain Monte Carlo (MCMC) methods to sample parameters from the posterior distribution. Our pipeline avoids the need for imputing missing values and is designed to produce more reliable fold-change estimates and uncertainty measures. We benchmark our method against existing approaches and demonstrate that it provides more accurate, stable, and robust estimates for differential expression analysis. Abstract Figure
Full text 2,230 characters · extracted from oa-doi-fallback · click to expand
Abstract Mass spectrometry-based label-free proteomics data often suffer from missing values, especially for low-abundance proteins or when a protein is completely absent in one condition. This makes it challenging to estimate fold changes reliably and perform downstream analyses. Traditional imputation methods often show inconsistent performance across datasets and they typically treat imputed values as fixed rather than uncertain. This can lead to an underestimation of variability in downstream analyses. To address those problems, we present a hierarchical model that accounts for both observed protein intensities and patterns of missing data. Missing values are modelled as left-censored observations below protein-specific detection limits, reflecting the limited sensitivity of the instrument, or being missing with the probability of an intensity dependent manner. Our proposed model captures structure at multiple levels: intensity-level measurements, group-level effects (e.g., experimental conditions), and protein-level variation. To estimate model parameters, we employ an empirical Bayes framework to infer hyperparameters across proteins and use Markov Chain Monte Carlo (MCMC) methods to sample parameters from the posterior distribution. Our pipeline avoids the need for imputing missing values and is designed to produce more reliable fold-change estimates and uncertainty measures. We benchmark our method against existing approaches and demonstrate that it provides more accurate, stable, and robust estimates for differential expression analysis. Competing Interest Statement The authors have declared no competing interest. 6 Data and Code availability The proteomics datasets analyzed during this study are publicly available in the PRIDE repository under identifiers PXD062621, PXD068192, PXD072249, PXD009815 and PXD004352. Our implementation is provided as the R package MissBayes, which is currently available on GitHub at: https://github.com/lmcdbd/MissBayes. The package supports multiple input formats including log-intensity matrices, data frames, and QFeatures objects. Source code and all analysis scripts used in this study are available at: https://github.com/lmcdbd/MissBayes-DataAnalysis.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00