Do Gene Perturbation Predictions Match Real Effects? A Computational Test Using Public Data

preprint OA: closed CC-BY-4.0
🔓 Open OA copy View at publisher

Abstract

Computational gene perturbation methods—such as differential expression from knockout databases, network propagation algorithms, or perturbation autoencoders—are widely used to infer the transcriptional consequences of disrupting a gene. However, it remains unclear how well these in silico predictions align with empirically observed effects, particularly when predictions are derived from models that may overfit to training data. This study proposes a computational benchmarking framework to address this gap. Using publicly available perturbation datasets (e.g., CRISPR screens, siRNA studies, or knockout transcriptomes) and matched gene function annotations (e.g., Gene Ontology or pathway databases), we will systematically compare predicted perturbation signatures against real biological effects. An adversarial validation approach will be employed to detect overfitting: if a prediction model cannot distinguish its training data from held-out experimental data, its generalizations are likely spurious. The primary outcome is a cross-method accuracy estimate, quantifying how often in silico perturbation profiles correctly predict direction, magnitude, or functional category of real gene disruption effects. This purely computational, meta-analytic study requires no new experiments and aims to provide a clear benchmark for researchers choosing perturbation prediction tools.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0