Interpretable machine learning classifiers implicate GPC6 in Parkinson's disease from single-nuclei midbrain transcriptomes

preprint OA: closed CC-BY-ND-4.0
📄 Open PDF Full text JSON View at publisher
AI-generated deep summary by claude@2026-06, 2026-06-24 · read from full text

The study developed an interpretable machine learning framework to distinguish post-mortem midbrain single-nucleus transcriptomes from Parkinson’s disease (PD) patients versus healthy controls, using classifiers whose decoded features identify gene expression signatures driving discrimination. Applying the approach to three publicly available snRNAseq datasets, the authors found cell-type-specific gene sets that consistently and accurately classified PD cells across all datasets. Targeted genomic analyses of the key genes implicated rare variants in GPC6, which the authors link to prior mechanisms involving intracellular accumulation of α-synuclein preformed fibrils, and they replicated the association in three independent case-control cohorts. The paper’s main caveat is that it uses openly available post-mortem and cross-cohort datasets, so model performance and variant associations are grounded in those specific data sources. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Parkinson's disease (PD) is a progressive and devastating neurodegenerative disease. An incomplete understanding of its genetic architecture remains a major barrier to the clinical translation of targeted therapeutics, necessitating novel approaches to uncover elusive genetic determinants. Single-cell and single-nuclear RNA sequencing (scnRNAseq) can help bridge this gap by profiling individual cells for disease-associated differential gene expression and nominating genes for targeted genomic analyses. Here, we introduce a machine learning framework to identify molecular features that characterize post-mortem brain cells from PD patients. We train classifiers to distinguish between PD and healthy cells, then decode the models to unravel the 'reasons' behind the classifications, revealing key genes expression signatures that characterize cells from the parkinsonian brain. Application of this framework to three publicly available snRNAseq datasets characterizing the post-mortem midbrain identified cell-type-specific gene sets that accurately classify PD cells across all datasets, demonstrating our approach's capacity to identify robust molecular markers of disease. Targeted genomic analyses of the key genes characterizing PD cells revealed a previously undescribed association between PD and rare variants in GPC6, a member of the heparan sulfate proteoglycan family, which have been implicated in the intracellular accumulation of alpha-synuclein preformed fibrils. We replicate this association in three separate case-control cohorts. Our method promises to enhance understanding of the genetic architecture in complex diseases like PD, representing a critical step toward targeted therapeutics. Our publicly available framework is readily applicable across diseases.
Full text 4,250 characters · extracted from oa-doi-fallback · click to expand
Abstract Parkinson’s disease (PD) is a progressive and devastating neurodegenerative disease. An incomplete understanding of its genetic architecture remains a major barrier to the clinical translation of targeted therapeutics, necessitating novel approaches to uncover elusive genetic determinants. Single-cell and single-nuclear RNA sequencing (scnRNAseq) can help bridge this gap by profiling individual cells for disease-associated differential gene expression and nominating genes for targeted genomic analyses. Here, we introduce a machine learning framework to identify molecular features that characterize post-mortem brain cells from PD patients. We train classifiers to distinguish between PD and healthy cells, then decode the models to unravel the ‘reasons’ behind the classifications, revealing key genes expression signatures that characterize cells from the parkinsonian brain. Application of this framework to three publicly available snRNAseq datasets characterizing the post-mortem midbrain identified cell-type-specific gene sets that accurately classify PD cells across all datasets, demonstrating our approach’s capacity to identify robust molecular markers of disease. Targeted genomic analyses of the key genes characterizing PD cells revealed a previously undescribed association between PD and rare variants in GPC6, a member of the heparan sulfate proteoglycan family, which have been implicated in the intracellular accumulation of α-synuclein preformed fibrils. We replicate this association in three separate case-control cohorts. Our method promises to enhance understanding of the genetic architecture in complex diseases like PD, representing a critical step toward targeted therapeutics. Our publicly available framework is readily applicable across diseases. Competing Interest Statement The authors have declared no competing interest. Funding Statement This work was supported by the Michael J. Fox Foundation [MJFF-021629 to EAF, SMKF, and RAT] and a Fonds d'Acceleration des Collaborations en Sante (FACS) grant from CQDM/MEI [to EAF]. Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The study only used openly available human data. SnRNAseq data prepared by Kamath et al. is available from the Single Cell Portal (https://singlecell.broadinstitute.org/single_cell/study/SCP1768). SnRNAseq data prepared by Wang et al. is available from the GEO with accession code GSE184950. SnRNAseq data prepared by Smajic et al. is available from the GEO with accession code GSE157783. ScRNAseq data chracterizing the iPSC-derrived DaNeurons prepared by Bressan et al. are available from the PPMI database (www.ppmi-info.org/access-dataspecimens/download-data), RRID:SCR 006431. The genomics data from the PD Genome Project and IPDGC Exome Sequencing Project are available from the Parkinson's Disease Variant Browser (https://pdgenetics.shinyapps.io/VariantBrowser/). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Footnotes ↵* Shared co-senior authorship michael.fiorini{at}mail.mcgill.ca; jialun.li{at}mail.mcgill.ca; ted.fon{at}mcgill.ca;

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-30T02:00:01.510937+00:00
License: CC-BY-ND-4.0