Clustering Gene Co-expression Using 2D Contour Analysis and Autoencoder-based Embedding

preprint OA: closed
📄 Open PDF Full text JSON View at publisher

Abstract

Identifying co-expressed genes is crucial for understanding biological processes; however, common methods such as Pearson and Spearman correlation rely on assumptions of linearity and monotonicity, respectively, that may not hold for complex biological data. To address these limitations, we propose a framework that captures visual features of gene expression profiles without relying solely on correlation-based methods. Our approach involves converting 3D gene expression data, which contain rich information on gene envelopes over time, into 2D contours that retain important visual information. We then train an autoencoder on the 2D contour images and generate embeddings from them, followed by clustering on the generated embeddings. We also introduce a new clustering algorithm, TAHC, based on hierarchical clustering, which performs better than existing methods in higher dimensions while using cosine similarity. We apply this framework to the Percellome database, which contains gene expression data with two variables—time and dosage—across various experimental conditions (e.g., tissue type). The resulting clusters exhibit good visual coherence, with an overall average Pearson correlation coefficient of 0.81, demonstrating the effectiveness of our approach.
Full text 1,450 characters · extracted from oa-doi-fallback · click to expand
Abstract Identifying co-expressed genes is crucial for understanding biological processes; however, common methods such as Pearson and Spearman correlation rely on assumptions of linearity and monotonicity, respectively, that may not hold for complex biological data. To address these limitations, we propose a framework that captures visual features of gene expression profiles without relying solely on correlation-based methods. Our approach involves converting 3D gene expression data, which contain rich information on gene envelopes over time, into 2D contours that retain important visual information. We then train an autoencoder on the 2D contour images and generate embeddings from them, followed by clustering on the generated embeddings. We also introduce a new clustering algorithm, TAHC, based on hierarchical clustering, which performs better than existing methods in higher dimensions while using cosine similarity. We apply this framework to the Percellome database, which contains gene expression data with two variables—time and dosage—across various experimental conditions (e.g., tissue type). The resulting clusters exhibit good visual coherence, with an overall average Pearson correlation coefficient of 0.81, demonstrating the effectiveness of our approach. Competing Interest Statement The authors have declared no competing interest. Footnotes nata{at}csl.sony.co.jp, ghosh{at}sbi.jp, kitano{at}sbi.jp, sucheendra{at}sbi.jp

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00