Abstract
Scalable proxies for 3D genome contacts - such as single-cell co-accessibility and deep learning predictions - have emerged as powerful alternatives to chromatin capture-based methods, but predictions systematically overestimate long-range interactions. Here we show how to correct this bias using distance-based penalty functions informed by Gaussian mixture modeling and polymer-physics scaling. Using Hi-C datasets from maize, rice, and soybean, we derive tissue-specific and global consensus penalties parameterized by multi-regime power-law exponents. Applying these corrections to scATAC-seq co-accessibility scores improves their distance profiles in concordance with Hi-C and reduces long-range false positives by an average of 73% with tissue-specific penalties and 66% with the global consensus. We provide open-source code and fitted parameters to support adoption in maize, rice, and soybean.
Full text
1,364 characters
· extracted from
oa-doi-fallback
· click to expand
Abstract
Scalable proxies for 3D genome contacts - such as single-cell co-accessibility and deep learning predictions - have emerged as powerful alternatives to chromatin capture-based methods, but predictions systematically overestimate long-range interactions. Here we show how to correct this bias using distance-based penalty functions informed by Gaussian mixture modeling and polymer-physics scaling. Using Hi-C datasets from maize, rice, and soybean, we derive tissue-specific and global consensus penalties parameterized by multi-regime power-law exponents. Applying these corrections to scATAC-seq co-accessibility scores improves their distance profiles in concordance with Hi-C and reduces long-range false positives by an average of 73% with tissue-specific penalties and 66% with the global consensus. We provide open-source code and fitted parameters to support adoption in maize, rice, and soybean.
Competing Interest Statement
The authors have declared no competing interest.
Footnotes
This version has been revised to include a finalized multi-component power-law penalty framework and expanded validation across soybean and rice datasets. Figures 2 and 3 have been updated.. Data and code availability sections were added to include GitHub and Zenodo repositories. Supplemental files, including Table S3 and penalty parameters, have been updated.
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.