{"paper_id":"0ff03da5-94c4-477e-ad7a-a5ed9b044f36","body_text":"MANUSCRIPT TITLE \n \nDeep learning framework ChIANet predicts protein-mediated chromatin architecture across \nfunctional contexts \n  \nHanyu Luo1, Renjie Wen1, Li Tang1, 2, Lingyi Chen1, Kailing Tang1 and Min Li1, * \n1School of Computer Science and Engineering, Central South University, Changsha 410083, \nChina. \n2Department of Genetics, Yale School of Medicine, New Haven, CT 06510, USA. \n \n* To whom correspondence should be addressed. Tel: +86 13975134596;  \nEmail: limin@mail.csu.edu.cn \n \n  \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nABSTRACT  \nThe spatial organization of the genome is dynamically shaped by chromatin -binding proteins, yet \nhow protein-mediated three-dimensional (3D) architectures are specified across functional contexts \nremains incompletely understood. Here we present ChIANet, a m ultimodal deep learning \nframework that enables de novo prediction of protein-mediated chromatin contact maps and loops \nfrom protein-binding profiles alone, using the reference genome sequence as a prior. By integrating \ntransformer-based long-range modeling with multi-task learning, ChIANet accurately reconstructs \nprotein-mediated 3D chromatin architectures and generalizes across diverse cellular contexts. \nSystematic application of ChIANet to CTCF, Cohesin and RNAPII across seven human cell types \nreveals that chromatin architectures follow conserved organizational principles while exhibiting \npronounced context -dependent reconfiguration: CTCF - and Cohesin -mediated interactions \npredominantly support stable structural frameworks, whereas RNAPII -mediated loops di splay \ngreater variability and are closely coupled to transcriptional programs and regulatory element \nactivity. Functional analyses further uncover distinct regulatory biases and super -enhancer \nassociations among the three proteins. Extending this framework  to cancer genomes, ChIANet \ncaptures RNAPII-mediated chromatin looping networks associated with extrachromosomal DNA \n(ecDNA), revealing highly connected, transcription -associated architectures within amplified \necDNA regions across multiple cancer cell types. Together, these results demonstrate that protein-\nmediated 3D genome organization is not determined by protein identity alone but is flexibly shaped \nby functional context, regulatory targets and cellular environment, establishing ChIANet as a \nunified and scalable approach for decoding context-dependent principles of genome folding. \n \n  \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nINTRODUCTION \nIn eukaryotic cells, chromatin is folded into a highly dynamic three-dimensional (3D) architecture \nwhose organization is hierarchically arranged yet flexibly reconfigured across cellular and \nregulatory contexts, thereby governing gene expression, transcrip tional coordination and cellular \nidentity 1-3. This organization emerges from the collective action of DNA -binding and structural \nproteins that dynamically form, stabilize and remodel chromatin loops 4,5. Rather than constituting \na uniform structural scaffold, protein -mediated chromatin loops bring distant genomic loci into \nspatial proximity in a context -dependent manner, enabling enhancer -promoter communication, \ninsulating regulatory domains and supporting diverse gene regulatory programs. Such spatial \nregulation is fundamental to lineage specification, developmental control and genome stability 6,7. \nAmong key regulators of chromatin architecture, the CCCTC -binding factor (CTCF) functions \nas a sequence -specific architectural anchor that contributes to the establishment of topologically \nassociating domain (TAD) boundaries  8,9. The Cohesin complex, a member of the structural \nmaintenance of chromosomes (SMC) family, cooperates with CTCF to extrude chromatin loops and \nstabilize higher-order domains, while also participating in regulatory interactions that vary across \ngenomic and ce llular contexts  10,11. RNA polymerase II (RNAPII), in contrast, mediates \ntranscription-associated chromatin loops by recruiting Mediator and transcription factors, forming \na dynamic regulatory layer that is closely coupled to gene activity and regulatory element usage 12,13. \nThese protein -mediated interactions are further shaped by chromatin states, including histone \nmodifications such as H3K27ac and H3K27me3, which mark active and repressive regulatory \nregions and modulate the accessibility, stability and functional outcome of chromatin contacts 14,15. \nChromatin conformation capture (3C) technologies and their derivatives, including Hi -C, have \nenabled genome-wide mapping of chromatin contact frequencies and revealed general principles of \n3D genome organization 16-18. Targeted assays such as ChIA-PET 19, HiChIP 20, and PLAC-seq 21 \nfurther integrate chromatin immunoprecipitation with proximity ligation to profile interactions \nassociated with specific proteins or histone modifications. However, these experimental approaches \nremain cost-intensive and laborious, requiring large cell numbers, deep sequencing and high-quality \nantibodies for each target protein  2. Their limited scalability across proteins and cell types has \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nconstrained systematic investigation of how protein -mediated chromatin architectures vary across \nfunctional contexts, leaving a substantial gap in our understanding of context-dependent 3D genome \nregulation. \nTo overcome these limitations, computational modeling has emerged as a powerful approach for \npredicting chromatin interactions 22,23. Deep learning-based models such as Akita 24, Orca 25, and \nC.Origami 26 have achieved impressive accuracy in reconstructing Hi -C contact maps from DNA \nsequence or epigenomic signals 27,28. Other frameworks, including Peakachu 29 and DeepChIA-PET \n30, predict loops from experimental Hi -C or ChIA -PET data. In parallel, another class of models \nformulates loop detection as an anchor-pair classification problem, in which two genomic loci are \nprovided as input and the model predicts whether they form a loop interaction  31-33. Despite these \nadvances, most existing approaches remain limited to data -dependent or protein-agnostic settings, \nfocusing primarily on generic architectural features rather than protein -mediated regulatory \nmechanisms. As a result, they provide limited insight into how distinct chromatin-binding proteins \norganize three-dimensional genome architecture across different functional contexts or cell types. \nMore recently, several protein-specific predictors—for example, models targeting CTCF- or YY1-\nmediated loops—have been developed to capture motif -driven interactions 34,35. However, these \nmethods are narrowly tailored to individual proteins and specific regulatory scenarios, lacking \nextensibility to diverse chromatin-associated factors or broader cellular environments. Consequently, \na unified computational framework capable o f de novo prediction of diverse protein -mediated \nchromatin architectures in a context-aware manner across cell types remains lacking. \nTo address this gap, we developed ChIANet, a protein -specific multimodal deep learning \nframework that integrates genomic sequence with protein -specific ChIP -seq profiles to predict \nchromatin contact maps and loops mediated by distinct chromatin regulators across functional \ncontexts. ChIANet couples a Transformer -based encoder, which captures long -range genomic \ndependencies from sequence and binding signals, wi th a multi -task decoder that jointly models \ncontact intensity and loop probability, thereby learnin g unified representations of structural \norganization and regulatory activity. Methodologically, ChIANet combines heteroscedastic \nuncertainty-weighted multi -task learning with genome -wide tiling, enabling stable training, \nconsistent resolution across genomic scales and efficient application to entire chromosomes. Unlike \nprevious factor-specific or Hi -C-dependent models, ChIANet supports de novo reconstruction of \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nprotein-mediated three-dimensional genome architecture in unseen cell types using only ChIP -seq \nprofiles as experimental input, without requiring Hi -C supervision or model retraining. By \nsystematically applying ChIANet to three representative chromatin regulators—CTCF, Cohesin and \nRNAPII—across seven human cell types, we reveal that protein-mediated chromatin architectures \nfollow conserved organizational rules while being flexibly reconfigured across regulatory and \ncellular contexts. We further demonstrate that this framework extends to cancer genomes, capturing \nRNAPII-mediated chromatin looping architectures associated with extrachromosomal DNA \n(ecDNA), which represents an extreme regulatory context characterized by copy -number \namplification and dense regulatory activity. \nRESULTS \nChIANet enables accurate and generalizable prediction of protein -mediated chromatin \ninteractions \nWe developed ChIANet, a multimodal encoder-decoder framework that integrates the reference \ngenome sequence with protein -specific ChIP-seq profiles to predict protein -mediated chromatin \norganization at genome scale (Fig. 1a, b; Supplementary Fig. 1). Within each 2.1-Mb window, local \nsequence and binding features are processed by stacked Conv1D residual blocks and an eight-layer \nTransformer encoder to capture long -range dependencies  36, followed by Conv2D decoding to \njointly generate contact maps and loop matrices under a heteroscedastic uncertainty-weighted multi-\ntask loss 37(Methods). \nFor each chromatin regulator (CTCF, Cohesin and RNAPII), we trained a dedicated model using \nGM12878 ChIA -PET data with a chromosome -level split (19/2/2 for training, validation and \ntesting), ensuring strict separation of genomic contexts during evaluation  (Fig. 1b). Genome-wide \npredictions were obtained by tiling overlapping windows and stitching regional outputs into \ncontinuous chromosomal maps (Methods). Because ChIANet conditions only on protein -specific \nChIP-seq signals and the invariant genome sequence, trained models can be directly applied to \nunseen cell types without retraining. Using this strategy, we generated de novo protein -mediated \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\ncontact maps and loop predictions across seven human cell types, establishing a unified framework \nfor cross-context comparative analyses (Fig. 1c). \nWe next systematically evaluated ChIANet ’s performance in reconstructing protein -mediated \ncontact maps. As no existing method explicitly models protein-specific interactions, we compared \nChIANet with three state -of-the-art Hi -C prediction models —Akita 24, Orca 25, and C.Origami \n26(Methods)—under identical preprocessing and evaluation settings (Methods) . Across all test-set \ngenomic windows, ChIANet consistently achieved higher Pearson and Spearman correlations and \nlower mean squared error compared with competing methods.  (Fig. 2a, Extended Data Fig. 1a). \nImportantly, ChIANet maintained stronger correlations at increasing genomic distances, indicating \nimproved modeling of long -range protein -mediated contacts  (Fig. 2 b, Extended Data Fig. 1b). \nChromosome-scale scatter analyses further confirmed the high concordance between predicted and \nexperimental contact intensities across all proteins (Supplementary Fig. 3). Representative regions \nfurther illustrate the accurate reconstruction of both local and distal interaction patterns observed in \nChIA-PET data (Fig. 2c, d, h; Extended Data Fig. 3). \nWe further assessed loop prediction performance by comparing ChIANet with  Peakachu 29 and \nDeepChIA-PET 30, the two genome-wide frameworks capable of direct loop inference (Methods). \nChIANet consistently achieved the highest AUC and AUPR at both window and chromosome scales \n(Fig. 2 e, f; Extended Data Fig. 2a, b) , and maintained robust performance under increasingly \nimbalanced classification settings (Extended Data Fig. 1d). Notably, whereas both baseline methods \nrequire Hi-C contact maps as input, ChIANet relies solely on sequence and protein -specific ChIP-\nseq profiles, demonstrating that accurate loop prediction can be achieved without Hi-C supervision. \nThe model also achieved the highest distance-stratified AUPR, reflecting its capacity to capture both \nproximal and distal looping patterns (Fig. 2g; Extended Data Fig. 1c, Extended Data Fig. 4). \nTogether, these results establish ChIANet as an accurate and generalizable framework for \npredicting protein-mediated chromatin interactions, providing a robust foundation for downstream \nanalyses of protein-specific regulatory architectures across functional contexts. \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nMultimodal integration improves protein-specific interaction prediction \nWhile CTCF binds to well-defined DNA motifs, other chromatin regulators such as Cohesin and \nRNAPII display weaker sequence specificity and encode their chromatin interactions more strongly \nthrough context-dependent chromatin features 8. To dissect the relative contributions of sequence \nand ChIP -seq features in protein -mediated interaction prediction, we conducted ablation \nexperiments by training two ChIANet variants with identical architectures but different inputs: a \nsequence-only model and a ChIP -seq-only model. Across all three proteins in GM12878, the full \nsequence + ChIP -seq model (original ChIANet) achieved the lowest convergence loss for both \ncontact map and loop prediction tasks (Extended Data Fig. 5a). Notably, the sequence-only model \nfor CTCF converged more stably, whereas those for Cohesin and RNAPII showed pronounced \nfluctuations during training, reflecting the limited ability of sequence information alone to capture \ninteraction patterns of proteins that are less motif-driven and more context-dependent. \nFor the contact map prediction task, the full model consistently achieved the highest Pearson \ncorrelation coefficients (PCC) and the lowest mean squared errors (MSE) across all proteins (Fig. \n2i; Extended Data Fig. 5b). In distance -stratified correlation analyses, ChIANet also showed the \nstrongest performance (Extended Data Fig. 5c). In contrast, the sequence -only model rapidly lost \ncorrelation beyond 50 bins, whereas the ChIP -seq-only model maintained moderate performance, \nhighlighting the necessity of ChI P-seq information for modeling protein -dependent long -range \ninteractions. For the loop prediction task, a similar trend was observed (Extended Data Fig. 5d). \nEven under increasingly imbalanced classification ratios, the full model maintained robust AUPR \nvalues across all proteins (Fig. 2j; Supplementary Fig. 4). At the chromosome scale (1:100 negative-\nto-positive ratio), ChIANet reached an AUPR of 0.683 for CTCF —representing improvements of \n+0.536 and +0.196 over the sequence-only and ChIP-seq-only models, respectively—and 0.364 for \nRNAPII (+0.226 and +0.068), and 0.670 for Cohesin (+0.523 and +0.140) (Extended Data Fig. 5e). \nTo visualize how multimodal integration enhances prediction fidelity, we examined \nrepresentative genomic regions. As shown in Supplementary Fig. 5 , the full ChIANet model \naccurately reconstructed clear protein-specific topological architectures and high-confidence loops \nconsistent with ChIA-PET data. In contrast, the ChIP-seq-only model produced blurred topological \ndomains with weakened distal conta cts, while the sequence -only model failed to recover any \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\ndiscernible structure or loop signal. Together, these analyses demonstrate that integrating nucleotide \nsequence with protein-specific ChIP-seq profiles enables ChIANet to capture distinct informational \nregimes underlying protein-mediated chromatin architectures, spanning motif-driven and context-\ndependent regulatory interactions. \nDe novo prediction of protein-mediated chromatin organization across cellular contexts \nHaving established the modality advantage of ChIANet, we next evaluated its ability to generalize \nacross cell types. We trained a CTCF-specific model exclusively on GM12878 cells and performed \nde novo predictions of CTCF-mediated chromatin architecture in H1 cells without any retraining. \nRemarkably, ChIANet faithfully reconstructed both local and long -range contact patterns that \nclosely mirrored the experimental ChIA-PET maps in H1 (Fig. 3a; Extended Data Figs. 3, 4). The \npredicted contact maps and loop structures captured pronounced chromatin rearrangements between \nGM12878 and H1, consistent with differential CTCF occupancy along the genome.  \nFor the contact map prediction task, ChIANet achieved comparable or higher correlations than \nsequence-based and multimodal Hi-C prediction models across all genomic windows, with median \nPearson and Spearman correlations exceeding 0.8 (Fig. 3b; Extended Data Fig. 6). For the loop \nprediction task, ChIANet also exhibited strong transferability, yielding the highest AUC and AUPR \nacross a wide range of positive -to-negative ratios (1:1 to 1:100) compared with Peakachu and \nDeepChIA-PET (Fig. 3c; Extended Data Fig. 7; Supplementary Fig. 6). Notably, while both baseline \nmodels require Hi-C contact maps as input features, ChIANet relies solely on DNA sequence and \nprotein-specific ChIP-seq tracks, underscoring its capacity to generalize to unseen cellular contexts \nwithout Hi-C guidance. \nTo further quantify global de novo  loop prediction accuracy, we performed Aggregate Peak \nAnalysis (APA) on the top -scoring predicted loops 38 (Fig. 3d, e; Methods). As the number of \npredicted loops increased from 1k to 50k, the enrichment metrics—P2LL and ZscoreLL—initially \nrose and peaked around 10,000 loops, indicating an optimal balance between precision and coverage \n(Fig. 3d). APA heatmaps showed strong focal enrichment centered at predicted loop anchors across \nall thresholds, consistent with interaction patterns observed in experimental Hi-C maps (Fig. 3e). \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nOverall, these results demonstrate that ChIANet can generalize across cell types to accurately \nreconstruct both contact maps and loop -level features of protein-mediated chromatin interactions. \nBy integrating DNA sequence and protein -specific ChIP -seq information, ChIANet captures \ntransferable organizational rules of protein-mediated chromatin architecture, enabling high-fidelity \nde novo prediction across distinct cellular contexts. \nProtein-specific yet coordinated chromatin architectures are conserved across cell types \nTo examine how distinct chromatin regulators jointly organize three -dimensional genome \narchitecture across cellular contexts, we applied ChIANet to predict protein-mediated contact maps \nand loops for CTCF, Cohesin and RNAPII across seven representative human cell types (GM12878, \nH1, A549, HeLa-S3, IMR90, HepG2 and K562). \nUsing H1 as an illustrative example, the predicted contact maps revealed distinct yet coordinated \norganizational patterns among the three proteins (Fig. 4a). CTCF - and Cohesin-mediated maps \nexhibited highly concordant domain -level structures, whereas RNAPI I-mediated interactions \ndisplayed more spatially diffuse patterns preferentially enriched at transcriptionally active regions. \nGenome-wide pairwise correlations confirmed this relationship, with consistently higher \nconcordance between CTCF and Cohesin compared with protein pairs involving RNAPII (Fig. 4b; \nExtended Data Fig. 8a). Importantly, global UMAP embedding of all 2.1-Mb windows demonstrated \nthat while CTCF and Cohesin occupied overlapping manifolds, RNAPII maps formed a partially \nseparated yet adjacent structural space, indicating that transcription -associated interactions are \norganized in close coordination with, rather than independently from, the architectural backbone \n(Fig. 4c; Extended Data Fig. 8c). These relationships were preserved across indi vidual \nchromosomes (Fig. 4d; Extended Data Fig. 8b) and across genomic distance scales, with stronger \nlong-range concordance for CTCF-Cohesin pairs (Fig. 4e; Extended Data Fig. 9a). Together, these \nresults indicate that architectural proteins (CTCF/Cohesin) cooperatively maintain domain -level \nstructure, whereas RNAPII contributes a transcription-associated layer. \nLoop-level analyses further highlighted both shared and protein-specific organizational principles. \nRNAPII-mediated loops were consistently shorter in genomic span, whereas CTCF- and Cohesin-\nmediated loops extended over broader distances (Fig. 4f; Extended  Data Fig. 9b). Substantial \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\noverlap was observed between CTCF and Cohesin loop sets, whereas RNAPII loops showed limited \nintersection, consistent with their preferential engagement in localized regulatory interactions (Fig. \n4g; Extended Data Fig. 9c).  Despite these differences, network topology analyses revealed \nconvergent global properties across all three proteins, including scale -free degree distributions \nindicative of hub-centered organization (Fig. 4h; Extended Data Fig. 10a). Across cell types, CTCF \nand Cohesin networks were enriched for high -degree hubs that contributed disproportionately to \nlong-range connectivity, whereas RNAPII networks were characterized by smaller k-core structures \nand locally clustered interaction communities (Fig. 4i; Extended Data Fig. 10b-d), reflecting a more \ndynamic and context-responsive mode of chromatin organization. \nCollectively, these analyses reveal a conserved yet flexible organizational framework in which \narchitectural proteins and transcription-associated regulators operate in a coordinated manner across \ndiverse cellular contexts. Rather than forming isolated str uctural layers, CTCF -, Cohesin - and \nRNAPII-mediated interactions collectively shape chromatin architecture through context-dependent \nreweighting of shared organizational principles, providing a mechanistic basis for both structural \nstability and regulatory plasticity across cell types. \nCross-cell-type conservation and variability of protein-mediated chromatin architecture \nBuilding on the coordinated yet protein -specific organizational patterns observed within \nindividual cell types, we next examined how protein-mediated chromatin architectures are preserved \nor reconfigured across distinct cellular contexts. Genome -wide corre lations between predicted \ncontact maps revealed that CTCF- and Cohesin-mediated interactions display substantially higher \ncross-cell-type concordance than RNAPII -mediated interactions (Fig. 5a), indicating that \narchitectural features associated with these proteins are more consistently maintained across cellular \nstates, whereas transcription -associated contacts exhibit pronounced context dependence. To \nquantify this behavior at higher resolution, we computed a conservation score for each genomic \nwindow across the seven cell types  (Methods). Chromosome-resolved analyses confirmed a clear \nstratification, with CTCF and Cohesin exhibiting uniformly high conservation and RNAPII showing \nmarkedly greater dispersion (Supplementary Fig. 9a). Visualization of these co nservation \nlandscapes along individual chromosomes further revealed extended domains of stable CTCF- and \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nCohesin-mediated organization interspersed with regions in which RNAPII -mediated interactions \nvaried substantially between cell types (Supplementary Fig. 9b). Representative loci were selected \nto illustrate conserved versus context-variable architectural patterns (Fig. 5b). \nWe next extended this analysis to the loop level. For each protein, a pan-cell set of predicted loops \nwas assembled, and loop conservation across cell types was quantified using a static score. In total, \n34,751 CTCF, 48,847 Cohesin and 46,093 RNAPII loops were identified. Based on their static \nscores, loops were stratified into stable and variable categories (Supplementary Fig. 10). The \nresulting distributions revealed a hierarchical pattern of conservation, with CTCF loops exhibiting \nthe greatest stability , followed by Cohesin and RNAPII (Fig. 5c), consistent with their distinct \norganizational roles. At the network level, loop stability was strongly associated with anchor \nconnectivity. Loops anchored at high-degree nodes were significantly more stable across cell types, \nparticularly for CTCF and Cohesin (Fig. 5d), suggesting that highly connected hubs serve as \nconserved organizational scaffolds. In contrast, lower -degree RNAPII loops exhibited greater \nvariability, reflecting their preferential engagement in cell-type-specific regulatory interactions. To \nlink loop conservation with transcriptional programs, genes were classified as housekeeping or cell-\ntype-specific based on RNA-seq expression profiles. Variable loops were significantly enriched in \nthe vicinit y of cell -type-specific genes, whereas stable loops preferentially associated with \nhousekeeping genes (Fig. 5e), indicating that context -dependent rewiring of chromatin loops \nunderlies transcriptional diversification across cell types. \nFinally, loop conservation was assessed using a probabilistic support metric that quantifies the \nnumber of cell types in which each loop is detected. CTCF and Cohesin loops were supported by \nsubstantially more cell types than RNAPII loops (Fig. 5f), with more than half of CTCF and Cohesin \nloops retained in at least four cell types, whereas the majority of RNAPII loops were restricted to \none or two. Together, these analyses demonstrate that protein -mediated chromatin architecture is \ngoverned by a balance between conserved structural scaffolding and context-dependent regulatory \nreconfiguration, with different proteins contributing distinctially to architectural stability and \ntranscriptional plasticity across cellular contexts. \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nProtein-mediated chromatin loops exhibit context-dependent functional biases and regulatory \npotential \nWe next examined how protein -mediated chromatin loops are functionally deployed across \ngenomic contexts and how their regulatory potential varies with protein identity and interaction \ntopology. At the genome -wide scale, CTCF - and Cohesin -mediated loops wer e preferentially \nassociated with chromatin domain boundaries, whereas RNAPII -mediated loops were enriched at \nregulatory elements, including promoters and enhancers (Fig. 6a; Extended Data Fig. 11a). \nImportantly, the degree of functional enrichment scaled with anchor connectivity: highly connected \nCTCF and Cohesin anchors showed progressively stronger association with domain boundaries, \nwhereas increasing RNAPII connectivity was accompanied by a pronounced shift toward promoter-\nlinked interactions, indicatin g context -dependent reweighting of loop function as network \ncomplexity increases. Classification of pan -cell-type loops further revealed distinct functional \ninteraction classes associated with different proteins. CTCF -mediated loops were enriched for \nstructural configurations, including a substantial fraction of silencer-promoter interactions, whereas \nRNAPII-mediated loops were dominated by promoter-promoter and enhancer-promoter interactions \n(Fig. 6b). These functional compositions were consistently observed across all additional cell types \nanalyzed (Extended Data Fig. 11b), indicating that protein -specific interaction biases represent \nconserved organizational tendencies rather than cell-type-specific artifacts. \nTo further contextualize these differences, we integrated chromatin -state annotations for cell -\ntype-specific loops. CTCF -specific anchors were most frequently localized within quiescent and \nrepressed chromatin states, consistent with an insulative or structural role, whereas RNAPII-specific \nanchors were predominantly associated with active transcription start sites and enhancer states (Fig. \n6c; Extended Data Fig. 12a). Quantitative enrichment analyses confirmed this dichotomy, with \nCTCF- and Cohesin-mediated loops showing strong enrichment at repressive chromatin features \nand RNAPII -mediated loops preferentially linked to transcriptionally active regions (Fig. 6d; \nExtended Data Fig. 12b). \nWe next connected loop architecture with transcriptional output by examining expression levels \nof genes associated with cell-type-specific loops. RNAPII- and Cohesin-mediated loops were linked \nto significantly higher gene expression compared with non -regulatory contacts, whereas CTCF -\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nmediated loops showed only modest effects (Fig. 6e; Extended Data Fig. 11c). Among promoter -\nenhancer loops, loop strength positively correlated with differential gene expression across cell \ntypes, with stronger associations observed for RNAPII and Cohesin than for CTCF (Extended Data \nFig. 13a). These results indicate that while CTCF -mediated loops primarily constrain chromatin \ntopology, Cohesin and RNAPII -mediated loops more directly modulate transcriptional \nresponsiveness in a context-dependent manner. \nWe further explored the relationship between protein-mediated loops and super-enhancers (SEs) \n39. Cell-type-specific loops mediated by Cohesin and RNAPII were significantly enriched in the \nvicinity of cell-type-specific SEs, whereas CTCF-mediated loops showed weaker and less consistent \nenrichment across cell types (Fig. 6f; Extended Data Fig. 13b). Thi s pattern suggests that \narchitectural scaffolding provided by CTCF is selectively coupled to transcriptional amplification \nthrough Cohesin- and RNAPII-mediated interactions in regulatory contexts associated with high \nenhancer activity. At the gene level, RNAPII-mediated loops targeted the largest and most distinct \nset of genes across cell types, whereas CTCF and Cohesin shared a substantial fraction of their \ntargets, consistent with their coordinated roles in organizing chromatin structure (Fig. 6g; Extended \nData Fig. 14a). Functional enrichment analyses of SE -associated loop genes further revealed \ncomplementary regulatory programs among the three proteins (Fig. 6h). In GM12878, CTCF -\nassociated SE loops were enriched for chromatin organization and architec tural maintenance, \nCohesin-associated loops for replication and cohesion -related pathways, and RNAPII -associated \nloops for immune activation and transcriptional execution. Similar trends were observed across \nadditional cell types (Extended Data Fig. 14b). \nCollectively, these analyses demonstrate that protein -mediated chromatin loops are not \nintrinsically defined by protein identity alone, but are functionally deployed in a context-dependent \nmanner. CTCF, Cohesin and RNAPII contribute distinct yet coordinate d looping programs that \ncollectively balance structural constraint and regulatory flexibility, providing a mechanistic link \nbetween three-dimensional genome organization and transcriptional control across cellular contexts. \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nContext-dependent reorganization of RNAPII-mediated chromatin architecture on ecDNA in \ncancer \nExtrachromosomal DNA (ecDNA) represents an extreme regulatory context in cancer genomes, \ncharacterized by focal copy-number amplification and dense clustering of regulatory elements 40-42. \nTo examine how protein -mediated chromatin architecture is reorganized under this context, we \nanalyzed RNAPII-mediated chromatin interactions in three cancer cell types with available ecDNA \ncontact maps generated by ChIA -Drop 43. Circular visualization of representative ecDNA regions \nrevealed pronounced copy -number amplification accompanied by dense occupancy of RNAPII, \nMED1 and H3K27ac, indicative of highly active transcriptional hubs (Fig. 7a). Consistently, ChIP-\nseq signals of RNAPII, MED1 and H3K27ac were strongly enriched around ecDNA -associated \ncontact anchors in PC3DM, with similar enrichment patterns observed across biological replicates \nin COLO320DM and GBM171DM (Fig. 7b; Supplementary Fig. 15a). Genome browser views \nfurther confirmed concordant spatial localization of ChIA -Drop in teraction signals and \ntranscription-associated epigenomic marks within ecDNA regions across all three cancer cell types \n(Supplementary Fig. 15b). \nMotivated by the close coupling between ecDNA contacts and transcription -associated \nepigenomic features, we applied ChIANet to predict RNAPII-mediated chromatin loops in the three \ncancer cell types. Predicted loops formed dense interaction networks within high-copy-number \necDNA regions, consistent with previously reported ecDNA -associated interaction structures \ndetected (Fig. 7c, d).  Quantitative enrichment analysis showed that ChIANet -predicted RNAPII \nloop anchors were significantly overrepresented at ecDNA contact regions compared with distance-\nmatched random RNAPII peak sets across all three cancer cell types (Fig. 7e), indicating preferential \ndeployment of RNAPII -mediated loops within ecDNA topology. Stratification of ChIA -Drop-\ndefined n -way Genome Extrusion Modules (nGEMs) by interaction strength further revealed a \npositive association between nGEM connectivity and MED1 and H3K27ac signal intensities in \nPC3DM and COLO320DM, whereas weaker trends were observed in GBM171DM (Supplementary \nFig. 12). \nDespite sharing a common RNAPII -centered architectural framework, ecDNA -associated \nlooping networks exhibited pronounced cancer-type-specific regulatory signatures. Motif analysis \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nof ChIANet -predicted RNAPII loop anchors revealed distinct transcription factor landscapes \ncoupled to expression across cancer cell types (Fig. 7f; Supplementary Fig. 15c). PC3DM showed \nenrichment of STA T1::STAT2, SP1/SP3 and KLF-family motifs, together with metabolic regulators \nsuch as SREBF2 and MLXIP . In COLO320DM, enriched motifs included TCF4 and additional GC-\nrich or architectural factors (for example MAZ, HMGA1 and SP1/SP3), consistent with a dominant \nWnt/ β -catenin-associated regulatory program. In G BM171DM, enriched motifs included \nimmediate-early and regulatory transcription factors such as EGR1 and ZEB1/ZEB2, together with \nSP1/SP3 and KLF -related motifs. Consistent with these motif -level distinctions, Gene Ontology \nenrichment analysis of genes associated with predicted RNAPII loop anchors identified both shared \nand cancer-type-specific functional programs (Extended Data Fig. 16). PC3DM was enriched for \ntranslation-related processes and chromatin organization, COLO320DM for chromatin and DNA \norganization and β-catenin-TCF complex assembly, whereas GBM171DM showed enrichment of \npost-transcriptional regulation, RNA catabolic processes and immune-related signaling.  \nTogether, these results demonstrate that ecDNA constitutes an extreme functional context in \nwhich RNAPII -mediated chromatin architecture is extensively reorganized, forming dense and \nhighly connected looping networks that couple copy -number amplification w ith transcriptional \nregulation in a cancer-type-specific manner. \nDISCUSSION \nUnderstanding how specific chromatin -binding proteins shape the three -dimensional genome \nremains a central challenge in regulatory genomics  44,45. Existing Hi-C-based approaches cannot \ndisentangle protein-specific contributions to genome folding, whereas sequence -only models lack \nthe regulatory context required to capture condition -dependent chromatin organization.  33. To \naddress these limitations, we developed ChIANet, a multimodal deep learning framework that \nintegrates DNA sequence and protein-specific ChIP-seq data to enable de novo prediction of protein-\nmediated chromatin interactions. By jointly modeling contact maps and chromatin loops within a \nunified encoder -decoder architecture, ChIANet accurately reconstructs genome -wide chromatin \norganization mediated by CTCF, Cohesin and RNAPII, capturing both local and long -range \narchitectural features and generalizing robustly across cell types without retraining. \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nSystematic application of ChIANet across seven human cell lines reveals a conserved hierarchical \norganization of the 3D genome in which distinct chromatin regulators play coordinated but \nnonredundant roles. Architectural proteins CTCF and Cohesin establish a stable structural backbone \nthat preserves domain-scale organization across cellular contexts, whereas RNAPII mediates a more \ndynamic, transcription-associated interaction layer. Loop-level analyses further support this division \nof labor: CTCF - and Cohesin-mediated loops are longer, more conserved and highly connected, \nwhile RNAPII -mediated loops are shorter, more variable and preferentially associated with \npromoters and enhancers. Integration with chromatin-state annotations, gene expression and super-\nenhancer landscapes demonstrates that Cohesin and RNAPII loops are tightly coupled to \ntranscriptional output, whereas CTCF primarily contributes to structural insulation. Together, these \nresults support a multilayered model of genome organization in which ar chitectural stability and \nregulatory flexibility are jointly encoded by different protein-mediated interaction networks. \nImportantly, extending ChIANet to cancer genomes carrying extrachromosomal DNA (ecDNA) \nreveals that this hierarchical organization is not fixed, but can be profoundly reorganized under \nextreme functional contexts. EcDNA represents a regulatory environment characterized by circular \ntopology, high copy number and dense clustering of regulato ry elements, in which canonical \nchromosomal constraints on genome folding are relaxed. In this context, RNAPII -mediated \ninteractions no longer constitute a secondary regulatory layer, but instead dominate chromatin \narchitecture, forming dense, highly inter connected looping networks that couple transcriptional \ncoactivator recruitment with amplified gene expression. Quantitative enrichment, nGEM \nstratification and epigenomic coupling analyses indicate that RNAPII assumes an architectural role \non ecDNA that is qualitatively distinct from its function on linear chromosomes. Moreover, despite \nsharing a common RNAPII -centered framework, ecDNA-associated interaction networks display \npronounced cancer-type-specific transcription factor and functional signatures, und erscoring how \nfunctional context and regulatory demand jointly reshape protein-mediated 3D genome organization. \nNotably, the conceptual design of ChIANet is not restricted to the specific chromatin regulators \nexamined here. By conditioning chromatin interaction prediction on protein -specific binding \nprofiles together with underlying genomic sequence, ChIANet provides a generalizable framework \nthat is, in principle, extensible to other chromatin-associated factors, including transcription factors, \ncofactors and chromatin modifiers, provided that appropriate binding and genomic context \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\ninformation are available. This flexibility positions ChIANet as a broadly applicable platform for \ninterrogating how diverse regulatory proteins contribute to three-dimensional genome organization \nacross distinct cellular states and biological contexts. \nTogether, these findings advance a context -dependent view of chromatin architecture in which \nprotein identity alone does not uniquely determine structural function. Instead, the architectural role \nof a given protein emerges from the interplay between its b iochemical activity, regulatory targets \nand the genomic environment in which it operates. By capturing both conserved organizational \nprinciples and context-specific architectural rewiring, ChIANet provides a scalable framework for \ndecoding how protein-mediated chromatin interactions adapt across cell types, regulatory states and \ndisease-associated genome configurations. Beyond descriptive modeling, the protein -conditioned \nnature of ChIANet naturally suggests future opportunities for in silico interrogation of genome \nfolding. In principle, systematic perturbation of protein binding landscapes, regulatory states or \ngenome configurations could enable predictive exploration of how changes in regulatory inputs \nreshape three-dimensional genome organization. While such applications remain to be explored, \nthese perspectives highlight the potential of ChIANet not only as a predictive model, but also as a \nhypothesis-generating framework for dissecting the functional logic that shapes the three -\ndimensional genome 46,47. \nMETHODS \nChIA-PET data processing \nWe used four ChIA -PET datasets in this study: CTCF (GM12878), CTCF (H1), Cohesin \n(GM12878) and RNAPII (GM12878) (Supplementary Table 1). All datasets are publicly available \nfrom ENCODE 48 (http://www.encodeproject.org/) and 4DN 49 (https://data.4dnucleome.org/). For \nCohesin-mediated chromatin interactions, we used the ChIA-PET dataset generated with antibodies \nagainst RAD21 and SMC1A, and processed them jointly as cohesin ChIA-PET to capture cohesin-\nassociated loops. Raw reads were  processed using the ChIA -PIPE 50 pipeline  \n(https://github.com/TheJacksonLaboratory/ChIA-PIPE) to generate contact maps and loop calls \naligned to the GRCh38 reference genome. For contact maps, interactions were called using raw \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\ncounts at 10 -kb resolution, and the matrices were log2 transformed with a pseudocount of 1 to \nstabilize variance.  \nLoop quality control and filtering \nLoop calls generated by ChIA-PIPE were further filtered to ensure high-confidence interactions. \nWe first removed inter -chromosomal loops to retain only intra -chromosomal interactions. Loops \nwith a supporting PET count (Score) < 5 were discarded to exclude weakly supported contacts. We \nalso excluded short -range loops (< 50 kb) to remove potential self -ligation or local background \nartifacts. The remaining high -confidence intra-chromosomal loops were used in all downstream \nanalyses. The number of loops before and after filtering for each dataset is summarized in \nSupplementary Table 2, and an example of quality control before and after filtering is shown in \nSupplementary Fig. 2. \nChIP-seq data processing \nWe collected 21 human ChIP -seq datasets from ENCODE  48 (http://www.encodeproject.org/) \ncorresponding to three proteins (CTCF, RAD21, RNAPII) across seven cell types. For each protein-\ncell-type combination, we downloaded all available isogenic replicate BAM files aligned to hg38 \nand merged replicates to obtain a unified signal track. To ensure signal comparability across cell \ntypes, we used deepTools 51 (bamCoverage) to generate genome -wide coverage files in bigWig \nformat with binSize  = 1 bp and RPGC (reads per genomic content) normalization. The resulting \ntracks were further log2 transformed with a pseudocount of 1 and used as input feature signals for \nChIANet. \nDNA sequence \nThe human reference genome (GRCh38/hg38) was downloaded from the UCSC Genome Browser \n52. We used the primary assembly and kept all nucleotide symbols present in the FASTA file. In \naddition to the canonical bases (A, C, G and T), positions annotated as unknown ( “N”) were \npreserved and treated as a separate category to avoid information loss in low -mappability or \nunresolved regions. For model input, the genome was converted to a five -channel one -hot \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nrepresentation corresponding to A, C, G, T and N, and the same reference sequence was used for all \ncell types and proteins to ensure a shared coordinate system. \nTraining data construction \nChIANet was trained on GM12878 for three proteins—CTCF, Cohesin and RNAPII—using one \nmodel per protein. For each model, inputs consisted of the DNA sequence and the corresponding \nprotein-specific ChIP-seq signal from a 2,097,152 -bp (2.1-Mb) genomic window. The prediction \ntargets were the protein -mediated contact map and loop matr ix for the same window.ChIA -PET \ncontact maps were initially generated at 10 -kb resolution and then downsampled to 8,192 -bp \n(256*256) resolution by bilinear interpolation to match th e model output grid. Loop targets were \nrepresented as a 256*256 binary matrix initialized to zeros, with bins overlapping ChIA-PET loop \nanchors set to 1. For both map and loop outputs, only the upper -triangular entries were used as \ntraining labels to avoid redundancy. \nTo create training samples, we tiled the genome with a sliding window of 2.1 Mb and a stride of \n50 kb. Windows overlapping centromeric or telomeric regions were excluded. Chromosomes were \nrandomly assigned to training, validation and test sets at the chromosome level: chromosomes 3 and \n15 were held out for validation, chromosomes 5 and 18 were held out for testing, and all remaining \nautosomes were used for training. This procedure yielded 4,672 training, 571 validation and 511 \ntest samples. \nModel architecture \nChIANet was implemented in PyTorch 53 and consists of three main components: a 1D \nconvolutional encoder, a Transformer module, and a multi -task convolutional decoder. DNA \nsequence and protein-specific ChIP-seq signals were concatenated along the channel dimension and \njointly fed into the encoder. \nThe encoder begins with a 1D convolution layer (kernel size = 11, stride = 2) to capture local \nmotif-like features. To reduce the input length from 2.1 Mb to 256 bins, the encoder applies 12 \nconvolutional modules, each composed of a residual block followed by a scaling block. The residual \nblock contains two 1D convolutions (kernel = 5, padding = 2), each followed by batch normalization \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nand ReLU activation, with skip connections to facilitate gradient flow and long -range feature \npropagation. The scaling block performs downsampling via a stride -2 convolution (kernel = 5) \nfollowed by batch normalization and ReLU. Hidden dimensions were progressively increased (32, \n32, 32, 32, 64, 64, 128, 128, 128, 128, 256, 256), resulting in a final encoder output of 256 bins * \n256 channels. \nThe Transformer module contains eight self-attention layers with eight attention heads and 256-\ndimensional hidden units. Each layer includes multi -head attention, feed -forward layers, layer \nnormalization and dropout (rate = 0.1). Sinusoidal positional enco ding was applied to preserve \ngenomic order. This module captures long-range dependencies between bins, enabling the model to \nlearn spatial interaction patterns from sequence and ChIP-seq features 36. \nThe decoder consists of five dilated residual convolutional blocks, with dilation rates set to 2, 4, \n8, 16 and 32 to ensure each output pixel has a receptive field covering the entire 2.1 Mb input. Each \nblock includes two 2D convolutions (kernel = 3), batc h normalization and ReLU activation with \nresidual connections. Two final 1*1 convolutions produce the outputs: one for the contact map and \none for the loop matrix. Both outputs are flattened to the upper -triangular part (length = 32,640). \nThe contact-map branch is trained using mean squared error (MSE) loss, and the loop branch using \nbinary cross-entropy (BCE) loss. A heteroscedastic uncertainty weighting strategy balances the two \nobjectives, and gradients are propagated end-to-end for joint optimization 37. \nModel training and prediction \nWe trained three protein -specific ChIANet models —CTCF, Cohesin and RNAPII —using \nGM12878 data. Each model jointly optimized two tasks: contact map prediction and loop prediction. \nTo balance the contributions of the two losses, we adopted a heteroscedastic uncertainty weighting \nstrategy 37: \n𝐿𝑡𝑜𝑡𝑎𝑙 = 1\n2𝜎𝑚𝑎𝑝2 𝐿𝑚𝑎𝑝 + 1\n2𝜎𝑙𝑜𝑜𝑝\n2 𝐿𝑙𝑜𝑜𝑝 + 𝑙𝑜𝑔𝜎𝑚𝑎𝑝 + 𝑙𝑜𝑔𝜎𝑙𝑜𝑜𝑝 \nwhere 𝐿𝑚𝑎𝑝  and 𝐿𝑙𝑜𝑜𝑝  denote the mean squared error (MSE) and binary cross -entropy (BCE) \nlosses, respectively. The uncertainty parameters 𝜎𝑚𝑎𝑝 and 𝜎𝑙𝑜𝑜𝑝 were jointly learned with model \nweights to adaptively reweight the two tasks during training. \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nTraining was performed using the Adam 54 optimizer (initial learning rate 2e-4, batch size = 16) \nfor 100 epochs. We employed a linear warm-up from 1e-5 followed by a cosine annealing schedule \nwith a minimum learning rate of 1e-6. An early stopping strategy was used to prevent overfitting: \ntraining was terminated if the total loss did not decrease for 10 consecutive epochs.  To enhance \ngeneralization, two forms of data augmentation were applied: (1) random shifts of the 2.1 -Mb \ngenomic window within ±0.36 Mb, and (2) reverse-complement flipping of the DNA sequence and \ncorresponding contact/loop matrices with a probability of 0.5. Each model was trained on a single \nNVIDIA A6000 GPU (~11 hours per model). \nFor in-cell predictions (GM12878), model inputs included the 2.1-Mb reference sequence and the \nprotein-specific ChIP-seq profile. For cross -cell de novo predictions, we replaced the GM12878 \nChIP-seq track with the corresponding track from the target cell type while keeping the reference \ngenome sequence unchanged. \nChromosome-scale prediction assembly \nTo generate chromosome-scale predictions, we tiled the genome using a sliding window of 2.1 \nMb with a stride of 262,144 bp (one -eighth of the window size). For each window, ChIANet \nproduced predictions for both the contact map and loop matrix. Adjacent regional outputs were then \nmerged to reconstruct the full chromosome-scale maps. To correct for regions of overlap between \nneighboring windows, we calculated the number of overlapping predictions per pixel and \nnormalized each pixel value by its corresponding overlap count. The resulting genome-wide contact \nmaps and loop matrices provide continuous, bias -corrected representations of protein -mediated \nchromatin interactions suitable for downstream quantitative and visualization analyses \n(Supplementary Fig. 7, 8). \nDistance-stratified correlation and AUPR analysis \nDistance-stratified metrics were computed to evaluate model performance across genomic \ndistances. For each predicted 2.1-Mb window, we measured the average performance as a function \nof distance i (in bins), corresponding to diagonals offset by i from the main diagonal of the contact \nmatrix. The distance -stratified correlation was defined as the Pearson correlation coefficient \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nbetween the predicted and experimental values along the i-th offset diagonal. Similarly, the distance-\nstratified AUPR was calculated by comparing the predicted and experimental loop matrices along \nthe same offset positions. For both metrics, values were averaged across all genomic windows to \nobtain genome-wide estimates of model performance as a function of genomic distance. \nComparison with existing methods \nDue to no existing framework specifically addresses genome-wide prediction of protein-mediated \nchromatin interactions, we benchmarked ChIANet against previously published models that \nrepresent the most relevant variants for each prediction task, using task-specific evaluation metrics \nto ensure fair comparison. \nFor the contact map prediction task, we compared ChIANet with three state -of-the-art Hi -C \nprediction models —Akita 24, Orca 25, and C.Origami 26—which are most closely aligned in \narchitecture and objective. Other methods such as DeepC 27 and Epiphany 28, which rely on multi-\nomics pretraining, and ChromaFold 55, which uses single-cell ATAC-seq as input, were excluded as \nthey are not directly comparable to our protein-specific setting. \nTo adapt Akita for this study, we modified the model’s input and output window size from 1 Mb \nto 2.1 Mb, and adjusted the output resolution to 256*256 to match ChIANet’s prediction scale. For \nOrca, we selected its 2.1 -Mb (8,192 -bp resolution) configuration. For C.Origami, we removed \nATAC-seq features and replaced CTCF ChIP -seq tracks with protein -specific ChIP-seq profiles \ncorresponding to CTCF, Cohesin, or RNAPII. All models were retrained from scratch using the \nsame GM12878 ChIA -PET datasets for the three  proteins to ensure a consistent training and \nevaluation setup. \nFor the loop prediction task, we focused on models capable of de novo , genome -wide loop \nprediction rather than anchor -pair classification, as most existing approaches require pre -defined \nloop candidates. We therefore compared ChIANet with Peakachu 29 and DeepChIA-PET 30, two \nframeworks capable of predicting loops directly from chromatin contact data. \nFor Peakachu, we used 10-kb-resolution Hi-C contact maps as input and the quality -controlled, \nprotein-mediated loops as training labels. For DeepChIA -PET, we adapted the original input and \noutput dimensions from 250 bins (10 kb) to 210 bins to match ChIANe t’s 2.1-Mb window. The \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nmodel was trained using paired Hi-C and protein-specific ChIP-seq signals as input, with the same \nfiltered ChIA-PET loops as supervision. \nAll models were trained and evaluated under identical genomic partitions and preprocessing \nprotocols described above to enable fair, protein-specific performance comparison. \nAggregate Peak Analysis (APA) of Predicted Loops \nTo evaluate the aggregate enrichment of predicted loops, we performed an Aggregate Peak \nAnalysis (APA) using loops generated by ChIANet. First, all loop predictions from individual 2.1-\nMb windows were merged to obtain chromosome -scale prediction matrices. From these genome-\nwide predictions, the top-scoring loops were selected at different TopN (N = 1k, 5k, 10k, 20k, 50k). \nEach set of predicted loop anchors was then mapped onto the corresponding experimental Hi -C \ncontact matrices of the same cell type. \nAPA was conducted using the Juicer toolkit, which computes two standard quantitative metrics:  \nP2LL (Peak-to-Lower-Left ratio):  the ratio of the central pixel intensity (representing the \npredicted loop) to the average intensity of the lower -left background region, reflecting local \nenrichment strength. \nZscoreLL: a standardized enrichment score calculated by normalizing the central signal to the \nvariance of the lower-left background, capturing the statistical significance of enrichment. \nAPA heatmaps were visualized as 2D intensity plots centered on loop anchors, with darker central \npixels indicating stronger enrichment of Hi -C contacts at predicted loop positions. Separate APA \nmaps were generated for each Top N threshold to assess loop qu ality and the trade -off between \nprecision and coverage. \nUMAP visualization of contact map manifolds \nFor each cell type, we embedded the genome -wide collection of predicted contact maps into a \ntwo-dimensional space for visualization. Specifically, for every 2.1-Mb window, the upper triangle \nof the ChIANet-predicted contact matrix (diagonal excluded) was vectorized to form a single feature \nvector per window. The resulting matrix of window -by-features was first reduced to 50 principal \ncomponents using PCA 56 implemented in scikit -learn (Python), and then further reduced to two \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\ndimensions using UMAP 57 implemented in the umap -learn Python package. PCA outputs were \npassed directly to UMAP, and a fixed random seed was set for reproducibility. The two-dimensional \nUMAP embeddings were used to visualize the distribution and separation of protein-specific contact \nmap windows within each cell type. \nContact-map conservation score \nFor each genomic window w and protein p, we computed a cross -cell conservation score by \naveraging the pairwise Pearson correlations of the predicted contact maps across the K = 7 cell types: \n𝑆𝑐𝑜𝑟𝑒𝑐𝑜𝑛𝑠\n(𝑝) (𝑤) = 2\n𝐾(𝐾 − 1) ∑ 𝑐𝑜𝑟𝑟(𝑀𝑤,𝑖\n(𝑝), 𝑀𝑤,𝑗\n(𝑝))\n1≤𝑖<𝑗≤𝐾\n \nHere 𝑐𝑜𝑟𝑟(·,·) is the same window -level PCC used elsewhere (computed on the upper triangle \nwith the standard distance range), and 𝑀𝑤,𝑖\n(𝑝) denotes the contact map for window w in cell type i. \nThis yields a single scalar summarizing how similarly a window is organized across cell types. \nPan-cell-type loop definition \nTo characterize protein -mediated chromatin loops across multiple cell types, we performed \ngenome-wide ChIANet predictions for seven representative human cell lines (GM12878, H1, A549, \nHeLa-S3, IMR90, HepG2, and K562). For each protein (CTCF, Cohesin, and R NAPII), the top \n10,000 predicted loops per cell type were combined to construct a pan -cell union set. Duplicate \nloops were merged based on genomic coordinates, and the corresponding prediction probability \nscores were retained for downstream analyses. \nThis process yielded 34,751 CTCF, 48,847 Cohesin, and 46,093 RNAPII loops, representing the \ncomprehensive cross-cell landscape of protein-mediated chromatin interactions. \nStatic score and classification of stable and variable loops \nTo quantify the cross-cell stability of each loop, we defined a static score based on the distribution \nof predicted loop probabilities across the seven cell types. For each loop, the prediction scores were \nfirst normalized to form a probability distribution 𝑓𝑗 over cell types. The divergence between this \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nempirical distribution and a uniform reference distribution 𝑞𝑗  was measured using Kullback -\nLeibler (KL) divergence: \n𝐷𝐾𝐿(𝑓 ∥ 𝑞) = ∑ 𝑓𝑗 log2(𝑓𝑗\n𝑞𝑗\n)\n𝑗\n \nThe static score was then defined as the inverse of this divergence, weighted by the mean predicted \nprobability: \n𝑆𝑐𝑜𝑟𝑒𝑠𝑡𝑎𝑡𝑖𝑐 = 1\n𝐷𝐾𝐿(𝑓 ∥ 𝑞) × 𝑚𝑒𝑎𝑛(𝑝𝑟𝑜𝑏) \nA higher static score indicates that a loop is consistently detected across multiple cell types (i.e., \nmore conserved). Within each protein’s pan-cell union set, the top 20% of loops by static score were \nclassified as stable loops, while the bottom 20% were designated as variable loops. \nDefinition of housekeeping and cell-type-specific genes and loop enrichment analysis \nRNA-seq data for the seven human cell types (GM12878, H1, A549, HeLa-S3, IMR90, HepG2, \nand K562) were obtained from the ENCODE  (http://www.encodeproject.org/) consortium \n(Supplementary Table 3). For each gene, expression levels across the seven cell types were used to \ncompute the same KL divergence -based variability metric as defined in Static score and \nclassification of stable and variable loops. Genes with low expression (minimum TPM ≤ 1) were \nexcluded. Genes within the top 10% of KL divergence (highly variable expression) were classified \nas cell-type-specific genes, whereas those within the bottom 10% (low variability) were designated \nas housekeeping genes. \nTo assess whether variable or stable loops were preferentially associated with specific gene \nclasses, we examined the overlap between loop anchors and gene bodies or promoter regions. For \neach loop set (variable vs. stable), we tabulated the number of loops overlapping cell-type-specific \nor housekeeping genes and compared these to non-overlapping loops. Two-sided Fisher’s exact tests \nwere used to evaluate the enrichment significance of each association. \nDefinition of genomic regulatory elements \nTo establish a unified reference set of regulatory annotations, we compiled four major classes of \ngenomic elements from ENCODE and GENCODE 58 resources (Supplementary Table 3). \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nContact domain boundary: Contact domain boundaries were aggregated from high-resolution \nHi-C data of GM12878, A549, HepG2, IMR90, and K562 (H1 and HeLa -S3 were unavailable). \nBoundaries from individual cell types were merged, and overlapping segments were unified to \ndefine a pan-cellular set of domain boundaries. \nEnhancer: Enhancers were derived from H3K27ac narrowPeak files across seven ENCODE cell \ntypes (GM12878, H1, A549, HeLa-S3, IMR90, HepG2, and K562). Peaks from all cell types were \ncombined, and overlapping regions were merged to generate a pan-cellular enhancer catalogue. \nSilencer: Silencers were similarly defined from H3K27me3 narrowPeak files across the six cell \ntypes (A549 were unavailable), followed by merging of overlapping regions to construct a \ncomprehensive silencer reference. \nPromoter: Promoters were defined as ±500 bp windows flanking the transcription start site \n(TSS) of each annotated gene in GENCODE v29.  \nChromatin state enrichment analysis of cell-type-specific loops \nChromatin state annotations for seven cell types were obtained from the Roadmap Epigenomics \nMapping Consortium 59 (15-state model; Supplementary Table 4). To simplify downstream analysis, \nthe 15 chromatin states were consolidated into eight major categories: transcription start site (TSS: \n1_TssA, 2_TssAFlnk), bivalent (BIV: 10_TssBiv, 11_BivFlnk), transcription (TX: 3_TxFlnk, 4_Tx, \n5_TxWk), repressive (REPRESS: 13_ReprPC, 14_ReprPCWk), repeat -associated (REPEAT: \n8_ZNF/Rpts), enhancer (ENH: 12_EnhBiv, 6_EnhG, 7_Enh), heterochromatic (HET: 9_Het), and \nquiescent (QUIES: 15_Quies). \nFor each cell type, cell -type-specific loops were defined by computing loop -specific z-scores \nbased on the proportion of loops present across other cell types. Loop probabilities were first logit-\ntransformed to approximate a normal distribution: \n𝑙𝑜𝑔𝑖𝑡(𝑝) = 𝑙𝑜𝑔( 𝑝\n1 − 𝑝) \nA pairwise comparison strategy was used to calculate a specificity z-score for each loop: \n𝑍𝑖 = 𝑃𝑖\n(𝑡𝑎𝑟𝑔𝑒𝑡) − 𝜇𝑖\n(𝑜𝑡ℎ𝑒𝑟𝑠)\n𝜎𝑖\n(𝑜𝑡ℎ𝑒𝑟𝑠)  \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nwhere 𝑃𝑖\n(𝑡𝑎𝑟𝑔𝑒𝑡) is the loop probability in the target cell type, and 𝜇𝑖\n(𝑜𝑡ℎ𝑒𝑟𝑠), 𝜎𝑖\n(𝑜𝑡ℎ𝑒𝑟𝑠) represent \nthe mean and standard deviation across the remaining cell types. Loops ranked within the top 5% \nof 𝑍𝑖 values were designated as cell-type-specific loops for that cell type. \nTo assess the enrichment of chromatin states in cell -type-specific loop anchors, we constructed \n2×2 contingency tables for each chromatin category. (1) anchors overlapping the chromatin state in \nspecific loops, (2) anchors not overlapping the state in specific loops,  (3) anchors overlapping the \nstate in non-specific loops, and (4) anchors not overlapping the state in non-specific loops. Statistical \nsignificance of enrichment or depletion was assessed using two-sided Fisher’s exact tests. \nIdentification and functional analysis of protein-mediated super-enhancer loops \nSuper-enhancers (SEs) are large clusters of enhancers densely occupied by transcriptional co -\nactivators, and they play central roles in establishing cell identity-specific gene expression programs \n60. To investigate the relationship between SEs and protein-mediated chromatin loops, we obtained \nSE annotations for seven human cell types from the SEdb database (v2.0) 61. \nCell-type-specific SEs were defined as those unique to a single cell type (i.e., not overlapping SE \nregions in any other cell type), and their distributions are shown in Supplementary Fig. 1 1a, c. \nShared SEs were defined as SEs present in at least three cell types, as determined from the SE -\nsharing curve (Supplementary Fig. 11b). \nProtein-mediated SE loops were identified by intersecting loop anchors with cell -type-specific \nSEs. For functional characterization, GREAT (Genomic Regions Enrichment of Annotations Tool, \nv4.0.4) 62 was applied to both SE regions and protein-mediated SE loops using default association \nparameters. Gene Ontology (GO) enrichment analyses were performed under the “Basal plus \nextension” association rule, and enriched biological processes were reported aft er multiple testing \ncorrection. \nEcDNA-associated chromatin contact definition \nEcDNA-associated chromatin contacts were obtained from processed RNAPII ChIA -Drop \ninteraction datasets for PC3DM, COLO320DM and GBM171DM cancer cells. Chromatin \ninteractions were represented as multiplex interaction complexes with genomic interaction anchors. \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nWe directly used the provided interaction anchors for downstream analyses. Chromatin contacts \nwere defined as ecDNA-associated if at least one interaction anchor overlapped annotated ecDNA \nregions. Contacts were classified as cis interactions (among ecDNA regions) or trans interactions \n(between ecDNA and chromosomal regions). For ecDNA -chromosomal interactions, only trans \ncontacts with one anchor overlapping an ecDNA region and the other anchor mapping to \nchromosomal loci were retained. Overlapping anchors across libraries were merged to define \ninteraction nodes, which were used for all subsequent analyses, including epigenomic signal \naggregation, enrichment analysis and comparison with ChIANet-predicted RNAPII loop anchors. \nnGEM stratification and epigenomic association analysis \nHigher-order chromatin interaction units were analyzed using n-way Genome Extrusion Modules \n(nGEMs) defined in the processed RNAPII ChIA-Drop datasets. Each nGEM represents a multiplex \nchromatin interaction complex comprising multiple genomic fragments ide ntified within a single \ninteraction unit. nGEMs were stratified by interaction complexity based on the number of \nparticipating fragments and grouped into deciles from low to high complexity. For each nGEM \ndecile, epigenomic signal intensities of MED1 and H3K27ac were aggregated at the corresponding \ninteraction anchors. Associations between nGEM complexity and epigenomic signals were \nevaluated using Spearman correlation across deciles and biological replicates. \nIdentification of ecDNA-associated super-enhancers \nEcDNA-associated super-enhancers (ecSEs) were identified by integrating MED1 ChIP -seq-\ndefined super-enhancer annotations with annotated ecDNA regions. Super-enhancer intervals were \nobtained from MED1 ChIP -seq data using the ROSE algorithm with default parameters. Super -\nenhancers whose genomic coordinates overlapped annotated ecDNA regions were classified as \necDNA-associated super -enhancers (ecSEs). Genomic overlap was determined by interval \nintersection, and all super-enhancers satisfying this criterion were retained as ecSEs for downstream \nanalyses. \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nEnrichment of predicted loop anchors at ecDNA contacts \nTo assess whether ChIANet -predicted RNAPII loop anchors were preferentially enriched at \necDNA-associated chromatin contacts, we compared their overlap with ecDNA contact anchors \nagainst a distance-matched random background. Genome-wide RNAPII ChIP-seq peaks were used \nas the background pool. For each set of predicted loop anchors, random RNAPII peaks were \nsampled to match the genomic distance distribution to ecDNA contact anchors, thereby controlling \nfor distance-dependent biases. \nEnrichment was evaluated using Fisher’s exact test by contrasting the number of predicted loop \nanchors overlapping ecDNA contact anchors with the corresponding overlap observed for the \ndistance-matched random RNAPII peaks. This procedure was repeated 50 times with independent \nrandom samplings. Odds ratios (ORs) were calculated for each iteration, and the distribution of ORs \nwas used to summarize enrichment across randomizations. \nMotif enrichment analysis of ecDNA-associated RNAPII loop anchors \nMotif enrichment analysis was performed on ChIANet -predicted RNAPII loop anchors that \noverlapped annotated ecDNA regions. Genomic sequences corresponding to these anchors were \nanalyzed using the Analysis of Motif Enrichment (AME) tool from the MEME suite, with all \nparameters set to default values. Motif enrichment was assessed against the JASPAR CORE 2026 \ntranscription factor binding motif database. To integrate motif enrichment with transcriptional \nactivity, RNA-seq expression data were used to annotate tr anscription factors corresponding to \nenriched motifs. Raw RNA -seq read counts were used as the measure of gene expression, and \ntranscription factors with expression counts greater than 300 were retained for downstream analysis. \nMotifs exhibiting statistically significant enrichment were identified based on a threshold of -log₂(p-\nvalue) > 20. Analyses focused on transcription factor binding motifs meeting both the enrichment \nand expression criteria. \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nGene Ontology enrichment analysis of ChIANet-predicted loop anchors \nGene Ontology (GO) enrichment analysis was performed to characterize functional programs \nassociated with ChIANet -predicted chromatin loop anchors in PC3DM, COLO320DM and \nGBM171DM cells. Predicted loop anchors for each cell type were analyzed separately. Lo op-\nassociated genes were assigned using the Genomic Regions Enrichment of Annotations Tool \n(GREAT, v4.0.4) by associating loop anchor regions with nearby genes. GREAT analyses were \nconducted under the basal plus extension association rule using default parameters. Enrichment was \nassessed for GO biological process terms, and statistically significant categories were identified \nafter multiple testing correction. All GO enrichment analyses were performed independently for \neach cancer cell type using the corresponding sets of predicted loop anchors. \n \nDATA AVAILABILITY \nMost of the ChIA-PET, ChIP-seq, RNA-seq and Hi-C datasets used in this study were obtained from \npublic resources, including the ENCODE portal and the 4D Nucleome (4DN) data portal, with \naccession codes provided in the corresponding Methods sections or Supplementary Information. \necDNA-associated chromatin interaction datasets for the three cancer cell types (PC3DM, \nCOLO320DM and GBM171DM) were obtained from the NCBI Gene Expression Omnibus under \naccession number GSE275060. \nCODE AVAILABILITY \nThe code for ChIANet is available at https://github.com/lhy0322/ChIANet. \n \nSUPPLEMENTARY DATA \nSupplementary Tables 1-4 and Figs. 1-12. \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\n \nACKNOWLEDGEMENTS \nThe authors thank colleagues for helpful discussions. Detailed acknowledgements and funding \ninformation will be provided upon journal submission. \n \nREFERENCES \n1. Davies, J.O., Oudelaar, A.M., Higgs, D.R. & Hughes, J.R. How best to identify \nchromosomal interactions: a comparison of approaches. Nature methods 14, 125-134 \n(2017). \n2. Jerkovic, I. & Cavalli, G. Understanding 3D genome organization by multidisciplinary \nmethods. Nature reviews Molecular cell biology 22, 511-528 (2021). \n3. Rowley, M.J. & Corces, V.G. Organizational principles of 3D genome architecture. Nature \nReviews Genetics 19, 789-800 (2018). \n4. Tang, Z. et al. CTCF-mediated human 3D genome architecture reveals chromatin \ntopology for transcription. Cell 163, 1611-1627 (2015). \n5. Weintraub, A.S. et al. YY1 is a structural regulator of enhancer-promoter loops. Cell 171, \n1573-1588. e28 (2017). \n6. Bulger, M. & Groudine, M. Functional and mechanistic diversity of distal transcription \nenhancers. Cell 144, 327-339 (2011). \n7. Lieberman-Aiden, E. et al. Comprehensive mapping of long-range interactions reveals \nfolding principles of the human genome. science 326, 289-293 (2009). \n8. Holwerda, S.J.B. & de Laat, W. CTCF: the protein, the binding partners, the binding sites \nand their chromatin loops. Philosophical Transactions of the Royal Society B: Biological \nSciences 368, 20120369 (2013). \n9. Rubio, E.D. et al. CTCF physically links cohesin to chromatin. Proceedings of the National \nAcademy of Sciences 105, 8309-8314 (2008). \n10. DeMare, L.E. et al. The genomic landscape of cohesin-associated chromatin interactions. \nGenome research 23, 1224-1234 (2013). \n11. Grubert, F. et al. Landscape of cohesin-mediated chromatin loops in the human \ngenome. Nature 583, 737-743 (2020). \n12. Li, X. & Fu, X.-D. Chromatin-associated RNAs as facilitators of functional genomic \ninteractions. Nature Reviews Genetics 20, 503-519 (2019). \n13. Struhl, K. Chromatin structure and RNA polymerase II connection: implications for \ntranscription. Cell 84, 179-182 (1996). \n14. Beacon, T.H. et al. The dynamic broad epigenetic (H3K4me3, H3K27ac) domain as a \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nmark of essential genes. Clinical epigenetics 13, 138 (2021). \n15. Cai, Y. et al. H3K27me3-rich genomic regions can function as silencers to repress gene \nexpression via chromatin interactions. Nature communications 12, 719 (2021). \n16. Denker, A. & De Laat, W. The second decade of 3C technologies: detailed insights into \nnuclear organization. Genes & development 30, 1357-1382 (2016). \n17. Simonis, M., Kooren, J. & De Laat, W. An evaluation of 3C-based methods to capture \nDNA interactions. Nature methods 4, 895-901 (2007). \n18. Van De Werken, H.J. et al. Robust 4C-seq data analysis to screen for regulatory DNA \ninteractions. Nature methods 9, 969-972 (2012). \n19. Fullwood, M.J. et al. An oestrogen-receptor-α-bound human chromatin interactome. \nNature 462, 58-64 (2009). \n20. Mumbach, M.R. et al. HiChIP: efficient and sensitive analysis of protein-directed genome \narchitecture. Nature methods 13, 919-922 (2016). \n21. Fang, R. et al. Mapping of long-range chromatin interactions by proximity ligation-\nassisted ChIP-seq. Cell research 26, 1345-1348 (2016). \n22. Piecyk, R.S., Schlegel, L. & Johannes, F. Predicting 3D chromatin interactions from DNA \nsequence using Deep Learning. Computational and Structural Biotechnology Journal 20, \n3439-3448 (2022). \n23. Zhang, Y. et al. Computational methods for analysing multiscale 3D genome \norganization. Nature Reviews Genetics 25, 123-141 (2024). \n24. Fudenberg, G., Kelley, D.R. & Pollard, K.S. Predicting 3D genome folding from DNA \nsequence with Akita. Nature methods 17, 1111-1117 (2020). \n25. Zhou, J. Sequence-based modeling of three-dimensional genome architecture from \nkilobase to chromosome scale. Nature genetics 54, 725-734 (2022). \n26. Tan, J. et al. Cell-type-specific prediction of 3D chromatin organization enables high-\nthroughput in silico genetic screening. Nature biotechnology 41, 1140-1150 (2023). \n27. Schwessinger, R. et al. DeepC: predicting 3D genome folding using megabase-scale \ntransfer learning. Nature methods 17, 1118-1124 (2020). \n28. Yang, R. et al. Epiphany: predicting Hi-C contact maps from 1D epigenomic signals. \nGenome Biology 24, 134 (2023). \n29. Salameh, T.J. et al. A supervised learning framework for chromatin loop detection in \ngenome-wide contact maps. Nature communications 11, 3428 (2020). \n30. Liu, T. & Wang, Z. DeepChIA-PET: Accurately predicting ChIA-PET from Hi-C and ChIP-\nseq with deep dilated networks. PLOS Computational Biology 19, e1011307 (2023). \n31. Hong, Z., Zeng, X., Wei, L. & Liu, X. Identifying enhancer–promoter interactions with \nneural network based on pre-trained DNA vectors and attention mechanism. \nBioinformatics 36, 1037-1043 (2020). \n32. Cao, F. et al. Chromatin interaction neural network (ChINN): a machine learning-based \nmethod for predicting chromatin interactions from DNA sequences. Genome biology \n22, 226 (2021). \n33. Tao, H. et al. Computational methods for the prediction of chromatin interaction and \norganization using sequence and epigenomic profiles. Briefings in bioinformatics 22, \nbbaa405 (2021). \n34. Kai, Y. et al. Predicting CTCF-mediated chromatin interactions by integrating genomic \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nand epigenomic features. Nature communications 9, 4221 (2018). \n35. Dao, F.-Y. et al. DeepYY1: a deep learning approach to identify YY1-mediated \nchromatin loops. Briefings in bioinformatics 22, bbaa356 (2021). \n36. Vaswani, A. et al. Attention is all you need. Advances in neural information processing \nsystems 30(2017). \n37. Kendall, A., Gal, Y. & Cipolla, R. Multi-task learning using uncertainty to weigh losses for \nscene geometry and semantics. in Proceedings of the IEEE conference on computer \nvision and pattern recognition 7482-7491 (2018). \n38. Pal, K., Forcato, M. & Ferrari, F. Hi-C analysis: from data generation to integration. \nBiophysical reviews 11, 67-78 (2019). \n39. Wang, X., Cairns, M.J. & Yan, J. Super-enhancers in transcriptional regulation and \ngenome organization. Nucleic acids research 47, 11481-11496 (2019). \n40. Bailey, C. et al. Origins and impact of extrachromosomal DNA. Nature 635, 193-200 \n(2024). \n41. Kraft, K. et al. Enhancer activation from transposable elements in extrachromosomal \nDNA. Nature cell biology, 1-11 (2025). \n42. Yan, X., Mischel, P. & Chang, H. Extrachromosomal DNA in cancer. Nature Reviews \nCancer 24, 261-273 (2024). \n43. Taghbalout, A. et al. Extrachromosomal DNA associates with nuclear condensates and \nreorganizes chromatin structures to enhance oncogenic transcription. Cancer cell 43, \n2191-2205. e6 (2025). \n44. Cao, J. et al. Three-dimensional regulation of transcription. Protein & cell 6, 241-253 \n(2015). \n45. Phair, R.D. et al. Global nature of dynamic protein-chromatin interactions in vivo: three-\ndimensional genome scanning and dynamic interaction networks of chromatin proteins. \nMolecular and cellular biology 24, 6393-6402 (2004). \n46. Zheng, Y. et al. A deep generative model for deciphering cellular dynamics and in silico \ndrug discovery in complex diseases. Nature Biomedical Engineering, 1-26 (2025). \n47. Miladinovic, D. et al. In silico biological discovery with large perturbation models. Nature \nComputational Science, 1-12 (2025). \n48. de Souza, N. The ENCODE project. Nature methods 9, 1046-1046 (2012). \n49. Dekker, J. et al. The 4D nucleome project. Nature 549, 219-226 (2017). \n50. Lee, B. et al. ChIA-PIPE: A fully automated pipeline for comprehensive ChIA-PET data \nanalysis and visualization. Science advances 6, eaay2078 (2020). \n51. Ramírez, F. et al. deepTools2: a next generation web server for deep-sequencing data \nanalysis. Nucleic acids research 44, W160 (2016). \n52. Karolchik, D. et al. The UCSC genome browser database. Nucleic acids research 31, 51-\n54 (2003). \n53. Imambi, S., Prakash, K.B. & Kanagachidambaresan, G. PyTorch. in Programming with \nTensorFlow: solution for edge computing applications 87-104 (Springer, 2021). \n54. Kingma, D.P. Adam: A method for stochastic optimization. arXiv preprint \narXiv:1412.6980 (2014). \n55. Gao, V.R. et al. ChromaFold predicts the 3D contact map from single-cell chromatin \naccessibility. Nature Communications 15, 9432 (2024). \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\n56. Mać kiewicz, A. & Ratajczak, W. Principal components analysis (PCA). Computers & \nGeosciences 19, 303-342 (1993). \n57. McInnes, L., Healy, J. & Melville, J. Umap: Uniform manifold approximation and \nprojection for dimension reduction. arXiv preprint arXiv:1802.03426 (2018). \n58. Frankish, A. et al. GENCODE 2021. Nucleic acids research 49, D916-D923 (2021). \n59. Bernstein, B.E. et al. The NIH roadmap epigenomics mapping consortium. Nature \nbiotechnology 28, 1045-1048 (2010). \n60. Pott, S. & Lieb, J.D. What are super-enhancers? Nature genetics 47, 8-12 (2015). \n61. Wang, Y. et al. SEdb 2.0: a comprehensive super-enhancer database of human and \nmouse. Nucleic Acids Research 51, D280-D290 (2023). \n62. McLean, C.Y. et al. GREAT improves functional interpretation of cis-regulatory regions. \nNature biotechnology 28, 495-501 (2010). \n \n  \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nFIGURE \n \nFig. 1: Overview of the ChIANet architecture and workflow.  (a) Protein-specific ChIP -seq \nprofiles, together with the reference genome sequence, are used as inputs to predict protein-mediated \nchromatin interactions. ChIANet generates contact maps and loop for each 2.1-Mb window, which \nare tiled along chromosomes and merged to obtain genome -scale predictions. (b) ChIANet is a \nmultimodal, multi-task encoder-decoder framework that integrates genomic sequence and protein-\nspecific ChIP -seq signals to predict protein -mediated chromatin organization. ChIANet jointly \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\noutputs a contact map and loop matrix under a heteroscedastic uncertainty -weighted loss. Models \nare trained with a 19/2/2 chromosome split (train/validation/test) to ensure unbiased evaluation. The \ntrained model supports de novo inference in unseen cell types from ChIP -seq input alone.  (c) \nPredicted chromatin interaction maps for CTCF, Cohesin and RNAPII across seven human cell \ntypes enable comparative analyses of protein -mediated chromatin architectures across functional \ncontexts. The same framework is app lied to cancer genomes to analyze RNAPII -mediated \nchromatin interactions associated with extrachromosomal DNA (ecDNA). \n  \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\n \nFig. 2: ChIANet accurately predicts protein-mediated contact maps and loop interactions. (a) \nModel performance comparison on GM12878 CTCF test sets. Pearson (left) and Spearman (right) \ncorrelations between predicted and experimental contact maps across all 2.1 -Mb windows, \nbenchmarked against Akita, Orca and C.Origami. Violin plots show the distribution across genomic \nwindows; horizontal lines indicate minimum, mean and maximum values. Statistical significance \nwas assessed using paired two -sided Wilcoxon signed-rank tests. (b) Distance-stratified Pearson \ncorrelations between predicted and experimental contact maps for different models.  c-d, \nRepresentative examples of experimental and model -predicted results for GM12878 CTCF. (c) \nContact maps and (d) loops shown for training (chr1), validation (chr3), and test (chr5) genomic \nregions (e) Genome-wide comparison of AUC and AUPR for loop prediction (positive-to-negative \nratio = 1:10) among Peakachu, DeepChIA -PET and ChIANet. (f) Chromosome-scale receiver \noperating characteristic (ROC; left) and precision-recall (PR; right) curves showing overall AUC \nand AUPR on the GM12878 CTCF test set.  (g) Distance-stratified area under the precision -recall \ncurve (AUPR) for loop prediction. (h) Protein-specific ChIP-seq input profiles corresponding to the \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nsame regions. (i) Comparison of model performance across different input modalities —ChIP-seq \nonly, sequence only, and sequence + ChIP-seq (full model)—on the contact map prediction task. \nBars show Pearson correlation coefficients (PCC, left) and mean squared error (MSE, right) on the \nvalidation and test sets.  (j) Comparison of window-level AUC and AUPR for the loop prediction \ntask under an imbalanced classification ratio (1:100, Neg:Pos). \n  \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\n \nFig. 3: De novo  prediction of cell type -specific, protein-mediated chromatin structure.  (a) \nExperimental and predicted CTCF -mediated contact maps and loop matrices for GM12878 (left) \nand H1 (center), together with their difference maps (GM12878 - H1, right). Blue/red in the \ndifference panels indicate contacts or loops that are stronger in GM12878 or H1, respectively. The \nbottom track shows the CTCF ChIP-seq signal used for conditioning. (b) Genome-wide comparison \nof contact map prediction performance on H1 CTCF. Violin plots show Pearson correlation (left) \nand Spearman rank correlation (right) bet ween predicted and experimental 2.1 -Mb windows (n = \n5,977), with ChIANet achieving the highest median concordance.  Statistical significance was \nassessed using paired two-sided Wilcoxon signed-rank tests. (c) Genome-wide comparison of loop \nprediction on H1 CTCF at a 1:10 positive -to-negative ratio. Violin plots show window-level AUC \n(left) and AUPR (right), indicating that ChIANet maintains competitive loop -level accuracy in de \nnovo prediction. (d) APA quality metrics for de novo predicted H1 CTCF loops as a function of the \nnumber of top-ranked loops retained. P2LL (solid line) and ZscoreLL (dashed line) were computed \non loop sets of increasing size (Top 1k to Top 50k). Both metrics peak around Top 10k, indicating \nthat this cutoff yields the best trade -off between loop strength and noise. (e) APA heatmaps of de \nnovo predicted H1 loops at different cutoffs (Top 1k, 5k, 10k, 20k and 50k), shown in a ±100-kb \nwindow around the loop anchor pair. \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\n \nFig. 4: Distinct 3D chromatin architectures mediated by different regulatory proteins in H1. \n(a) Representative predicted contact maps for CTCF-, Cohesin-, and RNAPII-mediated chromatin \ninteractions within the same genomic region (Chr4: 85.98 -88.08 Mb). CTCF and Cohesin maps \nexhibit sharply defined domain boundaries, whereas RNAPII shows diffuse, tran scription-\nassociated interaction patterns. (b) Genome-wide pairwise Pearson correlation of predicted contact \nmaps across all 2.1 -Mb windows, illustrating strong simila rity between CTCF and Cohesin, and \nweaker correlation between RNAPII and the other two factors. (c) UMAP embedding of all \npredicted contact maps across the genome, showing substantial overlap between CTCF and Cohesin \nclusters, while RNAPII forms a distinct group, reflecting a divergent interaction landscape. (d) \nChromosome-wise Pearson correlations between predicted maps of different proteins, with the \nhighest values consistently observed between CTCF and Cohesin, indicating conserved structural \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\ncoupling across chromosomes. (e) Distance-stratified correlations between predicted contact maps \nof different proteins. Correlations between CTCF -RNAPII or Cohesin -RNAPII drop markedly at \nlong genomic distances (>524 kb), whereas CTCF -Cohesin correlations remain stable, indicating \ntheir shared role in maintaining higher-order domain organization. (f) Distribution of loop lengths \nfor the top 10,000 predicted loops of each protein. RNAPII -mediated loops are generally shorter, \nconsistent with promoter -enhancer proximity, whereas CTCF and Cohesin loops span broader \ngenomic ranges. (g) Venn diagram showing the overlap of top predicted loops among the three \nproteins, highlighting substantial sharing between CTCF and Cohesin. (h) Complementary \ncumulative degree distribution (CCDF) of loop networks for each protein, revealing scale -free \nbehavior with RNAPII networks showing higher connectivity at low degrees and CTCF/Cohesin \nnetworks forming stronger high-degree hubs. (i) Neighbor interconnectivity as a function of l oop \nhub degree. CTCF and Cohesin display greater local clustering around high -degree nodes than \nRNAPII.\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\n \nFig. 5: Cross-cell-type conservation and variability of protein -mediated chromatin \norganization. (a) Genome-wide Pearson correlation coefficients (PCCs) between predicted contact \nmaps across seven human cell types for each protein. The lower triangle of the upper heatmap shows \nCTCF-mediated correlations, the upper triangle shows Cohesin, and the bottom pa nel shows \nRNAPII. (b) Representative regions on chromosome 13 (62.39 -66.58 Mb) illustrating cross -cell-\ntype differences in contact maps for CTCF, Cohesin, and RNAP II. (c) Distribution of loop static \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nscores quantifying the cross -cell-type conservation of predicted loops for each protein. (d) \nRelationship between node degree and mean loop stability across protein -specific loop networks. \n(e) Enrichment analysis comparing variable and stable  loops near cell type -specific versus \nhousekeeping genes (P values, Fisher’s exact test). (f) Fraction of predicted loops supported across \nincreasing numbers of cell types, estimated by a Poisson-Binomial model. \n  \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\n \nFig. 6: Functional bias and regulatory potential of protein-mediated chromatin interactions. \n(a) Functional enrichment of anchors from pan-cell-type CTCF-mediated loops for contact domain \nboundaries, enhancers, and promoters, stratified by the number of loop interactions per anchor. \nStatistical significance was assessed using two-sided Fisher’s exact tests (Number of interactions = \n34,751, *P<0.05, **P<0.01, ***P<0.001; NS, not significant). (b) Classification of loops mediated \nby CTCF, Cohesin, and RNAPII across all cell types into promoter -promoter interactions (PPI), \nenhancer-promoter interactions (EPI), enhancer -enhancer interactions (EEI), silencer -promoter \ninteractions (SPI), and other types . (c) Chromatin-state composition of cell -type-specific loop \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nanchors in GM12878. (d) Fold enrichment of chromatin states for cell -type-specific loop anchors \nmediated by CTCF, Cohesin, and RNAPII in GM12878 (n = 2,149, 2,626, and 2,033, respectively). \nEnrichment significance was evaluated using two -sided Fisher’s exact tests. (e) Differential gene \nexpression associated with genes linked to cell -type-specific loops in GM12878. P -values were \ncalculated by two -sided Wilcoxon rank -sum tests. (f) Enrichment of GM12878 -specific loops \nrelative to non-specific loops at cell-type-specific super-enhancers (SE-specific) and shared super-\nenhancers (SE -shared), with significance assessed by Fisher ’s exact test. (g) Overlap of genes \nassociated with CTCF-, Cohesin-, and RNAPII-mediated loops in GM12878, showing both shared \nand unique targets. (h) Canonical pathway enrichment analysis of genes associated with SE-linked \nloops mediated by each protein in GM12878, compared with SE regions themselves. \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\n \nFig. 7: RNAPII-mediated chromatin looping at ecDNA regions in cancer cells . (a) Circular \noverview of a representative ecDNA region. From inner to outer tracks, the circos plot shows \naverage copy-number profiles, followed by ChIP -seq signal intensities of RNAPII, H3K27ac and \nMED1. Gene annotations are shown in black . (b) Heatmaps showing ChIP-seq signal densities of \nRNAPII, MED1 and H3K27ac centered on ecDNA -associated contact anchors in PC3DM cells. \nResults from two biological replicates are shown (Rep-1, n = 10,562; Rep-2, n = 9,374). Signals are \ndisplayed within ±5 kb of anchor centers. Color scales indicate normalized ChIP -seq signal \nintensities. (c, d ) Genome browser views of two high -copy-number regions on chromosome 8. \nTracks show copy-number profiles, ChIA-drop interaction signals from two biological replicates, \nMED1, RNAPII and H3K27ac ChIP -seq signals, and ChIANet -predicted RNAPII -mediated \nchromatin loops. EcDNA -associated regions are highlighted in grey. Genomic coordinates are \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nindicated above each panel. (e) Enrichment analysis of ChIANet-predicted RNAPII loop anchors at \necDNA hubs compared with randomly sampled RNAPII peaks across PC3DM, COLO320DM and \nGBM171DM cells. Violin plots show odds ratios calculated using Fisher’s exact tests. The dashed \nline indicates an odds ratio of 1. (f) Motif analysis of ChIANet-predicted RNAPII loop anchors in \nPC3DM cells. Each point represents a transcription factor motif. Point size denotes motif \nenrichment significance (-log2 P value). Color intensity indicates RNA-seq-derived gene expression \nlevels (raw expression counts) of the corresponding transcription factors. The y -axis shows gene \nexpression levels. \n  \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nEXTENDED DATA \n \nExtended Data Fig. 1: ChIANet performance on RNAPII- and Cohesin-mediated chromatin \ninteractions. (a) Benchmarking ChIANet against Akita, Orca and C.Origami on GM12878 RNAPII \n(top) and GM12878 Cohesin (bottom). For each protein, Pearson correlation (left) and Spearman \nrank correlation (middle) were computed between predicted and experimental contact maps across \nall 2.1-Mb windows (n = 511). Right, window-level mean squared error (MSE) between predicted \nand experimental contact maps for the same models, show ing that ChIANet achieves comparable \nor lower reconstruction error . (b) Distance-stratified Pearson correlation between predicted and \nexperimental contact maps for GM12878 RNAPII (left) and GM12878 Cohesin (right), showing \nthat ChIANet maintains competitive or improved correlation at longer genomic distances. (c) \nDistance-stratified AUPR for loop prediction on GM12878 RNAPII (left) and GM12878 Cohesin \n(right), compared with Peakachu and DeepChIA -PET. (d) Genome-wide loop prediction under \ndifferent positive-to-negative ratios (1:1, 1:20 and 1:100) for GM12878 CTCF, GM12878 RNAPII \nand GM12878 cohesin. Violin plots show AUC (left in each panel) and AUPR (right in each panel) \nfor Peakachu, DeepChIA -PET and ChIANet, indicating that ChIANet remains robust under \nincreasingly imbalanced classification settings. \n  \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\n \nExtended Data Fig. 2: ChIANet performance on RNAPII- and Cohesin-mediated chromatin \ninteractions. (a) Receiver operating characteristic (ROC) curves comparing loop prediction \nperformance of Peakachu, DeepChIA -PET and ChIANet for GM12878 CTCF (left), RNAPII \n(middle) and Cohesin (right) at a 1:1 positive-to-negative ratio. ChIANet achieves the highest AUC \nacross all proteins. (b) Precision-recall (PR) curves under increasingly imbalanced conditions \n(positive-to-negative ratios of 1:1, 1:20 and 1:100) for the same proteins. ChIANet consistently \nmaintains superior AUPR compared with baseline models, demonstrati ng robustness to class \nimbalance. (c) Representative examples of experimental and predicted contact maps (top) and loop \nmatrices (bottom) for RNAPII (left) and Cohesin (right) in GM12878. Experimental contact maps \nand loop calls are shown in the upper panels, while ChIANet predictions are show n below, \nhighlighting accurate reconstruction of protein-mediated chromatin interactions at the 2.1-Mb scale. \n  \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\n \nExtended Data Fig. 3: Representative examples of protein -specific contact map predictions \nacross models. Comparison of predicted and experimental contact maps for randomly selected 2.1-\nMb genomic windows across CTCF-mediated chromatin interactions in GM12878 and H1 cell types. \nModels are grouped by input modality: sequence-based models (Akita, Orca), multi-modal models \n(C.Origami, ChIANet), and the corresponding experimental ChIA -PET contact maps. For each \nprotein and region, predictions from Akita, Orca, C.Origami, and ChIAN et are shown alongside \nexperimental ChIA -PET maps. ChIANet demonstrates stronger agreement with experimental \ncontact patterns, accurately recovering both short- and long-range structures, while maintaining cell-\ntype specificity in de novo predictions (bottom panels). \n  \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\n \nExtended Data Fig. 4: Representative examples of protein -specific loop predictions across \nmodels. Comparison of loop prediction outputs from Peakachu, DeepChIA-PET, and ChIANet for \nrandomly selected 2.1 -Mb genomic regions across CTCF -mediated chromatin interactions in \nGM12878 and H1 cell types. Models are grouped by input modality: Hi -C-derived models \n(Peakachu, DeepChIA -PET), ChIP -seq-derived model (ChIANet), and the corresponding \nexperimental ChIA-PET loops and contact maps. For each region, predicted loop matrices (red) are \nshown alongside experimental contact maps (blue). ChIANet pred ictions exhibit higher spatial \nprecision and cleaner topological boundaries compared to Hi -C-based models, capturing both \nprominent and distal loops visible in experimental data. De novo predictions in H1 demonstrate the \nmodel’s generalization ability across cell types using only sequence and protein -specific ChIP-seq \nas input. \n  \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\n \nExtended Data Fig. 5: Ablation analysis of ChIANet across proteins and evaluation of \nmultimodal contributions.  (a) Training dynamics of the sequence only, ChIP -seq only, and \nsequence + ChIP-seq (full) ChIANet models on GM12878 datasets for CTCF, RNAPII, and Cohesin. \nCurves show the contact map loss (left) and loop loss (right) across training epochs, indicating \nimproved convergence and lower task -specific losses when integrating both input mo dalities. (b) \nComparison of contact map prediction performance across models for RNAPII and Cohesin on the \nvalidation and test sets. Bars represent Pearson correlation coefficients (PCC, left) and mean squared \nerror (MSE, right), showing consistent performance gains when combining ChIP-seq and sequence \nfeatures. (c) Distance-stratified Pearson correlations between predicted and experimental contact \nmaps for CTCF, RNAPII and Cohesin, demonstrating that multimodal integration enhances long -\nrange contact modeling compared to single-modality inputs. (d) Distance-stratified area under the \nprecision-recall curve (AUPR) for loop prediction, showing that the combined model achieves \nhigher precision-recall performance across genomic distances. (e) Chromosome-scale comparisons \nof AUC and AUPR for loop prediction under a 1:100 negative-to-positive sampling ratio for CTCF, \nRNAPII, and Cohesin. The full model consistently achieves the highest overall accuracy, \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nhighlighting that both DNA sequence and ChIP -seq signals are essential for accurately modeling \nprotein-mediated chromatin interactions. \n  \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\n \nExtended Data Fig. 6: Genome-wide chromosome-level evaluation of de novo contact map \nprediction in H1 CTCF. a-c, Per-chromosome violin plots showing the distribution of ChIANet \nand baseline model performance (Akita, Orca, and C.Origami) across all 2.1-Mb windows in the de \nnovo prediction of H1 CTCF -mediated contact maps. Each violin represents the variability of \nwindow-level scores within each chromosome, with the number of evaluated windows (n) indicated \nabove each panel. ChIANet consistently achieves higher correlation and lo wer error across most \nchromosomes. (a) Pearson correlation (PCC). (b) Spearman rank correlation. (c) Mean squared error \n(MSE). (d) Distance-stratified Pearson correlation per chromosome, showing the decay of \npredictive accuracy with increasing genomic distance. ChIANet maintains stronger long -range \ncorrelation than sequence-based (Akita, Orca) and multimodal (C.Origami) models, demonstrating \nrobust genome-wide modeling of distal CTCF-mediated chromatin interactions. \n  \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\n \nExtended Data Fig. 7: Genome-wide chromosome-level evaluation of de novo loop prediction \nin H1 CTCF. a-d, Per-chromosome violin plots showing window -level AUPR for ChIANet and \nbaseline loop detectors (Peakachu, DeepChIA -PET) across all 2.1 -Mb windows in the de novo \nprediction setting. Each violin summarizes the distribution within one chromosome; the number of \nevaluated windows (n) is indicated above each panel. Class -imbalance is varied by the \npositive:negative ratio used to compute AUPR: (a) 1:1, (b) 1:10 , (c) 1:20, (d) 1:100. This figure \nprovides genome-wide, chromosome-resolved comparisons of AUPR under increasing imbalance \nfor the three models. \n  \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\n \nExtended Data Fig. 8: Cross-protein comparison of predicted chromatin architectures across \nmultiple cell types. (a) Chromosome-wise Pearson correlation matrices of predicted contact maps \nbetween CTCF-, Cohesin-, and RNAPII-mediated interactions across six representative cell types \n(GM12878, A549, HeLa-S3, IMR90, HepG2, and K562). Consistent with the GM12878 results, \nCTCF and Cohesin show the strongest pairwise correlations in all cell types, indicating a conserved \nstructural coupling between these two architectural proteins, whereas RNAPII -mediated maps \nexhibit lower similarity with either. (b) Genome-wide pairwise Pearson correlation distributions of \npredicted contact maps for each protein pair. The consistently higher CTCF -Cohesin correlation \ncompared to CTCF-RNAPII and Cohesin-RNAPII highlights that transcription-associated (RNAPII) \ninteractions form a distinct topolog ical regime from the architectural scaffolds of CTCF/Cohesin. \n(c) UMAP embedding of all predicted contact maps across the genome in each cell type. CTCF and \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nCohesin clusters show substantial overlap, whereas RNAPII maps form a clearly separated manifold, \nreinforcing the distinct 3D organization principles of structural versus transcriptional protein -\nmediated chromatin interactions. \n  \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\n \nExtended Data Fig. 9: Cross-cell-type comparison of protein-specific chromatin organization \npatterns. (a) Distance-stratified Pearson correlations between predicted contact maps for each pair \nof proteins (CTCF-Cohesin, CTCF-RNAPII, and Cohesin-RNAPII) across six representative cell \ntypes (GM12878, A549, HeLa -S3, IMR90, HepG2, and K562). Consistent across all c ontexts, \nCTCF and Cohesin maintain the highest correlation across distance scales, whereas correlations \ninvolving RNAPII decline sharply at long genom ic distances (>524 kb), indicating its distinct \ntranscription-associated interaction regime. (b) Loop distance distributions of the top 10,000 \npredicted loops for each protein across the six cell types. RNAPII -mediated loops are consistently \nshorter and more localized, whereas CTCF and Cohesin loops span broader genomic ranges, \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nreflecting their architectural scaffolding roles. (c) Venn diagrams showing overlap among the top \npredicted loops for CTCF, Cohesin, and RNAPII in each cell type. Across all cell types, CTCF and \nCohesin share substantial subsets of loops, while RNAPII loops largely occupy distinct regulatory \nneighborhoods, underscoring protein-specific topological organization principles conserved across \nthe genome. \n  \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\n \nExtended Data Fig. 10: Cross-cell-type network topology of protein-mediated chromatin loops. \n(a) Complementary cumulative degree distributions (CCDF, log -log scale) of loop interaction \nnetworks for CTCF -, Cohesin-, and RNAPII -mediated interactions across six representative cell \ntypes (GM12878, A549, HeLa -S3, IMR90, HepG2, and K562). All networks displ ay scale -free \ntopologies, with RNAPII showing slightly higher low-degree connectivity, consistent with its dense \ntranscription-associated loops. (b) Hub dominance curves showing the cumulative fraction of total \nloop edges explained by the top-ranked nodes (by degree). Across all cell types, CTCF and Cohesin \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nhubs dominate a larger proportion of total edges than RNAPII, reflecting their roles as structural \nanchors in chromatin organization. (c) k-core decomposition analysis of loop networks. The size of \nthe k-core decreases more gradually for CTCF and Cohesin compared with RNAPII, indicating that \narchitectural loops form more stable hierarchical cores, whereas RNAPII loops are more peripheral \nand transient. (d) Neighbor interconnectivity (ego -neighbor density) around loop hubs across \ndegrees. CTCF and Cohesin display higher local clustering around high -degree nodes, suggesting \ncooperative domain insulation and compartmentalization, while RNAPII exhibits weaker neighbor \ninterconnectivity, consistent with its dispersed regulatory contacts associated with transcript ional \nactivity. \n  \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\n \nExtended Data Fig. 11: Functional enrichment and regulatory classification of protein -\nmediated loops across multiple cell types.  (a) Functional enrichment of anchors from pan -cell-\ntype Cohesin - and RNAPII -mediated loops for contact domain boundaries, enhancers, and \npromoters, stratified by the number of loop interactions per anchor. Statistical significance was \nassessed using two-sided Fisher’s exact tests (Number of Cohesin interactions = 48,847, Number of \nRNAPII interactions = 46,093, * P<0.05, ** P<0.01, * **P<0.001; NS, not significant). (b) \nDistribution of loop categories mediated by CTCF, Cohesin, and RNAPII across six representative \ncell types (GM12878, H1, HeLa -S3, HepG2, IMR90, and K562). Loops were classified as \npromoter-promoter (PPI), enhancer -promoter (EPI), enhancer -enhancer (EEI), s ilencer-promoter \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\n(SPI), or other types based on genomic element annotations at loop anchors. (c) Differential gene \nexpression associated with genes linked to cell-type-specific loops mediated by CTCF, Cohesin, or \nRNAPII across the six cell types. Each violin represents the distribution of expression changes for \ngenes engaged in protein -mediated loops compared with genes not participating in promoter -\nenhancer loops (“No P-E”). Statistical significance was determined using two-sided Wilcoxon rank-\nsum tests. \n  \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\n \nExtended Data Fig. 12: Chromatin-state composition and enrichment of protein-specific loop \nanchors across additional cell types.  (a) Chromatin-state composition of cell -type-specific loop \nanchors mediated by CTCF, Cohesin, and RNAPII across six representative human cell types (H1, \nA549, HeLa-S3, HepG2, IMR90, and K562). Each pie chart shows the proportion of chromatin \nstates. (b) Fold enrichment of chromatin states at cell -type-specific loop anchors for CTCF -, \nCohesin-, and RNAPII -mediated loops in the same six cell types. Statis tical significance was \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nevaluated using two -sided Fisher ’s exact tests (* P<0.05, ** P<0.01, *** P<0.001; NS, not \nsignificant). \n  \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\n \nExtended Data Fig. 13: Functional coupling between loop strength, gene expression, and \nsuper-enhancer association across multiple cell types.  (a) Relationship between predicted loop \nstrength and the corresponding gene expression change for CTCF -, Cohesin -, and RNAPII -\nmediated loops. Each point represents a promoter-enhancer pair, with color indicating density. (b) \nFold enrichment of cell-type-specific versus non-specific loops in cell-type-specific or shared super-\nenhancers (SEs) across six representative c ell types (A549, H1, HeLa -S3, HepG2, IMR90, and \nK562). SE-specific loops were significantly enriched around cell -type-specific SEs, while shared \nSEs were more frequently linked by common loops. Statistical significance was assessed using two-\nsided Fisher’s exact tests (*P<0.05, **P<0.01, ***P<0.001; NS, not significant). \n  \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\n \nExtended Data Fig. 14: Overlap and canonical pathway enrichment of genes associated with \nprotein-mediated SE-linked loops across multiple cell types.  (a) Overlap of genes associated \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\nwith CTCF -, Cohesin-, and RNAPII -mediated loops across six representative human cell types \n(A549, H1, HeLa -S3, HepG2, IMR90, and K562). Venn diagrams depict the intersection among \ngenes linked to loops mediated by each protein, highlighting both shared and  protein-specific \nregulatory targets. (b) Canonical pathway enrichment analysis of genes associated with super -\nenhancer (SE)-linked loops mediated by CTCF, Cohesin, and RNAPII in the same six cell types. \nThe top enriched pathways are shown for each protein -specific SE-loop set as well as SE regions \nalone, revealing both common and protein-distinct biological processes. \n  \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\n \nExtended Data Fig. 15: Epigenomic and motif features of ecDNA -associated RNAPII \ninteractions across cancer cell types. (a) Heatmaps showing ChIP-seq signal densities of RNAPII, \nMED1 and H3K27ac centered on ecDNA -associated contact anchors in COLO320DM and \nGBM171DM cells. Results from two biological replicates are shown for each cell type \n(COLO320DM Rep-1, n = 23,585; Rep-2, n = 20,482; GBM171DM Rep-1, n = 1,937; Rep-2, n = \n2,155). Signals are displayed within ±5 kb of anchor centers. Color scales indicate normalized \nChIP-seq signal intensities. (b) Genome browser views of representative ecDNA-associated regions \nin PC3DM, COLO320DM and GBM171DM cells. Tracks show ChIA-drop interaction signals from \ntwo biological replicates, RNAPII, MED1 and H3K27ac ChIP -seq signals, and gene annotations . \n(c) Motif analysis of ChIANet -predicted RNAPII loop anchors in COLO320DM (top) and \nGBM171DM (bottom) cells. Each point represents a transcription factor motif. Point size denotes \nmotif enrichment significance (-log2 P value), and color intensity indicates RNA-seq-derived gene \nexpression levels ( raw expression counts) of the corresponding transcription factors. The y -axis \nshows gene expression levels. \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint \n\n \nExtended Data Fig. 16: Gene Ontology enrichment analysis of ChIANet -predicted loop \nanchors across cancer cell types. Gene Ontology (GO) biological process enrichment analysis of \nChIANet-predicted chromatin loop anchors in PC3DM (top), COLO320DM (middle) and \nGBM171DM (bottom) cells. Each point represents an enriched GO term. The x-axis shows -log10-\ntransformed binomial FDR Q values. Point color indicates binomial fold enrichment, and point size \nreflects the number of genes associated with each GO term. The red dashed line denotes the \nsignificance threshold used for enrichment analysis. \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 25, 2026. ; https://doi.org/10.64898/2026.02.24.707640doi: bioRxiv preprint","source_license":"CC-BY-4.0","license_restricted":false}