{"paper_id":"40b7be7d-6ee3-4dc5-9afe-659866185393","body_text":"1 \n \n \n \nHM-DyadCap – Capture and Mapping of 5-Hydroxymethyl-\ncytosine/5-Methylcytosine CpG Dyads in Mammalian DNA \nLena Engelharda§, Damian Schillera§, Marlon S. Zambrano-Milaa,b, Kotryna Keliuotytea, Benjamin Buchmullera, Shashank \nTiwarib, Jochen Imigb, Angela Simeonec, Christian Schröterb, Sidney Beckera,b* and Daniel Summerera* \naFaculty of Chemistry and Chemical Biology, TU Dortmund University, Otto -Hahn-Str. 4a, 44227 Dortmund (Germany). \n*Email: daniel.summerer@tu-dortmund.de and sidney.becker@mpi-dortmund.mpg.de \nbMax-Planck-Institute for Molecular Physiology, Otto-Hahn-Str. 11, 44227 Dortmund (Germany) \ncGenomic Services, Qiagen, Manchester, UK. \n§these authors contributed equally to this work. \n \nABSTRACT: 5-Methylcytosine (mC) and 5 -hydroxymethylcytosine (hmC)  are the main epigenetic modifications of \nmammalian DNA, and play crucial roles in cell differentiation, development, and tumorigenesis. Both modifications co-exist \nwith unmodified cytosine in palindromic CpG dyads in different symmetric and asymmetric combinations across the two DNA \nstrands, each having unique regulatory potential. To facilitate investigating the individual functions of such dyad modifications, \nwe report HM-DyadCap. This method employs an evolved methyl-CpG-binding domain (MECP2 HM) for the direct capture \nand sequencing of DNA fragments containing the CpG dyad hmC/mC. Binding studies reveal a high discrimination of MECP2 \nHM against off-target dinucleotides. We conduct comparative mapping experiments for mESC genomes with HM-DyadCap, \nstandard MethylCap employing wild type MECP2, as well as MeDIP and hMeDIP protocols . We find that MECP2 HM is \nblocked by hmC glucosylation , and conduct control enrichments with glucosylated genomes that indicate highly selective \nenrichment of hmC/mC dyads by MECP2 HM. Metagene profiles correlate hmC/mC marks with actively transcribed genes, \nand reveal global enrichment in gene bodies as well as depletion at transcription start sites. We anticipate that HM-DyadCap \nwill enable effective enrichment and mapping of hmC/mC marks with broad applicability for unravelling the function of this \ndyad in chromatin biology and cancer. \n \nINTRODUCTION \nThe epigenetic DNA modification 5 -methylcytosine (mC,  \nFig. 1a) is a central regulator of mammalian gene expression \nwith crucial roles in development, differentiation, and cancer \nformation1. mC is written and maintained over cell cycles by \nDNA methyltransferases (DNMTs) mainly within \npalindromic CpG dyads, of which 60 -80 % are methylated \nin somatic cells 2. mC can be read and converted into \ntranscriptionally repressive states by methyl -CpG-binding \ndomain (MBD) proteins 3, but can also recruit or repel \ntranscription factors and other chromatin proteins4. mC can \nfurther be oxidized by ten -eleven translocation (TET) \ndioxygenases, which generate CpG dyads containing 5 -\nhydroxymethylcytosine (hmC , Fig. 1a ), 5 -formylcytosine \n(fC) and 5-carboxylcytosine (caC). Whereas fC and caC are \nsubstrates of the thymine-DNA glycosylase (TDG)-initiated \nbase excision repair (BER) pathway leading to an active \ndemethylation5,6, hmC exhibits high stability, and occurs at \nhigh levels in embryonic stem cells (ESC) and neurons5-7. \nhmC thereby differs from mC in its physicochemical \nproperties, genomic distributions, dynamic changes during \ndevelopment, and protein interactions, and thus has potential \nto uniquely regulate chromatin-associated processes5,8,9. \nFigure 1. Mammalian cytosine modifications in the double -stranded CpG \ndyad. a) Structures of C, mC, hmC  (fC and caC not shown) . b) Both \ncytosines (Ca, Cb) in the CpG dyad can exist as C, mC or hmC. c) Possible \npathways for the creation of different hmC dyad  symmetries (hmC/mC \ndyad targeted in this study in grey box) by DNMT s, TETs, replication \n(Rep.) and BER.  \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted October 31, 2025. ; https://doi.org/10.1101/2025.10.29.685270doi: bioRxiv preprint \n\n2 \n \n \n \nFor example, whereas hmC is enriched in promoters and \ngene bodies in m ouse ESC (mESC) and neurons 10-12, it \nshows generally low levels and different genomic \ndistributions in cancer cells, making it an important cancer \nbiomarker13,14.  \nIn contrast to mC that can be kept in a strand-symmetric state \nin CpGs by maintenance DNMTs, the writing of hmC and \nother oxi-mCs by TETs occurs non -processively15, leading \nto different strand -symmetric and -asymmetric \ncombinations of C, mC and hmC in the CpG dyad (Fig. 1b). \nThe genomic landscape of these marks can further be shaped \nby BER, the inhibition of maintenance methylation 5, and \nTET-catalyzed oxidation of hemimethylated mC/C dyads 16 \n(Fig. 1c ). Each of the aforementioned dyads presents a \nphysicochemically unique signal in the DNA major \ngroove17, an important interaction surface for DNA binding \nproteins8. The question of how hmC has unique regulatory \nfunctions can thus only be answered conclusively by \nconsidering the hmC-modification symmetry of CpG dyads.  \nSequencing and mapping studies of hmC ha ve greatly \ncontributed to a better understanding of its functions, but \nestablished mapping methods do not simultaneously resolve \nC, mC and hmC in the same DNA duplexes 5,18. Very \nrecently, new strategies for the simultaneous sequencing of \nC, mC and hmC have been reported, offering potential for \nrefined maps with resolution for individual hmC dyad \nsymmetries. These employ protocols based on  restriction \nenzymes (DARESOME 19 and Dyad -seq20), on multiple \nconsecutive nucleobase conversions to achieve nucleotide \nresolution (e.g., EnIGMA 21, SCoTCH -Seq22,23, and \nSIMPLE-Seq24), or on direct  nanopore sequencing25. In two \ncases, strategies have been adapted/applied for mapping  \nindividual hmC  dyads. This revealed that in  mESC and \nmouse cerebellum genomes , the large majority of hmC \nresides in asymmetric hmC/mC dyads, whereas hmC/C and \nsymmetric hmC/hmC dyads are comparably rare 20,25. A  \nthird, very  recent mESC genome mapping study using \nSCoTCH-seq reported frequencies in the order \nhmC/C>hmC/mC>>hmC/hmC23. \nA deeper understanding of the individual hmC dyad´s roles \nin chromatin regulation during development and disease  \nrequires broader mapping studies across diverse tissues, \nwhich however are complicated by the low hmC levels \nfound in most tissues (e.g., cancer tissues) 6,7. Methods for \nthe effective enrichment of hmC -modified DNA could \nresolve this bottleneck by direct sequencing/mapping, or in \ncombination with aforementioned conversion -based \nsequencing. However,  current hmC-enrichment strategies \nrely on anti-hmC antibodies26-28 or T4-β-glucosyltransferase \n(T4 BGT) -catalyzed azidoglucose transfer to hmC  and \ncovalent capture by click chemistry29,30, both of which have \nnot been reported to be  dyad-specific31 (e.g., T4 BGT \neffectively glucosylates hmC in different dyad contexts, Fig. \nS1). \nTo address this bottleneck, we developed HM -DyadCap, a \nmethod that employs the evolved MECP2 variant MECP2 \nHM for selective enrichment and sequencing of hmC/mC \nCpG dyads  that are abundant marks  mESC and brain \ngenomes. In vitro binding studies show that MECP2 HM has \na high  specificity for its target dyad . We conduct \ncomparative mapping studies  with HM-DyadCap, \nMethylCap based on MECP2 wild type (MECP2 wt), as well \nas antibody-based methyl- and hydroxymethylcytosine -\nDNA-immunoprecipitation protocols ( MeDIP and \nhMeDIP), and observe method-dependent preferences in the \nenrichment of specific genomic features.  We discover a \nsensitivity of MECP2 HM to hmC glucosylation and   \nintroduce additional mapping controls with glucosylated \ngenomes that indicate a high hmC/mC selectivity of HM-\nDyadCap. Metagene profiles reveal enrichment of hmC/mC \nin gene bodies and depletion at transcription start sites, and \ncorrelate it with active transcription. HM-DyadCap offers a \nsimple and effective approach for enriching and mapping \nhmC/mC dyads, and we anticipate broad applica bility for \ncancer biomarker discovery and unravelling the function of \nhmC/mC in chromatin regulation.64 \n \nMATERIAL AND METHODS \nmESC culturing and gDNA isolation.  E14tg2a mouse \nembryonic stem cells32 were grown in dishes pre-coated with \n0.1 % gelatin (w/vol, Sigma) in GMEM supplemented with \n10% fetal bovine serum, sodium pyruvate, 50 μM β -\nmercaptoethanol, glutamax, non -essential amino acids (all \nfrom Gibco/ThermoFisher) and 10 ng/ml murine leukemia \ninhibitory factor (LIF, Protein Expression Facility, MPI \nDortmund). Cells were passaged every 2-3 days to maintain \ncultures between 10% and 90% confluency. For analysis, \nnear-confluent cultures were released from culture vessels \nwith trypsin, spun down, and snap-frozen in liquid nitrogen \nbefore further processing. gDNA was isolated using the \nMonarch Genomic DNA Purification Kit (NEB, T3010) \naccording to the manufacturer’s instructions. DNA \nconcentration and purity were assessed using Nanodrop. \ngDNA fragmentation. gDNA at a concentration of 10 ng/µl \nin 1xTE was sheared to an average fragment size of 200 bp \nusing a Bioruptor Pico sonication device (Diagenode). \nFragmentation was performed for 19 cycles of 30 s on, 30 s \noff at 4 °C. Sheared DNA was purified by ethanol  \nprecipitation overnight at -80 °C. dsDNA concentration was \nmeasured with a Quantus Fluorometer (Promega). Fragment \nsize distribution was confirmed using a TapeStation (Agilent \nTechnologies) using D1000 ScreenTapes.  Sheared DNA \nwas eith er subjected to glucosylation or directly used for \nIllumina library preparation.  \nGlucosylation of 5hmC. For each glucosylation reaction, 1 \nµg of sheared gDNA was used. Reactions were carried out \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted October 31, 2025. ; https://doi.org/10.1101/2025.10.29.685270doi: bioRxiv preprint \n\n3 \n \n \n \nin 1x Cutsmart buffer (NEB) supplemented with 40 µM \nUDP-Glucose and 20 U T4 -BGT (NEB, M0357) in a total \nvolume of 40 µl for 16 h at 37 °C. Subsequently, DNA was \npurified using the Monarch PCR & DNA Purification kit \nfollowing the manufacturers protocol for fragments < 2 kbp. \nDNA concentration was measured with a Quantus \nFluorometer (Promega). For glucosylated hmC binding \nstudies, 500 ng of 24-mer duplex DNA (hmC/mC: o2909 + \no3115; hmC/hmC: o2909 + o3115, Table S1) was incubated \nwith 10 U T4 -BGT (NEB) under conditions described \nabove. Subsequently, DNA was purified using the Oligo \nClean and Concentrator kit (Zymo Research) following the \nmanufacturers protocol. \nEnd repair and adapter ligation.  1 µg fragmented DNA, \neither glucosylated or non -glucosylated, was subjected to \nend repair and adapter ligation using the NEBNext Ultra II \nDNA Library Preparation Kit (NEB, E7103 ; version \n6.1_5/20). DNA purification and size selection was carried \nout using NEBNext Sample Purification Beads (NEB, \nE7103) according to the protocol with bead volumes \ncorresponding to a target fragment size of 200 bp. Elution \nwas carried out according to the manufacturers protocol, and \nDNA concentration was determined with the Quantus \nFluorometer (Promega). Fragment size distribution was \nverified using a TapeStation (Agilent Technologies) with \nD1000 ScreenTapes.  \nSpike-in probe preparation. Preparation of spike-in probes \nwas carried out as described  previously33. In brief, DNA \nduplexes carrying modified CpG dyads (“carriers”) were \nligated to adapter sequences containing general and unique \nprimer binding sites for qPCR quantitation. Carriers were \nprepared by annealing modified oligonucleotides (C/C: \no4728 + o4729, mC/mC: o4677 + o4727, hmC/mC: o4675 \n+ o4727) in 30 mM HEPES pH 7.5 and 100 mM KOAc. \nAnnealing was carried out by heating to 95 °C for 5 min \nfollowed by gradual cooling in a Dewar flask containing \nboiling water. Adapters were annealed at a final \nconcentration of 2.5 µM in the same way (o4371 + o4372, \no4373 + o4374, and o4392 + o4393) and subsequently \nextended with 25 mU/µl Klenow fragment (NEB) in the \npresence of 0.1 mM dNTPs for 20 min at 37 °C. To introduce \n5’ phosphorylation, primer o4123 was annealed to  the \nextended adapters and further extended with Klenow \nfragment for 30 min at 37 °C. 50 µl of crude adapters were \nligated with15 pmol of carriers at 16 °C overnight using 200 \nU T4 DNA ligase (NEB). Ligated spike-in probes were \npurified using the Macherey -Nagel PCR purification kit, \nwith the NTI buffer diluted to 16 % (v/v) with ddH2O. Probe \nconcentration was determined by qPCR against reference \nstandards (o4628 for o4374, o4627 for o4372 and o4629 for \no4393). qPCR primers o4368 and o4124 were premixed as a \n2 µM stock. Each reaction contained a 3 µl primer mix, 5 µl \n2x primaQuant SYBR Green Master Mix with ROX \n(Steinbrenner) and 2 µl of spike -in probe. Cycling \nconditions were 95 °C for 2 min, followed by 50 cycles of \n95 °C for 10 s and 60 °C for 30 s.  \nMECP2 wt and MECP2 HM expression and \npurification. GST-tagged MBP -MECP2 wt and MBP -\nMECP2 HM were recombinantly expressed in E. coli BL21-\nGold (DE3). A 300 ml LB culture supplemented with 1 mM \nMgCl2, 1 mM ZnSO4 and a suitable antibiotic was inoculated \nfrom a fresh overnight culture and grown at 37 °C with \nshaking at 220 rpm to an OD 600 of 0.5–0. 6. Cultures were \nchilled on ice and protein expression was induced by the \naddition of 1 mM IPTG. The expression proceeded \novernight at 30 °C with shaking at 150 rpm. The cultures \nwere harvested at 8,000 x g for 20 min at 4 °C, washed twice \nby resuspension in 50 ml cold 20 mM Tris -HCl (pH 8.0), \nand the resulting pellet was frozen until further processing. \nFor lysis, cell pellets were resuspended in 20 ml binding \nbuffer (20 mM Tris-HCl, 250 mM NaCl, 10% glycerol, 10 \nmM dithiothreitol (DTT), 5 mM imidazole, 0.1% Triton X-\n100, pH 8.0) supplemented with 1 mM PMSF. The \nsuspension was treated with 0.1 mg/ml lysozyme (Merck) \nand 1 U/ml DNase I (NEB) and incubated overnight at 4 °C. \nCells were disrupted by two rounds of sonication on ice (3 \nmin per run of alternating a 4 s ultrasonic wave pulse at 20 \n% amplitude and 30 s rest (Branson Digital Sonifier 450 Cell \nDisruptor). Cellular debris was removed by centrifugation at \n14,000 x g for 20 min at 4 °C and the cleared supernatant \nwas retained and purified over a Ni -NTA column (GE \nHealthcare) on an Äkta Purifier 10 FPLC system (GE \nHealthcare) using a gradient of imidazole (10 mM to 500 \nmM) in binding buffer. The fractions that contained the pure \nprotein were combined and dialyzed three times against \ndialysis buffer (20 mM HEPES, 100 mM NaCl, 10 % \nglycerol, adjusted to pH = 7.3, and 0.1 % Triton X -100) \nusing Slide -A-Lyzer dialysis cassettes (3.5 kDa MWCO, \nThermoFisher Scientific). The protein concentration was \ndetermined in triplicate using the Pierce BCA Protein Assay \n(ThermoFisher Scientific). Proteins were snap frozen in \nliquid nitrogen and stored at –80 °C at a concentration of 15 \nµM. \nElectrophoretic mobility shift assays . Complementary \noligonucleotide pairs (Table S1) were annealed by mixing \n1.5 μM of the FAM -labeled strand and 2.5 μM of the \nunlabeled strand in EMSA hybridization buffer (20 mM \nHEPES, 30 mM KCl, 1 mM EDTA, 1 mM (NH 4)2SO4, pH \n7.3), incubated at 95 °C for 5 min, and gradually cooled \ndown to RT in a Dewar filled with boiling water. The non -\nspecific dA:dT competitor duplex was prepared by \nannealing 24 -mer poly(A) with poly(T) oligos (o2968 + \no2969, Table S1) at an equimolar ratio of 40 μM. Prior to \nthe binding assay, 15 μM of purified recombinant proteins \nwere incubated with 0.25 μM TEV protease at 4 °C \novernight to remove the MBP tag. For glucosylated hmC \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted October 31, 2025. ; https://doi.org/10.1101/2025.10.29.685270doi: bioRxiv preprint \n\n4 \n \n \n \nbinding studies, varying concentrations of protein were \nincubated with 2 nM labeled duplex DNA in EMSA buffer \n(20 mM HEPES, 30 mM KCl, 1 mM EDTA, 1 mM \n(NH4)2SO4, 0.2 % Tween-20, pH 7.3) in presence of 3.4 µM \npoly(A):poly(T) competitor and 1 mM DTT in a total \nvolume of 15 µL. Reactions were incubated for 20 min at 21 \n°C. For assays using DNA probes  with CpG in random \nsequence contexts, 100 nM of protein was incubated with \n2 nM labeled probe under the same binding conditions. For \nantibody binding assays, 750 pM of labeled DNA probes \nwere incubated with either polyclonal anti -5hmC antibody \n(Active Motif, 39769; 1:100 and 1:200 dilutions, Fig. 4b-c), \nmonoclonal anti -5hmC antibodies (EpiGentek  (EG), A -\n1018; Cell Signaling  (CS), 51660; Diagenode  (DG), \nC15200200; each at 1:10 dilutions , Fig. 4a ), or 100 nM \nMECP2 HM protein (HM in Fig. 4a) in binding buffer (10 \nmM Tris, 100 mM KCl, 1 mM EDTA, 5 % glycerol, 0.1 \nmg/ml BSA, pH 7.5) containing 3.4 µM poly(A):poly(T) \ncompetitor in a total volume of 15 µL and incubated at 4 °C \novernight. After incubation, samples were mixed with 3 µL \nof 6x EMSA loading buffer (1.5x TBE, 40 % glycerol) and \nresolved on pre -run 15 % non -denaturing polyacrylamide \ngels. Electrophoresis was performed at 240 V for 45 min \n(MECP2 binding  assays) or 55 min (antibody binding \nassays) at 4 °C usin g Mini -PROTEAN vertical \nelectrophoresis units (Bio -Rad). Gel  fluorescence was \nrecorded using the 510 LP filter of a Typhoon FLA9500 \nlaser scanner (GE Healthcare) at 473 nm at 800 - 1000 V \namplification. The fraction of bound dsDNA was \ndetermined using ImageQuant TL v8.1 1D Gel Analysis (DE \nHealthcare), applying background subtraction and manual \npeak detection with approximately equal peak areas across \nall lanes. \nMass spectrometry. 20 µl of 0.1 µM DNA sample was \nanalyzed per injection in the Vanquish HPLC system \nconnected to an Orbitrap Exploris 120 ESI -MS (Thermo \nFisher Scientific). The sample components are separated \nwith a flow rate of 0.2 ml/min in a solvent system of 50 mM \nHFIP (1,1,1,3,3,3-Hexafluoro-2-propanol, 15 mM \ntriethylamine, pH = 9.0 (solvent A), and methanol (solvent \nB) at 80 °C on DNAPac RP 4 µm column (2.1 × 100 mm, \nThermo Scientific). A multistep gradient (3 % of solvent B \nin 0–2 min, followed by an increase from 3 to 10 % until 5 \nmin and 10 to 20  % until 20 min) was used for separation. \nMS analysis was performed in the negative mode with a \nspray voltage of 2500 V, flow of sheath gas, aux gas, and \nsweep gas maintained at 35, 7, and 0, respectively, and ion \ntransfer tube temperature and vaporizer tem perature set at \n300 °C and 275 °C, respectively. For MS1 analysis, the \nOrbitrap resolution of 120,000 was used, the scan range was \nset from 700 to 3000 m/z, RF lens at 70 %, with standard \nnormalized AGC target and data type set to profile mode. All \nMS data analysis was performed with Biopharma Finder \nv.05.1, automatic parameter values were set for component \ndetection. For oligonucleotide identification, output mass \nrange was set from 2 kDa to 12 kDa, a charge range of 3-20 \nwere considered with a minimum detected charge of 3. Of \nthe identified fragments, only those with a mass error less \nthan 10 ppm were considered for further analysis. \nGenomic enrichment of modified CpG  dyads. For \nenrichment, an equimolar mix of spike-in probes (6.04 fmol \neach) was added to 250 ng of sheared gDNA in Buffer B \n(from MethylCap Kit, Diagenode) in a total volume of 35.45 \nµl. From this mixture, 5.7 µl were set aside as input and the \nremaining 29.75 µl were used for the enrichment reaction. \nGST-tagged MBP-MECP2 wt and MBP-MECP2 HM were \nTEV-digested for 30 min at RT to remove MBP. TEV -\ndigested proteins were then added to the DNA mixture at a \nfinal concentration of 1.7 µM  in a total reaction volume of \n34 µl . Enrichment was performed according to the \nmanufacturer’s instruction s (MethylCap Kit Diagenode, \nversion 6, C02020010) following the high -salt elution \nprotocol. MECP2 wt and MECP2 HM enrichments were \ncarried out in separate reactions, each performed in three \ntechnical replicates per biological experiment. Eluted DNA \nwas purified using the Macherey Nagel PCR purification kit. \nSpike-in probe recovery was determined by qPCR against a \nstandard dilution series of spike -in probes. To distinguish \nindividual spike-ins, three different primer pairs were used \n(C/C: o4374+o4368, mC/mC:  o4372+o4368, hmC/mC: \no4393+o4368). qPCR was performed on the elution \nfractions. Reactions were carried out by initial denaturation \nat 95 °C for 2 min, followed by 50 cycles of 95 °C for 10 s \nand 60 °C for 30 s. The remaining elution samples were \nfurther processed and prepared for NGS. \nPreparation of Illumina libraries and NGS sequencing. \nEnriched DNA fragments were PCR amplified using the \nNEBNext Ultra II DNA Library Prep Kit for Illumina (NEB, \nversion 6.1_5/20) using NEBNext Multiplex Oligos (NEB, \nE73359). As amplification does not occur directly after \nadapter ligation but after the enric hment, DNA input \nconcentrations were treated as threefold higher for \ndetermining the appropriate number of PCR cycles. PCR \nclean-up was performed according to the “Cleanup of PCR \nReaction” protocol (version 6.1_5/20) using NEBNext \nSample Purification Beads. Fragment length distribution of \nthe final libraries was assessed with a Tapestation (Agilent \nTechnologies) with a D1000 ScreenTape according to the \nAgilent protocol. dsDNA concentration was measured with \nthe Quantus Fluorometer (Promega). Libraries were pooled \nin equimolar ratios and sequenced on a NovaSeq X -25B \ninstrument (Illumina) using paired-end reads, with a target \ndepth of 50 million reads per sample.  \nMeDIP and hMeDIP. For each immunoprecipitation, 1 µg \nof sheared gDNA was used as input. DIPs were performed \naccording to the MagMeDIP -seq Package V2 protocol \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted October 31, 2025. ; https://doi.org/10.1101/2025.10.29.685270doi: bioRxiv preprint \n\n5 \n \n \n \n(Diagenode, C02010041), using either the provided anti -\n5mC antibody for MeDIP or 0.6 µg per IP of the mouse \nmonoclonal anti -hmC antibody (Diagenode, C15200200) \nfor hMeDIP. Additionally, a control IP with 0.6 µg of mouse \nIgG (Diagenode, C15400001) was inclu ded. Illumina \nlibraries were prepared according to the Diagenode MeDIP-\nseq library preparation protocol. Final library concentrations \nwere quantified using the Quantus Fluorometer (Promega) \nand fragment size distributions were assessed with a \nTapeStation ( Agilent Technologies) using a D1000 \nScreenTape. Libraries were sequenced on a n Illumina \nNovaSeq X -25B instrument using paired -end 2x 150 bp \nreads with a target depth of 50 million reads per sample. \nData Analysis \nPaired-end sequencing libraries were processed using a \nstandardized and automated workflow implemented with \nSnakemake34. Initial quality assessment of raw reads was \ncarried out using FastQC 35, followed by adapter trimming \nand removal of low-quality bases using Trim Galore36. The \ncleaned reads were then aligned to the Mus musculus  \nreference genome mm10 using Bowtie2 37.  Retaining only \nproperly paired reads and removing PCR duplicates was \ndone using Samtools 38. To identify enriched regions, peak \ncalling was performed on deduplicated BAM files using \nMACS239. Technical reproducibility was assessed by \nconducting peak calling independently for each of the three \ntechnical replicates along with their respective input controls \nwithin each biological replicate. Reproducibility was \nevaluated across three independent replicates by generating \na consensus peak set, retaining peaks present in at least two \nout of the three replicates using bedtools 40. In the case of \nDIP-seq, enriched regions were obtained as previously \ndescribed using MACS2, with either IgG or input from \nmESCs serving as controls. True positive regions were \ndefined as those enriched regions identified for both IgG and \ninput controls, as described by Lentini et al., 2018. For signal \nvisualization, normalized BigWig files were generated using \nbamCoverage (deepTools v3.5.0 41), and final alignments \nwere visually inspected in the Integrative Genomics Viewer \n(IGV42). To assess the statistical significance of overlap \nbetween consensus peaks and annotated genomic features, \nthe Genomic Association Tester (GAT 43) was employed. \nMetagene profiles were generated to assess enrichment \nrelative to gene expression. Log₂ fold changes over input \nwere calculated across gene bodies using bamCompare \n(deepTools). Genes were ranked by expression, defined as \nthe geometric mean of FPKM from RNA -seq (E14 \nENCODE replicates 1 and 2), and grouped into highly \nexpressed (top 10%), lowly expressed (FPKM > 0.1 and \nbelow the top 10% threshold), and silenced (FPKM < 0.1). \nProfiles were computed with computeMatrix (deepTools) \nand visualized using plotHeatmap (deepTools). Additional \ndownstream analyses were performed  using bedtools, \ndeepTools, and custom Bash and R scripts. \nRESULTS AND DISCUSSION \nSelectivity of MECP2 wt and MECP2 HM in respect to \nCpG dyad-specific genomic enrichment \nEnrichment of genomic DNA fragments bearing specific \nepigenetic modifications greatly increase the efficiency of \nmapping studies by overcoming the need for whole genome \nsequencing.  For hmC, antibodies and T4 BGT are both \nbroadly used enrichment tools regardless of the specific CpG \ndyad symmetry. While MBD proteins are also popular \nprobes for genomic enrichment, they specifically recognize \nsymmetrically methylated (mC/mC) CpG dyads in native \ndsDNA, and are therefore limited to mC enrichment \nassays44. Indeed, the functional members of mammalian core \nfamily MBD proteins (MBD1, 2, 4 and MECP2) tend to be \nrepelled by hmC 45-49. To expand the scope of MBD -based \ntechnologies, we recently re-evolved the MBD domain of \nhuman MECP2 wt to switch its selectivity from the \ncanonical mC/mC CpG to TET-generated CpG dyads33,50. In \nthis course, we identified the mutant MBD domain MECP2 \nHM that selectively recognizes the dyad hmC/mC \n(mutations K109T, V122A, S134N 51). In electromobility \nshift assays (EMSA) , this mutant showed a similar on - \nversus main off-target selectivity (hmC/mC over mC/mC) as \nMECP2 wt (mC/mC over hmC/mC). Moreover, it exhibited \na higher discrimination as MECP2 wt against all other CpG \ndyads containing C, mC or hmC , particularly against the \nfrequent dyad modifications mC/C and hmC/C 23 (Fig. 2a, \ndata from ref 33). Importantly, MECP2 HM is based on a \nminimal MBD domain of MECP2 (aa 90-181 of the 486 aa \nfull-length protein) that does not contain AT -hook \nsequences, which have been shown to cause a weak \npreference for AT-rich sequence contexts around the target \nCpG52,53. Context-dependence has also been described for \nantibody-based DIP -Seq protocols that suffer from \nsignificant nonspecific enrichment of short repeat sequences \nvia IgG 54-56, a property that has not been observed in \nprevious enrichments using the MBD of MECP2 wt56.  \nTo further characterize the selectivity of MECP2 HM  in \nrespect to relevant genomic off-targets, we initially tested its \ninteraction with mCpA dinucleotides by EMSA. mCpA \noccurs at significant levels in neuronal and embryonic stem \ncells57,58, and it is bound by MECP2 wt (aa 1 -205)59. \nInterestingly, we observed significantly lower binding of \nMECP2 HM to both methylated and non-methylated CpA as \ncompared to MECP2 wt when present in an oligo dA/dT \nsequence context (Fig. 2b). Encouraged by this finding, we \nconducted broader EMSA studies covering other possible \noff-target dinucleotides. MECP2 wt has been shown to bind \nmCpA with a strong preference for a 3´ -A nucleotide \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted October 31, 2025. ; https://doi.org/10.1101/2025.10.29.685270doi: bioRxiv preprint \n\n6 \n \n \n \n(mCAA; a property not observed for CpG dyads60). To rule \nout such possible context dependencies, we employed \nrandom dsDNA probes in EMSA (5´ -…NNNXYNNN…-\n3´). Finally, in order to visualize  binding to even very low \naffinity off -targets, we  applied a high, 50 -fold excess of \nMECP2 over DNA. As expected, we observed mCpA off -\ntarget binding for both proteins under these  forcing \nconditions (Fig. 2c and S2). Hemi-methylated mCpG dyads \nwere bound with slightly lower and mCpT as well as mCpC \nwith much lower relative affinity, in agreement with \nprevious studies 60. Interestingly, MECP2 HM showed \nhigher selectivity over off-targets in all cases ( Fig. 2c and \nS2). We next tested nonmethylated CpN dinucleotides and \nobserved a generally much lower binding by both proteins, \nwith a different selectivity profile as observed for mCpN \n(Fig. 2d  and S2). Strikingly, CpN off -targets were again  \nbound weaker by  the HM mutant as compared to  the wt  \nprotein. For a comprehensive analysis of the so far \nuncharacterized dinucleotide selectivity of our evolved \nMECP2 HM, we finally tested probes containing all \nremaining, nonmethylated NpN dinucleotides.  \nFigure 2.  Characterization of off-target dinucleotide selectivity of \nMECP2 wt and HM . a) Selectivity of MECP2 wt and HM for indicated \nCpG modification symmetries. Shown are off - to on-target KD ratios from \nEMSA using dsDNA probes with a single CpG in an oligo dA/dT context \n(data from ref 33, color code as in Fig. 1a). b) EMSA with MECP2 wt and \nHM at indicated concentrations and 2 nM dsDNA probes with single CpA \nor mCpA in an oligo dA/dT context (arrow: bound probe). c -e) \nQuantification of EMSA with MECP2 wt or HM and dsDNA probes with \nindicated methylated or unmethylated dinucleotides in a random sequence \ncontext (*hemi-methylated CpG) . Shown are the % fractions of protein -\nbound probe normalized to the respective on-target of MECP2 wt (mC/mC) \nand HM (hmC/mC). Error bars show standard deviatio n from n=2.  \nWe observed a similar trend with particularly low relative \noff-target binding for MECP2 HM, corroborating the CpG-\nselectivity of this protein ( Fig. 2e and S2). In the light that \nMECP2 wt is a widely used enrichment probe for mC/mC \nCpGs, the observed high selectivity of MECP2 HM \nunderlines its suitability for enrichments of genomic \nhmC/mC CpG dyads with little potential for off -target \nenrichment.  \nWe next established a simple protocol for enriching \nhmC/mC-modified DNA fragments from mammalian \ngenomes (Fig. 3a). Briefly, gDNA is sheared via sonication \nand subjected to end -repair and adaptor ligation  for high \nthroughput sequencing  (Fig. 3a  and S3-6). After \npurification, three spike -in control DNAs containing a \nsequence with four CpG dyads (either unmodified, mC/mC- \nor hmC/mC-modified) are added in equimolar amounts.  \n  \nFigure 3. Enrichment and sequencing of modified CpG from mESC \ngenomes with MECP2 wt and HM.  a) Workflow of HM-DyadCap. b) \nRecovery of spike -in controls from genomic enrichments. c)  Spearman \ncorrelation of triplicate enrichment conditions . d) Venn diagrams of peaks \nobtained from enrichments with MECP2 wt and HM. e) Enriched or \ndepleted genomic features from enrichments with MECP2 wt and HM.  \n \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted October 31, 2025. ; https://doi.org/10.1101/2025.10.29.685270doi: bioRxiv preprint \n\n7 \n \n \n \nEach of these controls contains unique primer binding sites \nfor selective quantification by qPCR, but lack adaptors to not \ninterfere with downstream sequencing. For each enrichment, \n250 ng gDNA are incubated with MECP2 HM -GST fusion \nprotein immobilized on glutathione magnetic beads, and the \nDNA is eluted after washing. The samples are subjected to \nadaptor PCR with sample barcoding and then  purified and \nsequenced (Fig. 3a). Fragment size distribution is controlled \nafter adapter ligation and the final PCR, and enrichm ent \nsuccess is assessed by qPCR -measurement of mC/mC and \nhmC/mC spike-in recoveries normalized to the unmodified \nspike-in control directly after bead enrichment (Fig. 3b). To \ngenerate enrichment-based maps of hmC/mC CpG dyads in \na mammalian genome, we conducted enrichments with \nMECP2 HM and MECP2 wt  using mESC genomic DNA \n(E14tg2a cell line), a system where epigenetic regulation is \nparticularly essential. \nUsing spike -in oligonucleotides, we assessed the dyad \nselectivity of both MECP2 probes under enrichment \nconditions. Recovery of spike-ins showed low specificity for \nMECP2 wt (i.e., enrichment of both mC/mC and hmC/mC), \nwhereas MECP2 HM selectively enriched its hmC/mC \ntarget (Fig. 3b )33. Sequencing of these  libraries produced \nbetween ~63 million and 184 million reads (Table S2), with \nover 95% aligning to the mm10 mouse reference genome. \nTo assess the enrichment efficiency, we compared the \npercentage of reads within MACS2 -called peaks between \ninput and captured samples across MECP2 conditions (Fig. \nS7). In both cases, captured samples  showed higher \nenrichment than their respective inputs, with the stronge r \neffect observed for MECP2-HM. These results confirm the \nrobust enrichment achieved with our protocol.  \nTo evaluate reproducibility and overall similarity across \ndatasets, we calculated Spearman correlations across 100 kb \ngenomic bins and performed hierarchical clustering ( Fig. \n3c). Each of the two enrichment conditions was analyzed in \ntriplicates, and the replicates showed high correlations (r = \n1.00 for MECP2  wt and r ≥ 0.99 for MECP2  HM), \ndemonstrating the robustness of each condition. As \nexpected, the input libraries for MECP2 HM and MECP2 wt \nclustered together, but exhibited lower correlations with the \nenriched libraries (r ≈ 0.84  - 0.89). Additionally, heatmaps \nand average profiles showed sharp, reproducible enrichment \nof MECP2 wt and MECP2 HM at peak centers relative to \ncorresponding input ( Fig. S 8). Collectively, these results \nconfirm that MECP2 enrichment is robust, reproducible \nacross replicates, and clearly distinguishes the different \nexperimental conditions . Peak calling identified 91,481 \npeaks for MECP2  wt and 149,656 peaks for MECP2  HM \n(present in at least 2 of 3 technical replicates). Constructing \na consensus peak universe revealed that 24% of peaks were \nunique to MECP2 wt, 54% were unique to MECP2 HM, and \n22% were shared between the two conditions (Fig. 3d). We \nused GAT to assess peak overlap with genomic features \nrelative to background (workspace).  \nFigure 4. Comparison of MECP2 wt and HM maps with MeDIP and \nhMeDIP maps. a) EMSA analysis of monoc lonal anti-hmC antibodies \n(1:10 dilution) binding to  dsDNA oligos (0.75 pM) bearing a single \nhmC/mC dyad  (for data with hmC -modified ssDNA, see  Fig. S9 ). HM: \nMECP2 HM positive control, EG, DG, CS: antibody supplier (see  \nmethods). b) EMSA analysis of  polyclonal anti-hmC antibody  binding to \ndsDNA oligo (0.75 pM) containing the indicated CpG dyad modifications. \nAntibody dilutions on top. c) Bar diagram of quantification of EMSA data \nfrom Fig. 4b, error bars show standard deviation (n=2). d) Venn diagram \nfor peaks from MeDIP and MECP2 wt enrichments.  e) \nEnrichment/depletion of CpG islands for selected peak sets. f)  Venn \ndiagram for peaks from h MeDIP and MECP2 HM enrichments.  g-h) \nEnrichment/depletion of CpG islands and TTS for selected peak sets.  \n \nBoth MECP2 wt- and MECP2 HM-associated peaks were \nsignificantly enriched in 3′ -untranslated regions ( UTRs), \ncoding sequences (CDS), and transcription termination sites \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted October 31, 2025. ; https://doi.org/10.1101/2025.10.29.685270doi: bioRxiv preprint \n\n8 \n \n \n \n(TTS, ±1 kb). Notably, MECP2  HM peaks showed a \nsignificant enrichment for CpG islands compared to MECP2 \nwt. In contrast, both MECP2 wt and MECP2 HM peaks were \nsignificantly depleted (q < 0.05) in rRNA and tRNA genes, \nas well as in 5′ -UTRs and transcription start sites (TSS, ±1 \nkb) of protein-coding genes (Fig. S10). \nTo put HM-DyadCap into context with existing enrichment \nmethods, we next studied anti-mC and anti-hmC antibodies, \nbeing the enrichment probes  in MeDIP and hMeDIP  \nprotocols. We were first interested in assessing the dyad \nselectivities of anti -hmC antibodies and their potential use \nfor dyad-specific enrichments. Whereas antibodies for DNA \nmodifications typically require denaturation of the sample \nDNA (leading to loss of dyad information), some antibodies \noffer applications with native dsDNA according to the \nmanufacturers. We first tested three different monoclonal \nantibodies in EMSA with oligonucleotides containin g one \nhmCpG site in single -stranded form, or hybridized to a \ncomplement strand to afford a dsDNA with a single \nhmC/mC dyad. Whereas all antibodies bound the ssDNA \n(Fig S9), none of them showed binding to the dsDNA target \neven at high concentrations ( Fig. 4a). However, we \nidentified a polyclonal antibody that was able to bind this \ndsDNA target at high concentrations, and employed it in \nEMSA with dsDNA targets containing either an hmC/C, \nhmC/mC, hmC/hmC or mC/mC CpG dyad. We observed \nhigher affinity for hmC/hmC and hmC/mC dsDNA and only \nvery low or no affinity for hmC/C and mC/mC targets, \nindicating that this antibody is not suited for selective \nenrichment of individual hmC dyads (Fig 4b-c).  \nTo provide antibody -based reference maps for HM -\nDyadCap, and to facilitate comparison between the two \napproaches, we next conducted standard MeDIP and \nhMeDIP enrichments using denatured , single stranded \ngDNA of the same mESC batches as used above. After \nsequencing and mapping, w e corrected the dataset with a n \nIgG-only control enrichment according to Lentini et al. 54. \nWe then assessed the genome-wide overlap between regions \nenriched with MeDIP and MECP2  wt, respectively . This \nanalysis revealed that 35% of MECP2 wt peaks overlapped \nwith MeDIP. Consistent with this, peak comparison showed \nthat the majority of peaks were unique to M eDIP ( 56%), \nwhereas MECP2 wt contributed a smaller unique fraction \n(29%). A total of 30,717 peaks (15%) were shared between \nthe two datasets (Fig. 4d), indicating that MECP2 wt covers \na subset of genomic regions that is to a good part distinct \nfrom MeDIP. Although DIP-seq methods are important tools \nfor mapping modified DNA bases, they have limitations that \nmay account for the discrepancy of the two datasets.  These \ninclude bias toward low-CG regions61, overrepresentation of \nhighly modified regions62, nonspecific enrichment of short \ntandem repeats54, and (unlike MECP2 wt), an inability to \ndiscriminate between mC/mC and mC/C CpGs . A re -\nevaluation of DIP-seq approaches showed that 50 –99% of \nenriched regions in datasets without correction by an IgG -\nonly control were false positives, regardless of modification \ntype, cell type, or organism54. \nTo next investigate the genomic distribution of the peak sets \nof the MeDIP and MECP2 wt experiments, we assessed their \nenrichment across diverse genomic features . Both datasets \nshowed significant enrichment at CDS, 3′-UTRs, and TTS, \nwhile being consistently depleted at rRNA, tRNA genes, 5′-\nUTRs, and TSS (Fig. S11). Notably, the major discrepancy \nwas observed at CpG islands . To better study this \ndiscrepancy, we examined enrichment of selected peak \nsubsets (unique to MeDIP, unique to MECP2 wt, and \ncommon peaks for both treatments) . M eDIP peaks were \nmarkedly depleted  at CpG islands , whereas MECP2  wt \npeaks were not, with the former being more pronounced for \npeaks that were unique to MeDIP (Fig. 4e, such a differential \nenrichment of regions with high versus low  CpG densities \nhas previously been observed for the two methods 61,63,64). \nNext, comparison of distributions of hM eDIP and MECP2 \nHM peaks showed that 15% of MECP2 HM peaks \noverlapped with the hM eDIP peaks. Peak overlap analysis \nrevealed that a high number of  regions was unique to \nMECP2_HM (60%), while hMeDIP contributed a smaller \nfraction (30%, Fig. 4f). A total of 22,475 peaks were shared \nbetween the two datasets, suggesting that also MECP2 HM \ncovers a predominantly distinct subset of genomic regions \ncompared to hMeDIP. Notably, CpG islands displayed the \nmost pronounced global difference: MECP2 HM peaks were \nslightly enriched, whereas hMeDIP peaks were strongly \ndepleted. Importantly, these trends were mainly driven by \nthe peaks that are specific to each treatment (hMeDIP and \nMECP2 HM), whereas peaks that are common for both \ntreatments show ed an intermediary state (Fig. 4 g).  \nFurthermore, a t TTS, MECP2  HM specific peaks were \nsignificantly enriched , whereas hM eDIP-specific peaks \nshowed only a slight enrichment  (Fig. 4 h). Collectively, \nthese analyses indicate that MECP2 HM binding may be \nmore selectively enriched at regulatory regions  – including \nCpG islands, 3′-UTRs, and TTS – whereas hMeDIP captures \na distinct genomic subsets that may be functionally less \nspecific (Fig. S11 ). Again, whereas MECP2 HM shows \nselectivity for hmC/mC over other hmC -modified dyads  \n(Fig. 2a), hMeDIP is conducted with denatured DNA and \nexpected to enrich hmC regardless of dyad symmetry. \nMECP2 HM binds its target hmC/mC dyad with high \nselectivity over the off -target dyads mC/C, hmC/C, and \nhmC/hmC (>100 -, >125 - and ~50 -fold respectively; Fig. \n2a)33 that furthermore occur at low levels in mESC and other \ncharacterized genomes compared to mC/mC20,23,25. MECP2 \nHM also shows selectivity over mC/mC (10-fold, Fig. 2a33), \nbut this  off -target dyad is particularly abundant 20,23,25. We \nintroduced an additional control to assess  the hmC/mC \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted October 31, 2025. ; https://doi.org/10.1101/2025.10.29.685270doi: bioRxiv preprint \n\n9 \n \n \n \nselectivity of MECP2 HM across the genome . In EMSA \nexperiments, we found  that the glucosylation of hmC to \nglucosyl-hmC (ghmC, Fig. 5a and S1) by T4 BGT \neffectively blocks the binding of MECP2 HM (Fig. 5b and \nS12).  \nFigure 5. Glucosylation control indicates selective hmC/mC \nenrichment by MECP2 HM. a) Reaction scheme for T4 BGT -catalyzed \nglucosylation of hmC.  b) EMSA of MECP2 HM binding to dsDNA \ncontaining a single hmC/mC dyad and glucosylated or not glucosylated \nwith T4 BGT. c)  Venn diagram for peaks obtained from MECP2 HM \nenrichments with glucosylated or non -glucosyated mESC gDNA.  d) \nExample regions of enrichment signal tracks for non -glucosylated and \nglucosylated gDNA (shown are three technical replicates each) . e) Relative \npositioning of enrichment signal strengths in  genes (+/- 10 kb with respect \nto TSS and TTS) with high, medium or low expression.  \nSince MECP2 HM shows only very low binding to hmC off-\ntarget dyads and mC/mC cannot be glucosylated, enrichment \ndifferences between glucosylated and non -glucosylated \ngDNA should indicate hmC/mC dyads with high \nconfidence. Therefore, we glucosylated part of the same \nmESC gDNA batch used in prior experiments, and \nperformed MECP2 HM enrichment to verify if peak loss is \nindeed observed. A comparative peak analysis of these \nexperiments is summarized in the Venn diagram of Fig. 5c. \nIndeed, we observed a significant loss in identified peaks \nupon glucosylation: we identified only 42,194 peaks (+gluc) \nvs. 149,601 in the non -glucosylated enrichment ( -gluc), \nindicating that the large majority of peaks enriched by \nMECP2 HM indeed seems to originate from the selective \nenrichment of hmC/mC dyads. However, there was also \nsome overlap between both treatments (19%; 30,186 peaks), \npresumably due to residual background enrichment of \nclustered off-targets (Fig. 5c ). Fig 5d  shows integrative \ngenomics viewer ( IGV) signal tracks  for MECP2 HM \nenrichments with glucosylated and non-glucosylated mESC. \nThe data show consistent peak enrichment across technical \nreplicates in both cases , with both overlapping and \ncondition-specific regions (Fig. S14). We thereby observed \na striking loss of MECP2 HM signals at different loci for the \nglucosylated control , in agreement with the blocking of \nMECP2 HM by ghmC observed before (Fig. 5b).  \nThe enrichment of hmC within gene bodies has been linked \nto transcriptional activity in diverse tissues5,8,9. To study the \nrelative distribution of enrichment within genes on a global \nscale, we generated metagene read density profiles aligned \nto protein -coding genes, clustered according to their \nexpression levels ( Fig 5e). Profiles were highly consistent \nacross replicates, demonstrating reproducibility of the \nenrichment patterns (Fig. S13). Both MECP2 wt and HM \nshowed robust enrichment across bodies of active compared \nto silenced genes , with pronounced depletion at TSS and \nmild depletion at TTS. Enrichment also extended into \nupstream and downstream flanking regions, consistent with \nprior reports describing enrichment of hmC30 and hmC/mC23 \naround gene bodies of highly expressed genes.  In contrast, \nglucosylation of hmC markedly reduced MECP2 HM \nenrichment. Density plots and heatmaps revealed a \nsubstantial loss of signal across gene bodies, with increased \noccupancy near TSS and the depletion at TTS. Together, \nthese results indicate that MECP2 wt and HM associate with \ngene bodies and adjacent regulatory regions, but that the \ninteraction of MECP2 HM is significantly affected by hmC-\nglucosylation, in agreement with our in vitro EMSA data  \n(Fig. 5b). \nEnhancers are regulatory elements that control the \ntranscription of distal target genes and play essential roles in \ncell differentiation by establishing cell type –specific \ntranscriptional programs65. TET-mediated oxidation of mC \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted October 31, 2025. ; https://doi.org/10.1101/2025.10.29.685270doi: bioRxiv preprint \n\n10 \n \n \n \nhas been shown to modulate enhancer activity during early \nstages of differentiation 66, and enhancers consistently \nexhibit significant levels of hmC across diverse cell types65. \nTo assess how the binding of MECP2 variants is influenced \nby enhancer state, we analyzed the genome -wide \ndistribution of MECP2 variants across enhancer classes \ndefined by  Cruz-Molina et. al. 67. We examined MECP2 \noccupancy at four enhancer categories: active enhancers, \nprimed enhancers, poised enhancers, and poised-to-active \n(PoiAct) enhancers ( Fig. S 15). Distinct binding profiles \nwere observed across enhancer states. In general, MECP2 wt \nand MECP2  HM showed depletion at the center  of \nenhancers, while the MECP2 HM +gluc control showed \nenrichment. The strength of this effect varied between \nenhancer types and showed the most pronounced difference \nbetween MECP2 wt and HM for active and PoiAct \nenhancers, where MECP2 HM exhibited a lower relative \ndepletion at the enhancer center  than MECP2 wt . Input \ncontrols remained flat across all enhancer classes, \nconfirming signal specificity. Overall, these results indicate \nthat MECP2 binding is not uniform but is strongly shaped \nby enhancer state and MECP2 variant.  \n \nCONCLUSION \nIn this study, we report HM-DyadCap, a robust method for \nselectively capturing DNA fragments modified with \nasymmetric hmC/mC CpG dyads that are frequent marks in \nmESC and mouse brain tissue. Central to this approach is \nMECP2 HM, an engineered methyl -CpG-binding domain \nfor that EMSA studies reveal high target affinity and \nspecificity with respect to off -target dinucleotides . We \ndemonstrate that HM-DyadCap enables simple mapping of \nhmC/mC dyads in a mammalian genome, thus representing \nthe first dyad-specific hmC enrichment method.  \nComparative mapping studies with established protocols \nsuch as wild -type MECP2 -based MethylCap as well as  \nMeDIP and hMeDIP reveal differences in the capability of \nenriching particular genomic features, such as a high \nenrichment of CpG islands for MECP2 HM as compared to \nthe hMeDIP protocol.  The use of glucosylation controls \nfurther supports the selectivity of MECP2 HM for hmC/mC \ndyads. Our mapping experiments reveal that these dyads are \nenriched in gene bodies and depleted at transcription start \nsites of protein coding genes in mESC genomes , and that \nthey correlate with active transcription. Moreover, enhancer \nprofiling highlights a dynamic relationship between MECP2 \nbinding and enhancer state, implicating hmC/mC as a \npotentially important mark in enhancer regulation.  \nHM-DyadCap may serve as a valuable tool for studying \nhmC/mC´s regulatory function s and dynamics in \ndevelopment and disease. The ability of dyad -specific \nenrichment is particularly relevant for generating maps in \ngenomes that exhibit low hmC levels , and for large -scale \nstudies for that whole genome deep sequencing may be cost-\nprohibitive. HM -DyadCap will therefore be useful for \ncancer biomarker discovery via effective lower resolution \nmapping in cancer genomes  that have generally low hmC \nlevels. However, the method may also be combined with \nconversion-based sequencing methods to enable high -\nresolution maps without the need for whole genome deep \nsequencing. The MBD of MECP2 can be reengineered to \nrecognize dyads other than hmC/mC 50, promising an \nextension to other dyad marks with high relevance in \nchromatin regulation and cancer development. \n \nDATA AVAILABILITY \nThe processed RNA sequencing data reported in this paper \nwas obtained from ENCODE repository under the accession \ncodes: ENCFF827OZU and ENCFF898TDI. The \nsequencing datasets underlying this study are publicly \naccessible through ArrayExpress under the respec tive \naccession number: E-MTAB-15857 and E-MTAB-15961. \n \nSUPPLEMENTARY INFORMATION \nAssociated content 1: Tables S1-S2, Figures S1-S15 (PDF) \nThis material is available free of charge via the Internet.  \n \nREFERENCES \n \n1 Allis, C. D. & Jenuwein, T. The molecular hallmarks of epigenetic \ncontrol. Nat Rev Genet  17, 487 –500 (2016). \nhttps://doi.org/10.1038/nrg.2016.59  \n2 Bird, A. DNA methylation patterns and epigenetic memory. Genes Dev \n16, 6–21 (2002). https://doi.org/10.1101/gad.947102  \n3 Du, Q., Luu, P. L., Stirzaker, C. & Clark, S. J. Methyl -CpG-binding \ndomain proteins: readers of the epigenome. Epigenomics 7, 1051–1073 \n(2015). https://doi.org/10.2217/epi.15.39  \n4 Zhu, H., Wang, G. H. & Qian, J. Transcription factors as readers and \neffectors of DNA methylation. Nature Reviews Genetics  17, 551–565 \n(2016). https://doi.org/10.1038/nrg.2016.83  \n5 Wu, X. & Zhang, Y. TET -mediated active DNA demethylation: \nmechanism, function and beyond. Nat Rev Genet 18, 517–534 (2017). \nhttps://doi.org/10.1038/nrg.2017.33  \n6 Carell, T., Kurz, M. Q., Müller, M., Rossa, M. & Spada, F. Non -\ncanonical Bases in the Genome: The Regulatory Information Layer in \nDNA. Angew Chem Int Ed Engl  57, 4296 –4312 (2018). \nhttps://doi.org/10.1002/anie.201708228  \n7 Bachman, M. et al. 5-Hydroxymethylcytosine is a predominantly stable \nDNA modification. Nat Chem  6, 1049 –1055 (2014). \nhttps://doi.org/10.1038/nchem.2064  \n8 Pfeifer, G. P., Szabo, P. E. & Song, J. K. Protein Interactions at \nOxidized 5-Methylcytosine Bases. Journal of molecular biology  432, \n1718–1730 (2020). https://doi.org/10.1016/j.jmb.2019.07.039  \n9 Kriukiene, E., Tomkuviene, M. & Klimasauskas, S. 5 -\nHydroxymethylcytosine: the many faces of the sixth base of \nmammalian DNA. Chem Soc Rev  53, 2264 –2283 (2024). \nhttps://doi.org/10.1039/d3cs00858d  \n10 Jin, S. G., Wu, X., Li, A. X. & Pfeifer, G. P. Genomic mapping of 5 -\nhydroxymethylcytosine in the human brain. Nucleic Acids Res  39, \n5015–5024 (2011). https://doi.org/10.1093/nar/gkr120  \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted October 31, 2025. ; https://doi.org/10.1101/2025.10.29.685270doi: bioRxiv preprint \n\n11 \n \n \n \n11 Booth, M. J. et al. Quantitative sequencing of 5-methylcytosine and 5-\nhydroxymethylcytosine at single -base resolution. Science 336, 934 –\n937 (2012). https://doi.org/10.1126/science.1220671  \n12 Yu, M. et al. Base-resolution analysis of 5 -hydroxymethylcytosine in \nthe mammalian genome. Cell 149, 1368 –1380 (2012). \nhttps://doi.org/10.1016/j.cell.2012.04.027  \n13 Li, W. S. et al. 5-Hydroxymethylcytosine signatures in circulating cell-\nfree DNA as diagnostic biomarkers for human cancers. Cell research \n27, 1243–1257 (2017). https://doi.org/10.1038/cr.2017.121  \n14 Song, C. X.  et al.  5-Hydroxymethylcytosine signatures in cell -free \nDNA provide information about tumor types and stages. Cell research \n27, 1231–1242 (2017). https://doi.org/10.1038/cr.2017.106  \n15 Tamanaha, E., Guan, S. X., Marks, K. & Saleh, L. Distributive \nProcessing by the Iron(II)/α -Ketoglutarate-Dependent Catalytic \nDomains of the TET Enzymes Is Consistent with Epigenetic Roles for \nOxidized 5-Methylcytosine Bases. Journal of the American Chemical \nSociety 138, 9345–9348 (2016). https://doi.org/10.1021/jacs.6b03243  \n16 Crawford, D. J.  et al.  Tet2 Catalyzes Stepwise 5 -Methylcytosine \nOxidation by an Iterative and de novo Mechanism. J Am Chem Soc 138, \n730–733 (2016). https://doi.org/10.1021/jacs.5b10554  \n17 Szulik, M. W.  et al. Differential Stabilities and Sequence -Dependent \nBase Pair Opening Dynamics of Watson -Crick Base Pairs with 5 -\nHydroxymethylcytosine, 5 -Formylcytosine, or 5 -Carboxylcytosine. \nBiochemistry (2015). https://doi.org/10.1021/bi501534x  \n18 Booth, M. J., Raiber, E. A. & Balasubramanian, S. Chemical methods \nfor decoding cytosine modifications in DNA. Chem Rev  115, 2240 –\n2254 (2015). https://doi.org/10.1021/cr5002904  \n19 Viswanathan, R.  et al.  DARESOME enables concurrent profiling of \nmultiple DNA modifications with restriction enzymes in single cells \nand cell -free DNA. Sci Adv  9, eadi0197 (2023). \nhttps://doi.org/10.1126/sciadv.adi0197  \n20 Chialastri, A., Sarkar, S., Schauer, E. E., Lamba, S. & Dey, S. S. \nCombinatorial quantification of 5mC and 5hmC at individual CpG \ndyads and the transcriptome in single cells reveals modulators of DNA \nmethylation maintenance fidelity. Nat Struct Mol Biol  31, 1296–1308 \n(2024). https://doi.org/10.1038/s41594-024-01291-w \n21 Kawasaki, Y. et al. A Novel method for the simultaneous identification \nof methylcytosine and hydroxymethylcytosine at a single base \nresolution. Nucleic Acids Research  45 (2017). https://doi.org/ARTN \ne2410.1093/nar/gkw994 \n22 Fullgrabe, J. et al. Simultaneous sequencing of genetic and epigenetic \nbases in DNA. Nat Biotechnol  41, 1457 –1464 (2023). \nhttps://doi.org/10.1038/s41587-022-01652-0 \n23 Hardwick, J. S.  et al.  SCoTCH-seq reveals that 5 -\nhydroxymethylcytosine encodes regulatory information across DNA \nstrands. Proc Natl Acad Sci U S A  122, e2512204122 (2025). \nhttps://doi.org/10.1073/pnas.2512204122  \n24 Bai, D. et al. Simultaneous single-cell analysis of 5mC and 5hmC with \nSIMPLE-seq. Nat Biotechnol  43, 85 –96 (2025). \nhttps://doi.org/10.1038/s41587-024-02148-9 \n25 Halliwell, D. O., Honig, F., Bagby, S., Roy, S. & Murrell, A. Double \nand single stranded detection of 5 -methylcytosine and 5 -\nhydroxymethylcytosine with nanopore sequencing. Commun Biol  8, \n243 (2025). https://doi.org/10.1038/s42003-025-07681-0 \n26 Williams, K. et al. TET1 and hydroxymethylcytosine in transcription \nand DNA methylation fidelity. Nature 473, 343 –348 (2011). \nhttps://doi.org/10.1038/nature10066  \n27 Ficz, G.  et al.  Dynamic regulation of 5 -hydroxymethylcytosine in \nmouse ES cells and during differentiation. Nature 473, 398 –402 \n(2011). https://doi.org/10.1038/nature10008  \n28 Wu, H.  et al.  Genome-wide analysis of 5 -hydroxymethylcytosine \ndistribution reveals its dual function in transcriptional regulation in \nmouse embryonic stem cells. Genes Dev  25, 679 –684 (2011). \nhttps://doi.org/10.1101/gad.2036011  \n29 Robertson, A. B., Dahl, J. A., Ougland, R. & Klungland, A. Pull -down \nof 5-hydroxymethylcytosine DNA using JBP1 -coated magnetic beads. \nNat Protoc 7, 340–350 (2012). https://doi.org/10.1038/nprot.2011.443  \n30 Song, C. X. et al. Selective chemical labeling reveals the genome-wide \ndistribution of 5 -hydroxymethylcytosine. Nat Biotechnol  29, 68 –72 \n(2011). https://doi.org/10.1038/nbt.1732  \n31 Thomson, J. P.  et al.  Comparative analysis of affinity -based 5 -\nhydroxymethylation enrichment techniques (vol 41, pg e206, 2013). \nNucleic Acids Research  42, 4142 –4142 (2014). \nhttps://doi.org/10.1093/nar/gku089  \n32 Hooper, M., Hardy, K., Handyside, A., Hunter, S. & Monk, M. HPRT-\ndeficient (Lesch -Nyhan) mouse embryos derived from germline \ncolonization by cultured cells. Nature 326, 292 –295 (1987). \nhttps://doi.org/10.1038/326292a0  \n33 Buchmuller, B. C.  et al.  Evolved DNA Duplex Readers for Strand -\nAsymmetrically Modified 5 -Hydroxymethylcytosine/5-\nMethylcytosine CpG Dyads. J Am Chem Soc  144, 2987–2993 (2022). \nhttps://doi.org/10.1021/jacs.1c10678  \n34 Molder, F.  et al.  Sustainable data analysis with Snakemake. \nF1000Research 10, 33 (2021). \nhttps://doi.org/10.12688/f1000research.29032.3  \n35 Andrews, S. Babraham Bioinformatics - FastQC A Quality Control \ntool for High Throughput Sequence Data , \n<https://www.bioinformatics.babraham.ac.uk/projects/fastqc/ > \n(2010). \n36 Felix, K. Trim Galore - Babraham Bioinformatics , \n<https://www.bioinformatics.babraham.ac.uk/projects/trim_galore/ > \n(2012). \n37 B, L., C, T., M, P. & SL, S. Ultrafast and memory -efficient alignment \nof short DNA sequences to the human genome. Genome Biol 10 (2009). \nhttps://doi.org/10.1186/gb-2009-10-3-r25 \n38 Li, H.  et al.  The Sequence Alignment/Map format and SAMtools. \nBioinformatics 25, 2078 (2025). \nhttps://doi.org/10.1093/bioinformatics/btp352  \n39 J, F., T, L., B, Q., Y, Z. & XS, L. Identifying ChIP -seq enrichment \nusing MACS. Nature protocols  7 (2012). \nhttps://doi.org/10.1038/nprot.2012.101  \n40 Quinlan, A. R.  et al.  BEDTools: a flexible suite of utilities for \ncomparing genomic features. Bioinformatics 26, 841 –842 (2025). \nhttps://doi.org/10.1093/bioinformatics/btq033  \n41 F, R., F, D., S, D., BA, G. & T, M. deepTools: a flexible platform for \nexploring deep -sequencing data. Nucleic acids research  42 (2014). \nhttps://doi.org/10.1093/nar/gku365  \n42 JT, R.  et al.  Integrative genomics viewer. Nature biotechnology  29 \n(2011). https://doi.org/10.1038/nbt.1754  \n43 A, H., C, W., M, G., CP, P. & G, L. GAT: a simulation framework for \ntesting the association of genomic intervals. Bioinformatics (Oxford, \nEngland) 29 (2013). https://doi.org/10.1093/bioinformatics/btt343  \n44 Brinkman, A. B.  et al.  Whole-genome DNA methylation profiling \nusing MethylCap -seq. Methods 52, 232 –236 (2010). \nhttps://doi.org/10.1016/j.ymeth.2010.06.012  \n45 Hashimoto, H.  et al.  Recognition and potential mechanisms for \nreplication and erasure of cytosine hydroxymethylation. Nucleic Acids \nRes 40, 4841–4849 (2012). https://doi.org/10.1093/nar/gks155  \n46 Buchmuller, B. C., Kosel, B. & Summerer, D. Complete Profiling of \nMethyl-CpG-Binding Domains for Combinations of Cytosine \nModifications at CpG Dinucleotides Reveals Differential Read -out in \nNormal and Rett -Associated States. Sci Rep  10, 4053 (2020). \nhttps://doi.org/10.1038/s41598-020-61030-1 \n47 Jin, S. G., Kadam, S. & Pfeifer, G. P. Examination of the specificity of \nDNA methylation profiling techniques towards 5 -methylcytosine and \n5-hydroxymethylcytosine. Nucleic Acids Research  38 (2010). \nhttps://doi.org/ARTN e12510.1093/nar/gkq223 \n48 Otani, J. et al. Structural basis of the versatile DNA recognition ability \nof the methyl -CpG binding domain of methyl -CpG binding domain \nprotein 4. J Biol Chem  288, 6351 –6362 (2013). \nhttps://doi.org/10.1074/jbc.M112.431098  \n49 Valinluck, V.  et al.  Oxidative damage to methyl -CpG sequences \ninhibits the binding of the methyl -CpG binding domain (MBD) of \nmethyl-CpG binding protein 2 (MeCP2). Nucleic Acids Res 32, 4100–\n4108 (2004). https://doi.org/10.1093/nar/gkh739  \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted October 31, 2025. ; https://doi.org/10.1101/2025.10.29.685270doi: bioRxiv preprint \n\n12 \n \n \n \n50 Kosel, B.  et al.  Evolved Readers of 5 -Carboxylcytosine CpG Dyads \nReveal a High Versatility of the Methyl -CpG-Binding Domain for \nRecognition of Noncanonical Epigenetic Marks. Angew Chem Int Ed \nEngl 63, e202318837 (2024). https://doi.org/10.1002/anie.202318837  \n51 Singh, H. et al. Epigenetic CpG duplex marks probed by an evolved \nDNA reader via a well -tempered conformational plasticity. Nucleic \nAcids Res 51, 6495–6506 (2023). https://doi.org/10.1093/nar/gkad134  \n52 Klose, R. J.  et al.  DNA binding selectivity of MeCP2 due to a \nrequirement for A/T sequences adjacent to methyl -CpG. Mol Cell 19, \n667–678 (2005). https://doi.org/10.1016/j.molcel.2005.07.021  \n53 Lyst, M. J., Connelly, J., Merusi, C. & Bird, A. Sequence-specific DNA \nbinding by AT -hook motifs in MeCP2. Febs Letters 590, 2927–2933 \n(2016). https://doi.org/10.1002/1873-3468.12328 \n54 Lentini, A. et al. A reassessment of DNA -immunoprecipitation-based \ngenomic profiling. Nat Methods  15, 499 –+ (2018). \nhttps://doi.org/10.1038/s41592-018-0038-7 \n55 Lentini, A. & Nestor, C. E. Analyzing DNA -Immunoprecipitation \nSequencing Data. DNA Modifications  2198, 431 –439 (2021). \nhttps://doi.org/10.1007/978-1-0716-0876-0_31 \n56 Matarese, F., Pau, E. C. D. & Stunnenberg, H. G. 5 -\nHydroxymethylcytosine: a new kid on the epigenetic block? Molecular \nSystems Biology  7 (2011). https://doi.org/ARTN \n56210.1038/msb.2011.95 \n57 Guo, J. U.  et al. Distribution, recognition and regulation of non -CpG \nmethylation in the adult mammalian brain. Nat Neurosci 17, 215–222 \n(2014). https://doi.org/10.1038/nn.3607  \n58 Lister, R. et al. Global epigenomic reconfiguration during mammalian \nbrain development. Science 341, 1237905 (2013). \nhttps://doi.org/10.1126/science.1237905  \n59 Kinde, B., Gabel, H. W., Gilbert, C. S., Griffith, E. C. & Greenberg, \nM. E. Reading the unique DNA methylation landscape of the brain: \nNon-CpG methylation, hydroxymethylation, and MeCP2. P Natl Acad \nSci USA  112, 6800 –6806 (2015). \nhttps://doi.org/10.1073/pnas.1411269112  \n60 Lagger, S. et al. MeCP2 recognizes cytosine methylated tri -nucleotide \nand di -nucleotide sequences to tune transcription in the mammalian \nbrain. PLoS genetics  13 (2017). https://doi.org/ARTN \ne100679310.1371/journal.pgen.1006793  \n61 Nair, S. S.  et al.  Comparison of methyl -DNA immunoprecipitation \n(MeDIP) and methyl -CpG binding domain (MBD) protein capture for \ngenome-wide DNA methylation analysis reveal CpG sequence \ncoverage bias. Epigenetics-Us 6, 34 –44 (2011). \nhttps://doi.org/10.4161/epi.6.1.13313  \n62 Ko, M. et al. Impaired hydroxylation of 5 -methylcytosine in myeloid \ncancers with mutant TET2. Nature 468, 839 –843 (2010). \nhttps://doi.org/10.1038/nature09586  \n63 Li, S. & Tollefsbol, T. O. DNA methylation methods: Global DNA \nmethylation and methylomic analyses. Methods 187, 28 –43 (2021). \nhttps://doi.org/10.1016/j.ymeth.2020.10.002  \n64 Neary, J. L., Perez, S. M., Peterson, K., Lodge, D. J. & Carless, M. A. \nComparative analysis of MBD -seq and MeDIP -seq and estimation of \ngene expression changes in a rodent model of schizophrenia. Genomics \n109, 204–213 (2017). https://doi.org/10.1016/j.ygeno.2017.03.004  \n65 Heinz, S., Romanoski, C. E., Benner, C. & Glass, C. K. The selection \nand function of cell type -specific enhancers. Nature reviews. \nMolecular cell biology  16, 144 –154 (2015). \nhttps://doi.org/10.1038/nrm3949  \n66 Hon, G. C.  et al. 5mC oxidation by Tet2 modulates enhancer activity \nand timing of transcriptome reprogramming during differentiation. Mol \nCell 56, 286–297 (2014). https://doi.org/10.1016/j.molcel.2014.08.026  \n67 Cruz-Molina, S.  et al.  PRC2 Facilitates the Regulatory Topology \nRequired for Poised Enhancer Function during Pluripotent Stem Cell \nDifferentiation. Cell Stem Cell  20, 689 –705 e689 (2017). \nhttps://doi.org/10.1016/j.stem.2017.02.004  \nAUTHOR CONTRIBUTIONS \nL.E. and D. Sc. conducted wet lab experiments and analyzed \ndata. M.Z., A.S. and S.B. analyzed NGS data and \ncontributed to data interpretation. K.K. conducted EMSAs. \nB.B. developed methods. S.T. and J.I. supported the analysis \nof NGS data. C.S. conducted cell culture of  mESCs and \nprovided consulting. D. Su., S.B., M.Z., L.E. and K.K. wrote \nthe manuscript with input from all other authors. D. Su. and \nS.B. supervised the study. \n \nACKNOWLEDGEMENT \nWe thank the Faculty of Chemistry and Chemical Biology \nof the TU Dortmund University and the International Max-\nPlanck Research School for Living Matter for continuous \nsupport. We thank the members of the group for support. We \nthank Agnieszka Zelisko and Anne-Clémence Veillard from \nDiagenode for providing the hMeDIP antibody as well as for \ndiscussions and input regarding DIP experiments.  \nFUNDING \nFunded by the Deutsche Forschungsgemeinschaft \n(SU726/10-1), the Max -Planck Society, the CANTAR \nprogram “Netzwerke 2021” and the NRW returning scholars \nprogram, both being  initiatives of the Ministry of Culture \nand Science of the State of Northrhine Westphalia. Funded \nby the European Union (ERC, 101100794 COMBICODE). \nViews and opinions expressed are those of the author(s) only \nand do not necessarily reflect those of the European Union \nor the European Research Council. Neither the European \nUnion nor the granting authority can be held responsible for \nthem. \n \nCONFLICT OF INTEREST \nThe authors declare the following competing financial \ninterest: TU Dortmund University has filed a patent \napplication for the engineered MBD employed in the present \nstudy (PCT/EP2020/087979, pending). No further \ncompeting financial interests have been declared. \n \n \n \n \n \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted October 31, 2025. ; https://doi.org/10.1101/2025.10.29.685270doi: bioRxiv preprint","source_license":"CC-BY-4.0","license_restricted":false}