{"paper_id":"33f4217b-6dad-448c-8497-8f3c38499ae5","body_text":"Identifying crossovers in a cattle pangenome containing \nhaplotype-resolved assemblies from half-siblings \n \nAlexander S. Leonard1,* and Hubert Pausch1 \n \n1 Animal Genomics, ETH Zurich, Universitaetstrasse 2, 8092 Zurich \n* alleonard@ethz.ch \n \nKeywords: cattle, long reads, genome assembly, pangenomics, recombination, epigenetics \nAbstract \n \nBackground: \nRecombination of parental haplotypes is a fundamental biological process that ensures proper \nsegregation of homologous chromosomes and creates new combinations of alleles during \nmeiosis. Crossover events are typically detected from large-scale pedigree-based genetic \nstudies or linkage disequilibrium-based recombination maps, although these are generally \nlimited to SNPs. Increasing amounts of long read sequencing and haplotype-resolved \nassemblies offer an alternative approach to examining recombination events at basepair \nresolution, albeit with much smaller sample sizes. \n \nResults: \nHere, we analyse five high-quality genome assemblies from the Simmental cattle breed, \nincluding a newly assembled triobinned HiFi assembly of an Eringer x Simmental cross (N50 \nof 77 Mb and a k-mer quality value of 55.3). We integrate the five assemblies, of which two \noriginate from maternal half-siblings, into a reference-free Simmental-specific pangenome. \nBy considering path similarities in the pangenome, we were able to identify putative \ncrossover events in the haplotypes of the half-siblings, as well as a greater number of events \nrelative to the cousin due to an additional degree of generational separation. We validated the \npangenome approach with phased SNPs called from linear alignments of maternal short read \nsequencing, with 23 of 30 chromosomes having the same recombination predictions. We \nidentified 5 and 16.7 Mb of non-reference insertion sequences respectively shared or private \nto the half-siblings, enabling testing for recombination events beyond only SNP markers. We \nalso identified four differentially methylated CpG clusters from the 5mC signal of HiFi reads \nwhich allowed us to narrow the window containing the putative recombination event from 35 \nto 20 Mb within the longest run of homozygosity. \n \nConclusion: \nStructural variants and methylation information identified from long read sequencing and \ngenome assemblies may help identify recombination events in regions beyond those typically \ncalled from SNPs. Furthermore, while existing long read-based methylation calls can be \nnoisy and report unrealistic intermediate methylation levels, 5mC methylation appears to be a \npromising avenue for distinguishing haplotypes in the absence of genomic variation.  \nBackground \n \nRecombination is a crucial driver of genomic diversity, allowing the random exchange of \ngenetic segments between homologous chromosomes which creates new haplotype and allele \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 20, 2026. ; https://doi.org/10.64898/2026.02.20.706955doi: bioRxiv preprint \n\ncombinations. This process can improve selection of advantageous combinations and purge \ndeleterious alleles [1]. Recombination rates vary substantially along the chromosomes [2], \nwhere recombination hotspots are short genomic regions where genetic material is exchanged \nat much higher rates than elsewhere. These recombination hotspots enable rapid evolution of \ngenes, particularly those related to immunity [3]; oppositely, some regions may be repressed \nfor recombination, like supergenes [4]. A detailed understanding of recombination maps can \nimprove estimates of variant linkage disequilibrium used in imputation and other downstream \ngenomic analyses [5]. \n \nRecombination rates are commonly evaluated in cattle using sparse, SNP-based genotype \narrays in large-scale pedigreed mapping cohorts [2, 6–8] or linkage disequilibrium [9] \napproaches. While these methods often account for effective population size and \ndemographic history, their resolution is limited by the density of markers and are often \nlacking complex or repetitive genomic regions. Recombination events can also be detected \ndirectly from short read sequencing [10] or Hi-C reads [11, 12] from single gametes. \n \nRecombination is also a source of structural variation, particularly through nonallelic \nhomologous recombination. These rearrangements can lead to phenotypic disorders [13], \nalthough they can be challenging to interrogate with short read sequencing as they often fall \nin repetitive regions. Long reads and genome assemblies have improved the resolution of \ncomplex recombination events such as recurrent inversions [14] and transposable elements \n[15]. Pangenomes can be used to evaluate regions of similarity and differences across \ngenome assemblies without the explicit need for a reference [16]. Haplotype-resolved cattle \nassemblies have previously been integrated in pangenomes, but primarily for assessing gene \npresence/absence [17, 18] or putatively trait-associated structural variants [19, 20]. Recent \nadvances in pangenome alignment account for recombination [21, 22], utilising haplotype \ncombinations that are not explicitly present in the graph. \n \nVariation in methylation, primarily the 5mC marker, has also been associated with mediating \nor limiting recombination in arabidopsis [23], barley [24], and cannabis [25]. Although there \nis not a universal consensus, some studies report transgenerational inheritance of methylation \npatterns in mouse [9] and grape [26], as well as in human sperm for another methylation \nmarker m6A [27]. Methylation is also an increasingly promising source of haplotype-specific \ninformation, allowing improved phasing of sequencing reads or variants [28]. However, \nmethylation information has not yet been used to investigate recombination events in cattle. \n \nHere, we investigate crossover events in five highly contiguous and near-complete Simmental \ngenome assemblies. While the number of assemblies is far too small to confidently identify \nrecombination hot- or coldspots, we demonstrate the possibility of identifying crossover \nevents with structural variants or methylation. These structural variants and methylation \nmarkers, readily accessible from recent long read sequencing, may help identify crossover \nevents within SNP-based runs of homozygosity. \nMethods \n \nGenome assemblies \nWe used 4 publicly available Simmental assemblies (Table 1) and assembled an additional \ngenome from a Eringer x Simmental cross. DNA was prepared from blood samples of the F1 \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 20, 2026. ; https://doi.org/10.64898/2026.02.20.706955doi: bioRxiv preprint \n\nand cryopreserved semen samples of the Eringer sire. We generated 101 Gb of HiFi reads \nwith a read N50 of 16.4 Kb from one PacBio Revio 25M SMRT cell. We also generated 83 \nand 62 Gb of Illumina 2×150bp paired-end reads for the F1 and sire respectively, with 70 Gb \ndam short reads publicly available (SAMEA115121771). We followed the assembly and \nquality assessment approach described in [18]. Briefly, we assembled HiFi sequencing reads \nwith hifiasm [29] v.0.19.9 with parental k-mers from short read sequencing to produce \nhaplotype-resolved assemblies. We scaffolded the contigs to the cattle reference genome, \nARS-UCD2.0, and assessed gene completeness with compleasm [30] v0.2.6 and base-level \ncorrectness with merqury [31] v1.3. \n \nWe determined whole-genome pairwise similarity using mash [32] v2.3, with 10k sketches of \nk-mer length 31. We then created a tree using the UPGMA linkage algorithm from SciPy [33] \nv1.15.2. \n \nPangenome construction \nWe constructed per-chromosome pangenomes with pggb [16], using wfmash commit \n9c63d98, seqwish commit 0eb6468, smoothxg commit e93c623, and GFAffix v.0.2.0, with a \nsegment length of 5000 and an identity target of 0.975. \n \nJaccard distance-based crossover identification \nWe assessed binned Jaccard distances across the pangenome as described in [19]. Briefly, we \nbinned the “USA_1” Simmental haplotype into 1 Kb bins and extracted pangenomic \nsubgraphs corresponding to each bin using the extract command from odgi [34] v0.9. We \nthen calculated pairwise Jaccard distances on each subgraph using odgi similarity. \n \nWe created the Jaccard distance tree using the same UPGMA approach as for the mash \ndistances, instead using the median Jaccard distance over the whole genome for the input \nmatrix. \n \nSNP crossover identification \nWe aligned short sequencing reads from the Simmental dam (SAMEA115121771) of the \nhalf-sibling F1s against the “USA_1” Simmental haplotype as the reference using Strobealign \n(v0.15) [35]. We then called SNPs with DeepVariant [36] v1.8, using the “WGS” small \nmodel. We used BCFtools (v1.21) [37] to filter for biallelic SNPs and subset variants. \n \nTo derive variants for both half-siblings and the cousin assembly, we used vg deconstruct \n(v1.63.0) [38] on each chromosome pangenome, followed by vcfwave (v1.0.12) [39] to \nsimplify variant representation. \n \nPangenome alignment of short reads \nTo allow feasible mapping, individual chromosome graphs were simplified by only retaining \nnodes relevant to the reference path or the two half-sibling assemblies, followed by removing \nnon-reference stubs with “vg clip -s”. We further simplified the graphs with “vg simplify -i 0 \n-L 0.8 -k”. We then merged all chromosomes (autosomes and X) together using vg ids and vg \ncombine to produce a single whole-genome GFA. \n \nWe used vg autoindex with “-w sr-giraffe” and vg giraffe [40] in GAF mode with “--named-\ncoordinates” to align the Simmental dam short sequencing reads previously used in the SNP \nanalysis. We used vg surject to get linear-reference coordinates for the alignments for \ncomparisons and annotating pangenome plots made with BandageNG (v2025.5) [41]. \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 20, 2026. ; https://doi.org/10.64898/2026.02.20.706955doi: bioRxiv preprint \n\n \nMethylation crossover identification \nWe aligned HiFi sequencing reads for both half-sibling F1s using pbmm2 (v1.17.0) against \nthe “USA_1” Simmental haplotype as the reference. Reads were phased using triocanu (v2.3) \n[42] using k-mers from short read sequencing of their respective parents. We assessed per-\nCpG methylation only for the maternal assigned reads using pb-CpG-tools (v3.0.0) using a \nminimum coverage of 10 reads. We then took the absolute difference in CpG methylation at \nsites present in both F1s to assess methylation similarity. We used MethBat (v0.14.2) joint-\nsegment to identify regions of allele specific methylation. \nResults \nWe collected one primary and three haplotype-resolved Simmental genome assemblies from \nfour publicly available samples. An additional HiFi-based Simmental haplotype-resolved \nassembly was generated in this study from an Eringer x Simmental crossbred animal (Table \n1). Briefly, this assembly was extremely high quality (Supplementary Table 1), with an \nNG50 of 77 Mb, an estimated gene completeness of 99.64%, and an estimated base-level \ncorrectness of 99.9997% (QV~55.3). Three of the haplotyped-resolved assemblies were from \nclosely related cattle, with two half-siblings sharing the same dam and the other as their first \ncousin (sharing a grandsire). The other assembly from a Swiss Simmental individual was \ndistantly related to the three closely related ones, although separated by many generations. \nThe remaining assembly from an American Simmental cow had no discernible pedigree \nconnection with the other animals (Figure 1a). We confirmed the expected relationship of the \nanimals with a mash tree based on k-mers in the genome assemblies (Figure 1b), with the \nhalf-siblings and cousin clustering together. \n \nTable 1. Simmental genome assemblies used in pangenome construction. The USA_1 assembly is used as the Simmental \nreference genome. NG50 is calculated with respect to a 3 Gb genome size.  \nAssembly Related Accession Read technology Phasing NG50  \nHalfsib_1 Yes PRJEB42335 HiFi Haplotype-resolved 46.9 \nHalfsib_2 Yes this study HiFi Haplotype-resolved 76.7 \nCousin_1 Yes PRJEB42335 HiFi Haplotype-resolved 84.3 \nSwiss_1 Yes PRJEB72196 CLR Primary 15.4 \nUSA_1 No PRJNA677947 ONT Haplotype-resolved 70.6 \n \nDetecting recombination points in a pangenome \nWe used pggb to construct per-chromosome pangenomes across the 29 autosomes and X \nchromosome. Coincidentally, the four haplotype-resolved Simmental assemblies were all the \nmaternal haplotype of three male F1s and one female F1, and so we excluded the Y \nchromosome from consideration. Despite the graphs only containing five assemblies of the \nsame breed, we identified 443 Mb of “non-reference” sequence in total with respect to the \n“USA_1” reference haplotype (graph length of 3.09 Gb and reference length of 2.64 Gb). As \npreviously reported [19], we confirmed this is largely due to unaligned centromere “tips” in \nthe graph resulting from cattle chromosomes being acrocentric, with 259 Mb of the non-\nreference sequence contained in 60 massive nodes. The graphs otherwise contained 27.4 M \nnodes and 37.2 M edges (for an average degree of 1.37), comparable to other bovine graphs \ngenerated by pggb [43]. \n \nWe quantified genomic similarity between the five Simmental assemblies using the pairwise \nJaccard distance calculated over the per-assembly paths the graph for 1 Kb windows (with \nrespect to the “USA_1” reference). The distance is close to 0 when two assemblies take the \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 20, 2026. ; https://doi.org/10.64898/2026.02.20.706955doi: bioRxiv preprint \n\nsame paths throughout the chosen range and approaches 1 if they share no nodes (weighted \nby the length of the nodes). As such, we were able to construct another tree based on pairwise \ndistances (Figure 1c), recovering the topology of the previous k-mer-based tree. However, \nthe distance-based tree suggested a substantial (but consistent) higher genomic divergence, \nroughly three times the estimated mash divergence. This potentially resulted from poorly \nconstructed or complex tangles in the graph, leading to low node similarity despite sharing \nsimilar sequence, or structural variants larger than 1 Kb. There were (11.4±3.0)×103 of such 1 \nKb bins with a Jaccard distance above 0.9 (approximately 11 Mb or 0.4% of the reference \ngenome length), indicating such regions are comparatively rare. \n \n \n \n \nFigure 1. (a) Pedigree of samples derived from herdbooks where available. Simmental animals are highlighted in red and the \ngender of the F1s is indicated as male (M) or female (F). Swiss_1 is distantly related to the F1s through several generations \nof separation, while USA_1 has no presumed relationship to the others, and so their pedigrees are indicated with a dashed \nline. (b) Relationship tree based on the pairwise mash distance of assemblies. The cluster corresponding to the known \nrelated animals is coloured in grey. (c) Relationship tree based on the pangenome pairwise Jaccard distances of assemblies. \n \nThe haplotype assemblies constructed from the half-siblings were nearly identical to each \nother over large stretches of each chromosome but typically had at least one distinct break \nwhere the subsequent Jaccard distance was closer to that expected for unrelated assemblies \n(Figure 2a, Supplementary Figure 1). These breaks likely correspond to recombination \nevents, where the “maternal haplotype” assembled for each of the half-siblings corresponds \nto a different haplotype of the diploid maternal genome, changing from identical by descent \n(IBD) to non-IBD. We did not observe such a pattern in the unrelated assemblies, supporting \nthat this method is sensitive only to real regions of IBD. While we averaged the Jaccard \ndistance over 1 Kb windows to help mitigate small assembly errors, the true recombination \nresolution of the pangenome is limited only by the presence of heterozygous sequences. We \na)\nb) c)\nHalfsib_1 (F) Halfsib_2 (M) Cousin_1 (M)\nSire_1 Sire_2 Sire_3Dam_1 Dam_2\nGrandsire_1Granddam_1\nGreat...grandSire\nGranddam_2\nSwiss_1 (M)\n4 generations 5 generations\nUSA_1 (M)\n0\n0.2\n0.4\n0.6\nUSA_1 Cousin_1 Halfsib_1Halfsib_2Swiss_1 USA_1 Cousin_1 Halfsib_1Halfsib_2Swiss_1\n0\n0.05\n0.10\n0.15\n0.20\nMash divergence\nJaccard distance\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 20, 2026. ; https://doi.org/10.64898/2026.02.20.706955doi: bioRxiv preprint \n\nre-examined two putative recombination events on BTA9 shown in Figure 2a, confirming the \ndistinct change in heterozygous sequence between the half-siblings (Figure 2b) with the \npresence or absence of half-sibling-specific nodes. In one of those events, a 1.1 Kb \nheterozygous SV was within 7 Kb of the putative recombination point, further validating the \nheterozygous SNPs nearby and demonstrating SVs can also tag haplotype-of-origin.  \n \nBased on distinct breaks in Jaccard distance (e.g. several consecutive bins are zero followed \nby nonzero bins) between the assemblies of the half-siblings, we estimated a mean of 1.9 \nputative crossover events per chromosome (Supplementary Table 2), close to the \nexpectation of 2 meiotic crossovers for half-siblings (one crossover event for each dam-\noffspring pair). Crossovers were less frequent in shorter chromosomes, although we observed \nseveral longer chromosomes (10 and 12, respectively 102.5 and 88.5 Mb) without any \nputative recombination events. Detecting crossovers on the X chromosome was challenging \ndue to several regions of suboptimal graph construction in complex regions. Likewise, the \nCousin_1 assembly had multiple smaller stretches of near identical and average similarity \ncompared to the haplotypes of the half-siblings, corresponding to multiple putative \nrecombination events (2.9 and 2.8 events per chromosome respectively). The larger value \ncompared to between the half-sibling assemblies is expected, given the additional generation \nof separation, but is still an underestimate as we cannot detect recombination within the non-\nshared haplotype within the cousin. \n \nWe investigated the sensitivity of detecting crossover events with the pangenome approach \nby considering heterozygous SNPs called from short read sequencing of the half-siblings’ \ndam. Both half-sibling haplotypes overwhelmingly inherited the same allele up to the 1 Kb \nbin containing the putative crossover, and subsequently inherited opposite alleles (Figure 2c, \nSupplementary Figure 2), validating the predicted haplotype inheritance. Even the X \nchromosome graph prediction was validated, demonstrating the graph approach is robust \neven over complex regions. Several recombination events predicted by the graph-based \nJaccard distance were due to runs of homozygosity (RoHs), identified by the absence of \nheterozygous SNPs, and so the number of predicted crossover events per chromosome was \nreduced from 1.9 to 1.5 (Supplementary Table 2). However, predictions matched between \nthe graph and SNP approach for 23 out of the 30 chromosomes considered. Only \nchromosomes 5 and 17 had predictions differing by more than two predicted recombination \nevents, with the graph potentially overestimating 3 and 4 recombination events respectively. \n \nDue to the additional degree of relationship separation for the cousin assembly, it was \npossible to observe both IBD and non-IBD heterozygous SNPs in close proximity. However, \nwe still observed extended stretches of consistent phasing. Since X-chromosomal inheritance \nis different to autosomal inheritance, we observed more distinct phase switches between the \ncousin and the half-sibling assemblies (i.e. little mixing of homozygous alternate and \nheterozygous SNPs). The only exception was in the X-PAR, which displayed similar \nbehaviour to the autosomes as expected given the less restricted recombination in this region. \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 20, 2026. ; https://doi.org/10.64898/2026.02.20.706955doi: bioRxiv preprint \n\n \n \n \nFigure 2. (a) Path Jaccard distance binned every 100 Kb across four pairs of assemblies on BTA9. Long stretches at zero \nindicate near-identical paths, with putative recombination events marked by arrows. The Swiss_1 & USA_1 pair have no \nexpected relationship, hence have non-zero values along the entire chromosome. (b) Bandage plot for BTA9 subgraphs for \nthe first (top) and second (bottom) recombination events between the half-siblings. Arrows mark the node-level resolution of \nrecombination events, as determined by the presence of half-sibling specific nodes. (c) Heterozygous maternal SNPs binned \nevery 100 Kb where the two assemblies had the same (blue) or opposite (orange) genotype. Arrows mark the same \ncoordinates as predicted from recombination events in (a).  \nWe also investigated the sequence context of the putative recombination events and \ndifferences in the PRDM9 alleles, the gene associated with recombination hotspot \npositioning. Using the 15 bp long consensus “NCCNCCNTNNCCNCN”-motif of human \nPRDM9 allele A [44] (where N represents an unconstrained nucleotide), we identified 87,848 \npossible matches across the USA_1 reference genome. Out of the 92 putative recombination \nspots, the median distance to a binding motif was 11.1 Kb, with six spots within 1 Kb. \nHowever, the proximity was not statistically significant, with randomly permuted hotspots \nshowing a similar distribution of proximity to binding motifs (one-sided “less” Mann-\nWhitney U test p=0.154). The lack of significance might be due in part to the low specificity \nof the binding motif and possible human-cattle PRDM9 allele differences. We also examined \nthe allelic sequence of PRDM9, identifying the gene sequence from the annotated cattle \nreference genome ARS-UCD2.0 in the Simmental pangenome (Supplementary Figure 3). \nThe four Swiss Simmental assemblies had identical sequence across the 560 bp span \nconsisting of five 84 bp tandemly repeated motifs (with 81 and 59 bp partial motifs preceding \nand following the main repeats). The USA_1 assembly contained a single T-to-C substitution \nat the 81st base of the 3rd full tandem repeat motif. \n \n \nPangenome alignment distinguishes maternal runs of homozygosity \nWhen comparing the pangenome- and SNP-based crossovers, we noticed it was not possible \nto distinguish extended RoHs in the dam from the half-siblings inheriting the same maternal \nhaplotype. We identified three such extended RoHs of roughly 10 Mb or longer from \nmaternal SNPs called from short read sequencing. In regions of obvious recombination, \nheterozygous SNPs were substantially more common than heterozygous insertions or \ndeletions (indels). However, in all three RoHs, we observed heterozygous indels \noutnumbered heterozygous SNPs (Supplementary Table 3). The majority of these indels \nproved to be putative assembly errors, typically in homopolymers or other “off-by-one” \ndifferences in long SVs that could not be manually validated (Supplementary Figure 4). In \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 20, 2026. ; https://doi.org/10.64898/2026.02.20.706955doi: bioRxiv preprint \n\nseveral instances, indels from the non-Simmental parent of the crossbred F1s were \nerroneously included in the Simmental haplotypes, due to limited phasing power in the local \nregion (Supplementary Figure 5). Another instance of a 2.3 Kb insertion unique to the \nHalfsib_1 haplotype assembly was later identified in both half-sibling haplotypes, where the \nvariant was lost from the Halfsib_2 haplotype assembly during the vcfwave postprocessing of \nthe pangenome (Supplementary Figure 6). \n \nFor completeness, we investigated if using short sequencing reads from the dam aligned \ndirectly to the whole-genome pangenome, combing the 29 autosomes and X chromosome \nbuilt earlier into one graph, would distinguish maternal RoHs from when the haplotypes were \nIBD. We were able to manually verify several SVs that were either in one or both half-sibling \nassemblies, corresponding to a heterozygous or homozygous SV respectively. In the \nheterozygous case, we observed the dam short read graph alignments taking two distinct \npaths through the SV bubble (Figure 3a), while the homozygous alternate cases only had a \nsingle consistent path (Figure 3b). We also observed cases where the two assemblies shared \nthe same path, but the maternal sequencing also supported an alternative path, indicating the \ndam was heterozygous while both half-siblings inherited the same allele (Figure 3c). We \ngeneralised this approach by classifying individual alignment paths as congruent with the two \nhalf-sibling assemblies or not, enabling us to identify when maternal sequencing supported an \nalternative path to either assembled haplotype (Figure 3d, Supplementary Figure 7). We \nfound even coverage of read-to-pangenome alignments across the chromosomes, \ndemonstrating the above effect was not due to hard-to-align regions or other causes of \nalignment dropout. In combination with the path Jaccard distance, we could now classify \nregions as maternal RoHs, IBD, or non-IBD solely through the analysis of or alignment to the \npangenome without the need for any linear reference-based approaches. \n \n \n \nFigure 3. Maternal short read sequencing alignments to the graph indicate support for (a) heterozygous, (b) homozygous \nalternate, or (c) variants that are heterozygous in the dam but both half-siblings inherited the same allele. Blue and red \npaths respectively indicate the pangenome path for the Halfsib_1 and Halfsib_2 assemblies, while the black arrows indicate \nexplicit edge junctions supported from maternal short read sequencing aligned to the graph. (d) Incongruence ratio is the \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 20, 2026. ; https://doi.org/10.64898/2026.02.20.706955doi: bioRxiv preprint \n\nratio of maternal alignments that follow a graph path found in either half-sibling to those that don’t across 1Kb bins. A low \nincongruence ratio indicates that both maternal haplotypes are well represented in the graph, while a high value suggests \nthe half-siblings are IBD and represent only one of two distinct maternal haplotypes. Arrows mark the same recombination \nevents from Figure 2a. \nDistinguishing haplotypes with 5mC methylation \nWe investigated the 5-Methylcytosine (5mC) signal from the HiFi reads, hypothesising \ndifferent epigenetic modifications might enable to distinguish between homozygous \nhaplotypes. Both DNA samples used for HiFi sequencing to produce the half-sibling \nassemblies were extracted from blood and at similar ages, reducing possible bias in tissue-\nspecific or age-mediated methylation. We assessed 16.4 million reference CpG sites called in \nboth half-siblings (24.6M in Halfsib_1 and 19.3M in Halfsib_2). Of these, 12.3M were highly \nmethylated (>70%) in both, while 1.3M where unmethylated (<30%). Overwhelmingly, \nmethylation status was similar between both (Supplementary Figure 8), with 14.9M (91% \nof all shared CpGs) having an absolute difference of methylation status below 30%, with only \n0.2M (1.3% of all shared CpGs) sites with an absolute difference above 70% and considered \nto be differentially methylated sites (DMS). However, most of these sites involved mutations \nto the CpG motif, either removing or introducing a CpG motif rather than only modifying the \nepigenetic context and hence is strongly correlated with the heterozygous SNP analysis \nalready presented (Figure 4a). We were unable to identify a genome-wide signal of \nrecombination after only considering epigenetic changes to conserved CpG sites (Figure 4b). \n \n \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 20, 2026. ; https://doi.org/10.64898/2026.02.20.706955doi: bioRxiv preprint \n\n \nFigure 4. (a) The fraction of differentially methylated sites (DMS) to total CpGs in 100 Kb windows recovers the \nrecombination events identified with pangenome similarity and SNPs, with the arrows taken from Figure 2. However, the \nsignal is largely driven by the number of SNPs involving CpG motifs (either removing or adding), shown on an alternative y \naxis. (b) After excluding disrupted CpG motifs, the signal of recombination events is lost. (c-d) Mean 5mC methylation \nscores per CpG, averaged over all reads for each half-sibling, in a (c) expected IBD region on BTA22 and the (d) RoH on \nBTA17. Boxed regions indicate the detected allele-specific methylation between the two half-siblings. \n \nWe investigated genome-wide signals of allele-specific methylation (ASM), using MethBat \nsegment on a pseudo-diplotype of the two Simmental half-siblings to identity regions of \ninterest without limiting the analysis to a priori defined windows (i.e. CpG islands which are \nnot defined for the USA_1 reference). There were 207 ASM windows, ranging from 102 to \n3,267 bp, with a mean length of 901±611 bp. Of these, 139 (67%) occurred in non-IBD \nregions while 68 (33%) fell within IBD regions. Approximately 54% of the USA_1 reference \ngenome was within non-IBD regions (1.43 out of 2.64 Gb), leading to a statistically \nsignificant result that ASM was depleted in IBD regions (One-sided Fisher’s exact test \np=9.4×10-6). We confirmed the effect by also randomly permuting the windows 100 times \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 20, 2026. ; https://doi.org/10.64898/2026.02.20.706955doi: bioRxiv preprint \n\n(Supplementary Figure 9), calculating a statistically significant z-score of p=0.0039, \nindicating the real ASM windows occur within non-IBD regions more commonly than by \nchance. Despite the depletion of ASM in IBD regions, there are compelling cases of ASM \nwithin regions that are predicted to be IBD based on genetic sequence (Figure 4c). These \nmay indicate non-inheritance of 5mC methylation state at some loci or biases like cell types \nor developmental stage affecting methylation at these loci.  \n \nManually inspecting the alignments, we did not find evidence of a haplotype-switching event \n(i.e. were IBD between half-siblings before and after the RoH) in the long RoHs on BTA3 \nand BTA7. The BTA17 RoH (34.7 Mb long), however, changes from heterozygous to IBD \nbetween the half-siblings. This putatively suggests a crossover event occurred which was not \ndetectable from variation in the genomic sequence, although the methylation signals can be \ndominated by noise. We further manually identified four candidate regions within the BTA17 \nRoH with clustered CpGs displaying hypermethylation in the maternal reads of one half-\nsibling but hypomethylation in the other (Supplementary Table 4). The last ASM we \nobserved in this RoH was at 35.1 Mb, nearly 15 Mb after the start of the RoH, putatively \nnarrowing the window where a recombination event might have occurred (Figure 4d). We \nsimilarly identified two additional candidate regions within the smaller BTA3 RoH (9.5 Mb \nlong), where the half-siblings presumably remained in opposite phases. We only identified \none weak candidate region within the BTA7 RoH (45.2 Mb long), as expected given the half-\nsiblings presumably remained in the same phase. Without methylation information of the dam \nitself, we could not assess consistency in phase for differential methylation, instead only \nquantifying the magnitude of the absolute difference. \n \nWe were also able to use the 5mC methylation signal within regions containing opposite \nhaplotypes but limited number of heterozygous SNPs to further improve phasing. We used \nSNPs and the long reads to phase variants into 3,430 haploblocks, located within the \nheterozygous regions. With pomfret, we identified 10 possible joins supported by the \nmethylation signal in the absence of genomic variants (Supplementary Table 5, \nSupplementary Figure 10), with an average of 31±6 informative “meth-mer” sites flanking \nthe ambiguous phase block. Given the lack of phase blocks in homozygous regions, we were \nunable to assess methylation-based phasing in those regions. \nDiscussion \nWe investigated the feasibility of using genome assemblies in a pangenome for assessing \nrecombination events. While these types of data will unlikely reach the same scale of array-\ngenotyped cohorts, and thus be less able to predict recombination hot- or coldspots, we \ndemonstrate that structural variants and 5mC methylation can also be used to inform \nrecombination events. Through the reference-free pangenome approach, we also identified \n21.7 Mb of insertion sequence found in either of the half-sibling assemblies, 5 Mb of which \nwas in both haplotypes, which could contain phasing-relevant variation. However, \npangenome approaches can depend on the style and quality of the graph to accurately \nquantify relationships between all pairs of assemblies [43], and so more robust methods and \nvalidation may be needed to separate misassemblies from true recombination events. \n \nAlthough we did not observe any recombination events that could not be predicted from \nSNPs, potentially due to the small sample size (n=2 haplotypes), considering additional \nsources of genomic or epigenetic variation may help resolve large runs of homozygosity. In \nsome species, like European Bison, over half of the genome may be in (SNP-based) runs of \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 20, 2026. ; https://doi.org/10.64898/2026.02.20.706955doi: bioRxiv preprint \n\nhomozygosity but still contain comparable amounts of structural variants to other bovine \nspecies [45]. Variable number tandem repeats have a higher average mutation rate than SNPs \n[38] and are more likely to cause de novo mutations and interrupt RoHs, but are currently \noverlooked in recombination analyses. Although the assemblies examined here were also of \nhigh quality, homopolymer errors dominated putative heterozygous indels within RoHs. \nIncorporating sequencing reads with limited homopolymer bias for polishing or improved \nbasecalling may be useful to reduce false positives in these regions. Similarly, the pangenome \napproach is sensitive enough to identify gene conversion events or other short recombination \nevents, but the current specificity is limited by the assembly base-level accuracy. \n \nRead-to-graph alignments could also be used with an alternative pangenome design, \nincluding both haplotypes of a sire and mapping long sequencing reads from the germline, \ne.g. sperm [46, 47]. Aligning long reads to this pangenome could detect the large number of \nrecombination events present in sperm, similar to the path congruence analysis presented \nhere, noting where reads change onto a different haplotype path. In the intermediate future, \ncattle tend to have highly related pedigrees due to the extensive use of breeding bulls, and so \nreanalysis of existing male long read cohorts with known pedigree (e.g., [48]) may already \nprovide evidence of SV-recombination interactions. \n \nMethylation, specifically 5mC, is also becoming increasingly accessible directly from long \nread sequencing. We find compelling evidence for several instances of allele-specific \nmethylation or joining phaseblocks in the absence of SNPs. However, we found the initial \ndifferences in methylation signal are dominated by genomic mutations disrupting CpG \nmotifs, rather than epigenetic-only change, requiring careful comparison across samples. \nFurthermore, the presence of intermediate methylation calls (i.e. between 30% and 70% \nmethylated) further confounds analyses. However, improvements in sequencing chemistry \nand basecalling have already increased accuracy of 5mC methylation calls and enable the use \nof other methylation markers, which may provide additional evidence of haplotype-of-origin \nto detect recombination events. \n \nConclusions \nWhile massive pedigrees and genotype arrays are useful for confidently identifying \nrecombination hotspots, we show related haplotype-resolved assemblies and pangenomes can \naccurately identify recombination events with high resolution. This approach also can \nidentify haplotype status in the absence of heterozygous SNPs by incorporating structural \nvariants and methylation signals, further improving the resolution of recombination maps. \nHowever, these approaches will benefit from further accuracy improvements for both \nassembly base-level quality and single CpG methylation scores to minimise false positive \nrecombination events. \n \nList of abbreviations \nASM: Allele-specific methylation \nCpG: Cytosine-phosphate-guanine \nDMS: Differentially methylated sites \nIBD: Identity by descent \nQV: Quality value (in context of base-level accuracy) \nRoH: Run of homozygosity \nSNP: Single nucleotide polymorphism \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 20, 2026. ; https://doi.org/10.64898/2026.02.20.706955doi: bioRxiv preprint \n\nSV: Structural variant \nUPGMA: Unweighted pair group method with arithmetic mean \n \nEthics approval and consent to participate \nThe sampling of blood was approved by the veterinary office of the Canton of Zurich (animal \nexperimentation permit ZH137/2023). \n \nConsent for publication \nNot applicable \n \nAvailability of data and materials \nSequencing data for four Simmental genomes are public (PRJEB42335, PRJEB72196, and \nPRJNA677947), and the short sequencing reads of the Simmental dam of the Eringer x \nSimmental F1 (SAMEA115121771). New short sequencing reads for the Eringer sire of the \nF1 and the F1 itself have been made public (SAMEA121233886 and SAMEA121233887 \nrespectively at study accession PRJEB28191), as well as the long sequencing of the F1 \n(SAMEA121647171 at study accession PRJEB42335). All workflows related to these \nanalyses are available at https://github.com/ASLeonard/simbling. \n \nCompeting interests \nThe authors declare that they have no competing interests. \n \nFunding \nThis study was supported by the Swiss National Science Foundation (SNSF; grant ID: \n204654). The funding bodies were neither involved in the design of the study and collection, \nanalysis, and interpretation of data nor in writing the manuscript. \n \nAuthors' contributions \nASL assembled the non-public genomes, constructed the pangenomes, and conducted all \nanalyses. ASL and HP wrote the manuscript. \n \nAcknowledgements \nWe thank Cord Drögemüller for providing access to the SWISS_1 Simmental genome \nassembly. We also thank Xena Mapel for extracting the F1 and sire DNA for the Eringer x \nSimmental cross and Naveen Kadri for useful discussions regarding expected meiotic \ncrossovers in half-siblings. \n \n \nReferences \n \n1. Peñalba J V ., Wolf JBW. From molecules to populations: appreciating and estimating \nrecombination rate variation. Nature Reviews Genetics 2020 21:8. 2020;21:476–92. \nhttps://doi.org/10.1038/s41576-020-0240-1. \n2. Weng Z-Q, Saatchi M, Schnabel RD, Taylor JF, Garrick DJ. Recombination locations and \nrates in beef cattle assessed from parent-offspring pairs. Genet Sel Evol. 2014;46:34. \nhttps://doi.org/10.1186/1297-9686-46-34. \n3. Paigen K, Petkov P. Mammalian recombination hot spots: properties, control and \nevolution. Nat Rev Genet. 2010;11:221–33. https://doi.org/10.1038/nrg2712. \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 20, 2026. ; https://doi.org/10.64898/2026.02.20.706955doi: bioRxiv preprint \n\n4. Villoutreix R, Ayala D, Joron M, Gompert Z, Feder JL, Nosil P. Inversion breakpoints and \nthe evolution of supergenes. Mol Ecol. 2021;30:2738–55. \nhttps://doi.org/10.1111/MEC.15907. \n5. Hassan S, Surakka I, Taskinen M-R, Salomaa V , Palotie A, Wessman M, et al. High-\nresolution population-specific recombination rates and their effect on phasing and genotype \nimputation. European Journal of Human Genetics. 2021;29:615–24. \nhttps://doi.org/10.1038/s41431-020-00768-8. \n6. Shen B, Jiang J, Seroussi E, Liu GE, Ma L. Characterization of recombination features and \nthe genetic basis in multiple cattle breeds. BMC Genomics. 2018;19:304. \nhttps://doi.org/10.1186/s12864-018-4705-y. \n7. Ma L, O’Connell JR, VanRaden PM, Shen B, Padhi A, Sun C, et al. Cattle Sex-Specific \nRecombination and Genetic Control from a Large Pedigree Analysis. PLoS Genet. \n2015;11:e1005387. https://doi.org/10.1371/JOURNAL.PGEN.1005387. \n8. Sandor C, Li W, Coppieters W, Druet T, Charlier C, Georges M. Genetic Variants in REC8, \nRNF212, and PRDM9 Influence Male Recombination in Cattle. PLoS Genet. \n2012;8:e1002854. https://doi.org/10.1371/JOURNAL.PGEN.1002854. \n9. Mouresan EF, González-Rodríguez A, Cañas-Álvarez JJ, Munilla S, Altarriba J, Díaz C, et \nal. Mapping Recombination Rate on the Autosomal Chromosomes Based on the Persistency \nof Linkage Disequilibrium Phase Among Autochthonous Beef Cattle Populations in Spain. \nFront Genet. 2019;10:437238. https://doi.org/10.3389/fgene.2019.01170. \n10. Yang L, Gao Y , Li M, Park K-E, Liu S, Kang X, et al. Genome-wide recombination map \nconstruction from single sperm sequencing in cattle. BMC Genomics. 2022;23:181. \nhttps://doi.org/10.1186/s12864-022-08415-w. \n11. Malinsky M, Talbi M, Zhou C, Maurer N, Sacco S, Shapiro B, et al. Hi-reComb: \nconstructing recombination maps from bulk gamete Hi-C sequencing. bioRxiv. \n2025;:2025.03.06.641907. https://doi.org/10.1101/2025.03.06.641907. \n12. Sun H, Rowan BA, Flood PJ, Brandt R, Fuss J, Hancock AM, et al. Linked-read \nsequencing of gametes allows efficient genome-wide analysis of meiotic recombination. \nNature Communications 2019 10:1. 2019;10:4310-. https://doi.org/10.1038/s41467-019-\n12209-2. \n13. Carvalho CMB, Lupski JR. Mechanisms underlying structural variant formation in \ngenomic disorders. Nature Reviews Genetics 2016 17:4. 2016;17:224–38. \nhttps://doi.org/10.1038/nrg.2015.25. \n14. Porubsky D, Höps W, Ashraf H, Hsieh PH, Rodriguez-Martin B, Yilmaz F, et al. \nRecurrent inversion polymorphisms in humans associate with genetic instability and genomic \ndisorders. Cell. 2022;185:1986-2005.e26. https://doi.org/10.1016/J.CELL.2022.04.017. \n15. Derbyshire MC, Newman TE, Khentry Y , Michael PJ, Bennett SJ, Rijal Lamichhane A, et \nal. Recombination and transposition drive genomic structural variation potentially impacting \nlife history traits in a host-generalist fungal plant pathogen. BMC Biology 2025 23:1. \n2025;23:110-. https://doi.org/10.1186/S12915-025-02179-X. \n16. Garrison E, Guarracino A, Heumos S, Villani F, Bao Z, Tattini L, et al. Building \npangenome graphs. Nature Methods 2024 21:11. 2024;21:2008–12. \nhttps://doi.org/10.1038/s41592-024-02430-3. \n17. Crysnanto D, Leonard AS, Fang ZH, Pausch H. Novel functional sequences uncovered \nthrough a bovine multiassembly graph. Proc Natl Acad Sci U S A. 2021;118:2101056118. \nhttps://doi.org/10.1073/pnas.2101056118. \n18. Leonard AS, Crysnanto D, Fang ZH, Heaton MP, Vander Ley BL, Herrera C, et al. \nStructural variant-based pangenome construction has low sensitivity to variability of \nhaplotype-resolved bovine assemblies. Nat Commun. 2022;13:1–13. \nhttps://doi.org/10.1038/s41467-022-30680-2. \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 20, 2026. ; https://doi.org/10.64898/2026.02.20.706955doi: bioRxiv preprint \n\n19. Milia S, Leonard AS, Mapel XM, Ulloa SMB, Drögemüller C, Pausch H. Taurine \npangenome uncovers a segmental duplication upstream of KIT associated with \ndepigmentation in white-headed cattle. Genome Res. 2025;35:1041–52. \nhttps://doi.org/10.1101/GR.279064.124. \n20. Leonard AS, Mapel XM, Pausch H. Pangenome genotyped structural variation improves \nmolecular phenotype mapping in cattle. Genome Res. 2024;:gr.278267.123. \nhttps://doi.org/10.1101/GR.278267.123. \n21. Bonizzoni P, Monti DC, Vedova G Della, Riccardi B, Rizzi R, Siren J. RecAlign: A* \nrecombination-aware sequence to graph mapping. bioRxiv. 2025;:2025.01.18.633308. \nhttps://doi.org/10.1101/2025.01.18.633308. \n22. Jafarzadeh N, Eizenga JM, Paten B. An Efficient Graph Algorithm for Diploid Local \nAncestry Inference. bioRxiv. 2025;:2025.07.05.662656. \nhttps://doi.org/10.1101/2025.07.05.662656. \n23. Fernandes JB, Naish M, Lian Q, Burns R, Tock AJ, Rabanal FA, et al. Structural variation \nand DNA methylation shape the centromere-proximal meiotic crossover landscape in \nArabidopsis. Genome Biol. 2024;25:30. https://doi.org/10.1186/s13059-024-03163-4. \n24. Casale F, Arlt C, Kühl M, Li J, Engelhorn J, Hartwig T, et al. The role of methylation and \nstructural variants in shaping the recombination landscape of barley. bioRxiv. \n2024;:2024.07.22.604552. https://doi.org/10.1101/2024.07.22.604552. \n25. Stack GM, Quade MA, Wilkerson DG, Monserrate LA, Bentz PC, Carey SB, et al. \nComparison of Recombination Rate, Reference Bias, and Unique Pangenomic Haplotypes in \nCannabis sativa Using Seven De Novo Genome Assemblies. Int J Mol Sci. 2025;26:1165. \nhttps://doi.org/10.3390/IJMS26031165/S1. \n26. Cochetel N, V ondras A, Figueroa-Balderas R, Liou J, Peluso P, Cantu D. Phased \nepigenomics and methylation inheritance in a historical Vitis vinifera hybrid. bioRxiv. \n2025;:2025.05.27.656431. https://doi.org/10.1101/2025.05.27.656431. \n27. Tullius TW, Heuer RA, Bohaczuk SC, Mallory B, Dubocanin D, Ranchalis J, et al. \nProtamine lacunae preserve the paternal chromatin landscape in sperm. bioRxiv. \n2025;:2025.10.03.680364. https://doi.org/10.1101/2025.10.03.680364. \n28. Fu Y , Aganezov S, Mahmoud M, Beaulaurier J, Juul S, Treangen TJ, et al. MethPhaser: \nmethylation-based long-read haplotype phasing of human genomes. Nat Commun. \n2024;15:5327. https://doi.org/10.1038/s41467-024-49588-0. \n29. Cheng H, Concepcion GT, Feng X, Zhang H, Li H. Haplotype-resolved de novo assembly \nusing phased assembly graphs with hifiasm. Nat Methods. 2021;18:170–5. \nhttps://doi.org/10.1038/s41592-020-01056-5. \n30. Huang N, Li H. compleasm: a faster and more accurate reimplementation of BUSCO. \nBioinformatics. 2023;39. https://doi.org/10.1093/BIOINFORMATICS/BTAD595. \n31. Rhie A, Walenz BP, Koren S, Phillippy AM. Merqury: reference-free quality, \ncompleteness, and phasing assessment for genome assemblies. Genome Biol. 2020;21:245. \nhttps://doi.org/10.1186/s13059-020-02134-9. \n32. Ondov BD, Treangen TJ, Melsted P, Mallonee AB, Bergman NH, Koren S, et al. Mash: \nfast genome and metagenome distance estimation using MinHash. Genome Biol. \n2016;17:132. https://doi.org/10.1186/s13059-016-0997-x. \n33. Virtanen P, Gommers R, Oliphant TE, Haberland M, Reddy T, Cournapeau D, et al. SciPy \n1.0: fundamental algorithms for scientific computing in Python. Nat Methods. 2020;17:261–\n72. https://doi.org/10.1038/s41592-019-0686-2. \n34. Guarracino A, Heumos S, Nahnsen S, Prins P, Garrison E. ODGI: understanding \npangenome graphs. Bioinformatics. 2022;38:3319–26. \nhttps://doi.org/10.1093/BIOINFORMATICS/BTAC308. \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 20, 2026. ; https://doi.org/10.64898/2026.02.20.706955doi: bioRxiv preprint \n\n35. Sahlin K. Strobealign: flexible seed size enables ultra-fast and accurate read alignment. \nGenome Biol. 2022;23:1–27. https://doi.org/10.1186/S13059-022-02831-7/TABLES/2. \n36. Yun T, Li H, Chang PC, Lin MF, Carroll A, McLean CY . Accurate, scalable cohort variant \ncalls using DeepVariant and GLnexus. Bioinformatics. 2020;36:5582–9. \nhttps://doi.org/10.1093/bioinformatics/btaa1081. \n37. Danecek P, Bonfield JK, Liddle J, Marshall J, Ohan V , Pollard MO, et al. Twelve years of \nSAMtools and BCFtools. Gigascience. 2021;10:1–4. \nhttps://doi.org/10.1093/gigascience/giab008. \n38. Eslami Rasekh M, Hernández Y , Drinan SD, Fuxman Bass JI, Benson G. Genome-wide \ncharacterization of human minisatellite VNTRs: population-specific alleles and gene \nexpression differences. Nucleic Acids Res. 2021;49:4308–24. \nhttps://doi.org/10.1093/NAR/GKAB224. \n39. Liao WW, Asri M, Ebler J, Doerr D, Haukness M, Hickey G, et al. A draft human \npangenome reference. Nature. 2023;617:312–24. https://doi.org/10.1038/s41586-023-05896-\nx. \n40. Sirén J, Monlong J, Chang X, Novak AM, Eizenga JM, Markello C, et al. Pangenomics \nenables genotyping of known structural variants in 5202 diverse genomes. Science (1979). \n2021;374. https://doi.org/10.1126/science.abg8871. \n41. Wick RR, Schultz MB, Zobel J, Holt KE. Bandage: interactive visualization of de novo \ngenome assemblies. Bioinformatics. 2015;31:3350–2. \nhttps://doi.org/10.1093/BIOINFORMATICS/BTV383. \n42. Leonard AS, Crysnanto D, Mapel XM, Bhati M, Pausch H. Graph construction method \nimpacts variation representation and analyses in a bovine super-pangenome. Genome Biol. \n2023;24:1–24. https://doi.org/10.1186/s13059-023-02969-y. \n43. Patel A, Horton JR, Wilson GG, Zhang X, Cheng X. Structural basis for human PRDM9 \naction at recombination hot spots. Genes Dev. 2016;30:257–65. \nhttps://doi.org/10.1101/GAD.274928.115. \n44. Bortoluzzi C, Mapel XM, Neuenschwander S, Janett F, Pausch H, Leonard AS. Genome \nassembly of wisent (Bison bonasus) uncovers a deletion that likely inactivates the THRSP \ngene. Commun Biol. 2024;7:1–10. https://doi.org/10.1038/s42003-024-07295-y. \n45. Denis E, Grohs C, Donnadieu C, Iampietro C. Validated DNA isolation method ensuring \nsuccessful long-read sequencing of cattle semen genome. PLoS One. 2024;19:e0308011. \nhttps://doi.org/10.1371/JOURNAL.PONE.0308011. \n46. Xie H, Li W, Guo Y , Su X, Chen K, Wen L, et al. Long-read-based single sperm genome \nsequencing for chromosome-wide haplotype phasing of both SNPs and SVs. Nucleic Acids \nRes. 2023;51:8020–34. https://doi.org/10.1093/NAR/GKAD532. \n47. Mapel XM, Leonard AS, Pausch H. Molecular QTL are enriched for structural variants in \na cattle long-read cohort. Commun Biol. 2026;:2025.05.16.654493. \nhttps://doi.org/10.1038/s42003-026-09596-w. \n  \n \n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted February 20, 2026. ; https://doi.org/10.64898/2026.02.20.706955doi: bioRxiv preprint","source_license":"CC-BY-4.0","license_restricted":false}