SRmdup: A new finer duplicate-read filter improves phylogenetic analyses of low-coverage ancient human DNA

preprint OA: closed CC-BY-4.0

Abstract

Abstract Duplicate-read filtering is a fundamental step in high-throughput sequencing data analysis. For ancient human DNA with low coverage, short fragments, and significant end damage, a minimum coverage threshold is routinely used to keep only reliably observed nucleotides. Too coarse duplicate filters can make this approach ineffective. We demonstrate that the problem has misled previous analyses of Neanderthal Y-chromosome data and suggest SRmdup, a new finer filter for single-end low-coverage DNA.
Full text 55,257 characters · extracted from preprint-html · click to expand
SRmdup: A new finer duplicate-read filter improves phylogenetic analyses of low-coverage ancient human DNA | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Short Report SRmdup: A new finer duplicate-read filter improves phylogenetic analyses of low-coverage ancient human DNA Yuhan Zhao, Stefan Grünewald This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9482713/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Duplicate-read filtering is a fundamental step in high-throughput sequencing data analysis. For ancient human DNA with low coverage, short fragments, and significant end damage, a minimum coverage threshold is routinely used to keep only reliably observed nucleotides. Too coarse duplicate filters can make this approach ineffective. We demonstrate that the problem has misled previous analyses of Neanderthal Y-chromosome data and suggest SRmdup, a new finer filter for single-end low-coverage DNA. Duplicate reads Ancient human DNA Phylogenetic reconstruction Duplicate filtering Figures Figure 1 Figure 2 Background In high-throughput sequencing analysis, duplicate read processing is a routine preprocessing step for many downstream statistical inferences. Its purpose is to prevent duplicates from being counted as independent observations in subsequent analyses. For modern high-coverage datasets, duplicate read remnants typically manifest as locally inflated coverage. However, in low-coverage ancient DNA, researchers often use a minimum coverage threshold. If duplicate reads are not adequately identified, some sites that should not have reached the coverage threshold may do so due to duplicated observations and enter subsequent statistical analysis and phylogenetic inference. Common duplicate filtering methods typically identify and label duplicate reads based on read alignment coordinates, retaining a representative from a group of reads suspected of originating from the same original fragment. The most widely applied tools are Picard MarkDuplicates [ 1 ] and rmdup which is available in SAMtools [ 2 ] but its refined version bam-rmdup from the biohazard-tools [ 3 ] has been used for many ancient DNA studies. Although these methods are widely and successfully used for standard data processing, the filtering is too coarse for ancient DNA. Rmdup only considers a pair of reads to be duplicates if both, the start and end coordinates are identical. MarkDuplicates only requires the 5-prime start coordinates (the start positions for forward reads and the end positions for reverse reads) to be the same. In addition, the methods handle soft-clipped bases differently. Here we analyse the distribution of read-start and read-end differences for pairs of reads. We propose to consider a pair of reads to be potential duplicates, if the 5-prime start or end coordinates are identical. Further, we include pairs where both the start and end positions differ by one. Results and Discussion We first establish that currently many duplicates remain undetected by investigating the collection of Y-chromosome reads of the low-coverage Neanderthal sample Spy 94a that was used for phylogenetic tree reconstruction in Petr et al. [ 4 ]. We use the visualization tool Integrative Genomics Viewer (IGV) [ 5 ] to show some duplicate candidates. Figure 1 compares the reads of the published BAM file with the filtered version of Picard MarkDuplicates and the method SRmdup presented here. Petr et al. [ 4 ] ran bam-rmdup as a preprocessing step. The first example (Fig. 1 A) shows a collection of nine reads, for which 8 have the same end position and similar but not identical start positions. Given that there are no other overlapping reads, it is clear that duplicates are present. However, MarkDuplicates keeps all reads. The second example (Fig. 1 B) shows how not filtering duplicates can affect phylogenetic reconstruction. There are initially four reads which all contain two mutated positions compared to the human reference genome. MarkDuplicates keeps two and SRmdup only one. While inspecting examples by naked eye is enough to realise that a finer duplicate filter is needed, we examined the distribution of read start and read end differences for all pairs of reads. Figure 2 A shows this distribution for all read pairs where the start positions or the end positions differ by at most 20. Given that the capture region of the data set contains almost 7 million positions, it is not surprising that pairs where either the start or the end position differs by more than 20 are frequent. However, the leftmost column (identical start position) and the bottommost row (identical end position) are clearly over-represented, with the exception of their intersection where preprocessing already removed most read pairs. The decay within the column respectively row is consistent with the example in Fig. 1 A. Figure 2 B shows only pairs of reads where the start and end position differences both differ by at least one and at most 20. We use a linear instead of a logarithmic scale and observe that the pair (1,1) is much more frequent than any other pair. We therefore define a pair of reads to be potential duplicates, if the start positions or the end positions are identical or if start positions and end positions both differ by one. An exact definition of duplicate sets is given in the Methods section, and Supplementary Table S2 lists the counts underlying Fig. 2 A and 2 B. Figure 2 | A Distribution of start position and end position differences for pairs of reads where at least one of the differences is at most 20 on a logarithmic scale. B Distribution of start position and end position differences for pairs of reads where both differences are at least 1 and at most 20 on a linear scale C Depth of the covered positions of the Y chromosome for the original data set (blue) and after applying MarkDuplicates (green) and SRmdup (red). D Read length distribution before and after running MarkDuplicates (red and blue) and proportion of deleted reads by length (green) In order to demonstrate the effect of duplicate filtering on a phylogenetic analysis, we revisit the data set for the Y chromosome from Petr et al. [ 4 ]. Their main result was that Neanderthals acquired the modern human Y chromosome, thus all available Neanderthal Y-chromosome sequences are more similar to modern human than to Denisova. In addition, a phylogenetic tree on all available samples was constructed, and within the Neanderthal branch, the samples El Sidrón 1253 and Mezmaiskaya 2 were shown as sister taxa while the third sample Spy 94a was drawn as an outgroup to them. Identifying the correct rooted triple for these taxa is difficult, because Spy 94a has low coverage (0.8), and El Sidrón 1253 has a much shorter capture region than the other two Neanderthals (560 kB versus 6.9MB) [ 6 , 7 ]. As a result, there are only slightly more than 45k positions where all taxa have coverage at least three. All but 61 of them are invariant within Neanderthals. There are 58 positions where Spy 94a carries an alternative allele while the other Neanderthals share the modern human nucleotide and two positions (6,932,831 and 6,932,839, see Fig. 1 B) where both El Sidrón 1253 and Mezmaiskaya 2 but not Spy 94a carry an alternative allele. This distribution is not plausible, and indeed SRmdup removes almost all positions with private mutations of Spy as well as both parsimony informative sites. The median networks [ 8 ] with the number of observed mutations labelling the edges before and after filtering, visualized by Splitstree4 [ 9 ], are shown in Supplementary Fig. S1 . The filtered network displays four mutations for all Neanderthals and one private mutation for Spy and one for El Sidrón 1253. This is realistic, but we have to conclude that there was not enough data to resolve the Neanderthal tree. More recently, more Neanderthal Y chromosomes were sequenced, including a high coverage one from Chagyrskaya Cave [ 10 ]. The paper shows a new phylogenetic tree which includes all taxa from Petr et al. [ 4 ] as well as four newly sequenced individuals from the same cave. Now Spy 94a and Mezmaiskaya 2 are more related to each other than to El Sidrón 1253. The authors point out the same two parsimony informative sites mentioned above, and they argue that modern human contamination is likely the reason for the error. However, a single contaminated read would not have passed the minimum coverage threshold, so recognizing the duplicates would have been sufficient. We show that the new tree is well supported by the data, we separately analyzed alignments containing nucleotides for the new high-quality sample Chagyrskaya B, Mezmaiskaya 2, the human reference Hg19, and either Spy 94a or El Sidrón 1253. Median networks before and after running SRmdup are shown in Supplementary Fig. S2. Again, the pending edge leading to Spy 94a is extremely long before filtering, and incompatible parsimony informative sites indicate that it might be hard to decide which quartet is correct. Filtering yields a clear situation supporting only the quartet (Spy 94a, Mezmaiskaya 2), (Hg19, Chagyrskaya B), and the longest edge is incident with Hg19. For the other taxa set, filtering is less important because of better coverage, and the supported quartet is (Chagyrskaya B, El Sidrón 1253), (Mezmaiskaya 2, Hg19). From the two quartets, the rooted tree on all four Neanderthal taxa with two sister groups (Spy 94a, Mezmaiskaya 2) and (Chagyrskaya B, El Sidrón 1253) can be inferred. For the new data set, Skov et al. [ 10 ] published a fully resolved Y-chromosome tree on all their taxa. However, there are no parsimony-informative positions to resolve the tree within the samples from Chagyrskaya Cave, neither before or after running SRmdup. The tree is constructed using maximum likelihood with a molecular clock assumption. Given that the genetic distances between Hg19 and each of the Chagyrskaya samples vary more than the maximum distance between Chagyrskaya samples, this assumption is clearly violated. The published tree (((Chagyrskaya B, Chagyrskaya D), Chagyrskaya K), Chagyrskaya G) simply orders the taxa by their distance to the human reference. When there are no parsimony informative sites supporting a cluster, a bias that causes triangle inequality violations for the genetic distance can have a strong influence on the likelihood scores for different tree topologies. Therefore, there is not enough data to determine the correct tree within the Chagyrskaya cluster. Our filter is also relevant for the main result of [ 10 ]. For example, there are 87 positions listed as “fixed derived in Spy 94a”. This number is much lower than the 577 in our median network in Supplementary Fig. S1 , because the authors use a strand filter to deal with deaminated reads: all positions carrying a C > T mutation where all reads are from the forward strand and all positions carrying a G > A mutation where all reads are on the reversed strand are removed. The 87 remaining positions have an unrealistically low transition/transversion ratio of 0.53. However, after running SRmdup, the number of variants only present is Spy 94a decreases to seven. Now the genetic insights into the social organization of Neanderthals that the title of [ 10 ] reports are inferred from the much lower estimates for coalescent times for pairs of Chagyrskaya individuals for the Y chromosome than for mtDNA. Using SRmdup as a preprocessing step alone will not change this difference. However, we observe that, after filtering out deaminated positions, low-coverage individuals tend to have much lower genetic distance to the human reference than high-coverage ones. This is true for the data set from [ 4 ], for the issues with the phylogenetic tree within the Chagyrskaya cluster mentioned above, and for the average estimated coalescent time for the 6 pairs of higher coverage individuals (962 years) compared to the 15 pairs within all included individuals (446 years) in SI Table 8.3 in [ 10 ]. A systematic bias can occur, if during the preprocessing steps from the raw data to the sequence alignment, reads that are identical with the human reference are less likely to be filtered out than reads carrying one or more mutations. Candidates are the thresholds for the edit distance (useful to address contamination) and for the mapping quality. Conclusions This paper demonstrates that in previous publications duplicate reads for low-coverage ancient human DNA have not been filtered out sufficiently and this has a negative effect on downstream analyses, especially phylogenetic tree reconstruction. Our new method SRmdup has been shown to be useful for analysing Neanderthal Y-chromosome data. For high-coverage samples our filter will have a significant number of false positives (independent pairs of reads that fulfil our duplicate criteria). This will be a much smaller problem than the false negatives of coarser filters, but a model-based approach that computes duplicate probabilities for candidate sets of reads might address this issue. We also focus on single-end reads which tend to dominate ancient DNA data sets. In order to include paired reads, additional information needs to considered, but a modification of our method is feasible. Finally, our observations of uncorrected P-distances that are far away from consistent with any tree suggest that the preprocessing from the raw data to called genotypes might contain important biases for downstream analyses. In order to avoid wrong conclusions from invaluable data, we encourage the community to investigate how strong such biases are and how they can be prevented or corrected. Methods Duplicate comparisons were defined using corrected read starts and ends on hg19 coordinates derived from the input BAM file. Soft-clipped bases were counted as part of the read. For each read \(\:r\) , \(\:c\left(r\right)\) denotes the reference sequence, \(\:\sigma\:\left(r\right)\in\:\{+,-\}\) the strand, and \(\:I\left(r\right)\) ​ \(\:=\) ​ \(\:\left[a\left(r\right),b\left(r\right)\right]\:\) the corrected fragment interval, where \(\:a\left(r\right)\) and \(\:b\left(r\right)\) are the corrected start and end coordinates, respectively. Comparisons were restricted to reads with the same strand \(\:\sigma\:\left({r}_{i}\right)\) ​ \(\:=\) ​ \(\:\sigma\:\left({r}_{j}\right)\) . We considered four classes of duplicate relationships. First, reads were classified as exact-boundary duplicates when both corrected boundaries matched, i.e. \(\:a\left({r}_{i}\right)\) ​ \(\:=\) ​ \(\:a\left({r}_{j}\right)\) and \(\:b\left({r}_{i}\right)\) ​ \(\:=\) ​ \(\:b\left({r}_{j}\right)\) . Second, reads were classified as start-matched duplicates when they shared the same corrected start but had different corrected ends, i.e. \(\:a\left({r}_{i}\right)\) ​ \(\:=\) ​ \(\:a\left({r}_{j}\right)\) and \(\:b\left({r}_{i}\right)\ne\:b\left({r}_{j}\right)\) . Third, reads were classified as end-matched duplicates when they shared the same corrected end but had different corrected starts, i.e. \(\:b\left({r}_{i}\right)\) ​ \(\:=\) ​ \(\:b\left({r}_{j}\right)\) and \(\:a\left({r}_{i}\right)\ne\:a\left({r}_{j}\right)\) . Fourth, reads were classified as nearest neighbors when both corrected boundaries differed by no more than 1 bp, i.e. \(\:\mid\:a\left({r}_{i}\right)-a\left({r}_{j}\right)\mid\:=1\) and \(\:\mid\:b\left({r}_{i}\right)-b\left({r}_{j}\right)\mid\:=1\) . Together, these rules capture three levels of duplicate similarity: complete agreement at both boundaries, agreement at one boundary with displacement at the other, and small offsets at both boundaries. This formulation accommodates limited endpoint variation introduced by end repair, soft clipping, and alignment ambiguity, which are particularly relevant in low-coverage ancient DNA data. Duplicate sets were defined as the connected components of the graph induced by these pairwise rules, ensuring that set membership was independent of processing order. Within each duplicate set, one representative read was retained. Representative selection was prioritized by highest total base quality, then by longest effective length; remaining ties were resolved by retaining the read encountered first in the BAM file. This procedure collapses redundant observations while preferentially preserving the read with the most complete and reliable information for downstream analyses. Declarations Data availability The datasets supporting the conclusions of this article are available in the European Nucleotide Archive (ENA) repository under accession number PRJEB39390 and in the repository of the Max Planck Institute for Evolutionary Anthropology: http://ftp.eva.mpg.de/neandertal/ChagyrskayaOkladnikov/. Only the Y-chromosome BAM files were used in this study. No new raw sequencing data was generated in this study. The source code and processed files supporting the conclusions of this article are available at https://github.com/Zhao-YHan/SRmdup/ repository. Ethics approval and consent to participate Not applicable. Consent for publication Not applicable. Competing interests The authors declare that they have no competing interests. Funding Not applicable. Authors' contributions S.G. designed and supervised the project. Y.Z. performed the data analyses and wrote the scripts. Y.Z. and S.G. wrote the manuscript. Acknowledgements Not applicable. References Picard toolkit. [ https://github.com/broadinstitute/picard] Li H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, Marth G, Abecasis G, Durbin R, Subgroup GPDP. The Sequence Alignment/Map format and SAMtools. Bioinformatics. 2009;25:2078–9. Bam-rmdup. [ https://github.com/mpieva/biohazard-tools/] Petr M, Hajdinjak M, Fu Q, Essel E, Rougier H, Crevecoeur I, Semal P, Golovanova LV, Doronichev VB, Lalueza-Fox C, et al. The evolutionary history of Neanderthal and Denisovan Y chromosomes. Science. 2020;369:1653–6. Thorvaldsdóttir H, Robinson JT, Mesirov JP. Integrative Genomics Viewer (IGV): high-performance genomics data visualization and exploration. Brief Bioinform. 2013;14:178–92. Lippold S, Xu H, Ko A, Li M, Renaud G, Butthof A, Schröder R, Stoneking M. Human paternal and maternal demographic histories: insights from high-resolution Y chromosome and mtDNA sequences. Investig Genet. 2014;5:13. Wood R, Higham T, Torres T, Tisnérat-Laborde N, Valladas H, Ortiz J, Lalueza-Fox C, Sanchez-Moral S, Cañaveras J, Rosas A, et al. A new date for the Neanderthals from El Sidrón Cave (Asturias, northern Spain). Archaeometry. 2013;55:148–58. Bandelt HJ, Forster P, Sykes BC, Richards MB. Mitochondrial portraits of human populations using median networks. Genetics. 1995;141:743–53. Huson DH, Bryant D. Application of phylogenetic networks in evolutionary studies. Mol Biol Evol. 2006;23:254–67. Skov L, Peyrégne S, Popli D, Iasi LNM, Devièse T, Slon V, Zavala EI, Hajdinjak M, Sümer AP, Grote S, et al. Genetic insights into the social organization of Neanderthals. Nature. 2022;610:519–25. Additional Declarations No competing interests reported. Supplementary Files SupplementarymaterialofSRmdupAnewfinerduplicatereadfilterimprovesphylogeneticanalysesoflowcoverageancienthumanDNA.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9482713","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Short Report","associatedPublications":[],"authors":[{"id":627825112,"identity":"8005f862-63a0-477d-8dd5-d52dd204382f","order_by":0,"name":"Yuhan Zhao","email":"","orcid":"","institution":"Shanghai Institute of Nutrition and Health","correspondingAuthor":false,"prefix":"","firstName":"Yuhan","middleName":"","lastName":"Zhao","suffix":""},{"id":627825113,"identity":"a6cea191-b2ca-4dcf-90d5-725795314349","order_by":1,"name":"Stefan Grünewald","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAyElEQVRIiWNgGAWjYBAC+QYGhg8MDDYMbAw8RGoxOMbAOIOBIY0ULWxgLYeBTKK1yDcfbPi543weH/vZAww/ahjkzQlpkW9jS2zsPXO7mI0nL4Gx5xiD4c4GQnqO8Zg/4G27ndgmwWPAwNvAkGBwgKAW/o+Nf9vOgbUw/iVOCw9jM2/bAbAWZqJsMTiWZtgs25YM9sthmWMShhsIaZFvPvyw8W2bXZ58+9mDD9/U2MgTdhgUJIAIoGIJItXDtIyCUTAKRsEowAoAQvw+oh41GVoAAAAASUVORK5CYII=","orcid":"","institution":"Shanghai Institute of Nutrition and Health","correspondingAuthor":true,"prefix":"","firstName":"Stefan","middleName":"","lastName":"Grünewald","suffix":""}],"badges":[],"createdAt":"2026-04-21 10:24:47","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9482713/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9482713/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":107978402,"identity":"93bc8eeb-2b32-47b8-a365-5ac24cfe305d","added_by":"auto","created_at":"2026-04-28 08:04:25","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":6180727,"visible":true,"origin":"","legend":"\u003cp\u003eResidual duplicates in low-coverage ancient DNA persist after standard filtering, and can affect phylogenetically informative sites.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-9482713/v1/ddbd984b0dce46644ca95ee8.png"},{"id":107978401,"identity":"6d3a78d2-464b-446e-8e9c-2296a8899e54","added_by":"auto","created_at":"2026-04-28 08:04:25","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":5914983,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eA \u003c/strong\u003eDistribution of start position and end position differences for pairs of reads where at least one of the differences is at most 20 on a logarithmic scale. \u003cstrong\u003eB \u003c/strong\u003eDistribution of start position and end position differences for pairs of reads where both differences are at least 1 and at most 20 on a linear scale \u003cstrong\u003eC \u003c/strong\u003eDepth of the covered positions of the Y chromosome for the original data set (blue) and after applying MarkDuplicates (green) and SRmdup (red). \u003cstrong\u003eD \u003c/strong\u003eRead length distribution before and after running MarkDuplicates (red and blue) and proportion of deleted reads by length (green)\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-9482713/v1/72012be443c0346cea65e468.png"},{"id":108123729,"identity":"7039cb72-4460-4bbf-a5f7-545c8838bba5","added_by":"auto","created_at":"2026-04-29 14:55:57","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":12259323,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9482713/v1/2d6dccc8-a891-463d-bb32-e75ac2e3e2a9.pdf"},{"id":107978400,"identity":"2a39e59a-a0cd-4843-8619-e24154c4d049","added_by":"auto","created_at":"2026-04-28 08:04:25","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":1597740,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementarymaterialofSRmdupAnewfinerduplicatereadfilterimprovesphylogeneticanalysesoflowcoverageancienthumanDNA.docx","url":"https://assets-eu.researchsquare.com/files/rs-9482713/v1/97fa58392ca287ba17a886f3.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"SRmdup: A new finer duplicate-read filter improves phylogenetic analyses of low-coverage ancient human DNA","fulltext":[{"header":"Background","content":"\u003cp\u003eIn high-throughput sequencing analysis, duplicate read processing is a routine preprocessing step for many downstream statistical inferences. Its purpose is to prevent duplicates from being counted as independent observations in subsequent analyses. For modern high-coverage datasets, duplicate read remnants typically manifest as locally inflated coverage. However, in low-coverage ancient DNA, researchers often use a minimum coverage threshold. If duplicate reads are not adequately identified, some sites that should not have reached the coverage threshold may do so due to duplicated observations and enter subsequent statistical analysis and phylogenetic inference.\u003c/p\u003e \u003cp\u003eCommon duplicate filtering methods typically identify and label duplicate reads based on read alignment coordinates, retaining a representative from a group of reads suspected of originating from the same original fragment. The most widely applied tools are Picard MarkDuplicates [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e] and rmdup which is available in SAMtools [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e] but its refined version bam-rmdup from the biohazard-tools [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e] has been used for many ancient DNA studies. Although these methods are widely and successfully used for standard data processing, the filtering is too coarse for ancient DNA.\u003c/p\u003e \u003cp\u003eRmdup only considers a pair of reads to be duplicates if both, the start and end coordinates are identical. MarkDuplicates only requires the 5-prime start coordinates (the start positions for forward reads and the end positions for reverse reads) to be the same. In addition, the methods handle soft-clipped bases differently.\u003c/p\u003e \u003cp\u003eHere we analyse the distribution of read-start and read-end differences for pairs of reads. We propose to consider a pair of reads to be potential duplicates, if the 5-prime start or end coordinates are identical. Further, we include pairs where both the start and end positions differ by one.\u003c/p\u003e"},{"header":"Results and Discussion","content":"\u003cp\u003eWe first establish that currently many duplicates remain undetected by investigating the collection of Y-chromosome reads of the low-coverage Neanderthal sample Spy 94a that was used for phylogenetic tree reconstruction in Petr et al. [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. We use the visualization tool Integrative Genomics Viewer (IGV) [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e] to show some duplicate candidates. Figure\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e compares the reads of the published BAM file with the filtered version of Picard MarkDuplicates and the method SRmdup presented here. Petr et al. [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e] ran bam-rmdup as a preprocessing step.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe first example (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eA) shows a collection of nine reads, for which 8 have the same end position and similar but not identical start positions. Given that there are no other overlapping reads, it is clear that duplicates are present. However, MarkDuplicates keeps all reads. The second example (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eB) shows how not filtering duplicates can affect phylogenetic reconstruction. There are initially four reads which all contain two mutated positions compared to the human reference genome. MarkDuplicates keeps two and SRmdup only one.\u003c/p\u003e \u003cp\u003eWhile inspecting examples by naked eye is enough to realise that a finer duplicate filter is needed, we examined the distribution of read start and read end differences for all pairs of reads. Figure\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA shows this distribution for all read pairs where the start positions or the end positions differ by at most 20. Given that the capture region of the data set contains almost 7\u0026nbsp;million positions, it is not surprising that pairs where either the start or the end position differs by more than 20 are frequent. However, the leftmost column (identical start position) and the bottommost row (identical end position) are clearly over-represented, with the exception of their intersection where preprocessing already removed most read pairs. The decay within the column respectively row is consistent with the example in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eA.\u003c/p\u003e \u003cp\u003eFigure\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB shows only pairs of reads where the start and end position differences both differ by at least one and at most 20. We use a linear instead of a logarithmic scale and observe that the pair (1,1) is much more frequent than any other pair. We therefore define a pair of reads to be potential duplicates, if the start positions or the end positions are identical or if start positions and end positions both differ by one. An exact definition of duplicate sets is given in the Methods section, and Supplementary Table S2 lists the counts underlying Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA and \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e | \u003cb\u003eA\u003c/b\u003e Distribution of start position and end position differences for pairs of reads where at least one of the differences is at most 20 on a logarithmic scale. \u003cb\u003eB\u003c/b\u003e Distribution of start position and end position differences for pairs of reads where both differences are at least 1 and at most 20 on a linear scale \u003cb\u003eC\u003c/b\u003e Depth of the covered positions of the Y chromosome for the original data set (blue) and after applying MarkDuplicates (green) and SRmdup (red). \u003cb\u003eD\u003c/b\u003e Read length distribution before and after running MarkDuplicates (red and blue) and proportion of deleted reads by length (green)\u003c/p\u003e \u003cp\u003eIn order to demonstrate the effect of duplicate filtering on a phylogenetic analysis, we revisit the data set for the Y chromosome from Petr et al. [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. Their main result was that Neanderthals acquired the modern human Y chromosome, thus all available Neanderthal Y-chromosome sequences are more similar to modern human than to Denisova. In addition, a phylogenetic tree on all available samples was constructed, and within the Neanderthal branch, the samples El Sidr\u0026oacute;n 1253 and Mezmaiskaya 2 were shown as sister taxa while the third sample Spy 94a was drawn as an outgroup to them.\u003c/p\u003e \u003cp\u003eIdentifying the correct rooted triple for these taxa is difficult, because Spy 94a has low coverage (0.8), and El Sidr\u0026oacute;n 1253 has a much shorter capture region than the other two Neanderthals (560 kB versus 6.9MB) [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. As a result, there are only slightly more than 45k positions where all taxa have coverage at least three. All but 61 of them are invariant within Neanderthals. There are 58 positions where Spy 94a carries an alternative allele while the other Neanderthals share the modern human nucleotide and two positions (6,932,831 and 6,932,839, see Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eB) where both El Sidr\u0026oacute;n 1253 and Mezmaiskaya 2 but not Spy 94a carry an alternative allele. This distribution is not plausible, and indeed SRmdup removes almost all positions with private mutations of Spy as well as both parsimony informative sites. The median networks [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e] with the number of observed mutations labelling the edges before and after filtering, visualized by Splitstree4 [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e], are shown in Supplementary Fig. \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e. The filtered network displays four mutations for all Neanderthals and one private mutation for Spy and one for El Sidr\u0026oacute;n 1253. This is realistic, but we have to conclude that there was not enough data to resolve the Neanderthal tree.\u003c/p\u003e \u003cp\u003eMore recently, more Neanderthal Y chromosomes were sequenced, including a high coverage one from Chagyrskaya Cave [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. The paper shows a new phylogenetic tree which includes all taxa from Petr et al. [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e] as well as four newly sequenced individuals from the same cave. Now Spy 94a and Mezmaiskaya 2 are more related to each other than to El Sidr\u0026oacute;n 1253. The authors point out the same two parsimony informative sites mentioned above, and they argue that modern human contamination is likely the reason for the error. However, a single contaminated read would not have passed the minimum coverage threshold, so recognizing the duplicates would have been sufficient.\u003c/p\u003e \u003cp\u003eWe show that the new tree is well supported by the data, we separately analyzed alignments containing nucleotides for the new high-quality sample Chagyrskaya B, Mezmaiskaya 2, the human reference Hg19, and either Spy 94a or El Sidr\u0026oacute;n 1253. Median networks before and after running SRmdup are shown in Supplementary Fig. S2. Again, the pending edge leading to Spy 94a is extremely long before filtering, and incompatible parsimony informative sites indicate that it might be hard to decide which quartet is correct. Filtering yields a clear situation supporting only the quartet (Spy 94a, Mezmaiskaya 2), (Hg19, Chagyrskaya B), and the longest edge is incident with Hg19. For the other taxa set, filtering is less important because of better coverage, and the supported quartet is (Chagyrskaya B, El Sidr\u0026oacute;n 1253), (Mezmaiskaya 2, Hg19). From the two quartets, the rooted tree on all four Neanderthal taxa with two sister groups (Spy 94a, Mezmaiskaya 2) and (Chagyrskaya B, El Sidr\u0026oacute;n 1253) can be inferred.\u003c/p\u003e \u003cp\u003eFor the new data set, Skov et al. [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e] published a fully resolved Y-chromosome tree on all their taxa. However, there are no parsimony-informative positions to resolve the tree within the samples from Chagyrskaya Cave, neither before or after running SRmdup. The tree is constructed using maximum likelihood with a molecular clock assumption. Given that the genetic distances between Hg19 and each of the Chagyrskaya samples vary more than the maximum distance between Chagyrskaya samples, this assumption is clearly violated. The published tree (((Chagyrskaya B, Chagyrskaya D), Chagyrskaya K), Chagyrskaya G) simply orders the taxa by their distance to the human reference. When there are no parsimony informative sites supporting a cluster, a bias that causes triangle inequality violations for the genetic distance can have a strong influence on the likelihood scores for different tree topologies. Therefore, there is not enough data to determine the correct tree within the Chagyrskaya cluster.\u003c/p\u003e \u003cp\u003eOur filter is also relevant for the main result of [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. For example, there are 87 positions listed as \u0026ldquo;fixed derived in Spy 94a\u0026rdquo;. This number is much lower than the 577 in our median network in Supplementary Fig. \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e, because the authors use a strand filter to deal with deaminated reads: all positions carrying a C\u0026thinsp;\u0026gt;\u0026thinsp;T mutation where all reads are from the forward strand and all positions carrying a G\u0026thinsp;\u0026gt;\u0026thinsp;A mutation where all reads are on the reversed strand are removed. The 87 remaining positions have an unrealistically low transition/transversion ratio of 0.53. However, after running SRmdup, the number of variants only present is Spy 94a decreases to seven.\u003c/p\u003e \u003cp\u003eNow the genetic insights into the social organization of Neanderthals that the title of [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e] reports are inferred from the much lower estimates for coalescent times for pairs of Chagyrskaya individuals for the Y chromosome than for mtDNA. Using SRmdup as a preprocessing step alone will not change this difference. However, we observe that, after filtering out deaminated positions, low-coverage individuals tend to have much lower genetic distance to the human reference than high-coverage ones. This is true for the data set from [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e], for the issues with the phylogenetic tree within the Chagyrskaya cluster mentioned above, and for the average estimated coalescent time for the 6 pairs of higher coverage individuals (962 years) compared to the 15 pairs within all included individuals (446 years) in SI Table\u0026nbsp;8.3 in [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. A systematic bias can occur, if during the preprocessing steps from the raw data to the sequence alignment, reads that are identical with the human reference are less likely to be filtered out than reads carrying one or more mutations. Candidates are the thresholds for the edit distance (useful to address contamination) and for the mapping quality.\u003c/p\u003e"},{"header":"Conclusions","content":"\u003cp\u003eThis paper demonstrates that in previous publications duplicate reads for low-coverage ancient human DNA have not been filtered out sufficiently and this has a negative effect on downstream analyses, especially phylogenetic tree reconstruction. Our new method SRmdup has been shown to be useful for analysing Neanderthal Y-chromosome data.\u003c/p\u003e \u003cp\u003eFor high-coverage samples our filter will have a significant number of false positives (independent pairs of reads that fulfil our duplicate criteria). This will be a much smaller problem than the false negatives of coarser filters, but a model-based approach that computes duplicate probabilities for candidate sets of reads might address this issue. We also focus on single-end reads which tend to dominate ancient DNA data sets. In order to include paired reads, additional information needs to considered, but a modification of our method is feasible.\u003c/p\u003e \u003cp\u003eFinally, our observations of uncorrected P-distances that are far away from consistent with any tree suggest that the preprocessing from the raw data to called genotypes might contain important biases for downstream analyses. In order to avoid wrong conclusions from invaluable data, we encourage the community to investigate how strong such biases are and how they can be prevented or corrected.\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003eDuplicate comparisons were defined using corrected read starts and ends on hg19 coordinates derived from the input BAM file. Soft-clipped bases were counted as part of the read. For each read \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:r\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:c\\left(r\\right)\\)\u003c/span\u003e\u003c/span\u003e denotes the reference sequence, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\sigma\\:\\left(r\\right)\\in\\:\\{+,-\\}\\)\u003c/span\u003e\u003c/span\u003e the strand, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:I\\left(r\\right)\\)\u003c/span\u003e\u003c/span\u003e​\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:=\\)\u003c/span\u003e\u003c/span\u003e​\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\left[a\\left(r\\right),b\\left(r\\right)\\right]\\:\\)\u003c/span\u003e\u003c/span\u003ethe corrected fragment interval, where \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:a\\left(r\\right)\\)\u003c/span\u003e\u003c/span\u003eand \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:b\\left(r\\right)\\)\u003c/span\u003e\u003c/span\u003eare the corrected start and end coordinates, respectively. Comparisons were restricted to reads with the same strand \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\sigma\\:\\left({r}_{i}\\right)\\)\u003c/span\u003e\u003c/span\u003e​\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:=\\)\u003c/span\u003e\u003c/span\u003e​\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\sigma\\:\\left({r}_{j}\\right)\\)\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003cp\u003eWe considered four classes of duplicate relationships. First, reads were classified as exact-boundary duplicates when both corrected boundaries matched, i.e. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:a\\left({r}_{i}\\right)\\)\u003c/span\u003e\u003c/span\u003e​\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:=\\)\u003c/span\u003e\u003c/span\u003e​\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:a\\left({r}_{j}\\right)\\)\u003c/span\u003e\u003c/span\u003eand \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:b\\left({r}_{i}\\right)\\)\u003c/span\u003e\u003c/span\u003e​\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:=\\)\u003c/span\u003e\u003c/span\u003e​\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:b\\left({r}_{j}\\right)\\)\u003c/span\u003e\u003c/span\u003e. Second, reads were classified as start-matched duplicates when they shared the same corrected start but had different corrected ends, i.e. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:a\\left({r}_{i}\\right)\\)\u003c/span\u003e\u003c/span\u003e​\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:=\\)\u003c/span\u003e\u003c/span\u003e​\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:a\\left({r}_{j}\\right)\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:b\\left({r}_{i}\\right)\\ne\\:b\\left({r}_{j}\\right)\\)\u003c/span\u003e\u003c/span\u003e. Third, reads were classified as end-matched duplicates when they shared the same corrected end but had different corrected starts, i.e. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:b\\left({r}_{i}\\right)\\)\u003c/span\u003e\u003c/span\u003e​\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:=\\)\u003c/span\u003e\u003c/span\u003e​\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:b\\left({r}_{j}\\right)\\)\u003c/span\u003e\u003c/span\u003eand \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:a\\left({r}_{i}\\right)\\ne\\:a\\left({r}_{j}\\right)\\)\u003c/span\u003e\u003c/span\u003e. Fourth, reads were classified as nearest neighbors when both corrected boundaries differed by no more than 1 bp, i.e. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\mid\\:a\\left({r}_{i}\\right)-a\\left({r}_{j}\\right)\\mid\\:=1\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\mid\\:b\\left({r}_{i}\\right)-b\\left({r}_{j}\\right)\\mid\\:=1\\)\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003cp\u003eTogether, these rules capture three levels of duplicate similarity: complete agreement at both boundaries, agreement at one boundary with displacement at the other, and small offsets at both boundaries. This formulation accommodates limited endpoint variation introduced by end repair, soft clipping, and alignment ambiguity, which are particularly relevant in low-coverage ancient DNA data.\u003c/p\u003e \u003cp\u003eDuplicate sets were defined as the connected components of the graph induced by these pairwise rules, ensuring that set membership was independent of processing order. Within each duplicate set, one representative read was retained. Representative selection was prioritized by highest total base quality, then by longest effective length; remaining ties were resolved by retaining the read encountered first in the BAM file. This procedure collapses redundant observations while preferentially preserving the read with the most complete and reliable information for downstream analyses.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eData availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe datasets supporting the conclusions of this article are available in the European Nucleotide Archive (ENA) repository under accession number PRJEB39390 and in the repository of the Max Planck Institute for Evolutionary Anthropology: http://ftp.eva.mpg.de/neandertal/ChagyrskayaOkladnikov/. Only the Y-chromosome BAM files were used in this study. No new raw sequencing data was generated in this study. The source code and processed files supporting the conclusions of this article are available at https://github.com/Zhao-YHan/SRmdup/ repository.\u003c/p\u003e\u003cp\u003e\u003cem\u003eEthics approval and consent to participate\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eConsent for publication\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eCompeting interests\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eFunding\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eAuthors\u0026apos; contributions\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eS.G. designed and supervised the project. Y.Z. performed the data analyses and wrote the scripts. Y.Z. and S.G. wrote the manuscript.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eAcknowledgements\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003ePicard toolkit. [\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/broadinstitute/picard]\u003c/span\u003e\u003cspan address=\"https://github.com/broadinstitute/picard]\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, Marth G, Abecasis G, Durbin R, Subgroup GPDP. The Sequence Alignment/Map format and SAMtools. Bioinformatics. 2009;25:2078\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBam-rmdup. [\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/mpieva/biohazard-tools/]\u003c/span\u003e\u003cspan address=\"https://github.com/mpieva/biohazard-tools/]\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePetr M, Hajdinjak M, Fu Q, Essel E, Rougier H, Crevecoeur I, Semal P, Golovanova LV, Doronichev VB, Lalueza-Fox C, et al. The evolutionary history of Neanderthal and Denisovan Y chromosomes. Science. 2020;369:1653\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThorvaldsd\u0026oacute;ttir H, Robinson JT, Mesirov JP. Integrative Genomics Viewer (IGV): high-performance genomics data visualization and exploration. Brief Bioinform. 2013;14:178\u0026ndash;92.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLippold S, Xu H, Ko A, Li M, Renaud G, Butthof A, Schr\u0026ouml;der R, Stoneking M. Human paternal and maternal demographic histories: insights from high-resolution Y chromosome and mtDNA sequences. Investig Genet. 2014;5:13.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWood R, Higham T, Torres T, Tisn\u0026eacute;rat-Laborde N, Valladas H, Ortiz J, Lalueza-Fox C, Sanchez-Moral S, Ca\u0026ntilde;averas J, Rosas A, et al. A new date for the Neanderthals from El Sidr\u0026oacute;n Cave (Asturias, northern Spain). Archaeometry. 2013;55:148\u0026ndash;58.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBandelt HJ, Forster P, Sykes BC, Richards MB. Mitochondrial portraits of human populations using median networks. Genetics. 1995;141:743\u0026ndash;53.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuson DH, Bryant D. Application of phylogenetic networks in evolutionary studies. Mol Biol Evol. 2006;23:254\u0026ndash;67.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSkov L, Peyr\u0026eacute;gne S, Popli D, Iasi LNM, Devi\u0026egrave;se T, Slon V, Zavala EI, Hajdinjak M, S\u0026uuml;mer AP, Grote S, et al. Genetic insights into the social organization of Neanderthals. Nature. 2022;610:519\u0026ndash;25.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Duplicate reads, Ancient human DNA, Phylogenetic reconstruction, Duplicate filtering","lastPublishedDoi":"10.21203/rs.3.rs-9482713/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9482713/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eDuplicate-read filtering is a fundamental step in high-throughput sequencing data analysis. For ancient human DNA with low coverage, short fragments, and significant end damage, a minimum coverage threshold is routinely used to keep only reliably observed nucleotides. Too coarse duplicate filters can make this approach ineffective. We demonstrate that the problem has misled previous analyses of Neanderthal Y-chromosome data and suggest SRmdup, a new finer filter for single-end low-coverage DNA.\u003c/p\u003e","manuscriptTitle":"SRmdup: A new finer duplicate-read filter improves phylogenetic analyses of low-coverage ancient human DNA","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-04-28 08:04:15","doi":"10.21203/rs.3.rs-9482713/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"52cbfa5c-8248-4499-a3e0-697c9750c012","owner":[],"postedDate":"April 28th, 2026","published":true,"recentEditorialEvents":[{"type":"decision","content":"Rejected","date":"2026-04-29T14:41:59+00:00","index":"","fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-04-29T14:55:03+00:00","versionOfRecord":[],"versionCreatedAt":"2026-04-28 08:04:15","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9482713","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9482713","identity":"rs-9482713","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-06-05T02:00:03.366016+00:00
License: CC-BY-4.0