A novel DNA profiling method for individual sample matching and its applications to wildlife diagnostics

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract The use of DNA analysis to match individual biological samples is central to many wildlife diagnostic applications, and is particularly valuable in illegal wildlife trade investigations. However, although thousands of wildlife and plant species are affected by illegal trade, DNA profiling systems for individual identification exist for only a small number of taxa. This is primarily due to the need for extensive population reference data to enable conclusive interpretation of DNA profile matches, a key requirement in diagnostic contexts. Here, we introduce a novel approach that does not require reference data to interpret an individual profile match, and can be used with any set of genetic markers on a case-by-case basis. This approach, called ‘heterozygote analysis for reference-free probability assessment’ ( HARP ), is based on ultra-conservative match probability estimates using only heterozygote genotypes and assuming both alleles have a frequency of 0.5 in the population. HARP effectively eliminates the need for reference data and standard marker panels, and is therefore completely adaptable to any novel individual identification scenario. We demonstrate HARP using three case studies spanning a range of species and diagnostic scenarios.
Full text 65,308 characters · extracted from preprint-html · click to expand
A novel DNA profiling method for individual sample matching and its applications to wildlife diagnostics | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article A novel DNA profiling method for individual sample matching and its applications to wildlife diagnostics Kyle M Ewart, Rob Ogden This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9106203/v1 This work is licensed under a CC BY 4.0 License Status: Under Revision Version 1 posted 9 You are reading this latest preprint version Abstract The use of DNA analysis to match individual biological samples is central to many wildlife diagnostic applications, and is particularly valuable in illegal wildlife trade investigations. However, although thousands of wildlife and plant species are affected by illegal trade, DNA profiling systems for individual identification exist for only a small number of taxa. This is primarily due to the need for extensive population reference data to enable conclusive interpretation of DNA profile matches, a key requirement in diagnostic contexts. Here, we introduce a novel approach that does not require reference data to interpret an individual profile match, and can be used with any set of genetic markers on a case-by-case basis. This approach, called ‘heterozygote analysis for reference-free probability assessment’ ( HARP ), is based on ultra-conservative match probability estimates using only heterozygote genotypes and assuming both alleles have a frequency of 0.5 in the population. HARP effectively eliminates the need for reference data and standard marker panels, and is therefore completely adaptable to any novel individual identification scenario. We demonstrate HARP using three case studies spanning a range of species and diagnostic scenarios. Biological sciences/Biological techniques Biological sciences/Ecology Earth and environmental sciences/Ecology Biological sciences/Genetics Biological sciences/Plant sciences DNA profiling match probabilities wildlife diagnostics Figures Figure 1 Introduction The use of DNA analysis to match individual biological samples underpins many diagnostic applications, and is routinely utilised in forensic science [ 1 – 3 ], authenticity testing [ 4 , 5 ], and conservation genetics [ 6 – 8 ]. A key application of DNA-based individual identification is detecting, monitoring and investigating the illegal wildlife trade. Wildlife trafficking is a major driver of biodiversity loss, affecting thousands of animal and plant species worldwide [ 9 , 10 ]. Many wildlife trade investigations hinge on establishing whether biological samples originate from the same individual organism or from different individuals. For example, authorities may need to determine whether a seized wildlife part (e.g. a horn or skin) originated from a particular carcass, whether biological material on a suspect (e.g. blood) derives from a poached animal, or whether products encountered at different stages of a supply chain correspond to the same source (e.g. linking timber to a tree stump). In addition to generating a DNA profiles through genotyping samples using a panel of genetic markers, conclusive and justifiable interpretation of any resultant DNA profile match is essential. Two samples with identical DNA profiles do not necessarily derive from the same individual; different individuals can share the same profile by chance, particularly when only a limited number of genetic markers are examined. Accordingly, in diagnostic contexts, individual identity cannot be inferred from a DNA profile match alone, but must also incorporate an assessment of the strength of the match given the markers’ characteristics and the observed profile. This typically involves estimating the probability of observing the profile in the general population, known as the genotype match probability. Estimating this probability has received significant statistical refinement over the past forty years [ 11 ], but has always required large-scale population genetic data for the specific DNA profiling system in use, creating a significant barrier to the development of novel DNA profiling systems. The requirement to generate DNA profiles from population-wide reference samples using a specific marker panel incurs considerable research and development costs for each species. Collecting reference samples poses numerous practical and ethical challenges, particularly for wide-ranging, rare, or dangerous species, or for those requiring special permits. Even when reference samples are available, the design and development of a taxon-specific marker panel requires considerable resources, which is often not economically viable for species that are not tested in large numbers. Further, transferring an established DNA profiling system and associated reference data to a new laboratory introduces additional challenges; reference data may be unsuitable if the target population changes, and the system may require calibration across laboratories (e.g., using standard genotyping samples) to ensure compatibility. Consequently, DNA-based individual identification is largely restricted to abundant model species (e.g. humans and domestic animals) or those expected to undergo routine testing, and is generally infeasible for rarer, less-studied taxa. These limitations are particularly evident in wildlife trade applications, where most trafficked species lack established DNA profiling systems and reference datasets, and many laboratories have no access to the few systems that do exist. Here, we present a novel approach to DNA profiling that does not require population genetic data or any specific marker panel to interpret a profile match, hence can be applied to any species, and as a general method, is easily transferable across laboratories. Traditional profiling systems comprise relatively few independent genetic markers (typically < 25), with ample power to differentiate individuals due to the rarity of the genotype observed at each marker. However, it is now routine to generate hundreds to thousands of genome-wide markers from a DNA sample. We build on this capability and propose an updated approach for estimating match probabilities, which we term ‘heterozygote analysis for reference-free probability assessment’ ( HARP ). HARP considers only heterozygous site matches and applies ultra-conservative assumptions: that each marker is biallelic, with alleles present at equal frequencies in the population (i.e. p = q =0.5). The resulting DNA profile match probabilities are generated without the need for specific marker panels or associated population reference data. We demonstrate the utility of HARP using three case studies. The adaptability of HARP across species and use cases has the potential to transform the use of individual DNA profiling in wildlife diagnostics. Results HARP assumptions and discriminatory power HARP match probabilities are calculated without reference data by only considering heterozygous loci (see Methods and Supplementary Methods for the full derivation). HARP assumes two alleles per heterozygote locus, both with a population frequency of 0.5. This is conservative: when allele frequencies differ from 0.5, or when more than two alleles exist in the population, match probabilities decrease (i.e. become more powerful). HARP is underpinned by two additional key assumptions. First, markers are assumed to be in linkage disequilibrium (LD) to justify calculation of cumulative probabilities across independent loci. This assumption can be met by filtering markers based on genomic location or by carrying out LD testing. Second, markers are assumed to be in Hardy–Weinberg equilibrium (HWE). A conventional approach to account for HWE deviations when calculating match probabilities is to incorporate a theta ( θ ) correction, which addresses heterozygote deficiency (e.g. as a result of population structure or individual relatedness). However, incorporating θ > 0 into HARP decreases the resultant match probability, hence we conservatively assume θ = 0. A less common issue, but one that may lead to underestimation of heterozygote match probabilities, is heterozygote excess. This can arise through selection scenarios such as heterozygote advantage, deleterious recessive alleles, advantageous dominant alleles and disassortative mating (elaborated in the Supplementary Methods). These scenarios are accounted for by incorporating a correction factor, s , representing the proportion of the markers affected by scenarios leading to heterozygote excess. Applying HARP to a single heterozygous marker, conservatively assuming s = 0.1 (see Supplementary Methods for justification), results in a match probability of ~ 0.536; i.e. there is a 53.6% chance of observing this same heterozygote genotype in the general population. This can also be expressed as a likelihood ratio (LR): when matching a sample to a suspected donor, observing this genotype is 1.866 times more likely if the sample derives from the donor rather than from an unrelated individual. HARP probabilities can be multiplied across multiple heterozygous loci provided they segregate independently, yielding considerable discriminatory power despite the conservative assumptions underpinning the calculation. For example, for 25 independent heterozygote markers, assuming s = 0.1, the probability that another individual in the population shares this heterozygote marker profile is 5.9 million. Case studies We applied HARP to a hypothetical hippopotamus poaching case, in which a carcass is discovered, a hippo tooth (ivory) is recovered from a vehicle, and suspected poachers are subsequently detained in possession of a knife and clothing, displaying traces of blood (Fig. 1 ). Heterozygote profiles from the four hippopotamus whole genome datasets generated from the same individual (using independent library preparations; Supplementary Table S1 ) matched across 21–25 heterozygote single nucleotide polymorphism markers (SNPs), with no mismatches, yielding match probabilities ranging from 2.04x10 − 6 to 1.69x10 − 7 (Fig. 1 ). Profiles from different individuals, including first-degree relatives, were successfully excluded. This case study demonstrates the application of HARP when no marker panel is available. Notably, we took stringent measures to ensure markers were in LD by utilising only single SNPs from the longest scaffolds of the hippopotamus genome. Thousands of genome-wide SNPs could potentially be utilized, which would dramatically increase discriminatory power (provided markers were appropriately filtered or LD was accounted for). Increasing the number of markers compared also increases the risk of false mismatches; however, these can be accounted for using error rate thresholds or incorporating a mismatch probability into the likelihood framework. We also applied HARP to two species with established SNP profiling systems, but analysing the datasets without using population reference data. A 144 SNP panel for snow leopards [ 8 ] yielded an expected 15 heterozygotes between matching samples (see Supplementary Methods), corresponding to a HARP match probability of 8.63x10 − 5 , and a 128 SNP panel for bigleaf maple [ 12 ] yielded an expected 40 heterozygotes between matching samples, corresponding to a HARP match probability of 1.46x10 − 11 (Table 1 ). These examples represent applications to conservation genetic monitoring and timber trade authenticity testing, respectively. Table 1 Expected number of heterozygote matches for the snow leopard and bigleaf maple case studies and their associated match probabilities. The number of heterozygotes for each case study is based on the missingness and observed heterozygosity reported for the associated profiling systems (see Supplementary Methods). We assumed s = 0.1 for the HARP match probability calculations. Case study species Marker panel Expected number of matching heterozygotes Match probability Snow leopard 144 SNP panel genotyped using high-throughput amplicon sequencing 15 8.63 x 10 − 5 Bigleaf maple 128 SNP panel genotyped using the MassARRAY® iPLEX™ platform 40 1.46 x 10 − 11 Discussion Although DNA-based individual identification is widely applicable across wildlife diagnostic contexts, it has been implemented for only a small number of species. This is primarily due to the substantial research and development required to develop species-specific marker panels and associated reference data. HARP addresses this barrier by enabling individual identification analysis without the need for these prerequisites, and can therefore be applied to any species for which DNA profiling systems have not been developed. HARP will be particularly beneficial in wildlife forensic applications, as illustrated in the hippopotamus poaching case study. A lack of testing capacity likely hinders prosecution in many cases that hinge on matching samples. Unlike conventional forensic DNA profiling, which requires marker validation, the whole genome approach used in the hippopotamus case study does not explicitly characterise individual markers. While this may raise instinctive concerns within the forensic genetics community, we believe that any associated risks can be effectively mitigated through stringent and extremely conservative bioinformatic pipelines that only retain highly reliable SNPs, minimise LD, and correct for markers with potential excess heterozygosity through the use of a conservative value for s . We encourage further discussions and research into how genomic approaches and the application of HARP can be validated for forensic individualisation analyses. Beyond forensics, HARP is broadly applicable to any diagnostic context requiring individualisation, and could transform how DNA-based diagnostic systems are designed and managed for product traceability. Audit systems conducting one-up-one-down supply chain verification for any species, but particularly those with high value individuals, such as timbers, could apply this method without prior species-specific research and development. For example, HARP could be applied to match logs to tree stumps in illegal logging compliance testing involving multiple timber species, where individual trees are harvested and labelled within a concession. HARP is agnostic to how heterozygote markers are generated. When existing marker panels are used to generate heterozygote profiles, no inter-laboratory calibration is required (often necessary to enable shared reference datasets), facilitating straightforward transfer of profiling systems across laboratories. A whole genome sequencing approach, as illustrated in the hippopotamus case study, does not require a specific marker panel, and could be applied to any species. This approach relies on a relatively high-quality reference genome to confidently call SNPs, minimise false heterozygote calls, and evaluate linkage disequilibrium (LD); such resources are increasingly available through global reference sequencing initiatives (e.g. DToL, ERGA, EarthBioGenome). The per sample sequencing effort required for this HARP whole genome approach depends on genome size and sequence quality, and could be substantially reduced through enrichment or genome complexity reduction methods (e.g. [ 13 – 15 ]). In summary, HARP has the potential to revolutionise DNA profiling for matching biological samples. By eliminating the need for species-specific development and reference data, HARP paves the way for a dramatic increase in novel DNA profiling applications across forensic science, diagnostics, and research. Methods HARP approach Here we derive and justify a per marker match probability without the need for a population genetic reference database. At each step, we take a conservative approach for this derivation. Assuming two alleles (A and B) occur at a given marker, the maximum probability, P , of observing a heterozygous (AB) genotype is P (AB) = 0.5. The corresponding probabilities for AA and BB genotypes are P (AA) = P (BB) = 1, making homozygous loci uninformative in the absence of population data. While a conservative heterozygote genotype probability of P (AB) = 0.5 is less powerful than those typically calculated using population reference data (i.e. if p or q is 2 alleles), when using a relatively large number of markers in combination, it is possible to recover sufficient heterozygote genotype matches to produce extremely small match probabilities. HARP assumes HWE, which is violated through either deficiency or excess in observed heterozygote genotypes. Only heterozygote excess is relevant to HARP and is accounted for by estimating the proportion of markers affected and introducing an associated correction factor, s , as follows: $$P\left(AB\right)={0.5}^{\left(n\times\left(1-s\right)\right)}$$ where n is the number of DNA markers with matching heterozygous genotypes in the profile. This equation holds even when comparing highly related individuals (e.g. first-degree relatives). HARP also assumes markers are not linked, which is addressed by filtering markers to minimise LD. See the Supplementary Methods for further details of the HARP assumptions and corrections. Case studies HARP was applied to a hypothetical hippopotamus poaching case. DNA profile comparisons were made between samples from hippopotamus ivory, trace evidence on the knife and clothing, the carcass, and two other hippopotamus individuals (Fig. 1 ). The ivory, trace evidence and carcass were derived from the same individual, while the two additional individuals formed a trio (sire, dam and offspring) to assess the discriminatory power of the approach among closely related individuals. Each evidence sample was represented by an independently re-sequenced hippo genome (Supplementary Table S1 ), obtained from two studies [ 16 , 17 ]. This case study was chosen for two key reasons: 1) it represents a realistic wildlife forensic application of HARP ; 2) it involves a species for which no marker panel or associated reference data are available. Whole genome sequencing data were available for the hippopotamus samples involved in this hypothetical case, providing a large set of heterozygous markers available for HARP analyses. Details on the bioinformatic pipeline for this case study is outlined in the Supplementary Methods and relevant GitHub repository ( https://github.com/KyleEwart/HARP_WGS_pipeline ). Briefly, we aligned these short-read datasets to a hippopotamus genome assembly (GCA_030028035.1), called SNPs, applied stringent filtering to remove low quality SNP and potential false heterozygotes, then compared heterozygote calls across individuals. We also applied HARP to two additional case studies involving species that have established DNA profiling systems: 1) a 128 SNP panel for bigleaf maple, used to assess trade authenticity of this commonly traded timber species [ 12 ], and 2) a 144 SNP panel for snow leopards, developed for conservation genetic applications [ 8 ] (Table 1 ). In these analyses, the existing systems were treated as if suitable reference datasets were unavailable (or unsuitable) in order to illustrate how HARP would perform under such conditions. The expected number of heterozygote matches between profiles was estimated from reported heterozygosity and missingness in the respective studies (elaborated in the Supplementary Methods). Declarations Author contributions K.M.E. and R.O. came up with the initial concept; K.M.E. and R.O. developed the method; K.M.E. performed the analyses; K.M.E. and R.O. wrote the manuscript. Funding statement We are grateful for the support of a Marie Skłodowska-Curie Actions (MSCA) Postdoctoral Fellowship (K.M.E.). Conflict of interests The authors have no competing interests to declare. Data availability statement Analyses carried out in this study are based on previously generated and openly accessible datasets. The relevant studies and accession numbers are available within the paper and its Supplementary Information. References Evett, I. W., Buffery, C., Willott, G. & Stoney, D. A guide to interpreting single locus profiles of DNA mixtures in forensic cases. J. Forensic Sci. Soc. 31, 41–47 (1991). Meester, R. & Slooten, K. Probability and forensic evidence: theory, philosophy and applications (Cambridge Univ. Press, Cambridge, 2021). Harper, C., et al. Robust forensic matching of confiscated horns to individual poached African rhinoceros. Curr. Biol. 28, R13-R14 (2018). Ewart, K. M., et al. TigerBase : A DNA registration system to enhance enforcement and compliance testing of captive tiger facilities. Forensic Sci. Int. Genet. 74, 103149 (2025). De Bruyn, M., Dalton, D. L., Mwale, M., Ehlers, K. &Kotze, A. Development and validation of a novel forensic STR multiplex assay for blue ( Anthropoides paradiseus ), wattled ( Bugeranus carunculatus ), and grey-crowned crane ( Balearica regulorum ). Forensic Sci. Int. Genet. 73, 103100 (2024). Kalinowski, S. T., Taper, M. L. & Creel, S. Using DNA from non-invasive samples to identify individuals and census populations: an evidential approach tolerant of genotyping errors. Conserv. Genet. 7, 319–329 (2006). Sethi, S. A., et al. Accurate recapture identification for genetic mark–recapture studies with error-tolerant likelihood-based match calling and sample clustering. R. Soc. Open Sci. 3, 160457 (2016). Solari, K. A., et al. Next-generation snow leopard population assessment tool: multiplex‐PCR SNP panel for individual identification from faeces. Mol. Ecol. Resour. 24, e14074 (2024). Scheffers, B. R., Oliveira, B. F., Lamb, I. & Edwards, D. P. Global wildlife trade across the tree of life. Science 366, 71–76 (2019). Hughes, L. J., Morton, O., Scheffers, B. R. & Edwards, D. P. The ecological drivers and consequences of wildlife trade. Biol. Rev. 98, 775–791 (2023). Evett, I. W. & Weir, B. S. Interpreting DNA Evidence. (Sinauer, Sunderland, 1998). Dormontt, E. E., et al. Forensic validation of a SNP and INDEL panel for individualisation of timber from bigleaf maple ( Acer macrophyllum Pursch). Forensic Sci. Int. Genet. 46, 102252 (2020). Davey, J. W. & Blaxter, M. L. RADSeq: next-generation population genetics. Brief. Funct. Genom. 9, 416–423 (2010). Loose, M., Malla, S. & Stout, M. Real-time selective sequencing using nanopore technology. Nat. Methods 13, 751–754 (2016). Sinn, B. T., et al. ISSRseq: An extensible method for reduced representation sequencing. Methods Ecol. Evol. 13, 668–681 (2022). Árnason, Ú., Lammers, F., Kumar, V., Nilsson, M. A. & Janke, A. Whole-genome sequencing of the blue whale and other rorquals finds signatures for introgressive gene flow. Sci. Adv. 4, eaap9873 (2018). Bergeron, L. A. et al. Evolution of the germline mutation rate across vertebrates. Nature 615, 285–291 (2023). Additional Declarations No competing interests reported. Supplementary Files NatComharpSupplementaryInformationfinal.docx Cite Share Download PDF Status: Under Revision Version 1 posted Editorial decision: Revision requested 21 Apr, 2026 Reviews received at journal 14 Apr, 2026 Reviews received at journal 09 Apr, 2026 Reviewers agreed at journal 17 Mar, 2026 Reviewers agreed at journal 16 Mar, 2026 Reviewers invited by journal 16 Mar, 2026 Editor assigned by journal 16 Mar, 2026 Submission checks completed at journal 15 Mar, 2026 First submitted to journal 12 Mar, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9106203","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":607736940,"identity":"1381abb9-4a6b-4c88-b0b8-e19ad4558c86","order_by":0,"name":"Kyle M Ewart","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA5UlEQVRIie3RMcrCMBTA8VcCcYno+Ipgr5Di4KReJaUQryAo2A+hLh4gQ8FDCM5KQRdv0EUXXRzqJghqEMEt+dwc8p/e8ns8EgCX6wdDIMl7JCtgr4HaiPcmlIqvCeP/I76K/vYwyIP2fHbZn6EbAEphJA2MJhx2eZhtqoswgzhMUK6MpIlRil5aeIpWlw0GRAD2ExuZXr170VOUnW4MxnaiD0v1CxSRJpQwyDWxHObPDhMUm0esqGz5Gd+GKTsKI8FtvC7Lkewokh/K82AY1CqSGwnU9crPVm79FV3NfLjL5XK5AJ5HsD9HvJZb+gAAAABJRU5ErkJggg==","orcid":"","institution":"University of Edinburgh","correspondingAuthor":true,"prefix":"","firstName":"Kyle","middleName":"M","lastName":"Ewart","suffix":""},{"id":607736947,"identity":"5ca85ded-518b-41c5-9748-54863a8cb6de","order_by":1,"name":"Rob Ogden","email":"","orcid":"","institution":"University of Edinburgh","correspondingAuthor":false,"prefix":"","firstName":"Rob","middleName":"","lastName":"Ogden","suffix":""}],"badges":[],"createdAt":"2026-03-12 14:54:41","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9106203/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9106203/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":104998443,"identity":"67f09dbe-7887-4f09-b180-53caacd791d5","added_by":"auto","created_at":"2026-03-19 16:26:46","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":745414,"visible":true,"origin":"","legend":"\u003cp\u003eThe number of matches (green) and mismatches (red) between samples in the hypothetical hippopotamus poaching case study (lower diagonal) and their associated match probabilities (upper diagonal). For hippo A, four independent datasets were used to simulate results from different evidence items (hippo carcass, hippo ivory and trace evidence on weapons and clothing). We assumed \u003cem\u003es \u003c/em\u003e= 0.1 for the \u003cem\u003eHARP\u003c/em\u003ematch probability calculation. For details of the whole genome datasets, see Supplementary Methods.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-9106203/v1/28637cc937a31a42a6118059.png"},{"id":104998466,"identity":"6c279f67-85e2-4992-a6d5-0be271a7cba1","added_by":"auto","created_at":"2026-03-19 16:26:51","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1160003,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9106203/v1/446b1dd0-860e-4f55-b8a1-5afd7e9efe56.pdf"},{"id":104998410,"identity":"a4114200-4f01-4980-834d-6e6ece05b3ae","added_by":"auto","created_at":"2026-03-19 16:26:33","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":35897,"visible":true,"origin":"","legend":"","description":"","filename":"NatComharpSupplementaryInformationfinal.docx","url":"https://assets-eu.researchsquare.com/files/rs-9106203/v1/c1fcf646852205a3869f0404.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"A novel DNA profiling method for individual sample matching and its applications to wildlife diagnostics","fulltext":[{"header":"Introduction","content":"\u003cp\u003eThe use of DNA analysis to match individual biological samples underpins many diagnostic applications, and is routinely utilised in forensic science [\u003cspan additionalcitationids=\"CR2\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e], authenticity testing [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e], and conservation genetics [\u003cspan additionalcitationids=\"CR7\" citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. A key application of DNA-based individual identification is detecting, monitoring and investigating the illegal wildlife trade. Wildlife trafficking is a major driver of biodiversity loss, affecting thousands of animal and plant species worldwide [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. Many wildlife trade investigations hinge on establishing whether biological samples originate from the same individual organism or from different individuals. For example, authorities may need to determine whether a seized wildlife part (e.g. a horn or skin) originated from a particular carcass, whether biological material on a suspect (e.g. blood) derives from a poached animal, or whether products encountered at different stages of a supply chain correspond to the same source (e.g. linking timber to a tree stump).\u003c/p\u003e \u003cp\u003eIn addition to generating a DNA profiles through genotyping samples using a panel of genetic markers, conclusive and justifiable interpretation of any resultant DNA profile match is essential. Two samples with identical DNA profiles do not necessarily derive from the same individual; different individuals can share the same profile by chance, particularly when only a limited number of genetic markers are examined. Accordingly, in diagnostic contexts, individual identity cannot be inferred from a DNA profile match alone, but must also incorporate an assessment of the strength of the match given the markers\u0026rsquo; characteristics and the observed profile. This typically involves estimating the probability of observing the profile in the general population, known as the genotype match probability. Estimating this probability has received significant statistical refinement over the past forty years [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e], but has always required large-scale population genetic data for the specific DNA profiling system in use, creating a significant barrier to the development of novel DNA profiling systems.\u003c/p\u003e \u003cp\u003eThe requirement to generate DNA profiles from population-wide reference samples using a specific marker panel incurs considerable research and development costs for each species. Collecting reference samples poses numerous practical and ethical challenges, particularly for wide-ranging, rare, or dangerous species, or for those requiring special permits. Even when reference samples are available, the design and development of a taxon-specific marker panel requires considerable resources, which is often not economically viable for species that are not tested in large numbers. Further, transferring an established DNA profiling system and associated reference data to a new laboratory introduces additional challenges; reference data may be unsuitable if the target population changes, and the system may require calibration across laboratories (e.g., using standard genotyping samples) to ensure compatibility. Consequently, DNA-based individual identification is largely restricted to abundant model species (e.g. humans and domestic animals) or those expected to undergo routine testing, and is generally infeasible for rarer, less-studied taxa. These limitations are particularly evident in wildlife trade applications, where most trafficked species lack established DNA profiling systems and reference datasets, and many laboratories have no access to the few systems that do exist.\u003c/p\u003e \u003cp\u003eHere, we present a novel approach to DNA profiling that does not require population genetic data or any specific marker panel to interpret a profile match, hence can be applied to any species, and as a general method, is easily transferable across laboratories. Traditional profiling systems comprise relatively few independent genetic markers (typically\u0026thinsp;\u0026lt;\u0026thinsp;25), with ample power to differentiate individuals due to the rarity of the genotype observed at each marker. However, it is now routine to generate hundreds to thousands of genome-wide markers from a DNA sample. We build on this capability and propose an updated approach for estimating match probabilities, which we term \u0026lsquo;heterozygote analysis for reference-free probability assessment\u0026rsquo; (\u003cem\u003eHARP\u003c/em\u003e). \u003cem\u003eHARP\u003c/em\u003e considers only heterozygous site matches and applies ultra-conservative assumptions: that each marker is biallelic, with alleles present at equal frequencies in the population (i.e. \u003cem\u003ep\u003c/em\u003e\u0026thinsp;=\u0026thinsp;\u003cem\u003eq\u003c/em\u003e=0.5). The resulting DNA profile match probabilities are generated without the need for specific marker panels or associated population reference data. We demonstrate the utility of HARP using three case studies. The adaptability of \u003cem\u003eHARP\u003c/em\u003e across species and use cases has the potential to transform the use of individual DNA profiling in wildlife diagnostics.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eHARP \u003cem\u003eassumptions and discriminatory power\u003c/em\u003e\u003c/p\u003e \u003cp\u003e \u003cem\u003eHARP\u003c/em\u003e match probabilities are calculated without reference data by only considering heterozygous loci (see Methods and Supplementary Methods for the full derivation). \u003cem\u003eHARP\u003c/em\u003e assumes two alleles per heterozygote locus, both with a population frequency of 0.5. This is conservative: when allele frequencies differ from 0.5, or when more than two alleles exist in the population, match probabilities decrease (i.e. become more powerful).\u003c/p\u003e \u003cp\u003e \u003cem\u003eHARP\u003c/em\u003e is underpinned by two additional key assumptions. First, markers are assumed to be in linkage disequilibrium (LD) to justify calculation of cumulative probabilities across independent loci. This assumption can be met by filtering markers based on genomic location or by carrying out LD testing. Second, markers are assumed to be in Hardy\u0026ndash;Weinberg equilibrium (HWE). A conventional approach to account for HWE deviations when calculating match probabilities is to incorporate a theta (\u003cem\u003eθ\u003c/em\u003e) correction, which addresses heterozygote deficiency (e.g. as a result of population structure or individual relatedness). However, incorporating \u003cem\u003eθ\u003c/em\u003e\u0026thinsp;\u0026gt;\u0026thinsp;0 into \u003cem\u003eHARP\u003c/em\u003e decreases the resultant match probability, hence we conservatively assume \u003cem\u003eθ\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0. A less common issue, but one that may lead to underestimation of heterozygote match probabilities, is heterozygote excess. This can arise through selection scenarios such as heterozygote advantage, deleterious recessive alleles, advantageous dominant alleles and disassortative mating (elaborated in the Supplementary Methods). These scenarios are accounted for by incorporating a correction factor, \u003cem\u003es\u003c/em\u003e, representing the proportion of the markers affected by scenarios leading to heterozygote excess.\u003c/p\u003e \u003cp\u003eApplying \u003cem\u003eHARP\u003c/em\u003e to a single heterozygous marker, conservatively assuming \u003cem\u003es\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.1 (see Supplementary Methods for justification), results in a match probability of ~\u0026thinsp;0.536; i.e. there is a 53.6% chance of observing this same heterozygote genotype in the general population. This can also be expressed as a likelihood ratio (LR): when matching a sample to a suspected donor, observing this genotype is 1.866 times more likely if the sample derives from the donor rather than from an unrelated individual. \u003cem\u003eHARP\u003c/em\u003e probabilities can be multiplied across multiple heterozygous loci provided they segregate independently, yielding considerable discriminatory power despite the conservative assumptions underpinning the calculation. For example, for 25 independent heterozygote markers, assuming \u003cem\u003es\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.1, the probability that another individual in the population shares this heterozygote marker profile is \u0026lt;\u0026thinsp;1.7x10\u003csup\u003e\u0026minus;\u0026thinsp;7\u003c/sup\u003e, equivalent to a LR of \u0026gt;\u0026thinsp;5.9\u0026nbsp;million.\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eCase studies\u003c/h2\u003e \u003cp\u003eWe applied \u003cem\u003eHARP\u003c/em\u003e to a hypothetical hippopotamus poaching case, in which a carcass is discovered, a hippo tooth (ivory) is recovered from a vehicle, and suspected poachers are subsequently detained in possession of a knife and clothing, displaying traces of blood (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Heterozygote profiles from the four hippopotamus whole genome datasets generated from the same individual (using independent library preparations; Supplementary Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e) matched across 21\u0026ndash;25 heterozygote single nucleotide polymorphism markers (SNPs), with no mismatches, yielding match probabilities ranging from 2.04x10\u003csup\u003e\u0026minus;\u0026thinsp;6\u003c/sup\u003e to 1.69x10\u003csup\u003e\u0026minus;\u0026thinsp;7\u003c/sup\u003e (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Profiles from different individuals, including first-degree relatives, were successfully excluded. This case study demonstrates the application of \u003cem\u003eHARP\u003c/em\u003e when no marker panel is available. Notably, we took stringent measures to ensure markers were in LD by utilising only single SNPs from the longest scaffolds of the hippopotamus genome. Thousands of genome-wide SNPs could potentially be utilized, which would dramatically increase discriminatory power (provided markers were appropriately filtered or LD was accounted for). Increasing the number of markers compared also increases the risk of false mismatches; however, these can be accounted for using error rate thresholds or incorporating a mismatch probability into the likelihood framework.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWe also applied \u003cem\u003eHARP\u003c/em\u003e to two species with established SNP profiling systems, but analysing the datasets without using population reference data. A 144 SNP panel for snow leopards [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e] yielded an expected 15 heterozygotes between matching samples (see Supplementary Methods), corresponding to a \u003cem\u003eHARP\u003c/em\u003e match probability of 8.63x10\u003csup\u003e\u0026minus;\u0026thinsp;5\u003c/sup\u003e, and a 128 SNP panel for bigleaf maple [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e] yielded an expected 40 heterozygotes between matching samples, corresponding to a \u003cem\u003eHARP\u003c/em\u003e match probability of 1.46x10\u003csup\u003e\u0026minus;\u0026thinsp;11\u003c/sup\u003e (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). These examples represent applications to conservation genetic monitoring and timber trade authenticity testing, respectively.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eExpected number of heterozygote matches for the snow leopard and bigleaf maple case studies and their associated match probabilities. The number of heterozygotes for each case study is based on the missingness and observed heterozygosity reported for the associated profiling systems (see Supplementary Methods). We assumed \u003cem\u003es\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.1 for the \u003cem\u003eHARP\u003c/em\u003e match probability calculations.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCase study species\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMarker panel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eExpected number of matching heterozygotes\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMatch probability\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSnow leopard\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e144 SNP panel genotyped using high-throughput amplicon sequencing\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e15\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e8.63 x 10\u003csup\u003e\u0026minus;\u0026thinsp;5\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBigleaf maple\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e128 SNP panel genotyped using the MassARRAY\u0026reg; iPLEX\u0026trade; platform\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e40\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.46 x 10\u003csup\u003e\u0026minus;\u0026thinsp;11\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eAlthough DNA-based individual identification is widely applicable across wildlife diagnostic contexts, it has been implemented for only a small number of species. This is primarily due to the substantial research and development required to develop species-specific marker panels and associated reference data. \u003cem\u003eHARP\u003c/em\u003e addresses this barrier by enabling individual identification analysis without the need for these prerequisites, and can therefore be applied to any species for which DNA profiling systems have not been developed.\u003c/p\u003e \u003cp\u003e \u003cem\u003eHARP\u003c/em\u003e will be particularly beneficial in wildlife forensic applications, as illustrated in the hippopotamus poaching case study. A lack of testing capacity likely hinders prosecution in many cases that hinge on matching samples. Unlike conventional forensic DNA profiling, which requires marker validation, the whole genome approach used in the hippopotamus case study does not explicitly characterise individual markers. While this may raise instinctive concerns within the forensic genetics community, we believe that any associated risks can be effectively mitigated through stringent and extremely conservative bioinformatic pipelines that only retain highly reliable SNPs, minimise LD, and correct for markers with potential excess heterozygosity through the use of a conservative value for \u003cem\u003es\u003c/em\u003e. We encourage further discussions and research into how genomic approaches and the application of \u003cem\u003eHARP\u003c/em\u003e can be validated for forensic individualisation analyses.\u003c/p\u003e \u003cp\u003eBeyond forensics, \u003cem\u003eHARP\u003c/em\u003e is broadly applicable to any diagnostic context requiring individualisation, and could transform how DNA-based diagnostic systems are designed and managed for product traceability. Audit systems conducting one-up-one-down supply chain verification for any species, but particularly those with high value individuals, such as timbers, could apply this method without prior species-specific research and development. For example, \u003cem\u003eHARP\u003c/em\u003e could be applied to match logs to tree stumps in illegal logging compliance testing involving multiple timber species, where individual trees are harvested and labelled within a concession.\u003c/p\u003e \u003cp\u003e \u003cem\u003eHARP\u003c/em\u003e is agnostic to how heterozygote markers are generated. When existing marker panels are used to generate heterozygote profiles, no inter-laboratory calibration is required (often necessary to enable shared reference datasets), facilitating straightforward transfer of profiling systems across laboratories. A whole genome sequencing approach, as illustrated in the hippopotamus case study, does not require a specific marker panel, and could be applied to any species. This approach relies on a relatively high-quality reference genome to confidently call SNPs, minimise false heterozygote calls, and evaluate linkage disequilibrium (LD); such resources are increasingly available through global reference sequencing initiatives (e.g. DToL, ERGA, EarthBioGenome). The per sample sequencing effort required for this \u003cem\u003eHARP\u003c/em\u003e whole genome approach depends on genome size and sequence quality, and could be substantially reduced through enrichment or genome complexity reduction methods (e.g. [\u003cspan additionalcitationids=\"CR14\" citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]).\u003c/p\u003e \u003cp\u003eIn summary, \u003cem\u003eHARP\u003c/em\u003e has the potential to revolutionise DNA profiling for matching biological samples. By eliminating the need for species-specific development and reference data, \u003cem\u003eHARP\u003c/em\u003e paves the way for a dramatic increase in novel DNA profiling applications across forensic science, diagnostics, and research.\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003eHARP \u003cem\u003eapproach\u003c/em\u003e\u003c/p\u003e \u003cp\u003eHere we derive and justify a per marker match probability without the need for a population genetic reference database. At each step, we take a conservative approach for this derivation.\u003c/p\u003e \u003cp\u003eAssuming two alleles (A and B) occur at a given marker, the maximum probability, \u003cem\u003eP\u003c/em\u003e, of observing a heterozygous (AB) genotype is \u003cem\u003eP\u003c/em\u003e(AB)\u0026thinsp;=\u0026thinsp;0.5. The corresponding probabilities for AA and BB genotypes are \u003cem\u003eP\u003c/em\u003e(AA)\u0026thinsp;=\u0026thinsp;\u003cem\u003eP\u003c/em\u003e(BB)\u0026thinsp;=\u0026thinsp;1, making homozygous loci uninformative in the absence of population data. While a conservative heterozygote genotype probability of \u003cem\u003eP\u003c/em\u003e(AB)\u0026thinsp;=\u0026thinsp;0.5 is less powerful than those typically calculated using population reference data (i.e. if \u003cem\u003ep\u003c/em\u003e or \u003cem\u003eq\u003c/em\u003e is \u0026lt;\u0026thinsp;0.5, or if there are \u0026gt;\u0026thinsp;2 alleles), when using a relatively large number of markers in combination, it is possible to recover sufficient heterozygote genotype matches to produce extremely small match probabilities.\u003c/p\u003e \u003cp\u003e \u003cem\u003eHARP\u003c/em\u003e assumes HWE, which is violated through either deficiency or excess in observed heterozygote genotypes. Only heterozygote excess is relevant to \u003cem\u003eHARP\u003c/em\u003e and is accounted for by estimating the proportion of markers affected and introducing an associated correction factor, \u003cem\u003es\u003c/em\u003e, as follows:\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$P\\left(AB\\right)={0.5}^{\\left(n\\times\\left(1-s\\right)\\right)}$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003ewhere \u003cem\u003en\u003c/em\u003e is the number of DNA markers with matching heterozygous genotypes in the profile. This equation holds even when comparing highly related individuals (e.g. first-degree relatives).\u003c/p\u003e \u003cp\u003e \u003cem\u003eHARP\u003c/em\u003e also assumes markers are not linked, which is addressed by filtering markers to minimise LD. See the Supplementary Methods for further details of the \u003cem\u003eHARP\u003c/em\u003e assumptions and corrections.\u003c/p\u003e\n\u003ch3\u003eCase studies\u003c/h3\u003e\n\u003cp\u003e \u003cem\u003eHARP\u003c/em\u003e was applied to a hypothetical hippopotamus poaching case. DNA profile comparisons were made between samples from hippopotamus ivory, trace evidence on the knife and clothing, the carcass, and two other hippopotamus individuals (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). The ivory, trace evidence and carcass were derived from the same individual, while the two additional individuals formed a trio (sire, dam and offspring) to assess the discriminatory power of the approach among closely related individuals. Each evidence sample was represented by an independently re-sequenced hippo genome (Supplementary Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e), obtained from two studies [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. This case study was chosen for two key reasons: 1) it represents a realistic wildlife forensic application of \u003cem\u003eHARP\u003c/em\u003e; 2) it involves a species for which no marker panel or associated reference data are available. Whole genome sequencing data were available for the hippopotamus samples involved in this hypothetical case, providing a large set of heterozygous markers available for \u003cem\u003eHARP\u003c/em\u003e analyses.\u003c/p\u003e \u003cp\u003eDetails on the bioinformatic pipeline for this case study is outlined in the Supplementary Methods and relevant GitHub repository (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/KyleEwart/HARP_WGS_pipeline\u003c/span\u003e\u003cspan address=\"https://github.com/KyleEwart/HARP_WGS_pipeline\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). Briefly, we aligned these short-read datasets to a hippopotamus genome assembly (GCA_030028035.1), called SNPs, applied stringent filtering to remove low quality SNP and potential false heterozygotes, then compared heterozygote calls across individuals.\u003c/p\u003e \u003cp\u003eWe also applied \u003cem\u003eHARP\u003c/em\u003e to two additional case studies involving species that have established DNA profiling systems: 1) a 128 SNP panel for bigleaf maple, used to assess trade authenticity of this commonly traded timber species [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e], and 2) a 144 SNP panel for snow leopards, developed for conservation genetic applications [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e] (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). In these analyses, the existing systems were treated as if suitable reference datasets were unavailable (or unsuitable) in order to illustrate how \u003cem\u003eHARP\u003c/em\u003e would perform under such conditions. The expected number of heterozygote matches between profiles was estimated from reported heterozygosity and missingness in the respective studies (elaborated in the Supplementary Methods).\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAuthor contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eK.M.E. and R.O. came up with the initial concept; K.M.E. and R.O. developed the method; K.M.E. performed the analyses; K.M.E. and R.O. wrote the manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe are grateful for the support of a Marie Skłodowska-Curie Actions (MSCA) Postdoctoral Fellowship (K.M.E.).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflict of interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors have no competing interests to declare.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData availability statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAnalyses carried out in this study are based on previously generated and openly accessible datasets. The relevant studies and accession numbers are available within the paper and its Supplementary Information.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eEvett, I. W., Buffery, C., Willott, G. \u0026amp; Stoney, D. A guide to interpreting single locus profiles of DNA mixtures in forensic cases. \u003cem\u003eJ. Forensic Sci. Soc.\u003c/em\u003e 31, 41\u0026ndash;47 (1991).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMeester, R. \u0026amp; Slooten, K. \u003cem\u003eProbability and forensic evidence: theory, philosophy and applications\u003c/em\u003e (Cambridge Univ. Press, Cambridge, 2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHarper, C., et al. Robust forensic matching of confiscated horns to individual poached African rhinoceros. \u003cem\u003eCurr. Biol.\u003c/em\u003e 28, R13-R14 (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEwart, K. M., et al. \u003cem\u003eTigerBase\u003c/em\u003e: A DNA registration system to enhance enforcement and compliance testing of captive tiger facilities. \u003cem\u003eForensic Sci. Int. Genet.\u003c/em\u003e 74, 103149 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDe Bruyn, M., Dalton, D. L., Mwale, M., Ehlers, K. \u0026amp;Kotze, A. Development and validation of a novel forensic STR multiplex assay for blue (\u003cem\u003eAnthropoides paradiseus\u003c/em\u003e), wattled (\u003cem\u003eBugeranus carunculatus\u003c/em\u003e), and grey-crowned crane (\u003cem\u003eBalearica regulorum\u003c/em\u003e). \u003cem\u003eForensic Sci. Int. Genet.\u003c/em\u003e 73, 103100 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKalinowski, S. T., Taper, M. L. \u0026amp; Creel, S. Using DNA from non-invasive samples to identify individuals and census populations: an evidential approach tolerant of genotyping errors. \u003cem\u003eConserv. Genet.\u003c/em\u003e 7, 319\u0026ndash;329 (2006).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSethi, S. A., et al. Accurate recapture identification for genetic mark\u0026ndash;recapture studies with error-tolerant likelihood-based match calling and sample clustering. \u003cem\u003eR. Soc. Open Sci.\u003c/em\u003e 3, 160457 (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSolari, K. A., et al. Next-generation snow leopard population assessment tool: multiplex‐PCR SNP panel for individual identification from faeces. \u003cem\u003eMol. Ecol. Resour.\u003c/em\u003e 24, e14074 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eScheffers, B. R., Oliveira, B. F., Lamb, I. \u0026amp; Edwards, D. P. Global wildlife trade across the tree of life. \u003cem\u003eScience\u003c/em\u003e 366, 71\u0026ndash;76 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHughes, L. J., Morton, O., Scheffers, B. R. \u0026amp; Edwards, D. P. The ecological drivers and consequences of wildlife trade. \u003cem\u003eBiol. Rev.\u003c/em\u003e 98, 775\u0026ndash;791 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEvett, I. W. \u0026amp; Weir, B. S. \u003cem\u003eInterpreting DNA Evidence.\u003c/em\u003e (Sinauer, Sunderland, 1998).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDormontt, E. E., et al. Forensic validation of a SNP and INDEL panel for individualisation of timber from bigleaf maple (\u003cem\u003eAcer macrophyllum\u003c/em\u003e Pursch). \u003cem\u003eForensic Sci. Int. Genet.\u003c/em\u003e 46, 102252 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDavey, J. W. \u0026amp; Blaxter, M. L. RADSeq: next-generation population genetics. \u003cem\u003eBrief. Funct. Genom.\u003c/em\u003e 9, 416\u0026ndash;423 (2010).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLoose, M., Malla, S. \u0026amp; Stout, M. Real-time selective sequencing using nanopore technology. \u003cem\u003eNat. Methods\u003c/em\u003e 13, 751\u0026ndash;754 (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSinn, B. T., et al. ISSRseq: An extensible method for reduced representation sequencing. \u003cem\u003eMethods Ecol. Evol.\u003c/em\u003e 13, 668\u0026ndash;681 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e\u0026Aacute;rnason, \u0026Uacute;., Lammers, F., Kumar, V., Nilsson, M. A. \u0026amp; Janke, A. Whole-genome sequencing of the blue whale and other rorquals finds signatures for introgressive gene flow. \u003cem\u003eSci. Adv.\u003c/em\u003e 4, eaap9873 (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBergeron, L. A. et al. Evolution of the germline mutation rate across vertebrates. \u003cem\u003eNature\u003c/em\u003e 615, 285\u0026ndash;291 (2023).\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"npj-biodiversity","isNatureJournal":false,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"npjbiodivers","sideBox":"Learn more about [npj Biodiversity](https://www.nature.com/npjbiodivers/)","snPcode":"44185","submissionUrl":"https://mts-npjbiodivers.nature.com/cgi-bin/main.plex","title":"npj Biodiversity","twitterHandle":"@npjbiodiversity","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"npj","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"DNA profiling, match probabilities, wildlife diagnostics","lastPublishedDoi":"10.21203/rs.3.rs-9106203/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9106203/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe use of DNA analysis to match individual biological samples is central to many wildlife diagnostic applications, and is particularly valuable in illegal wildlife trade investigations. However, although thousands of wildlife and plant species are affected by illegal trade, DNA profiling systems for individual identification exist for only a small number of taxa. This is primarily due to the need for extensive population reference data to enable conclusive interpretation of DNA profile matches, a key requirement in diagnostic contexts. Here, we introduce a novel approach that does not require reference data to interpret an individual profile match, and can be used with any set of genetic markers on a case-by-case basis. This approach, called \u0026lsquo;heterozygote analysis for reference-free probability assessment\u0026rsquo; (\u003cem\u003eHARP\u003c/em\u003e), is based on ultra-conservative match probability estimates using only heterozygote genotypes and assuming both alleles have a frequency of 0.5 in the population. \u003cem\u003eHARP\u003c/em\u003e effectively eliminates the need for reference data and standard marker panels, and is therefore completely adaptable to any novel individual identification scenario. We demonstrate \u003cem\u003eHARP\u003c/em\u003e using three case studies spanning a range of species and diagnostic scenarios.\u003c/p\u003e","manuscriptTitle":"A novel DNA profiling method for individual sample matching and its applications to wildlife diagnostics","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-19 16:25:39","doi":"10.21203/rs.3.rs-9106203/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-04-21T15:49:18+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-15T00:52:44+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-09T15:10:09+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"300794574646949691799012848520912352952","date":"2026-03-17T15:54:57+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"19650096076876602010664885768835496292","date":"2026-03-16T22:40:21+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-03-16T22:34:23+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-03-16T07:54:52+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-03-16T02:39:27+00:00","index":"","fulltext":""},{"type":"submitted","content":"npj Biodiversity","date":"2026-03-12T14:48:33+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"npj-biodiversity","isNatureJournal":false,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"npjbiodivers","sideBox":"Learn more about [npj Biodiversity](https://www.nature.com/npjbiodivers/)","snPcode":"44185","submissionUrl":"https://mts-npjbiodivers.nature.com/cgi-bin/main.plex","title":"npj Biodiversity","twitterHandle":"@npjbiodiversity","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"npj","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"9d48bb72-43be-4550-bcd4-c6560da4af8a","owner":[],"postedDate":"March 19th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"in-revision","subjectAreas":[{"id":64668701,"name":"Biological sciences/Biological techniques"},{"id":64668702,"name":"Biological sciences/Ecology"},{"id":64668703,"name":"Earth and environmental sciences/Ecology"},{"id":64668704,"name":"Biological sciences/Genetics"},{"id":64668705,"name":"Biological sciences/Plant sciences"}],"tags":[],"updatedAt":"2026-04-21T15:55:14+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-19 16:25:39","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9106203","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9106203","identity":"rs-9106203","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00