Long-range PCR and Nanopore sequencing for localisation and phasing variants: an end-to-end clinical application workflow | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Long-range PCR and Nanopore sequencing for localisation and phasing variants: an end-to-end clinical application workflow Javad Jamshidi, Conor Rowntree, Shannon Fadaee, Futao Zhang, Ying Zhu, and 3 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7242084/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 19 Nov, 2025 Read the published version in BMC Medical Genomics → Version 1 posted 11 You are reading this latest preprint version Abstract Background: Next-Generation short-read sequencing has limited diagnostic utility in phasing distantly separated variants and analysing genomic regions with high homology. Determining the phase of variants from parental chromosomes is critical for accurate identification of compound heterozygosity. Long-read sequencing technology is able to overcome these limitations through the analysis of long haplotypes of separated variants. This study has developed and validated a robust, end-to-end workflow for phasing and localising variants using long-range PCR (LR-PCR) and targeted Nanopore sequencing for clinical implementation. Methods: NA24385 (HG002) reference DNA was used for all tests. Four PCR kits were tested to optimise LR-PCR for targets between 1 to 20 kb. Amplicons were barcoded and sequenced on Flongle flow cells, with up to eight amplicons on each flow cell. An in-house bioinformatic pipeline was developed to analyse the amplicons. This pipeline is capable of detecting chimeric reads (a known PCR artefact), and incorporating Clair3 for variant calling, and WhatsHap and HapCUT2 for phasing. Results: The UltraRun LongRange PCR Kit performed with a 90% success rate for DNA amplification up to 22 kb. All 15 tested heterozygous Single Nucleotide Variant (SNV) pairs, and 10 small InDels, with inter-variant distances from 5.8 to 21.4 kb, were phased with 100% concordance to known phase. Furthermore, SNV calling within six low-mappability genes demonstrated precision and sensitivity of 100% against benchmark data. The median proportion of chimeric reads was maintained at 2.80% (range 1.79–16.12%) under optimised conditions. Conclusions: This study establishes a reliable and affordable clinical diagnostic workflow for accurate phasing of variants separated by up to ~ 20 kb and for variant localisation in genomic regions not able to be sequenced by short-read sequencing. This integrated approach enables implementation in diagnostic settings to resolve complex genetic findings and improve variant interpretation. Nanopore Long-read sequencing Phasing Long-range PCR Figures Figure 1 Figure 2 Figure 3 Introduction Next generation sequencing (NGS) has revolutionised diagnostic genomics, with whole exome (WES) and whole genome (WGS) sequencing having become integral to clinical diagnosis. NGS has several inherent limitations related to short read sequencing (SRS), such as poor alignment in high homology regions of the human genome and the inability to phase variants more than a few hundred bases apart ( 1 ). Phasing refers to the process of determining the parental origin of each variant allele either located on the same chromosome ( in cis ) or on different chromosomes ( in trans ). Phasing can determine whether two alleles sequenced in a single individual are compound heterozygous ( in trans ) which is crucial for resolving the inheritance of variants in autosomal recessive conditions, particularly when parental segregation is not possible ( 2 ). In addition, when one parent is heterozygous for both variants or when one variant is de novo , parental sequencing cannot determine the phase. Poor alignment occurs in SRS within regions of the genome with large stretches of identical or nearly identical sequences that exist at multiple locations such as tandem repeats, transposable elements, paralogous genes and pseudogenes ( 3 ). In such regions short reads cannot be uniquely aligned to any one location, causing misalignment and poor mapping ( 4 ). Variants detected in regions with low mapping usually require confirmation using an alternative method, such as Sanger sequencing. However, designing specific Sanger primers may not be feasible due to high homology and sequencing length limitations, leaving limited diagnostic tools available to confirm the variants ( 3 ). Long-read sequencing (LRS) can overcome this issue as primers may be designed in unique sequence regions at a distance from the variant of interest. Recent advances in basecalling accuracy have enabled LRS to address the shortcomings of SRS, and for it to be implemented in diagnostic genomic testing ( 5 ). Oxford Nanopore Technologies (ONT) offers high accuracy, affordable and scalable LRS ( 6 ) using the smallest flow cell (Flongle Flow cells) for targeted sequencing, enabling financially feasible small-scale diagnostic assays. Although long-range PCR may amplify longer DNA sequences, challenges may arise with primer design and increased levels of chimeric reads, - a PCR artefact derived from two different biological sequences, affecting sequencing accuracy and phasing ( 7 , 8 ). Therefore, optimising long-range PCR and detection whilst minimising chimeric reads is crucial for successful sequencing applications for diagnostic purposes. This study has developed a robust long-range PCR protocol by comparing different PCR kits to maximise successful simultaneous amplification of long DNA targets (~ 1-20kb), a protocol to barcode and sequence multiple long amplicons on Flongle flow cells, and a bioinformatic pipeline to streamline and automate quality control, variant calling, localisation, and phasing from barcoded ONT amplicons. Combining these three steps has resulted in a reliable and affordable workflow for localising or phasing clinically relevant variants up to 20 kb apart, suitable for implementation in clinical diagnostics. Methods Samples and ethical considerations : NA24385 DNA (HG002 from GIAB) (Coriell Institute) was used to set up the PCRs, sequencing and evaluation of variant calling and phasing from amplicons. This project was reviewed by the South Eastern Sydney Local Health District HREC at a meeting of its Low and Negligible Risk Committee, which determined that the project did not involve any ethical risks in accordance with NSW Health Policy. Long-range PCR kit comparison Four different PCR kits with the ability to amplify long targets including Platinum SuperFi II PCR Master Mix (2X) (Invitrogen), LongAmp® Taq 2X Master Mix (New England Biolabs), Q5® Hot Start High-Fidelity 2X Master Mix (New England Biolabs), and UltraRun LongRange PCR Kit (Qiagen) were selected for assessment. Ten DNA regions of the human genome between 1–22 kb in size were chosen for amplification which contained at least one gene, and where each amplicon spanned a minimum of two exons. Primers for these regions were designed using NCBI primer blast (Primer IDs: A1-A10). PCR programs were configured according to the manufacturer’s recommendations for each kit. A single PCR program using a single annealing temperature and extension time was established for each kit to enable running samples simultaneously on a thermal cycler. All PCRs were run in 26 cycles to minimise chimeric reads. All PCRs were conducted in a final volume of 20 µl, using 150 ng of the same DNA sample (NA24385) and 0.5 µM of each forward and reverse primer. After PCR, the amplicons were analysed using the Agilent 4200 Tape Station System. A successful amplicon was defined as the presence of a clear band in the gel image (concentration > 2 ng/µl). Non-specific bands were defined as any additional bands, with a concentration > 2 ng/µl. A detailed description of PCR kit selection including PCR protocols, amplicons and primer sequences are presented in the supplementary document. Nanopore library preparation and sequencing : Barcoding genomic DNA using the Flongle protocol (Ligation Sequencing gDNA - Native Barcoding Kit 24 V14, SQK-NBD114.24) was adapted for amplicon library preparation. DNA repair and end-prep step was modified to accommodate the use of amplicons. More detail is available from the supplementary document. Nanopore libraries were prepared using the Ligation Sequencing Kit V14 (SQK-LSK114) and barcoded using Native Barcoding Kit 24 V14 (SQK-NBD114.24). Briefly, up to eight amplicons were multiplexed in each sequencing run with equimolar amounts of each amplicon. 0.75 µl Ultra II End-prep Enzyme Mix and 0.875 µl Ultra II End-prep Reaction Buffer (New England Biolabs #E7546) was added to 5–10 femtomole of each PCR product adjusted to 18.4 µl with nuclease free water. Reactions were incubated at 20°C for 5 minutes and 65°C for 5 minutes and then purified using AMPure XP beads (AXP). Native barcodes were ligated in a reaction with 7.5 µl end-prepped DNA, 2.5 µl Native Barcode, 10 µl Blunt/TA Ligase Master Mix (New England Biolabs #M0367), incubated at room temperature for 20 minutes. EDTA was then added to each reaction and samples were pooled and another AXP purification was performed. Sequencing adaptors were then ligated to the barcoded end-prepped amplicons in a reaction with 30 µl pooled barcoded sample, 5 µl Native Adapter (NA), 10 µl NEBNext Quick Ligation Reaction Buffer (5X) and 5 µl Quick T4 DNA Ligase (New England Biolabs #E6056) incubated for 20 minutes at room temperature, followed by washing with Long Fragment Buffer (LFB) and AXP clean up. The library was eluted in 7 µl of Elution Buffer (EB) and incubated at 37°C for 10 minutes. Ten femtomole of the library was sequenced using a Flongle flow cell (R10.4.1, FLO-FLG114) on a GridION (ONT) device. Basecalling and alignment Super accuracy basecalling (SUP) ( [email protected] ) was performed during sequencing using MinKNOW software (GridION Release 24.06.15 version) which uses dorado ( https://github.com/nanoporetech/dorado ) for basecalling and Minimap2 version 2.28 ( 9 ) to align the reads with human reference genome (build hg38) and output BAM files. The minimum phred score was set to ten. Bioinformatic analysis Samtools v1.21 ( 10 ) was employed for BAM file manipulation and quality control. For phasing analysis, reads with mapping quality (MAPQ) < 20 and those shorter than the distance between the two variants were excluded. For variant localisation analysis, reads shorter than 1000 bp, those with read identity < 80%, and MAPQ < 20 were filtered out. Variant calling from amplicons was performed using Clair3 v1.0.4 ( 11 ), while phasing was conducted using WhatsHap v2.3 ( 12 ) and HapCUT2 v1.3.4 ( 13 ). WhatsHap was additionally used to haplotag the BAM files for visualisation. A custom Python script was developed to identify variants, reads, and their corresponding phases. The chimeric read proportion was calculated as the percentage of reads classified as the opposite phase to that determined by WhatsHap and HapCUT2, representing the smaller fraction of discordant reads. All BAM files were visualised and manually inspected using the Integrative Genomics Viewer (IGV) ( 14 ) to confirm the variants localisation and phase assignment. A bioinformatics pipeline was developed to automate these analyses and generate reports containing both quality control metrics and phase/localisation details. A minimum threshold of 50 high-quality reads (as defined above) was set as the requirement for reliable phasing. All BAM files were down-sampled to 50 reads and re-evaluated to assess whether this threshold was sufficient for accurate phasing. Complete documentation of the bioinformatics pipeline and associated code are available in the corresponding GitHub repository ( https://github.com/j-jamshidi/ONT_amp_phase ). Phasing samples To assess the performance of the selected long-range PCR kit, sequencing protocol and phasing analysis, ten heterozygote variant pairs with a distance ranging from 6.8–21.4 kb were selected from the NA24385 sample. Primers were designed to span the two variants in each case (Primer IDs: B1-B10). All amplicons were sequenced, and the phases of the variants were determined and compared with the known phase. Details of the variants, primers and amplicons are provided in the supplementary material table S10. Chimeric read additional samples Eight targets were amplified using (a) Platinum SuperFi II kit (26 cycles) and (b) UltraRun LongRange PCR kit (28 cycles) to assess the impact of different PCR conditions on chimeric read formation. The percentage of chimeric reads were then compared against the UltraRun LongRange PCR kit (26 cycles) baseline condition with the Wilcoxon signed-rank test. A p-value < 0.05 was considered statistically significant. Low mappability genes To evaluate the suitability of the method to call and localise variants in low mappability regions of the human genome, eight genes ( TUBB2A , TUBB2B , CYP11B1 , SBDS , HBA1 , HBA2 , RBMX , and CEL ) were selected. These genes are associated with Mendelian disorders and contain regions of low mappability. The details of these genes and their associated phenotypes are summarised in Table S11. Primers were designed to amplify the entire length of each gene. HBA1 and HBA2 are in close proximity in the human genome, therefore, a single primer pair was designed to span the region covering both genes. Details of the primers and amplicons are provided in the supplementary table S12. Long-range PCR and sequencing were performed similarly to the phasing cases. The called SNVs and small indels from amplicon sequencing of these genes (excluding RBMX , which is located on the X chromosome and not included in the benchmarking VCF file) were compared with GIAB ground-truth variants (benchmark version v4.2.1) ( 15 ), using the vcfeval command from RTG Tools v3.13 ( https://github.com/RealTimeGenomics/rtg-tools ). Results Long-range PCR Among the four PCR kits tested, the UltraRun LongRange PCR Kit demonstrated the best performance, successfully amplifying 9 out of 10 amplicons (90% success rate). Platinum SuperFi II and LongAmp® Taq showed a similar success rate of 70% (7 out of 10). The Q5® Hot Start High-Fidelity kit had the lowest success rate, successfully amplifying only four amplicons. Therefore, the UltraRun LongRange PCR Kit was selected for further amplifications. Details on the performance of each kit is available in the supplementary material, table S4-S7. All ten primer pairs designed for phasing (B1-B10), amplified their target regions successfully, although one amplicon (B1) required the addition of Q Solution, a PCR additive included with the UltraRun LongRange PCR Kit to enhance amplification. Six of the seven PCR primers designed for the amplification of low-mappability genes produced a product ( TUBB2A and HBA1/HBA2 required Q solution), whereas CEL primers did not yield a product. Nanopore sequencing 15 amplicons that were amplified with UltraRun LongRange PCR kit in 26 cycles (5 amplicons that were used for testing different PCR kits and 10 that were designed specifically for phasing) were sequenced on Flongle flow cells. All amplicons had high depth of high quality (MAPQ < 20) reads covering both variants of interest (mean = 1920, median = 3248, standard deviation = 1716.8, range = 286–5483). Table 1 summarises the sequencing details. Table 1 List of amplicons sequenced for variant phasing Amplicon ID Amplicon Size Variant 1 (hg38) Variant 2 (hg38) Distance between variants (bp) HQ spanning reads* Chimeric reads % Known phase B1 8743 chr16:88,643,329 C > T chr16:88,650,126 A > G 6797 5483 13.97 cis A3 9625 chr1:154,579,450 C > T chr1:154,585,209 C > T 5759 5385 4.25 cis B2 10026 chr12:8,927,591 T > C chr12:8,936,676 A > C 9085 4413 16.12 trans B3 11521 chr12:6,936,914 G > A chr12:6,944,269 C > T 7355 1748 3.03 cis B4 13789 chr4:41,638,861 C > T chr4:41,651,881 G > A 13020 673 1.93 cis B5 14391 chr3:49,412,205 G > A chr3:49,425,854 G > A 13649 3165 2.62 cis B6 14466 chr21:42,375,588 G > A chr21:42,388,818 A > G 13230 4450 3.33 cis A4 15163 chr6:152,389,867 T > C chr6:152,403,191 G > T 13324 3349 2.51 cis A5 16084 chr1:236,544,737 T > C chr1:236,559,878 C > A 15141 3248 5.48 trans A6 16837 chr1:43,658,540 G > T chr1:43,667,233 A > G 8693 3792 2.24 cis B7 17828 chr21:42,372,760 C > T chr21:42,387,895 A > C 15135 906 2.32 cis A7 18097 chr1:6,629,764 A > G chr1:6,645,884 C > T 16120 3,633 1.79 cis B8 21530 chr3:48,566,437 G > A chr3:48,584,160 T > C 17723 709 6.77 cis B9 22097 chr9:130,407,156 A > G chr9:130,426,008 C > T 18852 2570 2.80 trans B10 22409 chr1:236,011,853 T > C chr1:236,033,263 A > G 21410 286 2.10 cis * HQ Spanning reads: High quality reads (MAPQ > 20) that span both variants of interest (i.e., variants to be phased). Additionally, all amplicons amplified using Platinum SuperFi II PCR Master Mix and those amplified with 28 cycles using the UltraRun LongRange PCR Kit were also sequenced. The sequencing details are available in supplementary documents Tables S13 and S14. The depth of sequencing for each amplicon targeting low mappability genes was > 500x after QC, see supplementary tables 15. Phasing Fifteen variant pairs in amplicons generated using the UltraRun LongRange PCR Kit were phased and compared with the known phase. All variants were SNVs, with inter-variant distances ranging from 5,759 to 21,410 bp. All 15 variant pairs were phased correctly and were concordant with the known phase (Table 1 ). Furthermore, to test the robustness of this pipeline for phasing small indels, ten indels were identified and phased in the existing amplicons. All the indels were also phased correctly (table S16). Moreover, re-evaluating all cases with only 50 high quality reads per sample did not compromise phasing accuracy. Figure 1 shows an IGV screenshot of the phased reads for the B7 amplicon, highlighting the positions of the selected variants used for phasing as well as the variants called within the amplicon and the benchmarking variants. Chimeric reads : The estimated proportion of chimeric reads in the fifteen phased amplicons ranged from 1.79–16.12% (mean = 4.75, median = 2.80, SD = 4.41). The percentage of chimeric reads for each amplicon is presented in Table 1 . Chimeric reads in different PCR conditions The UltraRun LongRange (26 cycles) baseline condition yielded a median chimeric read percentage of 3.64% across the eight samples. Compared to this baseline, amplification with Platinum SuperFi II (26 cycles) (median = 9.33%; W-statistic = 1.0, p = 0.017) and increasing the cycles to 28 (median = 6.56%; W-statistic = 1.0, p = 0.017) resulted in significantly higher chimeric read levels. See Fig. 2 and Table S17. Variant calling from low mappability genes amplicons : The SNVs called via the pipeline with Clair3 from the amplicons generated for low mappability genes ( TUBB2A , TUBB2B , CYP11B1 , SBDS , HBA1 and HBA2 , 64 variants in total ) showed 100% concordance with the benchmark (high confidence) variants (Precision = 1, Sensitivity = 1). Figure 3 illustrates short-read WGS and long-read amplicon sequencing of the TUBB2A gene. Discussion This study developed and validated an integrated, end-to-end clinical diagnostic workflow combining optimised LR-PCR, targeted ONT sequencing, and an in-house developed bioinformatic pipeline for variant phasing and localisation. The workflow achieved accurate phasing of variants up to 21.4 kb apart and precise variant localisation in low-mappability genomic regions, addressing critical limitations of current short-read approaches. Methodologically, the study presents a comprehensively optimised workflow. This rigorous approach involved comparing LR-PCR kits systematically, optimising PCR conditions, developing a modified Nanopore library preparation protocol for barcoding on Flongle flow cells, and adjusting the multiplexing number and molarity of amplicons. Complementing these laboratory improvements is an automated bioinformatic pipeline integrating established tools like Clair3 for variant calling, and WhatsHap and HapCUT2 for phasing, alongside custom scripts for quality control and specific analytical logic for phasing versus localisation tasks. This end-to-end development, tailored for a cost-constrained yet high-performance application, represents a significant advancement beyond existing tools. LR-PCR is a recognised versatile technique for amplifying large genomic segments ( 16 ), but its success is highly dependent on careful optimisation, limiting its applications in targeted long-read sequencing ( 17 ). In the current study, different LR-PCR kits were compared to optimise the workflow, with the UltraRun LongRange PCR Kit demonstrating the best performance by successfully amplifying nine out of ten targets. Both Platinum SuperFi II and LongAmp® Taq also performed well, each amplifying seven targets successfully. Notably, the three amplicons that failed with both kits were the longest in the panel (all > 18 kb), consistent with the difficulty of amplify ultra-long fragments. The Q5® Hot Start High-Fidelity kit demonstrated the weakest performance, with only four successful amplifications, likely due to the use of a fixed annealing temperature, despite the kit's recommendation for primer-specific annealing conditions. A universal annealing temperature and extension time is necessary to allow simultaneous amplification of samples, facilitate batching and streamlining the workflow, which are key considerations for implementation in a diagnostic laboratory. Although some regions slightly over 20 kb were amplified, limiting amplicons to 20 kb is recommended due to the technical challenges of amplifying longer targets ( 18 ). An internal analysis of a curated panel of 5,678 genes associated with Mendelian disorders used in our laboratory showed a median gene size of approximately 38 kb. In this dataset for any two random variants within a gene, there is a ~ 68% chance they would lie within a 20kb range (figure S1 and Supplementary Data for detailed analysis). 20kb should therefore be a sufficient length for diagnostic utility for the majority of variants requiring phasing in a clinical context. Variants > 20 kb apart could be potentially phased using overlapping amplicons, however, this approach adds complexity and may not always be feasible ( 20 ). While this strategy can be effective for targeted phasing of specific loci ( 21 ), it is less practical for general diagnostic testing. Given the generality and focus of this method on clinical implementation, we chose not to pursue this strategy to avoid unnecessary complexity. An important technical consideration in amplicon-based phasing is the formation of chimeric reads during PCR, which can falsely link variants on different haplotypes ( 7 ). Unlike microbial diversity studies, which benefit from established tools for filtering out chimeric reads ( 22 ), phasing applications lack robust methods for chimera removal, largely because the parental haplotypes may be nearly identical and sometimes differing only by the variants being phased ( 20 ). Furthermore, approximately 1.7% of reads in nanopore-based amplicon sequencing may contain post-amplification chimeric elements ( 23 ). Therefore, detecting and minimising chimeric reads is critical for reliable phasing using amplicon-based approaches. Previous studies have demonstrated that failure to account for chimeric reads can lead to incorrect phasing interpretations ( 7 ). The number of PCR cycles was reduced to 26 to reduce chimera formation, a strategy supported by earlier studies ( 7 , 20 , 24 ) and confirmed by this study, where increasing the cycle number to 28 led to higher chimeric reads. The UltraRun LongRange PCR Kit was observed to generate fewer chimeric reads than Platinum SuperFi II on the same samples, further supporting its selection. This difference may arise from the polymerases used or protocol variations; notably, UltraRun uses a two-step PCR protocol, which has been associated with lower chimera formation rates ( 25 ). Beyond PCR optimisation, the developed bioinformatic pipeline included a method for quantifying chimeric reads based solely on the two variants under phasing. This approach contrasts with other studies that relied on external software or manual review, requiring extra steps for chimeric detection ( 20 , 26 ). In the primary sample, chimeric read levels were between 1.79–16.12% (median 2.8%), consistent with previously reported figures (7,20). While a median rate of 2.8% allows for accurate phasing, the broad range highlights the potential for substantial chimera presence in specific amplicons. This variability underscores the importance of careful, amplicon-specific assessment of chimeric read rates, especially in diagnostic contexts where precision is critical. Notably, the developed pipeline maintained robust and accurate phasing performance even in samples with chimeric read levels as high as 27%, observed in Platinum SuperFi II amplicons (see Table S17). LR-PCR and nanopore sequencing have been employed for phasing variants ( 8 , 20 , 21 , 27 , 28 ) and localising those identified in low-mappability regions ( 29 , 30 ) in previous studies. However, these studies focused on specific genomic regions, such as haplotyping the ABCA4 locus ( 28 ), or the localisation of variants within the highly repetitive domains of TTN ( 29 ). This study has developed a robust and reliable workflow adaptable to any genomic region that can be feasibly amplified using LR-PCR. Notably, previous similar studies have typically used MinION flow cells for multiple amplicons in nanopore-based assays ( 8 , 31 , 32 ) while Flongle flow cells have been used often for single assay sequencing ( 27 , 33 , 34 ). Sequencing multiple barcoded amplicons on a Flongle flow cells was a key feature of our assay design, offering an affordable and practical solution, particularly for smaller diagnostic laboratories. While the workflow described here is adaptable to larger flow cells such as the MinION, achieving similar cost efficiency to Flongle would require a larger number of samples per run. In a diagnostic setting, accumulating sufficient samples for high-level multiplexing could lead to elevated turnaround times ( 35 ). Flongle-based multiplexing therefore currently offers the most efficient and affordable solution for clinically relevant applications and supports the broader adoption of targeted nanopore sequencing in clinical diagnostics. The median sequencing depth per amplicon in this study exceeded 3,000x, consistent with the multiplexing of more than eight amplicons per Flongle flow cell being technically feasible. However, variability in ONT flow cell performance relating to active pore count, and significant variation in depth, despite using equimolar quantities of PCR products ( 36 ) can affect stability and reliability, which are critical in diagnostic settings. Limiting the number of amplicons per run therefore ensures consistent performance, even when sequencing output is suboptimal. Additionally, as with MinION flow cells, higher multiplexing may delay turnaround time due to the need to batch more samples. Thus, using eight amplicons per run balances cost-efficiency, test robustness, and clinical practicality. The developed workflow has significant potential to enhance molecular diagnostics. Accurate phasing of variants is crucial for the interpretation of results in autosomal recessive Mendelian conditions, where it is necessary to determine if two heterozygous variants in a gene are in cis or in trans to confirm compound heterozygosity ( 2 , 37 ). Misinterpretation of phase can lead to incorrect diagnostic conclusions. By providing direct phasing from proband DNA, this workflow can clarify the clinical significance of co-occurring variants. Furthermore, the ability to reliably call variants in low-mappability regions is of high clinical importance. Many disease-associated genes contain segments with high homology to other genomic locations (e.g., pseudogenes) or repetitive sequences, which can lead to variants being missed or incorrectly called by standard NGS approaches ( 38 ). The successful validation of variant calling in genes like TUBB2A/TUBB2B , CYP11B1 , and HBA1 / HBA2 demonstrates the workflow's utility in these challenging but clinically important regions. The cost-effectiveness and targeted nature of this workflow positions it as an ideal "Tier 2" diagnostic test. It can be employed to resolve ambiguous findings from initial WES or WGS, such as phasing variants of uncertain significance (VUS) to aid in their classification, or to confirm variants detected in regions where NGS data quality is suboptimal. The routine availability of such assays could lead to faster, more accurate diagnoses and reduce the number of unresolved cases and ultimately improving patient care. While the study demonstrates considerable strengths, certain limitations should be acknowledged. The LR-PCR protocol, despite optimisation and selection of the best-performing kit, did not achieve universal amplification for all attempted targets. For example, primers designed for the CEL gene failed to yield an amplicon, and several other amplicons ( CYB , TUBB2A , HBA ) required the addition of a PCR additive (Q Solution) for successful amplification. This is consistent with certain genomic regions remaining challenging for LR-PCR amplification, potentially due to extreme GC content, secondary structures, or other sequence-specific features ( 16 ). Such targets might require further bespoke optimisation, alternative primer designs, or different amplification strategies. The proportion of chimeric reads, while generally low with the optimised protocol, showed variability across amplicons, with some exhibiting higher levels. Although not observed in this study, amplicons that consistently produce a high percentage of chimeric reads could compromise the accuracy of phasing or variant detection if not controlled stringently and bioinformatically flagged. This highlights the importance of establishing clear quality control thresholds for chimeric read proportions for each diagnostic target. This workflow proved robust for detecting SNVs and small indels, however, it was not designed for other variant types such as structural variants or short tandem repeats (STRs) and is therefore limited to SNVs and indels. Moreover, ONT sequencing has a higher error rate for indels than for SNVs ( 39 ), warranting extra caution when interpreting such variants. Detection of variants located within homopolymer regions also are challenging using ONT sequencing ( 40 ), placing them outside the scope of this workflow. Furthermore, although the workflow achieved a precision and sensitivity of 1 for variant calling in the selected genes containing low-mappability regions, it is important to note that this evaluation was based on partial genomic regions covering only 64 variants. In larger datasets involving more genes and samples, a decrease in precision and sensitivity is expected ( 41 ). The variable performance of ONT flow cells can introduce variability across sequencing batches, potentially impacting the stability and reliability of the assay. Although our workflow was designed to mitigate this through specific wet-lab (e.g., limiting amplicon numbers) and bioinformatic strategies, this intrinsic limitation remains. The bioinformatic pipeline assists in managing this by flagging suboptimal results, such as instances where fewer than 50 high-quality reads are available to support phase interpretation or variant localisation. Conclusions This study has developed and validated a robust, accurate, and cost-effective targeted long-read sequencing workflow using LR-PCR and ONT Flongle technology. By addressing the challenges of variant phasing and analysis of low-mappability genomic regions effectively, this workflow offers a significant advancement over conventional short-read NGS methods for specific, complex diagnostic questions. The comprehensive optimisation from sample preparation through to bioinformatic analysis provides a practical solution with the potential to enhance molecular diagnostics, clarify ambiguous genetic findings, and ultimately improve patient care by enabling more precise genetic diagnoses where current methodologies may fall short. Declarations Ethics approval and consent to participate This project was reviewed by the South Eastern Sydney Local Health District Human Research Ethics Committee (HREC) at a meeting of its Low and Negligible Risk (LNR) Committee, which determined that the project did not involve any ethical risks and was exempt from full ethical review, in accordance with the NHMRC National Statement (updated 2023) —Section 5.1.22 , the NHMRC Ethical Considerations in Quality Assurance and Evaluation Activities (2014) guidance and NSW Health Guideline GL2007_020 Human Research Ethics Committees - Quality Improvement and Ethics Review: A Practice Guide for NSW . The study was conducted in accordance with the Declaration of Helsinki. This study does not include human participants and only publicly available, de-identified human genomic data (HG002 sample from the Genome in a Bottle Consortium) was used; no informed consent was required. Consent for publication Not applicable Availability of data and materials The sequencing data used in this study were generated from the HG002 reference sample, which is publicly available through the Genome in a Bottle (GIAB) consortium and can be accessed at https://www.nist.gov/programs-projects/genome-bottle. All documentation, source code, and example data for the bioinformatics pipeline developed in this study are available in the associated GitHub repository: https://github.com/j-jamshidi/ONT_amp_phase Competing interests The authors declare that they have no competing interests. Funding This study was supported by the University of New South Wales Sydney and the Medical Research Futures Fund PreGen program (GHFMPACI000006). Authors' contributions JJ, TR and FH conceptualised and designed the study. JJ and CR performed the PCR and sequencing experiments. JJ, FZ, and YZ analysed the data and developed the bioinformatic pipeline. JJ wrote the draft of the manuscript, and TR, MB, SF, FH, FZ, YZ, and CR reviewed and edited the main text. TR and MB provided the funding and resources. All authors read and approved the final version of the manuscript. References Bonfiglio F, Legati A, Lasorsa VA, Palombo F, De Riso G, Isidori F, et al. Best practices for germline variant and DNA methylation analysis of second- and third-generation sequencing data. Hum Genomics. 2024 Nov 5;18(1):120. Tewhey R, Bansal V, Torkamani A, Topol EJ, Schork NJ. The importance of phase information for human genomics. Nat Rev Genet. 2011 Mar;12(3):215–23. Rojahn S, Hambuch T, Adrian J, Gafni E, Gileta A, Hatchell H, et al. Scalable detection of technically challenging variants through modified next-generation sequencing. Mol Genet Genomic Med. 2022;10(12):e2072. Mandelker D, Schmidt RJ, Ankala A, McDonald Gibson K, Bowser M, Sharma H, et al. Navigating highly homologous genes in a molecular diagnostic setting: a resource for clinical next-generation sequencing. Genet Med. 2016 Dec;18(12):1282–9. Warburton PE, Sebra RP. Long-Read DNA Sequencing: Recent Advances and Remaining Challenges. Annu Rev Genomics Hum Genet. 2023 Aug 25;24(Volume 24, 2023):109–32. Espinosa E, Bautista R, Larrosa R, Plata O. Advancements in long-read genome sequencing technologies and algorithms. Genomics. 2024 May 1;116(3):110842. Laver TW, Caswell RC, Moore KA, Poschmann J, Johnson MB, Owens MM, et al. Pitfalls of haplotype phasing from amplicon-based long-read sequencing. Sci Rep. 2016 Feb 17;6(1):21746. Maestri S, Maturo MG, Cosentino E, Marcolungo L, Iadarola B, Fortunati E, et al. A Long-Read Sequencing Approach for Direct Haplotype Phasing in Clinical Settings. Int J Mol Sci. 2020 Jan;21(23):9177. Li H. Minimap2: pairwise alignment for nucleotide sequences. Bioinforma Oxf Engl. 2018 Sep 15;34(18):3094–100. Li H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, et al. The Sequence Alignment/Map format and SAMtools. Bioinforma Oxf Engl. 2009 Aug 15;25(16):2078–9. Zheng Z, Li S, Su J, Leung AWS, Lam TW, Luo R. Symphonizing pileup and full-alignment for deep learning-based long-read variant calling. Nat Comput Sci. 2022 Dec;2(12):797–803. Martin M, Ebert P, Marschall T. Read-Based Phasing and Analysis of Phased Variants with WhatsHap. Methods Mol Biol Clifton NJ. 2023;2590:127–38. Edge P, Bafna V, Bansal V. HapCUT2: robust and accurate haplotype assembly for diverse sequencing technologies. Genome Res. 2016 Dec 9;gr.213462.116. Robinson JT, Thorvaldsdóttir H, Wenger AM, Zehir A, Mesirov JP. Variant Review with the Integrative Genomics Viewer. Cancer Res. 2017 Oct 31;77(21):e31–4. Krusche P, Trigg L, Boutros PC, Mason CE, De La Vega FM, Moore BL, et al. Best Practices for Benchmarking Germline Small Variant Calls in Human Genomes. Nat Biotechnol. 2019 May;37(5):555–60. Chowdhary MR, Gupta N. Long-range PCR in next generation sequencing: A low cost approach for large & complex genes. Indian J Med Res. 2023 Jun;157(6):591–2. Kee PS, Karunanathie H, Maggo SDS, Kennedy MA, Chua EW. Long-Range Polymerase Chain Reaction. In: Domingues L, editor. PCR: Methods and Protocols [Internet]. New York, NY: Springer US; 2023 [cited 2025 Jul 1]. p. 181–92. Available from: https://doi.org/10.1007/978-1-0716-3358-8_15 Maggi J, Koller S, Bähr L, Feil S, Kivrak Pfiffner F, Hanson JVM, et al. Long-Range PCR-Based NGS Applications to Diagnose Mendelian Retinal Diseases. Int J Mol Sci. 2021 Jan;22(4):1508. Piovesan A, Antonaros F, Vitale L, Strippoli P, Pelleri MC, Caracausi M. Human protein-coding genes and gene feature statistics in 2019. BMC Res Notes. 2019 Jun 4;12(1):315. McClinton B, Watson CM, Crinnion LA, McKibbin M, Ali M, Inglehearn CF, et al. Haplotyping Using Long-Range PCR and Nanopore Sequencing to Phase Variants: Lessons Learned From the ABCA4 Locus. Lab Invest. 2023 Aug 1;103(8):100160. Gueuning M, Thun GA, Wittig M, Galati AL, Meyer S, Trost N, et al. Haplotype sequence collection of ABO blood group alleles by long-read sequencing reveals putative A1 -diagnostic variants. Blood Adv. 2023 Mar 28;7(6):878–92. Callahan BJ, Grinevich D, Thakur S, Balamotis MA, Yehezkel TB. Ultra-accurate microbial amplicon sequencing with synthetic long reads. Microbiome. 2021 Jun 5;9(1):130. White R, Pellefigues C, Ronchese F, Lamiable O, Eccles D. Investigation of chimeric reads using the MinION. F1000Research. 2017 Aug 16;6:631. Namias A, Sahlin K, Makoundou P, Bonnici I, Sicard M, Belkhir K, et al. Nanopore sequencing of PCR products enables multicopy gene family reconstruction. Comput Struct Biotechnol J. 2023 Jan 1;21:3656–64. Qin Y, Wu L, Zhang Q, Wen C, Van Nostrand JD, Ning D, et al. Effects of error, chimera, bias, and GC content on the accuracy of amplicon sequencing. mSystems. 2023 Dec;8(6):e01025-23. Kroll F, Dimitriadis A, Campbell T, Darwent L, Collinge J, Mead S, et al. Prion protein gene mutation detection using long-read Nanopore sequencing. Sci Rep. 2022 May 18;12(1):8284. Durkie M, Watson CM, Winship P, Hogg AC, Nyanhete R, Cooley S, et al. The Common PKD1 p.(Ile3167Phe) Variant Is Hypomorphic and Associated with Very Early Onset, Biallelic Polycystic Kidney Disease. Hum Mutat. 2023;2023(1):5597005. Mc Clinton B, Crinnion LA, Ali M, Inglehearn C, Watson CM, Toomes C. Using Oxford Nanopore long-range sequencing to phase ABCA4. Invest Ophthalmol Vis Sci. 2021 Jun 21;62(8):1541. Perrin A, Van Goethem C, Thèze C, Puechberty J, Guignard T, Lecardonnel B, et al. Long-Reads Sequencing Strategy to Localize Variants in TTN Repeated Domains. J Mol Diagn. 2022 Jul 1;24(7):719–26. Watson CM, Dean P, Camm N, Bates J, Carr IM, Gardiner CA, et al. Long-read nanopore sequencing resolves a TMEM231 gene conversion event causing Meckel–Gruber syndrome. Hum Mutat. 2020;41(2):525–31. Holt GS, Batty LE, Alobaidi BKS, Smith HE, Oud MS, Ramos L, et al. Phasing of de novo mutations using a scaled-up multiple amplicon long-read sequencing approach. Hum Mutat. 2022;43(11):1545–56. Stockton JD, Nieto T, Wroe E, Poles A, Inston N, Briggs D, et al. Rapid, highly accurate and cost-effective open-source simultaneous complete HLA typing and phasing of class I and II alleles using nanopore sequencing. HLA. 2020;96(2):163–78. Jeck WR, Iafrate AJ, Nardi V. Nanopore Flongle Sequencing as a Rapid, Single-Specimen Clinical Test for Fusion Detection. J Mol Diagn. 2021 May 1;23(5):630–6. Watson CM, Holliday DL, Crinnion LA, Bonthron DT. Long-read nanopore DNA sequencing can resolve complex intragenic duplication/deletion variants, providing information to enable preimplantation genetic diagnosis. Prenat Diagn. 2022;42(2):226–32. Tafess K, Ng TTL, Lao HY, Leung KSS, Tam KKG, Rajwani R, et al. Targeted-Sequencing Workflows for Comprehensive Drug Resistance Profiling of Mycobacterium tuberculosis Cultures Using Two Commercial Sequencing Platforms: Comparison of Analytical and Diagnostic Performance, Turnaround Time, and Cost. Clin Chem. 2020 Jun 1;66(6):809–20. Whitford W, Hawkins V, Moodley KS, Grant MJ, Lehnert K, Snell RG, et al. Proof of concept for multiplex amplicon sequencing for mutation identification using the MinION nanopore sequencer. Sci Rep. 2022 May 20;12(1):8572. Guo MH, Francioli LC, Stenton SL, Goodrich JK, Watts NA, Singer-Berk M, et al. Inferring compound heterozygosity from large-scale exome sequencing data. Nat Genet. 2024 Jan;56(1):152–61. Claes KBM, Rosseel T, De Leeneer K. Dealing with Pseudogenes in Molecular Diagnostics in the Next Generation Sequencing Era. Methods Mol Biol Clifton NJ. 2021;2324:363–81. Santos R, Lee H, Williams A, Baffour-Kyei A, Lee SH, Troakes C, et al. Investigating the Performance of Oxford Nanopore Long-Read Sequencing with Respect to Illumina Microarrays and Short-Read Sequencing. Int J Mol Sci. 2025 May 8;26(10):4492. Olson ND, Wagner J, Dwarshuis N, Miga KH, Sedlazeck FJ, Salit M, et al. Variant calling and benchmarking in an era of complete human genome sequences. Nat Rev Genet. 2023 Jul;24(7):464–83. Nyaga DM, Tsai P, Gebbie C, Phua HH, Yap P, Le Quesne Stabej P, et al. Benchmarking nanopore sequencing and rapid genomics feasibility: validation at a quaternary hospital in New Zealand. NPJ Genomic Med. 2024 Nov 8;9:57. Additional Declarations No competing interests reported. Supplementary Files SupplementaryMaterialsBMCMG.docx Cite Share Download PDF Status: Published Journal Publication published 19 Nov, 2025 Read the published version in BMC Medical Genomics → Version 1 posted Editorial decision: Revision requested 08 Sep, 2025 Reviews received at journal 04 Sep, 2025 Reviews received at journal 28 Aug, 2025 Reviewers agreed at journal 08 Aug, 2025 Reviewers agreed at journal 08 Aug, 2025 Reviewers agreed at journal 05 Aug, 2025 Reviewers invited by journal 05 Aug, 2025 Editor assigned by journal 04 Aug, 2025 Editor invited by journal 31 Jul, 2025 Submission checks completed at journal 30 Jul, 2025 First submitted to journal 30 Jul, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7242084","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":496254293,"identity":"e367d344-9731-4701-b53b-56cacb3dbedd","order_by":0,"name":"Javad Jamshidi","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA+0lEQVRIiWNgGAWjYBACxgaGBCB1gIfhMPMBmKABsVrYEojTAgUHQLrgKvFrYW5gePjpxp87MnzHeb49+Nh2T56BvXmbBEPNYXwOS5bObXvGI3mYd7vhzLZiwwaeY2USDMfwakmQzm04zGNwmHebNG9bAmODRI6ZBAMbflt+5/wBaeF5Jv23LcG+Qf4NUMs/vFrSpHPYwFrYpBnbEhIbJHjMJBjb8GhpZkizzm07DPQLm5lkz7mE5DaetGKLxL50nFoM23uSbwMdZs93/vAziR9lCbb97Ic33vjwzRq3lmaeBFQRNhCRgKkSDoDRcACP9CgYBaNgFIwCIAAAxppTniQIocUAAAAASUVORK5CYII=","orcid":"","institution":"Neuroscience Research Australia","correspondingAuthor":true,"prefix":"","firstName":"Javad","middleName":"","lastName":"Jamshidi","suffix":""},{"id":496254294,"identity":"15d68035-bad9-45ad-94fc-0abeb0123951","order_by":1,"name":"Conor Rowntree","email":"","orcid":"","institution":"Neuroscience Research Australia","correspondingAuthor":false,"prefix":"","firstName":"Conor","middleName":"","lastName":"Rowntree","suffix":""},{"id":496254295,"identity":"4b329a90-ee48-4892-95fc-8b73b7d835d6","order_by":2,"name":"Shannon Fadaee","email":"","orcid":"","institution":"Prince of Wales Hospital","correspondingAuthor":false,"prefix":"","firstName":"Shannon","middleName":"","lastName":"Fadaee","suffix":""},{"id":496254296,"identity":"fe949d49-2f53-45d2-9d85-ab96427e17d9","order_by":3,"name":"Futao Zhang","email":"","orcid":"","institution":"Prince of Wales Hospital","correspondingAuthor":false,"prefix":"","firstName":"Futao","middleName":"","lastName":"Zhang","suffix":""},{"id":496254297,"identity":"2a778472-da85-4e66-97d0-968ba4bff11b","order_by":4,"name":"Ying Zhu","email":"","orcid":"","institution":"Prince of Wales Hospital","correspondingAuthor":false,"prefix":"","firstName":"Ying","middleName":"","lastName":"Zhu","suffix":""},{"id":496254301,"identity":"d180751d-f7a0-447b-9620-27cb00f4f59b","order_by":5,"name":"Michael Buckley","email":"","orcid":"","institution":"Prince of Wales Hospital","correspondingAuthor":false,"prefix":"","firstName":"Michael","middleName":"","lastName":"Buckley","suffix":""},{"id":496254303,"identity":"9c1b19f3-cd88-4eff-8a2a-bc343f17998b","order_by":6,"name":"Franki Hart","email":"","orcid":"","institution":"Prince of Wales Hospital","correspondingAuthor":false,"prefix":"","firstName":"Franki","middleName":"","lastName":"Hart","suffix":""},{"id":496254307,"identity":"aae3052a-23b4-497a-b807-36303c4c528a","order_by":7,"name":"Tony Roscioli","email":"","orcid":"","institution":"University of New South Wales","correspondingAuthor":false,"prefix":"","firstName":"Tony","middleName":"","lastName":"Roscioli","suffix":""}],"badges":[],"createdAt":"2025-07-29 10:23:03","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7242084/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7242084/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s12920-025-02251-z","type":"published","date":"2025-11-19T15:58:45+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":88573176,"identity":"f53198d3-15cd-4ace-8739-70fd9ab96779","added_by":"auto","created_at":"2025-08-08 00:11:36","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":435917,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eIGV screenshot of phased reads for the B7 amplicon.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis IGV screenshot illustrates phased long reads, coloured and grouped by haplotype (phase 1 and phase 2). From top to bottom, the tracks display: the reference gene, \u003cem\u003eTMPRSS3\u003c/em\u003e; high-confidence variants for benchmarking (HG002); variants called from the amplicon using Clair3; and the coverage and alignments from long-read sequencing of the B7 amplicon, grouped by phase. Two key variants, chr21:42,372,760 C\u0026gt;T and chr21:42,387,895 A\u0026gt;C, are highlighted by red boxes and arrows, demonstrating their cis configuration (on the same phase). The reads are down-sampled and small indels (\u0026lt;2 bp) are hidden to provide a clearer view.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7242084/v1/47dbfad796ea2ce9dcf7c35b.png"},{"id":88573178,"identity":"1d612328-7125-493e-9045-0f8e9689599f","added_by":"auto","created_at":"2025-08-08 00:11:36","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":176024,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eChimeric Reads Across Different PCR Conditions.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eBox plots display the median and interquartile range (IQR) of chimeric reads for each PCR condition, with individual data points overlaid. The UltraRun LongRange PCR Kit amplified samples for 26 cycles (reference group, blue) and 28 cycles (green). The Platinum SuperFi II PCR Master Mix amplified samples for 26 cycles (orange).\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7242084/v1/4fee46e987a7aa3bbf131f92.png"},{"id":88573175,"identity":"5abaa71f-e89d-4c91-b3dd-8852717ba144","added_by":"auto","created_at":"2025-08-08 00:11:36","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":100301,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eIGV screenshot of short-read WGS and long-read amplicon sequencing of the \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eTUBB2A\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003egene. \u003c/strong\u003eFrom top to bottom, the tracks display: the reference gene, \u003cem\u003eTUBB2A\u003c/em\u003e; high-confidence variants for benchmarking (HG002); coverage and alignments from short-read WGS of HG002 (from GIAB repository); coverage and alignments from long-read sequencing of the TUBB2A amplicon (down-sampled to 176 reads). The red arrow highlights a variant that appears in a low-mappability region in the short-read WGS. This variant is not present in the high-confidence call set and not a real variant. The variant is not present in the long-read amplicon sequencing, demonstrating the utility of this method for accurate variant localisation in low-mappability regions of the genome.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7242084/v1/8ca8486c360f28e6f82dd24e.png"},{"id":96650211,"identity":"07b70df5-91a6-4e21-8994-3097eeddfe8f","added_by":"auto","created_at":"2025-11-24 16:09:47","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1602825,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7242084/v1/561977bf-662b-41a5-9663-c94442f12a86.pdf"},{"id":88573319,"identity":"3f3594b8-4a3d-4f81-b83e-ba472b5dcbad","added_by":"auto","created_at":"2025-08-08 00:19:36","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":182935,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryMaterialsBMCMG.docx","url":"https://assets-eu.researchsquare.com/files/rs-7242084/v1/6a2596b42b827c3d05a1a813.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Long-range PCR and Nanopore sequencing for localisation and phasing variants: an end-to-end clinical application workflow","fulltext":[{"header":"Introduction","content":"\u003cp\u003eNext generation sequencing (NGS) has revolutionised diagnostic genomics, with whole exome (WES) and whole genome (WGS) sequencing having become integral to clinical diagnosis. NGS has several inherent limitations related to short read sequencing (SRS), such as poor alignment in high homology regions of the human genome and the inability to phase variants more than a few hundred bases apart (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e\u003cp\u003ePhasing refers to the process of determining the parental origin of each variant allele either located on the same chromosome (\u003cem\u003ein cis\u003c/em\u003e) or on different chromosomes (\u003cem\u003ein trans\u003c/em\u003e). Phasing can determine whether two alleles sequenced in a single individual are compound heterozygous (\u003cem\u003ein trans\u003c/em\u003e) which is crucial for resolving the inheritance of variants in autosomal recessive conditions, particularly when parental segregation is not possible (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e). In addition, when one parent is heterozygous for both variants or when one variant is \u003cem\u003ede novo\u003c/em\u003e, parental sequencing cannot determine the phase.\u003c/p\u003e\u003cp\u003ePoor alignment occurs in SRS within regions of the genome with large stretches of identical or nearly identical sequences that exist at multiple locations such as tandem repeats, transposable elements, paralogous genes and pseudogenes (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e). In such regions short reads cannot be uniquely aligned to any one location, causing misalignment and poor mapping (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e). Variants detected in regions with low mapping usually require confirmation using an alternative method, such as Sanger sequencing. However, designing specific Sanger primers may not be feasible due to high homology and sequencing length limitations, leaving limited diagnostic tools available to confirm the variants (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e). Long-read sequencing (LRS) can overcome this issue as primers may be designed in unique sequence regions at a distance from the variant of interest.\u003c/p\u003e\u003cp\u003eRecent advances in basecalling accuracy have enabled LRS to address the shortcomings of SRS, and for it to be implemented in diagnostic genomic testing (\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e). Oxford Nanopore Technologies (ONT) offers high accuracy, affordable and scalable LRS (\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e) using the smallest flow cell (Flongle Flow cells) for targeted sequencing, enabling financially feasible small-scale diagnostic assays.\u003c/p\u003e\u003cp\u003eAlthough long-range PCR may amplify longer DNA sequences, challenges may arise with primer design and increased levels of chimeric reads, - a PCR artefact derived from two different biological sequences, affecting sequencing accuracy and phasing (\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e). Therefore, optimising long-range PCR and detection whilst minimising chimeric reads is crucial for successful sequencing applications for diagnostic purposes.\u003c/p\u003e\u003cp\u003eThis study has developed a robust long-range PCR protocol by comparing different PCR kits to maximise successful simultaneous amplification of long DNA targets (~ 1-20kb), a protocol to barcode and sequence multiple long amplicons on Flongle flow cells, and a bioinformatic pipeline to streamline and automate quality control, variant calling, localisation, and phasing from barcoded ONT amplicons. Combining these three steps has resulted in a reliable and affordable workflow for localising or phasing clinically relevant variants up to 20 kb apart, suitable for implementation in clinical diagnostics.\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003e\u003cb\u003eSamples and ethical considerations\u003c/b\u003e:\u003c/p\u003e\u003cp\u003eNA24385 DNA (HG002 from GIAB) (Coriell Institute) was used to set up the PCRs, sequencing and evaluation of variant calling and phasing from amplicons. This project was reviewed by the South Eastern Sydney Local Health District HREC at a meeting of its Low and Negligible Risk Committee, which determined that the project did not involve any ethical risks in accordance with NSW Health Policy.\u003c/p\u003e\u003cp\u003e\u003cb\u003eLong-range PCR kit comparison\u003c/b\u003e\u003c/p\u003e\u003cp\u003eFour different PCR kits with the ability to amplify long targets including Platinum SuperFi II PCR Master Mix (2X) (Invitrogen), LongAmp® Taq 2X Master Mix (New England Biolabs), Q5® Hot Start High-Fidelity 2X Master Mix (New England Biolabs), and UltraRun LongRange PCR Kit (Qiagen) were selected for assessment. Ten DNA regions of the human genome between 1–22 kb in size were chosen for amplification which contained at least one gene, and where each amplicon spanned a minimum of two exons. Primers for these regions were designed using NCBI primer blast (Primer IDs: A1-A10). PCR programs were configured according to the manufacturer’s recommendations for each kit. A single PCR program using a single annealing temperature and extension time was established for each kit to enable running samples simultaneously on a thermal cycler. All PCRs were run in 26 cycles to minimise chimeric reads. All PCRs were conducted in a final volume of 20 µl, using 150 ng of the same DNA sample (NA24385) and 0.5 µM of each forward and reverse primer. After PCR, the amplicons were analysed using the Agilent 4200 Tape Station System. A successful amplicon was defined as the presence of a clear band in the gel image (concentration \u0026gt; 2 ng/µl). Non-specific bands were defined as any additional bands, with a concentration \u0026gt; 2 ng/µl. A detailed description of PCR kit selection including PCR protocols, amplicons and primer sequences are presented in the supplementary document.\u003c/p\u003e\u003cp\u003e\u003cb\u003eNanopore library preparation and sequencing\u003c/b\u003e:\u003c/p\u003e\u003cp\u003eBarcoding genomic DNA using the Flongle protocol (Ligation Sequencing gDNA - Native Barcoding Kit 24 V14, SQK-NBD114.24) was adapted for amplicon library preparation. DNA repair and end-prep step was modified to accommodate the use of amplicons. More detail is available from the supplementary document. Nanopore libraries were prepared using the Ligation Sequencing Kit V14 (SQK-LSK114) and barcoded using Native Barcoding Kit 24 V14 (SQK-NBD114.24). Briefly, up to eight amplicons were multiplexed in each sequencing run with equimolar amounts of each amplicon. 0.75 µl Ultra II End-prep Enzyme Mix and 0.875 µl Ultra II End-prep Reaction Buffer (New England Biolabs #E7546) was added to 5–10 femtomole of each PCR product adjusted to 18.4 µl with nuclease free water. Reactions were incubated at 20°C for 5 minutes and 65°C for 5 minutes and then purified using AMPure XP beads (AXP). Native barcodes were ligated in a reaction with 7.5 µl end-prepped DNA, 2.5 µl Native Barcode, 10 µl Blunt/TA Ligase Master Mix (New England Biolabs #M0367), incubated at room temperature for 20 minutes. EDTA was then added to each reaction and samples were pooled and another AXP purification was performed. Sequencing adaptors were then ligated to the barcoded end-prepped amplicons in a reaction with 30 µl pooled barcoded sample, 5 µl Native Adapter (NA), 10 µl NEBNext Quick Ligation Reaction Buffer (5X) and 5 µl Quick T4 DNA Ligase (New England Biolabs #E6056) incubated for 20 minutes at room temperature, followed by washing with Long Fragment Buffer (LFB) and AXP clean up. The library was eluted in 7 µl of Elution Buffer (EB) and incubated at 37°C for 10 minutes. Ten femtomole of the library was sequenced using a Flongle flow cell (R10.4.1, FLO-FLG114) on a GridION (ONT) device.\u003c/p\u003e\u003cp\u003e\u003cb\u003eBasecalling and alignment\u003c/b\u003e\u003c/p\u003e\u003cp\u003eSuper accuracy basecalling (SUP) (
[email protected]) was performed during sequencing using MinKNOW software (GridION Release 24.06.15 version) which uses dorado (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/nanoporetech/dorado\u003c/span\u003e\u003cspan address=\"https://github.com/nanoporetech/dorado\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) for basecalling and Minimap2 version 2.28 (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e) to align the reads with human reference genome (build hg38) and output BAM files. The minimum phred score was set to ten.\u003c/p\u003e\u003cp\u003e\u003cb\u003eBioinformatic analysis\u003c/b\u003e\u003c/p\u003e\u003cp\u003eSamtools v1.21 (\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e) was employed for BAM file manipulation and quality control. For phasing analysis, reads with mapping quality (MAPQ) \u0026lt; 20 and those shorter than the distance between the two variants were excluded. For variant localisation analysis, reads shorter than 1000 bp, those with read identity \u0026lt; 80%, and MAPQ \u0026lt; 20 were filtered out. Variant calling from amplicons was performed using Clair3 v1.0.4 (\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e), while phasing was conducted using WhatsHap v2.3 (\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e) and HapCUT2 v1.3.4 (\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e). WhatsHap was additionally used to haplotag the BAM files for visualisation. A custom Python script was developed to identify variants, reads, and their corresponding phases. The chimeric read proportion was calculated as the percentage of reads classified as the opposite phase to that determined by WhatsHap and HapCUT2, representing the smaller fraction of discordant reads. All BAM files were visualised and manually inspected using the Integrative Genomics Viewer (IGV) (\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e) to confirm the variants localisation and phase assignment. A bioinformatics pipeline was developed to automate these analyses and generate reports containing both quality control metrics and phase/localisation details. A minimum threshold of 50 high-quality reads (as defined above) was set as the requirement for reliable phasing. All BAM files were down-sampled to 50 reads and re-evaluated to assess whether this threshold was sufficient for accurate phasing.\u003c/p\u003e\u003cp\u003eComplete documentation of the bioinformatics pipeline and associated code are available in the corresponding GitHub repository (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/j-jamshidi/ONT_amp_phase\u003c/span\u003e\u003cspan address=\"https://github.com/j-jamshidi/ONT_amp_phase\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e\u003cp\u003e\u003cb\u003ePhasing samples\u003c/b\u003e\u003c/p\u003e\u003cp\u003eTo assess the performance of the selected long-range PCR kit, sequencing protocol and phasing analysis, ten heterozygote variant pairs with a distance ranging from 6.8–21.4 kb were selected from the NA24385 sample. Primers were designed to span the two variants in each case (Primer IDs: B1-B10). All amplicons were sequenced, and the phases of the variants were determined and compared with the known phase. Details of the variants, primers and amplicons are provided in the supplementary material table S10.\u003c/p\u003e\u003cp\u003e\u003cb\u003eChimeric read additional samples\u003c/b\u003e\u003c/p\u003e\u003cp\u003eEight targets were amplified using (a) Platinum SuperFi II kit (26 cycles) and (b) UltraRun LongRange PCR kit (28 cycles) to assess the impact of different PCR conditions on chimeric read formation. The percentage of chimeric reads were then compared against the UltraRun LongRange PCR kit (26 cycles) baseline condition with the Wilcoxon signed-rank test. A p-value \u0026lt; 0.05 was considered statistically significant.\u003c/p\u003e\u003cp\u003e\u003cb\u003eLow mappability genes\u003c/b\u003e\u003c/p\u003e\u003cp\u003eTo evaluate the suitability of the method to call and localise variants in low mappability regions of the human genome, eight genes (\u003cem\u003eTUBB2A\u003c/em\u003e, \u003cem\u003eTUBB2B\u003c/em\u003e, \u003cem\u003eCYP11B1\u003c/em\u003e, \u003cem\u003eSBDS\u003c/em\u003e, \u003cem\u003eHBA1\u003c/em\u003e, \u003cem\u003eHBA2\u003c/em\u003e, \u003cem\u003eRBMX\u003c/em\u003e, and \u003cem\u003eCEL\u003c/em\u003e) were selected. These genes are associated with Mendelian disorders and contain regions of low mappability. The details of these genes and their associated phenotypes are summarised in Table S11. Primers were designed to amplify the entire length of each gene. \u003cem\u003eHBA1\u003c/em\u003e and \u003cem\u003eHBA2\u003c/em\u003e are in close proximity in the human genome, therefore, a single primer pair was designed to span the region covering both genes. Details of the primers and amplicons are provided in the supplementary table S12. Long-range PCR and sequencing were performed similarly to the phasing cases. The called SNVs and small indels from amplicon sequencing of these genes (excluding \u003cem\u003eRBMX\u003c/em\u003e, which is located on the X chromosome and not included in the benchmarking VCF file) were compared with GIAB ground-truth variants (benchmark version v4.2.1) (\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e), using the \u003cem\u003evcfeval\u003c/em\u003e command from RTG Tools v3.13 (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/RealTimeGenomics/rtg-tools\u003c/span\u003e\u003cspan address=\"https://github.com/RealTimeGenomics/rtg-tools\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003e\u003cb\u003eLong-range PCR\u003c/b\u003e\u003c/p\u003e\u003cp\u003eAmong the four PCR kits tested, the UltraRun LongRange PCR Kit demonstrated the best performance, successfully amplifying 9 out of 10 amplicons (90% success rate). Platinum SuperFi II and LongAmp\u0026reg; Taq showed a similar success rate of 70% (7 out of 10). The Q5\u0026reg; Hot Start High-Fidelity kit had the lowest success rate, successfully amplifying only four amplicons. Therefore, the UltraRun LongRange PCR Kit was selected for further amplifications. Details on the performance of each kit is available in the supplementary material, table S4-S7. All ten primer pairs designed for phasing (B1-B10), amplified their target regions successfully, although one amplicon (B1) required the addition of Q Solution, a PCR additive included with the UltraRun LongRange PCR Kit to enhance amplification.\u003c/p\u003e\u003cp\u003eSix of the seven PCR primers designed for the amplification of low-mappability genes produced a product (\u003cem\u003eTUBB2A\u003c/em\u003e and \u003cem\u003eHBA1/HBA2\u003c/em\u003e required Q solution), whereas \u003cem\u003eCEL\u003c/em\u003e primers did not yield a product.\u003c/p\u003e\u003cp\u003e\u003cb\u003eNanopore sequencing\u003c/b\u003e\u003c/p\u003e\u003cp\u003e15 amplicons that were amplified with UltraRun LongRange PCR kit in 26 cycles (5 amplicons that were used for testing different PCR kits and 10 that were designed specifically for phasing) were sequenced on Flongle flow cells. All amplicons had high depth of high quality (MAPQ\u0026thinsp;\u0026lt;\u0026thinsp;20) reads covering both variants of interest (mean\u0026thinsp;=\u0026thinsp;1920, median\u0026thinsp;=\u0026thinsp;3248, standard deviation\u0026thinsp;=\u0026thinsp;1716.8, range\u0026thinsp;=\u0026thinsp;286\u0026ndash;5483). Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e summarises the sequencing details.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eList of amplicons sequenced for variant phasing\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"8\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAmplicon ID\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eAmplicon Size\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eVariant 1 (hg38)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eVariant 2 (hg38)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eDistance between variants (bp)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eHQ spanning reads*\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c7\"\u003e\u003cp\u003eChimeric reads %\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c8\"\u003e\u003cp\u003eKnown\u003c/p\u003e\u003cp\u003ephase\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eB1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e8743\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003echr16:88,643,329 C\u0026thinsp;\u0026gt;\u0026thinsp;T\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003echr16:88,650,126 A\u0026thinsp;\u0026gt;\u0026thinsp;G\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e6797\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e5483\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e13.97\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003ecis\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eA3\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e9625\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003echr1:154,579,450 C\u0026thinsp;\u0026gt;\u0026thinsp;T\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003echr1:154,585,209 C\u0026thinsp;\u0026gt;\u0026thinsp;T\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e5759\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e5385\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e4.25\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003ecis\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eB2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e10026\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003echr12:8,927,591 T\u0026thinsp;\u0026gt;\u0026thinsp;C\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003echr12:8,936,676 A\u0026thinsp;\u0026gt;\u0026thinsp;C\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e9085\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e4413\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e16.12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003etrans\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eB3\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e11521\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003echr12:6,936,914 G\u0026thinsp;\u0026gt;\u0026thinsp;A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003echr12:6,944,269 C\u0026thinsp;\u0026gt;\u0026thinsp;T\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e7355\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e1748\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e3.03\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003ecis\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eB4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e13789\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003echr4:41,638,861 C\u0026thinsp;\u0026gt;\u0026thinsp;T\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003echr4:41,651,881 G\u0026thinsp;\u0026gt;\u0026thinsp;A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e13020\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e673\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e1.93\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003ecis\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eB5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e14391\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003echr3:49,412,205 G\u0026thinsp;\u0026gt;\u0026thinsp;A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003echr3:49,425,854 G\u0026thinsp;\u0026gt;\u0026thinsp;A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e13649\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e3165\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e2.62\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003ecis\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eB6\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e14466\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003echr21:42,375,588 G\u0026thinsp;\u0026gt;\u0026thinsp;A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003echr21:42,388,818 A\u0026thinsp;\u0026gt;\u0026thinsp;G\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e13230\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e4450\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e3.33\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003ecis\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eA4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e15163\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003echr6:152,389,867 T\u0026thinsp;\u0026gt;\u0026thinsp;C\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003echr6:152,403,191 G\u0026thinsp;\u0026gt;\u0026thinsp;T\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e13324\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e3349\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e2.51\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003ecis\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eA5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e16084\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003echr1:236,544,737 T\u0026thinsp;\u0026gt;\u0026thinsp;C\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003echr1:236,559,878 C\u0026thinsp;\u0026gt;\u0026thinsp;A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e15141\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e3248\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e5.48\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003etrans\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eA6\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e16837\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003echr1:43,658,540 G\u0026thinsp;\u0026gt;\u0026thinsp;T\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003echr1:43,667,233 A\u0026thinsp;\u0026gt;\u0026thinsp;G\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e8693\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e3792\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e2.24\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003ecis\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eB7\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e17828\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003echr21:42,372,760 C\u0026thinsp;\u0026gt;\u0026thinsp;T\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003echr21:42,387,895 A\u0026thinsp;\u0026gt;\u0026thinsp;C\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e15135\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e906\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e2.32\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003ecis\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eA7\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e18097\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003echr1:6,629,764 A\u0026thinsp;\u0026gt;\u0026thinsp;G\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003echr1:6,645,884 C\u0026thinsp;\u0026gt;\u0026thinsp;T\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e16120\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e3,633\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e1.79\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003ecis\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eB8\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e21530\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003echr3:48,566,437 G\u0026thinsp;\u0026gt;\u0026thinsp;A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003echr3:48,584,160 T\u0026thinsp;\u0026gt;\u0026thinsp;C\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e17723\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e709\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e6.77\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003ecis\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eB9\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e22097\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003echr9:130,407,156 A\u0026thinsp;\u0026gt;\u0026thinsp;G\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003echr9:130,426,008 C\u0026thinsp;\u0026gt;\u0026thinsp;T\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e18852\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e2570\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e2.80\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003etrans\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eB10\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e22409\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003echr1:236,011,853 T\u0026thinsp;\u0026gt;\u0026thinsp;C\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003echr1:236,033,263 A\u0026thinsp;\u0026gt;\u0026thinsp;G\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e21410\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e286\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e2.10\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003ecis\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003csup\u003e*\u003c/sup\u003e HQ Spanning reads: High quality reads (MAPQ\u0026thinsp;\u0026gt;\u0026thinsp;20) that span both variants of interest (i.e., variants to be phased).\u003c/p\u003e\u003cp\u003eAdditionally, all amplicons amplified using Platinum SuperFi II PCR Master Mix and those amplified with 28 cycles using the UltraRun LongRange PCR Kit were also sequenced. The sequencing details are available in supplementary documents Tables S13 and S14. The depth of sequencing for each amplicon targeting low mappability genes was \u0026gt;\u0026thinsp;500x after QC, see supplementary tables 15.\u003c/p\u003e\u003cp\u003e\u003cb\u003ePhasing\u003c/b\u003e\u003c/p\u003e\u003cp\u003eFifteen variant pairs in amplicons generated using the UltraRun LongRange PCR Kit were phased and compared with the known phase. All variants were SNVs, with inter-variant distances ranging from 5,759 to 21,410 bp. All 15 variant pairs were phased correctly and were concordant with the known phase (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Furthermore, to test the robustness of this pipeline for phasing small indels, ten indels were identified and phased in the existing amplicons. All the indels were also phased correctly (table S16). Moreover, re-evaluating all cases with only 50 high quality reads per sample did not compromise phasing accuracy. Figure\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e shows an IGV screenshot of the phased reads for the B7 amplicon, highlighting the positions of the selected variants used for phasing as well as the variants called within the amplicon and the benchmarking variants.\u003c/p\u003e\u003cb\u003eChimeric reads\u003c/b\u003e:\u003c/p\u003e\u003cp\u003eThe estimated proportion of chimeric reads in the fifteen phased amplicons ranged from 1.79\u0026ndash;16.12% (mean\u0026thinsp;=\u0026thinsp;4.75, median\u0026thinsp;=\u0026thinsp;2.80, SD\u0026thinsp;=\u0026thinsp;4.41). The percentage of chimeric reads for each amplicon is presented in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e\u003cp\u003e\u003cb\u003eChimeric reads in different PCR conditions\u003c/b\u003e\u003c/p\u003e\u003cp\u003eThe UltraRun LongRange (26 cycles) baseline condition yielded a median chimeric read percentage of 3.64% across the eight samples. Compared to this baseline, amplification with Platinum SuperFi II (26 cycles) (median\u0026thinsp;=\u0026thinsp;9.33%; W-statistic\u0026thinsp;=\u0026thinsp;1.0, p\u0026thinsp;=\u0026thinsp;0.017) and increasing the cycles to 28 (median\u0026thinsp;=\u0026thinsp;6.56%; W-statistic\u0026thinsp;=\u0026thinsp;1.0, p\u0026thinsp;=\u0026thinsp;0.017) resulted in significantly higher chimeric read levels. See Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e and Table S17.\u003c/p\u003e\u003cp\u003e\u003cb\u003eVariant calling from low mappability genes amplicons\u003c/b\u003e:\u003c/p\u003e\u003cp\u003eThe SNVs called via the pipeline with Clair3 from the amplicons generated for low mappability genes (\u003cem\u003eTUBB2A\u003c/em\u003e, \u003cem\u003eTUBB2B\u003c/em\u003e, \u003cem\u003eCYP11B1\u003c/em\u003e, \u003cem\u003eSBDS\u003c/em\u003e, \u003cem\u003eHBA1\u003c/em\u003e and \u003cem\u003eHBA2\u003c/em\u003e, 64 variants in total\u003cem\u003e)\u003c/em\u003e showed 100% concordance with the benchmark (high confidence) variants (Precision\u0026thinsp;=\u0026thinsp;1, Sensitivity\u0026thinsp;=\u0026thinsp;1). Figure\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e illustrates short-read WGS and long-read amplicon sequencing of the \u003cem\u003eTUBB2A\u003c/em\u003e gene.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eThis study developed and validated an integrated, end-to-end clinical diagnostic workflow combining optimised LR-PCR, targeted ONT sequencing, and an in-house developed bioinformatic pipeline for variant phasing and localisation. The workflow achieved accurate phasing of variants up to 21.4 kb apart and precise variant localisation in low-mappability genomic regions, addressing critical limitations of current short-read approaches. Methodologically, the study presents a comprehensively optimised workflow. This rigorous approach involved comparing LR-PCR kits systematically, optimising PCR conditions, developing a modified Nanopore library preparation protocol for barcoding on Flongle flow cells, and adjusting the multiplexing number and molarity of amplicons. Complementing these laboratory improvements is an automated bioinformatic pipeline integrating established tools like Clair3 for variant calling, and WhatsHap and HapCUT2 for phasing, alongside custom scripts for quality control and specific analytical logic for phasing versus localisation tasks. This end-to-end development, tailored for a cost-constrained yet high-performance application, represents a significant advancement beyond existing tools.\u003c/p\u003e\u003cp\u003eLR-PCR is a recognised versatile technique for amplifying large genomic segments (\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e), but its success is highly dependent on careful optimisation, limiting its applications in targeted long-read sequencing (\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e). In the current study, different LR-PCR kits were compared to optimise the workflow, with the UltraRun LongRange PCR Kit demonstrating the best performance by successfully amplifying nine out of ten targets. Both Platinum SuperFi II and LongAmp\u0026reg; Taq also performed well, each amplifying seven targets successfully. Notably, the three amplicons that failed with both kits were the longest in the panel (all \u0026gt;\u0026thinsp;18 kb), consistent with the difficulty of amplify ultra-long fragments. The Q5\u0026reg; Hot Start High-Fidelity kit demonstrated the weakest performance, with only four successful amplifications, likely due to the use of a fixed annealing temperature, despite the kit's recommendation for primer-specific annealing conditions. A universal annealing temperature and extension time is necessary to allow simultaneous amplification of samples, facilitate batching and streamlining the workflow, which are key considerations for implementation in a diagnostic laboratory.\u003c/p\u003e\u003cp\u003eAlthough some regions slightly over 20 kb were amplified, limiting amplicons to 20 kb is recommended due to the technical challenges of amplifying longer targets (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e). An internal analysis of a curated panel of 5,678 genes associated with Mendelian disorders used in our laboratory showed a median gene size of approximately 38 kb. In this dataset for any two random variants within a gene, there is a\u0026thinsp;~\u0026thinsp;68% chance they would lie within a 20kb range (figure \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e and Supplementary Data for detailed analysis). 20kb should therefore be a sufficient length for diagnostic utility for the majority of variants requiring phasing in a clinical context.\u003c/p\u003e\u003cp\u003eVariants\u0026thinsp;\u0026gt;\u0026thinsp;20 kb apart could be potentially phased using overlapping amplicons, however, this approach adds complexity and may not always be feasible (\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e). While this strategy can be effective for targeted phasing of specific loci (\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e), it is less practical for general diagnostic testing. Given the generality and focus of this method on clinical implementation, we chose not to pursue this strategy to avoid unnecessary complexity.\u003c/p\u003e\u003cp\u003eAn important technical consideration in amplicon-based phasing is the formation of chimeric reads during PCR, which can falsely link variants on different haplotypes (\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e). Unlike microbial diversity studies, which benefit from established tools for filtering out chimeric reads (\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e), phasing applications lack robust methods for chimera removal, largely because the parental haplotypes may be nearly identical and sometimes differing only by the variants being phased (\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e). Furthermore, approximately 1.7% of reads in nanopore-based amplicon sequencing may contain post-amplification chimeric elements (\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e). Therefore, detecting and minimising chimeric reads is critical for reliable phasing using amplicon-based approaches.\u003c/p\u003e\u003cp\u003ePrevious studies have demonstrated that failure to account for chimeric reads can lead to incorrect phasing interpretations (\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e). The number of PCR cycles was reduced to 26 to reduce chimera formation, a strategy supported by earlier studies (\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e) and confirmed by this study, where increasing the cycle number to 28 led to higher chimeric reads. The UltraRun LongRange PCR Kit was observed to generate fewer chimeric reads than Platinum SuperFi II on the same samples, further supporting its selection. This difference may arise from the polymerases used or protocol variations; notably, UltraRun uses a two-step PCR protocol, which has been associated with lower chimera formation rates (\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eBeyond PCR optimisation, the developed bioinformatic pipeline included a method for quantifying chimeric reads based solely on the two variants under phasing. This approach contrasts with other studies that relied on external software or manual review, requiring extra steps for chimeric detection (\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e). In the primary sample, chimeric read levels were between 1.79\u0026ndash;16.12% (median 2.8%), consistent with previously reported figures (7,20). While a median rate of 2.8% allows for accurate phasing, the broad range highlights the potential for substantial chimera presence in specific amplicons. This variability underscores the importance of careful, amplicon-specific assessment of chimeric read rates, especially in diagnostic contexts where precision is critical. Notably, the developed pipeline maintained robust and accurate phasing performance even in samples with chimeric read levels as high as 27%, observed in Platinum SuperFi II amplicons (see Table S17).\u003c/p\u003e\u003cp\u003eLR-PCR and nanopore sequencing have been employed for phasing variants (\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e, \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e, \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e, \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e) and localising those identified in low-mappability regions (\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e, \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e) in previous studies. However, these studies focused on specific genomic regions, such as haplotyping the \u003cem\u003eABCA4\u003c/em\u003e locus (\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e), or the localisation of variants within the highly repetitive domains of \u003cem\u003eTTN\u003c/em\u003e (\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e). This study has developed a robust and reliable workflow adaptable to any genomic region that can be feasibly amplified using LR-PCR.\u003c/p\u003e\u003cp\u003eNotably, previous similar studies have typically used MinION flow cells for multiple amplicons in nanopore-based assays (\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e, \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e, \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e) while Flongle flow cells have been used often for single assay sequencing (\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e, \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e, \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e). Sequencing multiple barcoded amplicons on a Flongle flow cells was a key feature of our assay design, offering an affordable and practical solution, particularly for smaller diagnostic laboratories.\u003c/p\u003e\u003cp\u003eWhile the workflow described here is adaptable to larger flow cells such as the MinION, achieving similar cost efficiency to Flongle would require a larger number of samples per run. In a diagnostic setting, accumulating sufficient samples for high-level multiplexing could lead to elevated turnaround times (\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e). Flongle-based multiplexing therefore currently offers the most efficient and affordable solution for clinically relevant applications and supports the broader adoption of targeted nanopore sequencing in clinical diagnostics.\u003c/p\u003e\u003cp\u003eThe median sequencing depth per amplicon in this study exceeded 3,000x, consistent with the multiplexing of more than eight amplicons per Flongle flow cell being technically feasible. However, variability in ONT flow cell performance relating to active pore count, and significant variation in depth, despite using equimolar quantities of PCR products (\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e) can affect stability and reliability, which are critical in diagnostic settings. Limiting the number of amplicons per run therefore ensures consistent performance, even when sequencing output is suboptimal. Additionally, as with MinION flow cells, higher multiplexing may delay turnaround time due to the need to batch more samples. Thus, using eight amplicons per run balances cost-efficiency, test robustness, and clinical practicality.\u003c/p\u003e\u003cp\u003eThe developed workflow has significant potential to enhance molecular diagnostics. Accurate phasing of variants is crucial for the interpretation of results in autosomal recessive Mendelian conditions, where it is necessary to determine if two heterozygous variants in a gene are \u003cem\u003ein cis\u003c/em\u003e or \u003cem\u003ein trans\u003c/em\u003e to confirm compound heterozygosity (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e). Misinterpretation of phase can lead to incorrect diagnostic conclusions. By providing direct phasing from proband DNA, this workflow can clarify the clinical significance of co-occurring variants. Furthermore, the ability to reliably call variants in low-mappability regions is of high clinical importance. Many disease-associated genes contain segments with high homology to other genomic locations (e.g., pseudogenes) or repetitive sequences, which can lead to variants being missed or incorrectly called by standard NGS approaches (\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e). The successful validation of variant calling in genes like \u003cem\u003eTUBB2A/TUBB2B\u003c/em\u003e, \u003cem\u003eCYP11B1\u003c/em\u003e, and \u003cem\u003eHBA1\u003c/em\u003e/\u003cem\u003eHBA2\u003c/em\u003e demonstrates the workflow's utility in these challenging but clinically important regions. The cost-effectiveness and targeted nature of this workflow positions it as an ideal \"Tier 2\" diagnostic test. It can be employed to resolve ambiguous findings from initial WES or WGS, such as phasing variants of uncertain significance (VUS) to aid in their classification, or to confirm variants detected in regions where NGS data quality is suboptimal. The routine availability of such assays could lead to faster, more accurate diagnoses and reduce the number of unresolved cases and ultimately improving patient care.\u003c/p\u003e\u003cp\u003eWhile the study demonstrates considerable strengths, certain limitations should be acknowledged. The LR-PCR protocol, despite optimisation and selection of the best-performing kit, did not achieve universal amplification for all attempted targets. For example, primers designed for the \u003cem\u003eCEL\u003c/em\u003e gene failed to yield an amplicon, and several other amplicons (\u003cem\u003eCYB\u003c/em\u003e, \u003cem\u003eTUBB2A\u003c/em\u003e, \u003cem\u003eHBA\u003c/em\u003e) required the addition of a PCR additive (Q Solution) for successful amplification. This is consistent with certain genomic regions remaining challenging for LR-PCR amplification, potentially due to extreme GC content, secondary structures, or other sequence-specific features (\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e). Such targets might require further bespoke optimisation, alternative primer designs, or different amplification strategies.\u003c/p\u003e\u003cp\u003eThe proportion of chimeric reads, while generally low with the optimised protocol, showed variability across amplicons, with some exhibiting higher levels. Although not observed in this study, amplicons that consistently produce a high percentage of chimeric reads could compromise the accuracy of phasing or variant detection if not controlled stringently and bioinformatically flagged. This highlights the importance of establishing clear quality control thresholds for chimeric read proportions for each diagnostic target.\u003c/p\u003e\u003cp\u003eThis workflow proved robust for detecting SNVs and small indels, however, it was not designed for other variant types such as structural variants or short tandem repeats (STRs) and is therefore limited to SNVs and indels. Moreover, ONT sequencing has a higher error rate for indels than for SNVs (\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e), warranting extra caution when interpreting such variants. Detection of variants located within homopolymer regions also are challenging using ONT sequencing (\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e), placing them outside the scope of this workflow. Furthermore, although the workflow achieved a precision and sensitivity of 1 for variant calling in the selected genes containing low-mappability regions, it is important to note that this evaluation was based on partial genomic regions covering only 64 variants. In larger datasets involving more genes and samples, a decrease in precision and sensitivity is expected (\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eThe variable performance of ONT flow cells can introduce variability across sequencing batches, potentially impacting the stability and reliability of the assay. Although our workflow was designed to mitigate this through specific wet-lab (e.g., limiting amplicon numbers) and bioinformatic strategies, this intrinsic limitation remains. The bioinformatic pipeline assists in managing this by flagging suboptimal results, such as instances where fewer than 50 high-quality reads are available to support phase interpretation or variant localisation.\u003c/p\u003e"},{"header":"Conclusions","content":"\u003cp\u003eThis study has developed and validated a robust, accurate, and cost-effective targeted long-read sequencing workflow using LR-PCR and ONT Flongle technology. By addressing the challenges of variant phasing and analysis of low-mappability genomic regions effectively, this workflow offers a significant advancement over conventional short-read NGS methods for specific, complex diagnostic questions. The comprehensive optimisation from sample preparation through to bioinformatic analysis provides a practical solution with the potential to enhance molecular diagnostics, clarify ambiguous genetic findings, and ultimately improve patient care by enabling more precise genetic diagnoses where current methodologies may fall short.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis project was reviewed by the South Eastern Sydney Local Health District Human Research Ethics Committee (HREC) at a meeting of its Low and Negligible Risk (LNR) Committee, which determined that the project did not involve any ethical risks and was exempt from full ethical review, in accordance with the \u003cem\u003eNHMRC National Statement (updated 2023) \u0026mdash;Section 5.1.22\u003c/em\u003e, the \u003cem\u003eNHMRC Ethical Considerations in Quality Assurance and Evaluation Activities (2014) guidance\u003c/em\u003e and \u003cem\u003eNSW Health Guideline GL2007_020 Human Research Ethics Committees - Quality Improvement and Ethics Review: A Practice Guide for NSW\u003c/em\u003e. The study was conducted in accordance with the Declaration of Helsinki. This study does not include human participants and only publicly available, de-identified human genomic data (HG002 sample from the Genome in a Bottle Consortium) was used; no informed consent was required.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe sequencing data used in this study were generated from the HG002 reference sample, which is publicly available through the Genome in a Bottle (GIAB) consortium and can be accessed at https://www.nist.gov/programs-projects/genome-bottle.\u003c/p\u003e\n\u003cp\u003eAll documentation, source code, and example data for the bioinformatics pipeline developed in this study are available in the associated GitHub repository: https://github.com/j-jamshidi/ONT_amp_phase\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study was supported by the University of New South Wales Sydney and the Medical Research Futures Fund PreGen program (GHFMPACI000006).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026apos; contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eJJ, TR and FH conceptualised and designed the study. JJ and CR performed the PCR and sequencing experiments. JJ, FZ, and YZ analysed the data and developed the bioinformatic pipeline. JJ wrote the draft of the manuscript, and TR, MB, SF, FH, FZ, YZ, and CR reviewed and edited the main text. TR and MB provided the funding and resources. All authors read and approved the final version of the manuscript.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eBonfiglio F, Legati A, Lasorsa VA, Palombo F, De Riso G, Isidori F, et al. Best practices for germline variant and DNA methylation analysis of second- and third-generation sequencing data. Hum Genomics. 2024 Nov 5;18(1):120. \u003c/li\u003e\n\u003cli\u003eTewhey R, Bansal V, Torkamani A, Topol EJ, Schork NJ. The importance of phase information for human genomics. Nat Rev Genet. 2011 Mar;12(3):215\u0026ndash;23. \u003c/li\u003e\n\u003cli\u003eRojahn S, Hambuch T, Adrian J, Gafni E, Gileta A, Hatchell H, et al. Scalable detection of technically challenging variants through modified next-generation sequencing. Mol Genet Genomic Med. 2022;10(12):e2072. \u003c/li\u003e\n\u003cli\u003eMandelker D, Schmidt RJ, Ankala A, McDonald Gibson K, Bowser M, Sharma H, et al. Navigating highly homologous genes in a molecular diagnostic setting: a resource for clinical next-generation sequencing. Genet Med. 2016 Dec;18(12):1282\u0026ndash;9. \u003c/li\u003e\n\u003cli\u003eWarburton PE, Sebra RP. Long-Read DNA Sequencing: Recent Advances and Remaining Challenges. Annu Rev Genomics Hum Genet. 2023 Aug 25;24(Volume 24, 2023):109\u0026ndash;32. \u003c/li\u003e\n\u003cli\u003eEspinosa E, Bautista R, Larrosa R, Plata O. Advancements in long-read genome sequencing technologies and algorithms. Genomics. 2024 May 1;116(3):110842. \u003c/li\u003e\n\u003cli\u003eLaver TW, Caswell RC, Moore KA, Poschmann J, Johnson MB, Owens MM, et al. Pitfalls of haplotype phasing from amplicon-based long-read sequencing. Sci Rep. 2016 Feb 17;6(1):21746. \u003c/li\u003e\n\u003cli\u003eMaestri S, Maturo MG, Cosentino E, Marcolungo L, Iadarola B, Fortunati E, et al. A Long-Read Sequencing Approach for Direct Haplotype Phasing in Clinical Settings. Int J Mol Sci. 2020 Jan;21(23):9177. \u003c/li\u003e\n\u003cli\u003eLi H. Minimap2: pairwise alignment for nucleotide sequences. Bioinforma Oxf Engl. 2018 Sep 15;34(18):3094\u0026ndash;100. \u003c/li\u003e\n\u003cli\u003eLi H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, et al. The Sequence Alignment/Map format and SAMtools. Bioinforma Oxf Engl. 2009 Aug 15;25(16):2078\u0026ndash;9. \u003c/li\u003e\n\u003cli\u003eZheng Z, Li S, Su J, Leung AWS, Lam TW, Luo R. Symphonizing pileup and full-alignment for deep learning-based long-read variant calling. Nat Comput Sci. 2022 Dec;2(12):797\u0026ndash;803. \u003c/li\u003e\n\u003cli\u003eMartin M, Ebert P, Marschall T. Read-Based Phasing and Analysis of Phased Variants with WhatsHap. Methods Mol Biol Clifton NJ. 2023;2590:127\u0026ndash;38. \u003c/li\u003e\n\u003cli\u003eEdge P, Bafna V, Bansal V. HapCUT2: robust and accurate haplotype assembly for diverse sequencing technologies. Genome Res. 2016 Dec 9;gr.213462.116. \u003c/li\u003e\n\u003cli\u003eRobinson JT, Thorvaldsd\u0026oacute;ttir H, Wenger AM, Zehir A, Mesirov JP. Variant Review with the Integrative Genomics Viewer. Cancer Res. 2017 Oct 31;77(21):e31\u0026ndash;4. \u003c/li\u003e\n\u003cli\u003eKrusche P, Trigg L, Boutros PC, Mason CE, De La Vega FM, Moore BL, et al. Best Practices for Benchmarking Germline Small Variant Calls in Human Genomes. Nat Biotechnol. 2019 May;37(5):555\u0026ndash;60. \u003c/li\u003e\n\u003cli\u003eChowdhary MR, Gupta N. Long-range PCR in next generation sequencing: A low cost approach for large \u0026amp; complex genes. Indian J Med Res. 2023 Jun;157(6):591\u0026ndash;2. \u003c/li\u003e\n\u003cli\u003eKee PS, Karunanathie H, Maggo SDS, Kennedy MA, Chua EW. Long-Range Polymerase Chain Reaction. In: Domingues L, editor. PCR: Methods and Protocols [Internet]. New York, NY: Springer US; 2023 [cited 2025 Jul 1]. p. 181\u0026ndash;92. Available from: https://doi.org/10.1007/978-1-0716-3358-8_15\u003c/li\u003e\n\u003cli\u003eMaggi J, Koller S, B\u0026auml;hr L, Feil S, Kivrak Pfiffner F, Hanson JVM, et al. Long-Range PCR-Based NGS Applications to Diagnose Mendelian Retinal Diseases. Int J Mol Sci. 2021 Jan;22(4):1508. \u003c/li\u003e\n\u003cli\u003ePiovesan A, Antonaros F, Vitale L, Strippoli P, Pelleri MC, Caracausi M. Human protein-coding genes and gene feature statistics in 2019. BMC Res Notes. 2019 Jun 4;12(1):315. \u003c/li\u003e\n\u003cli\u003eMcClinton B, Watson CM, Crinnion LA, McKibbin M, Ali M, Inglehearn CF, et al. Haplotyping Using Long-Range PCR and Nanopore Sequencing to Phase Variants: Lessons Learned From the \u003cem\u003eABCA4\u003c/em\u003e Locus. Lab Invest. 2023 Aug 1;103(8):100160. \u003c/li\u003e\n\u003cli\u003eGueuning M, Thun GA, Wittig M, Galati AL, Meyer S, Trost N, et al. Haplotype sequence collection of \u003cem\u003eABO\u003c/em\u003e blood group alleles by long-read sequencing reveals putative \u003cem\u003eA1\u003c/em\u003e-diagnostic variants. Blood Adv. 2023 Mar 28;7(6):878\u0026ndash;92. \u003c/li\u003e\n\u003cli\u003eCallahan BJ, Grinevich D, Thakur S, Balamotis MA, Yehezkel TB. Ultra-accurate microbial amplicon sequencing with synthetic long reads. Microbiome. 2021 Jun 5;9(1):130. \u003c/li\u003e\n\u003cli\u003eWhite R, Pellefigues C, Ronchese F, Lamiable O, Eccles D. Investigation of chimeric reads using the MinION. F1000Research. 2017 Aug 16;6:631. \u003c/li\u003e\n\u003cli\u003eNamias A, Sahlin K, Makoundou P, Bonnici I, Sicard M, Belkhir K, et al. Nanopore sequencing of PCR products enables multicopy gene family reconstruction. Comput Struct Biotechnol J. 2023 Jan 1;21:3656\u0026ndash;64. \u003c/li\u003e\n\u003cli\u003eQin Y, Wu L, Zhang Q, Wen C, Van Nostrand JD, Ning D, et al. Effects of error, chimera, bias, and GC content on the accuracy of amplicon sequencing. mSystems. 2023 Dec;8(6):e01025-23. \u003c/li\u003e\n\u003cli\u003eKroll F, Dimitriadis A, Campbell T, Darwent L, Collinge J, Mead S, et al. Prion protein gene mutation detection using long-read Nanopore sequencing. Sci Rep. 2022 May 18;12(1):8284. \u003c/li\u003e\n\u003cli\u003eDurkie M, Watson CM, Winship P, Hogg AC, Nyanhete R, Cooley S, et al. The Common PKD1 p.(Ile3167Phe) Variant Is Hypomorphic and Associated with Very Early Onset, Biallelic Polycystic Kidney Disease. Hum Mutat. 2023;2023(1):5597005. \u003c/li\u003e\n\u003cli\u003eMc Clinton B, Crinnion LA, Ali M, Inglehearn C, Watson CM, Toomes C. Using Oxford Nanopore long-range sequencing to phase ABCA4. Invest Ophthalmol Vis Sci. 2021 Jun 21;62(8):1541. \u003c/li\u003e\n\u003cli\u003ePerrin A, Van Goethem C, Th\u0026egrave;ze C, Puechberty J, Guignard T, Lecardonnel B, et al. Long-Reads Sequencing Strategy to Localize Variants in \u003cem\u003eTTN\u003c/em\u003e Repeated Domains. J Mol Diagn. 2022 Jul 1;24(7):719\u0026ndash;26. \u003c/li\u003e\n\u003cli\u003eWatson CM, Dean P, Camm N, Bates J, Carr IM, Gardiner CA, et al. Long-read nanopore sequencing resolves a TMEM231 gene conversion event causing Meckel\u0026ndash;Gruber syndrome. Hum Mutat. 2020;41(2):525\u0026ndash;31. \u003c/li\u003e\n\u003cli\u003eHolt GS, Batty LE, Alobaidi BKS, Smith HE, Oud MS, Ramos L, et al. Phasing of de novo mutations using a scaled-up multiple amplicon long-read sequencing approach. Hum Mutat. 2022;43(11):1545\u0026ndash;56. \u003c/li\u003e\n\u003cli\u003eStockton JD, Nieto T, Wroe E, Poles A, Inston N, Briggs D, et al. Rapid, highly accurate and cost-effective open-source simultaneous complete HLA typing and phasing of class I and II alleles using nanopore sequencing. HLA. 2020;96(2):163\u0026ndash;78. \u003c/li\u003e\n\u003cli\u003eJeck WR, Iafrate AJ, Nardi V. Nanopore Flongle Sequencing as a Rapid, Single-Specimen Clinical Test for Fusion Detection. J Mol Diagn. 2021 May 1;23(5):630\u0026ndash;6. \u003c/li\u003e\n\u003cli\u003eWatson CM, Holliday DL, Crinnion LA, Bonthron DT. Long-read nanopore DNA sequencing can resolve complex intragenic duplication/deletion variants, providing information to enable preimplantation genetic diagnosis. Prenat Diagn. 2022;42(2):226\u0026ndash;32. \u003c/li\u003e\n\u003cli\u003eTafess K, Ng TTL, Lao HY, Leung KSS, Tam KKG, Rajwani R, et al. Targeted-Sequencing Workflows for Comprehensive Drug Resistance Profiling of Mycobacterium tuberculosis Cultures Using Two Commercial Sequencing Platforms: Comparison of Analytical and Diagnostic Performance, Turnaround Time, and Cost. Clin Chem. 2020 Jun 1;66(6):809\u0026ndash;20. \u003c/li\u003e\n\u003cli\u003eWhitford W, Hawkins V, Moodley KS, Grant MJ, Lehnert K, Snell RG, et al. Proof of concept for multiplex amplicon sequencing for mutation identification using the MinION nanopore sequencer. Sci Rep. 2022 May 20;12(1):8572. \u003c/li\u003e\n\u003cli\u003eGuo MH, Francioli LC, Stenton SL, Goodrich JK, Watts NA, Singer-Berk M, et al. Inferring compound heterozygosity from large-scale exome sequencing data. Nat Genet. 2024 Jan;56(1):152\u0026ndash;61. \u003c/li\u003e\n\u003cli\u003eClaes KBM, Rosseel T, De Leeneer K. Dealing with Pseudogenes in Molecular Diagnostics in the Next Generation Sequencing Era. Methods Mol Biol Clifton NJ. 2021;2324:363\u0026ndash;81. \u003c/li\u003e\n\u003cli\u003eSantos R, Lee H, Williams A, Baffour-Kyei A, Lee SH, Troakes C, et al. Investigating the Performance of Oxford Nanopore Long-Read Sequencing with Respect to Illumina Microarrays and Short-Read Sequencing. Int J Mol Sci. 2025 May 8;26(10):4492. \u003c/li\u003e\n\u003cli\u003eOlson ND, Wagner J, Dwarshuis N, Miga KH, Sedlazeck FJ, Salit M, et al. Variant calling and benchmarking in an era of complete human genome sequences. Nat Rev Genet. 2023 Jul;24(7):464\u0026ndash;83. \u003c/li\u003e\n\u003cli\u003eNyaga DM, Tsai P, Gebbie C, Phua HH, Yap P, Le Quesne Stabej P, et al. Benchmarking nanopore sequencing and rapid genomics feasibility: validation at a quaternary hospital in New Zealand. NPJ Genomic Med. 2024 Nov 8;9:57. \u003c/li\u003e\n\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"bmc-medical-genomics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"mgnm","sideBox":"Learn more about [BMC Medical Genomics](http://bmcmedgenomics.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/mgnm/default.aspx","title":"BMC Medical Genomics","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Nanopore Long-read sequencing, Phasing, Long-range PCR","lastPublishedDoi":"10.21203/rs.3.rs-7242084/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7242084/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground:\u003c/h2\u003e\u003cp\u003eNext-Generation short-read sequencing has limited diagnostic utility in phasing distantly separated variants and analysing genomic regions with high homology. Determining the phase of variants from parental chromosomes is critical for accurate identification of compound heterozygosity. Long-read sequencing technology is able to overcome these limitations through the analysis of long haplotypes of separated variants. This study has developed and validated a robust, end-to-end workflow for phasing and localising variants using long-range PCR (LR-PCR) and targeted Nanopore sequencing for clinical implementation.\u003c/p\u003e\u003ch2\u003eMethods:\u003c/h2\u003e\u003cp\u003eNA24385 (HG002) reference DNA was used for all tests. Four PCR kits were tested to optimise LR-PCR for targets between 1 to 20 kb. Amplicons were barcoded and sequenced on Flongle flow cells, with up to eight amplicons on each flow cell. An in-house bioinformatic pipeline was developed to analyse the amplicons. This pipeline is capable of detecting chimeric reads (a known PCR artefact), and incorporating Clair3 for variant calling, and WhatsHap and HapCUT2 for phasing.\u003c/p\u003e\u003ch2\u003eResults:\u003c/h2\u003e\u003cp\u003eThe UltraRun LongRange PCR Kit performed with a 90% success rate for DNA amplification up to 22 kb. All 15 tested heterozygous Single Nucleotide Variant (SNV) pairs, and 10 small InDels, with inter-variant distances from 5.8 to 21.4 kb, were phased with 100% concordance to known phase. Furthermore, SNV calling within six low-mappability genes demonstrated precision and sensitivity of 100% against benchmark data. The median proportion of chimeric reads was maintained at 2.80% (range 1.79\u0026ndash;16.12%) under optimised conditions.\u003c/p\u003e\u003ch2\u003eConclusions:\u003c/h2\u003e\u003cp\u003eThis study establishes a reliable and affordable clinical diagnostic workflow for accurate phasing of variants separated by up to ~\u0026thinsp;20 kb and for variant localisation in genomic regions not able to be sequenced by short-read sequencing. This integrated approach enables implementation in diagnostic settings to resolve complex genetic findings and improve variant interpretation.\u003c/p\u003e","manuscriptTitle":"Long-range PCR and Nanopore sequencing for localisation and phasing variants: an end-to-end clinical application workflow","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-08-08 00:11:31","doi":"10.21203/rs.3.rs-7242084/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-09-08T05:10:06+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-09-04T13:01:35+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-08-28T15:34:57+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"64709660050729591273581332295745890651","date":"2025-08-08T18:11:16+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"86031172648720450609195540831736544258","date":"2025-08-08T04:33:35+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"221665443037639002172011996064157849120","date":"2025-08-05T12:51:54+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-08-05T12:43:51+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-08-04T12:51:00+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2025-07-31T04:41:49+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-07-31T00:36:41+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Medical Genomics","date":"2025-07-31T00:33:46+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"bmc-medical-genomics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"mgnm","sideBox":"Learn more about [BMC Medical Genomics](http://bmcmedgenomics.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/mgnm/default.aspx","title":"BMC Medical Genomics","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"395beace-61fe-4dcf-bfca-7ea5290c9587","owner":[],"postedDate":"August 8th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2025-11-24T16:04:09+00:00","versionOfRecord":{"articleIdentity":"rs-7242084","link":"https://doi.org/10.1186/s12920-025-02251-z","journal":{"identity":"bmc-medical-genomics","isVorOnly":false,"title":"BMC Medical Genomics"},"publishedOn":"2025-11-19 15:58:45","publishedOnDateReadable":"November 19th, 2025"},"versionCreatedAt":"2025-08-08 00:11:31","video":"","vorDoi":"10.1186/s12920-025-02251-z","vorDoiUrl":"https://doi.org/10.1186/s12920-025-02251-z","workflowStages":[]},"version":"v1","identity":"rs-7242084","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7242084","identity":"rs-7242084","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.